Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

502 results about "Speech generation" patented technology

Speech generation. Speech generation and recognition are used to communicate between humans and machines. Rather than using your hands and eyes, you use your mouth and ears. This is very convenient when your hands and eyes should be doing something else, such as: driving a car, performing surgery, or (unfortunately) firing your weapons at the enemy.

Intelligent real-time interactive question-answering system based on virtual digital human

The invention provides an intelligent real-time interactive question-answering system based on a virtual digital human, and belongs to the technical field of voice signal processing and voice recognition, and the system comprises a data acquisition module which receives a voice or text interaction request input by a user, collects the expression dynamic parameter sequence and limb movement sequence data of the user in real time, and transmits the data to a user interaction module; obtaining a standardized voice feature vector and structured text data; the cross-modal fusion module is used for constructing an interactive feature matrix; the behavior decision module outputs a decision instruction set; the knowledge retrieval module is used for generating an answer text with emotional adaptability and voice features; and the voice generation module is used for generating a mouth shape animation key frame, a micro expression parameter sequence and a limb action track of the virtual digital human, generating a voice response in combination with the answer text and the voice characteristics, and pushing the voice response to the user terminal. According to the method, the interaction experience and adaptability of the virtual digital human are remarkably improved.
Owner:XIAMEN DUOXIANG ANIMATION CO LTD

Methods and systems of text-conditioned audio-visual speech generation with multi-modal latent diffusion models

Methods, systems, and computer programs are presented for audio-visual speech generation with multi-modal latent diffusion models. One method includes encoding raw audio signals and video frames into respective latent spaces using audio and visual autoencoders. A text transcript is processed into phoneme sequences using a text transcript processor. The audio and visual latent spaces are conditioned using the text transcript and a conditioning variable. Joint distributions of the visual and audio latent spaces, text transcript, and conditioning variable are learned using a multi-modal latent diffusion model. The model adds noise to the latent audio-visual representations and predicts the noise through denoising neural networks. An inverted diffusion process is utilized to generate diverse speech content and speaker characteristics, resulting in realistic audio-visual speech. The technology presented provides a novel approach to conditional speech generation with potential applications in speech synthesis, voice conversion, and speech recognition.
Owner:TENSORTYPE INC

Voice generation method and device, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes such as financial science and technology, medical health and the like, and discloses a speech generation method, device, equipment and medium. And inputting a text speech language model to generate an intermediate code in combination with a speech code extracted based on a codebook generation mode, extracting a speaker feature vector in the prompt speech, decoding the intermediate code and the speaker feature vector by a generative adversarial decoder, and outputting a target speech. According to the method, fine-grained text representation is established by fusing character semantics and pinyin pronunciation information, and voice cloning is completed through unified voice language modeling and an adversarial generation mechanism in combination with codebook-driven acoustic coding and speaker personality characteristics, so that the naturalness, similarity and pronunciation accuracy of generated voices are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Knowledge distillation-based text-to-voice method, apparatus and device, and medium

The invention relates to the technical field of voice processing, can be applied to business scenes in the fields of medical health, financial science and technology, barrier-free service and the like, and discloses a text-to-voice method based on knowledge distillation, which comprises the following steps: carrying out standardization processing on an input text to generate a standard text sequence; the lightweight text encoder encodes the standard text sequence to generate a text implicit vector; the non-autoregressive acoustic feature prediction module maps the text implicit vector into a student acoustic feature sequence, and calculates alignment loss through knowledge distillation; and performing structured pruning and parameter quantization based on alignment loss, generating an acoustic feature sequence by the optimized model, and converting the acoustic feature sequence into a voice waveform by a vocoder. According to the method, through knowledge distillation, pruning optimization and parameter quantification, the reasoning speed and the cross-equipment adaptability are improved while the model size and the calculation requirement are reduced, so that the TTS system can realize efficient, low-delay and low-power-consumption voice generation in a resource-constrained environment.
Owner:PING AN TECH (SHENZHEN) CO LTD

Training and speech generation methods and apparatuses for speech generation model, electronic device, computer-readable storage medium, and computer program product

The present application provides training and speech generation methods and apparatuses for a speech generation model, an electronic device, a computer-readable storage medium, and a computer program product. The method comprises: obtaining a first speech generation model; obtaining sample data of a plurality of modalities; on the basis of a prompt image sequence and speech text, respectively calling a plurality of encoders to perform encoding, so as to obtain a multi-modal encoding vector sequence; on the basis of the multi-modal encoding vector sequence, calling a decoder to perform decoding, so as to obtain decoded text; determining a probability distribution for the decoded text and the sample data of the plurality of modalities, and determining a target loss on the basis of the probability distribution; and on the basis of the target loss, updating parameters of the decoder and at least one of the encoders, wherein the updated decoder and the plurality of updated encoders are configured to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to a prompt image sequence of a target object.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Environmental acoustic simulation voice generation method, apparatus and device, and medium

ActiveCN120612917ASpeech recognitionSpeech synthesisPhonetic environmentData set
The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses an environmental acoustics simulation voice generation method, device and equipment and a medium, and the method comprises the steps: obtaining and separating mixed voice data, and generating original voice content and original environmental acoustics information; converting the original voice content into first text information; determining a target environment acoustic tag in combination with the original environment acoustic tag, the first text information and the target geographical location information; acquiring target environmental acoustic information from a preset sound data set based on the tag, and adjusting the amplitude characteristic of the target environmental acoustic information to match the original environmental acoustic information; and synthesizing the adjusted target environment acoustic information and the original voice content into simulated voice data. According to the method, the target geographic position information is introduced to participate in acoustic feature determination and amplitude adjustment, so that the generated simulated voice data is more consistent in geographic semantics and acoustic performance, and the authenticity and concealment of voice environment disguise are effectively improved.
Owner:PING AN TECH (BEIJING) CO LTD

Hierarchical emotional speech generation method and device, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a hierarchical emotional speech generation method, which comprises the following steps: acquiring an input text and an emotional speech sample, extracting a text embedding feature from the input text, extracting a Mel spectrum feature from the emotional speech sample, and obtaining a text embedding feature; the method comprises the following steps: extracting phoneme-level, word-level and statement-level sentiment distribution characteristics through a hierarchical sentiment distribution extraction module, carrying out time dimension alignment, generating a multi-level sentiment guidance matrix, and inputting the multi-level sentiment guidance matrix and text embedding characteristics into a sentiment synthesis module to generate a target Mel spectrum; and finally, converting the target Mel spectrum into target voice through a vocoder. The emotion control is expanded from the statement level to the fine-grained level of phonemes, words and the like, and the multi-level emotion guidance matrix is combined for generation, so that the generated speech is finer and more abundant in emotion expression and conforms to the context, and the naturalness of speech generation and the accuracy of emotion transmission are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Three-dimensional digital human generation method and system capable of voice interaction

The invention belongs to the technical field of three-dimensional reconstruction, and discloses a three-dimensional digital human generation method and system capable of voice interaction. According to the invention, brand new speaking audios in different languages are automatically generated according to different languages of the input target text and the sampled human voice audios; the sequential stability and detail reduction capability of three-dimensional human motion are guaranteed by using multi-model joint estimation and a sequential loss function, and facial expression details and hand postures in the image can be accurately estimated. After the high-precision three-dimensional human body model is obtained through estimation, human body action and expression generation is carried out based on voice driving, accurate synchronization of actions and expressions generated through voice is achieved, and facial expression movement and body posture movement, namely a whole-body three-dimensional human body model, conforming to brand-new speaking audio are accurately generated; and finally, rendering the whole-body three-dimensional human body model into a real digital human capable of voice interaction by using a three-dimensional neural rendering model. According to the invention, the realization of single person picture input, high-precision three-dimensional digital person generation and voice interaction is facilitated.
Owner:NANJING UNIV OF SCI & TECH

Speech generation method and device based on rhythm characteristics, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to service system platforms of medical health, financial science and technology and the like, and discloses a speech generation method, device, equipment and medium based on rhythm characteristics, and the method comprises the steps: obtaining an input text and a corresponding reference speech, extracting a Mel spectrum, and generating speech representation; extracting overall rhythm features through a multi-head attention mechanism, combining with a text embedding feature input condition flow matching module, generating intermediate rhythm features, and executing threshold adjustment processing to obtain a fine rhythm feature set; and fusing the fine rhythm feature set with the text embedded features to generate an acoustic parameter sequence, and converting the acoustic parameter sequence into a voice waveform through a pre-trained neural vocoder. According to the method, rhythm feature generation is optimized through condition flow matching, the flexibility of rhythm expression is enhanced in combination with threshold adjustment, and the matching degree of the text and the rhythm is improved through cross-modal attention fusion, so that the generated voice is more natural, diversified and controllable.
Owner:PING AN TECH (SHENZHEN) CO LTD

Privacy information desensitization method and system for voice generation type large model

The invention discloses a privacy information desensitization method and system for a voice generation type large model, and relates to the technical field of artificial intelligence. Input voice data is discretized, Gaussian noise disturbance sensitive features are injected, and a cross-modal voice generation model is constructed in combination with three-stage training; and meanwhile, a cross-modal privacy enhancement mechanism is applied to detect and fuzzify sensitive information in real time in an output stage. According to the invention, the adaptability of a large-scale voice generation model in a privacy protection scene is improved, and comprehensive protection of user privacy is realized. And moreover, the capability of extracting user privacy information by an adversarial attacker is effectively limited, and the risk of sensitive information leakage in the data transmission, storage and generation process of the voice generation model is reduced. On the premise that privacy is ensured, the voice generation model can still keep high-quality generation performance, generated voice output has high naturalness and accuracy, and actual application requirements are met.
Owner:ZHEJIANG UNIV

Natural language generation

Techniques for determining when speech is directed at another individual of a dialog, and storing a representation of such user-directed speech for use as context when processing subsequently-received system-directed speech are described. A system receives audio data and / or video data and determines therefrom that speech in the audio data is user-directed. Based on this, the system determine whether the speech is able to be used to perform an action by the system. If the speech is able to be used to perform an action, the system stores a natural language representation of the speech. Thereafter, when the system receives system-directed speech, the system generates a rewrite of a natural language representation of the system-directed speech based on the previously-received user-directed speech. The system then determines output data responsive to the system-directed speech using the rewritten natural language representation.
Owner:AMAZON TECH INC

Speech synthesis model training method, speech synthesis method, electronic device, and storage medium

The present disclosure provides a speech synthesis model training method. The method can comprise: acquiring initial training data; selecting a plurality of first text segments, second text segments, and third text segments from coherent text, and acquiring first audio segments, second audio segments, and third audio segments from coherent audio; detecting whether the plurality of audio segments and text segments meet a stitching condition; on the basis of the text sequence, stitching, by alternatively arranging text and audio, the plurality of text segments and audio segments meeting the stitching condition to generate combined training data, so as to generate a combined training data set; on the basis of the combined training data set, training the initial speech generation model to obtain a trained speech synthesis model, wherein during the training, prosodic, tonal, and / or emotional features from the second text segments and / or the third text segments are extracted. Further disclosed is a speech synthesis method. The present disclosure allows for use of the context of the text and audio to realize text comprehension to capture prosody, tonality, and emotion in the context.
Owner:SHANGHAI XIYU JIZHI TECH CO LTD

Robot control method and device based on pulse neural network, equipment and medium

The invention relates to the technical field of robot control, and discloses a spiking neural network-based robot control method, which comprises the following steps of: preprocessing collected multi-modal data to obtain an emotion pulse signal; inputting the emotion pulse signal into a pre-constructed emotion pulse neural network, and outputting a comprehensive emotion pulse; inputting the comprehensive emotion pulse into a central pattern generator, and outputting a behavior rhythm; acquiring an environment feedback signal generated by executing the behavior rhythm, and adjusting a connection weight according to the environment feedback signal; optimizing the comprehensive emotion pulse according to the connection weight, and converting the optimized comprehensive emotion pulse into emotion interaction voice information; and adjusting the emotion intensity according to the optimized comprehensive emotion pulse and the multi-modal data. According to the method, data are collected through the multi-mode sensor to generate emotion pulses, the emotion pulses are input into the central mode generator to generate behavior rhythms after SNN processing, weight optimization, emotional speech generation and emotional steady-state control setting are combined with the STDP algorithm, and the efficiency of a robot service scene is improved.
Owner:SHENZHEN ZHONGSHEN ZHIHUI TECHNOLOGY CO LTD

Voice generation method and device, medium, electronic equipment and program product

The invention belongs to the technical field of artificial intelligence, and particularly relates to a voice generation method, a voice generation device, a computer readable medium, electronic equipment and a computer program product. The method comprises the following steps: acquiring a natural language instruction, wherein the natural language instruction is used for describing a presentation effect of voice in a natural language; a sample pair having a similar voice presentation effect with the natural language instruction is obtained, the sample pair comprises an audio sample and a text sample, and the text sample is used for describing the voice presentation effect of the audio sample; combining the natural language instruction with the sample pair to obtain a multi-modal cue word; and generating voice according to the multi-mode prompt word. According to the invention, the control precision and flexibility of voice generation can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Voice generation method and device based on distribution prediction, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business system platforms of communication, medical health, financial science and technology and the like, and discloses a speech generation method, device, equipment and medium based on distribution prediction. And inputting the context features and the historical waveform feature sequence into a next distribution prediction module, predicting probability distribution for generating a voice waveform at a next moment, generating voice waveform points based on the probability distribution, updating the historical waveform feature sequence, performing loop execution until a termination condition is met, and finally performing combination to generate a target voice waveform. According to the method, context coding and historical waveform features are combined, the long-distance dependence modeling capability is improved, the voice coherence and naturalness are ensured, voice generation is optimized through distributed prediction, the reasoning efficiency is improved, and the real-time synthesis requirement is met.
Owner:PING AN TECH (SHENZHEN) CO LTD

Emotion recognition method and device in film and television play dubbing

The invention relates to an emotion recognition method and device in film and television play dubbing, and the method comprises the steps: obtaining a film and television play audio of a source language and a to-be-dubbed text of a target language; the movie and television play audio is processed by using an emotion recognition model corresponding to the trained source language, and target emotion features are extracted; wherein the emotion recognition model is based on an audio pre-training sub-model and a feature extraction sub-model, and is obtained by training marked training data; the feature extraction sub-model is used for extracting audio features, and the audio pre-training sub-model is used for carrying out feature coding on the extracted audio features to generate emotional features; and performing voice synthesis on the to-be-dubbed text based on the target emotional characteristics by using a voice generation model corresponding to the trained target language, and generating dubbing of the target language corresponding to the movie and television play audio. Through the dubbing method and device, the original emotion is reserved in dubbing, cross-language voice seamless conversion is achieved, and the dubbing achieves the original taste and flavor effect.
Owner:YOUKU CULTURE TECH (BEIJING) CO LTD

Automated video generation

Disclosed systems and methods convert user-supplied textual content into animated videos. Textual input is received and processed to identify narrative elements such as characters, settings, and events. These elements are then transformed into visual scene components. Still images generated based on these components are subsequently animated in line with the narrative context. The system can automatically implement storytelling techniques adapted to incorporate neuroscience principles. The system also synthesizes speech for dialogues or narrations using voice synthesis technology that considers emotional markers, tone, and pace. Generated media and metadata are stored in a data storage system that maintains data integrity and enables efficient retrieval. Users can interact with an export interface to choose video resolution, format, and sharing options. A feedback system employing machine learning algorithms collects and analyzes user feedback for real-time adjustments to the generated animated video.
Owner:RIVERS DORINE

Voice generation method and device based on voice style adaptation, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes of medical health, financial science and technology, debate and the like, and discloses a speech generation method and device based on speech style adaptation, equipment and a medium, and the method comprises the steps: obtaining a target text, a target speaker speech and a reference style speech, extracting a phoneme feature sequence, acoustic features and rhythm coding information, and performing feature fusion to generate a fusion coding vector; and processing the fusion coding vector through a style adaptation module, generating an acoustic code containing a target style feature, generating a Mel-frequency spectrum feature based on the acoustic code, and inputting the Mel-frequency spectrum feature into a pre-training vocoder to generate a voice waveform. The voice style matching and rhythm control capability is improved through multi-feature fusion, the generalization capability of unseen speakers is enhanced by optimizing the style adaptation module, high-naturalness voice is generated in combination with the Mel spectrum and the vocoder, and the voice synthesis definition and adaptability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Device for Creating Digital Persona

A device for creating digital persona, which includes a data collection module responsible for collecting personality data of a target object, a personality training module utilizing a large language model and the personality data to train and generate a virtual personality model with personality characteristics of the target object, thereby generating a virtual personality consistent with the personality characteristics of the target object, an appearance video generation module and voice generation module respectively utilizing face replacement and voice cloning technologies to extract the pictures and sounds from the personality data to generate a virtual personality model with the appearance and voice characteristics of the target object, a lip synchronization module, utilizing lip synchronization technology to ensure that the mouth shape and voice of the digital persona are synchronized, and an interactive module, providing an interactive interface that allows users to interact with the virtual personality and receive responses from it.
Owner:MORPHUSAI CO LTD

Voice generation method and device based on pseudo-autoregression modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice generation method, device and equipment based on pseudo-autoregression modeling and a medium, and the method comprises the steps: obtaining a training sample containing a text sequence, a prompt voice segment and a target semantic token sequence; performing continuous fragment mask training on the text-to-semantic model to obtain a pseudo-autoregression trained text-to-semantic model; generating candidate speech output by using the text-to-semantic model and the initial semantic-to-acoustic model which are subjected to pseudo-autoregression training, and constructing a preference data pair; updating the semantics-to-acoustics model based on the preference data pair to obtain a preference optimized semantics-to-acoustics model; and generating target voice output based on the target text and the target prompt voice. According to the method, the time sequence modeling capability of the model is enhanced through pseudo-autoregression training, and the voice generation quality is directly optimized through the preference data pair, so that the voice alignment precision and the subjective listening feeling performance are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

High-fidelity end-to-end acoustic modeling method combining speech intelligibility reconstruction and autoregressive feedback optimization

PendingCN121565148ASpeech recognitionFrequency spectrumWeighting filter
The invention belongs to the technical field of speech recognition and acoustic modeling, and particularly relates to a high-fidelity end-to-end acoustic modeling method combining speech intelligibility reconstruction and autoregressive feedback optimization, which comprises the following steps of: firstly, establishing an intelligibility loss function based on a human speech intelligibility model; subjective definition features of the voice are reconstructed through a frequency spectrum reconstruction network, intelligibility related features are extracted and embedded in combination with a perceptual weighted filter bank, and the intelligibility related features and Mel frequency spectrum features are input into an acoustic encoder in parallel; based on this, by introducing a voice intelligibility reconstruction and autoregressive feedback mechanism, while the conciseness of an end-to-end voice modeling framework is maintained, systematic improvement of voice sharpness, naturalness and stability is realized, and the method can be widely applied to scenes such as intelligent voice assistants, voice transcription, virtual anchors, voice restoration, cross-language voice generation and the like. And the method has extremely high practical application value and popularization prospect.
Owner:GUANGDONG UNIV OF TECH

Voice generation method and device

The embodiment of the invention provides a voice generation method and device, computer equipment, a computer readable storage medium and a computer program product, and belongs to the field of audio processing. The voice generation method comprises the following steps: acquiring a content text and an emotion description text; determining an emotion weight array according to the emotion description text; determining a target emotion vector according to the emotion weight array and a plurality of basic emotion vectors; the target emotion vector and the content text serve as model input, first target voice is generated through a pre-trained audio synthesis model, and the first target voice comprises text content in the content text and emotion features in the emotion description text. According to the technical scheme provided by the embodiment of the invention, emotion control on the first target voice can be realized by utilizing the emotion description text, so that the voice generation stability and the emotion controllability and accuracy in the voice generation process are improved.
Owner:SHANGHAI HODE INFORMATION TECH CO LTD

Multi-speech synthesis model bearing method and device based on virtual GPU

The invention provides a multi-speech synthesis model bearing method and device based on a virtual GPU, and relates to the technical field of graphics processing units, and the method comprises the steps: carrying out the virtualization processing of a physical graphics processing unit, dividing the physical graphics processing unit into a plurality of virtual processing units with independent video memories and calculation quotas, and combining a resource scheduling mechanism, and deploying the speech synthesis language model instances in a plurality of service containers, and constructing a plurality of speech synthesis model bearing units. After a voice synthesis request is accessed, the scheduling module carries out load balancing according to the request connection number of each bearing unit, the request is distributed to a target bearing unit with the minimum connection number, and a voice generation task is completed by a virtual processing unit bound with the target bearing unit. According to the invention, resource division can be carried out on the physical graphic processing unit, and efficient operation of the multi-speech synthesis model is realized.
Owner:ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD

Synthetic speech generation with flexible emotion control

Disclosed are apparatuses, systems, and techniques that may use machine learning for generating artificial speech. The techniques include generating a synthetic speech using a machine learning model-readable speech embedding associated with a target degree of an emotion and obtained by combining a plurality of reference speech embeddings associated with respective reference degrees of the emotion.
Owner:NVIDIA CORP

Voice action synchronization method and device of virtual character, equipment and storage medium

The embodiment of the invention discloses a voice action synchronization method and device for a virtual character, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the voice generation of an answer text, obtaining an answer voice sequence and the attribute information of each phoneme, determining a standard mouth shape identifier of the corresponding phoneme based on the phoneme identifier of each phoneme, and obtaining a voice action synchronization result; generating a mouth shape animation sequence based on the standard mouth shape identifier of each phoneme according to the time range of each phoneme; determining an emotional action sequence and an emotional action starting timestamp corresponding to the emotional label of the answer text, and generating a supplementary animation sequence based on the emotional action sequence and the emotional action starting timestamp; and synchronizing the answer voice sequence, the mouth shape animation sequence and the supplementary animation sequence according to the time sequence to obtain a synchronization relationship, and driving the virtual response character of the user based on the synchronization relationship, the mouth shape animation sequence and the supplementary animation sequence when the answer voice sequence is played, so that the voice action synchronization function of the virtual character is realized.
Owner:SHANGHAI JIACHE INFORMATION TECH CO LTD

Marketing verbal skill generation method and device, equipment, storage medium and program product

The embodiment of the invention provides a marketing verbal skill generation method and device, equipment, a storage medium and a program product, and relates to the field of financial science and technology and artificial intelligence. The method comprises the following steps: acquiring multi-modal data; performing cross-modal feature fusion on the voice data and the text data to generate a joint feature vector; based on the joint feature vector, through an attention mechanism, a multi-dimensional emotion evaluation result is generated, and the multi-dimensional emotion evaluation result comprises an evaluation result of at least one preset emotion dimension; and according to the multi-dimensional emotion evaluation result, generating and pushing a marketing verbal skill corresponding to the multi-dimensional emotion evaluation result. According to the method provided by the invention, through joint modeling of voice and text features and in combination with an attention mechanism, a key emotion region is focused, so that the accuracy of emotion scoring is remarkably improved, and the attitude of a customer can be reflected more truly; according to the emotion evaluation result, the optimization suggestion is generated, the generated verbal skill accurately matches the demand of the customer, the response efficiency of the marketing verbal skill is improved, and the customer emotion improvement rate is greatly improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Retrieval enhancement generation-based verbal skill generation method and device, equipment and storage medium

The invention provides a verbal skill generation method and device based on retrieval enhancement generation, equipment and a storage medium, the method converts a service verbal skill into a vector and stores the vector in a vector database, when a customer problem exists, vector retrieval is directly carried out in the verbal skill vector database without recalculating all verbal skill vectors, and the user experience is improved. According to the method and the system, the vectors are extracted, labels are added to the vectors, and accurate screening can be performed during vector retrieval according to the verbal skill labels and specific features of customer problems, so that the retrieval efficiency is improved, the service response speed of customer service personnel is improved, the customer waiting time is shortened, and the customer service efficiency is improved. And then the accuracy of the target service answer is further improved based on the context corresponding to the retrieved verbal skill customer question in combination with the RAG model. The method is applied to business scenes of customer service systems in the financial business or medical field and the like for replying customer problems, and the retrieval efficiency and the retrieval accuracy of the customer service knowledge base in the financial business or medical health field can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Digital human tour guide voice generation method, system and device and storage medium

The invention relates to the field of intelligent speech synthesis and emotion calculation, and discloses a digital human tour guide speech generation method, system and device and a storage medium, the digital human tour guide speech generation method comprises the following steps: S1, constructing a user mental model comprising an initial knowledge state of a user for a knowledge graph; s2, on the basis of the model, predicting and evaluating candidate narrative paths, planning an optimal path and determining an expected mental state; s3, multi-modal explanation content is generated and broadcasted according to the optimal path; s4, collecting real-time feedback of the user to obtain a real mental state; and S5, comparing the real state with the expected state, calculating a prediction deviation, and dynamically calibrating the user mental model for subsequent planning according to the prediction deviation. According to the method, prospective path planning is carried out by constructing the mental model, closed-loop calibration and robustness evaluation are combined, and personalized explanation which is accurate, stable and free of lag adjustment is achieved.
Owner:NANJING NICEBRIDGE INFORMATION TECH CO LTD

Real-time intention recognition method and system based on streaming incremental reasoning

The invention relates to the technical field of voice processing, in particular to a real-time intention recognition method and system based on streaming incremental reasoning, and the method comprises the following steps: receiving the voice input of a user through a voice collection module, slicing the voice into a plurality of audio frames, and carrying out the recognition of the intention of the user through an incremental large language model module by adopting an Early-Exit reasoning mechanism; side outlets are arranged at multiple levels of the model, incremental reasoning is performed on token streams based on a QLoRA4-bit quantization technology, prediction results of multiple tokens are smoothed by using a stream ASR decoding module and an accumulative fusion module, and a stable final label is generated; the method has the beneficial effects that by combining streaming ASR decoding with incremental large language model reasoning, intention recognition and risk assessment can be immediately performed on each speech token after the speech token is generated. Through an Early-Exit reasoning mechanism, under the condition of high confidence, the system can output a fraud intention in advance in a middle layer in the reasoning process and stop subsequent calculation, and unnecessary calculation overhead is reduced.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Large model real-time voice interaction method and device based on autoregression voice synthesis

The invention provides a large-model real-time voice interaction method and device based on autoregressive voice synthesis, and the method comprises the steps: obtaining a voice instruction marked with a target text response and a target voice response, enabling a voice encoder to code the voice instruction into voice representation, and enabling a voice adapter to carry out the dimension reduction and feature conversion of the original voice representation; the large language model generates a hidden state according to the converted voice representation and samples the hidden state to obtain a text sequence; and processing the text sequence by adopting a text-voice language model based on an autoregression Transform structure, generating a voice marking sequence in a streaming manner, and converting the voice marking sequence into a voice signal through a vocoder. According to the method provided by the invention, the naturalness and fluency of speech synthesis are greatly improved while high real-time performance is ensured. The optimized voice decoding architecture effectively reduces the voice generation delay and improves the response speed of the voice interaction system.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI