Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

8 results about "Conversational speech" patented technology

Conversational speech is also known as Informal speech. Speakers regularly apply informal speech with friends and relatives, in daily conversations and in personal letters. Informal speech can cover informal text messages and different types of written communication.

Method, system, electronic device and storage medium for interactive speech recognition

PendingCN122337186AModelSimConversational speech
This invention provides a method, system, electronic device, and storage medium for interactive speech recognition. The method includes: recognizing the dialogue speech input by a user in the current round; inputting text hypotheses into an interactive automated simulation framework; performing semantic-based intent allocation between the text hypotheses and the recognition results of the previous round; if the text hypotheses are determined to be correction instructions for the recognition results of the previous round, using the correction instructions to infer and correct the previous round's recognition results to obtain a corrected result for the previous round; if the text hypotheses are determined to be new dialogue content, using a user simulator to simulate the user's correction behavior to generate correction prompts, and the interactive speech recognizer using the correction prompts to infer and correct the text hypotheses of the current round to obtain a corrected recognition result for the current round. In modeling the speech recognition task, this invention utilizes a large text model and an interactive automated simulation framework to achieve real-time error correction based on user feedback.
Owner:SHANGHAI JIAOTONG UNIV

A high-dynamic-environment intercom voice enhancement method and system based on intelligent noise reduction

ActiveCN122050413BStationary noiseNoise
The application provides a high-dynamic-environment intercom voice enhancement method and system based on intelligent noise reduction. In response to a release event of a push-to-talk button, a two-state noise dictionary is constructed based on impulsive noise components and non-stationary noise components in background noise in a current high-dynamic environment. In response to a press event of the push-to-talk button, when it is detected that there is a transient region matching the impulsive noise components in the noisy intercom audio signal, online updating of the two-state noise dictionary is triggered to obtain an online updating dictionary. The noisy intercom audio signal is reconstructed based on the online updating dictionary to obtain an initial enhanced voice signal. Envelope reconstruction is performed on the voice segment with abnormal zero-crossing rate in the initial enhanced voice signal to generate a final enhanced voice signal. The technical scheme provided by the application can enhance conversation voice in a high-dynamic environment with non-stationary noise and impulsive noise.
Owner:SHENZHEN AIQISHI INTELLIGENT TECHNOLOGY CO LTD

Real-time full-duplex conversational speech-oriented streaming conversational state prediction method and system

Embodiments of the present application provide a kind of real-time full duplex voice dialogue-oriented streaming dialogue state prediction method and system.The method comprises: the processed speech mark sequence is input to the dialogue state prediction model of multimodal full duplex, carries out streaming dialogue state prediction, adopts staggered prediction mechanism, the current speech feature, text recognition result and dialogue state mark are modeled on time dimension, for realizing that text recognition result explicitly participates in the judgment process of dialogue state prediction, obtains the full duplex dialogue state mark of streaming continuous prediction;The interactive state of user is judged in real time based on the full duplex dialogue state mark of continuous prediction, to determine the control behavior of downstream dialogue.The embodiments of the present application reduce the overall computing overhead and real-time inference pressure of system, more accurately identify complex interactive behaviors such as user sentence completion, pause and interruption, etc.Improve the overall response performance and user experience of full duplex voice interaction system.
Owner:SHANGHAI JIAOTONG UNIV

Voice dialogue method and device, terminal device and storage medium

Embodiments of the application disclose a voice dialogue method and device, terminal equipment and a storage medium. The method comprises: obtaining user state information, wherein the user state information comprises first state information corresponding to a first user of the first terminal equipment and / or second state information corresponding to a second user, the second user being in a close relationship with the first user; generating dialogue content according to the first state information and / or the second state information through a dialogue model; obtaining a voiceprint feature of a target user, and generating dialogue voice according to the voiceprint feature and the dialogue content, and playing the dialogue voice. The embodiment can enhance the accompanying effect of AI voice chat provided for the user, and improve the human-computer interaction of the terminal equipment.
Owner:GUANGDONG XIAOTIANCAI TECH CO LTD

Interactive AI plush toy device

This not only enables natural conversations that respond to the user's intent and dialogue history, but also allows for inexpensive and highly reproducible manufacturing. In addition, it provides an interactive device that can handle highly personalized responses and respect privacy. [Solution] The system comprises a voice input means, a voice output means, a small, general-purpose computer, a control program executed by the computer, and a network communication means. The control program processes the user's voice information acquired by the voice input means, transmits the voice information or data based on the voice information to an external artificial intelligence service via an API, and controls the voice output means based on the response information received from the artificial intelligence service, thereby collaborating with the external artificial intelligence service to perform dialogue with the user. The voice input means, voice output means, and computer are integrally arranged inside a stuffed animal or a similar external component.
Owner:福元 秀

Dialogue speech synthesis method and device, computer device and storage medium

PendingCN122435917ASynthesis methodsConversational speech
The application relates to the technical field of artificial intelligence, and discloses a dialogue speech synthesis method and device, computer equipment and a storage medium, which comprise the following steps: obtaining a dialogue text sequence; wherein the dialogue text sequence comprises dialogue text word units and at least two speakers with speaker labels; obtaining speaker embedding features of each speaker based on the speaker labels; constructing speaker index data by indexing the dialogue text word units based on the speaker labels; wherein the speaker index data is used to represent the mapping relationship between the dialogue text word units and the speakers; generating target speaker embedding sequences based on the speaker index data and the speaker embedding features; and performing speech synthesis based on the target speaker embedding sequences and the dialogue text word units to obtain dialogue speech data between the speakers. The application can be applied to the dialogue speech synthesis scene of financial technology and medical health, and the accuracy of dialogue speech synthesis is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Selective disablement of noise cancelation for conversations

A noise canceling disablement system is provided to enables a user to hear select conversational speech directed at the user while wearing a noise cancelling hearable device. The disablement system automatically at least partial disables the noise canceling feature of the hearable device in response to recognizing conversational speech of a speaking person within a detected conversation zone of the user. In some cases, triggering of the noise canceling disablement further requires the speaking person to be identified by the system as significant person of the user.
Owner:SONY GROUP CORP

Multi-modal based companion robot dialogue quality evaluation method and computer device

This application provides a multimodal dialogue quality assessment method and computer device for companion robots, relating to the field of robot control technology. The method includes: collecting multimodal raw data of the current dialogue turn during the interaction process; extracting semantic features, speech prosody features, facial expression features, and physiological emotion features based on dialogue speech, facial expression images, and user physiological signals; obtaining semantic relevance scores, emotional consistency scores, and contextual coherence scores based on facial expression images, semantic features, speech prosody features, facial expression features, physiological emotion features, response text, and historical text scores; and obtaining the current dialogue quality assessment result of the companion robot based on the semantic relevance score, emotional consistency score, and contextual coherence score. This application improves the accuracy of the current dialogue quality assessment result of the companion robot, thereby improving the service quality of the companion robot.
Owner:CHONGQING PHOENIX TECHNOLOGY CO LTD