Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

96 results about "Speech interaction" patented technology

Sound duplicating and low-delay streaming speech synthesis method and system based on ultra-short sample

The invention provides a voice replication and low-delay streaming voice synthesis method and system based on an ultra-short sample, relates to the technical field of artificial intelligence, and is suitable for intelligent interaction, outbound service and multi-mode communication scenes. Deep personalized customization of intelligent voice interaction and real-time generation of ultra-low delay are realized, and bidirectional streaming interaction of a system level is supported, so that the fluency and response speed of dialogues are improved; an ultra-short sample sound duplicating module specially designed for processing ultra-short audio samples and a speech synthesis engine with ultra-low delay and bidirectional streaming output capability are integrated, and the capabilities of the ultra-short sample sound duplicating module and the speech synthesis engine are applied to a real-time and interactive intelligent speech interaction process to form a complete and efficient solution. The industrial pain point is directly solved in a targeted manner, and the method has important commercial application value.
Owner:GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD

Method for processing cross-modal question answerning based on large model, apparatus and storage medium

A method for processing cross-modal question answering based on large model, an apparatus, and a storage medium are suggested, which relates to the field of artificial intelligence technologies such as speech interaction processing, large models, machine learning and natural language processing. The specific implementation includes: performing an activity detection on a target speech input by a user; in response to detecting a pause in the inputting of the target speech, obtaining a first text corresponding to a first input speech before the moment of the pause in the target speech; performing a text response processing using a pre-trained speech question answering processing system based on the first text and the first input speech.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Multi-modal voice interaction large model training method and system based on voice acoustic feature regulation and control, terminal equipment and medium

ActiveCN120954388ASpeech recognitionSpeech synthesisSpeech comprehensionModal voice
The invention discloses a multi-modal voice interaction large model training method and system based on voice acoustic feature regulation and control, terminal equipment and a medium, and relates to the technical field of multi-modal voice interaction.The method comprises the steps that a text token of a text training sample is obtained, a corresponding voice token is constructed, and pre-training data used for converting the text token into the voice token is obtained; in combination with the multi-modal input sample and the pre-training data, constructing fine-tuning training data for speech understanding and dialogue generation; constructing and pre-training a basic model by using the pre-training data; a multi-modal voice interaction large model is constructed based on a pre-training basic model, and fine tuning data is used for training, so that voice acoustic features can be regulated and controlled based on multi-modal input, and voice is output. According to the method, through alignment and staged training of the text token and the voice token, fine regulation and control of voice acoustic characteristics are realized, long voice continuity and interaction naturalness are improved, and a model is efficiently endowed with voice interaction capability of controllable timbre and emotion.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Transform-based voice correction system and method

PendingCN121096357ABiological modelsSpeech recognitionSpectral subtractionFeature extraction
The invention provides a Transform-based voice correction system and method, and the system comprises a voice preprocessing module which is used for carrying out the noise suppression and feature extraction of an input southern Fujian dialect voice signal, carrying out the noise suppression through a frequency spectrum subtraction method, and extracting 40-dimensional MFCC features; the voice splicing module is used for solving the problem of voice input fragmentation; the dialect speech recognition module adopts an improved Transform encoder structure; and the character correction module adopts a character level Transform language model. According to the invention, more accurate and smooth voice interaction experience is provided for the user in the dialect area in southern Fujian. Meanwhile, the technical framework has good expandability, can adapt to other dialects and language variants, and provides important support for industrial application of the dialect voice technology.
Owner:ZHANGZHOU THIRD HOSPITAL (ZHANGZHOU LONGWEN HOSPITAL)

Airborne flight safety enhancement system and method based on multi-modal interaction

The invention discloses an airborne flight safety enhancement system and method based on multi-modal interaction, and relates to the field of crossing of avionics systems and artificial intelligence, and the method comprises the steps that an environment sensing module generates structured environment information; the airborne equipment monitoring module performs real-time health state analysis; the large language model module comprises a computing platform sub-module, a language model sub-module, a knowledge base sub-module, a voice recognition sub-module and a voice synthesis sub-module; and the output module is used for displaying the decision support information generated by the large language model module on a cockpit display device. The method has the advantages that through deep cooperation of multi-mode perception fusion, distillation large language model intelligent decision, natural voice interaction and airborne health monitoring, on the premise of ensuring airworthiness safety, the flight safety is remarkably improved, the pilot load is reduced, and the decision efficiency is optimized.
Owner:上海柘飞航空科技有限公司

Emotional voice interaction method and device based on humanoid robot

The invention provides an emotional voice interaction method and device based on a humanoid robot, and relates to the field of robot interaction. The method comprises the following steps: acquiring voice data of a target user through a humanoid robot, and extracting acoustic features and semantic features in the voice data; inputting the acoustic features and the semantic features into a cross-modal fusion emotion classification model, and outputting an emotion recognition result corresponding to the target user; determining an emotion recognition misjudgment risk, and correcting an emotion recognition result based on the emotion recognition misjudgment risk; inputting the corrected emotion recognition result and the semantic feature into an emotion voice understanding interaction model, and outputting an emotion voice interaction result; and controlling the humanoid robot to execute a corresponding multi-modal interaction response based on the emotion voice interaction result. The problem that the voice interaction experience is greatly reduced due to the fact that the voice interaction content output by an existing humanoid robot is difficult to fit the real emotional state of the elderly user in multiple aspects such as text expression and voice emotion is solved.
Owner:BEIJING SHENMOU TECH CO LTD

User feedback for speech interactions

An interactive system may be implemented in part by an audio device located within a user environment, which may accept speech commands from a user and may also interact with the user by means of generated speech. In order to improve performance of the interactive system, a user may use a separate device, such as a personal computer or mobile device, to access a graphical user interface that lists details of historical speech interactions. The graphical user interface may be configured to allow the user to provide feedback and / or corrections regarding the details of specific interactions.
Owner:AMAZON TECH INC

A method and apparatus for echo cancellation and noise reduction

Embodiments of the present application disclose a method and device for echo cancellation and noise reduction. In the embodiments of the present application, a first near-end speech and a first far-end speech are obtained, and after pre-processing, delay detection and short-time Fourier transform, a fourth near-end speech and a fourth far-end speech are generated, and linear echo cancellation is performed to generate an error signal and a linear echo signal of the linear echo cancellation; the fourth near-end speech, the fourth far-end speech, the error signal and the linear echo signal of the linear echo cancellation are input into a two-stage neural network multi-task fusion model to be trained to determine a joint loss function; the two-stage neural network multi-task fusion model is generated according to the joint loss function, and then model compression is performed to generate a target two-stage neural network multi-task fusion model, which is used to realize echo cancellation and noise reduction. Through the above method, acoustic echo cancellation and noise reduction are realized under the condition of extremely low signal-to-echo ratio, and the speech interaction and communication performance of a terminal device are improved.
Owner:ZHEJIANG FUTURE ELF ARTIFICIAL INTELLIGENCE TECH CO LTD

SEMANTIC ENTITY RECOGNITION METHOD AND DEVICE, LANGUAGE INTERACTION SYSTEM, VEHICLE AND MEDIUM

This application relates to a semantic entity recognition method and a semantic entity recognition device, a speech interaction system, a vehicle, and a computer-readable storage medium. The semantic entity recognition method includes: preprocessing speech instruction information to obtain a core-mentioned word contained within the speech instruction information; and searching a scenario-semantically directed structure graph based on the core-mentioned word to obtain a target entity object that matches the core-mentioned word, wherein the scenario-semantically directed structure graph is constructed from a scenario entity object in a current scenario and semantically associated words of different granularities of the scenario entity object.
Owner:BYD CO LTD

An ethical AI integration framework for smart home systems that use voice assistants

System (100) for integrating ethical artificial intelligence for smart home systems using voice assistants, wherein the system (100) comprises: a speech interaction module (104) that is used to receive and is configured to process voice-based user commands; an analysis module (106) configured to interpret the processed speech commands based on context parameters; Ethics assessment modules (108-1) configured to make automated decisions regarding bias, Assess transparency and data protection; a smart home control module (108-2) configured to perform automated actions based on ethically evaluated decisions; and a monitoring module (110) configured to monitor system behavior and ensure ongoing compliance with predefined ethical standards.
Owner:DR VISHWANATH KARAD MIT WORLD PEACE UNIV PUNE

Teaching speech interaction-oriented education metadata processing method and device, and storage medium

PendingCN121071590AData processing applicationsArtificial lifeEngineeringEducational metadata
The invention relates to a teaching speech interaction-oriented education metadata processing method and device, and a storage medium, and the method comprises the steps: carrying out the slicing and sentence segmentation of classroom teacher and student speech corpus data, and carrying out the sampling and aggregation to form a plurality of sentence blocks; based on a preset labeling instruction and the knowledge graph, performing word-level labeling of teaching behaviors and knowledge points on the sentence blocks by utilizing an intelligent agent, and generating a behavior container through majority voting; based on a preset discrimination instruction, the labeled sentence tag and the behavior definition of the tag, performing anti-fact verification by utilizing an intelligent agent, judging whether the sentence tag is related to the behavior definition or not, and if so, storing the sentence tag as a final labeling result in a preset graph database; and if not, triggering an error correction process, generating a new label by the intelligent agent based on the correction instruction, and selecting the label with the highest category overlapping degree as a result by matching with the positive example set and sorting similarity. And the correction result is fed back to the analysis process, so that iterative optimization update of the labeling judgment and correction module is realized.
Owner:SHANGHAI NORMAL UNIVERSITY +1

Speech recognition method and electronic device

PCT designated stageWO2026045574A1Speech recognitionPhonetic environmentAudio electronics
A speech recognition method and an electronic device. The method comprises: an electronic device acquiring audio which is input by a user (S801); on the basis of the audio, the electronic device detecting a speech environment when the user inputs the audio, and determining a speech interaction mode corresponding to the audio (S802); and the electronic device performing speech recognition on the audio on the basis of a speech recognition scheme corresponding to the speech interaction mode, so as to obtain a speech recognition result (S803). By detecting the environment in which the user inputs the audio, the electronic device can determine that the user is in a noisy, quiet, or another speech environment, and can then determine the speech interaction mode. The electronic device can execute corresponding speech recognition on the acquired audio by means of the speech recognition scheme corresponding to the speech interaction mode, thereby improving the accuracy of speech recognition.
Owner:HUAWEI TECH CO LTD

A speech recognition method

The application provides a speech recognition method, which comprises: obtaining feature data of audio to be recognized; inputting the feature data into an acoustic model to obtain a time sequence label matrix corresponding to the feature data; decoding the time sequence label matrix through a first language model to obtain a plurality of decoding paths and corresponding probability scores thereof, and determining the decoding paths with the top N probability scores as N first decoding results, wherein N is a positive integer; determining a corresponding target intent field based on the N first decoding results, a previous round of speech interaction field and a current scene field; determining a corresponding second language model based on the target intent field, re-computing probability values of the decoding paths with the top N probability scores through the second language model, and generating second decoding results; and determining a speech recognition result of the audio to be recognized based on the second decoding results. The application improves the recognition efficiency while ensuring the speech recognition accuracy.
Owner:HUBEI QIGUANG TECHNOLOGY CO LTD

Speech interruption decision method and system based on multi-granularity semantic completeness prediction

PendingCN122313981ASemantic featureData mining
This invention discloses a speech interruption decision-making method and system based on multi-granularity semantic integrity prediction, belonging to the field of speech interaction technology. The method includes: extracting multi-granularity semantic features and combining them into a current multi-granularity semantic feature vector; calculating the offset of the current multi-granularity semantic feature vector with a standard semantic pause model to obtain a semantic offset feature vector; performing similarity matching and weighted voting between the semantic offset feature vector and a historical interruption decision case library to obtain the current predicted interruption confidence; and generating an interruption command when the current predicted interruption confidence exceeds a dynamic threshold. This invention solves the problems of existing technologies where speech interruption decisions rely on simple energy thresholds or semantic integrity probabilities, lack utilization of user expression habits and contextual experience, and have low decision accuracy by introducing a standard semantic pause model and a historical case analogy reasoning mechanism, thus achieving more intelligent and accurate speech interruption judgment.
Owner:GUANGZHOU JIUSI INTELLIGENT TECH CO LTD

Cross-modal attention fusion method and device based on voiceprint features and language semantics

The application relates to the technical field of speech recognition, and discloses a cross-modal attention fusion method and device based on voiceprint features and language semantics, which comprises the following steps: performing a feature extraction operation on speech data to obtain voiceprint features; based on a sentiment perception mask mechanism and a context sentiment memory unit, fusing sentiment features into text data to obtain semantic features; constructing a cross-modal relationship graph of the voiceprint features and the semantic features; determining a time sequence dependency relationship between the voiceprint features and the semantic features based on the edge weights generated by the cross-modal relationship graph; projecting the voiceprint features and the semantic features into the same semantic space based on the time sequence dependency relationship, and performing a feature reconstruction fusion operation based on a reconstruction loss function to obtain voiceprint and semantic fusion features; and adjusting a dialogue strategy determined based on the voiceprint and semantic fusion features; and using the adjusted dialogue strategy to generate interactive response data with historical interactive data, so that a high-accuracy and personalized speech interaction experience is realized.
Owner:GUANGDONG GUANGXIN COMM SERVICES COMPANY

Automobile air-conditioner control method and air-conditioner control system

Provided in the present disclosure are an automobile air-conditioner control method and an air-conditioner control system. The air-conditioner control system comprises a speech interaction device, an infotainment host, a central operation module and a zone controller, wherein the speech interaction device collects user's speech, and sends the speech to the infotainment host; the infotainment host generates an air-conditioner control instruction on the basis of the speech, and sends the air-conditioner control instruction to the central operation module; the central operation module determines, on the basis of the instruction type of the air-conditioner control instruction, whether a vehicle meets an air-conditioner control condition, and if the vehicle meets the air-conditioner control condition, the central operation module sends the air-conditioner control instruction to the zone controller; and on the basis of the air-conditioner control instruction, the zone controller controls an automobile air conditioner to operate. In this way, an air-conditioner control instruction is accurately generated by identifying user's speech, whether a vehicle meets an air-conditioner control condition is verified, and on the basis of the rational air-conditioner control instruction that passes the verification, the automobile air-conditioner is controlled to operate, such that an air conditioner in a vehicle can be conveniently controlled by means of the user's speech, thereby avoiding the driving safety risk caused by means of manual adjustment.
Owner:CHINA FAW CO LTD

Voice noise reduction method and device, earphone, medium and program product

The invention discloses a voice noise reduction method and device, an earphone, a medium and a program product. The method comprises the following steps: acquiring voice sequence data; performing at least one self-attention operation on the voice sequence data; obtaining the output of the self-attention mechanism; noise reduction is carried out according to the output of the self-attention mechanism; the sg attention operation comprises the following steps: carrying out linear processing on the input sequence data by utilizing a preset query-key matrix and a preset value matrix to obtain a query-key vector and a value vector; the query-key matrix corresponds to a product between a preset query matrix and a preset key matrix; obtaining a target attention score according to the transpose of the voice sequence data and the query-key vector; carrying out Softmax operation on the target attention score to obtain an attention weight; and performing matrix multiplication on the attention weight and the value vector to obtain the output of the self-attention operation. In this way, the harsh requirement of the embedded device for high real-time performance in a voice interaction scene can be met.
Owner:BESTECHNIC SHANGHAI CO LTD

Multi-language natural voice interaction method and system for coffee robot

PendingCN121768397AMeet individual needsImprove consumer experienceSpeech recognitionEngineeringAcoustics
The invention relates to the technical field of voice services, in particular to a multi-language natural voice interaction method and system for a coffee robot, and the method comprises the steps: collecting a voice signal of a user, carrying out the recognition of the language type of the voice signal, obtaining the type information of the voice signal, activating a corresponding language recognition model according to the type information, and carrying out the recognition of the type information; performing semantic analysis processing on the voice signal, generating language text information corresponding to the voice signal, performing coffee customization demand analysis on a user based on the language text information to obtain coffee customization information of the user, initiating a voice interaction behavior to the user according to the coffee customization information to obtain an interaction feedback instruction of the user, and sending the interaction feedback instruction to the user. And a corresponding coffee beverage is produced according to the interaction feedback instruction and the coffee customization information.
Owner:SHENZHEN CHUANGJIE INTELLIGENT TECHNOLOGY CO LTD

Call method, and electronic device, storage medium and product

The present disclosure relates to the technical field of computers, and particularly relates to a call method, and an electronic device, a storage medium and a product. The call method comprises: in an interaction interface of a first object and a second object and on the basis of the input of the first object, determining a target scenario mode from among one or more scenario modes configured for the second object (S102), wherein each scenario mode is configured with a speech feature; and controlling the second object to perform speech interaction with the first object on the basis of the speech feature of the target scenario mode (S104).
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Expression control method and system for driving lip, tooth and tongue linkage of humanoid robot based on anatomy, facial topological structure and language pronunciation rules

The invention discloses a humanoid robot expression control method and system based on anatomy and a facial topological structure, relates to the technical field of bionic robots and voice interaction, and aims to solve the problems that the mouth shape and pronunciation of an existing robot are not matched, the expression is rigid, and no natural micro-expression exists in a non-pronunciation state. A facial topological structure model matched with a robot is constructed based on human anatomy, a lip-tooth-tongue time sequence synchronous linkage control instruction is generated in combination with a language pronunciation rule, facial expressions, micro expressions and eye movement data matched with pronunciation are synchronously generated, meanwhile, an independent expression control instruction can be generated in a non-pronunciation state, and the control precision is improved. And finally, full-scene full-face natural expression control of the robot is realized. The humanoid reality and man-machine interaction naturalness of the humanoid robot are greatly improved, and the method can be widely applied to various bionic service robots and virtual digital humans.
Owner:孙量

An interaction method, apparatus, product, electronic device, and medium

The present disclosure provides an interaction method, device, product, electronic equipment and medium, relating to the technical field of artificial intelligence. The method comprises: receiving input information of a user, generating a control result according to the associated information of the input information; injecting the control result into the decoding process of a speech interaction model used to generate target speech as a constraint condition of the model output; generating target speech matched with the control result based on the speech interaction model, forming a mandatory constraint on the generation direction by injecting the control result through the model, enabling the model to follow the control result when generating target speech, actively regulating the speech generation process, and making the generated target speech matched with the interaction demand of the input information, thereby improving the pertinence and rationality of human-computer interaction.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Voice interaction method and device based on thinking chain, computer equipment and medium

The invention relates to the technical field of voice generation, and particularly discloses a voice interaction method and device based on a thinking chain, computer equipment and a medium. The voice feature vectors are directly received through the large language model to generate the thinking chain, accumulative errors caused by a traditional series framework are avoided, semantic accuracy is improved, reply voice quality is improved, the initial text thinking chain capable of being edited by the user is generated through the large language model, and user experience is improved. And the initial text thinking chain is corrected according to the user input content to obtain the target text thinking chain, so that the user can modify the voice in the voice generation process, personalized customization is realized, and the reply voice quality is further improved. The method is applied to a voice interaction system of financial or medical services such as voice product explanation, remote medical consultation, medical education and training and the like, the text content, the voice emotion and the style modified by the user can be obtained through the text thinking chain, voice editing is achieved, and the user experience is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech data processing method and system based on large language model

The application discloses a speech data processing method and system based on a large language model, relates to the technical field of speech data processing, and comprises the following steps: performing multi-channel feature decomposition on a received original speech signal, and constructing an acoustic state representation tensor; constructing a semantic candidate distribution space, generating multiple sets of semantic hypothesis vectors, and constructing a semantic evolution path graph; generating a semantic uncertainty function representing semantic ambiguity and speech disturbance sensitivity; dynamically constructing a reasoning depth control parameter and inputting the same to a multi-layer reasoning path scheduling unit of the large language model, constructing an intention structure vector, and mapping the intention structure vector into a structured semantic output. The technical problems that in the prior art, under a complex acoustic environment, it is difficult to accurately and effectively separate acoustic features, leading to low speech understanding accuracy of high ambiguity, and lacking dynamic reasoning ability to cope with semantic uncertainty risks are solved, and the technical effects of improving semantic understanding precision, ambiguity resolution ability of speech interaction, and reducing business misjudgment rate and risk are achieved.
Owner:GUANGDONG JINWAN INFORMATION TECH CO LTD

Intelligent evaluation method, device and equipment for classroom teacher-student interaction and storage medium

The invention belongs to the technical field of education informatization, and particularly discloses an intelligent evaluation method and device for classroom teacher-student interaction, equipment and a storage medium. According to the invention, classroom teacher and student speech interaction characteristics are determined according to the identification result; determining classroom teacher-student non-speech interaction characteristics according to a detection result; generating a classroom interaction quantization matrix according to the classroom teacher-student speech interaction features and the classroom teacher-student non-speech interaction features, and constructing a classroom interaction feature tensor according to the classroom interaction quantization matrix; and based on the target tensor clustering model, performing intelligent evaluation on classroom teacher-student interaction according to the classroom interaction feature tensor. Through the above mode, the classroom teacher-student speech interaction characteristics and the classroom teacher-student non-speech interaction characteristics are determined according to the classroom teaching video to be evaluated, and intelligent evaluation is carried out based on the target tensor clustering model which can reveal the high-dimensional interaction relationship while maintaining the integrity of the data structure. Therefore, the accuracy and fairness of teacher-student interaction evaluation can be effectively improved.
Owner:HUAZHONG NORMAL UNIV

Speech interaction method, and interaction device, electronic device and storage medium

The embodiments of the present disclosure relate to a speech interaction method, and an interaction device, an electronic device and a storage medium. The method proposed herein comprises: in response to a preset operation for an interaction device, controlling an audio collection unit of the interaction device to start collecting a first speech signal, wherein the interaction device is suitable for being worn on a finger of a user; acquiring the first speech signal collected by the audio collection unit; and sending the first speech signal to a terminal device connected to the interaction device, such that the terminal device uses the first speech signal to generate a second speech signal.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Voice interaction method and device based on primitive context information and time sequence alignment

The invention relates to the technical field of man-machine interaction and artificial intelligence, and particularly discloses a voice interaction method and device based on primitive context information and time sequence alignment. When a trigger event is detected, voice is collected in real time, the context of a target primitive is analyzed, semantic binding and fusion data generation are carried out after the time sequence overlapping rate reaches the standard, a target agent processes the fusion data, constructs prompt words and transmits the prompt words to a large language model, a structured interaction response instruction is generated, and finally the interaction response instruction is executed. And interaction is realized. According to the invention, the context information of the target primitive is acquired in real time, so that the problem of context missing in traditional voice interaction is solved. When the user triggers the interaction event, the context information and the semantic environment corresponding to the target primitive can be accurately recognized, so that the voice instruction is associated with the interaction focus, the high-quality interaction response prompt word is constructed, the accurate interaction response instruction is finally generated, and the interaction efficiency and accuracy are improved.
Owner:SHENZHEN REALMAX TECHNOLOGY CO LTD

Multi-mode emotional voice interaction method and system

The invention discloses a multi-mode emotional voice interaction method and system, and belongs to the technical field of artificial intelligence interaction. The invention aims to solve the problems of voice interaction mechanization and emotional expression and environment separation of the existing intelligent terminal (such as a robot with a body and a virtual digital human). The system is used as a middleware engine independent of a hardware body and a general large model, and receives multi-modal perception data (visual scenes, user expressions and environment contexts) of an intelligent terminal and a text intention of the general large model through a standard interface; adopting a multi-head attention mechanism to carry out adaptive weighted fusion on multi-modal features, and dynamically capturing an interaction relationship between modals; and introducing an anthropomorphic degree dynamic adjustment mechanism, calculating anthropomorphic degree parameters according to scene social attributes, and adaptively adjusting acoustic control parameters (fundamental frequency, energy and duration). And finally, driving a speech synthesis engine to generate anthropomorphic speech with emotional infection. The system further comprises an end-cloud collaborative architecture and a multi-modal conflict resolution mechanism, and various intelligent interaction terminals can be endowed in a low-cost and standardized mode, so that the intelligent interaction terminals have natural and robust emotional interaction capability.
Owner:DEEP DIMENSION QUADRANT (SHIJIAZHUANG) TECHNOLOGY CO LTD

Voice interaction method, device and electronic equipment

The application provides a speech interaction method and device and electronic equipment, and relates to the technical field of speech processing. The method comprises the following steps: acquiring speech information input by a user, and acquiring historical intention text of the user; inputting the speech information into a speech encoder of a spoken language understanding model to obtain acoustic coding features output by the speech encoder; inputting the historical intention text into a text encoder of the spoken language understanding model to obtain text coding features output by the text encoder; inputting the acoustic coding features and the text coding features into an intention recognition module of the spoken language understanding model to obtain an intention recognition result output by the intention recognition module, so as to be used for speech interaction. The application can improve the accuracy of the intention recognition result and accurately acquire the real intention of the user.
Owner:HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Intelligent accompanying-oriented emotional anthropomorphic multi-mode voice interaction large model system

The invention provides an intelligent accompanying-oriented emotional personification multi-modal voice interaction large model system, which relates to the field of language processing and is characterized by comprising the following modules: a user interaction module, a multi-modal data fusion module, a database module, a data analysis module and a function processing module, the user interaction module comprises touch interaction, voice interaction, text interaction and local interaction. The intelligent accompanying toy has the advantages that by integrating core technologies such as multi-modal data fusion, hierarchical memory and RAG, dynamic emotion response and edge-cloud collaboration, the problems of emotion interaction templating, lack of role consistency, high response delay, insufficient utilization of complex modals, weak long-term memory ability and the like of an existing intelligent accompanying toy are systematically solved; and the reality sense, the individuation degree and the real-time performance of interaction are remarkably improved, so that high-quality accompanying experience with higher emotional value and immersion is provided for the user.
Owner:SUZHOU PINGPING MAINLAND TECHNOLOGY CO LTD