Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

115 results about "Speech technology" patented technology

Speech technology relates to the technologies designed to duplicate and respond to the human voice. They have many uses. These include aid to the voice-disabled, the hearing-disabled, and the blind, along with communication with computers without a keyboard. They enhance game software and aid in marketing goods or services by telephone.

Off-line and on-line mixed use method of AI voice

The invention relates to the technical field of AI voice, and particularly discloses an off-line and on-line mixed use method of AI voice, which comprises the following six steps of: evaluating a network connection state to obtain a network reliability index; analyzing voice data input by the user and preamble information, determining the complexity of voice intention and generating a demand priority; and then cooperatively selecting a voice processing mode according to the network reliability index, the demand priority and the decision rule. If the mode is an off-line mode, extracting a voice feature vector, classifying responses by using a local model, obtaining a processing mode confidence coefficient according to user feedback, and if the processing mode confidence coefficient is lower than a threshold value, triggering an on-line supplementary response and updating a rule; and if the mode is an online mode, feature vectors are extracted to be matched with a cloud knowledge base, a fine classification result is obtained, and accurate response is performed by means of a cloud large model. The whole process is combined with network conditions and user requirements, and a processing mode is flexibly decided, so that high-quality AI voice interaction service is provided.
Owner:SHENZHEN MAICHIRUI SOFTWARE CO LTD

Model processing method, voice interaction method, electronic device, and storage medium

Provided is a model processing method, a voice interaction method, an electronic device and a storage medium, relating to fields of artificial intelligence, big data and voice technologies. The model processing method includes: obtaining a candidate question set of each initial sample data in M initial sample data, wherein the initial sample data includes m rounds of question-and-answer between an object and an agent, the candidate question set includes a next round of question set corresponding to the mth round of question in the m rounds of question-and-answer; obtaining M training sample data based on the M initial sample data, the candidate question set and label data of each initial sample data, wherein the label data includes a target question to be generated by the agent in the (m+1)th round; and training a model to be trained by using the M training sample data to obtain a target model.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Customizable voice messaging platform

The present disclosure describes a customized voice messaging platform. The customized voice message placement can enhance engagement using personalized audio content. The platform can generate authentic-sounding voice messages using Al-driven processes, including text-to-speech (TTS) technology and audio concatenation, to create personalized audio content. The platform supports various delivery channels such as SMS, email, podcasts, and streaming services. The platform can incorporation visual elements like brand logos and animations into personalized messages. The platform can provide campaign creation functionality, enabling users to create campaigns deliver customized messages across multiple channels. The platform can provide message suggestions, automated testing, and other features. The platform can support rules for determining when to send messages and / or the content of messages.
Owner:ROBIN VOICE INC

Target positioning method, device, electronic device and storage medium based on large model

The present application discloses a target positioning method, device, electronic device, and storage medium based on a large model, and relates to the fields of computer technology, particularly to large models, voice technology, computer vision, deep learning, and the like. The solution is as follows: receiving a positioning request sent by a target terminal, the positioning request including a target image and a voice command; extracting first object information of the object to be positioned from the voice command, performing target detection on the target image based on the first object information, and obtaining a detection result; intercepting an object image of the candidate object from the target image based on the position information of the candidate object in the target image; determining the target object from the candidate objects using the large model based on the object image, the position information of the candidate object in the target image, and the first object information; and sending the position information of the target object in the target image to the terminal, so that the target terminal determines the position information of the target object relative to the target terminal based on the position information of the target object in the target image.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Automatic content generation method and device, electronic equipment and storage medium

The embodiment of the invention discloses an automatic content generation method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence content generation, the method comprises the following steps: obtaining a demand text, using a large language model to carry out deep semantic analysis on the demand text to obtain a script, and using a script generation model to generate a video script; performing semantic understanding through an A I technology and an image generation algorithm to generate an image, and converting a script into voice through a text-to-voice technology and a deep learning model; performing style analysis on the script, the image and the voice to extract style features, and performing style fusion according to the style features and style requirements in the demand text; and aligning the script, the image and the voice according to a time sequence, and editing and integrating the script, the image and the voice by using a video synthesis technology to generate a plurality of versions of video contents. According to the method, the problems of low generation efficiency, poor cross-modal content fusion and poor style uniformity in the prior art are solved.
Owner:ZHAOLIAN CONSUMER FINANCE CO LTD

Transform-based voice correction system and method

The invention provides a Transform-based voice correction system and method, and the system comprises a voice preprocessing module which is used for carrying out the noise suppression and feature extraction of an input southern Fujian dialect voice signal, carrying out the noise suppression through a frequency spectrum subtraction method, and extracting 40-dimensional MFCC features; the voice splicing module is used for solving the problem of voice input fragmentation; the dialect speech recognition module adopts an improved Transform encoder structure; and the character correction module adopts a character level Transform language model. According to the invention, more accurate and smooth voice interaction experience is provided for the user in the dialect area in southern Fujian. Meanwhile, the technical framework has good expandability, can adapt to other dialects and language variants, and provides important support for industrial application of the dialect voice technology.
Owner:ZHANGZHOU THIRD HOSPITAL (ZHANGZHOU LONGWEN HOSPITAL)

Voice processing method, device, equipment and storage medium

The present disclosure provides a speech processing method, apparatus, device and storage medium, which relate to the field of artificial intelligence, and in particular to the field of speech technology. The specific implementation scheme is: based on the parameter characteristics of the streaming acoustic processing module in the acoustic model, an acoustic processing filling block is obtained; the acoustic processing filling block is added to the i-th data block to be acoustically processed to obtain the i-th target acoustic processing data block; wherein the i-th data block to be acoustically processed is obtained by slicing the data to be processed into n data blocks to be acoustically processed; the i is a natural number not greater than n, and the n is a natural number not less than 2; the i-th target acoustic processing data block is input into the streaming acoustic processing module in the acoustic model for streaming acoustic processing to obtain the i-th streaming acoustic processing result.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Alarm method based on telephone scene and alarm information broadcasting robot

The invention provides an alarm method based on a telephone scene and an alarm information broadcasting robot, relates to the technical field of voice, and solves the technical problem that manual monitoring of alarm information is not timely in the prior art. The method comprises the following steps: monitoring alarm information, and generating alarm voice according to the alarm information; sending a voice request message to an operation and maintenance terminal; and if the operation and maintenance terminal responds to the voice request message, sending and broadcasting alarm voice to the operation and maintenance terminal. The method is used in a system alarm process.
Owner:KEXUN JIALIAN INFORMATION TECH CO LTD

Speech processing methods, apparatus, devices, storage media, and computer program products

This disclosure discloses speech processing methods, apparatus, devices, storage media, and computer program products, relating to the field of artificial intelligence, particularly speech technology and deep learning. Specifically, in response to the detection of an external sound card device being connected, the following operations are performed: disabling the voice wake-up recognition link; adapting the parameters of the connected external sound card device; and, in response to the completion of parameter adaptation of the external sound card device, restarting the voice wake-up recognition link to enable speech processing via the external sound card device.
Owner:BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD

Schedule reminding method, device and equipment based on vehicle terminal and storage medium

The present disclosure provides a schedule reminding method and device based on a vehicle terminal, equipment and a storage medium, relates to the field of artificial intelligence, in particular to the field of Internet of Vehicles, intelligent cockpit, natural language processing, voice technology, computer vision, Internet of Things, intelligent search technology. The specific implementation scheme is: in response to obtaining the starting state information of the vehicle to which the vehicle terminal belongs, logging in a first office account through a vehicle mobile office application, obtaining the to-be-processed work schedule of the first office account, issuing a schedule reminder, in response to receiving a play instruction issued for the schedule reminder, verifying the identity of the user who issues the play instruction, and in response to the identity verification of the user being passed, obtaining and outputting the work schedule information of the first office account. The technical scheme logs in the office account first and then performs identity verification, which can ensure that the user knows about new schedules in time and ensure the security of the schedule information.
Owner:APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD

Industrial internet software intelligent customer service interaction method based on multi-modal deep fusion

The invention relates to the technical field of industrial internet, artificial intelligence and human-computer interaction, in particular to an industrial internet software intelligent customer service interaction method based on multi-modal deep fusion. According to the method, an intelligent customer service system integrating natural language processing, a voice technology, computer vision and augmented reality (AR) is constructed aiming at the characteristics of complex operation, multiple user groups, high problem speciality and the like in an industrial scene. The core of the method is that operation questions, interface states and potential faults of engineers in software use are accurately understood through multi-modal input; an accurate solution is generated in combination with the knowledge graph of the industrial software and the real-time data of the equipment; a 3D virtual expert image is driven, and immersive and scene-based guidance is provided by fusing voice explanation, interface marking guidance, AR superposition demonstration and operation demonstration videos. The method has the working condition self-adaptive capability, can identify the professional level (such as a green hand / expert mode) of a user, and adjusts the explanation depth; and cross-region collaboration is supported, a plurality of users are allowed to synchronously share a guidance picture of a virtual expert, and efficient remote collaboration troubleshooting is realized. The method solves the technical problems that traditional industrial software is slow in customer service response, abstract in guidance, difficult to process complex problems on site and the like, and operation and maintenance efficiency and user experience are remarkably improved.
Owner:SHANGHAI CAIJIANG INTELLIGENT TECH CO LTD

Model training method, speech recognition method, and related devices

The present disclosure provides a model training method, a speech recognition method and related devices, which are related to the technical field of data processing, and in particular to the technical fields of artificial intelligence, computer vision, speech technology, intelligent search and the like. The specific implementation scheme is as follows: inputting a mouth shape sample sequence into a mouth shape processing model to obtain a first dictionary code prediction result predicted based on the mouth shape sample sequence; determining a loss value based on the first dictionary code prediction result and a second dictionary code prediction result; the second dictionary code prediction result is determined based on a target text corresponding to the mouth shape sample sequence; adjusting model parameters of the mouth shape processing model based on the loss value to obtain a mouth shape interpretation model; wherein the mouth shape interpretation model is used to assist speech recognition. The mouth shape interpretation model trained by the method can adapt to any speech recognition network and can realize plug and play.
Owner:APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD

Interaction method and device, intelligent agent, equipment, medium and program product

The invention provides an interaction method, an interaction device, an intelligent agent, equipment, a medium and a program product, and relates to the technical field of artificial intelligence, in particular to the technical field of man-machine interaction, computer vision and voice. According to the specific implementation scheme, in response to received multi-modal information input by a target object to a virtual object, intention recognition is conducted on the multi-modal information, and text information representing the intention of the current round is determined; based on the text information and historical visual perception information crossing time domains with the text information, response analysis is carried out, response information is obtained, and the historical visual perception information is obtained by carrying out visual perception on historical multi-modal information input by the target object in historical rounds; and broadcasting the response information through the virtual object.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Speech restoration model training method and device, and speech restoration method and device

The invention provides a training method of a voice restoration model, and a voice restoration method and device, and relates to the technical field of voice. The method is applied to the electronic equipment, and comprises the following steps: under the condition that damaged voice data exists in original voice data, obtaining damaged description information of the damaged voice data; wherein the damage description information is used for representing the damage reason of the damaged voice data; on the basis of the damaged description information, obtaining voice restoration features used for restoring the damaged voice data; taking the voice features corresponding to the damaged voice data as input data, taking the voice restoration features as prior information, and taking the sample voice features as target output data to train a voice restoration model, and obtaining a trained voice restoration model; wherein the sample voice features are used for representing voice features when the damaged voice data are not damaged. Based on the above scheme, the generalization and the restoration capability of the voice restoration model can be improved, and the restoration effect of the damaged voice data can be improved.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

High-performance zero-sample text-to-speech conversion method and system based on vLLM acceleration

The invention discloses a high-performance zero-sample text-to-speech method and system based on vLLM acceleration, and belongs to the technical field of intelligent speech, and the system comprises a dynamic streaming sentence segmentation module which carries out the intelligent segmentation of an input text, and obtains a to-be-converted text; the multi-reference audio fusion module is used for receiving a plurality of reference audios of a speaker and combining weight fusion to obtain a final audio coding feature; the session management and control module is used for judging whether the to-be-converted text is a new session request according to the session ID of the to-be-converted text, and if the to-be-converted text is the new session request, distributing a speaker ID and an associated audio signaling feature to the to-be-converted text; if not, searching a speaker ID (Identity) and an audio conditioning feature; and the text-to-voice module is used for performing voice synthesis on the text to be converted by adopting the IndexTTS model accelerated by the vLLM. According to the method, the reasoning performance can be remarkably improved, the audio quality and the response time delay stability in a multi-round dialogue scene are ensured, and the natural continuity of long text synthesis is ensured.
Owner:JIANGSU HAOBAI INFORMATION SERVICE CO LTD

Bluetooth connection voice prompter

This utility model relates to the field of Bluetooth voice technology, specifically a Bluetooth-connected voice prompt device, including a speaker, a sound hole, and a protective component. The sound hole is located on the outer arc surface of the speaker. A light guide ring is mounted on the upper surface of the speaker, and the light guide ring is arc-shaped. A button is fixedly mounted on the upper surface of the speaker. The protective component is located on the outer arc surface of the speaker and includes a guide rail. The guide rail is fixedly connected to the left side of the speaker near the sound hole and is rectangular in shape. A guide block is slidably connected to the inner wall of the guide rail, and a baffle is fixedly connected to the side of the guide block away from the guide rail. This utility model, by incorporating the protective component, can effectively prevent dust, hair, debris, and other foreign objects from entering the sound hole. If these foreign objects enter the interior, they may adhere to the speaker diaphragm and other sound-producing components, affecting sound propagation and sound quality, and may even damage the diaphragm, shortening the lifespan of the voice prompt device.
Owner:SHENZHEN LONGQIANSHU INVESTMENT SERVICES CO LTD

A quick order opening method based on voice

The present invention discloses a method for rapid ordering based on voice, which belongs to the field of intelligent ordering of goods and solves the problem of how to rapidly order goods based on voice technology, thereby improving ordering efficiency. After defining a database and an ordering language template, the present invention adopts voice wake-up technology, voice recognition technology and semantic recognition SDK technology to convert the acquired voice into a target character string; data entities are extracted from the target character string according to the ordering language template and matched with the database; finally, the matched information in the database is integrated and displayed on the software interface; the user confirms the information through voice and completes the order. The present invention can accurately obtain the user's ordering demand information, and no longer uses the method of inputting keywords or manual selection to search and order goods, thereby realizing rapid ordering of goods based on voice technology and improving the user's ordering efficiency. At the same time, it improves the user's software interaction experience and has won praise from users.
Owner:HEFEI YINGYUN INFORMATION TECH

Method and system for generating expressive audio based on text data

The invention relates to the technical field of audio processing, in particular to a method and system for generating expressive audio based on text data. The method comprises the following steps: firstly, acquiring to-be-converted audio needing to be converted, then splitting a to-be-converted text to obtain a plurality of corresponding text segments, and determining sound effect labels corresponding to the text segments; obtaining sound effect audios corresponding to the sound effect labels, and performing voice conversion on the to-be-converted text to generate a first target audio; then determining an insertion time period of each sound effect audio in the first target audio, and synthesizing each sound effect audio and the first target audio based on the insertion time period of the sound effect audio in the first target audio to obtain a second target audio; according to the method, the text-to-speech technology and the sound effect synthesis technology are combined, so that the immersion and richness of the finally obtained second target audio are better, and the auditory experience of the user is enhanced.
Owner:HUNAN MANGO INTELLIGENT MEDIA TECH DEV CO LTD

Analog interview processing method and device based on digital human and electronic equipment

The invention discloses a simulation interview processing method and device based on a digital human and electronic equipment, and relates to the technical field of computers, in particular to the artificial intelligence fields of deep learning, large models, voice technologies and the like. The specific implementation scheme is as follows: in response to a detected trigger operation of a job seeker on an interview start control in a simulation interview interface, displaying a digital person interview interface, and displaying a digital person interviewer in the digital person interview interface; the job hunting information of the job hunters is sent to the server side; receiving an interview problem of the simulation interview sent by the server; wherein the interview problem is generated by utilizing a large model according to the job application information; an interview question is asked to the job seeker through a digital person interviewer; and obtaining an answer result of the job seeker to the interview question.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

3D digital human lip shape driving method and device, electronic equipment and storage medium

The present disclosure provides a 3D digital human lip shape driving method and device, electronic equipment and storage medium, which relates to the technical field of computer vision. The method comprises: obtaining input text information; converting the text information into a phoneme sequence, audio data and timestamp information based on text-to-speech (TTS) technology; deleting corresponding silent phonemes in the phoneme sequence according to the timestamp information; performing a preset multiple sampling on the phoneme sequence after the deletion processing to obtain a bs animation coefficient sequence; and generating a 3D digital human lip shape animation according to the bs animation coefficient sequence, the audio data, a preset phoneme lip shape mapping table and a preset optimization of special phonemes. The present disclosure improves the robustness and fluency of 3D digital human lip shape driving.
Owner:CHINA TELECOM CORP LTD

Audio training dataset screening method based on phoneme matching pronunciation table

This invention belongs to the field of speech technology, specifically relating to a method for selecting audio training datasets based on phoneme matching pronunciation tables. Addressing the problems of high redundancy, high acquisition and annotation costs, and difficulty in verifying content integrity in existing technologies, the following solution is proposed: A target phoneme set is constructed according to the needs of the target task or domain; a pronunciation table is constructed or obtained and stored in key-value pair format; the original audio-text pair dataset is obtained, and the text is preprocessed; the preprocessed text is converted into corresponding phoneme sequences in the pronunciation table, and each phoneme sequence is analyzed; based on preset selection criteria, sample pairs in the original audio-text pair dataset are selected, retaining those that meet the selection criteria to form a reduced dataset; the reduced dataset is output. This solution reduces annotation costs while ensuring data comprehensiveness, and is suitable for processing datasets for lightweight edge speech synthesis models and fine-tuning of speech models.
Owner:HANGZHOU JUNTONG FUTURE TECHNOLOGY CO LTD

A controller and chip

The present disclosure provides a kind of controller and chip, it is related to integrated circuit technical field, more particularly to chip technology and voice technology.The specific implementation scheme is: a kind of controller, configuration is in chip, the controller includes multiple storage and interconnection interface and configuration interface;Wherein, the configuration interface is configured to the working mode of the multiple storage and interconnection interface, wherein the working mode includes storage interface mode and interconnection interface mode;Each storage and interconnection interface is configured based on the configuration of the configuration interface, and it is externally connected memory in the storage interface mode, or it is externally connected other chip in the interconnection interface mode.The present disclosure can realize multi-port storage and multi-chip interconnection in the same interface in the controller of chip by configuration, on the premise of controllable cost, not only can expand the required storage capacity, but also can improve overall computing power through inter-chip interconnection, meet different product requirements.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Speech synthesis model training method, speech synthesis method, speech synthesis device and electronic equipment

The invention provides a speech synthesis model training method, a speech synthesis method, a speech synthesis device and electronic equipment, and relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, natural language processing, speech technologies, large models and the like. The specific implementation scheme is as follows: acquiring training data; training samples in the training data comprise style description sample texts, input sample texts, output sample voices and style dimension data of the output sample voices; obtaining an initial speech synthesis model; training the speech synthesis model according to the style description sample text, the input sample text, the output sample speech and the style dimension data to obtain a trained speech synthesis model; wherein the electronic equipment can perform training processing on the speech synthesis model in combination with the style dimension data, so that more features can be considered when the speech synthesis model is trained, and the accuracy of the trained speech synthesis model can be improved.
Owner:BAIDU INT TECH (SHENZHEN) CO LTD

Industrial control instruction interaction method for grinding machine

According to the industrial control instruction interaction method for the grinding machine, whether the preset conditions are met or not is checked through the first downlink instruction, for example, whether the authority requirement and the operation confirmation requirement are met or not, then the operation is executed through the second downlink instruction, hierarchical authority control is achieved, and the misoperation risk of the grinding machine is reduced. In addition, instruction interaction is preferably carried out based on a Modbus TCP protocol, high coupling of an interaction protocol and a voice technology can be realized, and complex control logic such as parameter setting can be supported.
Owner:DONGGUAN YIFU MASCH TECH CO LTD

Audio-based device fault detection method, electronic device and storage medium

Provided is an audio-based device fault detection method, an electronic device, and a storage medium, relating to the field of data processing and in particular to technical fields of deep learning and voice technology. The method includes: obtaining initial audio data collected by a drone for a target device; preprocessing the initial audio data to obtain audio data to be detected; performing feature extraction on the audio data to be detected to obtain an audio feature of the audio data to be detected; constructing an information graph based on the audio feature; and obtaining a fault detection result for the target device based on the information graph and a graph neural network model.
Owner:ZHEJIANG HENGYI PETROCHEMICAL CO LTD

Video outbound method, device, equipment, medium and product

The invention discloses a video outbound method, device and equipment, a medium and a product, and relates to the field of artificial intelligence, and the method comprises the steps: determining a user background picture based on an outbound mobile phone number in a video call instruction; obtaining all video frames of a preset video; the preset video corresponds to the video call instruction; for any video frame, taking the background picture as a background, taking the video frame as a foreground, and performing pixel-level fusion on the background picture and the video frame to obtain a new video frame corresponding to the video frame; synthesizing the new video frames corresponding to the video frames into a video stream; converting the target character into voice by using a TTS character-to-voice technology; the target character corresponds to a preset video; and synthesizing the voice and the video stream to obtain a to-be-pushed video, and pushing the to-be-pushed video to a user side by using the VoLTE technology, so that the video to be pushed can be automatically generated, and the method and the device are compatible with various mobile phone models.
Owner:CHINA UNITED NETWORK COMM CO LTD JIANGXI BRANCH

Method, device and equipment for processing audio file based on large model and storage medium

The invention provides a method, a device and equipment for processing an audio file based on a large model, and a storage medium, and relates to the technical field of artificial intelligence such as voice technology and natural language processing. The specific implementation scheme is as follows: acquiring an input audio file in an interactive interface based on a large model; based on the large model, identifying the audio file, and obtaining and displaying identification information in the interactive interface; and based on the large model, analyzing the identification information, and obtaining and displaying an analysis result in the interactive interface. According to the technology disclosed by the invention, the functions of the AI product can be effectively enriched; the accuracy and the effectiveness of processing the audio file by the AI product can be improved; and the operation complexity of the user in the interaction process based on the large model can be effectively reduced, and the user experience is further effectively improved.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Image-to-speech system, device, method, and response output system

The present invention provides image-to-speech technology that can be used more suitably. An image-to-speech system according to the present invention includes a device that is carried by or worn by a user and performs speaking of objects from images captured by a camera. The device repeatedly automatically captures images from video of the camera at predetermined timings (step S2). The captured images are analyzed to acquire information including text representing objects in the images (step S3). An object and text to be spoken are decided in a predetermined determination, on the basis of the acquired information (step S4). Speaking of the text representing the object that is decided, is automatically repeated from the device at predetermined timings (step S5).
Owner:MAXELL LTD

Audio recall method, model training method, device and electronic equipment

The present disclosure provides an audio recall method, a model training method, a device and an electronic device, relates to the technical field of cloud computing, in particular to the technical field of deep learning, intelligent search and voice technology, and the audio recall method comprises: acquiring a first audio; segmenting the first audio to obtain N audio segments, any two adjacent audio segments in the N audio segments partially coincide, and N is an integer greater than 1; recalling a second audio corresponding to each of the N audio segments from a sample pool.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Voice wake-up method and device

The invention relates to the technical field of voice, and discloses a voice wake-up method and device, and the method comprises the steps: obtaining the living body indication information and spatial information, detected by voice interaction equipment based on target perception, of a target user relative to the voice interaction equipment, and carrying out the voice wake-up according to the target information of the target user relative to each piece of voice interaction equipment; and determining at least one target activation device in each voice interaction device, and controlling each target activation device to respond to a voice instruction sent by the target user. It can be seen that the probability that the voice interaction device is awakened by non-user sound sources such as environment sound can be reduced, the nearby awakening and nearby interaction performance of the voice interaction device and the user response accuracy are improved, and the voice interaction performance between the user and the device and the use experience of the user on the voice interaction device are improved.
Owner:FOSHAN VIOMI ELECTRICAL TECH