Voice interaction method and device, equipment and storage medium

The voice interaction data and status information are obtained through voice interaction devices, and the target message data is automatically played, which solves the problem that smart home appliances cannot play messages without sensory, improves the success rate of message reception and reduces the operation burden, and expands the intelligence of smart home appliances.

CN120452443APending Publication Date: 2025-08-08SHENZHEN TCL NEW-TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510599782.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing smart home appliances cannot effectively play messages for users without sensory and non-invasive conditions, resulting in users who may miss viewing messages, affecting the interaction effect of messages.

Method used

The voice interaction device responds to the voice interaction request, obtains the voice interaction data and the status information of the target interaction object. When the preset message triggering conditions are met, the target message data associated with the target interaction object is automatically played, including voiceprint feature matching and status recognition to determine the playback conditions.

Benefits of technology

It realizes the effective playback of messages for users without any sense, improves the success rate of message reception, reduces the operation burden of users to listen to messages, and expands the intelligence level of smart home appliances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452443A_ABST
    Figure CN120452443A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction method and device, equipment and a storage medium, and the voice interaction method comprises the steps: obtaining voice interaction data of a voice interaction request in response to the voice interaction request; obtaining object state information of a target interaction object corresponding to the voice interaction request; and when the voice interaction data and the object state information meet a preset message triggering condition, playing target message data associated with the target interaction object. According to the technical scheme, the message interaction operation can be performed by using the intelligent household electrical appliance, the success rate of receiving the message by the user is improved, the operation burden of listening to the message by the user can be effectively reduced, and the intelligent degree of the intelligent household electrical appliance is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a voice interaction method, apparatus, device and storage medium. Background Art

[0002] At present, with the continuous advancement of science and technology, home appliances are gradually developing in the direction of intelligence, and more and more smart home appliances are beginning to integrate voice recognition and voice control functions. However, although existing smart home appliances have integrated voice recognition and voice control functions, they still rely on paper records, text messages or emails to leave messages. However, these message methods cannot ensure that users can effectively read messages in a non-invasive and senseless manner, which may cause users to miss the messages and affect the message interaction effect. Summary of the Invention

[0003] The embodiments of the present application provide a voice interaction method, apparatus, device and storage medium, which aim to solve the technical problem in the prior art that smart home appliances cannot effectively play messages for users without the users noticing.

[0004] In one aspect, an embodiment of the present application provides a voice interaction method, the voice interaction method comprising the following steps:

[0005] Responding to a voice interaction request, obtaining voice interaction data of the voice interaction request;

[0006] Obtaining object state information of a target interaction object corresponding to the voice interaction request;

[0007] When the voice interaction data and the object status information meet a preset message triggering condition, target message data associated with the target interaction object is played.

[0008] In a possible implementation of the present application, when the voice interaction data and the object status information meet a preset message triggering condition, playing the target message data associated with the target interaction object includes:

[0009] Extracting voiceprint features based on the target wake-up word in the voice interaction data to obtain target voiceprint features;

[0010] Matching the target voiceprint feature with a preset voiceprint feature to obtain a target identity corresponding to the target voiceprint feature;

[0011] If the object status information and the target identity meet the preset message triggering condition, the target message data associated with the target identity is played.

[0012] In a possible implementation of the present application, if the object status information and the target identity meet a preset message triggering condition, playing the target message data associated with the target identity includes:

[0013] If the object status information is the first object status information and there is a first candidate message associated with the target identity, obtaining a voiceprint confirmation password associated with the preset message triggering condition;

[0014] Acquire updated voice data associated with the voiceprint confirmation password, and updated voiceprint features in the updated voice data;

[0015] If the updated voiceprint feature is the same as the target voiceprint feature, the first candidate message is determined as the target message data, and the target message data is played.

[0016] In a possible implementation of the present application, if the object status information and the target identity meet a preset message triggering condition, playing the target message data associated with the target identity includes:

[0017] If the object status information is the first object status information and the target identity meets the preset message triggering condition, generating a message acquisition request according to the device identification information and the target identity;

[0018] The target message data associated with the target identity identifier in the message database is read according to the message acquisition request, and the target message data is played.

[0019] In a possible implementation of the present application, before playing the target message data associated with the voice interaction data, the method further includes:

[0020] Responding to a message generation request, obtaining message input data associated with the message generation request;

[0021] Candidate message data is generated according to the message pointing identifier and message content data in the message input data.

[0022] In a possible implementation of the present application, the message input data includes voice input data;

[0023] The generating of candidate message data according to the message pointing identifier and message content data in the message input data includes:

[0024] If the message input data is voice input data, performing audio recognition on the voice input data to obtain a message pointing identifier corresponding to the voice input data;

[0025] Performing activity detection on the voice input data to obtain message content data of the voice input data;

[0026] Candidate message data is generated according to the message pointing identifier and the message content data.

[0027] In a possible implementation of the present application, performing activity detection on the voice input data to obtain message content data of the voice input data includes:

[0028] Performing feature extraction on the voice input data to obtain voice input features;

[0029] Calculating the window feature value of each sliding analysis window according to the speech input feature;

[0030] Determining a target analysis window corresponding to the speech input data according to the window feature value and a preset window detection threshold;

[0031] The target analysis window is used to extract message content data from the voice input data.

[0032] In a possible implementation of the present application, the message pointing identifier includes a first message identifier and a second message identifier, and the candidate message data includes a first candidate message and a second candidate message;

[0033] The generating of candidate message data according to the message pointing identifier and the message content data includes:

[0034] If the message pointing identifier is the first message identifier, generating a first candidate message according to the preset voiceprint feature corresponding to the first message identifier and the message content data;

[0035] If the message pointing identifier is a second message identifier, a second candidate message is generated according to the first message identifier and the message content data.

[0036] In one possible implementation of the present application, the message input data includes text input data;

[0037] The generating of candidate message data according to the message pointing identifier and the message content data includes:

[0038] If the message input data is text input data, then obtaining text content data in the text input data;

[0039] generating synthesized message data according to the text content data and the target message tone;

[0040] Candidate message data is generated according to the synthesized message data and the message pointing identifier.

[0041] In another aspect, the present application provides a voice interaction device, comprising:

[0042] A data acquisition module is configured to respond to a voice interaction request and acquire voice interaction data of the voice interaction request;

[0043] A state acquisition module is configured to obtain object state information of a target interaction object corresponding to the voice interaction request;

[0044] The message playing module is configured to play target message data associated with the voice interaction data when the voice interaction data meets a preset message triggering condition.

[0045] On the other hand, the present application also provides a voice interaction device, the voice interaction device comprising:

[0046] one or more processors;

[0047] Memory; and

[0048] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps of the voice interaction method.

[0049] On the other hand, the present application also provides a computer-readable storage medium on which a computer program is stored, and the computer program is loaded by a processor to execute the steps in the voice interaction method.

[0050] In this application, by responding to a voice interaction request, the voice interaction data of the voice interaction request is obtained; the object status information of the target interaction object corresponding to the voice interaction request is obtained; and when the voice interaction data and the object status information meet the preset message triggering condition, the target message data associated with the target interaction object is played. When the target interaction object and the smart home appliance perform voice interaction, the voice interaction data of the target interaction object can be obtained, and the object status information of the target interaction object can be detected. When the voice interaction data and the object status information meet the preset message triggering condition indicating that the corresponding target message data can be effectively listened to, the target message data associated with the target interaction object is automatically played, thereby realizing the message interaction operation using the smart home appliance, improving the success rate of users receiving messages, and being able to effectively reduce the operational burden of users listening to messages, thereby expanding the intelligence level of the smart home appliance. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0052] Figure 1 This is a schematic diagram of a scenario of the voice interaction method according to an embodiment of the present application;

[0053] Figure 2 This is a flow chart of an embodiment of the voice interaction method in the embodiment of the present application;

[0054] Figure 3 A flowchart of an embodiment of generating message data in the voice interaction method provided in an embodiment of the present application;

[0055] Figure 4 A flowchart of an embodiment of generating candidate message data based on text input data in a voice interaction method provided in an embodiment of the present application;

[0056] Figure 5 A schematic diagram of the structure of an embodiment of the voice interaction device provided in the embodiment of the present application;

[0057] Figure 6 A schematic diagram of the structure of an embodiment of the voice interaction device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0058] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0059] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.

[0060] In this application, the word "exemplary" is used to mean "serving as an example, illustration, or illustration." Any embodiment described in this application as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. The following description is given to enable any person skilled in the art to make and use the invention. In the following description, details are listed for the purpose of explanation. It should be understood that one of ordinary skill in the art will recognize that the invention can be practiced without these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0061] At present, with the continuous advancement of science and technology, home appliances are gradually developing in the direction of intelligence, and more and more smart home appliances are beginning to integrate voice recognition and voice control functions. However, although existing smart home appliances have integrated voice recognition and voice control functions, they still rely on paper records, text messages or emails to leave messages. However, these message methods cannot ensure that users can effectively read messages in a non-invasive and senseless manner, which may cause users to miss the messages and affect the message interaction effect.

[0062] Based on this, the present application proposes a voice interaction method, apparatus, device and computer-readable storage medium to solve the technical problem in the prior art that smart home appliances cannot effectively play messages for users without the users noticing.

[0063] The voice interaction method in the embodiment of the present invention is applied to a voice interaction device, and the voice interaction device is arranged in a voice interaction device. The voice interaction device is provided with one or more processors, memories, and one or more applications, wherein the one or more applications are stored in the memories and are configured to be executed by the processor to implement the voice interaction method; wherein the voice interaction device can be an intelligent terminal, such as a mobile phone, a tablet computer, a network device, and a smart computer, etc.; optionally, the voice interaction device can also be a server, or a service cluster composed of multiple servers.

[0064] like Figure 1 As shown, Figure 1 This is a scene diagram of the voice interaction method of the embodiment of the present application. The voice interaction scene in the embodiment of the present invention includes a voice interaction device 100 (a voice interaction device is integrated in the voice interaction device 100), a cloud server 200 and a mobile message terminal 300. A computer-readable storage medium corresponding to the voice interaction method is running in the voice interaction device 100 to execute the steps of the voice interaction method. Among them, the cloud server 200 is a cloud server for storing message data and transmitting messages during the voice interaction process. The mobile message terminal 300 can be a mobile terminal for inputting messages to any voice interaction device 100. Optionally, the mobile message terminal 300 can be an intelligent terminal such as a mobile phone, a tablet computer, a network device, a server and a smart computer.

[0065] Optionally, in other scenarios, the voice interaction scenario may also include only the voice interaction device 100.

[0066] It is understandable that Figure 1 The voice interaction devices in the voice interaction method scenario shown, or the devices included in the voice interaction devices, do not constitute a limitation on the embodiments of the present invention. That is, the number and type of voice interaction devices included in the voice interaction method scenario, or the number and type of devices included in each device, do not affect the overall implementation of the technical solution in the embodiments of the present invention, and can all be regarded as equivalent replacements or derivatives of the technical solution claimed to be protected in the embodiments of the present invention.

[0067] The voice interaction device 100 in the embodiment of the present invention is mainly used to: respond to a voice interaction request and obtain voice interaction data of the voice interaction request;

[0068] Obtaining object state information of a target interaction object corresponding to the voice interaction request;

[0069] When the voice interaction data and the object status information meet a preset message triggering condition, target message data associated with the target interaction object is played.

[0070] The voice interaction device 100 in the embodiment of the present invention can be an independent smart home appliance, such as an air conditioner, a television, a refrigerator and other household appliances, or other smart home devices such as a car, a door lock and a smart speaker.

[0071] The embodiments of the present application provide a voice interaction method, apparatus, device, and computer-readable storage medium, which are described in detail below.

[0072] It will be understood by those skilled in the art that Figure 1 The application environment shown in is only one of the application scenarios related to the present application solution and does not constitute a limitation on the application scenario of the present application solution. Other application environments may also include Figure 1 More or fewer voice interaction devices, or voice interaction network connection relationships shown, such as Figure 1 Only one voice interaction device is shown in the figure. It can be understood that the scenario of the voice interaction method can also include one or more voice interaction devices, which is not limited here. The voice interaction device 100 can also include a memory for storing voice interaction data and other data.

[0073] It should be noted that Figure 1 The scenario diagram of the voice interaction method shown is only an example. The scenario of the voice interaction method described in the embodiment of the present invention is to more clearly illustrate the technical solution of the embodiment of the present invention, and does not constitute a limitation on the technical solution provided by the embodiment of the present invention.

[0074] Based on the scenarios of the above-mentioned voice interaction method, various embodiments of the voice interaction method disclosed in the present invention are proposed.

[0075] like Figure 2 As shown, Figure 2 This is a flow chart of an embodiment of a voice interaction method in an embodiment of the present application. The voice interaction method includes the following steps 201 to 203:

[0076] 201. Respond to a voice interaction request and obtain voice interaction data of the voice interaction request;

[0077] The voice interaction method in this embodiment is applied to voice interaction devices. The type and number of voice interaction devices are not specifically limited. That is, the voice interaction device can be one or more smart home appliance terminals. In a specific embodiment, the voice interaction device is a smart home appliance or other smart device. For example, the voice interaction device can be a smart home appliance such as an air conditioner, a television, a refrigerator, a smart door lock, and a smart speaker with voice interaction function.

[0078] When message data exists and a pre-set interaction request is made with the target interaction object, the voice interaction device can seamlessly play target message data that meets preset message trigger conditions to the target interaction object, thereby ensuring that the user can effectively receive the target message data and effectively reducing the user's operational burden of listening to messages. The target interaction object is the user object that conducts voice interaction with the voice interaction device. The target message data is the voice message data associated with the target interaction object that is stored by the voice interaction device or received from a cloud server or mobile message terminal.

[0079] That is, after receiving the voice message data sent by the cloud server and / or mobile message terminal, the voice interaction device can trigger the message playback detection process when receiving the voice interaction request triggered by the target interaction object, and then identify whether to play the specified message data in the subsequent steps.

[0080] Specifically, during operation, the voice interaction device can respond to a voice interaction request and perform voice interaction with the user, and obtain the voice interaction data of the voice interaction request during the voice interaction process, and then determine whether to play the specified message data based on the voice interaction data in the subsequent steps, so as to realize the seamless playback of the associated message data to the target interaction object. Among them, the voice interaction request is for the target interaction object to wake up the voice interaction device through voice interaction and drive the voice interaction device to perform a specific interactive operation. For example, in a specific embodiment, the target interaction object wakes up the voice interaction device by speaking the wake-up word corresponding to the voice interaction device, and speaks the relevant operation instructions when the voice interaction device is in the awake state to trigger the voice interaction request. For another example, in another specific embodiment, the wake-up word of the voice interaction device is "Xiao T Xiao T", and the target interaction object can trigger the voice interaction request for the voice interaction device by speaking "Xiao T Xiao T + operation instruction".

[0081] Specifically, the voice interaction device is pre-installed with a microphone component for receiving a voice interaction request and obtaining voice interaction data corresponding to the voice interaction request, wherein the voice interaction data is voice data carrying a voice interaction instruction to be executed by the target interaction object to drive the voice interaction device.

[0082] Optionally, the voice interaction instruction may be a regular interaction instruction and a message retrieval instruction. The regular interaction instruction may be a voice interaction instruction other than a message interaction instruction. For example, the regular interaction instruction may be a voice interaction instruction in which the target interaction object drives the voice interaction device to perform function switching, parameter adjustment, mode adjustment, and device startup and shutdown via voice interaction. The message retrieval instruction is a voice interaction instruction in which the target interaction object drives the voice interaction device to broadcast relevant messages via voice interaction.

[0083] That is, the voice interaction device can trigger the message playback process when the target interaction object conducts regular voice interaction. In addition, the voice interaction device can also trigger the message playback process when the target interaction object actively obtains the message data, so as to achieve the seamless and non-invasive playback of the associated target message data to the target interaction object.

[0084] 202. Obtain object state information of a target interaction object corresponding to the voice interaction request;

[0085] Specifically, after receiving a voice interaction request and executing the corresponding voice interaction instruction, the voice interaction device also obtains object status information of the target interaction object corresponding to the voice interaction request to ensure that the target message data can be fully received by the target interaction object. The object status information indicates whether the target interaction object is currently able to effectively listen to the message data.

[0086] Optionally, the object status information includes first object status information for which the message information can be completely obtained and second object status information for which the message information cannot be completely obtained.

[0087] Specifically, the voice interaction device can identify the state of the target interactive object through an image acquisition component or a posture detection component, thereby determining the object state information of the target interactive object. For example, the voice interaction device can collect the object posture data of the target interactive object through an image acquisition component or a posture detection component, and use the object posture data to perform object state identification, thereby determining the object state information of the target interactive object. Wherein, the image acquisition component and / or posture detection component can be set by the voice interaction device itself, or set by other terminals, and the voice interaction device communicates and interacts with other terminals to obtain the object posture data collected by the image acquisition component and / or posture detection component set by the other terminals.

[0088] Optionally, the voice interaction device can use the image acquisition component to obtain regional image information in a preset area, and use a preset image segmentation model to segment object image information in the regional image information, perform posture recognition on the object image information, obtain object posture data of the target interactive object, and match the object posture data with reference posture information corresponding to each object state information, determine the reference posture information whose similarity with the object posture data is greater than a preset similarity threshold as the target posture information, and determine the object state information corresponding to the target posture state information as the object state information of the target interactive object.

[0089] Optionally, in a specific embodiment, the reference posture information associated with the first object state information may be standing posture information, sitting posture information, or other posture information indicating that the user is focused on interacting with the voice interaction device. The reference posture information associated with the second object state information may be sleeping posture information, walking posture information, or undetected posture information indicating that the user is not focused on interacting with the voice interaction device.

[0090] 203. When the voice interaction data and the object status information meet a preset message triggering condition, play the target message data associated with the target interaction object.

[0091] Specifically, after obtaining the voice interaction data and object status information of the target interaction object, the voice interaction device further determines whether the voice interaction data and object status information meet the preset message trigger conditions. When the voice interaction data and object status information meet the preset message trigger conditions, the target message data associated with the target interaction object is obtained and the target message data is played.

[0092] Optionally, after the voice interaction device obtains the object status information of the target interaction object, if it determines that the object status information is the second object status information, it determines that the target interaction object is currently unable to concentrate on listening to the message data. When there is no clear message playback instruction in the voice interaction data of the target interaction object, the message data is not played to the target interaction object.

[0093] Optionally, after the voice interaction device obtains the object status information of the target interaction object, if it determines that the object status information is the first object status information, it determines that the target interaction object is currently able to effectively listen to the message data, and then the voice interaction device further determines whether the voice interaction data meets the preset message trigger condition, and when the voice interaction data meets the preset message trigger condition, it obtains and plays the target message data associated with the target interaction object, thereby achieving a seamless, non-invasive and effective playback of the message data to the user. Among them, the preset message trigger condition is a trigger condition strategy that characterizes the driving voice interaction data to play the message data. Optionally, in a specific embodiment, the preset message trigger condition is when the object status information is the first object status information and there is message data associated with the target identity identifier corresponding to the voice interaction data.

[0094] The target message data is voice message data generated by the message recipient through voice input or text input via a voice interaction device or mobile message terminal. Optionally, in a specific embodiment, the target message data includes first target message data and second target message data. The first target message data is targeted message data received by a specific interaction partner. The second target message data is general message data received by any interaction partner.

[0095] Specifically, the candidate message data to be played in the voice interaction device can be the first candidate message and the second candidate message data. When there is a first candidate message that is only played to a specified user, after the voice interaction device determines that the object status information of the target interaction object is the first object status information, the voice interaction device further determines whether the voice interaction data meets the preset message trigger condition. That is, the voice interaction device extracts the voiceprint based on the target wake-up word in the voice interaction data to determine the target voiceprint feature of the target interaction object. Among them, the target wake-up word is the voice data corresponding to the specified wake-up word preset for triggering the voice interaction device to perform voice interaction. For example, in a specific embodiment, the target wake-up word is set to "Xiaot T Xiaot T". In addition, the target wake-up word can also be customized and modified by the user.

[0096] Specifically, after being awakened by the target interaction object, the voice interaction device obtains the voice interaction data corresponding to the voice interaction request, extracts the wake-up word from the voice interaction data, determines the target wake-up word in the voice interaction data, and extracts the audio feature of the target wake-up word to obtain the target voiceprint feature of the target interaction object.

[0097] Specifically, after the voice interaction device obtains the target voiceprint feature of the target interaction object, the voice interaction device matches the target voiceprint feature with the preset voiceprint feature to determine the target identity identifier corresponding to the target voiceprint feature. That is, the voice interaction device compares the target voiceprint feature with the preset voiceprint feature, and determines the identity identifier associated with the preset voiceprint feature that is the same as the target voiceprint feature as the target identity identifier of the target interaction object.

[0098] Among them, the preset voiceprint feature is the voiceprint feature associated with a specific identity identifier stored in the preset voiceprint database of the voice interaction device. Each preset voiceprint feature corresponds to an object identity identifier. The target identity identifier is the identity identifier information that uniquely represents a specific speaker. Optionally, the identity identifier information can be any one or more of the identity tags, user nicknames, and user names configured by the user or the voice interaction device. For example, in a specific embodiment, the identity identifier information can be an identity tag, such as dad / mom / grandpa / grandma, etc. In another specific embodiment, the identity identifier information can be the user name, such as Zhang San / Li Si / Wang Wu, etc. In addition, in another specific embodiment, the identity identifier information can also be the user nickname customized by the user.

[0099] Specifically, after obtaining the target identity of the target interaction object, the voice interaction device determines whether the target identity meets the preset message trigger conditions based on the target identity, and then determines whether to play the corresponding voice message data. In other words, after determining the target identity, the voice interaction device accesses a preset message database to confirm whether there is a first candidate message associated with the target identity in the message database. If there is a first candidate message associated with the target identity, the first candidate message is determined as the target message data.

[0100] Optionally, when the voice interaction device determines that the object status information of the target interaction object is the first object status information, and there is a first candidate message associated with the target identity identifier in the candidate message data, the voice interaction device also guides the target interaction object to perform voiceprint confirmation according to a pre-set voiceprint confirmation password.

[0101] The voiceprint confirmation password is a preset password that directs the target interactive object to output updated voice data. In one embodiment, the voiceprint confirmation password includes a random confirmation password and a custom confirmation password. The random confirmation password is a random password generated by a cloud server or voice interactive device to remind the target interactive object to read along. For example, in one embodiment, the random confirmation password is a random password composed of a preset password template and a randomly generated verification code: "You have unread messages. Say

[5634] to play them for you."

[5634] is the randomly generated verification code.

[0102] In addition, in another specific embodiment, the custom confirmation password is a follow-up password information pre-defined by the target interaction object or the target message object. For example, in a specific embodiment, the custom confirmation password can be "You have unread messages, and they will be played for you after you say the password."

[0103] Specifically, after obtaining the voiceprint confirmation password, the voice interaction device plays the voiceprint confirmation password to the target interaction object, obtains the updated voice data read aloud by the target interaction object based on the voiceprint confirmation instruction, and performs voiceprint extraction on the updated voice data to obtain the updated voiceprint features in the updated voice data.

[0104] Optionally, if the voice interaction device determines that the updated voiceprint feature is the same as the target voiceprint feature, the first candidate message is determined as the target message data, and the target message data is played.

[0105] Optionally, in other embodiments, if the voice interaction device determines that the object status information is the first object status information and the target identity meets the preset message trigger condition, a message acquisition request is generated based on the device identification information and the target identity to drive the voice interaction device or the cloud server to read the target message data associated with the target identity in the message database according to the message acquisition request, and use the audio output module of the voice interaction device to play the target message data. Optionally, the message database can be set as a cloud database on the cloud server, or it can be set locally on the voice interaction device when the voice interaction device hardware meets the corresponding storage requirements.

[0106] Optionally, in other embodiments, after the voice interaction device determines that there is a second candidate message that can be listened to by any interactive object, when the object status information of the target interactive object is the first object status information and the target interactive object sends voice interaction data, the voice interaction device determines that the target interactive object meets the preset message conditions, and generates a message acquisition request corresponding to the second candidate message. Based on the message acquisition request, the second candidate message is obtained from the message database, and the second candidate message is played as the target message data.

[0107] Optionally, in another specific embodiment, the voice interaction device further receives a message replay request from a target interaction object, determines a target voiceprint feature of the target interaction object based on the message replay request, and plays historical message data associated with the target interaction object based on the target voiceprint feature. The historical message data may be the first target message data and / or the second target message data.

[0108] In this embodiment, the voice interaction device obtains the voice interaction data of the voice interaction request by responding to the voice interaction request; obtains the object status information of the target interaction object corresponding to the voice interaction request; and plays the target message data associated with the target interaction object when the voice interaction data and the object status information meet the preset message triggering condition. This is achieved by obtaining the voice interaction data of the target interaction object and detecting the object status information of the target interaction object when the target interaction object and the smart home appliance perform voice interaction. When the voice interaction data and the object status information meet the preset message triggering condition indicating that the corresponding target message data can be effectively listened to, the target message data associated with the target interaction object is automatically played, thereby realizing the message interaction operation using the smart home appliance, improving the success rate of users receiving messages, and effectively reducing the operational burden of users listening to messages, thereby expanding the intelligence level of the smart home appliance.

[0109] like Figure 3 As shown, Figure 3This is a flow chart of an embodiment of generating message data in the voice interaction method provided in an embodiment of the present application. Specifically, in this embodiment, the voice interaction method further includes steps 301 to 302:

[0110] 301. Respond to a message generation request and obtain message input data associated with the message generation request;

[0111] 302. Generate candidate message data according to the message pointing identifier and message content data in the message input data.

[0112] Based on the above embodiments, in this embodiment, the voice interaction device can generate candidate message data through the voice interaction device or mobile message terminal before responding to the voice interaction request and triggering the message playback process and playing the target message data when executing the voice interaction request.

[0113] Specifically, during operation, the terminal responds to a message generation request, which is an operational request that drives the terminal to generate candidate message data. The triggering method for this message generation request is not specifically limited herein. Specifically, the message generation request can be triggered by the target message recipient waking up the terminal with voice and inputting the corresponding message input data. Furthermore, the message generation request can also be triggered by clicking the message button on the preset message interface.

[0114] Specifically, after receiving the message generation request, the terminal also obtains the message input data associated with the message generation request. The message input data is the input data of the target message recipient during the message generation interaction. In one specific embodiment, the message input data includes voice input data or text input data. The voice input data refers to the voice audio data input by the target message recipient using voice as the input method during the message generation process. The text input data refers to the text data input by the target message recipient using text as the input method during the message generation process.

[0115] Optionally, in a specific embodiment, the message input data may be voice input data. That is, after waking up a voice interaction device or mobile message terminal, the target message recipient uses voice input to make an appointment to interact with the terminal. The terminal then uses a voice recognition model to perform voice recognition on the target message recipient's input data to obtain the voice input data. Optionally, the voice recognition model may be an ASR (Automatic Speech Recognition) recognition model.

[0116] Specifically, after receiving the message input data, the terminal generates candidate message data based on the message reference identifier and message content data in the message input data. The message reference identifier represents the target user to whom the candidate message data is to be addressed. That is, the target message recipient can enter the message reference identifier during the message generation process to determine who will receive the message. The message reference identifier includes a first message reference identifier and a second message reference identifier. The first message reference identifier is a message reference identifier that directs the message to a specific listener. For example, the first message reference identifier can be any one or more of a configured identity tag, a user nickname, and a user name. For example, in one embodiment, the first message reference identifier can be an identity tag, such as "Dad / Mom / Grandpa / Grandma." In another embodiment, the first message reference identifier can be a user name, such as "Zhang San / Li Si / Wang Wu." Furthermore, in another embodiment, the first message reference identifier can be a user nickname customized by the user. Optionally, in one embodiment, after waking up the terminal and entering the message flow, the target message recipient voice-enters "Leave a message to User A" into the terminal. The terminal then recognizes the message reference identifier as "User A."

[0117] In addition, the message pointing identifier also includes a second message identifier for leaving a message to all family user objects.

[0118] Specifically, after determining the message pointing identifier, the terminal also maps and binds the message pointing identifier and the voiceprint information of the target pointing object corresponding to the message pointing identifier to obtain a preset voiceprint feature, and when the target pointing object triggers the message playback, the preset voiceprint feature corresponding to the message pointing identifier is matched with the target voiceprint feature of the target interactive object to determine whether the candidate message data associated with the target pointing object is played as the target message data.

[0119] The message content data is a portion of the message input data that represents the actual message content.

[0120] Specifically, after obtaining the voice input data, the terminal creates an initial message file in the message database, performs audio recognition on the voice input data, matches the voice input data with each preset pointing identifier, and uses the preset pointing identifier whose similarity exceeds the pointing identifier similarity as the message pointing identifier corresponding to the voice input data to determine the object to be listened to corresponding to the voice input data, and maps and binds the message pointing identifier and the voiceprint information of the object to be listened to.

[0121] Specifically, the terminal also performs audio activity detection on the voice input data to extract the message content data from the voice input data. That is, the terminal performs feature extraction on the voice input data to obtain voice input features corresponding to the voice input data. The voice input features may be any one or more of the following: audio short-time energy, audio zero-crossing rate, and Mel-frequency cepstral coefficients associated with the voice input data.

[0122] Specifically, the terminal also calculates the window feature value of each sliding analysis window based on the voice input feature, that is, the terminal pre-sets each sliding analysis window for the voice input data, and calculates the voice feature value within each sliding analysis window as the window feature value of the sliding analysis window.

[0123] Specifically, the terminal also pre-sets a window detection threshold for distinguishing voice activity from non-voice activity, and determines the target analysis window corresponding to the voice input data based on the window characteristic value and the preset window detection threshold, that is, the terminal determines the sliding analysis window whose window characteristic value is greater than or equal to the preset window detection threshold as the target analysis window, wherein the target analysis window is a sliding analysis window that characterizes the presence of voice activity.

[0124] Specifically, after determining the target analysis window, the terminal uses the target analysis window to extract the message content data from the voice input data. That is, after determining the target analysis window, the terminal extracts the window audio data in each target analysis window, performs voice content recognition on the window audio data, and obtains the message content data.

[0125] Optionally, in a specific embodiment, the terminal can also input the acquired voice input data into a cloud server, which then performs audio activity detection and natural language recognition on the voice input data to obtain the message content data in the voice input data, and generates candidate message data based on the message content data and the message pointing identifier. When any voice interaction device triggers a message playback request, the corresponding candidate message data is sent to the voice interaction device as the target message data for playback. This reduces the terminal's computing power requirements and improves terminal operation smoothness.

[0126] Specifically, when generating candidate message data, the terminal also generates different candidate message data according to different message pointing identifiers.

[0127] Optionally, if the terminal determines that the message pointing identifier is the first message identifier, it generates a first candidate message that can only be listened to by the interactive object pointed to by the first message identifier based on the preset voiceprint feature corresponding to the first message identifier and the message content data, that is, the terminal associates and maps the preset voiceprint feature and the message content data and stores it as the first candidate message, and in the subsequent message playback process, determines whether to play the first candidate message as the target message data through the preset voiceprint feature and the target voiceprint feature of the target interactive object.

[0128] Optionally, if the terminal determines that the message pointing identifier is a second message identifier, a second candidate message that can be answered by any user is generated according to the second message identifier and the message content data.

[0129] Specifically, after the terminal or cloud server obtains the message pointing identifier and the message content data, it generates candidate message data based on the message pointing identifier and the message content data. That is, the terminal obtains the candidate message file corresponding to the message pointing identifier, stores the message content data into the candidate message file, obtains the candidate message data, and stores the candidate message data in the message database of the cloud server.

[0130] In this embodiment, the terminal responds to a message generation request and obtains the message input data associated with the message generation request; and generates candidate message data based on the message direction identifier and message content data in the message input data. This enables the voice interaction device and the mobile message terminal cloud server to generate different types of message data based on the message direction identifier and message content data, thereby optimizing the message generation effect.

[0131] like Figure 4 As shown, Figure 4 This is a flow chart of an embodiment of generating candidate message data based on text input data in the voice interaction method provided in an embodiment of the present application. Specifically, in this embodiment, the voice interaction method includes steps 401 to 403:

[0132] 401. If the message input data is text input data, obtain text content data in the text input data;

[0133] 402. Generate synthesized message data based on the text content data and the target message tone;

[0134] 403. Generate candidate message data according to the synthesized message data and the message pointing identifier.

[0135] Based on the above embodiment, in this embodiment, the target message object can also use the voice interaction device or mobile message terminal to input text input data as message input data, and drive the voice interaction device or mobile message terminal to generate candidate message data that can be played by voice based on the text input data.

[0136] Specifically, during operation, the terminal responds to a message generation request, wherein the message generation request is an operation request that drives the terminal to generate candidate message data. The triggering method of the message generation request is not specifically limited herein. The message generation request can be triggered by the target message recipient clicking a message button on a preset message interface to trigger the message generation request.

[0137] Specifically, after triggering a message generation request, the terminal parses the message generation request and determines the text message type corresponding to the message generation request, wherein the text message type is the message type corresponding to the message recipient determined in the text input message scenario. In a specific embodiment, the text message type includes a first message type for a specified recipient and a second message type that can be received by any object.

[0138] Specifically, after determining the text message type, the terminal further receives text data input by the target message recipient and obtains text content data in the text input data, wherein the text content data is the message content data representing the target message recipient.

[0139] Specifically, after obtaining the text content data for candidate message data, the terminal also generates synthesized message data based on the text content data and the target message timbre. The synthesized message data is message audio data representing the text content data converted into a speech pattern using the target message timbre. The target message timbre is the timbre data used for speech synthesis of the text content data. The target message timbre can be a preset timbre selected by the target message recipient or a cloned timbre of the target message recipient.

[0140] Specifically, after generating the synthesized message data, the terminal generates candidate message data based on the synthesized message data and the message pointing identifier. Specifically, the terminal determines the target recipient of the message to be listened to based on the message pointing identifier, creates a candidate message folder corresponding to the message pointing identifier on a cloud server or local storage area, uploads the synthesized message data to the candidate message folder, and stores it, thereby obtaining the candidate message data.

[0141] Optionally, in a specific embodiment, the terminal can also input the acquired text input data into a cloud server, which then performs audio synthesis on the text input data to obtain synthesized message data corresponding to the text input data. The server then generates candidate message data based on the synthesized message data and a message pointing identifier, and when any voice interaction device triggers a message playback request, sends the corresponding candidate message data to the voice interaction device as the target message data for playback. This reduces the computing power requirements of the terminal and improves terminal operation smoothness.

[0142] In this embodiment, the terminal determines if the message input data is text input data, obtains the text content data from the text input data, generates synthesized message data based on the text content data and the target message timbre, and generates candidate message data based on the synthesized message data and the message pointing identifier. This allows for the generation of voice message data using multiple input methods, enriching the message generation scenarios and increasing the fun of message interaction.

[0143] In order to better implement the voice interaction method in the embodiment of the present application, based on the voice interaction method, the embodiment of the present application also provides a voice interaction device, such as Figure 5 As shown, Figure 5 This is a structural diagram of an embodiment of a voice interaction device provided in an embodiment of the present application. Specifically, the voice interaction device 500 includes:

[0144] The data acquisition module 501 is configured to respond to a voice interaction request and acquire voice interaction data of the voice interaction request;

[0145] A state acquisition module 502 is configured to acquire object state information of a target interaction object corresponding to the voice interaction request;

[0146] The message playing module 503 is configured to play target message data associated with the voice interaction data when the voice interaction data meets a preset message triggering condition.

[0147] In a possible implementation of this embodiment, when the voice interaction data and the object status information meet a preset message triggering condition, the voice interaction device plays the target message data associated with the target interaction object, including:

[0148] Extracting voiceprint features based on the target wake-up word in the voice interaction data to obtain target voiceprint features;

[0149] Matching the target voiceprint feature with a preset voiceprint feature to obtain a target identity corresponding to the target voiceprint feature;

[0150] If the object status information and the target identity meet the preset message triggering condition, the target message data associated with the target identity is played.

[0151] In a possible implementation of this embodiment, if the object status information and the target identity meet the preset message triggering condition, the voice interaction device plays the target message data associated with the target identity, including:

[0152] If the object status information is the first object status information and there is a first candidate message associated with the target identity, obtaining a voiceprint confirmation password associated with the preset message triggering condition;

[0153] Acquire updated voice data associated with the voiceprint confirmation password, and updated voiceprint features in the updated voice data;

[0154] If the updated voiceprint feature is the same as the target voiceprint feature, the first candidate message is determined as the target message data, and the target message data is played.

[0155] In a possible implementation of this embodiment, if the object status information and the target identity meet the preset message triggering condition, the voice interaction device plays the target message data associated with the target identity, including:

[0156] If the object status information is the first object status information and the target identity meets the preset message triggering condition, generating a message acquisition request according to the device identification information and the target identity;

[0157] The target message data associated with the target identity identifier in the message database is read according to the message acquisition request, and the target message data is played.

[0158] In a possible implementation of this embodiment, before the voice interaction device plays the target message data associated with the voice interaction data, the method further includes:

[0159] Responding to a message generation request, obtaining message input data associated with the message generation request;

[0160] Candidate message data is generated according to the message pointing identifier and message content data in the message input data.

[0161] In a possible implementation of this embodiment, the message input data in the voice interaction device includes voice input data;

[0162] The generating of candidate message data according to the message pointing identifier and message content data in the message input data includes:

[0163] If the message input data is voice input data, performing audio recognition on the voice input data to obtain a message pointing identifier corresponding to the voice input data;

[0164] Performing activity detection on the voice input data to obtain message content data of the voice input data;

[0165] Candidate message data is generated according to the message pointing identifier and the message content data.

[0166] In a possible implementation of this embodiment, the voice interaction device performs activity detection on the voice input data to obtain message content data of the voice input data, including:

[0167] Performing feature extraction on the voice input data to obtain voice input features;

[0168] Calculating the window feature value of each sliding analysis window according to the speech input feature;

[0169] Determining a target analysis window corresponding to the speech input data according to the window feature value and a preset window detection threshold;

[0170] The target analysis window is used to extract message content data from the voice input data.

[0171] In a possible implementation of this embodiment, the message pointing identifier in the voice interaction device includes a first message identifier and a second message identifier, and the candidate message data includes a first candidate message and a second candidate message;

[0172] The generating of candidate message data according to the message pointing identifier and the message content data includes:

[0173] If the message pointing identifier is the first message identifier, generating a first candidate message according to the preset voiceprint feature corresponding to the first message identifier and the message content data;

[0174] If the message pointing identifier is a second message identifier, a second candidate message is generated according to the first message identifier and the message content data.

[0175] In a possible implementation of this embodiment, the message input data in the voice interaction device includes text input data;

[0176] The generating of candidate message data according to the message pointing identifier and the message content data includes:

[0177] If the message input data is text input data, then obtaining text content data in the text input data;

[0178] generating synthesized message data according to the text content data and the target message tone;

[0179] Candidate message data is generated according to the synthesized message data and the message pointing identifier.

[0180] In this embodiment, the voice interaction device obtains the voice interaction data of the voice interaction request by responding to the voice interaction request; obtains the object status information of the target interaction object corresponding to the voice interaction request; and plays the target message data associated with the target interaction object when the voice interaction data and the object status information meet the preset message triggering condition. This is achieved by obtaining the voice interaction data of the target interaction object and detecting the object status information of the target interaction object when the target interaction object conducts voice interaction with the smart home appliance. When the voice interaction data and the object status information meet the preset message triggering condition indicating that the corresponding target message data can be effectively listened to, the target message data associated with the target interaction object is automatically played, thereby realizing the message interaction operation using the smart home appliance, improving the success rate of users receiving messages, and effectively reducing the operational burden of users listening to messages, thereby expanding the intelligence level of the smart home appliance.

[0181] The embodiment of the present invention also provides a voice interaction device, such as Figure 6 As shown, Figure 6 This is a schematic diagram of the structure of an embodiment of the voice interaction device provided in the embodiments of the present application.

[0182] The voice interaction device integrates any one of the voice interaction devices provided in the embodiments of the present invention, and the voice interaction device includes:

[0183] one or more processors;

[0184] Memory; and

[0185] One or more applications, wherein the one or more applications are stored in the memory and configured so that the processor executes the steps of the voice interaction method described in any of the above-mentioned voice interaction method embodiments.

[0186] Specifically, the voice interaction device may include one or more processing core processors 601, one or more computer-readable storage media memories 602, a power supply 603, an input unit 604, and other components. Those skilled in the art will understand that Figure 6 The structure of the voice interaction device shown in the figure does not constitute a limitation on the voice interaction device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0187] Processor 601 is the control center of the voice interaction device. It uses various interfaces and lines to connect the various parts of the entire voice interaction device. By running or executing software programs and / or modules stored in memory 602 and calling data stored in memory 602, it performs various functions of the voice interaction device and processes data, thereby monitoring the voice interaction device as a whole. Optionally, processor 601 may include one or more processing cores; preferably, processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly handles wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into processor 601.

[0188] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the voice interaction device, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.

[0189] The voice interaction device also includes a power supply 603 for supplying power to various components. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 603 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0190] The voice interaction device may further include an input unit 604, which may be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0191] Although not shown, the voice interaction device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 601 in the voice interaction device will load the executable files corresponding to the processes of one or more application programs into the memory 602 according to the following instructions, and the processor 601 will run the application programs stored in the memory 602 to implement various functions as follows:

[0192] Responding to a voice interaction request, obtaining voice interaction data of the voice interaction request;

[0193] Obtaining object state information of a target interaction object corresponding to the voice interaction request;

[0194] When the voice interaction data and the object status information meet a preset message triggering condition, target message data associated with the target interaction object is played.

[0195] To this end, an embodiment of the present invention provides a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps of any of the voice interaction methods provided in the embodiments of the present invention. For example, the computer program loaded by the processor may execute the following steps:

[0196] Responding to a voice interaction request, obtaining voice interaction data of the voice interaction request;

[0197] Obtaining object state information of a target interaction object corresponding to the voice interaction request;

[0198] When the voice interaction data and the object status information meet a preset message triggering condition, target message data associated with the target interaction object is played.

[0199] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of other embodiments above and will not be repeated here.

[0200] In specific implementation, the above units or structures can be implemented as independent entities, or can be arbitrarily combined to implement as the same or several entities. The specific implementation of the above units or structures can refer to the previous method embodiments and will not be repeated here.

[0201] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0202] The above is a detailed introduction to a voice interaction method provided in an embodiment of the present application. Specific embodiments are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A voice interaction method, characterized in that: The voice interaction method includes: Responding to a voice interaction request, obtaining voice interaction data of the voice interaction request; Obtaining object state information of a target interaction object corresponding to the voice interaction request; When the voice interaction data and the object status information meet a preset message triggering condition, target message data associated with the target interaction object is played.

2. The voice interaction method according to claim 1, characterized in that: When the voice interaction data and the object status information meet a preset message triggering condition, playing the target message data associated with the target interaction object includes: Extracting voiceprint features based on the target wake-up word in the voice interaction data to obtain target voiceprint features; Matching the target voiceprint feature with a preset voiceprint feature to obtain a target identity corresponding to the target voiceprint feature; If the object status information and the target identity meet the preset message triggering condition, the target message data associated with the target identity is played.

3. The voice interaction method according to claim 2, characterized in that: If the object status information and the target identity meet a preset message triggering condition, playing the target message data associated with the target identity includes: If the object status information is the first object status information and there is a first candidate message associated with the target identity, obtaining a voiceprint confirmation password associated with the preset message triggering condition; Acquire updated voice data associated with the voiceprint confirmation password, and updated voiceprint features in the updated voice data; If the updated voiceprint feature is the same as the target voiceprint feature, the first candidate message is determined as the target message data, and the target message data is played.

4. The voice interaction method according to claim 2, wherein: If the object status information and the target identity meet a preset message triggering condition, playing the target message data associated with the target identity includes: If the object status information is the first object status information and the target identity meets the preset message triggering condition, generating a message acquisition request according to the device identification information and the target identity; The target message data associated with the target identity identifier in the message database is read according to the message acquisition request, and the target message data is played.

5. The voice interaction method according to claim 1, wherein: Before playing the target message data associated with the voice interaction data, the method further includes: Responding to a message generation request, obtaining message input data associated with the message generation request; Candidate message data is generated according to the message pointing identifier and message content data in the message input data.

6. The voice interaction method according to claim 5, characterized in that: The message input data includes voice input data; The generating of candidate message data according to the message pointing identifier and message content data in the message input data includes: If the message input data is voice input data, performing audio recognition on the voice input data to obtain a message pointing identifier corresponding to the voice input data; Performing activity detection on the voice input data to obtain message content data of the voice input data; Candidate message data is generated according to the message pointing identifier and the message content data.

7. The voice interaction method according to claim 5, characterized in that: The performing activity detection on the voice input data to obtain message content data of the voice input data includes: Performing feature extraction on the voice input data to obtain voice input features; Calculating the window feature value of each sliding analysis window according to the speech input feature; Determining a target analysis window corresponding to the speech input data according to the window feature value and a preset window detection threshold; The target analysis window is used to extract message content data from the voice input data.

8. The voice interaction method according to claim 6, characterized in that: The message pointing identifier includes a first message identifier and a second message identifier, and the candidate message data includes a first candidate message and a second candidate message; The generating of candidate message data according to the message pointing identifier and the message content data includes: If the message pointing identifier is the first message identifier, generating a first candidate message according to the preset voiceprint feature corresponding to the first message identifier and the message content data; If the message pointing identifier is a second message identifier, a second candidate message is generated according to the first message identifier and the message content data.

9. The voice interaction method according to claim 5, characterized in that: The message input data includes text input data; The generating of candidate message data according to the message pointing identifier and the message content data includes: If the message input data is text input data, then obtaining text content data in the text input data; generating synthesized message data according to the text content data and the target message tone; Candidate message data is generated according to the synthesized message data and the message pointing identifier.

10. A voice interaction device, characterized in that: The voice interaction device includes: A data acquisition module is configured to respond to a voice interaction request and acquire voice interaction data of the voice interaction request; A state acquisition module is configured to obtain object state information of a target interaction object corresponding to the voice interaction request; The message playing module is configured to play target message data associated with the voice interaction data when the voice interaction data meets a preset message triggering condition.

11. A voice interaction device, characterized in that: The voice interaction device includes: one or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps of the voice interaction method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps of the voice interaction method according to any one of claims 1 to 9.