Voice interaction method and device, computer equipment and storage medium

Through voice interaction methods and generative large models, sensor devices can directly interact with users, solving the problems of high cost and poor user experience of traditional sensor devices, and realizing efficient utilization of sensor data and improved user experience.

CN121459797APending Publication Date: 2026-02-03BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411046465.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing sensor devices merely digitize traditional components, resulting in high costs and a poor user experience, failing to effectively improve users' quality of life.

Method used

By using voice interaction, user commands are received, generative large models are used to generate descriptive labels, and answer information is played, enabling direct interaction between sensor data and users, thereby improving the convenience of data utilization and user experience.

Benefits of technology

Sensor data is no longer just a substitute for mechanical components, but rather serves as parameters for interactive devices, providing targeted answers and greatly enhancing user experience and device connectivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459797A_ABST
    Figure CN121459797A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction method and device, computer equipment and a storage medium. The method comprises the steps that a user instruction input by a user through a voice assistant is received, and the user instruction comprises event problem information associated with scene information; answer information corresponding to the event question information is played through the voice assistant, and the answer information is used for describing sensing data collected by a sensor in a scene indicated by the scene information. According to the voice interaction method provided by the invention, the application value of sensor data can be greatly improved, and the actual experience of a user is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of sensor data processing technology, and in particular to a voice interaction method, device, computer equipment, and storage medium. Background Technology

[0002] With the gradual maturation and promotion of IoT technology, sensor devices such as smart cameras and smart door locks with voice or peephole functions have gradually entered people's daily lives and are becoming important means for users to understand the surrounding environment of their homes and their own life trajectory.

[0003] In related technologies, the design and application ideas of sensor devices, as mentioned above, are usually limited to digitizing traditional components, such as using cameras to replace traditional peepholes or speakers to replace traditional doorbells. However, while these digitized traditional components are more expensive, they do not bring much substantial change to users' lives, and the user experience is not ideal. Summary of the Invention

[0004] To overcome the problems existing in related technologies, this disclosure provides a voice interaction method, apparatus, computer device, and storage medium.

[0005] The first aspect of this disclosure provides a voice interaction method, the method comprising:

[0006] Receive user commands input by the user through a voice assistant, wherein the user commands include event and problem information associated with scene information;

[0007] The voice assistant plays the answer information corresponding to the event question information, wherein the answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

[0008] Optionally, the method further includes:

[0009] The event question information is sent to the cloud, so that the cloud can determine the answer information corresponding to the event question information based on the recorded description tags, and return the answer information to the interactive device;

[0010] The cloud is used to receive sensing data collected by the sensors and generate a description tag for the sensing data based on the sensing data. The description tag is used to characterize the content of the sensing data.

[0011] Optionally, the event issue information and the description label include text-based data, and the sensing data includes data in at least one of the following modalities:

[0012] Image modality, video modality, audio modality.

[0013] A second aspect of this disclosure provides a voice interaction method, the method comprising:

[0014] Receive sensor data collected by the sensor, generate a description label based on the sensor data, and record the description label, wherein the description label is used to characterize the content of the sensor data;

[0015] In response to receiving event question information sent by the interactive device, the system determines the answer information corresponding to the event question information based on the recorded description tags, and returns the answer information to the interactive device so that the interactive device can play the answer information through a voice assistant. The event question information is associated with scene information and is indicated by user commands input by the user based on the voice assistant of the interactive device. The answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

[0016] Optionally, determining the answer information corresponding to the event question information based on the recorded description tags includes:

[0017] The description tag corresponding to the event question information is determined from the recorded description tags, and the answer information corresponding to the event question information is determined based on the generative large model according to the description tag corresponding to the event question information, wherein the length of the description tag is less than the length of the answer information.

[0018] Optionally, generating descriptive labels based on the sensing data includes:

[0019] The sensor data is input into a generative large model to obtain descriptive labels for the sensor data.

[0020] Optionally, inputting the sensing data into a generative large model to obtain descriptive labels for the sensing data includes:

[0021] The sensor data is input into a generative large model so that the generative large model responds to the inclusion of anomalous events in the sensor data and determines the descriptive labels of the sensor data based on the content of the anomalous events.

[0022] Optionally, determining the descriptive label based on the content of the anomalous event includes:

[0023] Based on the content of the abnormal event, a multi-level descriptive label is determined from high to low, wherein the multi-level descriptive label from high to low is used to characterize the content of the abnormal event;

[0024] And / or, determine the description label based on the start time, end time, and content of the abnormal event.

[0025] Optionally, the generative large model includes a multimodal generative large model;

[0026] The event problem information and the description label include text-based data, and the sensor data includes data in at least one of the following modalities:

[0027] Image modality, video modality, audio modality.

[0028] Optionally, the multimodal generative large model includes a multimodal generative large model obtained by reinforcement training based on the application environment of the sensor.

[0029] Optionally, the method further includes:

[0030] In response to receiving event problem information sent by the interactive device, an upload command is sent to the sensor in the scene indicated by the scene information associated with the event problem information, so that the sensor sends the currently collected sensing data to the cloud.

[0031] Optionally, the method further includes:

[0032] Receive currently collected sensor data uploaded by the sensor;

[0033] The sensor data currently collected by the sensor and the event problem information are input into the generative large model so that the generative large model can determine the description label of the sensor data based on the sensor data currently collected by the sensor and the event problem information.

[0034] The answer information is determined based on the description label, and the answer information is returned to the interactive device.

[0035] A third aspect of this disclosure provides a voice interaction device, the device comprising:

[0036] The instruction receiving module is used to receive user instructions input by the user through a voice assistant, wherein the user instructions include event and problem information associated with scene information;

[0037] The answer playback module is used to play the answer information corresponding to the event question information through the voice assistant, wherein the answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

[0038] Optionally, the device further includes:

[0039] The cloud interaction module is used to send the event question information to the cloud, so that the cloud can determine the answer information corresponding to the event question information based on the recorded description tags, and return the answer information to the interaction device;

[0040] The cloud is used to receive sensing data collected by the sensors and generate a description tag for the sensing data based on the sensing data. The description tag is used to characterize the content of the sensing data.

[0041] Optionally, the event problem information and the description label include text modal data, and the sensing data includes data in at least one of the following modalities: image modal, video modal, and audio modal.

[0042] A fourth aspect of this disclosure provides a voice interaction device, the device comprising:

[0043] The tag determination module is used to receive sensing data collected by the sensor, generate a description tag based on the sensing data, and record the description tag, wherein the description tag is used to characterize the content of the sensing data;

[0044] The answer return module is used to respond to receiving event question information sent by the interactive device, determine the answer information corresponding to the event question information according to the recorded description tag, and return the answer information to the interactive device so that the interactive device can play the answer information through the voice assistant. The event question information is associated with scene information and is indicated by user commands input by the user based on the voice assistant of the interactive device. The answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

[0045] Optionally, when determining the answer information corresponding to the event question information based on the recorded description tags, the method is used to:

[0046] The description tag corresponding to the event question information is determined from the recorded description tags, and the answer information corresponding to the event question information is determined based on the generative large model according to the description tag corresponding to the event question information, wherein the length of the description tag is less than the length of the answer information.

[0047] Optionally, when generating the description label based on the sensing data, the step is used to:

[0048] The sensor data is input into a generative large model to obtain descriptive labels for the sensor data.

[0049] Optionally, when inputting the sensing data into a generative large model to obtain descriptive labels for the sensing data, the method is used to:

[0050] The sensor data is input into a generative large model so that the generative large model responds to the inclusion of anomalous events in the sensor data and determines the descriptive labels of the sensor data based on the content of the anomalous events.

[0051] Optionally, when determining the description tag based on the content of the anomalous event, the method is used to:

[0052] Based on the content of the abnormal event, a multi-level descriptive label is determined from high to low, wherein the multi-level descriptive label from high to low is used to characterize the content of the abnormal event;

[0053] And / or, determine the description label based on the start time, end time, and content of the abnormal event.

[0054] Optionally, the generative large model includes a multimodal generative large model;

[0055] The event problem information and the description label include text modal data, and the sensor data includes data in at least one of the following modalities: image modal, video modal, and audio modal.

[0056] Optionally, the multimodal generative large model includes a multimodal generative large model obtained by reinforcement training based on the application environment of the sensor.

[0057] Optionally, the device further includes:

[0058] The sensor control module is used to respond to receiving event problem information sent by the interactive device, and send an upload command to the sensor in the scene indicated by the scene information associated with the event problem information, so that the sensor sends the currently collected sensing data to the cloud.

[0059] Optionally, the device further includes a response module for performing the following steps:

[0060] Used to receive currently acquired sensor data uploaded by the sensor;

[0061] The sensor data currently collected by the sensor and the event problem information are input into the generative large model so that the generative large model can determine the description label of the sensor data based on the sensor data currently collected by the sensor and the event problem information.

[0062] The answer information is determined based on the description label, and the answer information is returned to the interactive device.

[0063] The fifth aspect of this disclosure provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the method described in the first or second aspect.

[0064] A sixth aspect of this disclosure provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in the first or second aspect.

[0065] The seventh aspect of this disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in the first or second aspect.

[0066] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:

[0067] In this embodiment, the interactive device receives user commands input by a user through a voice assistant, including event question information associated with scene information. The voice assistant then plays the answer information corresponding to the event question information. Based on this, the interactive device can effectively become a real-world assistant for the user, with the sensors acting as the sensory organs of this assistant. The data collected by the sensors is no longer used merely as an electronic replacement for traditional mechanical components (such as a peephole), but rather as parameters for interaction between the interactive device and the user. The collected data can be returned to the user in the form of answer information, specifically addressing questions raised by the user during daily use. This fundamentally changes the way sensor data is utilized and the way sensors interact with users, greatly improving the ease of acquiring sensor data and its utilization in users' lives, enhancing user experience and inter-device connectivity.

[0068] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0069] The accompanying drawings, which are incorporated in and form part of this disclosure, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0070] Figure 1 This is a flowchart illustrating a voice interaction method through some exemplary embodiments.

[0071] Figure 2 These are application scenario diagrams illustrating a voice interaction method through some exemplary embodiments.

[0072] Figure 3 These are step diagrams illustrating a voice interaction method using some exemplary embodiments.

[0073] Figure 4 These are multi-party interaction diagrams illustrating a voice interaction method using some exemplary embodiments.

[0074] Figure 5 This is a block diagram illustrating a voice interaction device through some exemplary embodiments.

[0075] Figure 6 This is a block diagram illustrating another voice interaction device through some exemplary embodiments.

[0076] Figure 7 These are hardware structure diagrams of a computer device illustrating some exemplary embodiments. Detailed Implementation

[0077] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0078] With the gradual maturation and promotion of IoT technology, sensor devices such as smart cameras and smart door locks with voice or peephole functions have gradually entered people's daily lives and are becoming important means for users to understand the surrounding environment of their homes and their own life trajectory.

[0079] In related technologies, the design and application ideas of sensor devices, as mentioned above, are usually limited to digitizing traditional components, such as using cameras to replace traditional peepholes or speakers to replace traditional doorbells. However, while these digitized traditional components are more expensive, they do not bring much substantial change to users' lives, and the user experience is not ideal.

[0080] In view of the above, this disclosure provides a voice interaction method, apparatus, computer device, and storage medium.

[0081] The embodiments of this disclosure will now be described in detail from both the interactive device and cloud (server) perspectives.

[0082] The first aspect of this disclosure provides a voice interaction method applicable to interactive devices. Please refer to [link to relevant documentation]. Figure 1 It may include steps S101 and S102.

[0083] Step S101: Receive user instructions input by the user through the voice assistant, wherein the user instructions include event problem information associated with the scene information.

[0084] A voice assistant can be a built-in function of an interactive device, that is, a function that provides services to users through voice interaction, in order to achieve a functional experience similar to a real assistant. For example, the interactive device can be an electronic device equipped with voice interaction functions (i.e., a voice assistant), such as a smart speaker, a mobile phone, or an in-vehicle system. In the above steps, the voice assistant can determine the event or question information indicated by the user's speech based on speech recognition technology, and this event or question information is associated with the context information.

[0085] The following provides an illustrative example of the concepts of scenario information, scenario, event question information, answer information, and application environment that appear in this disclosure.

[0086] Please see Figure 2 The example illustrates an optional application architecture. The sensors include cameras (e.g., smart locks with peephole functionality, or dedicated cameras) for capturing images of the area outside a user's home. These cameras are connected to the cloud via a gateway (e.g., a router, switch, etc., which primarily manages network connectivity; it may be omitted or replaced by other types of devices). In addition, the gateway connects to interactive devices, such as smart speakers, mobile phones, and computers, for interacting with the user, receiving user commands, and presenting information to the user.

[0087] It should be understood that in this example, the cloud can include any device with model deployment and computing capabilities. This can be a server room deployed on the public internet, a host deployed on a local area network, or even another module deployed in the same device as the sensor. Similarly, the interactive device can also be another module deployed in the same device as the sensor (for example, a smart door lock can also have audio interaction or image display functions). In the following description of the implementation methods, the cloud, interactive device, and sensor may be described as independent terminals, or expressions such as "cloud-sensor architecture" may appear. However, this disclosure does not limit their specific forms of existence. When the same device has the functions provided by the cloud, interactive device, and sensor, the device can be considered as the cloud, interactive device, and sensor simultaneously; that is, the device has functional modules corresponding to the cloud, interactive device, and sensor.

[0088] Sensors have application environments. For example, smart door locks are typically designed to be mounted on front doors to identify visitors, people lingering at the entrance, and changes in objects. These specific scenarios arise from the application environment of smart door locks, and in these scenarios, sensors have adaptive functional requirements. A scenario can include elements such as time, location, and person identity. These elements can be used to determine the content of one or more events, or to indicate the sensing content collected by one or more sensors, such as "13:00," "courier," and "front door." Scenario information is the information composed of scenario elements, such as "Was a courier present at the front door at 13:00?" Event question information is question information related to scenario information, such as "Was a courier present at 1 PM?" Upon receiving this event question information, the interactive device can determine the corresponding answer, such as "Yes, at 1:32 PM, the courier placed a package at the front door." The answer information can be determined based on sensor data parsing algorithms, deep learning models, etc., or it can be uploaded to the cloud so that the cloud can determine the answer information and return it to the interactive device (which will be explained in detail later).

[0089] Step S102: Play the answer information corresponding to the event question information through the voice assistant, wherein the answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

[0090] For example, in the scenario described above, the sensors in the context of the "Will a delivery person arrive at the door at 13:00?" scenario could be smart door locks or human body sensors installed at the entrance door. The specific sensors included can be determined based on user pre-configuration, the device's automatic discovery function, or, of course, directly indicated by the event question information (e.g., the user directly asking about the sensor information collected by a particular sensor at a certain time). This will not be elaborated further here. After determining the answer information corresponding to the event question information, the interactive device can present the answer information describing the sensor data to the user through various methods such as auditory and visual means, for example, by playing the answer information based on the aforementioned voice assistant function.

[0091] In the above embodiments, the interactive device receives user commands input by the user through a voice assistant, including event question information associated with scene information, and then plays the answer information corresponding to the event question information through the voice assistant. Based on this, the interactive device can effectively become a real-world assistant for the user, with the sensors acting as the sensory organs of this assistant. The data collected by the sensors is no longer used merely as an electronic replacement for traditional mechanical components (such as a peephole), but rather as parameters for interaction between the interactive device and the user. The collected content can be returned to the user in the form of answer information, specifically addressing the problems raised by the user during daily use. This fundamentally changes the way sensor data is utilized and the way sensors interact with users, greatly improving the ease of acquiring sensor data and its utilization in users' lives, enhancing user experience and the interoperability between devices.

[0092] Furthermore, the method may also include:

[0093] The event question information is sent to the cloud, so that the cloud can determine the answer information corresponding to the event question information based on the recorded description tags, and return the answer information to the interactive device;

[0094] The cloud is used to receive sensing data collected by the sensors and generate a description tag for the sensing data based on the sensing data. The description tag is used to characterize the content of the sensing data.

[0095] For example, the cloud can receive sensor data collected by sensors, which can be sent by the sensors in real time or at regular intervals, or sent by the sensors in response to a certain instruction. In one example, receiving sensor data collected by sensors may include: in response to receiving event problem information sent by an interactive device, sending an upload instruction to the sensor in the scene indicated by the scene information associated with the event problem information, so that the sensor sends the currently collected sensor data to the cloud. The sensing commands can indicate event-related information. For example, if a user sends a message to a smart speaker asking "Has the delivery person arrived at the door yet?", the smart speaker or the cloud can analyze the message and determine that it is a sensing command (i.e., there is a need for the sensor to upload sensing data). Then, an upload command can be sent to the camera to start reporting the currently collected image information to the cloud. This upload process can have a preset execution duration (e.g., the camera uploads images collected within the next 10 seconds) or it can stop in response to a cutoff command issued by the cloud. The timing of the cutoff command can be determined by the cloud. For example, if the amount of sensing data received by the cloud in subsequent steps has reached the level required to determine the result (e.g., it has been determined that a delivery person is at the door), a cutoff command can be issued. This implementation reduces the device's power consumption and bandwidth consumption while providing users with controllable and purposeful sensing data collection, greatly enhancing the freedom of user interaction and improving the user experience.

[0096] Next, the cloud can obtain descriptive labels based on sensor data parsing algorithms, or input the sensor data into a generative large model to obtain descriptive labels, which are used to characterize the content of the sensor data.

[0097] Generative large-scale models refer to generative models based on large-scale corpora, such as GPT (Generative Pre-trained Transformer) and VQ-VAE (Vector Quantized Variational Autoencoder). Optionally, the generative large-scale models in this disclosure can be multimodal generative large-scale models. Multimodality, also known as joint modality, can refer to a combination of different perception modes, input methods, or media forms. Multimodal data can be data composed of at least two modalities such as speech, image, and text. These modalities can be interconnected and complementary, thereby jointly representing the deep information corresponding to the multimodal data.

[0098] When a generative large model can be a multimodal generative large model, the sensing data can contain data from multiple modalities, for example in Figure 2In the example shown, the sensor data uploaded by the camera to the cloud can include image information captured by the camera, audio information captured by the microphone module deployed on the camera, and human recognition information captured by the human body sensor deployed outside the door, etc., to generate descriptive labels. Furthermore, when a user issues a sensor command indicating event-related information, this event-related information can also be used as input to the generative big model, thereby enabling the generative big model to generate descriptive labels that are more closely aligned with the user's needs. That is, inputting the sensor data into the generative big model can include: inputting the sensor data and the event-related information indicated by the sensor command into the generative big model, so that the generative big model determines the descriptive labels based on the sensor data and the event-related information. Continuing the previous example, when a user issues a sensor command to the smart speaker asking "Has the delivery person arrived at the door?", the cloud can parse the data to determine that the user's goal is to determine whether a delivery person is outside the door. At this time, the generative big model can determine the descriptive result corresponding to "Is there a delivery person outside the door?" based on the two inputs: "Is there a delivery person outside the door?" and the sensor data uploaded by the sensor, thereby further improving the user experience. In summary, the event problem information and the description tags include text modal data, and the sensor data includes data in at least one of the following modalities: image modal, video modal, and audio modal.

[0099] This disclosure also provides various embodiments regarding label generation methods, which can be found in the second aspect and will not be repeated in the embodiments of the first aspect.

[0100] The second aspect of this disclosure provides a voice interaction method that can be applied to the cloud (or other devices capable of performing the same function). Please refer to [link to relevant documentation]. Figure 3 It may include steps S301 and S302.

[0101] Step S301: Receive sensing data collected by the sensor, generate a description label based on the sensing data, and record the description label, wherein the description label is used to characterize the content of the sensing data.

[0102] For example, the cloud can receive sensor data collected by sensors, which can be sent by the sensors in real time or at regular intervals, or sent by the sensors in response to a certain instruction. In one example, receiving sensor data collected by sensors may include: in response to receiving event problem information sent by an interactive device, sending an upload instruction to the sensor in the scene indicated by the scene information associated with the event problem information, so that the sensor sends the currently collected sensor data to the cloud.

[0103] The sensing commands can indicate event-related information. For example, if a user sends a message to a smart speaker asking "Has the delivery person arrived at the door yet?", the smart speaker or the cloud can analyze the message and determine that it is a sensing command (i.e., there is a need for the sensor to upload sensing data). Then, an upload command can be sent to the camera to start reporting the currently collected image information to the cloud. This upload process can have a preset execution duration (e.g., the camera uploads images collected within the next 10 seconds) or it can stop in response to a cutoff command issued by the cloud. The timing of the cutoff command can be determined by the cloud. For example, if the amount of sensing data received by the cloud in subsequent steps has reached the level required to determine the result (e.g., it has been determined that a delivery person is at the door), a cutoff command can be issued. This implementation reduces the device's power consumption and bandwidth consumption while providing users with controllable and purposeful sensing data collection, greatly enhancing the freedom of user interaction and improving the user experience.

[0104] Next, the cloud can obtain descriptive labels based on sensor data parsing algorithms. Alternatively, the sensor data can be input into a generative large model to obtain descriptive labels. That is, generating descriptive labels based on the sensor data can include: inputting the sensor data into a generative large model to obtain descriptive labels for the sensor data, wherein the descriptive labels are used to characterize the content of the sensor data.

[0105] Generative large-scale models refer to generative models based on large-scale corpora, such as GPT (Generative Pre-trained Transformer) and VQ-VAE (Vector Quantized Variational Autoencoder). Optionally, the generative large-scale models in this disclosure can be multimodal generative large-scale models. Multimodality, also known as joint modality, can refer to a combination of different perception modes, input methods, or media forms. Multimodal data can be data composed of at least two modalities such as speech, image, and text. These modalities can be interconnected and complementary, thereby jointly representing the deep information corresponding to the multimodal data.

[0106] When a generative large model can be a multimodal generative large model, the sensing data can contain data from multiple modalities, for example in Figure 2In the example shown, the sensor data uploaded by the camera to the cloud can include image information captured by the camera, audio information captured by the microphone module deployed on the camera, and human recognition information captured by the human body sensor deployed outside the door, etc., to generate descriptive labels. Furthermore, when a user issues a sensor command indicating event-related information, this event-related information can also be used as input to the generative big model, thereby enabling the generative big model to generate descriptive labels that are more closely aligned with the user's needs. That is, inputting the sensor data into the generative big model can include: inputting the sensor data and the event-related information indicated by the sensor command into the generative big model, so that the generative big model determines the descriptive labels based on the sensor data and the event-related information. Continuing the previous example, when a user issues a sensor command to the smart speaker asking "Has the delivery person arrived at the door?", the cloud can parse the data to determine that the user's goal is to determine whether a delivery person is outside the door. At this time, the generative big model can determine the descriptive result corresponding to "Is there a delivery person outside the door?" based on the two inputs: "Is there a delivery person outside the door?" and the sensor data uploaded by the sensor, thereby further improving the user experience. In summary, the event problem information and the description tags include text modal data, and the sensor data includes data in at least one of the following modalities: image modal, video modal, and audio modal.

[0107] In the process of generating descriptive labels based on the generative large model described above, if the sensor data is uploaded by the sensor after receiving event question information sent by the interactive device, the sensor data can be processed and descriptive labels obtained based on the following steps: receiving the currently collected sensor data uploaded by the sensor; inputting the currently collected sensor data and the event question information into the generative large model, so that the generative large model determines the descriptive labels of the sensor data based on the currently collected sensor data and the event question information; determining the answer information based on the descriptive labels, and returning the answer information to the interactive device. Details can be found in similar processes described in other parts of this disclosure. Based on this, the descriptive labels generated by the generative large model can be more targeted. For example, when a user asks "Is there a courier coming to my door?", the cloud can directly obtain the image currently collected by the smart lock and use scene elements such as "courier" as guiding information to specifically determine the descriptive labels, greatly improving the hit rate of the generated descriptive label results for the user's question (because the generative model can specifically generate descriptive labels for scene elements in the user's question information). It should be understood that the above process is not necessarily only applicable to the "current real-time upload" process. If the cloud or other devices store sensor information, the cloud can also retrieve this stored sensor information and generate descriptive tags based on the above steps, which will not be elaborated here.

[0108] In the aforementioned process, users can interact with the interactive terminal via voice (e.g., by asking questions to a smart speaker), or via touch, text input, etc. (e.g., by operating applications deployed on mobile devices). The subsequent process of "sensor uploading sensor data and cloud-based determination of descriptive labels" can be understood as the camera executing tasks according to user instructions. Combined with the generative big data model, the overall framework comprised of sensors and the generative big data model effectively acts as an assistant to the user, replacing their own senses to perform tasks according to user instructions and determine the descriptive labels for the collected data. This fundamentally changes the relationship between users and sensors, greatly enhancing the role of sensors in users' daily lives and increasing the amount of information provided by sensors and the relevance of that information to user needs.

[0109] Step S302: In response to receiving event question information sent by the interactive device, determine the answer information corresponding to the event question information according to the recorded description tag, and return the answer information to the interactive device so that the interactive device can play the answer information through the voice assistant. The event question information is associated with scene information and is indicated by user commands input by the user based on the voice assistant of the interactive device. The answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

[0110] Optionally, in this embodiment of the disclosure, the aforementioned multimodal generative large model includes a multimodal generative large model obtained through reinforcement training based on the application environment of the sensor. For example, in Figure 2 In the scenario shown, the sensor (i.e., the camera) is dedicated to capturing video data outside the user's home door. For common user groups, events that frequently occur outside the door may include "delivery person arriving", "stranger passing by the door", "stranger lingering at the door", "someone placing an object at the door", etc. Based on the above application scenarios and the common needs and events in these scenarios, the generative large model can be reinforced and trained in advance to obtain a universally anthropomorphic multimodal generative large model at a conventional training cost, thereby improving the accuracy of the multimodal generative large model in generating descriptive labels.

[0111] Once the descriptive labels are obtained, they can be presented directly to the user or stored, details of which will be elaborated later. It should be understood that this disclosure does not limit the specific meaning of the concept of "label," which can be keywords, descriptive statements, or other information used to characterize the content of sensor data, determined according to the model's own capabilities.

[0112] In summary, in this embodiment, the content collected by the sensor becomes the input to the generative large model, allowing the data collected by the sensor to be represented in the form of descriptive labels. This makes it easy to record and report to the user, fundamentally changing the application form of sensor data. Sensor data is transformed from traditional environmental detection data into parameters for generating descriptive labels. Even if the user does not participate in viewing the sensor data, the content contained in the data collected by the sensor will be summarized by the generative large model, thus serving as the basis for reporting the surrounding environment and the user's life trajectory to the user, greatly enhancing the application value of sensor data and the user's actual experience. Furthermore, in this embodiment, the descriptive labels for the sensor data collected by the sensor are actively determined by the generative large model, rather than simply determining the truth or falsehood of a preset event (i.e., not limited to judging whether a preset event A has occurred and notifying the user if it has). This greatly improves the richness of the events, ensuring that the final descriptive labels are determined based on the actual situation collected by the sensor and the model's capabilities (rather than based on preset events). This makes the generated descriptive labels closer to real-life situations, increasing the amount of information that can be reported when reporting the surrounding environment and the user's life trajectory to the user, further enhancing the user experience.

[0113] In some embodiments of this disclosure, for the steps after obtaining the description tags, for example, determining the answer information corresponding to the event question information based on the recorded description tags may include: determining the description tag corresponding to the event question information from the recorded description tags, and determining the answer information corresponding to the event question information based on the description tag corresponding to the event question information according to the generative large model, wherein the length of the description tag is less than the length of the answer information.

[0114] In other words, the descriptive tag can be a keyword or other descriptive phrase. For example, in the previous example, if the courier actually arrives at 3 PM, the descriptive tag determined by the generative big data model can be either "courier" or "3:00 PM" (of course, the generative big data model can also choose to directly generate descriptive text, or generate both descriptive tags and descriptive text simultaneously; furthermore, descriptive tags and their corresponding descriptive text can have a binding relationship, and multiple descriptive tags for the same event or the same instruction can also have a binding relationship). Then, the cloud can use a text-to-text model or other models (optionally, this model can be a generative big data model based on an attention mechanism, such as BERT).

[0115] (Bibliographic Encoder Representations from Transformers), Transformer, etc.) generate descriptive text based on descriptive tags. For example, based on the two tags "courier" and "15:00" which represent personnel and time information respectively and have a binding relationship, the descriptive text is determined to be "Hello, at 3:00 PM, a courier will arrive at the door." The interactive device then uses TTS (Text To Speech) technology (the voice assistant can have a built-in TTS algorithm) to read the descriptive text to the user, or directly displays the descriptive text on the display module of the interactive device.

[0116] In other words, after determining the descriptive tags corresponding to the event / question information, the generative large model generates answer information based on the descriptive tags (and may also combine them with the event / question information). Based on this, a more human-like and richer answer can be presented to the user, while also taking into account the storage capacity occupied by the descriptive tags. Since users typically do not pay attention to all events collected by sensors, but rather to a small portion of them—for example, a camera might detect dozens of people passing by the door on a single day—for user safety and other considerations, the generative large model can generate and record descriptive tags for each person passing by the door, but only for non-routine events (such as a delivery person arriving; details on this will be discussed later). This achieves a balance between device power consumption, storage space, and user experience.

[0117] As another example, the method may further include: recording the description tags. The step of generating description text based on the description tags may include:

[0118] In response to receiving a search instruction from a user, the system determines the description tag corresponding to the answer information indicated by the search instruction from the recorded description tags, generates the description text based on the description tag corresponding to the answer information, and sends the description text back to the user.

[0119] Continuing with the previous example, when the courier arrives at 3 PM, the user may not ask any questions or inform the cloud to announce the courier's arrival. In this case, the cloud can simply record the descriptive tags "courier" and "3:00 PM". However, when the user returns home from get off work, they might ask, "When did the courier arrive?" The cloud can then treat this question as a search command, identify "3:00 PM" as the tag associated with "courier" from the recorded tags, generate descriptive text based on this information, and present it to the user. As can be seen, the implementation method provided in this disclosure closely aligns with users' actual needs, significantly enhancing the auxiliary role of sensors in users' lives. This makes the sensor-cloud architecture a human-like assistant integrated into users' lives, greatly improving the user experience compared to traditional application and processing methods (e.g., in related technologies, users may only be able to select "please inform me when someone passes by the door" from preset events and then passively receive frequent event notifications).

[0120] In some embodiments of this disclosure, regarding the method of generating description tags, the step of inputting the sensing data into a generative large model to obtain description tags for the sensing data may include: inputting the sensing data into a generative large model so that the generative large model responds to the sensing data including anomalous events, and determines the description tags for the sensing data based on the content of the anomalous events.

[0121] In this context, "abnormal events" can be understood as the opposite of "normal events." "Normal events" can be identified by enabling generative models to recognize "abnormal events" through methods such as recording the frequency of occurrence, reinforcing the training of generative models, and receiving preset normal events input by manufacturers or users. A typical "normal event" is the natural movement of objects within the sensor's range, such as the natural swaying of power lines outside a door or plants. If this event occurs frequently and the user hasn't asked about it, it's more likely to be identified as a "normal event." Conversely, an event like "a delivery person arriving at the door" is more likely to be of interest to the user and is more likely to be identified as an "abnormal event" and recorded. It should be understood that "determining whether an event has occurred" and "identifying abnormal events" can be considered capabilities of the model itself, which can be determined based on the model's structure, the reinforcement training it undergoes during training, and preset prompts or events provided by manufacturers or users.

[0122] After identifying an abnormal event, for example, determining the descriptive tags based on the content of the abnormal event may include: determining a multi-level descriptive tag from high to low based on the content of the abnormal event, wherein the multi-level descriptive tag from high to low is used to characterize the content of the abnormal event.

[0123] The scope of events represented by multi-level descriptive tags, from high to low, can gradually decrease. Continuing the previous example, when a courier arrives at 3 PM, the descriptive tags can be sequentially defined as "Person," "Person's Location," and "Courier's Location." When a user subsequently asks about this event, the cloud can quickly recognize that the information the user wants—"When did the courier arrive?"—is actually related to "Person's Location," significantly accelerating subsequent searches. Furthermore, when a user asks "Was anyone parked at my door today?", the cloud can report all events corresponding to the tags under the "Person's Location" tag to the user, improving the orderliness of tag storage and making its storage structure more closely resemble real-world applications.

[0124] As another example, determining the description tag based on the content of the anomalous event may include: determining the description tag based on the start time, end time, and content of the anomalous event.

[0125] For example, the descriptive tags recorded for such abnormal events can be summary tags. For instance, if someone parks a bicycle outside at 1 PM and moves it at 4 PM, the cloud can record two descriptive tags: "Item" - "Bicycle" - "Placement" - "13:00" and "Item" - "Bicycle" - "Removal" - "16:00". "13:00" and "16:00" represent the start and end times of the event, while "bicycle" can be considered the event content. If a user later asks, "Was there a bicycle parked outside at 2 PM?", the cloud can confirm from the two recorded tags that the "bicycle" event occurred at 2 PM and has not yet ended. It can then confirm that the bicycle is still parked outside and generate descriptive text based on this result. This further enhances the balance between reducing device operating and storage pressure and improving user experience.

[0126] A brief introduction to the voice interaction method applied to the sensor is also provided here. The sensor can send the sensor-collected data to the cloud, so that the cloud can input the sensor data into a generative large model to obtain descriptive labels, which represent the content of the sensor data.

[0127] The process of the sensor sending sensor data to the cloud can be routine, periodic (i.e., sending sensor data to the cloud once every certain period of time), or it can be uploaded in response to other sensors detecting preset events (for example, if a human body sensor is deployed outside the door, the camera can upload the captured image to the cloud in response to the human body sensor recognizing a person passing by), or it can be uploaded in response to an upload command sent by the cloud or a sensor command sent by the user. Specific details regarding the above-mentioned voice interaction method applied to the sensor end can be found in the relevant sections of the first aspect of the implementation method, and will not be repeated here.

[0128] Next, combine Figure 4 The multi-party interaction diagram shown provides a comprehensive description of some embodiments of this disclosure.

[0129] First, in S401, the sensor can send sensing data to the cloud in response to some triggering conditions (as described above). In S402, the cloud uses a multimodal large model to determine and record descriptive labels that can characterize the content of the sensing data.

[0130] Next, the user may issue a user command indicating event problem information. This event problem information can be collected by the interactive device and sent to the cloud in S403. After receiving the event problem information, the cloud can determine the corresponding description tag for the event problem information based on the scene information associated with the event problem information from the recorded description tags (S404), and then generate answer information based on the retrieved description tags (S405), and return the answer information to the interactive device (S406). Finally, the interactive device can present the answer information in S407 to provide the user with feedback on the event problem information.

[0131] Corresponding to the embodiments of the foregoing methods, this disclosure also provides embodiments of the apparatus and the terminal to which it is applied.

[0132] A third aspect of this disclosure provides a voice interaction device; please refer to [link / reference]. Figure 5 The device includes:

[0133] The instruction receiving module 501 is used to receive user instructions input by the user through a voice assistant, wherein the user instructions include event problem information associated with scene information;

[0134] The answer playback module 502 is used to play the answer information corresponding to the event question information through the voice assistant, wherein the answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

[0135] Optionally, the device further includes:

[0136] The cloud interaction module is used to send the event question information to the cloud, so that the cloud can determine the answer information corresponding to the event question information based on the recorded description tags, and return the answer information to the interaction device;

[0137] The cloud is used to receive sensing data collected by the sensors and generate a description tag for the sensing data based on the sensing data. The description tag is used to characterize the content of the sensing data.

[0138] Optionally, the event problem information and the description label include text modal data, and the sensing data includes data in at least one of the following modalities: image modal, video modal, and audio modal.

[0139] The fourth aspect of this disclosure provides a voice interaction device; please refer to [link / reference]. Figure 6 The device includes:

[0140] The tag determination module 601 is used to receive sensing data collected by the sensor, generate a description tag based on the sensing data, and record the description tag, wherein the description tag is used to characterize the content of the sensing data;

[0141] The answer return module 602 is used to respond to receiving event question information sent by the interactive device, determine the answer information corresponding to the event question information according to the recorded description tag, and return the answer information to the interactive device so that the interactive device can play the answer information through the voice assistant. The event question information is associated with scene information and is indicated by user commands input by the user based on the voice assistant of the interactive device. The answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

[0142] Optionally, when determining the answer information corresponding to the event question information based on the recorded description tags, the method is used to:

[0143] The description tag corresponding to the event question information is determined from the recorded description tags, and the answer information corresponding to the event question information is determined based on the generative large model according to the description tag corresponding to the event question information, wherein the length of the description tag is less than the length of the answer information.

[0144] Optionally, when generating the description label based on the sensing data, the step is used to:

[0145] The sensor data is input into a generative large model to obtain descriptive labels for the sensor data.

[0146] Optionally, when inputting the sensing data into a generative large model to obtain descriptive labels for the sensing data, the method is used to:

[0147] The sensor data is input into a generative large model so that the generative large model responds to the inclusion of anomalous events in the sensor data and determines the descriptive labels of the sensor data based on the content of the anomalous events.

[0148] Optionally, when determining the description tag based on the content of the anomalous event, the method is used to:

[0149] Based on the content of the abnormal event, a multi-level descriptive label is determined from high to low, wherein the multi-level descriptive label from high to low is used to characterize the content of the abnormal event;

[0150] And / or, determine the description label based on the start time, end time, and content of the abnormal event.

[0151] Optionally, the generative large model includes a multimodal generative large model;

[0152] The event problem information and the description label include text modal data, and the sensor data includes data in at least one of the following modalities: image modal, video modal, and audio modal.

[0153] Optionally, the multimodal generative large model includes a multimodal generative large model obtained by reinforcement training based on the application environment of the sensor.

[0154] Optionally, the device further includes:

[0155] The sensor control module is used to respond to receiving event problem information sent by the interactive device, and send an upload command to the sensor in the scene indicated by the scene information associated with the event problem information, so that the sensor sends the currently collected sensing data to the cloud.

[0156] Optionally, the device further includes a response module for performing the following steps:

[0157] Used to receive currently acquired sensor data uploaded by the sensor;

[0158] The sensor data currently collected by the sensor and the event problem information are input into the generative large model so that the generative large model can determine the description label of the sensor data based on the sensor data currently collected by the sensor and the event problem information.

[0159] The answer information is determined based on the description label, and the answer information is returned to the interactive device.

[0160] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0161] The fifth aspect of this disclosure provides a computer program product including a computer program / instructions that, when executed by a processor, implement the method as described in the first or second aspect.

[0162] For the device embodiments and computer program product embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. Furthermore, the device embodiments described above are merely illustrative; the modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without any inventive effort.

[0163] Sixthly, the embodiments of the voice interaction device provided in this disclosure can be applied to computer devices. See also... Figure 7 The illustration exemplifies a hardware schematic of a computer device. For example, device 700 could be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0164] Device 700 may include one or more of the following components: processing component 701, memory 702, power supply component 703, multimedia component 704, audio component 705, input / output (I / O) interface 706, sensor component 707, and communication component 708.

[0165] Processing component 701 typically controls the overall operation of device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 701 may include one or more processors 709 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 701 may include one or more modules to facilitate interaction between processing component 701 and other components. For example, processing component 701 may include a multimedia module to facilitate interaction between multimedia component 704 and processing component 701.

[0166] Memory 702 is configured to store various types of data to support the operation of device 700. Examples of this data include instructions for any application or method operating on device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0167] The power supply component 703 provides power to the various components of the device 700. The power supply component 703 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 700.

[0168] Multimedia component 704 includes a screen that provides an output interface between the device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, swipe, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 704 includes a front-facing camera and / or a rear-facing camera. When the device 700 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0169] Audio component 705 is configured to output and / or input audio signals. For example, audio component 705 includes a microphone (MIC) configured to receive external audio signals when device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 702 or transmitted via communication component 708. In some embodiments, audio component 705 also includes a speaker for outputting audio signals.

[0170] I / O interface 706 provides an interface between processing component 701 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0171] Sensor assembly 707 includes one or more sensors for providing state assessments of various aspects of device 700. For example, sensor assembly 707 can detect the on / off state of device 700, the relative positioning of components such as the display and keypad of device 700, changes in the position of device 700 or a component of device 700, the presence or absence of user contact with device 700, the orientation or acceleration / deceleration of device 700, and temperature changes of device 700. Sensor assembly 707 may also include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 707 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 707 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0172] Communication component 708 is configured to facilitate wired or wireless communication between device 700 and other devices. Device 700 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G or 5G, or combinations thereof. In one exemplary embodiment, communication component 708 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 708 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IRDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0173] In an exemplary embodiment, device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the voice interaction method of the computer device described above.

[0174] In a seventh aspect, this disclosure also provides a non-transitory computer-readable storage medium including instructions, such as a memory 702 including instructions, which can be executed by a processor 709 of device 700 to complete the voice interaction method of the computer device. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0175] The foregoing has described specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0176] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention applied herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0177] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

[0178] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A voice interaction method, characterized in that, The method includes: Receive user commands input by the user through a voice assistant, wherein the user commands include event and problem information associated with scene information; The voice assistant plays the answer information corresponding to the event question information, wherein the answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

2. The voice interaction method according to claim 1, characterized in that, The method further includes: The event question information is sent to the cloud, so that the cloud can determine the answer information corresponding to the event question information based on the recorded description tags, and return the answer information to the interactive device; The cloud is used to receive sensing data collected by the sensors and generate a description tag for the sensing data based on the sensing data. The description tag is used to characterize the content of the sensing data.

3. The voice interaction method according to claim 2, characterized in that, The event problem information and the description label include text modal data, and the sensor data includes data in at least one of the following modalities: image modal, video modal, and audio modal.

4. A voice interaction method, characterized in that, The method includes: Receive sensor data collected by the sensor, generate a description label based on the sensor data, and record the description label, wherein the description label is used to characterize the content of the sensor data; In response to receiving event question information sent by the interactive device, the system determines the answer information corresponding to the event question information based on the recorded description tags, and returns the answer information to the interactive device so that the interactive device can play the answer information through a voice assistant. The event question information is associated with scene information and is indicated by user commands input by the user based on the voice assistant of the interactive device. The answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

5. The voice interaction method according to claim 4, characterized in that, Determining the answer information corresponding to the event question information based on the recorded description tags includes: The description tag corresponding to the event question information is determined from the recorded description tags, and the answer information corresponding to the event question information is determined based on the generative large model according to the description tag corresponding to the event question information, wherein the length of the description tag is less than the length of the answer information.

6. The voice interaction method according to claim 4, characterized in that, The step of generating descriptive labels based on the sensing data includes: The sensor data is input into a generative large model to obtain descriptive labels for the sensor data.

7. The voice interaction method according to claim 6, characterized in that, The step of inputting the sensing data into a generative large model to obtain descriptive labels for the sensing data includes: The sensor data is input into a generative large model so that the generative large model responds to the inclusion of anomalous events in the sensor data and determines descriptive labels for the sensor data based on the content of the anomalous events.

8. The voice interaction method according to claim 7, characterized in that, The step of determining the descriptive label based on the content of the anomalous event includes: Based on the content of the anomalous event, a multi-level descriptive label is determined from high to low, wherein the multi-level descriptive label is used to characterize the content of the anomalous event; and / or, The description label is determined based on the start time, end time, and content of the abnormal event.

9. The voice interaction method according to claim 6, characterized in that, The generative large model includes a multimodal generative large model; The event problem information and the description label include text modal data, and the sensor data includes data in at least one of the following modalities: image modal, video modal, and audio modal.

10. The voice interaction method according to claim 9, characterized in that, The multimodal generative large model includes a multimodal generative large model obtained by reinforcement training based on the application environment of the sensor.

11. The voice interaction method according to claim 4, characterized in that, The method further includes: In response to receiving event problem information sent by the interactive device, an upload command is sent to the sensor in the scene indicated by the scene information associated with the event problem information, so that the sensor sends the currently collected sensing data to the cloud.

12. The voice interaction method according to claim 11, characterized in that, The method further includes: Receive currently collected sensor data uploaded by the sensor; The sensor data currently collected by the sensor and the event problem information are input into the generative large model so that the generative large model can determine the description label of the sensor data based on the sensor data currently collected by the sensor and the event problem information. The answer information is determined based on the description label, and the answer information is returned to the interactive device.

13. A voice interaction device, characterized in that, The device includes: The instruction receiving module is used to receive user instructions input by the user through a voice assistant, wherein the user instructions include event and problem information associated with scene information; The answer playback module is used to play the answer information corresponding to the event question information through the voice assistant, wherein the answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

14. A voice interaction device, characterized in that, The device includes: The tag determination module is used to receive sensing data collected by the sensor, generate a description tag based on the sensing data, and record the description tag, wherein the description tag is used to characterize the content of the sensing data; The answer return module is used to respond to receiving event question information sent by the interactive device, determine the answer information corresponding to the event question information according to the recorded description tag, and return the answer information to the interactive device so that the interactive device can play the answer information through the voice assistant. The event question information is associated with scene information and is indicated by user commands input by the user based on the voice assistant of the interactive device. The answer information is used to describe the sensor data collected by the sensor in the scene indicated by the scene information.

15. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method described in any one of claims 1 to 12.

16. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 12.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method as described in any one of claims 1 to 12.