Interaction method and device of intelligent device, storage medium and electronic device
By integrating multimodal data acquisition, including voice and images, the problem of low success rate in voice interaction has been solved, achieving more efficient human-computer interaction.
Patent Information
- Application Number
- CN202111662830.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2041-12-30
AI Technical Summary
Existing technologies that use voice to interact with smart devices have a low success rate because they cannot accurately acquire and recognize voice data.
When interaction parameters cannot be obtained, target interaction data and reference data are fused to determine the interaction parameters. Multimodal data acquisition and fusion operations, such as combining voice and image data, are used to assist in determining the interaction parameters.
It improves the success rate of human-computer interaction, solves the problem of interaction failure in situations such as high noise or inaccurate pronunciation, and enhances the accuracy and reliability of interaction.
Smart Images

Figure CN116418611B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of device interaction, and in particular, to a device interaction method and apparatus, a storage medium, and an electronic device. BACKGROUND
[0002] Smart home devices generally have human-computer interaction capabilities (i.e., the device and the user interact), and a voice-based interaction mode is a commonly used human-computer interaction mode at present. The user controls the smart home device to perform a corresponding interaction operation through a voice instruction, for example, to play the weather.
[0003] However, the voice-based interaction mode requires the voice data to be parsed to obtain the user's demand. If there is too much environmental noise or inaccurate pronunciation (e.g., there is an accent), it will cause the voice data to be unable to be accurately obtained and recognized, thereby causing the human-computer interaction to fail.
[0004] Therefore, the voice-based interaction mode with the smart device in the related art has the problem of low success rate of human-computer interaction due to the inability to accurately obtain and recognize the voice data. SUMMARY
[0005] Embodiments of the present application provide a smart device interaction method and apparatus, a storage medium, and an electronic device to at least solve the technical problem of low success rate of human-computer interaction due to the inability to accurately obtain and recognize voice data in the related art voice-based interaction mode with the smart device.
[0006] According to an aspect of an embodiment of the present application, a smart device interaction method is provided, including: obtaining target interaction data issued by a use object, wherein the target interaction data is first modal interaction data, and the target interaction data is used to trigger a first device to perform a first interaction operation; in a case where an interaction parameter corresponding to the first interaction operation is not obtained according to the target interaction data, obtaining target reference data corresponding to the target interaction data, wherein the target reference data is second modal reference data, and the target reference data is used to assist in determining the interaction parameter corresponding to the first interaction operation; performing a fusion operation on the target interaction data and the target reference data to obtain a first interaction parameter corresponding to the first interaction operation; and controlling the first device to perform the first interaction operation according to the first interaction parameter.
[0007] In an example embodiment, the performing a fusion operation on the target interaction data and the target reference data to obtain a first interaction parameter corresponding to the first interaction operation comprises: performing a fusion operation on a first feature vector of the target interaction data and a second feature vector of the target reference data to obtain a target fusion feature vector; and obtaining the first interaction parameter corresponding to the first interaction operation using the target fusion feature vector.
[0008] In an example embodiment, the obtaining the target reference data corresponding to the target interaction data comprises: obtaining the target reference data corresponding to the target interaction data according to a data acquisition time of the target interaction data.
[0009] In an example embodiment, the obtaining the target reference data corresponding to the target interaction data comprises: obtaining the target reference data collected by a second device within a first time period before the data acquisition time; or obtaining the target reference data collected by the second device within a second time period after the data acquisition time; or obtaining the target reference data collected by the second device within a third time period containing the data acquisition time.
[0010] In an example embodiment, the obtaining the target interaction data emitted by the use object comprises: simultaneously starting a plurality of collection components to collect data in a case where it is detected that a target component of the first device is used, wherein each collection component in the plurality of collection components is used to collect data of one modality; and determining that the target interaction data is obtained in a case where interaction information is identified from first collection data collected by a target collection component in the plurality of collection components, wherein the target collection component is a collection component corresponding to the first modality, and the target interaction data is the first collection data.
[0011] In an example embodiment, the obtaining the target reference data corresponding to the target interaction data comprises: obtaining collection data collected by other collection components in the plurality of collection components except for the target collection component to obtain the target reference data, wherein the other collection components are collection components corresponding to the second modality.
[0012] In an example embodiment, after the target interaction data issued by the use object is acquired, the method further comprises: in a case where a second interaction parameter corresponding to the first interaction operation is acquired according to the target interaction data, acquiring a current environment parameter in which the use object is located; in a case where the second interaction parameter does not match the current environment parameter, updating the second interaction parameter using the current environment parameter to obtain an updated second interaction parameter; and controlling the first device to perform the first interaction operation according to the updated second interaction parameter.
[0013] According to another aspect of the embodiments of the present application, an interaction device of an intelligent device is also provided, which comprises: a first acquisition unit configured to acquire target interaction data issued by a use object, wherein the target interaction data is first-modal interaction data, and the target interaction data is used to trigger a first device to perform a first interaction operation; a second acquisition unit configured to, in a case where an interaction parameter corresponding to the first interaction operation is not acquired according to the target interaction data, acquire target reference data corresponding to the target interaction data, wherein the target reference data is second-modal reference data, and the target reference data is used to assist in determining the interaction parameter corresponding to the first interaction operation; a first execution unit configured to perform a fusion operation on the target interaction data and the target reference data to obtain a first interaction parameter corresponding to the first interaction operation; and a second execution unit configured to control the first device to perform the first interaction operation according to the first interaction parameter.
[0014] In an example embodiment, the first execution unit comprises: an execution module configured to perform a fusion operation on a first feature vector of the target interaction data and a second feature vector of the target reference data to obtain a target fusion feature vector; and an identification module configured to acquire the first interaction parameter corresponding to the first interaction operation using the target fusion feature vector.
[0015] In an example embodiment, the second acquisition unit comprises: a first acquisition module configured to acquire the target reference data corresponding to the target interaction data according to a data acquisition time of the target interaction data.
[0016] In an example embodiment, the first acquisition module comprises: a first acquisition submodule configured to acquire the target reference data collected by a second device within a first time period before the data acquisition time; or a second acquisition submodule configured to acquire the target reference data collected by the second device within a second time period after the data acquisition time; or a third acquisition submodule configured to acquire the target reference data collected by the second device within a third time period containing the data acquisition time.
[0017] In an example embodiment, the first obtaining unit comprises: a starting module, configured to start a plurality of collecting components to collect data simultaneously upon detecting that the target component of the first device is in use, wherein each of the plurality of collecting components is configured to collect data of one modality; and a determining module, configured to determine that the target interaction data is obtained upon identifying the interaction information from first collected data collected by a target collecting component, wherein the target collecting component is a collecting component corresponding to the first modality, and the target interaction data is the first collected data.
[0018] In an example embodiment, the second obtaining unit comprises: a second obtaining module, configured to obtain collected data collected by other collecting components of the plurality of collecting components except the target collecting component, to obtain the target reference data, wherein the other collecting components are collecting components corresponding to the second modality.
[0019] In an example embodiment, the apparatus further comprises: a third obtaining unit, configured to, after obtaining the target interaction data emitted by the use object, obtain a current environment parameter of the use object upon obtaining a second interaction parameter corresponding to the first interaction operation according to the target interaction data; an updating unit, configured to update the second interaction parameter using the current environment parameter to obtain an updated second interaction parameter upon the second interaction parameter not matching the current environment parameter; and a third executing unit, configured to control the first device to perform the first interaction operation according to the updated second interaction parameter.
[0020] According to another aspect of the embodiments of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is configured to execute the interaction method of the smart device when running.
[0021] According to another aspect of the embodiments of the present application, an electronic device is provided, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the interaction method of the smart device through the computer program.
[0022] In the embodiment of the present application, when the interaction parameters corresponding to the interaction operation cannot be obtained from the target interaction data, the interaction parameters of the interaction operation are obtained by fusing the interaction data and the reference data, and the corresponding device is controlled to perform the interaction operation. The target interaction data is obtained, wherein the target interaction data is the interaction data of the first mode, and the target interaction data is used to trigger the first device to perform the first interaction operation; in the case that the interaction parameters corresponding to the first interaction operation are not obtained according to the target interaction data, the target reference data corresponding to the target interaction data is obtained, wherein the target reference data is the reference data of the second mode, and the target reference data is used to assist in determining the interaction parameters corresponding to the first interaction operation; the target interaction data and the target reference data are fused to obtain the first interaction parameters corresponding to the first interaction operation; and the first device is controlled to perform the first interaction operation according to the first interaction parameters. If the interaction parameters required for performing the interaction operation cannot be obtained from the interaction data of one mode, the reference data of the same mode or different modes can be fused to obtain the interaction parameters of the interaction operation, and then the interaction operation is performed based on the obtained interaction parameters. For the voice interaction or non-voice interaction scene, the purpose of improving the success rate of obtaining the interaction parameters can be achieved, and the technical effect of improving the success rate of human-computer interaction is achieved, thereby solving the technical problem of low success rate of human-computer interaction caused by the inability to accurately obtain and recognize voice data in the related art. BRIEF DESCRIPTION OF DRAWINGS
[0023] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate an embodiment consistent with the present application and, together with the description, serve to explain the principles of the application.
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0025] Figure 1 is a schematic diagram of a hardware environment of an optional interaction method of a smart device according to an embodiment of the present application;
[0026] Figure 2 is a flowchart of an optional interaction method of a smart device according to an embodiment of the present application;
[0027] Figure 3 is a schematic diagram of an optional interaction method system of a smart device according to an embodiment of the present application;
[0028] Figure 4is a structural block diagram of an optional interaction device of a smart device according to an embodiment of the present application;
[0029] Figure 5 is a structural block diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to make the personnel in the art better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person of ordinary skill in the art without creative labor should belong to the protection scope of the present application.
[0031] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0032] According to an aspect of an embodiment of the present application, an interaction method of a smart device is provided. Optionally, in the present embodiment, the above-mentioned interaction method of a smart device can be applied to a hardware environment composed of a terminal device 102 and a server 104 as shown in the figure. Figure 1 As shown in the figure, the server 104 is connected with the terminal device 102 through a network, which can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal, and a database can be set on the server or independently of the server, which is used to provide data storage services for the server 104. Figure 1 The above-mentioned network can include but is not limited to at least one of the following: wired network, wireless network. The above-mentioned wired network can include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network, and the above-mentioned wireless network can include but is not limited to at least one of the following: WIFI (Wireless Fidelity, Wireless Fidelity), Bluetooth. The terminal device 102 can not be limited to PC, mobile phone, tablet computer, etc.
[0033]
[0034] The interaction method of the smart device in the embodiments of the present application can be executed by the server 104, or can be executed by the terminal device 102, or can be executed by the server 104 and the terminal device 102 jointly. Wherein, the terminal device 102 executing the interaction method of the smart device in the embodiments of the present application can also be executed by the client installed thereon.
[0035] Taking the interaction method of the smart device in the embodiments executed by the server 104 as an example, Figure 2 is a flow diagram of an optional interaction method of a smart device according to the embodiments of the present application, as Figure 2 shown, the flow of the method can include the following steps:
[0036] Step S202, obtaining target interaction data issued by a use object, wherein the target interaction data is first modal interaction data, and the target interaction data is used to trigger the first device to execute a first interaction operation.
[0037] The interaction method of the smart device in the embodiments can be applied in the scene of interaction between the terminal device, the smart home device or other devices and the user. The smart home device mentioned above can be a smart home device located in the user's home, which can be an electronic device installed with a smart chip such as a smart TV or a smart refrigerator. The type of smart device in the embodiments is not limited.
[0038] When the use object (corresponding to the user using the smart device, which can be an object representing the user of the smart device) wants to use the first device, the terminal device can be sent target interaction data. The terminal device can be the first device, or other devices associated with the first device. The terminal device can collect the target interaction data through the collection device thereon and send it to the server. The server can obtain the target interaction data mentioned above. Here, the target interaction data can be first modal interaction data, which can be used to trigger the first device to execute a first interaction operation.
[0039] It should be noted that "modal" refers to a source or form of information, for example, "modal" can refer to human touch, hearing, vision, smell, etc., and for example, "modal" can be a medium of voice, video, text and other information. The first modal can be the way the use object issues the interaction data, which can be the data type, such as voice data, gesture data, etc. The embodiments are not limited in this regard.
[0040] Optionally, various interference factors can exist in the collected target interaction data, if the target interaction data is directly forwarded to the server, the resource consumption of the server can be greatly increased, thereby causing the delay of the process of the interaction between the intelligent device and the use object to be high. In the embodiment, after the target interaction data is collected, the terminal device can perform preliminary processing on the target interaction data to remove the interference data in the target interaction data, and then send the processed target interaction data to the server.
[0041] For example, when the target interaction data is target voice data issued by the use object, and there is abnormal voice data of other people in the target voice data, the target voice data can be subjected to noise reduction operation to eliminate the noise in the target voice data. When the target interaction data is target gesture data issued by the use object, and there is abnormal gesture data of other people in the target gesture data, the abnormal gesture data in the target gesture data can be removed to eliminate the interference in the target gesture data.
[0042] It should be noted that the first device can be a smart home device (for example, a smart refrigerator, etc.) in the same geographical environment as the use object, can be a terminal device (for example, a smart phone, etc.) corresponding to the use object, and can be other smart devices (for example, a smart navigation robot, etc.). The first interaction operation can be a query operation, a purchase operation, and can be other operations, which are not limited in the embodiment,
[0043] For example, after the user issues the voice data of “call me up at 8 o'clock tomorrow morning” to the smart phone, the smart phone can obtain the interaction voice data, and determine that the user needs to set an alarm clock at 8 o'clock tomorrow morning. Here, the operation of setting the alarm clock is the first interaction operation.
[0044] It should be noted that the first interaction operation can include one or more operations. For example, after the user issues the voice instruction of “take me to XXX place” to the smart navigation robot, the smart robot can first perform a query operation to determine the route from the current location to the XXX place, and then perform a navigation operation to guide the user to the XXX place according to the determined route.
[0045] In step S204, the target reference data corresponding to the target interaction data is obtained in the case where the interaction parameter corresponding to the first interaction operation is not obtained from the target interaction data, wherein the target reference data is reference data of a second mode, and the target reference data is used to assist in determining the interaction parameter corresponding to the first interaction operation.
[0046] In this embodiment, after obtaining the target interaction data, the server can obtain the interaction parameter corresponding to the interaction operation to be performed according to the target interaction data. Here, the interaction parameter corresponding to the first interaction operation can be various. For the interaction parameter that must be specified for performing the first interaction operation, the corresponding parameter value needs to be identified from the target interaction data. For the interaction parameter that is not necessarily specified for performing the first interaction operation, if the corresponding parameter value is identified from the target interaction data, the first interaction operation can be performed using the identified parameter value from the target interaction data. If the corresponding parameter value is not identified from the target interaction data, a default parameter value can be used to perform the first interaction parameter.
[0047] If the interaction parameter corresponding to the first interaction operation is not obtained (part of it can be obtained, but it is not enough to perform the submission interaction operation) according to the target interaction data, the server can obtain target reference data corresponding to the target interaction data. The target reference data is reference data of a second modality, and the target reference data is used to assist in determining the interaction parameter corresponding to the first interaction operation.
[0048] It should be noted that the first modality and the second modality can be the same modality, or different modalities. For example, when a user coughs while issuing an interaction voice "I want to eat something" to an AI (Artificial Intelligence, artificial intelligence) assistant on his smart device (for example, a smart phone). The AI assistant recognizes the coughing sound according to the current sound, and recognizes the user's facial expression (an example of target reference data) according to the image, and combines the recognized voice interaction query to comprehensively judge the user's intention: coughing to eat medicine. Alternatively, the second modality can be multiple modalities, and each modality in the multiple modalities can be different. For example, the multiple modalities can include an image modality or a gesture modality.
[0049] Step S206, performing a fusion operation on the target interaction data and the target reference data to obtain the first interaction parameter corresponding to the first interaction operation.
[0050] In this embodiment, after obtaining the target interaction data and the target reference data, the server can control the first device to perform the first interaction operation according to the target interaction data and the target reference data. Alternatively, the server can perform a fusion operation on the target interaction data and the target reference data to obtain the first interaction parameter corresponding to the first interaction operation.
[0051] Optionally, the above fusion operation can be extracting interaction information from the target interaction data, extracting reference information from the target reference data, and fusing the interaction information and the reference information to obtain fused information, and determining the fused information as the first interaction parameter corresponding to the first interaction operation. The interaction information and the reference information can be feature vectors, feature values, or other forms, which are not limited in the embodiment.
[0052] For example, when the user finds that the milk in the smart refrigerator is expired, the user can issue a voice interaction data of "help me reorder two bottles" to the smart refrigerator. The smart refrigerator can obtain the voice interaction data and determine that the interaction operation corresponding to the voice interaction data is a reorder operation, but cannot determine the target of the reorder. At this time, the image data of the user at this time can be synchronously obtained, and when it is identified that the object held by the user is milk, it can be determined that the target of the reorder is milk. The brand, model, and the like of the milk can also be determined.
[0053] In step S208, the first device is controlled to perform the first interaction operation according to the first interaction parameter.
[0054] After obtaining the first interaction parameter, the server can control the first device to perform the first interaction operation matching the first interaction parameter. For example, the server generates a control instruction corresponding to the first interaction parameter according to the first interaction parameter, and sends the control instruction to the first device to control the first device to perform the first interaction operation according to the first interaction parameter. For another example, the server directly sends the first interaction parameter to the first device, and the first device generates a control instruction according to the first interaction parameter after receiving the first interaction parameter, and performs the first interaction operation according to the control instruction, which is not limited in the embodiment. For example, when the server obtains that the target of the reorder operation is a certain kind of milk, the smart refrigerator can be controlled to perform the reorder operation.
[0055] Illustratively, the AI assistant identifies the current environment through the camera and other perception sensors on the hardware. When it is identified that it is an indoor space (such as a shopping mall), the user asks the AI assistant a voice interaction question: "How to get to the subway station?". The AI assistant comprehensively judges that the user's intention is to find an indoor traffic map according to the environment information and the recognized user voice interaction query.
[0056] In another environment, the AI assistant identifies the current environment through the camera and other perception sensors on the hardware. When it is identified that it is an outdoor environment, the user asks the AI assistant a voice interaction question: "How to get to the subway station?". The AI assistant comprehensively judges that the user's intention is to find an outdoor traffic map according to the environment information and the recognized user voice interaction query.
[0057] Through the steps S202 to S208, the target interaction data issued by the use object is acquired, the target interaction data is the interaction data of the first mode, and the target interaction data is used to trigger the first device to perform the first interaction operation; in a case where the interaction parameter corresponding to the first interaction operation is not acquired according to the target interaction data, the target reference data corresponding to the target interaction data is acquired, the target reference data is the reference data of the second mode, and the target reference data is used to assist in determining the interaction parameter corresponding to the first interaction operation; a fusion operation is performed on the target interaction data and the target reference data to obtain the first interaction parameter corresponding to the first interaction operation; and the first device is controlled to perform the first interaction operation according to the first interaction parameter, thereby solving the technical problem in the related art that the success rate of human-computer interaction is low due to the failure to accurately acquire and recognize voice data in the manner of interacting with the intelligent device through voice, and improving the success rate of human-computer interaction.
[0058] In one example embodiment, performing the fusion operation on the target interaction data and the target reference data to obtain the first interaction parameter corresponding to the first interaction operation comprises:
[0059] S11, performing a fusion operation on the first feature vector of the target interaction data and the second feature vector of the target reference data to obtain a target fusion feature vector;
[0060] S12, identifying the first interaction parameter corresponding to the first interaction operation by using the target fusion feature vector.
[0061] In this embodiment, the target server performs a fusion operation on the feature vectors corresponding to the target interaction data and the target reference data, that is, the target server can perform a fusion operation on the first feature vector of the target interaction data and the second feature vector of the target reference data to obtain a target fusion feature vector.
[0062] The process of fusing the first feature vector and the second feature vector can be splicing the first feature vector and the second feature vector at a target position, for example, the second feature vector can be spliced after the first feature vector to obtain a target fusion feature vector, but this way can result in a too long fused feature vector, which is not conducive to subsequent identification of the first interaction parameter corresponding to the first interaction operation by using the target fusion feature vector. Alternatively, the first feature vector and the second feature vector can be fused by using a multi-layer convolutional neural network.
[0063] The multi-layer convolutional neural network comprises a feature extractor composed of a convolutional layer and a subsampling layer. In the convolutional layer of the convolutional neural network, a neuron is connected to only part of the neurons of the adjacent layer. In a convolutional layer of the convolutional neural network, a plurality of feature planes are usually included, each of which is composed of a plurality of neurons arranged in a rectangle, and the neurons of the same feature plane share weights, which are the convolutional kernels. The convolutional kernels can be initialized in the form of a random decimal matrix, and the convolutional kernels will learn reasonable weights in the training process of the network. The direct benefit of shared weights (convolutional kernels) is to reduce the connections between the layers of the network and reduce the risk of overfitting, so that the fused target fusion feature vector not only retains the elements of the first feature vector and the second feature vector, but also does not make the fused feature vector too long.
[0064] Through the embodiment, the first feature vector of the target interaction data and the second feature vector of the target reference data are used for the fusion operation, which can simplify the operation of fusing the target interaction data and the target reference data and accelerate the fusion speed.
[0065] In an example embodiment, the target reference data corresponding to the target interaction data is obtained, comprising:
[0066] S21, obtaining the target reference data corresponding to the target interaction data according to the data acquisition time of the target interaction data.
[0067] Since the interaction operation required by the user to use the object to control the smart device can be different in different scenarios. For example, when the user needs to buy milk, a voice instruction can be issued to control the smart refrigerator to perform the ordering operation of buying 3 bottles of milk, and if the user needs to buy apples, a voice instruction can be issued to control the smart refrigerator to perform the ordering operation of buying 3 kg of apples. Therefore, the reference data obtained is matched with the current interaction scenario.
[0068] In the embodiment, the reference data matched with the current interaction data can be obtained based on the correlation of time (or, time and space). Alternatively, the server can obtain the target reference data corresponding to the target interaction data according to the data acquisition time of the target interaction data,
[0069] For example, when the server obtains the voice instruction "help me order two bottles again" at the first time, the target of ordering is not recognized from "help me order two bottles again", and the image collected by the camera inside the smart refrigerator corresponding to the first time can be obtained, so as to identify the target of ordering as milk.
[0070] Through the embodiment, the target reference data corresponding to the target interaction data is obtained according to the data acquisition time of the target interaction data, which can improve the accuracy of the execution of the device interaction operation.
[0071] In one example embodiment, the target reference data corresponding to the target interaction data is acquired, including:
[0072] S31, acquiring the target reference data collected by the second device in a first time period before the data acquisition time; or,
[0073] S32, acquiring the target reference data collected by the second device in a second time period after the data acquisition time; or,
[0074] S33, acquiring the target reference data collected by the second device in a third time period containing the data acquisition time.
[0075] In the embodiment, the target reference data can be collected by the second device, and the second device can be the same device as the first device or a different device. The second device can be a smart home device, a terminal device, or other devices. The target server can acquire the reference data collected by the second device in a time period. The time period can have various relationships with the data acquisition time.
[0076] As an optional implementation, the target reference data can be the data collected by the second device in a first time period before the data acquisition time. The first time period can be a time period corresponding to the effective reference data found from the data acquisition time.
[0077] As another optional implementation, the target reference data can be the data collected by the second device in a second time period after the data acquisition time. The second time period can be a time period corresponding to the effective reference data found from the data acquisition time.
[0078] As yet another optional implementation, the target reference data can be the data collected by the second device in a third time period containing the data acquisition time. The second time period can be a time period corresponding to the effective reference data found from the data acquisition time.
[0079] According to the embodiment, the effective reference data collected by the second device is found from the data acquisition time, and the credibility of the acquired reference data can be improved.
[0080] In one example embodiment, the target interaction data emitted by the use object is acquired, including:
[0081] S41, in the case of detecting that the target component of the first device is used, simultaneously starting a plurality of acquisition components to perform data acquisition, wherein each acquisition component in the plurality of acquisition components is used to acquire data of one modality;
[0082] S42, in the case that the interaction information is identified in the first acquisition data acquired from the target acquisition component in the plurality of acquisition components, determining that the target interaction data is acquired, wherein the target acquisition component is the acquisition component corresponding to the first modality, and the target interaction data is the first acquisition data.
[0083] In the related art, when the smart interaction function needs to be used, a wake-up word or the like is needed to wake up the function. In the embodiment, the scene of automatic wake-up can be set in advance, for example, when a certain component of the smart device is used, the wake-up interaction function can be directly triggered, without the need for additional operation of the user, so that the smart interaction can be realized. For the first device, the interaction function can be automatically woken up when the target component of the first device is used. The first device can be a smart refrigerator, a smart television or the like, and the target component of the first device can be a component of the smart refrigerator, the smart television or the like, for example, when the first device is a smart refrigerator, the target component can be a refrigerator door of the smart refrigerator, or a touch screen on the smart refrigerator, and the like, which is not limited in the embodiment.
[0084] The first device can be provided with a plurality of acquisition components, and when the interaction function of the first device is woken up, the plurality of acquisition components can be simultaneously started to perform data acquisition, and each acquisition component is used to acquire data of one modality, for example, when the acquisition component is a microphone, data of a voice modality of a user can be acquired, and when the acquisition component is a camera, data of an image modality of the user can be acquired. In the case that the target component of the first device is detected to be used, the target server can simultaneously start the plurality of acquisition components to perform data acquisition, and obtain acquisition data of a plurality of modalities. Optionally, in the case that a predetermined operation is performed on the target component or other components, the plurality of acquisition components are controlled to stop data acquisition, for example, adjusted to a dormant state to end the data acquisition.
[0085] For example, when the user opens the refrigerator door of the smart refrigerator, the microphone in the smart refrigerator is started to acquire the voice of the user, and the camera in the smart refrigerator is simultaneously started to acquire the image of the user. When the user closes the refrigerator door of the smart refrigerator, the microphone and the camera in the smart refrigerator are closed to stop the data acquisition.
[0086] The automatic starting interaction function can be started after obtaining authorization of the user. For example, the user can pre-configure whether to start the automatic starting interaction function. Alternatively, the automatic interaction function can be started, and the user is prompted by voice or other prompt information whether to authorize data collection. After obtaining the authorization of the user, the data collection is started.
[0087] There can be multiple target components in the first device. Alternatively, multiple collection components can be started simultaneously to collect data when it is detected that any one of the multiple target components is used. Alternatively, the multiple target components can be specified. Only when it is detected that the specified target component is used, the multiple collection components are started simultaneously to collect data.
[0088] After starting the multiple collection components to collect data, the target server can obtain the data collected by each collection component, and identify the data collected by each collection component. If the interaction information is identified from the first collection data collected by the target collection component, it is determined that the target interaction data is obtained, the target collection component is the collection component corresponding to the first modality, and the target interaction data is the first collection data.
[0089] The process of identifying the interaction information can be identifying the interaction keywords or other interaction key information in the collection data. The keywords can be words used by the user to express the purpose or question in the interaction process. For example, when the word “order” in the voice data “This item, help me order two bottles” is identified, the word is an interaction keyword, and the target server can determine that the target interaction data is identified.
[0090] It should be noted that since multiple collection components are started to collect data, a large amount of storage space is required to store the collected data. In order to save the storage space occupied by the collection data, the collected data can be stored in segments during the collection of the data, it is analyzed whether the data contains interaction data, the original collected data is deleted, and the newly collected data is stored, thereby reducing the occupation of the storage space. For example, the storage time interval can be set to 30s, and the original data stored is deleted every 30s.
[0091] Through the embodiment, when it is detected that the target component of the first device is used, the multiple collection components are automatically started to collect data, which can ensure the smoothness of the interaction and improve the user experience.
[0092] In one example embodiment, the target reference data corresponding to the target interaction data is obtained, including:
[0093] S51, acquire the acquisition data acquired by the other acquisition components in the plurality of acquisition components except for the target acquisition component, to obtain target reference data, wherein the other acquisition components are acquisition components corresponding to the second modality.
[0094] When the first acquisition data is determined to be the interaction data, the target server can store the acquisition data acquired by the other acquisition components in the plurality of acquisition components except for the target acquisition component, and acquire the stored data when the data is needed as reference data, or the acquisition data acquired by the other acquisition components can be directly used as reference data.
[0095] Optionally, the target server can acquire the acquisition data acquired by the other acquisition components in the plurality of acquisition components except for the target acquisition component, and perform data fusion on the acquired acquisition data and the target interaction data to determine the interaction operation to be performed, taking the acquired acquisition data as target reference data.
[0096] Through the embodiment, the acquisition data acquired by part of the plurality of acquisition components started by the fusion of the voice information stream and the image information stream of the user interaction in a natural and non-inductive manner is taken as reference data to determine the interaction operation to be performed, which utilizes the correlation between the acquisition data in the time dimension and the space dimension, and improves the accuracy of the interaction operation recognition.
[0097] In one example embodiment, after the target interaction data issued by the user is acquired, the above method further includes:
[0098] S61, acquire the current environment parameter in which the user is located, in a case where the second interaction parameter corresponding to the first interaction operation is acquired according to the target interaction data;
[0099] S62, update the second interaction parameter using the current environment parameter to obtain an updated second interaction parameter, in a case where the second interaction parameter does not match the current environment parameter;
[0100] S63, control the first device to perform the first interaction operation according to the updated second interaction parameter.
[0101] In the embodiment, if the interaction parameter corresponding to the first interaction operation, i.e., the second interaction parameter, can be acquired according to the target interaction data, the first device can be directly controlled to perform the first interaction operation according to the second interaction parameter. Considering that the interaction operation expected by the user is different in different scenarios, the interaction parameter with the default parameter value included in the second interaction parameter may not meet the user's expectation. Therefore, the way of performing the interaction operation according to the default parameter value for the unrecognized interaction parameter has the problem that the matching degree between the execution result of the interaction operation and the execution expectation of the interaction operation is low.
[0102] For example, when the user asks the AI assistant in a voice interaction: "How to go to XX place?", since the specified transportation is not recognized, that is, the parameter value of the parameter of the specified transportation is not specified, any transportation can be used by default to reach the XX place, and the terminal device can show the user all the transportation that can reach the XX place.
[0103] However, if the user is in an outdoor environment (for example, in a subway station), the more desirable transportation is the subway, and if the user is in an outdoor environment, the more desirable transportation is the bus.
[0104] In this embodiment, after obtaining the second interaction parameter, the server can obtain the current environment parameter in which the use object is located. The current environment parameter can be obtained in various ways, for example, by image acquisition through an image acquisition device on the terminal device, and identifying the acquired environment image to determine the current environment parameter, and for example, the current environment parameter can be determined by sound acquisition through a microphone array on the terminal device, and identifying the acquired environment sound.
[0105] If the second interaction parameter does not match the current environment parameter, the second interaction parameter can be updated using the current environment parameter to obtain an updated second interaction parameter. The update method can be to update the parameter value of the interaction parameter corresponding to the current environment parameter to a parameter value that matches the current environment parameter. After updating the second interaction parameter, the terminal device can control the first device to perform the first interaction operation according to the updated second interaction parameter.
[0106] Through this embodiment, the use of environment parameters to update the interaction parameters of interaction operations can improve the matching degree of the way the interaction operation is performed and the expected performance, thereby improving the user's experience.
[0107] The interaction method of the intelligent device in the embodiments of the present application will be explained and described below in combination with optional examples. In this optional example, the intelligent device is a smart refrigerator, a smart navigation robot, and the like.
[0108] With the general human-computer interaction capability of home appliances in the smart home ecosystem, interaction based on modal data such as voice and image is becoming the mainstream human-computer interaction method. However, both voice interaction and visual interaction have single modal scene problems (for example, environmental noise problem of voice interaction, image occlusion problem of face recognition).
[0109] To solve the single modality scene problem, in this optional example, a multi-modal fusion identity recognition and semantic understanding scheme is provided, which fuses the speech information stream and the image information stream, trains the multi-modal data in the home scene, and enables the home appliance device to perceive through multi-modalities like a person, and then learn and understand the user's intention and interact with the user.
[0110] In this optional example, the visual interaction, voice interaction, gesture interaction, and other interaction methods can be fused for human-computer interaction, which enriches the interaction method of the home appliance device and the user and improves the humanization of the home appliance device. In the natural interaction between the user and the home appliance device, the user does not need to deliberately enter information, the speech information stream and the image information stream of the user are fused in multi-modalities, the multi-modal perception information is comprehensively used for user identity recognition and natural language understanding, and more information is obtained through multi-round interaction with the user, which can improve the accuracy of natural language understanding. Here, the multi-modal perception information can include environmental information, user information (expression, gesture, sound, etc.).
[0111] The interaction method of the intelligent device in this optional example can be applied to the human-computer interaction system as shown in Figure 3 The system can be divided into three modules, which are:
[0112] 1) Multi-modal fusion perception module
[0113] The multi-modal fusion perception module can realize single modality perception to multi-modality perception, can still work normally and obtain effective information even if a certain modality fails or is missing, greatly improves the robustness of the model, and improves the comprehensiveness and accuracy of the perceived user and environmental information. For example, in the daily interaction between the user and the home appliance device, the identity space and the speaking information of the speaker are processed in the form of a voiceprint model clustering, to confirm the identity of the speaker. At the same time, the image information of the speaker is associated with the voiceprint model, the voice and image perception information are fused and input, multi-modal perception information learning is performed, and a specific user can be identified through voice or image information, to meet the scene needs of personalized device control, personalized information broadcast, personalized service recommendation, etc.
[0114] 2) Multi-modal fusion understanding module
[0115] The multi-modal fusion understanding module can complement the visual information and the voice information, the user does not need to say all the information required for semantic understanding, can use abbreviations and references, and can achieve more natural interaction of the user and more accurate understanding of the user's intention by the machine.
[0116] 3) Multi-modal fusion interaction module
[0117] The multi-modal fusion interaction module can support fusion of visual interaction, voice interaction, gesture interaction and the like, enriches the interaction mode between the household appliance and the human, and improves the humanization degree of the household appliance.
[0118] Through the optional example, the voice information stream and the image information stream of the user interaction can be naturally and unobtrusively fused, the household appliance can conveniently understand the intention of the human and interact with the human, and therefore the intelligent service capability of the household appliance is improved, and the human-computer interaction efficiency and the satisfaction of the user are improved.
[0119] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0120] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, and of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM (Read-Only Memory), a RAM (Random Access Memory), a magnetic disk, or an optical disk), and includes a number of instructions to make a terminal device (which can be a mobile phone, a computer, a server, or a network device) execute the methods described in the embodiments of the present application.
[0121] According to another aspect of the embodiments of the present application, an interaction device of a smart device for implementing the interaction method of the smart device is also provided. Figure 4 is a structural block diagram of an optional interaction device of a smart device according to the embodiments of the present application, as shown in the figure, the device can include: Figure 4
[0122] The first acquisition unit 402 is configured to acquire target interaction data emitted by a use object, wherein the target interaction data is interaction data of a first modality, and the target interaction data is used to trigger the first device to execute a first interaction operation;
[0123] The second acquisition unit 404 is connected with the first acquisition unit 402, configured to acquire target reference data corresponding to the target interaction data in a case where the interaction parameter corresponding to the first interaction operation is not acquired according to the target interaction data, wherein the target reference data is reference data of a second modality, and the target reference data is used for assisting in determining the interaction parameter corresponding to the first interaction operation;
[0124] The first execution unit 406 is connected with the second acquisition unit 404, configured to perform a fusion operation on the target interaction data and the target reference data to obtain the first interaction parameter corresponding to the first interaction operation.
[0125] The second execution unit 408 is connected with the first execution unit 406, configured to control the first device to perform the first interaction operation according to the first interaction parameter.
[0126] It should be noted that the first acquisition unit 402 in this embodiment can be configured to perform the step S202, the second acquisition unit 404 in this embodiment can be configured to perform the step S204, the first execution unit 406 in this embodiment can be configured to perform the step S206, and the second execution unit 408 in this embodiment can be configured to perform the step S208.
[0127] Through the above modules, the target interaction data issued by the use object is acquired, wherein the target interaction data is interaction data of a first modality, and the target interaction data is used for triggering the first device to perform the first interaction operation; in a case where the interaction parameter corresponding to the first interaction operation is not acquired according to the target interaction data, the target reference data corresponding to the target interaction data is acquired, wherein the target reference data is reference data of a second modality, and the target reference data is used for assisting in determining the interaction parameter corresponding to the first interaction operation; a fusion operation is performed on the target interaction data and the target reference data to obtain the first interaction parameter corresponding to the first interaction operation; and the first device is controlled to perform the first interaction operation according to the first interaction parameter, thereby solving the technical problem in the related art that the success rate of human-computer interaction is low due to the failure to accurately acquire and recognize voice data in the mode of interacting with the intelligent device through voice, and improving the success rate of human-computer interaction.
[0128] In one example embodiment, the first execution unit includes:
[0129] The execution module is configured to perform a fusion operation on the first feature vector of the target interaction data and the second feature vector of the target reference data to obtain a target fusion feature vector.
[0130] The recognition module is configured to acquire the first interaction parameter corresponding to the first interaction operation by using the target fusion feature vector.
[0131] In an example embodiment, the second obtaining unit comprises:
[0132] The first obtaining module is configured to obtain the target reference data corresponding to the target interaction data according to the data obtaining time of the target interaction data.
[0133] In an example embodiment, the first obtaining module comprises:
[0134] The first obtaining sub-module is configured to obtain the target reference data collected by the second device within a first time period before the data obtaining time; or
[0135] The second obtaining sub-module is configured to obtain the target reference data collected by the second device within a second time period after the data obtaining time; or
[0136] The third obtaining sub-module is configured to obtain the target reference data collected by the second device within a third time period containing the data obtaining time.
[0137] In an example embodiment, the first obtaining unit comprises:
[0138] The starting module is configured to start the plurality of collecting components to collect data simultaneously in a case where it is detected that the target component of the first device is used, wherein each collecting component in the plurality of collecting components is configured to collect data of one modality;
[0139] The determining module is configured to determine that the target interaction data is obtained in a case where the interaction information is identified from the first collecting data collected by the target collecting component in the plurality of collecting components, wherein the target collecting component is a collecting component corresponding to the first modality, and the target interaction data is the first collecting data.
[0140] In an example embodiment, the second obtaining unit comprises:
[0141] The second obtaining module is configured to obtain collecting data collected by other collecting components in the plurality of collecting components except the target collecting component, to obtain the target reference data, wherein the other collecting components are collecting components corresponding to the second modality.
[0142] In an example embodiment, the device further comprises:
[0143] The third obtaining unit is configured to obtain the current environmental parameter in which the user is located in a case where the second interaction parameter corresponding to the first interaction operation is obtained according to the target interaction data after the target interaction data emitted by the user is obtained.
[0144] an updating unit, configured to update the second interaction parameter using the current environment parameter to obtain an updated second interaction parameter in a case where the second interaction parameter does not match the current environment parameter;
[0145] a third execution unit, configured to control the first device to perform the first interaction operation according to the updated second interaction parameter.
[0146] It should be noted that the above modules and the examples and application scenarios realized by the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules as part of the device can run in the hardware environment as shown in Figure 1 The hardware environment includes a network environment.
[0147] According to another aspect of the embodiments of the present application, a storage medium is also provided. Optionally, in the present embodiment, the storage medium can be used to execute the program code of the interaction method of any one of the intelligent devices in the embodiments of the present application.
[0148] Optionally, in the present embodiment, the storage medium can be located on at least one of the plurality of network devices in the network shown in the above embodiments.
[0149] Optionally, in the present embodiment, the storage medium is configured to store program code for executing the following steps:
[0150] S1, obtaining target interaction data issued by a use object, wherein the target interaction data is first modal interaction data, and the target interaction data is used to trigger the first device to perform a first interaction operation;
[0151] S2, in a case where the interaction parameter corresponding to the first interaction operation is not obtained according to the target interaction data, obtaining target reference data corresponding to the target interaction data, wherein the target reference data is second modal reference data, and the target reference data is used to assist in determining the interaction parameter corresponding to the first interaction operation;
[0152] S3, performing a fusion operation on the target interaction data and the target reference data to obtain the first interaction parameter corresponding to the first interaction operation;
[0153] S4, controlling the first device to perform the first interaction operation according to the first interaction parameter.
[0154] Optionally, the specific examples in the present embodiment can refer to the examples described in the above embodiments, which will not be described herein.
[0155] Optionally, in the embodiment, the storage medium can include, but is not limited to, a U disk, a ROM, a RAM, a mobile hard disk, a magnetic disk or an optical disk, and various storage program codes.
[0156] According to still another aspect of the embodiments of the present application, an electronic device for implementing the interaction method of the intelligent device is also provided, which can be a server, a terminal, or a combination thereof.
[0157] Figure 5 is a structural block diagram of an optional electronic device according to the embodiments of the present application, as shown in Figure 5 The electronic device includes a processor 502, a communication interface 504, a memory 506 and a communication bus 508, wherein the processor 502, the communication interface 504 and the memory 506 complete mutual communication through the communication bus 508, and
[0158] The memory 506 is configured to store a computer program.
[0159] The processor 502 is configured to execute the computer program stored in the memory 506, and implement the following steps.
[0160] S1, obtaining target interaction data sent by a user, wherein the target interaction data is interaction data of a first mode, and the target interaction data is used to trigger a first device to execute a first interaction operation;
[0161] S2, in a case where interaction parameters corresponding to the first interaction operation are not obtained according to the target interaction data, obtaining target reference data corresponding to the target interaction data, wherein the target reference data is reference data of a second mode, and the target reference data is used to assist in determining the interaction parameters corresponding to the first interaction operation;
[0162] S3, performing a fusion operation on the target interaction data and the target reference data to obtain the first interaction parameters corresponding to the first interaction operation;
[0163] S4, controlling the first device to execute the first interaction operation according to the first interaction parameters.
[0164] Optionally, the communication bus can be a PCI (Peripheral Component Interconnect, Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture, Extended Industry Standard Architecture) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 5 only one thick line is used, but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the electronic device and other devices.
[0165] The memory can include a RAM and can also include a non-volatile memory, for example, at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.
[0166] As an example, the aforementioned memory 506 can include, but is not limited to, the first acquisition unit 402, the second acquisition unit 404, the first execution unit 406 and the second execution unit 408 in the interaction device of the aforementioned smart device. In addition, other module units in the interaction device of the aforementioned smart device can also be included, but are not limited to, which will not be described herein.
[0167] The aforementioned processor can be a general-purpose processor, which can include, but is not limited to, a CPU (Central Processing Unit), a NP (Network Processor), etc. It can also be a DSP (Digital Signal Processing), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0168] Optionally, the specific examples in the present embodiment can refer to the examples described in the above embodiments, which will not be described herein.
[0169] Those skilled in the art can understand that, Figure 5 The structure shown is only schematic, and the device implementing the interaction method of the aforementioned smart device can be a terminal device, which can be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, etc. Figure 5 It does not limit the structure of the aforementioned electronic device. For example, the electronic device can further include more or less components (such as a network interface, a display device, etc.) than Figure 5 shown, or have a different configuration from Figure 5 shown.
[0170] Those skilled in the art can understand that all or part of the steps of various methods in the above embodiments can be completed by a program instructing the terminal device related hardware, and the program can be stored in a computer readable storage medium, which can include a flash disk, a ROM, a RAM, a magnetic disk or an optical disk, etc.
[0171] The serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0172] The integrated units in the above embodiments, if realized in the form of software function units and sold or used as independent products, can be stored in the above computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or the whole or part of the technical solutions can be embodied in the form of software products, which are stored in the storage medium and include a plurality of instructions for causing one or more computer devices (which can be personal computers, servers or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0173] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0174] In the several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Of course, the above device embodiment is only illustrative, and the division of the units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, unit or module, and can be electrical or other forms.
[0175] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the scheme provided in the embodiments.
[0176] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or at least two units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or software function unit.
[0177] The above merely describes the preferred embodiments of the present application, and it should be pointed out that, for those skilled in the art, some improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as the protection scope of the present application.
Claims
1. An interaction method of a smart device, characterized in that, The method comprises: acquiring target interaction data emitted by a use object, wherein the target interaction data is first modal interaction data, and the target interaction data is used to trigger a first device to perform a first interaction operation; in a case where an interaction parameter corresponding to the first interaction operation is not acquired according to the target interaction data, acquiring target reference data corresponding to the target interaction data, wherein the target reference data is second modal reference data, and the target reference data is used to assist in determining the interaction parameter corresponding to the first interaction operation; performing a fusion operation on the target interaction data and the target reference data to obtain a first interaction parameter corresponding to the first interaction operation; wherein the fusion operation is extracting interaction information from the target interaction data, extracting reference information from the target reference data, and then fusing the interaction information and the reference information to obtain fusion information, and determining the fusion information as the first interaction parameter corresponding to the first interaction operation; controlling the first device to perform the first interaction operation according to the first interaction parameter; after acquiring the target interaction data emitted by the use object, the method further comprises: in a case where a second interaction parameter corresponding to the first interaction operation is acquired according to the target interaction data, acquiring a current environment parameter in which the use object is located; in a case where the second interaction parameter does not match the current environment parameter, updating the second interaction parameter using the current environment parameter to obtain an updated second interaction parameter, wherein the manner of updating the second interaction parameter is updating a parameter value in the second interaction parameter corresponding to the current environment parameter to a parameter value matching the current environment parameter; controlling the first device to perform the first interaction operation according to the updated second interaction parameter.
2. The method of claim 1, wherein, The method of performing a fusion operation on the target interaction data and the target reference data to obtain a first interaction parameter corresponding to the first interaction operation comprises: performing a fusion operation on a first feature vector of the target interaction data and a second feature vector of the target reference data to obtain a target fusion feature vector; acquiring the first interaction parameter corresponding to the first interaction operation using the target fusion feature vector.
3. The method of claim 1, wherein, The method of acquiring target reference data corresponding to the target interaction data comprises: acquiring the target reference data corresponding to the target interaction data according to a data acquisition time of the target interaction data.
4. The method of claim 3, wherein, The method of acquiring target reference data corresponding to the target interaction data comprises: acquiring the target reference data collected by a second device within a first time period before the data acquisition time; or acquiring the target reference data collected by a second device within a second time period after the data acquisition time; or acquiring the target reference data collected by a second device within a third time period containing the data acquisition time.
5. The method of claim 1, wherein, The method of acquiring target interaction data emitted by a use object comprises: In a case where it is detected that the target component of the first device is used, a plurality of acquisition components are simultaneously started to collect data, wherein each of the plurality of acquisition components is configured to collect data of one modality; In a case where interaction information is identified in first acquisition data collected by a target acquisition component from the plurality of acquisition components, it is determined that the target interaction data is acquired, wherein the target acquisition component is an acquisition component corresponding to the first modality, and the target interaction data is the first acquisition data.
6. The method of claim 5, wherein, The target reference data corresponding to the target interaction data is acquired, including: Acquisition data collected by other acquisition components from the plurality of acquisition components except the target acquisition component is obtained to obtain the target reference data, wherein the other acquisition components are acquisition components corresponding to the second modality.
7. An interaction device for a smart device, the interaction device comprising: Including: The first acquisition unit is configured to acquire target interaction data issued by a use object, wherein the target interaction data is interaction data of a first modality, and the target interaction data is used to trigger a first device to perform a first interaction operation; The second acquisition unit is configured to acquire target reference data corresponding to the target interaction data in a case where an interaction parameter corresponding to the first interaction operation is not acquired according to the target interaction data, wherein the target reference data is reference data of a second modality, and the target reference data is used to assist in determining the interaction parameter corresponding to the first interaction operation; The first execution unit is configured to perform a fusion operation on the target interaction data and the target reference data to obtain a first interaction parameter corresponding to the first interaction operation; wherein the fusion operation is to extract interaction information from the target interaction data, extract reference information from the target reference data, fuse the interaction information and the reference information to obtain fusion information, and determine the fusion information as the first interaction parameter corresponding to the first interaction operation; The second execution unit is configured to control the first device to perform the first interaction operation according to the first interaction parameter. The first acquisition unit is further configured to acquire a current environment parameter in which the use object is located in a case where a second interaction parameter corresponding to the first interaction operation is acquired according to the target interaction data, and update the second interaction parameter using the current environment parameter in a case where the second interaction parameter does not match the current environment parameter to obtain an updated second interaction parameter, wherein the second interaction parameter is updated by updating a parameter value in the second interaction parameter corresponding to the current environment parameter to a parameter value matching the current environment parameter; and control the first device to perform the first interaction operation according to the updated second interaction parameter.
8. A computer readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program executes the method of any one of claims 1 to 6 when executed. 9.An electronic device comprising a memory and a processor, the electronic device characterized by, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 6 by using the computer program.
Citation Information
Patent Citations
Multimodal input-based interactive method and device
CN106997236A
Method and device for processing information
CN112309387A