Intelligent equipment control method and device

By acquiring multimodal information from smart devices, combining audio and video features, and using state models to determine user intent, the problem of unnatural human-computer interaction and false wake-up in existing technologies has been solved, achieving a more efficient and accurate human-computer interaction experience.

CN121143884APending Publication Date: 2025-12-16BEIJING ORION STAR TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511139233.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

The human-computer interaction methods of existing smart devices rely on specific wake words, resulting in unnatural interaction processes, frequent false wake-ups/missed wake-ups, and interference from multiple devices, which affects the user experience.

Method used

By acquiring multimodal information, including audio and video features, and combining it with a state model to determine user intent, the system can decide whether to wake up the smart device and execute commands, thereby reducing erroneous operations and improving the accuracy and robustness of system responses.

Benefits of technology

It enables more natural and accurate human-computer interaction, reduces misoperation, and improves user experience and the level of intelligent interaction of smart devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121143884A_ABST
    Figure CN121143884A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent device control method and device, and the method comprises the steps: obtaining first multi-mode information corresponding to a first instruction in a first environment, and the first multi-mode information comprises the audio and video information of a target user issuing the first instruction; determining state information when the target user issues the first instruction based on the first multi-mode information, wherein the state information comprises direction information when the target user issues the first instruction; when it is determined that the state information meets a wakeup condition, wakeup of the intelligent device is triggered, the first instruction is executed, and the wakeup condition comprises that it is determined that the target user issues the first instruction to the intelligent device. The source direction of the first instruction is determined in combination with the multi-modal information, so that whether the intelligent equipment is awakened or not is determined, and in the process of executing the first instruction, the intelligent equipment does not need to be awakened by an awakening word, man-machine interaction is more natural, misoperation is effectively reduced, and the accuracy and intelligence are higher.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the application relates to the technical field of intelligent devices, in particular to a control method and device for controlling an intelligent device. BACKGROUND

[0002] With the rapid development of intelligent devices, the intelligent devices are used in more and more occasions. For example, the intelligent devices are often triggered to perform corresponding functions through human-computer interaction. The current human-computer interaction process usually depends on a specific wake-up word to activate the device, and there are problems such as unnatural interaction process, frequent false wake-up / misplaced wake-up, and multi-device interference, which seriously affect the user experience.

[0003] Therefore, there is an urgent need for an efficient and accurate human-computer interaction method with high user experience. SUMMARY

[0004] The embodiment of the application provides a control method and device for controlling an intelligent device, so as to provide an efficient and accurate human-computer interaction method with high user experience.

[0005] In a first aspect, the application provides a control method for an intelligent device, comprising:

[0006] obtaining first multi-modal information corresponding to a first instruction in a first environment, the first multi-modal information comprising audio and video information of a target user issuing the first instruction; determining state information of the target user when issuing the first instruction based on the first multi-modal information, the state information comprising direction information of the target user when issuing the first instruction; when the state information meets a wake-up condition, triggering the wake-up of an intelligent device and executing the first instruction, the wake-up condition comprising determining that the target user issues the first instruction towards the intelligent device.

[0007] In the above technical solution, the direction of the user sending the first instruction is determined by combining the multi-modal information such as audio and video of the user, so that the process of determining whether to wake up the intelligent device and execute the corresponding first instruction based on the direction of the first instruction does not need a wake-up word to wake up the intelligent device, making the human-computer interaction more natural. The fusion of audio and video features enables the system to more accurately determine whether the user has an interaction intention, effectively reduces false operations, improves the accuracy and robustness of system response, can more easily track one person in a group for long-context interaction, is more in line with human interaction, improves the human-computer interaction experience, and improves the intelligence.

[0008] As an implementation manner, the determining of the state information of the target user when issuing the first instruction based on the first multi-modal information comprises:

[0009] The audio information in the first multi-modal information is processed based on a preset audio front-end algorithm, and a sound feature of the target user is determined, the sound feature including a sound angle feature of the target user; the video information in the first multi-modal information is picture extracted based on a preset picture extraction algorithm, and a first number of target pictures are obtained; a visual feature of the target user is determined based on the first number of target pictures, the visual feature including one or more of a face angle of the target user, a coordinate of the target user, a distance between the target user and the smart device, and lip movement information of the target user; and state information of the target user when issuing the first instruction is determined based on the sound feature and the visual feature.

[0010] As an example, the sound feature in the embodiments of the present application can also include a timbre feature of the target user, so that sound localization can be performed more targetedly based on the timbre feature of the target user. For example, the embodiments of the present application can pre-store a timbre feature of user A, and when the smart device identifies a timbre feature consistent with user A, further determination of the corresponding sound angle feature is triggered.

[0011] As an implementation manner, the state information of the target user when issuing the first instruction is determined based on the first multi-modal information, including:

[0012] State information of the first environment is obtained, and an environment type of the first environment is determined based on the state information of the first environment, the environment type including one or more of an environment scene, an environment atmosphere, and an environment noise level; a first processing manner is determined based on a mapping relationship between the environment type and the processing manner of the first multi-modal information; and the first processing manner is applied to process the first multi-modal information to determine the state information of the target user when issuing the first instruction.

[0013] As an example, the embodiments of the present application set a relatively complex first processing manner for a complex environment, for example, the collected multi-modal information can be processed in segments; and a relatively simple and convenient first processing manner can be set for a simple environment, for example, the collected multi-modal information can be processed as a whole.

[0014] As an implementation manner, the state information of the target user when issuing the first instruction is determined based on the first multi-modal information by applying the first processing manner, including:

[0015] The audio information and the video information in the obtained first multi-modal information are labeled every first time interval; the audio information and the video information under different labels are separated by user to determine at least one target user; and the state information of each target user when issuing the first instruction is determined.

[0016] As an implementation manner, the separating the at least one target user from the audio information and the video information under different tags comprises:

[0017] Based on the mapping relationship between the user and the timbre, and / or the mapping relationship between the user and the appearance, the users involved in different tags are separated to obtain at least one target user.

[0018] As an example, the embodiments of the present application can process multiple target users in parallel when multiple users issue the first instruction to the smart device, and determine the state information of each target user when issuing the first instruction.

[0019] As an example, the embodiments of the present application can fuse the first instructions issued by the multiple target users, for example, the content of the multi-round interaction context of different users can be processed, so that the smart device control is more accurate and efficient.

[0020] As an example, the embodiments of the present application can accurately separate the voice part of each person.

[0021] As an implementation manner, the determining the state information of each target user when issuing the first instruction comprises:

[0022] Aligning the features of the video information and the audio information in the same time period based on the tags of the audio information and the tags of the video information; obtaining second multi-modal information in which the audio-video feature combination in at least one time period is combined together; determining the state information of each target user when issuing the first instruction based on at least one second multi-modal information.

[0023] As an implementation manner, the determining the state information of each target user when issuing the first instruction based on at least one second multi-modal information comprises:

[0024] Inputting the at least one second multi-modal information into a state model for state judgment, the state model being obtained by training based on multi-modal data and corresponding state information data; determining the state information of each target user when issuing the first instruction based on the output result of the state model.

[0025] In a second aspect, the embodiments of the present application provide a device comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method in the first aspect and any design thereof.

[0026] In a third aspect, the embodiments of the present application provide a state model, which is used to execute the method in the first aspect and any design thereof. As an implementation manner, the state model is obtained by training based on multi-modal data and corresponding state information data.

[0027] As an example, the multi-modal data used for training the state model can be obtained in the following manner, i.e., the audio information such as sound source can be extracted every threshold time length A, and M pictures are extracted from the video information at the same time to extract visual features, then the video features and sound features are used as multi-modal data under the corresponding time length, and the multi-modal data under the corresponding time length is associated with the actual state information, and the whole is used as a training data. Repeat the corresponding operation to obtain multiple training data, obtain a training data set, and train the state model based on the training data set.

[0028] As an example, when the state model is actually applied to intelligent device control, related data generated in the actual control process can also be recorded, such as using the correct state information determined based on the multi-modal data of the user actually performing intelligent device control as positive training data set to optimize the state model, or using the error state information determined based on the multi-modal data of the user actually performing intelligent device control as negative training data set to optimize the state model, etc., which is not limited herein.

[0029] As an example, the input information of the state model includes but is not limited to sound angle, noise angle, person coordinate direction, person coordinate in picture, man-machine distance, lip movement, and other feature information.

[0030] As an example, the output state information of the state model includes but is not limited to the following three state classifications:

[0031] State classification 1: environmental noise state.

[0032] State classification 2: the target user speaks to the intelligent device when issuing the first instruction.

[0033] State classification 3: the target user speaks to someone else when issuing the first instruction.

[0034] As an example, the state model can be an end-side model. For example, the state model can be a 2-layer fully connected fc classification model, and the parameter amount is only about 7k.

[0035] As an example, in order to increase accuracy, when training the state model, a loss function can be set, and the loss function during training is different in weight for the plurality of state classifications.

[0036] In a fourth aspect, an apparatus for controlling is provided, including:

[0037] The acquisition module is configured to acquire first multi-modal information corresponding to the first instruction in the first environment, the first multi-modal information including audio and video information of a target user issuing the first instruction;

[0038] The processing module is configured to determine state information of the target user when issuing the first instruction based on the first multi-modal information, the state information including directional information of the target user when issuing the first instruction; and trigger the wake-up of the smart device and execute the first instruction when the state information meets a wake-up condition, the wake-up condition including determining that the target user is facing the smart device when issuing the first instruction.

[0039] As an implementation manner, the processing module is specifically configured to:

[0040] process audio information in the first multi-modal information based on a preset audio front-end algorithm to determine a voice feature of the target user, the voice feature including a voice angle feature of the target user; perform picture extraction on video information in the first multi-modal information based on a preset picture extraction algorithm to obtain a first number of target pictures; determine a visual feature of the target user based on the first number of target pictures, the visual feature including one or more of a face angle of the target user, coordinates of the target user, a distance between the target user and the smart device, and lip movement information of the target user; and determine the state information of the target user when issuing the first instruction based on the voice feature and the visual feature.

[0041] As an example, the voice feature can further include a voice timbre feature of the target user.

[0042] As an implementation manner, the processing module is specifically configured to:

[0043] determine state information of the first environment based on the state information of the first environment; determine an environment type of the first environment based on the state information of the first environment, the environment type including one or more of an environment scene, an environment atmosphere, and an environment noise level; determine a first processing manner based on a mapping relationship between the environment type and a processing manner of the first multi-modal information; and process the first multi-modal information by applying the first processing manner to determine the state information of the target user when issuing the first instruction.

[0044] As an implementation form, the processing module is specifically configured to:

[0045] The acquired audio information and video information in the first multi-modal information are labeled every first time interval; user separation is performed on the audio information and video information under different labels to determine at least one target user; and state information when the first instruction is issued by each target user is determined.

[0046] As an implementation form, the processing module is specifically configured to:

[0047] Based on a mapping relationship between the user and the timbre, and / or a mapping relationship between the user and the appearance, the users involved in different labels are separated to obtain at least one target user.

[0048] As an example, the embodiments of the present application can be processed in parallel for multiple target users, i.e., there are multiple users issuing the first instruction to the smart device, and the state information when the first instruction is issued by each target user can be determined respectively.

[0049] As an example, the embodiments of the present application can perform fusion processing on the first instructions respectively issued by the multiple target users, for example, the content of the multi-round interaction context of different users can be processed, so that the smart device control is more accurate and efficient.

[0050] As an example, the embodiments of the present application can accurately separate the voice part of each person.

[0051] As an implementation form, the processing module is specifically configured to:

[0052] Based on the label of the audio information and the label of the video information, the video information and the audio information in the same time period are aligned in feature; second multi-modal information in which audio-video feature combinations in at least one time period are obtained; and based on the at least one second multi-modal information, the state information when the first instruction is issued by each target user is determined.

[0053] As an implementation form, the processing module is specifically configured to:

[0054] The at least one second multi-modal information is input to a state model for state judgment, the state model is obtained by training based on multi-modal data and corresponding state information data; and based on the output result of the state model, the state information when the first instruction is issued by each target user is determined.

[0055] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium including a computer program, when the computer program is run on an electronic device, the computer program is used to make the electronic device execute the method in the first aspect and any design thereof.

[0056] In a sixth aspect, an embodiment of the present application provides a computer program product including a computer program, the computer program is stored in a computer readable storage medium; when a processor of an electronic device reads the computer program from the computer readable storage medium, the processor executes the method in the first aspect and any design thereof.

[0057] In a seventh aspect, an embodiment of the present application provides a chip, the chip includes a computer program and a processor, and the chip is used to execute the method in the first aspect and any design thereof.

[0058] In addition, the technical effects brought by the second aspect to the seventh aspect and any design thereof can be referred to the technical effects brought by the different design manners in the first aspect, which will not be repeated here.

[0059] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent from the description, or can be learned by practice of the present application. The objects and other advantages of the present application will be realized and achieved by means of the structures particularly pointed out in the written description and claims hereof as well as the appended drawings. BRIEF DESCRIPTION OF DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without any creative effort.

[0061] Figure 1 A schematic diagram of a system architecture of an intelligent device provided by an embodiment of the present application;

[0062] Figure 2 A schematic diagram of a method of intelligent device control provided by an embodiment of the present application;

[0063] Figure 3 A schematic diagram of a flow of intelligent device control provided by an embodiment of the present application;

[0064] Figure 4 A schematic diagram of the structure of a control device provided by an embodiment of the present application;

[0065] Figure 5 A schematic diagram of the structure of another control device provided by an embodiment of the present application;

[0066] Figure 6 A hardware component structure diagram of a control device in an embodiment of the present application. DETAILED DESCRIPTION

[0067] For the convenience of those skilled in the art, some terms related to the present application will be explained first.

[0068] The term "and / or" in the embodiments of the present application describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.

[0069] The application scenarios described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems. In the description of the present application, unless otherwise specified, "multiple" means two or more. In the description of the embodiments of the present application, "first", "second", etc. are used only for the purpose of distinguishing the description, and cannot be understood as indicating or implying relative importance, nor can it be understood as indicating or implying order.

[0070] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0071] With the rapid development of intelligent devices, the use of intelligent devices is more and more, for example, the intelligent device is often triggered to execute corresponding functions through human-computer interaction. The current human-computer interaction process usually relies on a specific wake-up word to activate the device, which has problems such as unnatural interaction process, frequent false wake-up / missed wake-up, and multi-device interference, which seriously affects the user experience.

[0072] For example, one human-computer interaction mode is a lightweight wake-up mode based on hotword detection, which continuously monitors environmental sound by using a low-power chip and detects a specific wake-up word by using a local model. However, this mode is excessively dependent on the wake-up word, leading to an artificial and unnatural interaction process, being easily disturbed by environmental noise, resulting in false wake-up and missed wake-up, and being unable to determine whether the speech is directed to the device, especially in a multi-person conversation scenario. For another example, another human-computer interaction mode is a continuous speech monitoring and keyword retrieval mode, which continuously identifies user speech by using an intelligent system and retrieves possible instruction keywords in real time from the identification result, and no longer depends on a wake-up word, such as a partial intelligent customer service and a vehicle-mounted voice system. However, this mode has high power consumption and high requirements for chip resources, and has weak conversation context, is likely to misjudge environmental conversation as an instruction, is unable to distinguish a real user intention, and has poor system speech intention recognition capability.

[0073] Generally, the related art generally depends on a specific wake-up word to activate a device, and has problems such as an unnatural interaction process, frequent false wake-up, frequent missed wake-up, and multi-device interference, which seriously affect user experience. Meanwhile, a single voice signal lacks context understanding of a user intention and is difficult to accurately determine an interaction intention. For example, these shortcomings result in poor intelligence of a human-computer interaction system and greatly reduce the user experience of overall human-computer interaction.

[0074] Therefore, there is an urgent need for an efficient, convenient, and accurate human-computer interaction mode to better improve user experience in a human-computer interaction process.

[0075] Therefore, there is an urgent need for an efficient, convenient, and accurate human-computer interaction mode to better improve user experience in a human-computer interaction process.

[0076] The intelligent device control method in the embodiments of the present application can be implemented in a software, hardware, or combination of software and hardware manner. In the following, the electronic device is taken as an execution subject for example, and the intelligent device control method of the embodiments of the present application is introduced.

[0077] Reference is made to Figure 1As shown, it is a possible smart device structure schematic diagram applicable to the embodiments of the present application. The smart device 100 includes radio frequency (RF) circuit 110, memory 120, input unit 130, wireless fidelity (WiFi) module 170, display unit 140, sensor 150, audio circuit 160, processor 180, and power supply 190, etc.

[0078] Wherein, those skilled in the art can understand that, Figure 1 The smart device 100 structure shown in the figure is only an example and is not limited, and the smart device 100 can also include more or less components than the figure, or combine certain components, or different component arrangement.

[0079] The RF circuit 110 can be used for receiving and sending signals in the process of transmitting information or calling. Usually, the RF circuit includes but is not limited to antenna, at least one amplifier, transceiver, coupler, low noise amplifier (LNA), duplexer, etc. In addition, the RF circuit 110 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to global system for mobile communication (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), long term evolution (LTE), email, short message service (SMS), etc.

[0080] Wherein, the memory 120 can be used to store software programs and modules, and the processor 180 executes various function applications and data processing by running the software programs and modules stored in the memory 120. The memory 120 can mainly include a program storage area and a data storage area, wherein the program storage area can store operating systems, application programs required by at least one function (such as sound playing function, image playing function, etc.), etc. In addition, the memory 120 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid state memory device.

[0081] The input unit 130 can be configured to receive input digital or character information and to generate key signals related to user settings and function controls of the intelligent device 100. Specifically, the input unit 130 can include a touch panel 131, a camera device 132, an audio device 133, and other input devices 134. The camera device 132 can take a picture of a user to transmit a face image of the user to the processor 150 for face recognition. The touch panel 131, also referred to as a touch screen, can collect a touch operation (such as an operation of a user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 131) on or near the touch panel 131 and drive a corresponding connection device according to a pre-set program. Optionally, the touch panel 131 can include two parts, a touch detection device and a touch controller. The touch detection device detects a touch position of a user and detects a signal caused by a touch operation and transmits the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts the touch information into touch coordinates, and sends the touch coordinates to the processor 180. The touch controller can also receive a command from the processor 180 and execute the command. In addition, the touch panel 131 can be implemented in various types, such as a resistive type, a capacitive type, an infrared type, and a surface acoustic wave type. The audio device 133 can include, but is not limited to, a microphone and can obtain a sound feature, such as an angle of a user's voice. In addition to the touch panel 131, the camera device 132, and the audio device 133, the input unit 130 can further include other input devices 134. Specifically, the other input devices 134 can include one or more of, but are not limited to, a physical keyboard, a function key (such as a volume control key, an on-off key, etc.), a trackball, a mouse, a joystick, and the like.

[0082] The display unit 140 can be configured to display information input by a user or information provided to the user and various menus. The display unit 140 can include a display panel 141, which can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like. Further, the touch panel 131 can cover the display panel 141. When the touch panel 131 detects a touch operation on or near the touch panel 131, the touch panel 131 transmits the touch operation to the processor 180 to determine a type of the touch event, and then the processor 180 provides a corresponding visual output on the display panel 141 according to the type of the touch event.

[0083] In addition, the intelligent device 100 can further include at least one sensor 150, such as a posture sensor, a distance sensor, and other sensors.

[0084] In addition, in the embodiments of the present application, as the sensor 150, other sensors such as a barometer, a hygrometer, a thermometer, and an infrared sensor can also be configured, which will not be described here.

[0085] The light sensor can include an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 141 according to the brightness of ambient light, and the proximity sensor can turn off the display panel 141 and / or the backlight when the smart device 100 is moved to the ear.

[0086] The audio circuit 160, the speaker 161, and the microphone 162 can provide an audio interface between the user and the smart device 100. Further, the audio device 133 can cover the audio circuit 160, and when the audio device 133 acquires audio data, the acquired audio data can be converted into an electrical signal and transmitted to the processor 180 for audio feature recognition.

[0087] The WiFi belongs to a short-distance wireless transmission technology, and the smart device 100 can help the user to provide wireless broadband Internet access through the WiFi module 170. Although Figure 1 The WiFi module 170 is shown, but it can be understood that it does not belong to the essential components of the smart device 100, and can be omitted as needed without changing the essence of the application.

[0088] The processor 180 is the control center of the smart device 100, and connects all parts of the smart device 100 through various interfaces and lines, executes the software programs and / or modules stored in the storage 120 and calls the data stored in the storage 120, executes various functions and processes data of the smart device 100, and thus monitors the whole smart device 100. Optionally, the processor 180 can include one or more processing units; preferably, the processor 180 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, the user interface, and the application program, and the modem processor mainly processes wireless communication.

[0089] It can be understood that the above-mentioned modem processor can also not be integrated into the processor 180.

[0090] Optionally, the smart device 100 can also include at least one motor 190. Since the smart device 100 is a power-consuming device, the motor 190 can be a small motor, and according to the power size that the motor can provide, multiple motors 190 can be configured for the smart device.

[0091] The smart device 100 further includes a power supply (not shown in the figure) for supplying power to each component.

[0092] Preferably, the power supply can be logically connected with the processor 180 through a power management system, so that the power management system can realize the functions of managing charging, discharging, power consumption management, etc. Although not shown, the smart device 100 can also include a Bluetooth module, a headset interface, etc., which will not be described here.

[0093] It should be noted that the smart device 100 in the embodiment of the application can be a smart robot, a smart sound box, etc., or other smart devices such as a smart phone, a PAD, etc.

[0094] Based on the above description, Figure 2 An exemplary smart device control method provided by the embodiment of the application is shown, which can be controlled by a human-computer interaction manner, and is also applicable to other execution manners, which will not be limited here.

[0095] As Figure 2 shown, the flow specifically includes:

[0096] In step 201, first multi-modal information corresponding to a first instruction in a first environment is obtained, and the first multi-modal information includes audio and video information of a target user issuing the first instruction.

[0097] As an example, the first multi-modal information of the embodiment of the application includes, but is not limited to, audio information of a user and video information of the user. For example, the first multi-modal information can also include state information of the first environment, etc.

[0098] As an example, the first instruction of the embodiment of the application can be triggered by a user through a voice, etc., that is, obtaining the first multi-modal information corresponding to the first instruction can also be understood as obtaining the first multi-modal information of the user issuing the first instruction.

[0099] In some implementations, the microphone array can be used to receive sound to obtain the audio information in the first multi-modal information. In addition, the audio information in the first multi-modal information can also be processed through a preset audio front-end algorithm, so as to determine the sound characteristics of the target user. Optionally, the sound characteristics of the embodiment of the application include, but are not limited to, the timbre characteristics of the target user, the sound angle characteristics of the target user, etc.

[0100] As an example, the microphone array in the embodiment of the application can use a multi-microphone array ring array, such as a 6-microphone array ring array, to receive sound in all directions. For example, the spatial positions of each microphone in the array are different, so that the electrical signals of a single sound source in different microphones have different time delays, so that the sound source positioning can be realized through a delay and sum (DAS) algorithm.

[0101] In some implementations, the embodiment of the present application can acquire video information in the first multi-modal information through a camera. In addition, the embodiment of the present application can perform picture extraction on the video information in the first multi-modal information based on a preset picture extraction algorithm, acquire a first number of target pictures, and determine the visual feature of the target user based on the first number of target pictures.

[0102] Optionally, the visual feature includes one or more of a face angle of the target user, a coordinate of the target user, a distance between the target user and the smart device, and lip movement information of the target user. After the audio information and the video information are acquired, the embodiment of the present application can determine the state information of the target user when the first instruction is issued based on the sound feature in the audio information and the visual feature in the video information.

[0103] For example, the embodiment of the present application can extract audio information such as sound sources every 100 ms. M pictures are extracted from the acquired video information at the same time to extract visual features, and then the video features and the sound features are input into a state model. In addition, in order to align with the audio, the embodiment of the present application extracts sound features every 100 ms while extracting pictures at the same time from the video. In order to ensure the smoothness of the extracted features, the visual features at the current time are extracted based on the current picture and the previous (M-1) pictures. Optionally, the embodiment of the present application can use a residual network (resnet) structure model, the input is M pictures, and the output is the visual feature of the person in the current picture.

[0104] Further, when the first multi-modal information is acquired in the first environment, the embodiment of the present application can acquire the second multi-modal information in time periods, and then obtain the first multi-modal information based on at least one second multi-modal information, so as to determine the state information of the target user when the first instruction is issued based on the first multi-modal information. Alternatively, the embodiment of the present application can acquire the first multi-modal information, divide the first multi-modal information into multiple time periods, and determine the corresponding second multi-modal information in each time period, so as to determine the state information of each target user when the first instruction is issued based on at least one second multi-modal information.

[0105] For example, the embodiment of the present application can continuously acquire audio information and video information of A time periods, align audio features and visual features of the same time period, and acquire second multi-modal information in which audio and video features of the same time period are combined. For example, audio information and video information of A1-A5 time periods are acquired respectively, then the audio information and the video information of the A1 time period are combined in features to obtain multi-modal information of the A1 time period, and the audio information and the video information of the A2-A5 time period are combined in features to obtain multi-modal information corresponding to the A2-A5 time period. The multi-modal information corresponding to the A1-A5 time period is summarized to obtain the first multi-modal information.

[0106] For example, the embodiment of the present application can acquire audio information and video information of B time periods in a targeted manner, each time period in the B time periods can be a part or all of non-continuous time periods, then, audio features and visual features of the same time period are aligned to obtain second multi-modal information in which audio and video features of the same time period are combined. For example, audio information and video information of B1, B2, B4, and B7 time periods are acquired respectively, then the audio information and the video information of the B1 time period are combined in features to obtain multi-modal information of the B1 time period, and the audio information and the video information of the B2, B4, and B7 time period are combined in features to obtain multi-modal information corresponding to the B2, B4, and B7 time period. The multi-modal information corresponding to the B1, B2, B4, and B7 time period is summarized to obtain the first multi-modal information.

[0107] As an example, the time period selected in the embodiment of the present application can be a time period that meets a time period selection condition. The time period selection condition includes, but is not limited to, the presence of a keyword in the time period, the presence of a key timbre in the time period, and the time period being a key time period.

[0108] For example, assuming that the entire detection duration can be divided into C1-C5 time periods, and the C2, C3, and C4 time periods detect the keyword pre-configured in the embodiment of the present application, such as the words of turning on, turning off, and adjusting, then the C2, C3, and C4 time periods can be determined as selected time periods.

[0109] For another example, assuming that the entire detection duration can be divided into C1-C5 time periods, and the C3, C4, and C5 time periods detect the key timbre pre-configured in the embodiment of the present application, such as the timbre of user 1, then the C3, C4, and C5 time periods can be determined as selected time periods. Optionally, the embodiment of the present application can set a user-specific control, for example, for the exclusive space of user 1 (such as the home of user 1, or the bedroom, etc.), only user 1 can control the smart device in the exclusive space.

[0110] For another example, assuming that the entire detection duration can be divided into C1-C5 time periods, and C2 and C3 time periods are determined as high-frequency time periods for the user to control the smart device through long-term user usage habit record learning or based on user usage time period configuration, C2 and C3 time periods can be determined as the selected time periods.

[0111] For another example, assuming that the entire detection duration can be divided into C1-C5 time periods, and C2, C3 and C4 time periods detect the keywords preconfigured in the embodiments of the present application, and C3, C4 and C5 time periods detect the key timbres preconfigured in the embodiments of the present application, C3 and C4 time periods with both keywords and key timbres can be determined as the selected time periods. For example, in the home of user 1, user 2 says "It's so hot, turn on the air conditioner", at this time, although there is a keyword, since user 2 is not a stored key timbre, the air conditioner opening operation can not be performed, when user 1 responds to user 2 and replies, such as user 1 saying "Yes, it is really hot, turn on the air conditioner", at this time, there are both keywords and key timbres, triggering the air conditioner opening operation.

[0112] It should be noted that the time period selection condition in the embodiments of the present application can be learned and updated at any time according to the user's usage, so as to better adapt to the user's usage demand and improve the user experience. In addition, in order to better avoid information loss, the embodiments of the present application can also analyze the entire detection duration and focus on analyzing the selected time period.

[0113] As an example, the embodiments of the present application can perform time period division or selection operation based on time stamp.

[0114] As an example, the length of the time period in the embodiments of the present application can be set according to actual conditions, for example, when the environment is noisy, in order to effectively improve the accuracy of the obtained multi-modal information, the time period can be set longer, when the environment is less noisy, the time period can be set shorter; for another example, in order to save the energy consumption of the smart device, the time period can be set shorter at night and longer during the day, and the specific division of the length of the time period can also be performed according to different time scenarios, which is not limited herein.

[0115] The embodiments of the present application will be described below Figure 2 The steps described above.

[0116] In step 202, state information of the target user when issuing the first instruction is determined based on the first multi-modal information, and the state information includes direction information of the target user when issuing the first instruction.

[0117] As an example, the embodiments of the present application can input the acquired first multi-modal information to a state model trained based on a neural network to perform state determination, so as to determine the direction of the first instruction. For example, the state model trained based on the neural network determines whether the first instruction is speaking to the robot or speaking to someone else, etc. In addition, the state model trained based on the neural network can also determine the state of the environmental noise based on the first multi-modal information, which is not limited here.

[0118] In some implementations, in order to better improve the accuracy of intelligent device control and the flexibility of information processing, the embodiments of the present application can also acquire the state information of the first environment, so as to better process the first multi-modal information based on the state information of the first environment.

[0119] For example, the embodiments of the present application can determine the environment type of the first environment based on the state information of the first environment, the environment type including one or more of environment scene, environment atmosphere, environment noise level, and environment theme; determine the first processing mode based on the mapping relationship between the environment type and the processing mode of the first multi-modal information; and apply the first processing mode to process the first multi-modal information to determine the state information when the target user issues the first instruction.

[0120] The environment scene includes but is not limited to open air, indoor, etc.; the environment atmosphere includes but is not limited to warm, excited, etc.; the environment noise level includes but is not limited to quiet, noisy, and can specifically set a noise level, etc.; and the environment theme includes but is not limited to home, office, coffee shop, etc. It can be understood that the above description of the environment type is only an example of the present application, and can be set according to actual conditions, which is not limited here.

[0121] When determining the first processing mode based on the environment type, the related information of the entire environment can be comprehensively analyzed, such as setting different weights for the environment scene, the environment atmosphere, the environment noise level, etc., so as to obtain the comprehensive analysis result of the first environment based on the weight assignment, and determine the first processing mode based on the comprehensive analysis result of the first environment. Alternatively, when determining the first processing mode based on the environment type, the first processing mode can also be determined based on the information of a certain aspect of the environment type.

[0122] As an example, the embodiments of the present application can determine the current scene as a simple scene through the analysis result of the first environment. For example, based on the analysis result of the first environment, when it is determined that the first environment is a simple single-person scene, the embodiments of the present application can determine the person located in the first environment as the target user, analyze the first multi-modal information corresponding to the target user, and thus control the intelligent device.

[0123] Optionally, the embodiments of the present application can also limit the target user according to actual needs, for example, only user A is set as the target user in the first environment, when only user B exists in the first environment, since user B is not the target user in the first environment, the multi-modal information analysis processing is not performed on user B. In other words, user B does not have the right to control the intelligent device in the first environment, therefore, in order to save system overhead, the multi-modal information of user B does not need to be collected and analyzed. Wherein, how to determine the user identity in the first environment can take various ways, such as face recognition, judging according to user's behavior habits, etc., which are not limited here.

[0124] As an example, the embodiments of the present application can determine whether the current scene is a multi-user complex scene through the analysis result of the first environment. When based on the multi-user complex scene, in order to better and more accurately control the intelligent device, the embodiments of the present application can separate the detected first multi-modal information, determine at least one target user, and analyze the first multi-modal information corresponding to the target user, thereby controlling the intelligent device.

[0125] As an example, the embodiments of the present application can label the audio information and video information in the acquired first multi-modal information every first time interval; separate the users based on the audio information and video information under different labels, and determine at least one target user; determine the state information of each target user when issuing the first instruction. Optionally, the embodiments of the present application can separate the users involved in different labels based on the mapping relationship between the user and the tone, and / or the mapping relationship between the user and the appearance, to obtain at least one target user.

[0126] As an example, the label labeled by the embodiments of the present application can be a timestamp. For example, the embodiments of the present application can label the first multi-modal information with a timestamp label every first time interval, thereby dividing the first multi-modal information into multi-time period multi-modal information based on the timestamp label, and performing multi-modal information analysis for different time periods.

[0127] As an example, the label punched by the embodiment of the application can be a user quantity label. For example, the embodiment of the application can punch a user quantity single label on a time period in which only one user voice is detected in the first multi-modal information, and punch a user quantity multiple label on a time period in which multiple user voices are detected; or the embodiment of the application can punch a user quantity label every first time interval, for example, determine the user quantity label corresponding to a first time period based on the number of user voices in the first time period. It can be understood that the first time interval in the embodiment of the application can be a variable value or a fixed value, which is not limited here.

[0128] As an example, the embodiment of the application can be processed in parallel for multiple target users, that is, when multiple users issue a first instruction to a smart device, the state information when each target user issues the first instruction can be determined respectively, so that the first instruction issued by each target user is executed in turn or simultaneously. For example, assuming that there are multiple smart devices in the room, such as air conditioner A, air conditioner B, television A, smart curtain A, smart curtain B, etc., and it is currently determined that there are target user 1 and target user 2 by analyzing the detected multi-modal information, the embodiment of the application can simultaneously analyze and process the multi-modal information corresponding to the target user 1 and the target user 2, such as determining that the first instruction issued by the target user 1 is to turn on the air conditioner, and the state when the first instruction is issued is to speak to air conditioner A, and determining that the first instruction issued by the target user 2 is to close the curtain, and speak to curtain B, then the embodiment of the application can simultaneously close curtain B and open air conditioner A, or execute in turn, and the specific execution order can be set according to the actual situation.

[0129] As an example, the embodiment of the application can be processed in parallel for multiple target users, that is, when multiple users issue a first instruction to a smart device, the first instruction issued by each target user can be fused and processed, so as to more accurately control the smart device. For example, assuming that there are multiple smart devices in the room, such as air conditioner A, fan A, smart window A, television A, smart curtain A, smart curtain B, etc., and it is currently determined that there are target user 1 and target user 2 by analyzing the detected multi-modal information, the embodiment of the application can simultaneously analyze and process the multi-modal information corresponding to the target user 1 and the target user 2, such as analyzing and processing the collected multi-modal information to obtain the interaction content of the target user 1 and the target user 2 as follows:

[0130] The target user 1 points to the air conditioner A with a finger and says, "It is very hot, and we can turn on the air conditioner for the meeting." The target user 2 says, "The air conditioner is just in front of my seat, and the blowing feels uncomfortable on my shoulders. We can turn on the window for the meeting." The target user 1 says to the fan A, "We can turn on the fan for the meeting." The target user 2 says, "OK, let's turn on the fan." In this case, the embodiment of the present application determines to turn on the fan A by the multi-round interaction context content of the target user 1 and the target user 2, so as to more accurately and efficiently control the intelligent device.

[0131] The embodiment of the present application is described below Figure 2 The step.

[0132] In step 203, when it is determined that the state information meets the wake-up condition, the wake-up of the intelligent device is triggered, and the first instruction is executed.

[0133] As an example, the wake-up condition in the embodiment of the present application includes determining that the target user issues the first instruction to the intelligent device, such as speaking to the robot, so as to activate the artificial interaction system, execute the first instruction, and perform the subsequent ASR, intent recognition, TTS, and other voice interaction services.

[0134] As an example, the embodiment of the present application can analyze the first multi-modal information to obtain the sound angle feature of the target user, the face angle feature of the target user, the distance between the target user and the intelligent device, the coordinates of the target user, the feature of the lip movement of the target user, and the like, to comprehensively determine the state of the target user, such as whether the target user is speaking, whether the target user is speaking to the intelligent device or other users, and whether the target user is speaking to a specific intelligent device.

[0135] Through the above method, the embodiment determines the direction of the user sending the first instruction by combining the multi-modal information such as audio and video of the user, so that the process of determining whether to wake up the intelligent device and execute the corresponding first instruction based on the direction of the first instruction does not need to use the wake-up word to wake up the intelligent device, making the human-computer interaction more natural. The fusion of audio and video features enables the system to more accurately determine whether the user has an interaction intention, effectively reduces the misoperation, improves the accuracy and robustness of the system response, can more easily track one of multiple people for long context interaction, is more in line with human interaction, improves the human-computer interaction experience, and improves the intelligence.

[0136] In order to better introduce the above-mentioned intelligent device control method, the related details involved in the embodiment of the present application are described in detail based on a scene example below, and the specific examples are not limited to the following examples:

[0137] As shown in Figure 3 Fig. 1 is a schematic diagram of a control process of an intelligent device provided by an embodiment of the present application, which can include the following steps:

[0138] In step 301, an audio feature of a first instruction issued by a target user is acquired, and a sound angle feature is extracted.

[0139] In step 302, a visual feature of the first instruction issued by the target user is acquired, and a visual angle feature is extracted.

[0140] As an example, the visual angle feature described in the embodiment of the present application includes, but is not limited to, a face angle, a coordinate, a distance, a lip movement, and the like of the target user issuing the first instruction.

[0141] In step 303, the sound angle feature and the visual angle feature are aligned in a time axis to obtain multi-modal information.

[0142] As an example, the embodiment of the present application can separately use a separate sound acquisition device to acquire audio and a separate image acquisition device to acquire images, combine the acquired audio and images to obtain first multi-modal information, and then further align the sound angle feature and the visual angle feature of the first multi-modal information in a time axis to obtain second multi-modal information, i.e., the second multi-modal information is further optimized and determined based on the first multi-modal information. Alternatively, the embodiment of the present application can directly acquire audio and video information based on a collection device with audio and video functions, i.e., acquire first multi-modal information, and then perform audio processing and video processing on the audio information in the first multi-modal information to obtain second multi-modal information, i.e., the second multi-modal information is further optimized and determined based on the first multi-modal information. Alternatively, the embodiment of the present application can separately use a separate sound acquisition device to acquire audio and perform feature processing on the acquired audio to obtain a sound angle feature, and a separate image acquisition device to acquire images and perform feature processing on the acquired images to obtain a visual angle feature, and then align the sound angle feature and the visual angle feature in a time axis to obtain multi-modal information, i.e., the first multi-modal information and the second multi-modal information are the same at this time.

[0143] The multi-modal information described in the embodiment of the present application can include one frame of data or multiple frames of data, and the number of data frames included can be set according to actual conditions.

[0144] In step 304, the multi-modal information is input into a state model for state judgment.

[0145] In step 305, it is determined whether the target user speaks to the smart device when issuing the first instruction based on the output result of the state model. If yes, step 306 is performed, and if no, step 307 is performed.

[0146] In step 306, the smart device corresponding to the first instruction is woken up, and the first instruction is executed.

[0147] In step 307, the current process is ended.

[0148] Through the multi-modal wake-up-free interaction technology in the embodiments of the present application, the naturalness and intelligent level of human-computer interaction can be significantly improved, and the user can realize the "speak and use" non-sensing interaction experience. By fusing visual features such as mouth movement and face orientation, the system can more accurately determine whether the user has an interaction intention, effectively reduce misoperation, and improve the accuracy and robustness of system response. In addition, the above technical solutions in the embodiments of the present application effectively meet the development trend of active sensing and context understanding of intelligent terminals and service robots, and are an important basis for promoting the next-generation robot intelligent interaction system.

[0149] Based on the same inventive concept, the embodiments of the present application also provide a control device. The control device can perform the steps or operations in the intelligent device control method provided by any one of the above embodiments, and can achieve the same technical effects. In this embodiment, the structure of the control device can be as shown in the Figure 4 The control device can perform the steps or operations in the intelligent device control method provided by any one of the above embodiments, and can achieve the same technical effects. In this embodiment, the structure of the control device can be as shown in the

[0150] The memory 401 is used to store the computer programs executed by the processor 402. The memory 401 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and programs required for running instant messaging functions, etc.; and the data storage area can store various instant messaging information and operation instruction sets, etc.

[0151] The memory 401 can be a volatile memory such as a random-access memory (RAM); the memory 401 can also be a non-volatile memory such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 401 can be any other medium capable of carrying or storing a desired computer program in the form of an instruction or a data structure and capable of being accessed by a computer, but is not limited thereto. The memory 401 can be a combination of the above memories.

[0152] The processor 402 can include one or more central processing units (CPUs) or digital processing units, etc. The processor 402 is configured to invoke the computer program stored in the memory 401 to implement the intelligent device control method.

[0153] The communication module 404 is configured to communicate with the control device and other servers.

[0154] The specific connection medium between the memory 401, the communication module 403 and the processor 402 is not limited in the embodiments of the present application. In the embodiments of the present application, the memory 401 and the processor 402 are connected through a bus 404, and the bus 404 is described by a thick line in the embodiments of the present application. Figure 4 The connection mode between other components is only schematically described, and is not limited. The bus 404 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of description, the bus 404 is only described by a thick line in the embodiments of the present application, but it is not described that there is only one bus or only one type of bus. Figure 4 Figure 4 The memory 401 stores a computer storage medium, and the computer storage medium stores computer executable instructions. The computer executable instructions are used to implement the intelligent device control method of the embodiments of the present application. The processor 402 is configured to execute the intelligent device control method.

[0155] Based on the same inventive concept, the embodiments of the present application provide another control device 500.

[0156] The embodiments of the present application provide a structure diagram of a control device. As shown in the figure, the device includes:

[0157] Figure 5 The embodiments of the present application provide a structure diagram of a control device. As shown in the figure, the device includes: Figure 5 The acquisition module 501 is configured to acquire first multi-modal information corresponding to a first instruction in a first environment, and the first multi-modal information includes audio and video information of a target user who issues the first instruction.

[0158] The processing module 502 is configured to determine state information of the target user when the target user issues the first instruction based on the first multi-modal information, and the state information includes direction information of the target user when the target user issues the first instruction. When it is determined that the state information satisfies a wake-up condition, the wake-up of the intelligent device is triggered, and the first instruction is executed. The wake-up condition includes determining that the target user issues the first instruction towards the intelligent device.

[0159] As an implementation manner, the processing module 502 is specifically configured to:

[0160]

[0161] ​​The audio information in the first multi-modal information is processed based on a preset audio front-end algorithm, and a sound feature of the target user is determined, the sound feature including a sound angle feature of the target user; the video information in the first multi-modal information is picture extracted based on a preset picture extraction algorithm, and a first number of target pictures are obtained; a visual feature of the target user is determined based on the first number of target pictures, the visual feature including one or more of a face angle of the target user, a coordinate of the target user, a distance between the target user and the smart device, and lip movement information of the target user; and state information of the target user when the first instruction is issued is determined based on the sound feature and the visual feature.

[0162] As an implementation manner, the sound feature further includes a timbre feature of the target user.

[0163] As an implementation manner, the processing module 502 is specifically configured to:

[0164] obtain state information of the first environment; determine an environment type of the first environment based on the state information of the first environment, the environment type including one or more of an environment scene, an environment atmosphere, and an environment noise level; determine a first processing manner based on a mapping relationship between the environment type and a processing manner of the first multi-modal information; and process the first multi-modal information by applying the first processing manner, to determine the state information of the target user when the first instruction is issued.

[0165] As an implementation manner, the processing module 502 is specifically configured to:

[0166] label the audio information and the video information in the obtained first multi-modal information at every first time interval; separate users based on the audio information and the video information under different labels, to determine at least one target user; and determine the state information of each target user when the first instruction is issued.

[0167] As an implementation manner, the processing module 502 is specifically configured to:

[0168] Separate the users involved in different labels based on a mapping relationship between the users and timbres, and / or a mapping relationship between the users and appearances, to obtain at least one target user.

[0169] As an implementation manner, the processing module 502 is specifically configured to:

[0170] Align features of the video information and the audio information of the same time period based on the label of the audio information and the label of the video information; obtain second multi-modal information in which audio-video features of at least one time period are combined together; and determine state information of each target user when the first instruction is issued based on the at least one second multi-modal information.

[0171] As an implementation form, the processing module 502 is specifically configured to:

[0172] input the at least one second multi-modal information into a state model for state determination, the state model being obtained by training based on multi-modal data and corresponding state information data; and determine the state information of each target user when the first instruction is issued based on an output result of the state model.

[0173] Based on the same inventive concept, the embodiments of the present application provide a computer readable storage medium, and a computer program product, which comprises computer program code. When the computer program code is run on a computer, the computer executes the control method of the intelligent device as any one of the foregoing embodiments. Since the principle of solving problems of the computer readable storage medium is similar to that of the transaction identification method, the implementation of the computer readable storage medium can be referred to the implementation of the method, and the repeated parts will not be described herein.

[0174] The control device 600 according to this embodiment of the present application will be described below with reference to Figure 6 The control device 600 is only an example and should not limit the functions and use range of the embodiments of the present application. Figure 6 The control device 600 is in the form of a general control device. The components of the control device 600 can include but are not limited to the at least one processing unit 601, the at least one storage unit 602, and a bus 603 connecting different system components including the storage unit 1002 and the processing unit 1001.

[0175] As shown in Figure 6 , the control device 600 is in the form of a general control device. The components of the control device 600 can include but are not limited to the at least one processing unit 601, the at least one storage unit 602, and a bus 603 connecting different system components including the storage unit 1002 and the processing unit 1001.

[0176] The bus 603 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a processor or a local bus using any of a variety of bus structures.

[0177] The storage unit 602 can include a readable medium in the form of volatile memory, such as random access memory (RAM) 621 and / or cache memory 622, and can further include read only memory (ROM) 623. The storage unit 602 can also include a program / utility 625 having a set of programs / modules 624, including an operating system, one or more application programs, other program modules, and program data, each of which or a combination of which can include implementation of a network environment as in each of these examples or some combination thereof.

[0178] The control device 600 can also communicate with one or more external devices 604 such as a keyboard or a pointing device, among other devices, which can enable a user to interact with the control device 600 and / or any devices (e.g., a router, a modem, and so on) that enable the control device 600 to communicate with one or more other computing devices. Such communication can occur via an input / output (I / O) interface 605. Still yet, the control device 600 can communicate with one or more networks, such as a local area network (LAN), a wide area network (WAN), and / or the Internet, through a network adapter 606. As Figure 6 illustrated, the network adapter 606 can communicate with the other components of the control device 600 through the bus 603. It should be understood that although not shown, other hardware and / or software components could be used in conjunction with the control device 600. These include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0179] Based on the same inventive concept, the embodiments of the present application also provide a smart device control system, which can include the foregoing smart device and a state model trained based on a neural network. The smart device can execute the control method executed by the control device in any one of the foregoing embodiments.

[0180] The embodiments of the present application also provide a computer program product. The methods in the present application can be implemented, wholly or partially, by software, hardware, firmware, or any combination thereof. When implemented by software, the methods can be implemented in the form of a computer program product, wholly or partially. The computer program product includes one or more computer programs or instructions. When the computer programs or instructions are loaded and executed on a computer, the processes or functions described in the present application are executed, wholly or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, a core network device, an OAM, or other programmable devices.

[0181] The computer readable storage medium can be an implementation of the computer program product, i.e., the embodiments of the present application also provide a computer readable storage medium including a computer program, which, when executed by a processor, implements any of the intelligent device control methods described above.

[0182] The computer program or instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer program or instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center through wired or wireless manner. The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium, for example, a floppy disk, a hard disk, a magnetic tape; or an optical medium, for example, a digital video disc; or a semiconductor medium, for example, a solid state disk. The computer readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile storage media.

[0183] The above embodiments of the present application are introduced from the perspective of the electronic device as the execution subject. In order to implement the functions in the above embodiments of the present application, the electronic device can include hardware structures and / or software modules, and implement the above functions in the form of hardware structures, software modules, or hardware structures plus software modules. Whether a certain function in the above functions is implemented in the form of hardware structure, software module, or hardware structure plus software module depends on specific application and design constraints of the technical solutions.

[0184] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0185] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart

[0186] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart

[0187] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for functionally implementing the steps listed in the flowchart

[0188] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A method of controlling an intelligent device, the method comprising: The method comprises: obtaining first multi-modal information corresponding to a first instruction in a first environment, the first multi-modal information comprising audio and video information of a target user issuing the first instruction; determining state information of the target user when issuing the first instruction based on the first multi-modal information, the state information comprising directional information of the target user when issuing the first instruction; when the state information meets a wake-up condition, triggering the wake-up of a smart device and executing the first instruction, the wake-up condition comprising determining that the target user is facing the smart device when issuing the first instruction.

2. The method of claim 1, wherein, The determination of the state information of the target user when issuing the first instruction based on the first multi-modal information comprises: processing audio information in the first multi-modal information based on a preset audio front-end algorithm to determine a voice feature of the target user, the voice feature comprising a voice angle feature of the target user; performing picture extraction on video information in the first multi-modal information based on a preset picture extraction algorithm to obtain a first number of target pictures; determining a visual feature of the target user based on the first number of target pictures, the visual feature comprising one or more of a face angle of the target user, a coordinate of the target user, a distance between the target user and the smart device, and lip movement information of the target user; determining the state information of the target user when issuing the first instruction based on the voice feature and the visual feature.

3. The method according to claim 1 or 2, characterized in that, The determination of the state information of the target user when issuing the first instruction based on the first multi-modal information comprises: obtaining state information of the first environment; determining an environment type of the first environment based on the state information of the first environment, the environment type comprising one or more of an environment scene, an environment atmosphere, and an environment noise level; determining a first processing mode based on a mapping relationship between the environment type and the processing mode of the first multi-modal information; processing the first multi-modal information using the first processing mode to determine the state information of the target user when issuing the first instruction.

4. The method of claim 3, wherein, The processing of the first multi-modal information using the first processing mode to determine the state information of the target user when issuing the first instruction comprises: labeling audio information and video information obtained in the first multi-modal information at every first time interval; separating users based on audio information and video information under different labels to determine at least one target user; determining the state information of each target user when issuing the first instruction.

5. The method of claim 4, wherein, The separation of users based on audio information and video information under different labels to determine at least one target user comprises: separating users involved in different labels based on a mapping relationship between users and voice timbres and / or a mapping relationship between users and appearances to obtain at least one target user.

6. The method according to claim 4 or 5, characterized in that, The determination of the state information of each target user when issuing the first instruction comprises: aligning the video information and the audio information in the same time period based on the labels of the audio information and the labels of the video information. Obtain second multi-modal information in which audio and video features in at least one time period are combined together; Determine state information of each target user when the first instruction is issued based on the at least one second multi-modal information.

7. The method of claim 6, wherein, The determination of the state information of each target user when the first instruction is issued based on the at least one second multi-modal information comprises: Input the at least one second multi-modal information into a state model for state judgment, the state model being obtained by training based on multi-modal data and corresponding state information data; Determine the state information of each target user when the first instruction is issued based on an output result of the state model.

8. A control device characterized by comprising: Comprise: An acquisition module, configured to acquire first multi-modal information corresponding to a first instruction in a first environment, the first multi-modal information comprising audio and video information of a target user issuing the first instruction; A processing module, configured to determine state information of the target user when the first instruction is issued based on the first multi-modal information, the state information comprising direction information of the target user when the first instruction is issued; when the state information meets a wake-up condition, trigger wake-up of an intelligent device and execute the first instruction, the wake-up condition comprising determining that the target user is facing the intelligent device when the first instruction is issued.

9. An apparatus, comprising: Comprise: A memory, configured to store program instructions; A processor, configured to call the program instructions stored in the memory, and execute the method according to any one of claims 1 to 7 according to the program.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions, and the computer executable instructions are used to make the computer execute the method according to any one of claims 1 to 7.