Human-computer interaction system based on artificial intelligence

By designing a human-computer interaction system based on artificial intelligence, collecting user iris and visual data, analyzing user intentions, and adjusting smart homes according to user behavior, the shortcomings of existing systems in user intention analysis and equipment adjustment are solved, and the accuracy and user experience of the system are improved.

CN119937368AInactive Publication Date: 2025-05-06XIQIAO TECH NANJING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411862315.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for existing smart home systems to accurately analyze user intentions when users use them, and they are prone to misjudgment or unresponsiveness, which affects the user experience, and cannot adjust the equipment according to users' usage habits.

Method used

A human-computer interaction system based on artificial intelligence is designed to collect user's iris-reflected infrared data, visual images and environmental data through the data acquisition module. The human-computer interaction module analyzes the user's action command intentions, and adjusts the smart home according to the user's behavior through the device control module.

Benefits of technology

It improves the accuracy of the system, can more accurately judge the device selected by the user, reduces misjudgment and unresponsiveness, improves the user experience, and adjusts the equipment according to the user's usage habits to provide more comfortable services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937368A_ABST
    Figure CN119937368A_ABST
Patent Text Reader

Abstract

The invention discloses a man-machine interaction system based on artificial intelligence, which comprises a data acquisition module, a man-machine interaction module and an equipment control module, and is characterized in that the data acquisition module is used for collecting reflected infrared data of user irises, shooting visual images in a room and collecting environment data in the room; the man-machine interaction module is used for analyzing and selecting a smart home which a user wants to adjust and analyzing an action instruction intention of the user, and the equipment control module is used for analyzing a scene and an environment where the user is located and adjusting the smart home according to behaviors of the user. The data acquisition module, the man-machine interaction module and the equipment control module are in communication connection with one another, the data acquisition module comprises an infrared device module, a visual module and a sensor module, and the infrared device module is used for irradiating infrared rays to eyeballs of a user to acquire iris characteristics of the user. The method has the characteristics of improving judgment accuracy and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction, and in particular to a human-computer interaction system based on artificial intelligence. Background Art

[0002] With the rapid development of Internet technology, household appliances in people's homes are gradually becoming intelligent. Smart homes connect various appliances in the home through the Internet of Things technology and control them through intelligent networking. At present, most smart homes are controlled by voice, but when family members are resting, voice control will affect the family's rest, so visual command control has emerged. However, since people's body behaviors are very random, users may inadvertently make command actions. The existing technology cannot accurately analyze the user's intentions, and it is easy to make misjudgments or no responses. The user's intentions may not match the action commands, affecting the user experience. Moreover, since the user's use of the device is highly random, the existing technology cannot adjust the device according to the user's usage habits, so that the user has a better user experience. Therefore, it is very necessary to design an artificial intelligence-based human-computer interaction system that improves judgment accuracy and user experience. Summary of the invention

[0003] The purpose of the present invention is to provide a human-computer interaction system based on artificial intelligence to solve the problems raised in the above background technology.

[0004] In order to solve the above technical problems, the present invention provides the following technical solutions: an artificial intelligence-based human-computer interaction system, comprising a data acquisition module, a human-computer interaction module and a device control module, characterized in that: the data acquisition module is used to collect the reflected infrared data of the user's iris, capture the visual image in the room and collect the environmental data in the room, the human-computer interaction module is used to analyze and select the smart home that the user wants to adjust, and analyze the user's action instruction intention, the device control module is used to analyze the scene and environment where the user is located and adjust the smart home according to the user's behavior, and the data acquisition module, the human-computer interaction module and the device control module are connected to communicate with each other;

[0005] The device selection module includes a device analysis submodule, an iris recognition submodule and an action recognition submodule. The device analysis submodule is used to analyze the location of the device, the iris recognition submodule is used to recognize the characteristic data of the user's eye iris, and the action recognition submodule is used to further determine whether the user intends to adjust the device;

[0006] The instruction analysis module includes a feature extraction submodule, an instruction judgment submodule and a scene selection submodule. The feature extraction submodule is used to extract frames of the video and extract user behavior features in each frame image. The instruction judgment submodule is used to compare and judge whether the extracted user behavior features conform to the action instructions. The scene selection submodule is used to analyze the surrounding environment and user behavior features and select the corresponding scene according to the analysis results.

[0007] According to the above technical solution, the data acquisition module includes an infrared device module, a visual module and a sensor module. The infrared device module is used to irradiate infrared rays to the user's eyeballs to collect the user's iris features, the visual module is used to collect visual images of the user's behavior and actions in the room, and the sensor module is used to collect environmental data in the room.

[0008] According to the above technical solution, the human-computer interaction module includes a device selection module and an instruction analysis module. The device selection module is used to identify the user's device adjustment instructions and analyze the user's iris data to select the corresponding device. The instruction analysis module is used to analyze the user's behavioral actions and determine whether the user has the intention to adjust the device.

[0009] According to the above technical solution, the device control module includes an environment analysis module and a behavior analysis module. The environment analysis module is used to detect the surrounding environment, and the behavior analysis module is used to analyze the user's action instructions in the current scene.

[0010] According to the above technical solution, the device control module also includes a device adjustment module, and the device adjustment module is used to adjust the operating state of the device according to the user's action instructions.

[0011] According to the above technical solution, the operation method of the human-computer interaction system mainly includes the following steps:

[0012] Step S1: anchoring the user and the smart home in the room through the visual module, capturing the visual image in the room in real time, irradiating the user's eyeballs with infrared rays through the infrared device module, collecting the user's iris data, and entering the environmental data in the room into the system through the sensor module;

[0013] Step S2: When determining the device that the user wants to adjust, the system sends an electrical signal to trigger the device selection module to start, start analyzing the angle of the user facing the device and the iris data, and determine the device that the user wants to adjust based on the analysis results;

[0014] Step S3: After determining the device that the user wants to adjust, the system starts the instruction analysis module, starts analyzing the scene the user is in, and determines whether the action instruction made by the user meets the signal triggering standard in the current scene, and adjusts the device if so;

[0015] Step S4: When the user uses the system, the system collects the user's usage habits, starts the habit analysis module, begins to analyze the user's usage habits, and adjusts the device according to the user's usage habits.

[0016] According to the above technical solution, step S2 further includes the following steps:

[0017] Step S21: retrieve the visual image of the room taken by the visual module, scan and identify the device outline feature nodes and the human head feature nodes in the image, construct the device outline model and the human head model according to the device outline feature nodes and the human head feature nodes, respectively, establish a three-dimensional coordinate system in the room model and measure the coordinates of the device and the coordinates of the human head model, measure the angle between the direction of the human head facing and the vertical plane where the device is located, and the angle between the direction of the human head facing and the horizontal plane, and calculate the coordinates of the intersection of the human head model and the plane where the device is located by the formula:

[0018]

[0019] Z3=Z2

[0020] Wherein, (X3, Y3, Z3) represents the coordinates of the intersection point, (X1, Y1, Z1) represents the coordinates of the human head model, Z2 represents the vertical axis coordinate of the vertical plane where the device is located, θ1 represents the angle between the human body facing vector and the vertical plane where the device is located, and θ2 represents the angle between the human body facing vector and the horizontal plane;

[0021] Step S22: Calculate the distance L between the current intersection and the device using the distance formula between two points. If the distance L is greater than the system-set threshold, the device corresponding to the distance is marked as a non-selected state. Otherwise, the device is marked as an alternative state. Retrieve all the devices in the alternative state, sort the distance L in ascending order, mark the device with the smallest distance as the user's pre-selected device, retrieve the user's iris data and motion data, determine whether the user's iris features and motion features meet the conditions, and control the device based on the determination result.

[0022] According to the above technical solution, step S22 further includes the following steps:

[0023] Step S221: retrieve iris data, build a user's eye iris model based on the iris data, build an infrared irradiation point in the user's eye iris model according to the coordinates of the device, use infrared irradiation to measure the angle between the iris and the infrared, if the angle between the iris and the infrared is greater than a first threshold and less than a second threshold, mark the current device as a user selected device, otherwise mark the current device as a non-selected device, retrieve all user selected devices, identify whether there is a user pre-selected tag in the user selected device, if there is a user pre-selected tag in the current user selected device, mobilize the current user selected device to standby state, otherwise the system has no response;

[0024] Step S222: retrieve the visual image of the user's action, scan and identify the user's body contour feature nodes, build a body contour model based on the user's body contour feature nodes, and compare the body contour model with the trigger model in the system. If the similarity between the body contour model and the trigger model in the system is greater than a threshold, the system triggers the current device to start, otherwise the system continues to verify.

[0025] According to the above technical solution, step S3 further includes the following steps:

[0026] Step S31: retrieve the current visual image, identify the location tag in the visual image, read the scene setting in the location tag, retrieve the adjustment action instruction corresponding to the device according to the scene setting, scan and identify the contour feature nodes of the user in the current visual image, and when the number of contour feature nodes of the user detected in the visual image is less than the threshold, retrieve the infrared thermal image, scan and identify the human body contour nodes, if no human body contour nodes are detected, mark that there is no user in the current scene, and the system does not make any adjustments, otherwise mark that there is a user in the current scene, and the system turns on the light to shoot the visual image using the visual module;

[0027] Step S32: Identify each frame of the visual image, crop the visual image according to the first frequency set by the system, scan the cropped image, identify the action instruction features in the image, and compare the identified action instruction features with the action instruction features set by the system. When the similarity is less than the minimum threshold set by the system, the system does not respond. When the similarity is greater than the maximum threshold set by the system, crop the visual image according to the second frequency set by the system, scan and identify the action instruction features in the image, compare all the collected action instruction features with the corresponding action instruction features in the system database, and calculate the average similarity of the current action instruction features through the formula Where, i = 1, 2, 3...n, represents the average similarity of the current action instruction features, S iIt represents the similarity of each action command feature, β represents the error coefficient of the average similarity of the action command features. If the average similarity of the current action command features is greater than the system set threshold, the device is called to execute the current action command, otherwise the system does not respond.

[0028] According to the above technical solution, in step S4, the historical usage records of the device and the environmental data and visual image data during use in the system database are retrieved, the user's device usage habit characteristics, environmental characteristics during use, and user behavior habit characteristics during use are extracted, the above characteristics are used as samples to train the user habit model, the environmental data in the room collected by the sensor module is retrieved, the temperature, humidity and light data in the current room are identified, the light, humidity and temperature data in the room are imported into the user habit model, and the user's habitual light, humidity and temperature data are compared. If they do not match, the corresponding device is adjusted for adjustment, otherwise no adjustment is made. When the user wakes up, the visual image is retrieved, the user behavior characteristics in the visual image are identified, and the user behavior habit characteristics extracted by the system are compared. If the similarity is greater than the threshold, the device is controlled to be turned on according to the user's usage habits.

[0029] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: the present invention, by converting an image into a room model, can accurately measure the coordinates of the device and the human head, and then accurately calculate the coordinates of the user's gaze focus point, and by calculating the distance between the intersection and the device, it is determined whether the distance between the intersection and the device meets the conditions, and the factors affecting the system's judgment can be eliminated, and then the device selected by the user can be more accurately judged, which greatly improves the accuracy of the system. By using iris recognition to perform a secondary verification of the device, the system selects the device that has passed the dual verification as the selected device, which can ensure that the system will not make a misjudgment when selecting a device, resulting in a response error that affects the user's experience, thereby improving the accuracy of the system and the user experience. By detecting vision Whether the number of contour nodes in the visual module meets the number set by the system can avoid incomplete visual images taken by the visual module due to light reasons, resulting in the inability to recognize complete action instructions, and can avoid errors in the extracted action instruction features, which leads to system judgment errors, greatly improving the accuracy of the system. By making preliminary judgments on the action instructions and further judging them, the judgment of the action instructions can be more accurate, avoiding misjudgment and wrong judgment, further improving the accuracy of the system. By analyzing the user's usage habits and training the user habit model, the system can adjust the smart home in the room according to the user's usage habits, provide users with more comfortable services, and greatly improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0031] Figure 1 It is a schematic diagram of the system module composition of the present invention. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] See also Figure 1 , the present invention provides a technical solution: a human-computer interaction system based on artificial intelligence, including a data acquisition module, a human-computer interaction module and a device control module, characterized in that: the data acquisition module is used to collect the reflected infrared data of the user's iris, capture the visual image in the room and collect the environmental data in the room, the human-computer interaction module is used to analyze and select the smart home that the user wants to adjust, and analyze the user's action instruction intention, the device control module is used to analyze the scene and environment where the user is located and adjust the smart home according to the user's behavior, and the data acquisition module, the human-computer interaction module and the device control module are connected to each other in communication;

[0034] The device selection module includes a device analysis submodule, an iris recognition submodule and a motion recognition submodule. The device analysis submodule is used to analyze the location of the device, the iris recognition submodule is used to identify the characteristic data of the user's iris, and the motion recognition submodule is used to further determine whether the user intends to adjust the device.

[0035] The instruction analysis module includes a feature extraction submodule, an instruction judgment submodule and a scene selection submodule. The feature extraction submodule is used to extract frames of the video and extract user behavior features in each frame image. The instruction judgment submodule is used to compare and judge whether the extracted user behavior features conform to the action instructions. The scene selection submodule is used to analyze the surrounding environment and user behavior characteristics, and select the corresponding scene according to the analysis results.

[0036] The data acquisition module includes an infrared device module, a visual module and a sensor module. The infrared device module is used to irradiate infrared rays to the user's eyes to collect the user's iris features. The visual module is used to collect visual images of the user's behavior and actions in the room. The sensor module is used to collect environmental data in the room.

[0037] The human-computer interaction module includes a device selection module and an instruction analysis module. The device selection module is used to identify the user's device adjustment instructions and analyze the user's iris data to select the corresponding device. The instruction analysis module is used to analyze the user's behavioral actions and determine whether the user has the intention to adjust the device.

[0038] The device control module includes an environment analysis module and a behavior analysis module. The environment analysis module is used to detect the surrounding environment, and the behavior analysis module is used to analyze the user's action instructions in the current scene.

[0039] The device control module also includes a device adjustment module, which is used to adjust the operating state of the device according to the user's action instructions.

[0040] The operation method of the human-computer interaction system mainly includes the following steps:

[0041] Step S1: anchoring the user and the smart home in the room through the visual module, capturing the visual image in the room in real time, irradiating the user's eyeballs with infrared rays through the infrared device module, collecting the user's iris data, and entering the environmental data in the room into the system through the sensor module;

[0042] Step S2: When determining the device that the user wants to adjust, the system sends an electrical signal to trigger the device selection module to start, start analyzing the angle of the user facing the device and the iris data, and determine the device that the user wants to adjust based on the analysis results;

[0043] Step S3: After determining the device that the user wants to adjust, the system starts the instruction analysis module, starts analyzing the scene the user is in, and determines whether the action instruction made by the user meets the signal triggering standard in the current scene, and adjusts the device if so;

[0044] Step S4: When the user uses the system, the system collects the user's usage habits, starts the habit analysis module, begins to analyze the user's usage habits, and adjusts the device according to the user's usage habits.

[0045] Step S2 further comprises the following steps:

[0046] Step S21: retrieve the visual image of the room taken by the visual module, scan and identify the device outline feature nodes and the human head feature nodes in the image, construct the device outline model and the human head model according to the device outline feature nodes and the human head feature nodes, respectively, establish a three-dimensional coordinate system in the room model and measure the coordinates of the device and the coordinates of the human head model, measure the angle between the direction of the human head facing and the vertical plane where the device is located, and the angle between the direction of the human head facing and the horizontal plane, and calculate the coordinates of the intersection of the human head model and the plane where the device is located by the formula:

[0047]

[0048] Z3=Z2,

[0049] In the formula, (X3, Y3, Z3) represents the coordinates of the intersection, (X1, Y1, Z1) represents the coordinates of the human head model, Z2 represents the vertical axis coordinate of the vertical plane where the device is located, θ1 represents the angle between the human body facing vector and the vertical plane where the device is located, and θ2 represents the angle between the human body facing vector and the horizontal plane. By converting the image into a room model, the coordinates of the device and the human head can be accurately measured, and then the coordinates of the user's gaze focus point can be accurately calculated;

[0050] Step S22: Calculate the distance L between the current intersection and the device using the distance formula between two points. If the distance L is greater than the threshold set by the system, the device corresponding to the distance will be marked as a non-selected state. Otherwise, the device will be marked as an alternative state. Retrieve all the devices in the alternative state, sort the distance L in ascending order, mark the device with the smallest distance as the user's pre-selected device, retrieve the user's iris data and motion data, determine whether the user's iris features and motion features meet the conditions, control the device based on the determination result, and determine whether the distance between the intersection and the device meets the conditions by calculating the distance between the intersection and the device. This can eliminate factors that affect the system's determination, and thus more accurately determine the device selected by the user, greatly improving the accuracy of the system.

[0051] Step S22 further includes the following steps:

[0052] Step S221: retrieve iris data, build a user's eye iris model based on the iris data, build an infrared irradiation point in the user's eye iris model according to the coordinates of the device, use infrared irradiation to measure the angle between the iris and the infrared, if the angle between the iris and the infrared is greater than the first threshold and less than the second threshold, mark the current device as a user selected device, otherwise mark the current device as a non-selected device, retrieve all user selected devices, identify whether there is a user pre-selected tag in the user selected device, if there is a user pre-selected tag in the current user selected device, mobilize the current user selected device to standby mode, otherwise the system has no response, and use iris recognition to perform secondary verification on the device, the system selects the device that has passed the double verification as the selected device, which can ensure that the system does not make a misjudgment when selecting a device, resulting in a response error that affects the user experience, thereby improving the accuracy of the system and the user experience;

[0053] Step S222: retrieve the visual image of the user's action, scan and identify the user's body contour feature nodes, build a body contour model based on the user's body contour feature nodes, and compare the body contour model with the trigger model in the system. If the similarity between the body contour model and the trigger model in the system is greater than a threshold, the system triggers the current device to start, otherwise the system continues to verify.

[0054] Step S3 further comprises the following steps:

[0055] Step S31: retrieve the current visual image, identify the location tag in the visual image, read the scene setting in the location tag, retrieve the corresponding adjustment action instruction of the device according to the scene setting, scan and identify the contour feature nodes of the user in the current visual image, and when the number of contour feature nodes of the user detected in the visual image is less than the threshold, retrieve the infrared thermal image, scan and identify the human body contour nodes, if the human body contour nodes cannot be detected, it is marked that there is no user in the current scene, and the system does not make any adjustments, otherwise it is marked that there is a user in the current scene, and the system turns on the light to shoot the visual image using the visual module, and by detecting whether the number of contour nodes in the visual module meets the number set by the system, it can avoid that the visual image shot by the visual module is incomplete due to light reasons, resulting in the inability to identify the complete action instruction, and it can avoid that the extracted action instruction features or errors occur, resulting in errors in system judgment, thereby greatly improving the accuracy of the system;

[0056] Step S32: Identify each frame of the visual image, crop the visual image according to the first frequency set by the system, scan the cropped image, identify the action instruction features in the image, and compare the identified action instruction features with the action instruction features set by the system. When the similarity is less than the minimum threshold set by the system, the system does not respond. When the similarity is greater than the maximum threshold set by the system, crop the visual image according to the second frequency set by the system, scan and identify the action instruction features in the image, compare all the collected action instruction features with the corresponding action instruction features in the system database, and calculate the average similarity of the current action instruction features through the formula Where, i = 1, 2, 3...n, represents the average similarity of the current action instruction features, S i It represents the similarity of each action instruction feature, β represents the error coefficient of the average similarity of the action instruction feature. If the average similarity of the current action instruction feature is greater than the system set threshold, the device is called to execute the current action instruction, otherwise the system does not respond. By making a preliminary judgment on the action instruction and then making further judgment, the judgment of the action instruction can be more accurate, avoiding misjudgment and wrong judgment, and further improving the accuracy of the system.

[0057] In step S4, the historical usage records of the device and the environmental data and visual image data during use in the system database are retrieved, the user's device usage habit characteristics, environmental characteristics during use, and user behavior habit characteristics during use are extracted, and the above characteristics are used as samples to train the user habit model, and the environmental data in the room collected by the sensor module are retrieved to identify the temperature, humidity and light data in the current room, and the light, humidity and temperature data in the room are imported into the user habit model, and compared with the user's habitual light, humidity and temperature data. If they do not match, the corresponding device is adjusted, otherwise no adjustment is made. When the user wakes up, the visual image is retrieved, the user behavior characteristics in the visual image are identified, and compared with the user behavior habit characteristics extracted by the system. If the similarity is greater than the threshold, the device is controlled to be turned on according to the user's usage habits. By analyzing the user's usage habits to train the user habit model, the system can adjust the smart home in the room according to the user's usage habits, provide users with more comfortable services, and greatly improve the user's usage experience.

[0058] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0059] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A human-computer interaction system based on artificial intelligence, comprising a data acquisition module, a human-computer interaction module and a device control module, characterized in that: The data acquisition module is used to collect the user's iris data, capture the visual images in the room and collect the environmental data in the room. The human-computer interaction module is used to analyze and select the smart home that the user wants to adjust and analyze the user's action instruction intention. The device control module is used to analyze the scene and environment where the user is and adjust the smart home according to the user's behavior. The data acquisition module, the human-computer interaction module and the device control module are connected to each other for communication; The human-computer interaction module includes a device selection module and an instruction analysis module. The device selection module is used to identify the user's device adjustment instruction and analyze the user's iris data to select the corresponding device. The instruction analysis module is used to analyze the user's behavior and determine whether the user has the intention to adjust the device. The device selection module includes a device analysis submodule, an iris recognition submodule and an action recognition submodule. The device analysis submodule is used to analyze the location of the device, the iris recognition submodule is used to recognize the characteristic data of the user's eye iris, and the action recognition submodule is used to further determine whether the user intends to adjust the device; The instruction analysis module includes a feature extraction submodule, an instruction judgment submodule and a scene selection submodule. The feature extraction submodule is used to extract frames from the video and extract user behavior features in each frame image. The instruction judgment submodule is used to compare and judge whether the extracted user behavior features conform to the action instructions. The scene selection submodule is used to analyze the surrounding environment and user behavior features and select the corresponding scene according to the analysis results. The data acquisition module includes an infrared device module, a visual module and a sensor module. The infrared device module is used to irradiate infrared rays to the user's eyeballs to collect the user's iris data. The visual module is used to collect the user's behavior and action visual images in the room. The sensor module is used to collect the environment data in the room. The operation method of the human-computer interaction system mainly includes the following steps: Step S1: anchoring the user and the smart home in the room through the visual module, capturing the visual image in the room in real time, irradiating the user's eyeballs with infrared rays through the infrared device module, collecting the user's iris data, and entering the environmental data in the room into the system through the sensor module; Step S2: When determining the device that the user wants to adjust, the system sends an electrical signal to trigger the device selection module to start, start analyzing the angle of the user facing the device and the iris data, and determine the device that the user wants to adjust based on the analysis results; Step S3: After determining the device that the user wants to adjust, the system starts the instruction analysis module, starts analyzing the scene the user is in, and determines whether the action instruction made by the user meets the signal triggering standard in the current scene, and adjusts the device if so; Step S4: When the user uses the system, the system collects the user's usage habits, starts the habit analysis module, starts analyzing the user's usage habits, and adjusts the device according to the user's usage habits; The step S2 further comprises the following steps: Step S21: retrieve the visual image of the room taken by the visual module, scan and identify the device outline feature nodes and the human head feature nodes in the image, construct the device outline model and the human head model according to the device outline feature nodes and the human head feature nodes, respectively, establish a three-dimensional coordinate system in the room model and measure the coordinates of the device and the coordinates of the human head model, measure the angle between the direction of the human head facing and the vertical plane where the device is located, and the angle between the direction of the human head facing and the horizontal plane, and calculate the coordinates of the intersection of the human head model and the plane where the device is located by the formula: Z3=Z2 Wherein, (X3, Y3, Z3) represents the coordinates of the intersection point, (X1, Y1, Z1) represents the coordinates of the human head model, Z2 represents the vertical axis coordinate of the vertical plane where the device is located, θ1 represents the angle between the human body facing vector and the vertical plane where the device is located, and θ2 represents the angle between the human body facing vector and the horizontal plane; Step S22: Calculate the distance L between the current intersection and the device using the distance formula between two points. If the distance L is greater than the threshold set by the system, mark the device corresponding to the distance as a non-selected state. Otherwise, mark the device as a candidate state. Retrieve all the devices in the candidate state, sort the distance L in ascending order, mark the device with the smallest distance as the user's pre-selected device, retrieve the user's iris data and motion data, determine whether the user's iris features and motion features meet the conditions, and control the device according to the determination result. The step S22 further comprises the following steps: Step S221: retrieve iris data, build a user's eye iris model based on the iris data, build an infrared irradiation point in the user's eye iris model according to the coordinates of the device, use infrared irradiation to measure the angle between the iris and the infrared, if the angle between the iris and the infrared is greater than a first threshold and less than a second threshold, mark the current device as a user selected device, otherwise mark the current device as a non-selected device, retrieve all user selected devices, identify whether there is a user pre-selected tag in the user selected device, if there is a user pre-selected tag in the current user selected device, mobilize the current user selected device to standby state, otherwise the system has no response; Step S222: Retrieve the visual image of the user's action, scan and identify the user's human body contour feature nodes, build a human body contour model based on the user's human body contour feature nodes, and compare the human body contour model with the trigger model in the system. If the similarity between the human body contour model and the trigger model in the system is greater than a threshold, the system triggers the current device to start, otherwise the system continues to verify; The step S3 further comprises the following steps: Step S31: retrieve the current visual image, identify the location tag in the visual image, read the scene setting in the location tag, retrieve the adjustment action instruction corresponding to the device according to the scene setting, scan and identify the contour feature nodes of the user in the current visual image, and when the number of contour feature nodes of the user detected in the visual image is less than the threshold, retrieve the infrared thermal image, scan and identify the human body contour nodes, if no human body contour nodes are detected, mark that there is no user in the current scene, and the system does not make any adjustments, otherwise mark that there is a user in the current scene, and the system turns on the light to shoot the visual image using the visual module; Step S32: Identify each frame of the visual image, crop the visual image according to the first frequency set by the system, scan the cropped image, identify the action instruction features in the image, and compare the identified action instruction features with the action instruction features set by the system. When the similarity is less than the minimum threshold set by the system, the system does not respond. When the similarity is greater than the maximum threshold set by the system, crop the visual image according to the second frequency set by the system, scan and identify the action instruction features in the image, compare all the collected action instruction features with the corresponding action instruction features in the system database, and calculate the average similarity of the current action instruction features through the formula Where, i = 1, 2, 3...n, represents the average similarity of the current action instruction features, S i It represents the similarity of each action command feature, β represents the error coefficient of the average value of the action command feature. If the average similarity of the current action command feature is greater than the system set threshold, the device is called to execute the current action command, otherwise the system does not respond.