Human brain and machine interaction system and method

Through the trigger module, tracking module, determination module, intent analysis module and feedback module, combined with audio positioning and eye tracking technology, the problem of accurate positioning and intention recognition of interactive objects in multi-person scenarios is solved, and the accuracy and reliability of intelligent interaction are achieved.

CN120704528APending Publication Date: 2025-09-26SHENZHEN MICHOI SECURITY TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510817579.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In complex environments, existing intelligent robots find it difficult to accurately identify interaction objects, especially in multi-person scenarios, where interaction targets are not accurately determined.

Method used

It uses a trigger module, tracking module, determination module, intent analysis module and feedback module, combined with audio positioning technology and eye tracking, to determine the intention of the interactive object and output feedback through line of sight tracking and voice positioning.

Benefits of technology

Accurate identification of interactive objects and effective interaction are achieved in complex environments, improving the accuracy and reliability of intelligent interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704528A_ABST
    Figure CN120704528A_ABST
Patent Text Reader

Abstract

The invention provides a human brain and machine interaction system and method. The system comprises a trigger module, a tracking module, a determination module, an intention analysis module and a feedback module. Wherein the triggering module is used for analyzing an interaction environment to determine whether an interaction environment analysis result meets a triggering condition or not; if yes, triggering tracking; the tracking module performs target sight tracking; the determination module determines an interaction object according to the target sight tracking result; the intention analysis module is used for tracking eyeballs of the interaction object in the interaction process, comprehensively analyzing interaction information and eyeball tracking information of the interaction object and determining the intention of the interaction object; and the feedback module is used for outputting corresponding feedback according to the intention of the interaction object. According to the human brain and machine interaction system and method, accurate determination of an interaction target in a complex environment is realized, and effective intelligent interaction is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction technology, and in particular to a human brain-machine interaction system and method. Background Art

[0002] Human-computer interaction, or human-machine interaction, essentially refers to the interaction between humans and computers, or more broadly, humans and "machines containing computers." Human-computer interaction is a core technology in intelligent robots. While existing intelligent robots have developed sophisticated interaction techniques for single-person scenarios, in complex environments, such as those with multiple people, accurately identifying the interacting party remains a pressing technical challenge. Summary of the Invention

[0003] One of the purposes of the present invention is to provide a human brain-machine interaction system and method to achieve accurate determination of interaction targets in complex environments and ensure effective intelligent interaction.

[0004] An embodiment of the present invention provides an interaction system between the human brain and a machine, comprising: a trigger module, a tracking module, a determination module, an intention analysis module, and a feedback module; wherein the trigger module is used to analyze the interaction environment to determine whether the interaction environment analysis result meets the trigger condition; when the condition is met, tracking is triggered; the tracking module performs target line of sight tracking; the determination module determines the interaction object based on the target line of sight tracking result; the intention analysis module is used to track the eyeballs of the interaction object during the interaction process, comprehensively analyze the interaction information and eye tracking information of the interaction object, and determine the intention of the interaction object; the feedback module is used to output corresponding feedback based on the intention of the interaction object.

[0005] Preferably, the triggering conditions include: multiple people existing in a pre-configured interaction area, and / or multiple people existing in a positioning area positioned using audio positioning technology.

[0006] Preferably, the human brain and machine interaction system further includes:

[0007] The audio positioning module is used to locate the interactive voice and determine the positioning area where the interactive object is located.

[0008] Preferably, the audio positioning module performs the following operations:

[0009] Constructing a positioning space according to the positions of the at least two configured audio receiving modules;

[0010] Based on the time difference of the audio received by the audio receiving module, a positioning area is determined from the positioning space.

[0011] Preferably, the determination module determines the interaction object based on the target sight tracking result and performs the following operations:

[0012] The interactive object is determined based on the matching conditions between the gaze landing objects of each object to be analyzed in the first object set of the target gaze tracking result and the relevant objects in the second object set corresponding to the interactive voice.

[0013] Preferably, the steps of constructing the first object set are as follows:

[0014] Monitor the dwell time of the sight point;

[0015] When the single stay time at the same location is greater than a preset first time threshold, and / or when the user stays at the same location for a preset number of times, the target at the stay location is taken as a sight point object in the first object set.

[0016] Preferably, determining the interactive object according to the matching between the gaze landing object of each object to be analyzed in the first object set of the target gaze tracking result and each related object in the second object set corresponding to the interactive voice includes:

[0017] Determine a matching value based on the subjective value corresponding to the matched sight point object and the importance coefficient corresponding to the related object;

[0018] The target with the largest matching value is extracted as the interaction object.

[0019] Preferably, the human brain and machine interaction system further includes: an intention assistance module for tracking the eyes of the interactive object during the interaction, and combining the eye tracking results and interaction data to determine the interactive operation.

[0020] Preferably, the human brain and machine interaction system further includes: an interaction end judgment module, which is used to determine whether to end the interaction based on the eye tracking result of the interaction object after the interaction operation is performed.

[0021] The present invention also provides a method for interaction between a human brain and a machine, comprising:

[0022] Analyze the interaction environment to determine whether the interaction environment analysis results meet the trigger conditions;

[0023] When satisfied, trigger tracking;

[0024] Tracking the target's sight;

[0025] Determine the interaction object based on the target gaze tracking results;

[0026] Track the eyes of the interactive object during the interaction;

[0027] Comprehensively analyze the interaction information and eye tracking information of the interactive object to determine the intention of the interactive object;

[0028] Output corresponding feedback based on the intention of the interactive object.

[0029] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0030] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0032] Figure 1 Schematic diagram of a human brain-machine interaction system according to an embodiment of the present invention;

[0033] Figure 2 Schematic diagram of a method for interaction between the human brain and a machine in an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0035] Example 1

[0036] The embodiment of the present invention provides a human brain and machine interaction system, such as Figure 1 As shown, it includes: a trigger module 1, a tracking module 2, a determination module 3, an intention analysis module 4 and a feedback module 5; wherein, the trigger module 1 is used to analyze the interactive environment to determine whether the interactive environment analysis result meets the trigger condition; when it meets the condition, the tracking is triggered; the tracking module 2 performs target sight tracking; the determination module 3 determines the interactive object based on the target sight tracking result; the intention analysis module 4 is used to track the eyeballs of the interactive object during the interaction process, comprehensively analyze the interactive information and eye tracking information of the interactive object, and determine the intention of the interactive object; the feedback module 5 is used to output corresponding feedback based on the intention of the interactive object.

[0037] The human brain and machine interaction system of this embodiment integrates audio and image processing technologies to realize human-computer interaction. During the interaction process, it tracks the eyes of the interactive object, comprehensively analyzes the interactive information and eye tracking information of the interactive object, determines the intention of the interactive object, and ensures the accuracy of the interactive feedback.

[0038] Among them, the trigger conditions include: the presence of multiple people in the pre-configured interaction area. The pre-configured interaction area is specifically configured according to the actual situation. Generally, it can be configured as a preset range (any value between 1m-5m) centered on the intelligent robot; it can also be configured as a fan-shaped area with a preset angle in front of the intelligent robot, for example: an area within a 120-degree fan-shaped area within a distance of 3 meters; the acquisition of the interaction environment can be obtained by analyzing the data detected by detection sensors such as infrared detection sensors; the analysis of the interaction environment is to determine the number of people around the robot and their distribution;

[0039] The determination module determines the interaction object based on the target gaze tracking results and performs the following operations:

[0040] The interactive object is determined based on the matching conditions between the gaze landing objects of each object to be analyzed in the first object set of the target gaze tracking result and the relevant objects in the second object set corresponding to the interactive voice.

[0041] The steps for constructing the first object set are as follows:

[0042] Monitor the dwell time of the sight point;

[0043] When the single stay time at the same location is greater than the preset first time threshold, and / or when staying at the same location for a preset number of times, the target at the stay location is taken as a sight point object in the first object set. The specific steps for monitoring the sight point location include: first, performing image analysis processing on the human eye area to obtain the movement of the eyeball and the position of the pupil; according to the position of the pupils of both eyes in the eyes, the sight direction is determined from the pre-configured sight determination table; in the pre-configured environmental space, a human body model is mapped at the personnel position and extended according to the sight direction, and the point in contact with the object in the corresponding extension direction is taken as the sight point, and the position of this point is the sight point location; the first time threshold includes any value from 200ms to 10s; the preset number includes any value from 2 times to 10 times;

[0044] The interactive object is determined based on a matching condition between the gaze landing object of each object to be analyzed in the first object set of the target gaze tracking result and each related object in the second object set corresponding to the interactive voice, including:

[0045] Determine a matching value based on the subjective value corresponding to the matched sight point object and the importance coefficient corresponding to the related object; the matching value is the product of the subjective value corresponding to the matched sight point object and the importance coefficient corresponding to the related object;

[0046] The target with the largest matching value is extracted as the interaction object.

[0047] Among them, the subjective value is obtained by analyzing the monitoring data during the target gaze tracking process. For example, the subjective value can be allocated according to the total duration of the gaze on the object where the gaze falls and the total duration of the gaze tracking; the important coefficient is to configure the second object set for the interactive voice and to configure the association for each configured related object one by one, and it is based on the analysis and configuration of professionals based on the association between the interactive voice and each related object.

[0048] Example 2

[0049] An embodiment of the present invention provides an interaction system between the human brain and a machine, comprising: a trigger module, a tracking module, a determination module, an intention analysis module and a feedback module; wherein the trigger module is used to analyze the interaction environment to determine whether the interaction environment analysis result meets the trigger condition; when the condition is met, tracking is triggered; the tracking module performs target line of sight tracking; the determination module determines the interaction object based on the target line of sight tracking result; the intention analysis module is used to track the eyeballs of the interaction object during the interaction process, comprehensively analyze the interaction information and eye tracking information of the interaction object, and determine the intention of the interaction object; the feedback module is used to output corresponding feedback based on the intention of the interaction object.

[0050] The triggering conditions include: the presence of multiple people in the positioning area located by audio positioning technology.

[0051] In order to achieve audio positioning, the human brain and machine interaction system also includes:

[0052] The audio positioning module is used to locate the interactive voice and determine the positioning area where the interactive object is located.

[0053] The audio positioning module performs the following operations:

[0054] Constructing a positioning space according to the positions of the at least two configured audio receiving modules;

[0055] Based on the time difference of the audio received by the audio receiving module, the positioning area is determined from the positioning space. By pre-configuring a time difference and angle range table, the corresponding angle range is extracted from the table based on the time difference received by the audio receiving module, and the positioning area around the intelligent robot is determined based on the angle range; for example, if the time difference is 0, the corresponding angle range is a sector area of ​​plus or minus 5 degrees in front of or behind the robot;

[0056] The determination module determines the interaction object based on the target gaze tracking results and performs the following operations:

[0057] The interactive object is determined based on the matching conditions between the gaze landing objects of each object to be analyzed in the first object set of the target gaze tracking result and the relevant objects in the second object set corresponding to the interactive voice.

[0058] The steps for constructing the first object set are as follows:

[0059] Monitor the dwell time of the sight point;

[0060] When the single stay time at the same location is greater than a preset first time threshold, and / or when the user stays at the same location for a preset number of times, the target at the stay location is taken as a sight point object in the first object set.

[0061] The interactive object is determined based on a matching condition between the gaze landing object of each object to be analyzed in the first object set of the target gaze tracking result and each related object in the second object set corresponding to the interactive voice, including:

[0062] Determine a matching value based on the subjective value corresponding to the matched sight point object and the importance coefficient corresponding to the related object;

[0063] The target with the largest matching value is extracted as the interaction object.

[0064] Example 3

[0065] An embodiment of the present invention provides an interaction system between the human brain and a machine, comprising: a trigger module, a tracking module, a determination module, an intention analysis module and a feedback module; wherein the trigger module is used to analyze the interaction environment to determine whether the interaction environment analysis result meets the trigger condition; when the condition is met, tracking is triggered; the tracking module performs target line of sight tracking; the determination module determines the interaction object based on the target line of sight tracking result; the intention analysis module is used to track the eyeballs of the interaction object during the interaction process, comprehensively analyze the interaction information and eye tracking information of the interaction object, and determine the intention of the interaction object; the feedback module is used to output corresponding feedback based on the intention of the interaction object.

[0066] The triggering conditions include: multiple people are present in a pre-configured interactive area, and multiple people are present in a positioning area located using audio positioning technology.

[0067] In order to achieve audio positioning, the human brain and machine interaction system also includes:

[0068] The audio positioning module is used to locate the interactive voice and determine the positioning area where the interactive object is located.

[0069] The audio positioning module performs the following operations:

[0070] Constructing a positioning space according to the positions of the at least two configured audio receiving modules;

[0071] Based on the time difference of the audio received by the audio receiving module, a positioning area is determined from the positioning space.

[0072] The determination module determines the interaction object based on the target gaze tracking results and performs the following operations:

[0073] The interactive object is determined based on the matching conditions between the gaze landing objects of each object to be analyzed in the first object set of the target gaze tracking result and the relevant objects in the second object set corresponding to the interactive voice.

[0074] The steps for constructing the first object set are as follows:

[0075] Monitor the dwell time of the sight point;

[0076] When the single stay time at the same location is greater than a preset first time threshold, and / or when the user stays at the same location for a preset number of times, the target at the stay location is taken as a sight point object in the first object set.

[0077] The interactive object is determined based on a matching condition between the gaze landing object of each object to be analyzed in the first object set of the target gaze tracking result and each related object in the second object set corresponding to the interactive voice, including:

[0078] Determine a matching value based on the subjective value corresponding to the matched sight point object and the importance coefficient corresponding to the related object;

[0079] The target with the largest matching value is extracted as the interaction object.

[0080] Example 4

[0081] An embodiment of the present invention provides an interaction system between the human brain and a machine, comprising: a trigger module, a tracking module, a determination module, an intention analysis module and a feedback module; wherein the trigger module is used to analyze the interaction environment to determine whether the interaction environment analysis result meets the trigger condition; when the condition is met, tracking is triggered; the tracking module performs target line of sight tracking; the determination module determines the interaction object based on the target line of sight tracking result; the intention analysis module is used to track the eyeballs of the interaction object during the interaction process, comprehensively analyze the interaction information and eye tracking information of the interaction object, and determine the intention of the interaction object; the feedback module is used to output corresponding feedback based on the intention of the interaction object.

[0082] In addition, the human brain and machine interaction system also includes: an interaction end judgment module, which is used to determine whether to end the interaction based on the eye tracking results of the interaction object after the interaction operation is performed.

[0083] When determining the interactive object, due to interference from other factors, the interactive object may not be accurately determined. In order to cope with this situation, the human brain and machine interaction system also includes: an interactive object analysis and adjustment module; the interactive object analysis and adjustment module performs the following operations: after outputting the first voice for interaction, if the second voice of the determined interactive object is not received within a preset time or the second voice is not corresponding to the first voice, the interactive object adjustment analysis mode is entered; after entering the interactive object adjustment analysis mode, the regional image of the determined interactive object is obtained, and the line of sight analysis and gesture analysis are performed on the regional image; it is determined whether the line of sight direction and the gesture direction are pointing to a third-party target; when any of them points to a third-party target or both point to the same third-party target, the pointed third-party target is used as the target for interactive object adjustment; when neither points to a third-party target or points to different third-party targets, the pointing targets of the line of sight directions and / or gesture directions of other people within a preset range are obtained; the pointing targets are counted, and the pointing target with the largest statistical number is used as the target for interactive object adjustment. The correspondence between the first and second voices can be analyzed using a pre-configured analysis library. A count is incremented if the third-party target is in the direction of any other person's gaze or gesture. By analyzing the behavior of the identified interaction partner and surrounding people, the correct interaction partner can be more accurately identified and adjusted in the event of an incorrect interaction partner.

[0084] Example 5

[0085] The present invention also provides a method for interaction between human brain and machine, such as Figure 2 As shown, including:

[0086] Step 1: Analyze the interaction environment to determine whether the interaction environment analysis results meet the trigger conditions;

[0087] Step 2: When the conditions are met, trigger tracking; track the target's sight line;

[0088] Step 3: Determine the interaction object based on the target gaze tracking results;

[0089] Step 4: Track the eyes of the interacting object during the interaction process;

[0090] Step 5: Comprehensively analyze the interaction information and eye tracking information of the interaction object to determine the intention of the interaction object;

[0091] Step 6: Output corresponding feedback based on the intention of the interactive object.

[0092] Among them, the trigger conditions include: the presence of multiple people in the pre-configured interaction area. The pre-configured interaction area is specifically configured according to the actual situation. Generally, it can be configured as a preset range (any value between 1m-5m) centered on the intelligent robot; it can also be configured as a fan-shaped area with a preset angle in front of the intelligent robot, for example: an area within a 120-degree fan-shaped area within a distance of 3 meters; the acquisition of the interaction environment can be obtained by analyzing the data detected by detection sensors such as infrared detection sensors; the analysis of the interaction environment is to determine the number of people around the robot and their distribution;

[0093] Among them, according to the target line of sight tracking results, the interactive object is determined and the following operations are performed:

[0094] The interactive object is determined based on the matching conditions between the gaze landing objects of each object to be analyzed in the first object set of the target gaze tracking result and the relevant objects in the second object set corresponding to the interactive voice.

[0095] The steps for constructing the first object set are as follows:

[0096] Monitor the dwell time of the sight point;

[0097] When the single stay time at the same location is greater than the preset first time threshold, and / or when staying at the same location for a preset number of times, the target at the stay location is taken as a sight point object in the first object set. The specific steps for monitoring the sight point location include: first, performing image analysis processing on the human eye area to obtain the movement of the eyeball and the position of the pupil; according to the position of the pupils of both eyes in the eyes, the sight direction is determined from the pre-configured sight determination table; in the pre-configured environmental space, a human body model is mapped at the personnel position and extended according to the sight direction, and the point in contact with the object in the corresponding extension direction is taken as the sight point, and the position of this point is the sight point location; the first time threshold includes any value from 200ms to 10s; the preset number includes any value from 2 times to 10 times;

[0098] The interactive object is determined based on a matching condition between the gaze landing object of each object to be analyzed in the first object set of the target gaze tracking result and each related object in the second object set corresponding to the interactive voice, including:

[0099] Determine a matching value based on the subjective value corresponding to the matched sight point object and the importance coefficient corresponding to the related object; the matching value is the product of the subjective value corresponding to the matched sight point object and the importance coefficient corresponding to the related object;

[0100] The target with the largest matching value is extracted as the interaction object.

[0101] Among them, the subjective value is obtained by analyzing the monitoring data during the target gaze tracking process. For example, the subjective value can be allocated according to the total duration of the gaze on the object where the gaze falls and the total duration of the gaze tracking; the important coefficient is to configure the second object set for the interactive voice and to configure the association for each configured related object one by one, and it is based on the analysis and configuration of professionals based on the association between the interactive voice and each related object.

[0102] When determining an interaction object, due to interference from other factors, the interaction object may not be accurately determined. To address this situation, the human brain and machine interaction method further includes: after outputting a first voice for interaction, if a second voice of the determined interaction object is not received within a preset time or the second voice does not correspond to the first voice, entering an interaction object adjustment analysis mode; after entering the interaction object adjustment analysis mode, obtaining a regional image of the determined interaction object, performing eye gaze analysis and gesture analysis on the regional image; determining whether the eye gaze direction and gesture direction are pointing to a third-party target; when either of them points to a third-party target or both point to the same third-party target, using the pointed third-party target as the target for interaction object adjustment; when neither of them points to a third-party target or points to different third-party targets, obtaining the eye gaze direction and / or gesture direction of other persons within a preset range; counting the pointed targets, and using the pointed target with the largest statistical number as the target for interaction object adjustment. The correspondence between the first voice and the second voice can be analyzed using a pre-configured analysis library; and when the third-party target is the pointed target of any other person's eye gaze direction or gesture direction, the statistical number is increased by one. By analyzing the behavior of the identified interaction object and the behavior of the surrounding people, it is possible to more accurately determine the correct interaction object and make adjustments when the interaction object is incorrectly determined.

[0103] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A human brain and machine interaction system, characterized in that: include: Trigger module, tracking module, determination module, intention analysis module and feedback module; among them, the trigger module is used to analyze the interactive environment to determine whether the interactive environment analysis results meet the trigger conditions; when satisfied, trigger tracking; the tracking module performs target line of sight tracking; the determination module determines the interactive object based on the target line of sight tracking results; the intention analysis module is used to track the eyeballs of the interactive object during the interaction process, comprehensively analyze the interactive information and eye tracking information of the interactive object, and determine the intention of the interactive object; the feedback module is used to output corresponding feedback based on the intention of the interactive object.

2. The human brain-machine interaction system according to claim 1, wherein: The triggering conditions include: multiple people existing in a pre-configured interaction area, and / or multiple people existing in a positioning area positioned using audio positioning technology.

3. The human brain-machine interaction system according to claim 2, wherein: Also includes: The audio positioning module is used to locate the interactive voice and determine the positioning area where the interactive object is located.

4. The human brain-machine interaction system according to claim 3, wherein: The audio positioning module performs the following operations: Constructing a positioning space according to the positions of the at least two configured audio receiving modules; Based on the time difference of the audio received by the audio receiving module, a positioning area is determined from the positioning space.

5. The human brain-machine interaction system according to claim 1, wherein: The determination module determines the interaction object based on the target gaze tracking results and performs the following operations: The interactive object is determined based on the matching conditions between the gaze landing objects of each object to be analyzed in the first object set of the target gaze tracking result and the relevant objects in the second object set corresponding to the interactive voice.

6. The human brain-machine interaction system according to claim 5, wherein: The steps for constructing the first object set are as follows: Monitor the dwell time of the sight point; When the single stay time at the same location is greater than a preset first time threshold, and / or when the user stays at the same location for a preset number of times, the target at the stay location is taken as a sight point object in the first object set.

7. The human brain-machine interaction system according to claim 5, wherein: Determining the interaction object based on a match between the gaze landing object of each object to be analyzed in the first object set of the target gaze tracking result and each related object in the second object set corresponding to the interactive voice includes: Determine a matching value based on the subjective value corresponding to the matched sight point object and the importance coefficient corresponding to the related object; The target with the largest matching value is extracted as the interaction object.

8. The human brain-machine interaction system according to claim 1, wherein: Also includes: The interaction end judgment module is used to determine whether to end the interaction based on the eye tracking results of the interaction object after the interaction operation is performed.

9. A method for interaction between a human brain and a machine, characterized in that: include: Analyze the interaction environment to determine whether the interaction environment analysis results meet the trigger conditions; When satisfied, trigger tracking; Tracking the target's sight; Determine the interaction object based on the target gaze tracking results; Track the eyes of the interactive object during the interaction; Comprehensively analyze the interaction information and eye tracking information of the interactive object to determine the intention of the interactive object; Output corresponding feedback based on the intention of the interactive object.

10. The method for human brain-machine interaction according to claim 9, wherein: The triggering conditions include: multiple people existing in a pre-configured interaction area, and / or multiple people existing in a positioning area positioned using audio positioning technology.