Robot man-machine interaction intention recognition method and system based on cooperation of Internet of Things

By collecting multimodal perception signals and environmental context data through IoT collaboration and combining them with the robot's task status for intent matching, the problem of low intent recognition accuracy in complex environments caused by single-modal perception is solved, thus improving the naturalness and accuracy of human-computer interaction.

CN121456828AInactive Publication Date: 2026-02-03SHENZHEN HUAAISHENG TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202512009402.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-02-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex and dynamic environments, single-modal perception methods lead to a decrease in the accuracy of human-computer interaction intent recognition, affecting the naturalness and accuracy of the interaction.

Method used

Through IoT collaboration, multimodal perception signals are collected, and combined with environmental context data to generate a final contextual intent candidate set. Intent matching is then performed in conjunction with the robot's task status to identify human-computer interaction intents.

Benefits of technology

It improves the naturalness and accuracy of human-computer interaction, and avoids misidentification and invalid response caused by environmental interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456828A_ABST
    Figure CN121456828A_ABST
Patent Text Reader

Abstract

The invention provides a robot man-machine interaction intention recognition method and system based on Internet of Things cooperation, and the method comprises the steps: collecting a multi-modal original sensing signal of a user in an interaction process based on an Internet of Things sensing terminal disposed in a physical space, and obtaining a multi-modal sensing sequence, on the basis of the multi-modal sensing sequence, identifying a fragment containing an interactive behavior actively initiated by a user to obtain a target sensing fragment; taking the target sensing fragment as an interaction event, and acquiring current environment state information based on an Internet of Things environment sensing unit in a spatial position area indicated by the interaction event to obtain environment context data; generating a final context intention candidate set representing the user interaction intention based on the interaction event and the environment context data; and performing interaction intention matching based on the final context intention candidate set in combination with the current task state of the robot and the executable action to obtain a man-machine interaction intention recognition result. According to the invention, the naturalness and the interaction accuracy of human-computer interaction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method and system for recognizing human-computer interaction intentions in robots based on Internet of Things (IoT) collaboration. Background Technology

[0002] In current robotic human-computer interaction systems, a primary existing method is to recognize user intent based on a single sensor modality (e.g., speech recognition or visual gesture recognition). This method typically relies on a preset instruction set for a fixed scenario, determining the interaction intent by matching user input with a pre-stored template.

[0003] However, in complex and dynamic environments, when users express their intentions in diverse, unstructured ways or when there is environmental interference (such as background noise, lighting changes, occlusion, etc.), single-modal perception can easily lead to a significant decrease in the accuracy of intention recognition, resulting in human-computer interaction response errors or delays, which seriously affect the naturalness and accuracy of the interaction. Summary of the Invention

[0004] This invention provides a robot human-computer interaction intent recognition method and system based on Internet of Things collaboration, in order to improve the naturalness and accuracy of human-computer interaction.

[0005] In a first aspect, the present invention provides a method for recognizing robot human-computer interaction intentions based on Internet of Things (IoT) collaboration, comprising: Based on the collection of multimodal raw sensing signals of users during the interaction process by IoT sensing terminals deployed in physical space, a multimodal sensing sequence is obtained, and based on the multimodal sensing sequence, segments containing user-initiated interactive behaviors are identified to obtain target sensing segments. Using the target perception segment as an interaction event, the IoT environmental perception unit within the spatial location area indicated by the interaction event collects the current environmental state information to obtain environmental context data; Based on the interaction events and the environmental context data, a final context intent candidate set representing the user's interaction intent is generated; Based on the final contextual intent candidate set, the robot's current task state and executable actions are combined to perform interaction intent matching, and the human-computer interaction intent recognition result is obtained.

[0006] Secondly, the present invention also provides a robot human-computer interaction intent recognition system based on Internet of Things (IoT) collaboration, applied to the robot human-computer interaction intent recognition method based on IoT collaboration as described in the first aspect; the robot human-computer interaction intent recognition system based on IoT collaboration includes: The behavior perception module is used to collect multimodal raw perception signals of users during the interaction process based on IoT sensing terminals deployed in physical space, obtain multimodal perception sequences, and identify segments containing user-initiated interactive behaviors based on the multimodal perception sequences to obtain target perception segments. The environmental perception module is used to collect current environmental state information from IoT environmental perception units within the spatial location area indicated by the target perception segment as an interaction event, and obtain environmental context data. The intent recognition module is used to generate a final contextual intent candidate set representing the user's interaction intent based on the interaction event and the environmental context data; The intent matching module is used to perform interactive intent matching based on the final context intent candidate set combined with the robot's current task state and executable actions to obtain human-computer interaction intent recognition results.

[0007] Thirdly, the present invention also provides an electronic device, comprising: a memory for storing computer software programs; and a processor for reading and executing the computer software programs, thereby realizing the robot human-computer interaction intent recognition method based on Internet of Things collaboration as described in the first aspect above.

[0008] Fourthly, the present invention also provides a non-transitory computer-readable storage medium storing a computer software program, which, when executed by a processor, implements the robot human-computer interaction intent recognition method based on Internet of Things collaboration as described in the first aspect above.

[0009] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the robot human-computer interaction intent recognition method based on Internet of Things collaboration as described in the first aspect above.

[0010] The robot human-computer interaction intent recognition method based on IoT collaboration provided in this invention identifies segments containing user-initiated interactive behaviors based on multimodal perception sequences, obtaining target perception segments. Therefore, the initial behavior expressing the user's true intent is accurately captured through these target perception segments, avoiding false responses to non-interactive background signals and overcoming the misidentification problem caused by the inability to distinguish between effective interaction and environmental noise. Using the target perception segments as interaction events, environmental context data obtained based on the current environmental state information within the spatial location area indicated by the interaction event provides physical context information related to the location of the user's behavior. This allows intent understanding to no longer rely solely on the user's behavior itself, but rather to be embedded in the actual environmental context. Based on the interaction event and environmental context data, a final contextual intent candidate set representing the user's interaction intent is generated. Therefore, by fusing user-initiated behavior with real-time environmental state, ambiguity caused by vague or ambiguous behavior expressions (such as a pointing action potentially corresponding to multiple objects) is effectively resolved, improving the semantic accuracy of intent candidates. Based on the final contextual intent candidate set, the robot's current task state and executable actions are combined to match the interaction intent, and the human-computer interaction intent recognition result is obtained. This ensures that the recognized intent is not only semantically reasonable, but also executable within the robot's capabilities. It avoids invalid responses caused by intents deviating from the actual state of the system, solves the problem of low intent recognition accuracy caused by relying on a single modality perception in complex dynamic environments, and improves the naturalness and accuracy of human-computer interaction. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating the robot human-computer interaction intent recognition method based on Internet of Things collaboration provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of the robot human-computer interaction intent recognition system based on Internet of Things collaboration provided in an embodiment of the present invention; Figure 3 An embodiment diagram of the electronic device provided in this invention; Figure 4 An embodiment diagram of a computer-readable storage medium provided in accordance with the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] Optionally, see Figure 1 , Figure 1This is a flowchart illustrating the robot human-computer interaction intent recognition method based on IoT collaboration provided by the present invention. In this embodiment of the invention, the execution entity of the robot human-computer interaction intent recognition method based on IoT collaboration is the intent recognition system. Therefore, the robot human-computer interaction intent recognition method based on IoT collaboration includes: Step 10: Collect multimodal raw sensing signals from users during the interaction process based on IoT sensing terminals deployed in the physical space to obtain multimodal sensing sequences, and identify segments containing user-initiated interactive behaviors based on the multimodal sensing sequences to obtain target sensing segments.

[0014] Optionally, the intent recognition system relies on IoT sensing terminals deployed in the physical space to collect multimodal raw sensing signals generated by users during their interaction with the environment or robots.

[0015] Physical space refers to the physical area where users actually engage in activities, including but not limited to family living rooms, office areas, and industrial workshops.

[0016] IoT sensing terminals refer to hardware devices that have signal acquisition capabilities and are connected to the Internet of Things. Their deployment method is to install them in a fixed location or move them around in key locations where users may interact, based on the interaction scenario requirements of the physical space. For example, they can be deployed next to the sofa in the living room, near the desk in the office area, or around the workstation in the workshop.

[0017] Multimodal raw sensory signals refer to unprocessed raw signals from different sensory dimensions, including visual sensory signals, auditory sensory signals, tactile sensory signals, and motion sensory signals. Visual sensory signals are images or videos that reflect a user's appearance features and posture. Auditory sensory signals are audio signals such as voices and sounds produced by the user's actions. Tactile sensory signals are signals such as pressure and vibration generated when a user touches a specific object. Motion sensory signals are signals such as displacement and speed generated by the user's body movement and limb movements.

[0018] Furthermore, the intent recognition system synchronously processes the acquired multimodal raw sensing signals, aligning the raw sensing signals of different modalities in chronological order to obtain a multimodal sensing sequence.

[0019] Among them, the multimodal sensing sequence refers to a set of signals arranged in chronological order, consisting of aligned sensing signals of each modality. The sampling interval in the time dimension is determined according to the hardware performance of the IoT sensing terminal and the real-time requirements of the interaction scenario. For example, the sampling interval can be set to 10 milliseconds to 50 milliseconds to ensure that the signal changes during the user interaction process can be fully captured.

[0020] Furthermore, the intent recognition system performs segmentation and interactive behavior recognition on the multimodal perception sequence to filter out segments containing user-initiated interactive behaviors, thus obtaining the target perception segment.

[0021] User-initiated interaction refers to actions that users consciously take to achieve specific needs, aiming to elicit environmental feedback or interact with the robot, such as waving their hand, issuing voice commands, touching the interactive panel, or walking towards a specific device. This is different from unconscious natural actions (such as raising a hand at random or touching something unintentionally).

[0022] Segment segmentation refers to dividing a multimodal sensing sequence into multiple continuous signal segments according to a preset time length or signal change nodes. The preset time length can be set according to the duration of common interactive behaviors, such as a time window of 1 to 3 seconds. Signal change nodes refer to the moments when the amplitude, frequency, and other characteristics of the sensing signal undergo significant changes.

[0023] Interactive behavior recognition refers to the intention recognition system extracting features from each segmented signal fragment to determine whether the fragment contains user-initiated interactive behavior. The extracted features include amplitude changes, frequency distribution, and temporal characteristics of each modal signal. If the features of a certain signal fragment meet the preset threshold for active interactive behavior features, then the fragment is determined to contain user-initiated interactive behavior, i.e., the target perception fragment.

[0024] In one embodiment, the IoT sensing terminals deployed in the physical space of a family living room include: a high-definition camera installed in the center of the living room ceiling (for collecting visual sensing signals), a sound collector placed on the side of the sofa (for collecting auditory sensing signals), a pressure sensor embedded in the surface of the smart coffee table in the living room (for collecting tactile sensing signals), and an infrared displacement sensor installed at the entrance of the living room (for collecting motion sensing signals).

[0025] When a user enters the living room, an infrared displacement sensor collects motion sensing signals (displacement change signals) generated by the user's movement. The user then walks to the smart coffee table and touches its surface; a pressure sensor collects tactile sensing signals (pressure change signals) generated by this touch. Simultaneously, a high-definition camera captures visual sensing signals (continuous image frames) of the user raising and touching the table. No significant audio signal is detected by the sound sensor. The intent recognition system synchronizes and aligns these motion sensing signals, tactile sensing signals, and visual sensing signals in chronological order, setting a sampling interval of 20 milliseconds to obtain a multimodal perception sequence.

[0026] The intent recognition system segments the multimodal perception sequence into multiple signal segments based on a preset time length of 2 seconds. For a signal segment containing the user's action of touching the coffee table, the system extracts the limb movement features of the visual perception signal (changes in posture such as raising and touching the hand), the pressure features of the tactile perception signal (pressure value increasing from 0 to a preset threshold), and the displacement features of the motion perception signal (changes in displacement as the user moves from the entrance to the coffee table). If these features are determined to meet the preset threshold for active interaction behavior features, the signal segment is identified as the target perception segment.

[0027] Step 20: Using the target perception segment as an interaction event, the IoT environmental sensing unit within the spatial location area indicated by the interaction event collects the current environmental state information to obtain environmental context data.

[0028] Optionally, the intent recognition system treats the target perception segment as an interaction event and clarifies the spatial location area indicated by the interaction event. Here, an interaction event refers to the behavioral event represented by the signal segment corresponding to the user's proactive interaction. The spatial location area indicated by the interaction event refers to the specific physical space area where the user is located when the proactive interaction occurs. This area is determined based on the deployment location of the IoT sensing terminal that collects the multimodal raw sensing signals corresponding to the target perception segment, and the spatial location features contained in the multimodal sensing signals (such as environmental references in visual sensing signals, the endpoint of displacement trajectories in motion sensing signals, etc.). The extent of this area can be set according to the influence range of the interaction and the actual scenario requirements, such as a circular area with a radius of 0.5 meters to 2 meters centered on the point where the user's interaction occurs.

[0029] Furthermore, the intent recognition system invokes IoT environmental sensing units deployed within the spatial location area to collect current environmental status information, thereby obtaining environmental context data.

[0030] Among them, the Internet of Things (IoT) environmental sensing unit refers to IoT devices deployed in physical space for collecting environmental status parameters. Its type is determined according to the needs of collecting environmental status information, including but not limited to temperature sensors, humidity sensors, light sensors, air quality sensors, sound intensity sensors, distance sensors, etc. The device maintains a communication connection with the intent recognition system and can respond to data collection commands in real time.

[0031] Current environmental state information refers to the various environmental physical parameters and environmental scene characteristics within the spatial location area at the same time node or time window where the interactive event occurs; environmental context data refers to the collection of current environmental state information collected by the IoT environmental sensing unit and preliminarily sorted (including noise reduction and format standardization), which is used to characterize the environmental background when the interactive event occurs.

[0032] In one embodiment, the intent recognition system uses the identified target perception segment (the signal segment of the user touching the smart coffee table) as an interaction event. Based on the deployment location of the pressure sensor (embedded on the surface of the smart coffee table) that collects the target perception segment, and the environmental reference object (smart coffee table) in the visual perception signal, the system determines that the spatial location area indicated by the interaction event is a circular area with a radius of 1 meter centered on the smart coffee table.

[0033] The IoT environmental sensing units deployed within the spatial location area include: temperature and humidity sensors installed on the side of the smart coffee table, light sensors installed on the ceiling above the coffee table, and air quality sensors placed inside the coffee table drawers.

[0034] The intent recognition system sends a data acquisition command to the IoT environmental sensing unit to collect current environmental status information synchronized with the time window of the interaction event (i.e., the 2-second time window when the user touches the coffee table). Specifically, this includes: the temperature in the area collected by the temperature sensor is 25 degrees Celsius, the humidity in the area collected by the humidity sensor is 50%, the light intensity in the area collected by the light sensor is 300 lux, and the PM2.5 concentration in the area collected by the air quality sensor is 15 micrograms per cubic meter.

[0035] The intent recognition system performs noise reduction (removing random interference signals from the sensor-collected data) and format standardization (unifying parameters of different units into a preset format) on the environmental state information such as temperature, humidity, light intensity, and PM2.5 concentration collected above to obtain environmental context data. The environmental context data characterizes the environmental background state of the spatial location area when the user touches the smart coffee table.

[0036] Step 30: Based on the interaction events and environmental context data, generate a final contextual intent candidate set that represents the user's interaction intent.

[0037] Optionally, the intent recognition system generates a final contextual intent candidate set representing the user's interaction intent based on the interaction events and environmental context data, as described in steps 301 to 303.

[0038] The final context intent candidate set includes multiple candidate intents, each representing the user's interaction intent in the context of the environment.

[0039] Step 40: Based on the final contextual intent candidate set, the robot's current task state and executable actions are combined to perform interaction intent matching to obtain the human-computer interaction intent recognition result.

[0040] Optionally, the intent recognition system performs interaction intent matching based on each candidate intent in the final context intent candidate set, combined with the robot's current task state and executable actions, to obtain the human-computer interaction intent recognition result, as described in steps 404 to 404.

[0041] The embodiments of the present invention improve the naturalness and accuracy of human-computer interaction.

[0042] Optionally, the processes of steps 301 to 303 include: Step 301: Based on the timestamp information and spatial location information of the user's active interaction behavior contained in the interaction event, determine the spatiotemporal anchor point of the user interaction, and extract user behavior features from the multimodal perception sequence associated with the interaction event based on the spatiotemporal anchor point of the user interaction.

[0043] Optionally, the intent recognition system extracts the timestamp information and spatial location information of the user's active interaction behavior from the interaction event, and determines the spatiotemporal anchor point of the user interaction based on these two types of information.

[0044] The timestamp information refers to the time data corresponding to the start time, duration, and end time of the user's proactive interactive behavior. This data is synchronously recorded by the IoT sensing terminal that collects multimodal raw sensing signals, and the time base is consistent with the system time of the intent recognition system.

[0045] Spatial location information refers to the specific spatial coordinates or area range when a user actively initiates an interactive behavior. This information comes from the spatial features contained in the multimodal perception sequence corresponding to the interactive event (such as the positioning of environmental reference objects in visual perception signals, the endpoint of displacement trajectory in motion perception signals, and the deployment location association positioning of IoT sensing terminals, etc.).

[0046] User interaction spatiotemporal anchor points refer to the set of information that uniquely identifies the time and space intersection of a user's proactive interaction behavior by associating and fusing timestamp information with spatial location information.

[0047] Furthermore, the intent recognition system extracts user behavior features from the multimodal perception sequence associated with interaction events based on the spatiotemporal anchor points of user interaction.

[0048] Among them, the multimodal perception sequence associated with the interaction event refers to the complete multimodal perception sequence in step 10 that includes the user's active initiation of the interaction behavior (i.e., the original multimodal perception sequence that generates the target perception segment).

[0049] User behavior characteristics refer to various characteristic parameters that can reflect the essential attributes of a user's proactive interactive behavior, including but not limited to visual behavior characteristics (such as body posture, movement amplitude, movement frequency, facial expressions, etc.), auditory behavior characteristics (such as tone of voice commands, speech rate, semantic keywords, sound characteristics of actions, etc.), motor behavior characteristics (such as body movement speed, movement direction, limb joint movement angle, etc.), and tactile behavior characteristics (such as touch pressure, touch duration, touch frequency, etc.).

[0050] Optionally, the feature extraction process in this embodiment of the invention is as follows: the intent recognition system first selects signal segments in the corresponding spatiotemporal range of the multimodal perception sequence based on the spatiotemporal anchor points of user interaction, then performs submodal feature extraction on the signal segments, and finally fuses the features extracted from each modality to obtain a set of user behavior features of a unified dimension. During the extraction process, the temporal synchronization and spatial consistency of each modality feature must be ensured.

[0051] In the aforementioned example of a family living room scenario, the interaction event is "user touches the smart coffee table." The intent recognition system extracts timestamp information from this interaction event: the user's active touch action begins at 14:30:05, lasts for 2 seconds, and ends at 14:30:07. It also extracts spatial location information: based on the deployment location of the pressure sensor embedded in the smart coffee table (the center area of ​​the coffee table surface) and the image features of "user's hand touching the center of the coffee table surface" in the visual perception signal collected by the high-definition camera, the spatial location is determined to be "the center area of ​​the smart coffee table surface in the living room (coordinates are X=3.2 meters, Y=2.5 meters, Z=0.75 meters in the living room floor coordinate system)." The intent recognition system fuses the above timestamp information with the spatial location information to obtain the user's interaction spatiotemporal anchor point: from 14:30:05 to 14:30:07, the center area of ​​the smart coffee table surface in the living room (X=3.2 meters, Y=2.5 meters, Z=0.75 meters).

[0052] Based on this spatiotemporal anchor point of user interaction, the intent recognition system filters signal segments corresponding to the spatiotemporal range from the associated multimodal perception sequences, and then extracts user behavior features: visual behavior features are extracted from visual perception signals (the user's hand lifts from the side of the body to the coffee table, the lifting angle is 60 degrees, and the palm faces down when the hand touches the table); tactile behavior features are extracted from tactile perception signals (the peak touch pressure is 5 Newtons, the touch duration is 2 seconds, and the pressure remains stable without significant fluctuations); motion behavior features are extracted from motion perception signals (the user's hand movement speed is 0.3 meters per second, and the movement direction is vertically downward pointing towards the coffee table); no effective behavior features are extracted from auditory perception signals (no speech or obvious movement sounds). Finally, the system fuses the visual, tactile, and motion behavior features to obtain the user behavior feature set corresponding to this interaction event.

[0053] Step 302: Based on the action type, action duration and action direction in the user behavior features, obtain the semantic description information of user behavior, and predict the preliminary interaction tendency category expressed by the user during the interaction process based on the semantic description information of user behavior.

[0054] Optionally, the intent recognition system performs semantic parsing on user behavior features and generates semantic descriptions of user behavior based on the parsing results. Semantic parsing refers to the process of transforming abstract user behavior features into textual descriptions with clear semantic meanings that can be understood by the intent recognition system. The parsing is based on a pre-defined mapping rule base between behavior features and semantic descriptions. This rule base contains the correlation between the feature parameter ranges of various common user interaction behaviors and their corresponding semantic descriptions.

[0055] User behavior semantic description information refers to information that accurately describes the specific manifestations and core characteristics of user-initiated interactive behaviors. Its content must fully cover the key characteristics of user behavior. For example, "The user touches the center area of ​​the smart coffee table with a palm down for 2 seconds with a pressure of 5 Newtons, and moves his hand vertically downwards from the side of his body to the table at a speed of 0.3 meters per second."

[0056] Furthermore, the intent recognition system predicts the initial interaction tendency category expressed by the user during the interaction process based on the generated semantic description information of user behavior.

[0057] The preliminary interaction tendency category refers to the initial classification of user interaction intent, representing the direction of the user's interaction behavior, rather than the specific content of the intent. Examples include "device control tendency", "information query tendency", "help-seeking tendency", and "no clear tendency (misoperation)".

[0058] Optionally, the prediction process in this embodiment of the invention is as follows: the intent recognition system matches the semantic description information of user behavior with a preset rule base for determining interaction tendency categories. The rule base contains behavioral semantic features corresponding to different interaction tendency categories. If the core features in the semantic description information of user behavior meet the determination conditions of a certain interaction tendency category, then the category is determined as a preliminary interaction tendency category. If there are multiple categories that meet the conditions, then all categories that meet the conditions are taken as preliminary interaction tendency categories to form a set of preliminary interaction tendency categories.

[0059] In one embodiment, the intent recognition system matches the user behavior features extracted in step 301 with a preset behavior feature-semantic description mapping rule base to generate user behavior semantic description information: "From 14:30:05 to 14:30:07, the user, with palm down, continuously touches the center area of ​​the smart coffee table in the living room with a pressure of 5 Newtons for 2 seconds. The hand moves vertically downwards from the side of the body to the table at a speed of 0.3 meters per second, without accompanying voice or obvious movement sounds." Further, the intent recognition system matches this user behavior semantic description information with a preset interaction tendency category judgment rule base. The judgment condition corresponding to "device control tendency" in the rule base is "the user performs continuous touch operation on the smart device, the touch pressure is stable, there is no help-seeking voice, and there are no obvious erroneous operation characteristics (such as excessively short touch duration or drastic pressure fluctuations)." The current user behavior semantic description information fully meets this judgment condition, and there are no other interaction tendency categories that meet the conditions. Therefore, the intent recognition system predicts the initial interaction tendency category as "device control tendency."

[0060] Step 303: Based on the semantic description information of user behavior, the preliminary interaction tendency category and the environmental context data, determine the final context intent candidate set.

[0061] Optionally, the intent recognition system determines the final contextual intent candidate set based on the semantic description information of user behavior, the preliminary interaction tendency category and environmental context data, as in steps 3031 to 3035.

[0062] This invention combines user behavior semantic description information, preliminary interaction tendency categories and environmental context data to generate a final contextual intent candidate set, achieving deep integration of user behavior and environmental scene, ensuring that the generated candidate intent can closely fit the current interaction scene, avoiding misjudgment of intent detached from the environmental background, and improving the naturalness and accuracy of human-computer interaction.

[0063] Optionally, the processes of steps 3031 to 3035 include: Step 3031: Based on the physical object state information and spatial layout information in the environmental context data, determine the static semantic elements of the environment, and based on the static semantic elements of the environment, parse out the interactive physical objects and their corresponding interactive functional attributes in the current spatial area.

[0064] Optionally, the intent recognition system extracts physical object state information and spatial layout information from environmental context data, and determines environmental static semantic elements based on these two types of information.

[0065] Among them, physical object state information refers to the information in the environmental context data that characterizes the inherent state and existence state of various physical objects within the current spatial area. The inherent state includes fixed attributes such as the material, size, and type of the physical object, while the existence state includes variable attributes such as the on / off state and working mode state of the physical object. Spatial layout information refers to the information in the environmental context data that characterizes the relative positional relationships, distribution arrangement, and spatial structural characteristics of various physical objects within the current spatial area.

[0066] Environmental static semantic elements refer to the set of semantic information that can comprehensively characterize the static environmental features within the current spatial area after semantically fusing the physical object state information and spatial layout information. Static environmental features refer to environmental features that do not change rapidly over time.

[0067] Furthermore, the intent recognition system, based on static semantic elements of the environment, parses out interactive physical objects and their corresponding interactive functional attributes within the current spatial area.

[0068] Interactive physical objects refer to physical objects within the current spatial area that can respond to user interaction and perform corresponding functions, including but not limited to smart home appliances, smart interactive panels, and smart sensors; interactive functional attributes refer to the specific functional characteristics of interactive physical objects that can be triggered by user interaction, such as the lighting function, desktop temperature adjustment function, and network connection function of a smart coffee table.

[0069] Optionally, the parsing process in this embodiment of the invention is as follows: the intent recognition system first filters out physical objects with interactive response capabilities from the static semantic elements of the environment and determines them as interactive physical objects; For each interactive physical object, the corresponding functional feature information is extracted from the static semantic elements of the environment. Combined with the preset physical object-functional attribute mapping rule library, the interactive functional attributes corresponding to the physical object are parsed to form the association mapping relationship of "interactive physical object-interactive functional attribute".

[0070] In one embodiment, the environmental context data is "the current spatial area (a circular area with a radius of 1 meter centered on the smart coffee table) has a temperature of 25 degrees Celsius, humidity of 50%, light intensity of 300 lux, and PM2.5 concentration of 15 micrograms per cubic meter". Combined with the preset environmental information of this area, the physical object status information is extracted: the smart coffee table (made of wood, with dimensions of 120 cm * 60 cm * 75 cm, currently in standby mode), the ceiling light fixture (made of metal, currently in the on state), and the temperature sensor (made of plastic, currently in the working state). The spatial layout information is extracted: the smart coffee table is located in the center of the area, the ceiling light fixture is located 1.8 meters directly above the smart coffee table, and the temperature sensor is embedded in the side of the smart coffee table. The intent recognition system integrates the above information semantics to obtain the static semantic elements of the environment: "The current space area is the area around the smart coffee table in the living room. The center of the area is a wooden smart coffee table in standby mode (120 cm * 60 cm * 75 cm). There is a metal lighting fixture in the on state 1.8 meters directly above it, and a plastic temperature sensor in working state is embedded on the side."

[0071] Based on the static semantic elements of the environment, the intent recognition system analyzes the interactive physical objects and interactive functional attributes: from the static semantic elements of the environment, the physical objects with interactive response capabilities are selected as smart coffee tables and ceiling lighting fixtures.

[0072] Furthermore, the intent recognition system, combined with a pre-set physical object-functional attribute mapping rule base, parses the interactive functional attributes corresponding to the smart coffee table as "turn on desktop lighting, adjust desktop temperature, start wireless charging, connect to Bluetooth", and the interactive functional attributes corresponding to the ceiling light fixture as "adjust brightness, switch lighting color, turn off lighting", forming an associated mapping relationship: smart coffee table - [turn on desktop lighting, adjust desktop temperature, start wireless charging, connect to Bluetooth], ceiling light fixture - [adjust brightness, switch lighting color, turn off lighting].

[0073] Step 3032: Based on the dynamic environmental parameters in the environmental context data, identify the environmental events occurring in the current spatial area and their impact on the environmental events generated by user interaction.

[0074] Optionally, the intent recognition system extracts dynamic environmental parameters from the environmental context data and identifies environmental events occurring within the current spatial area based on these dynamic environmental parameters.

[0075] Dynamic environmental parameters refer to environmental parameters in the environmental context data that characterize changes over time within the current spatial region, including but not limited to the rate of change of temperature, humidity, light intensity, sound intensity, and air quality parameters. Environmental events refer to events representing changes in the environmental state occurring within the current spatial region, as characterized by changes in dynamic environmental parameters, such as rapid temperature increases, sudden drops in light intensity, and sudden increases in sound intensity.

[0076] Optionally, the recognition process in this embodiment of the invention is as follows: the intent recognition system first performs trend analysis on the extracted dynamic environmental parameters and calculates the amount and rate of change of the dynamic environmental parameters within a preset time window; then, the calculation results are matched with a preset environmental event recognition threshold library. If the amount or rate of change of a certain dynamic environmental parameter exceeds the corresponding recognition threshold, it is determined that an environmental event corresponding to that parameter is occurring in the current spatial area; if the changes of multiple dynamic environmental parameters all exceed the corresponding thresholds, multiple environmental events that are occurring are identified, forming an environmental event set.

[0077] Furthermore, the intent recognition system analyzes the impact of identified environmental events on user interactions. This impact refers to the interaction between environmental events and user interaction behavior, specifically whether environmental events will prompt user interaction, whether they will affect the user's target selection, and whether they will change the way user interaction behavior is executed.

[0078] Optionally, the analysis process in this embodiment of the invention is as follows: The intent recognition system, based on a preset environmental event-interaction influence mapping rule base, and combined with the characteristics of the current user interaction behavior (extracted from the semantic description information of user behavior), determines the specific influence of each environmental event on the user interaction. For example, a "rapid temperature rise event" may prompt the user to perform an interaction behavior to adjust the temperature, that is, the environmental event and the user interaction behavior have a "prompting" influence relationship; if the environmental event has no obvious influence on the user interaction, it is determined to be a "no influence" relationship, and finally a "environmental event-user interaction" influence relationship mapping is formed.

[0079] In one embodiment, the intent recognition system extracts dynamic environmental parameters from environmental context data: the light intensity in the current spatial area drops from 300 lux to 200 lux within 10 seconds, a change rate of 10 lux / second; the changes in temperature, humidity, and PM2.5 concentration within a preset 10-second time window are all less than preset thresholds. The intent recognition system calculates the light intensity change rate as 10 lux / second and matches it with a preset environmental event recognition threshold library (the light intensity change rate threshold is 5 lux / second). Since this change rate exceeds the threshold, the environmental event occurring in the current spatial area is identified as a "sudden drop in light intensity event". Combining the interactive behavior characteristics of "user touching the smart coffee table" in the semantic description information of user behavior, the influence relationship is analyzed based on the preset environmental event-interaction influence mapping rule base: the interaction influence corresponding to the "sudden drop in light intensity event" in the rule base is "may prompt the user to perform interactive behavior related to adjusting light". The object of the current user interaction behavior is the smart coffee table, and the smart coffee table has the functional attribute of being associated with lighting adjustment. Therefore, it is determined that there is an environmental event influence relationship between the "sudden drop in light intensity event" and the user's current interaction behavior, which "prompts the user to select an interactive target related to light adjustment".

[0080] Step 3033: Based on the semantic description information of user behavior and the static semantic elements of the environment, establish the spatial proximity and functional matching association between user actions and interactive physical objects.

[0081] Optionally, the intent recognition system establishes the association between user actions and interactive physical objects based on user behavior semantic description information and static environmental semantic elements, from two dimensions: spatial proximity and functional matching. Here, user actions refer to the specific interactive actions represented in the user behavior semantic description information, such as touching, waving, or speaking; spatial proximity refers to the closeness between the spatial location where the user action occurs and the spatial location of the interactive physical object; and functional matching refers to the degree of compatibility between the characteristics of the user action and the interactive functional attributes of the interactive physical object.

[0082] Optionally, the spatial proximity association establishment process is as follows: the intent recognition system extracts the specific spatial location of the user action from the semantic description information of the user behavior, and extracts the spatial location of each interactive physical object from the static semantic elements of the environment; calculates the straight-line distance between the location of the user action and the location of each interactive physical object; compares the calculated straight-line distance with a preset spatial proximity threshold, and if the straight-line distance is less than or equal to the preset threshold, it is determined that the user action and the interactive physical object have a spatial proximity association; the preset spatial proximity threshold is set according to the needs of different interaction scenarios, for example, the spatial proximity threshold for touch actions is 0.5 meters, and the spatial proximity threshold for waving actions is 1 meter.

[0083] Optionally, the functional matching association establishment process is as follows: The intent recognition system extracts the features of user actions from the semantic description information of user behavior, including action type, action amplitude, action duration, etc.; from the "interactive physical object-interactive functional attribute" association mapping relationship obtained in step 3031, the system extracts the trigger action features corresponding to the interactive functional attributes of each interactive physical object; the user action features are matched with the trigger action features of each interactive physical object, and if the matching degree between the two exceeds the preset functional matching threshold, it is determined that the user action and the interactive physical object have a functional matching association; the preset functional matching threshold is set according to the similarity requirements between the action features and the trigger action features, for example, a matching degree of 80% or above is considered to indicate the existence of a functional matching association.

[0084] In one embodiment, the user behavior semantic description information is: "From 14:30:05 to 14:30:07, the user, with palm down, continuously touches the center area of ​​the smart coffee table in the living room with a pressure of 5 Newtons for 2 seconds. The hand moves vertically downwards from the side of the body to the table at a speed of 0.3 meters per second, without accompanying voice or obvious movement sounds." From this, the user action is extracted as: a continuous touch action, the action occurring at the center area of ​​the smart coffee table in the living room (coordinates X=3.2 meters, Y=2.5 meters, Z=0.75 meters), and the action characteristics are "palm down, lasting 2 seconds, pressure 5 Newtons." The spatial locations of the interactive physical objects in the environmental static semantic elements are: the smart coffee table (X=3.2 meters, Y=2.5 meters, Z=0.75 meters) and the ceiling light fixture (X=3.2 meters, Y=2.5 meters, Z=2.55 meters).

[0085] Establish spatial proximity association: The straight-line distance between the user's action location and the smart coffee table is calculated to be 0 meters, and the straight-line distance between the user's action location and the ceiling light fixture is 1.8 meters; the preset spatial proximity threshold for touch-type actions is 0.5 meters. Therefore, it is determined that the user's action has a spatial proximity association with the smart coffee table, but no spatial proximity association with the ceiling light fixture.

[0086] Establishing a functional matching association: Extracting user action features "palm down, duration 2 seconds, pressure 5 Newtons", the trigger action features of each interactive physical object are as follows: the trigger action features corresponding to the interactive functional attributes of the smart coffee table are "touch, duration 1 second or more, pressure 3 Newtons or more", and the trigger action features of the ceiling lighting fixture are "wave, voice command". Matching the user action features with the trigger action features of the two objects, the matching degree with the smart coffee table is 90% (exceeding the 80% functional matching threshold), and the matching degree with the ceiling lighting fixture is 0%. Therefore, it is determined that there is a functional matching association between the user action and the smart coffee table. Finally, a dual association of spatial proximity and functional matching between the user action and the smart coffee table is established.

[0087] Step 3034: Based on the preliminary interaction tendency category and environmental dynamic semantic elements, analyze whether the user's interaction behavior responds to the current environmental event, and generate the first-level interaction intent.

[0088] Optionally, the intent recognition system analyzes whether the user's interaction behavior responds to current environmental events based on the initial interaction tendency category and dynamic semantic elements of the environment, and then generates the first-level interaction intent.

[0089] Among them, environmental dynamic semantic elements refer to the set of semantic information containing the environmental events occurring in the current spatial region and the characteristics of dynamic environmental parameter changes, obtained by fusing environmental events with extracted dynamic environmental parameters. Responding to the current environmental event means that the goal of the user's interaction behavior matches the needs of the environmental state change represented by the current environmental event; that is, the user's interaction behavior is to cope with or resolve the impact of the current environmental event. First-level interaction intent refers to the user interaction intent with preliminary direction obtained based on environmental event response analysis, which provides guidance for subsequently determining the final contextual intent candidate set.

[0090] Optionally, the analysis and generation process of this embodiment of the invention is as follows: the intent recognition system first clarifies the intent direction corresponding to the preliminary interaction tendency category, and combines the environmental events in the environmental dynamic semantic elements to determine the potential user interaction needs corresponding to the environmental events; then, the features of the user interaction behavior (extracted from the semantic description information of user behavior) are compared with the interaction behavior features corresponding to the potential interaction needs to determine whether the user interaction behavior matches the potential interaction needs.

[0091] If a match is found, the user's interaction behavior is determined to be in response to the current environmental event. A first-level interaction intent is generated based on the initial interaction tendency category and the environmental event. For example, if the initial interaction tendency category is "device control tendency" and the environmental event is "sudden drop in light intensity event", then the first-level interaction intent is "control the device to adjust the light". If a mismatch is found, the user's interaction behavior is determined to be in response to the current environmental event. A first-level interaction intent is generated only based on the initial interaction tendency category. For example, "control the device to perform non-environmental response functions".

[0092] In one embodiment, the initial interaction tendency category is "device control tendency," and the environmental dynamic semantic element is "a sudden drop in light intensity is occurring in the current spatial area, with the light intensity decreasing from 300 lux to 200 lux within 10 seconds, at a rate of 10 lux / second." The potential user interaction need corresponding to this environmental event is "adjusting the light intensity to increase the brightness of the space." The user interaction behavior feature extracted from the user behavior semantic description information is "touching the smart coffee table, without voice commands, with continuous and stable movements."

[0093] The intent recognition system compares the user's interactive behavior characteristics with the interactive behavior characteristics corresponding to the potential interactive need of "adjusting the light intensity" (such as touching a device with light adjustment function, issuing a voice command to adjust the light intensity, etc.), and the two match. Combining the initial interaction tendency category "device control tendency" and the environmental event "sudden drop in light intensity event", the system determines that the user's interactive behavior is responding to the current environmental event, and thus generates the first-level interaction intent as "control the smart coffee table to perform the light adjustment function".

[0094] Step 3035: Based on spatial proximity, functional matching association, environmental event influence relationship, first-level interaction intent and interactive function attributes, determine the final context intent candidate set.

[0095] Optionally, the intent recognition system determines the final context intent candidate set based on spatial proximity, functional matching association, environmental event influence relationship, first-level interaction intent, and interactive functional attributes, as in steps 30351 to 30353.

[0096] The embodiments of the present invention achieve a comprehensive integration of user behavior features, static environment features, dynamic environment features and initial intent direction, ensuring that the generated candidate intents are accurately matched with the current interaction scenario, effectively avoiding the problem of misjudgment of intent caused by being out of context, and improving the naturalness and accuracy of human-computer interaction.

[0097] Optionally, the processes of steps 30351 to 30353 include: Step 30351: Generate the second-level interactive intent based on spatial proximity and functional matching association.

[0098] Optionally, the intent recognition system generates a second-level interactive intent based on spatial proximity and functional matching association. The second-level interactive intent refers to an interactive intent with a clear object and function orientation, formed by combining the relationship between the user's action and the interactive physical object with the interactive functional attributes of the interactive physical object. Spatial proximity refers to the closeness between the spatial location where the user's action occurs and the spatial location of the interactive physical object, while functional matching association refers to the adaptation relationship between the user's action features and the trigger action features corresponding to the interactive functional attributes of the interactive physical object. Both together constitute the core basis for generating the second-level interactive intent.

[0099] Optionally, the generation process in this embodiment of the invention is as follows: The intent recognition system determines interactive physical objects that have a dual association with the user's action (i.e., both spatial proximity and functional matching association exist simultaneously) from the spatial proximity and functional matching association results, and uses these as target interactive physical objects; if there are multiple target interactive physical objects, a set of target interactive physical objects is formed. Subsequently, the interactive functional attributes corresponding to each target interactive physical object are extracted, and the degree of adaptation between the user's action features and the interactive functional attributes is combined to match the corresponding interactive functional attributes for each target interactive physical object. Finally, based on the correspondence between "user-target interactive physical object-interactive functional attribute", a second-level interactive intent is generated. Each second-level interactive intent clearly includes the interactive object and the function to be triggered, such as "controlling the smart coffee table to turn on the desktop lighting function" or "controlling the smart speaker to play music function". If the same target interactive physical object has multiple adapted interactive functional attributes, multiple second-level interactive intents corresponding to different functions are generated, forming a set of second-level interactive intents.

[0100] In one embodiment, the spatial proximity and functional matching association results obtained in step 3033 are as follows: the user action (continuous touch action) has a dual association with the smart coffee table, but no association with the ceiling lighting fixture. Therefore, the target interactive physical object is the smart coffee table. The interactive functional attributes of the smart coffee table parsed in step 3031 are "turn on desktop lighting, adjust desktop temperature, start wireless charging, and connect to Bluetooth".

[0101] The intent recognition system extracts the aforementioned interactive functional attributes and matches them with the user's action features (palm down, duration 2 seconds, pressure 5 Newtons) and the degree of fit between these features and the trigger action features of each functional attribute. All interactive functional attributes of the smart coffee table are triggered by the following actions: touch, duration of 1 second or more, and pressure of 3 Newtons or more. The user's action characteristics match the trigger action characteristics of these functional attributes by 90%. Therefore, based on the correspondence between "user-smart coffee table-interactive functional attributes", a second-level set of interactive intents is generated: "control the smart coffee table to turn on the desktop lighting function", "control the smart coffee table to adjust the desktop temperature function", "control the smart coffee table to start the wireless charging function", and "control the smart coffee table to connect to Bluetooth".

[0102] Step 30352: Perform cross-validation of the first-level and second-level interaction intents to exclude interaction intents that are contradictory or logically conflicting, and obtain a preliminary candidate set of contextual intents.

[0103] Optionally, the intent recognition system performs cross-validation of first-level and second-level interaction intents to eliminate conflicting or logically contradictory intents, thus obtaining a preliminary candidate set of contextual intents. Cross-validation refers to verifying whether there is logical consistency between the first-level and second-level interaction intents, i.e., whether the direction of the second-level interaction intent matches the direction of the first-level interaction intent. Conflicts or logical contradictions refer to situations where the core elements of the two intents, such as intent direction, scope of interaction objects, and functional type, are mutually opposed or incompatible. For example, if the first-level interaction intent is "control the device to adjust the lighting," while the second-level interaction intent is "control the device to adjust the temperature," then there is a logical conflict in their functional types.

[0104] Optionally, the verification and screening process in this embodiment of the invention is as follows: The intent recognition system identifies the core elements of the first-level interaction intent, including intent direction (such as light adjustment, temperature adjustment, device control, etc.), associated environmental events (such as a sudden drop in light intensity, a rise in temperature, etc.), and the range of potential interaction objects (such as devices with light adjustment functions, devices with temperature adjustment functions, etc.). Then, the logical consistency between the core elements (interaction object, function type) of each second-level interaction intent and the core elements of the first-level interaction intent is analyzed one by one: if the core elements of both are completely consistent or without conflict, the second-level interaction intent is determined to have passed cross-validation; if the core elements of both are contradictory or logically conflicting, the second-level interaction intent is determined to have failed cross-validation and is excluded. Finally, all second-level interaction intents that have passed cross-validation are integrated to form a preliminary context intent candidate set.

[0105] In one embodiment, the first-level interaction intent is "control the smart coffee table to perform the lighting adjustment function," and its elements are: the intent direction is "lighting adjustment," the associated environmental event is "sudden drop in light intensity event," the interaction object is "smart coffee table," and the function type is "lighting adjustment related function." The second-level interaction intent set includes: "control the smart coffee table to turn on the desktop lighting function," "control the smart coffee table to adjust the desktop temperature function," "control the smart coffee table to start the wireless charging function," and "control the smart coffee table to connect to Bluetooth."

[0106] The intent recognition system underwent cross-validation: The logical consistency between each second-level interaction intent and the first-level interaction intent was analyzed one by one. Specifically, the function type of "controlling the smart coffee table to turn on the desktop lighting" is related to light adjustment, perfectly matching the core elements of the first-level interaction intent, and thus passed validation. The function type of "controlling the smart coffee table to adjust the desktop temperature" is temperature adjustment, which logically conflicts with the "light adjustment" direction of the first-level interaction intent, and therefore failed validation. The functions of "controlling the smart coffee table to start the wireless charging function" and "controlling the smart coffee table to connect to Bluetooth" are both unrelated to light adjustment and have no connection to the associated "sudden drop in light intensity event," thus exhibiting logical conflicts and failing validation. After excluding the failed interaction intents, the validated "controlling the smart coffee table to turn on the desktop lighting" was integrated to obtain a preliminary contextual intent candidate set: {controlling the smart coffee table to turn on the desktop lighting}.

[0107] Step 30353: Based on the user behavior semantic description information, environmental event influence relationship and interactive function attributes corresponding to each interaction intent in the preliminary context intent candidate set, semantic completion is performed on each interaction intent to obtain the final context intent candidate set.

[0108] Optionally, the intent recognition system performs semantic completion on each interactive intent based on the user behavior semantic description information, environmental event influence relationship and interactive functional attributes corresponding to each interactive intent in the preliminary context intent candidate set, to obtain the final context intent candidate set.

[0109] Optionally, the completion process in this embodiment of the invention is as follows: For each interaction intent in the preliminary context intent candidate set, the intent recognition system first extracts key elements from the user behavior semantic description information corresponding to the interaction intent (such as interaction action type, action duration, occurrence time, occurrence location, etc.), key elements from the environmental event influence relationship (such as associated environmental events, the way events affect the interaction, etc.), and key elements from the interactive function attributes (such as the specific function's purpose, the effect after the function is triggered, etc.). Then, these key elements are added to the original interaction intent, so that the completed interaction intent clearly covers the core dimensions such as "user-interaction action-interaction object-interaction function-interaction background (environmental event)". Finally, all completed interaction intents are integrated to form the final context intent candidate set.

[0110] Through multi-dimensional verification and optimization, the embodiments of the present invention generate a final context intent candidate set that is accurate, complete, and adaptable to different scenarios, thereby improving the naturalness and accuracy of human-computer interaction.

[0111] Optionally, the processes of steps 401 to 404 include: Step 401: Based on the user behavior features associated with each candidate intent in the final context intent candidate set, determine the user's limb movement trajectory sequence, and based on the user's limb movement trajectory sequence, construct a set of dynamic occupied areas of the user in the physical space.

[0112] Optionally, the intent recognition system first extracts the user behavior features associated with each candidate intent from the final context intent candidate set. These user behavior features refer to the core feature parameters corresponding to the candidate intent that reflect the user's interactive actions, including the user's body posture, action amplitude, action trajectory, and spatiotemporal information of the action. The user behavior features originate from the set of user behavior features extracted and associated with the candidate intents in step 301.

[0113] Furthermore, based on user behavior characteristics, the intent recognition system determines the user's limb movement trajectory sequence. This sequence refers to the chronological arrangement of the coordinates of the user's participating limb parts (such as hands and arms) in physical space as they change over time during the interaction. The determination process is as follows: the intent recognition system extracts key time points of limb movements and their corresponding limb position coordinates from the user's behavior characteristics. The selection interval for key time points is determined based on the duration and frequency of the user's actions. For example, if the action lasts for 2 seconds and has a high frequency of change, the selection interval is set to 100 milliseconds, resulting in 21 key time points. The limb position coordinates corresponding to each key time point are then arranged chronologically to form the user's limb movement trajectory sequence.

[0114] Furthermore, the intent recognition system constructs a set of dynamically occupied regions in physical space based on the user's limb movement trajectory sequence. The dynamically occupied region refers to the physical space range occupied by the user's limb at a specific time point. The delineation of this range needs to consider the actual size of the limb and the error redundancy during the movement process. For example, using the limb's position coordinates as the center, an error redundancy range of 5 cm is extended outwards according to the actual size of the limb (e.g., a hand size of 15 cm * 8 cm * 5 cm), forming a cubic or irregular polyhedral spatial region. The set of dynamically occupied regions refers to the collection of user dynamically occupied regions corresponding to different time points in chronological order. Each region is associated with a corresponding timestamp to characterize the spatial occupancy status of the user's limb at different times.

[0115] Step 402: Based on the task path planning information and the robotic arm's working envelope in the current task state, determine the robot's current operating space region. The robot's current operating space region includes the task movement channel, the reachable range of the end effector, and the preset safety buffer boundary.

[0116] Optionally, the intent recognition system determines the robot's current operating space region based on the task path planning information in the robot's current task state and the working envelope of the robotic arm. The robot's current task state refers to the task or working state the robot is executing at the moment of interaction intent matching, including task type, task progress, task path planning information, and robotic arm working parameters. The task path planning information refers to the robot's pre-set movement route information to complete the current task, including the starting point, ending point, nodes along the way, direction, and speed. The robotic arm working envelope refers to the closed surface formed by all spatial points reachable by the end effector during the robot's movement, determined by structural parameters such as the joint range of motion and link length. The robot's current operating space region refers to the entire physical space range that the robot needs to occupy or may reach during the execution of the current task; this region is the core space range for ensuring the robot's safe task execution and enabling human-robot interaction.

[0117] The robot's current operating space specifically includes the task movement channel, the reachable range of the end effector, and the preset safety buffer boundary. The task movement channel refers to the space occupied by the robot's body and moving parts when moving along the planned task path. Its width is determined by the robot's dimensions, typically the maximum width of the robot plus a safety margin of 50 to 100 centimeters on each side. The reachable range of the end effector refers to the spatial area covered by the working envelope of the robotic arm, i.e., the area consisting of all spatial points that the end effector can reach. The preset safety buffer boundary is a buffer area set outside the task movement channel and the reachable range of the end effector to prevent collisions between the robot and the surrounding environment or users. Its width is determined based on the robot's movement speed and the safety requirements of the interaction scenario, typically 30 to 50 centimeters.

[0118] Optionally, the determination process in this embodiment of the invention is as follows: extracting task path planning information and robotic arm working parameters from the robot's current task state, and calculating the robotic arm working envelope based on the robotic arm working parameters; secondly, defining the task movement channel according to the task path planning information and the robot's body size; then determining the reachable range of the end effector according to the robotic arm working envelope; finally, defining a preset safety buffer boundary outside the task movement channel and the reachable range of the end effector, and integrating the task movement channel, the reachable range of the end effector, and the preset safety buffer boundary to obtain the robot's current operating space area.

[0119] Step 403: Based on the user's dynamically occupied area set and the robot's current operating space area under the condition of time alignment, perform spatial intersection calculation to obtain the spatial intersection result.

[0120] Optionally, the intent recognition system performs spatial intersection calculation based on the user's dynamically occupied area set and the robot's current operating space area obtained in step 402, under time alignment conditions, to obtain the spatial intersection result. Here, the time alignment condition refers to matching the timestamp corresponding to each area in the user's dynamically occupied area set with the time dimension of the robot's current operating space area, ensuring that the spatial intersection calculation is performed within the same time range; the spatial intersection calculation refers to determining whether there is an overlapping spatial portion between the user's dynamically occupied area and the robot's current operating space area at the same time node. If so, the spatial range of the overlapping portion is calculated; otherwise, the intersection is empty; the spatial intersection result is a set composed of the spatial intersection judgment results at each time node and the overlapping range (if any).

[0121] Optionally, the calculation process in this embodiment of the invention is as follows: First, determine the time range of the user's dynamically occupied area set and the task execution time range corresponding to the robot's current operating space area, and select the overlapping time range of the two as the calculation time window; second, extract each area in the user's dynamically occupied area set that is within the calculation time window according to the timestamp, and at the same time determine the robot's current operating space area corresponding to each timestamp (if the robot is in a moving state, the robot position corresponding to the timestamp needs to be determined according to the task path planning information, and then the spatial coordinates of the robot's current operating space area are adjusted); finally, compare the spatial range of the user's dynamically occupied area and the robot's current operating space area corresponding to each timestamp to determine whether there is an overlap. If there is, calculate the spatial coordinate range of the overlapping area. After sorting the judgment result of each timestamp and the overlapping range (if present), the spatial intersection result is obtained.

[0122] Step 404: Based on the spatial intersection result, intent matching is performed to obtain the human-computer interaction intent recognition result.

[0123] Optionally, the intent recognition system performs interactive intent matching based on the spatial intersection results to obtain human-computer interaction intent recognition results, as described in steps 4041 to 4044.

[0124] The embodiments of the present invention improve the accuracy of human-computer interaction intent recognition through deep fusion and adaptation verification in the spatiotemporal dimensions, thereby enhancing the naturalness and accuracy of human-computer interaction.

[0125] Optionally, the processes of steps 4041 to 4044 include: Step 4041: Based on the spatial intersection result, determine whether there are regions with temporal overlap and spatial intersection, obtain a set of potential disturbance events, and determine whether the duration of each disturbance event in the set of potential disturbance events exceeds the preset minimum effective disturbance threshold based on the spatial intersection volume, and obtain the disturbance judgment result.

[0126] Optionally, the intent recognition system determines whether there are temporally overlapping and spatially intersecting regions based on the spatial intersection results, thus obtaining a set of potential disturbance events. Temporal overlap refers to the timestamp corresponding to the user's dynamically occupied area and the timestamp corresponding to the robot's current operating space area being in the same time interval, with the smallest unit of this time interval consistent with the time interval of the user's dynamically occupied area, such as 100 milliseconds. Spatially intersecting regions refer to the physical spatial portion where the user's dynamically occupied area and the robot's current operating space area overlap at the same timestamp. Potential disturbance events refer to events that may interfere with or affect the robot's current task execution due to the temporal overlap and spatial intersection between the user's dynamically occupied area and the robot's current operating space area. The set of potential disturbance events is a collection of all events that meet the "temporally overlapping and spatially intersecting" condition, with each potential disturbance event associated with a corresponding occurrence time window (i.e., the time period of temporal overlap) and spatial overlap range (i.e., the range of the spatially intersecting area).

[0127] Furthermore, the intent recognition system determines whether the duration and spatial intersection volume of each potential disturbance event in the set of potential disturbance events exceed a preset minimum effective disturbance threshold, thus obtaining a disturbance judgment result. The duration refers to the length of the time overlap period corresponding to each potential disturbance event, calculated by subtracting the start timestamp from the end timestamp of the event. The spatial intersection volume refers to the size of the spatial intersection region corresponding to each potential disturbance event, calculated using the corresponding volume calculation formula based on the shape of the spatial intersection region (e.g., cube, sphere, irregular polyhedron, etc.). For example, the volume of a cube region is length * width * height, and the volume of a sphere region is four-thirds multiplied by pi multiplied by the cube of the radius (pi is taken as 3.14). The preset minimum effective disturbance threshold refers to... Pre-defined thresholds for distinguishing whether a disturbance event has a real impact include a minimum duration threshold and a minimum spatial intersection volume threshold. These thresholds are determined based on the robot's movement speed, body structure, task type, and safety requirements of the interaction scenario. For example, the minimum duration threshold is set to 300 milliseconds, and the minimum spatial intersection volume threshold is set to 100 cubic centimeters. The disturbance judgment result refers to the judgment result of whether each potential disturbance event exceeds the preset minimum effective disturbance threshold, specifically divided into two cases: "exceeds the threshold (constitutes an effective disturbance)" and "does not exceed the threshold (does not constitute an effective disturbance)".

[0128] Optionally, the specific processing procedure of this embodiment of the invention is as follows: the intent recognition system extracts each disturbance event in the potential disturbance event set one by one, and calculates its duration and spatial intersection volume respectively; the calculated duration is compared with a preset minimum duration threshold, and the spatial intersection volume is compared with a preset minimum spatial intersection volume threshold; if the duration of a disturbance event exceeds the preset minimum duration threshold and the spatial intersection volume exceeds the preset minimum spatial intersection volume threshold, then the event is determined to exceed the preset minimum effective disturbance threshold and constitute a valid disturbance; if any parameter does not exceed the corresponding threshold, then the event is determined not to exceed the preset minimum effective disturbance threshold and does not constitute a valid disturbance; the judgment results of all potential disturbance events are sorted and summarized to obtain the disturbance judgment result.

[0129] Step 4042: Based on the disturbance judgment results, the disturbance events that constitute effective spatial disturbances are selected to obtain a subset of effective disturbance events. Then, based on the occurrence time of each disturbance event in the subset of effective disturbance events, the candidate intents in the corresponding time window in the final context intent candidate set are traced back to obtain the intent tracing results.

[0130] Optionally, the intent recognition system filters out disturbance events constituting effective spatial disturbances based on the disturbance judgment results, obtaining a subset of effective disturbance events. Effective spatial disturbances refer to disturbance events that can actually interfere with or affect the robot's current task execution, i.e., disturbance events that "exceed the threshold" in the disturbance judgment results. The subset of effective disturbance events refers to a set consisting of all disturbance events constituting effective spatial disturbances, where each event retains core information such as its corresponding occurrence time window and spatial overlap range.

[0131] Furthermore, the intent recognition system, based on the occurrence time of each perturbation event in the subset of valid perturbation events, traces back to the candidate intents within the corresponding time window in the final context intent candidate set, obtaining the intent tracing result. Here, the occurrence time refers to the overlapping time period (i.e., the occurrence time window) corresponding to the valid perturbation events. The time window refers to the user interaction time window that overlaps with the occurrence time window of the valid perturbation events; tracing back refers to the process of searching and matching candidate intents with overlapping time windows in the final context intent candidate set according to the occurrence time windows of the valid perturbation events; the intent tracing result is a set consisting of all candidate intents that match the occurrence time windows of the valid perturbation events. If multiple valid perturbation events exist, and different events trace back to the same candidate intent, then that candidate intent is retained only once in the result; if there are no events in the subset of valid perturbation events, or no candidate intent matches the occurrence time window of the event, then the intent tracing result is an empty set.

[0132] Optionally, the specific processing procedure of this embodiment of the invention is as follows: the intent recognition system extracts each perturbation event in the subset of valid perturbation events one by one and determines its occurrence time window; for each occurrence time window, it traverses all candidate intents in the final context intent candidate set and extracts the user interaction time window associated with each candidate intent; it determines whether two time windows overlap, and if they overlap (overlap duration is not less than 100 milliseconds), the candidate intent is determined as the backtracking intent corresponding to the current valid perturbation event; after deduplicating all backtracking intents corresponding to valid perturbation events, the backtracking intents are sorted and summarized to obtain the intent backtracking result.

[0133] Step 4043: Based on the intent backtracking results, candidate intents that are aligned with the time of the effective perturbation event are identified as perturbation-related intents, and the number of perturbation-related intents is determined.

[0134] Optionally, based on the intent backtracking results, the intent recognition system identifies candidate intents that are time-aligned with the effective disturbance event as disturbance-related intents. Here, time alignment means that the overlap between the user interaction time window associated with the candidate intent and the occurrence time window of the effective disturbance event is no less than 80% of the duration of the effective disturbance event, ensuring a high degree of temporal consistency between the user interaction behavior corresponding to the candidate intent and the effective disturbance event. Disturbance-related intents refer to interaction intents that are closely temporally related to the effective disturbance event and may be generated by the user intrusion into the robot's operating space.

[0135] Furthermore, the intent recognition system determines the number of intents associated with the perturbation. The number of intents refers to the total number of perturbation-associated intents. If the intent backtracking result is an empty set, the number of intents is zero; if the intent backtracking result contains multiple non-repeating candidate intents, the number of intents is the number of corresponding candidate intents; if the intent backtracking result contains only one candidate intent, the number of intents is one.

[0136] The specific processing steps are as follows: The intent recognition system performs time alignment verification on the candidate intents in the intent backtracking results, judges the time alignment of each candidate intent with the corresponding valid disturbance event, filters out the candidate intents that meet the time alignment conditions, and identifies them as disturbance-related intents; the identified disturbance-related intents are deduplicated to obtain the number of intents of disturbance-related intents.

[0137] Step 4044: Perform interaction intent matching based on the number of intents to obtain the human-computer interaction intent recognition result.

[0138] Optionally, the intent recognition system performs interaction intent matching based on the number of intents to obtain the human-computer interaction intent recognition result, as specifically in steps 40441 to 40445.

[0139] The embodiments of the present invention improve the accuracy and reliability of human-computer interaction intent recognition through spatiotemporal correlation and multi-round screening, effectively avoid the problem of intent misjudgment caused by spatial interference, ensure the smoothness of human-computer collaborative interaction, and thus improve the naturalness and accuracy of human-computer interaction.

[0140] Optionally, the processes of steps 40441 to 40445 include: Step 40441: If the number of intents is a preset number, then the perturbation-related intents are determined as the human-computer interaction intent recognition results.

[0141] Optionally, the preset quantity in this embodiment of the invention is 1. Therefore, if the number of intentions is equal to 1, the unique perturbation associated intention is directly determined as the human-computer interaction intention recognition result.

[0142] Step 40442: If the number of intentions is greater than the preset number, then based on the user limb movement trajectory and robot operation space boundary corresponding to each disturbance-related intention, determine the approach rate of each disturbance-related intention, and determine the disturbance-related intention corresponding to the largest approach rate as the priority intention.

[0143] Optionally, if the number of intentions is greater than 1, the intention recognition system determines the approach rate of each disturbance-related intention based on the user's limb movement trajectory and the robot's operating space boundary, and determines the disturbance-related intention corresponding to the maximum approach rate as the priority intention.

[0144] Among them, the user limb movement trajectory refers to the sequence formed by arranging the position coordinates of the user's limb parts in the physical space that change over time and correspond to each disturbance-related intent in chronological order (derived from the trajectory information associated with the user's dynamic occupied area set constructed in step 401); the robot operating space boundary refers to the outer contour boundary of the robot's current operating space area determined in step 402, which is composed of the outer contour of the task movement channel, the reachable range of the end effector, and the preset safety buffer boundary; the approach rate refers to the instantaneous speed at which the user's limb moves towards the robot operating space boundary along the movement trajectory, and its magnitude reflects the urgency of the user's interaction behavior, with a larger value indicating a more urgent user interaction need; the priority intent refers to the intent that corresponds to the most urgent user interaction need among multiple disturbance-related intents.

[0145] The specific processing steps are as follows: The intent recognition system retrieves the user's limb movement trajectory corresponding to each disturbance-related intent one by one, extracts multiple consecutive key time nodes and corresponding limb position coordinates near the robot's operating space boundary in the trajectory; for each key time node, the straight-line distance between the user's limb position at that node and the nearest point on the robot's operating space boundary is calculated; based on the time interval between two adjacent key time nodes and the corresponding change in straight-line distance, the instantaneous velocity of the user's limb moving towards the robot's operating space boundary is calculated, i.e., the approach rate (calculated by dividing the change in straight-line distance between two adjacent time nodes by the time interval; if the distance gradually decreases, the approach rate is positive, and if the distance increases, it is negative; a positive value is taken as the effective approach rate); the effective approach rates corresponding to each disturbance-related intent are filtered, and the largest effective approach rate is selected as the approach rate of that intent; the approach rates of all disturbance-related intents are compared, and the disturbance-related intent with the largest approach rate is selected and determined as the priority intent.

[0146] Step 40443: Based on the endpoint position of the user's limb movement trajectory associated with the priority intent, determine whether it is located within a preset sub-region of the robot's operating space. The preset sub-region includes the path entrance, the area above the grasping point, and the emergency stop sensing area.

[0147] Optionally, the intent recognition system determines whether the endpoint of the user's limb movement trajectory associated with the priority intent is located within a preset sub-region of the robot's operating space, thus completing the adaptation verification of the priority intent. Here, the endpoint of the user's limb movement trajectory refers to the limb position coordinates of the last critical time node in the user's limb movement trajectory corresponding to the priority intent; this position represents the final landing point of the user's interactive action. The preset sub-region of the robot's operating space refers to a sub-space area with specific functional significance pre-defined within the robot's current operating space area. These areas are key areas for user-robot interaction, specifically including the path entrance, the area above the grasping point, and the emergency stop sensing area. The path entrance refers to the sub-region corresponding to the entrance position into the core operating area in the robot's movement channel for performing the task. The area above the grasping point refers to the sub-region directly above the target point where the robot's end effector performs the grasping operation. The emergency stop sensing area refers to the sub-space corresponding to the sensing area set within the robot's body or operating space to trigger the emergency stop function.

[0148] The specific processing steps are as follows: The intent recognition system retrieves the user's limb movement trajectory associated with the priority intent and extracts the coordinate information of the trajectory's endpoint; it obtains the range information of the robot's current operating space area determined in step 402, as well as the specific spatial coordinate range of the preset sub-regions (path entrance, above the gripping point, and emergency stop sensing area). (This range is preset according to the robot's task type, robotic arm operation requirements, and safety design requirements. For example, the path entrance sub-region is set as a rectangular space with a width of 100 cm and a height of 80 cm, the above the gripping point sub-region is set as a spherical space with a radius of 30 cm centered on the gripping point, and the emergency stop sensing area is set as a spherical space with a radius of 20 cm centered on the emergency stop button.) The coordinates of the user's limb movement trajectory endpoint are compared with the spatial coordinate ranges of the three preset sub-regions to determine whether the endpoint falls within the coordinate range of any one of the preset sub-regions. If it falls within any one of the ranges, it is determined to be "located in a preset sub-region"; otherwise, it is determined to be "not located in a preset sub-region".

[0149] Step 40444: If it is located in a preset sub-region, the priority intention will be determined as the result of human-computer interaction intention recognition.

[0150] Optionally, if the judgment result of step 40443 is "the end point of the user's limb movement trajectory is located in a preset sub-region", the intent recognition system will prioritize the intent to determine the human-computer interaction intent recognition result.

[0151] Step 40445: If it is not located in the preset sub-region, then redetermine the perturbation association intent until the perturbation association intent located in the preset sub-region is determined or only one perturbation association intent remains, and obtain the human-computer interaction intent recognition result.

[0152] Optionally, if the judgment result of step 40443 is "the end position of the user's limb movement trajectory is not located in the preset sub-region", the intent recognition system will redetermine the perturbation-related intent until a perturbation-related intent located in the preset sub-region is determined or only one perturbation-related intent remains, and finally obtain the human-computer interaction intent recognition result. Here, redetermining the perturbation-related intent refers to the cyclical process of removing the verified priority intents that do not match the preset sub-region from the current remaining perturbation-related intents, and then re-selecting priority intents from the remaining perturbation-related intents and performing adaptation verification; the remaining perturbation-related intents refer to the set of perturbation-related intents that have not been removed and have not completed the preset sub-region adaptation verification after one round of screening.

[0153] The specific processing procedure is as follows: The intent recognition system removes the priority intents that do not currently match the preset sub-regions from the perturbation-related intent set to obtain an updated perturbation-related intent set; the number of intents in the updated perturbation-related intent set is determined, and if the number of intents is zero, a message is displayed indicating that there are no valid interaction intents.

[0154] If the number of intents is 1, the only remaining perturbation-related intent is directly identified as the human-computer interaction intent recognition result; if the number of intents is greater than 1, return to step 40442, recalculate the proximity rate of each intent based on the updated perturbation-related intent set, determine the new priority intent, and then perform the preset sub-region adaptation verification in step 40443; repeat the above process of elimination, re-filtering, and verification until the perturbation-related intent located in the preset sub-region is filtered out (and identified as the recognition result) or only one remaining perturbation-related intent is identified (and identified as the recognition result).

[0155] The embodiments of the present invention, through hierarchical verification and cyclic screening, ultimately determine the human-computer interaction intent recognition result, which can accurately match the user's real interaction needs and adapt to the robot's operating space characteristics and task execution requirements, thereby improving the naturalness and accuracy of human-computer interaction.

[0156] The following describes the robot human-computer interaction intent recognition system based on IoT collaboration provided by the present invention. The robot human-computer interaction intent recognition system based on IoT collaboration described below can be referred to in correspondence with the robot human-computer interaction intent recognition method based on IoT collaboration described above.

[0157] Reference Figure 2 , Figure 2 This is a schematic diagram of the structure of the robot human-computer interaction intent recognition system based on Internet of Things (IoT) collaboration provided by the present invention. The robot human-computer interaction intent recognition system based on IoT collaboration includes: The behavior perception module 210 is used to collect multimodal raw perception signals of users during the interaction process based on IoT sensing terminals deployed in the physical space, obtain multimodal perception sequences, and identify segments containing user-initiated interactive behaviors based on the multimodal perception sequences to obtain target perception segments. The environmental perception module 220 is used to collect current environmental state information from IoT environmental perception units within the spatial location area indicated by the target perception segment as an interaction event, and obtain environmental context data. The intent recognition module 230 is used to generate a final contextual intent candidate set representing the user's interaction intent based on interaction events and environmental context data; The intent matching module 240 is used to perform interactive intent matching based on the final context intent candidate set, combined with the robot's current task state and executable actions, to obtain the human-computer interaction intent recognition result.

[0158] The embodiments of the present invention improve the naturalness and accuracy of human-computer interaction.

[0159] Please see Figure 3 , Figure 3An embodiment diagram of an electronic device provided in accordance with the present invention. For example... Figure 3 As shown, an embodiment of the present invention provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, it implements the processes of steps 10 to 40.

[0160] Please see Figure 4 , Figure 4 An embodiment diagram of a computer-readable storage medium provided in accordance with an embodiment of the present invention is shown. Figure 4 As shown, this embodiment provides a computer-readable storage medium 400 on which a computer program 311 is stored. When the computer program 311 is executed by a processor, it implements the processes of steps 10 to 40.

[0161] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the robot human-computer interaction intention recognition method based on Internet of Things collaboration provided by the above methods, which includes steps 10 to 40.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for recognizing human-robot interaction intention based on Internet of Things cooperation, characterized in that, The method comprises the following steps: Collecting multi-modal original perception signals of a user in an interaction process based on Internet of Things (IoT) perception terminals deployed in a physical space to obtain a multi-modal perception sequence, and identifying a segment containing a user-initiated interaction behavior based on the multi-modal perception sequence to obtain a target perception segment; Taking the target perception segment as an interaction event, collecting current environmental state information based on IoT environment perception units in a spatial position region indicated by the interaction event to obtain environmental context data; Generating a final context intent candidate set representing a user interaction intent based on the interaction event and the environmental context data; Performing interaction intent matching based on the final context intent candidate set in combination with a current task state and executable actions of a robot to obtain a human-robot interaction intent recognition result. 2.The method of claim 1, wherein, The step process of determining the human-robot interaction intent recognition result comprises: Determining a user limb movement trajectory sequence based on user behavior features associated with each candidate intent in the final context intent candidate set, and constructing a dynamic occupation region set of the user in the physical space based on the user limb movement trajectory sequence; Determining a current operation space region of the robot based on task path planning information in the current task state and a robot working envelope; the current operation space region of the robot includes a task movement channel, an end effector reachable range, and a preset safety buffer boundary; Performing spatial intersection calculation under a time alignment condition based on the dynamic occupation region set of the user and the current operation space region of the robot to obtain a spatial intersection result; Performing interaction intent matching based on the spatial intersection result to obtain a human-robot interaction intent recognition result. 3.The method of claim 2, wherein, The interaction intent matching based on the spatial intersection result to obtain a human-robot interaction intent recognition result comprises: Determining whether there is a region with time overlap and spatial intersection based on the spatial intersection result to obtain a set of potential disturbance events, and determining whether each disturbance event in the set of potential disturbance events exceeds a preset minimum effective disturbance threshold based on a duration and a spatial intersection volume of the disturbance event to obtain a disturbance judgment result; Filtering out disturbance events constituting effective spatial disturbances based on the disturbance judgment result to obtain a subset of effective disturbance events, and backtracking the occurrence time of each disturbance event in the subset of effective disturbance events to a candidate intent in a corresponding time window in the final context intent candidate set to obtain an intent backtracking result; Determining a disturbance-associated intent based on the intent backtracking result, and determining the number of intents of the disturbance-associated intent; Performing interaction intent matching based on the number of intents to obtain a human-robot interaction intent recognition result.

4. The method of claim 3, wherein the method further comprises: The interaction intent matching based on the number of intents to obtain a human-robot interaction intent recognition result comprises: If the number of intents is a preset number, the disturbance-associated intent is determined as the human-robot interaction intent recognition result. If the number of the intentions is greater than the preset number, proximity rates of the disturbance-associated intentions are determined based on user limb movement trajectories corresponding to the disturbance-associated intentions and boundaries of a robot operating space, and a disturbance-associated intention corresponding to the maximum proximity rate is determined as a priority intention; It is judged whether the end position of the user limb movement trajectory associated with the priority intention is located in a preset sub-region of the robot operating space based on the end position; the preset sub-region includes a path entrance, an area above a grasping point, and an emergency stop sensing area; If the end position is located in the preset sub-region, the priority intention is determined as a human-robot interaction intention recognition result; If the end position is not located in the preset sub-region, the disturbance-associated intentions are re-determined until a disturbance-associated intention located in the preset sub-region or only one disturbance-associated intention is determined, and a human-robot interaction intention recognition result is obtained. 5.The method of claim 1, wherein, The generating of the final context intention candidate set representing the user interaction intention based on the interaction event and the environment context data comprises: A user interaction space-time anchor point is determined based on timestamp information and spatial position information of a user-initiated interaction behavior in the interaction event, and user behavior features are extracted from a multi-modal perception sequence associated with the interaction event based on the user interaction space-time anchor point; Action type, action duration, and action directionality in the user behavior features are used to obtain user behavior semantic description information, and a preliminary interaction tendency category expressed by the user in the interaction process is predicted based on the user behavior semantic description information; The final context intention candidate set is determined based on the user behavior semantic description information, the preliminary interaction tendency category, and the environment context data; each candidate intention in the final context intention candidate set represents an interaction intention of the user in the environment context.

6. The method of claim 5, wherein the method further comprises: The determining of the final context intention candidate set based on the user behavior semantic description information, the preliminary interaction tendency category, and the environment context data comprises: Physical object state information and spatial layout information in the environment context data are used to determine environment static semantic elements, and an interactable physical object in a current space region and a corresponding interactable functional attribute of the interactable physical object are analyzed based on the environment static semantic elements; Dynamic environment parameters in the environment context data are used to identify an environment event occurring in the current space region and an environment event influence relationship of the environment event on the user interaction; The user behavior semantic description information and the environment static semantic elements are used to establish a spatial proximity and functional matching relationship between a user action and an interactable physical object; The preliminary interaction tendency category and the environment dynamic semantic elements are used to analyze whether the user interaction behavior responds to a current environment event, and a first-level interaction intention is generated; The spatial proximity, the functional matching relationship, the environment event influence relationship, the first-level interaction intention, and the interactable functional attribute are used to determine the final context intention candidate set.

7. The method of claim 6, wherein the method further comprises: The determining of the final context intention candidate set based on the spatial proximity, the functional matching association, the environmental event influence relationship, the first level interaction intention and the interactive function attribute comprises: generating a second level interaction intention based on the spatial proximity and the functional matching association; performing intention cross verification based on the first level interaction intention and the second level interaction intention, excluding interaction intentions with contradictions or logical conflicts, to obtain a preliminary context intention candidate set; performing semantic completion on each interaction intention based on the user behavior semantic description information, the environmental event influence relationship and the interactive function attribute corresponding to each interaction intention in the preliminary context intention candidate set, to obtain the final context intention candidate set.

8. An Internet of Things cooperation-based machine-human interaction intention recognition system, characterized in that, The method for recognizing a human-robot interaction intention based on Internet of Things cooperation according to any one of claims 1 to 7; The system for recognizing a human-robot interaction intention based on Internet of Things cooperation comprises: a behavior perception module configured to collect multi-modal original perception signals of a user in an interaction process based on Internet of Things perception terminals deployed in a physical space, to obtain a multi-modal perception sequence, and to identify a target perception segment containing a user-initiated interaction behavior based on the multi-modal perception sequence; an environment perception module configured to collect current environmental state information based on Internet of Things environmental perception units in a spatial location area indicated by the interaction event, to obtain environmental context data, with the target perception segment as an interaction event; an intention recognition module configured to generate a final context intention candidate set representing a user interaction intention based on the interaction event and the environmental context data; an intention matching module configured to perform interaction intention matching based on the final context intention candidate set in combination with a current task state and executable actions of a robot, to obtain a human-robot interaction intention recognition result.

9. An electronic device comprising: a memory configured to store a computer software program; a processor configured to read and execute the computer software program, wherein the processor, when executing the computer software program, implements the method for recognizing a human-robot interaction intention based on Internet of Things cooperation according to any one of claims 1 to 7.

10. A non-transitory computer readable storage medium having stored therein a computer software program, characterized in that, The computer software program, when executed by the processor, implements the method for recognizing a human-robot interaction intention based on Internet of Things cooperation according to any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-modal interaction data processing method and system for children picture book reading

    CN121859270A