Multimodal perception-driven scene adaptive response method

By constructing multi-layer scene graphs and multimodal perception sensors, the robot can achieve dynamic perception and adaptive response in complex and changing scenarios, improving the perception accuracy and adaptability of the response, and solving the problem of insufficient response of robots in complex scenarios in existing technologies.

CN120516716BActive Publication Date: 2025-10-03广州云趣信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510971240.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-03
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

Existing robots lack multimodal dynamic perception and adaptive response capabilities in complex and changing scenarios, and find it difficult to cope with task interruptions, command reconstruction, or emergency response caused by scenario changes.

Method used

By constructing a multi-layer scene graph, including the spatial layer, the interaction layer, and the state layer, integrating motion and static recognition, interaction fitting analysis, and a tolerant response mechanism, activating multimodal perception sensors for multimodal perception, establishing dynamic and static obstacle identification, analyzing interaction tasks and performing fitting compensation, and configuring tolerant response attention to achieve scene adaptive response.

Benefits of technology

The robot's dynamic perception accuracy and adaptability and accuracy of response actions in multiple scenarios have been improved, ensuring that it can flexibly respond to changes and efficiently perform tasks in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120516716B_ABST
    Figure CN120516716B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal perception-driven scene adaptive response method, which relates to the field of data processing technology, including: activating a multimodal perception sensor to perform multimodal perception and establish a multi-layer scene graph; performing dynamic and static channel identification on the multi-layer scene graph and establishing dynamic and static obstacle identification; executing interactive attention at the interactive layer within the multi-layer scene graph and establishing interactive tasks; performing interactive impact fitting analysis within the multi-layer scene graph based on the interactive tasks and dynamic and static obstacle identification, establishing fitting compensation, correcting the interactive tasks, and establishing the robot's timed response actions; using the results of the interactive impact fitting analysis to configure tolerant response attention, and performing attention association on the timed response actions to perform scene adaptive response management. The present invention solves the technical problem in the prior art that robots lack multimodal dynamic perception and adaptive response capabilities in complex and changing scenes, achieving the technical effect of improving the dynamic perception accuracy and the adaptability and accuracy of response actions in multiple scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a scene adaptive response method driven by multimodal perception. Background Art

[0002] In complex service scenarios like government affairs, users often express their needs through voice, gestures, or mixed interactions, often taking into account dynamic factors such as emotions, location, and service processes. Service robots, while performing their tasks, require multimodal perception and dynamic understanding of user intent, scene state, and interactive behaviors to achieve semantically appropriate responses. However, existing robots often rely on single-channel perception and static rule matching, lacking the ability to integrate and model multi-source information such as voice, gestures, and states, making it difficult to cope with demands such as task interruption, command reconstruction, or emergency response brought about by changing scenarios. Summary of the Invention

[0003] This application provides a multimodal perception-driven scene adaptive response method, which is used to solve the technical problem in the existing technology that robots lack multimodal dynamic perception and adaptive response capabilities in complex and changing scenes.

[0004] In view of the above problems, the present application provides a scene adaptive response method driven by multimodal perception.

[0005] This application provides a multimodal perception-driven scene adaptive response method, the method comprising:

[0006] Activate the multimodal perception sensor, perform multimodal perception of the scene, and establish a multi-layer scene graph, which includes a spatial layer, an interaction layer, and a state layer; after performing dynamic and static channel identification on the multi-layer scene graph, establish dynamic and static obstacle identification; after performing interactive attention on the interactive layer within the multi-layer scene graph, establish an interactive task; perform an interactive impact fitting analysis within the multi-layer scene graph based on the interactive task and the dynamic and static obstacle identification, and establish a fitting compensation; after correcting the interactive task using the fitting compensation, establish the robot's timing response action; configure tolerant response attention using the interactive impact fitting analysis results, and perform attention association on the timing response action based on the tolerant response attention to perform scene adaptive response management.

[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0008] The present application activates a multimodal perception sensor, performs multimodal perception of a scene, and establishes a multi-layer scene graph, the multi-layer scene graph including a spatial layer, an interaction layer, and a state layer; after performing dynamic and static channel identification on the multi-layer scene graph, a dynamic and static obstacle identification is established; after performing interactive attention on the interactive layer within the multi-layer scene graph, an interactive task is established; based on the interactive task and the dynamic and static obstacle identification, an interactive impact fitting analysis is performed within the multi-layer scene graph to establish a fitting compensation; after correcting the interactive task using the fitting compensation, a time-series response action of the robot is established; using the interactive impact fitting analysis results to configure a tolerant response attention, the time-series response action is associated with the tolerant response attention to perform scene adaptive response management. The present invention solves the technical problem in the prior art that robots lack multimodal dynamic perception and adaptive response capabilities in complex and changing scenes. By constructing a multi-layer scene graph including a spatial layer, an interaction layer, and a state layer, and integrating dynamic and static identification, interactive fitting analysis, and a tolerant response mechanism, the present invention achieves the technical effect of improving the accuracy of dynamic perception and the adaptability and accuracy of response actions in multiple scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0010] Figure 1 A schematic diagram of the process flow of the multimodal perception-driven scenario adaptive response method provided in an embodiment of the present application;

[0011] Figure 2 A schematic diagram of the process of establishing a multi-layer scene graph in the multimodal perception-driven scene adaptive response method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0012] This application provides a multimodal perception-driven scene adaptive response method to solve the technical problem in the existing technology that robots lack multimodal dynamic perception and adaptive response capabilities in complex and changeable scenes. By constructing a multi-layer scene graph including a spatial layer, an interaction layer, and a state layer, and integrating dynamic and static recognition, interactive fitting analysis, and a tolerant response mechanism, the application achieves the technical effect of improving the dynamic perception accuracy and the adaptability and accuracy of response actions in multiple scenes.

[0013] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0014] It should be noted that any variations of the terms "include" and "have" are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.

[0015] Examples, such as Figure 1 As shown, the present application provides a multimodal perception-driven scene adaptive response method, the method comprising:

[0016] Step S100: Activate the multimodal perception sensor, perform multimodal perception of the scene, and establish a multi-layer scene graph, which includes a spatial layer, an interaction layer, and a state layer.

[0017] In an embodiment of the present application, the positioning sensor in the multimodal perception sensor is first activated, and the robot obtains its own real-time position data in the target scene, and uploads the data to the cloud to obtain a scene feedback template that matches the current position. Based on the template, the scene association sensor is started, and the target scene is perceived and updated in combination with multi-source data such as images, voice, and position information collected by the multimodal perception sensor. By fusing the perception data with the feedback template, a multi-layer scene graph is finally established. The multi-layer scene graph is constructed for service scenarios such as government affairs processing and human-computer interaction consultation, and integrates feature information such as voice content, action instructions, and service area distribution. It can dynamically reflect changes in the user's language expression and operation instructions during the service process. The multi-layer scene graph consists of a spatial layer, an interaction layer, and a state layer. The spatial layer represents the spatial distribution and geometric relationship of each entity in the scene, the interaction layer represents the interaction information between operable objects or between people, and the state layer reflects the real-time changing state information such as temperature, humidity, light, and warning status in the environment.

[0018] Further, such as Figure 2 As shown, in the method provided by the embodiment of the application, the method of activating the multimodal perception sensor, performing multimodal perception of the scene, and establishing a multi-layer scene graph also includes:

[0019] Activate the positioning sensor in the multimodal perception sensor to obtain real-time location data; after uploading the real-time location data to the cloud, receive the scene feedback template from the cloud; establish a scene association sensor, use the scene feedback template as a basic template, and perform scene multimodal perception update through the scene association sensor and the multimodal perception sensor to establish the multi-layer scene graph.

[0020] In an embodiment of the present application, the positioning sensor in the multimodal perception sensor, such as GPS or inertial navigation device, is first activated, and a position acquisition method is used to obtain the real-time position coordinates and orientation information of the target entity in three-dimensional space, thereby obtaining real-time position data.

[0021] Next, the real-time location data is uploaded to the cloud via wireless communication methods (such as Wi-Fi or cellular networks). The cloud then uses spatial indexing to retrieve and return a scene feedback template corresponding to the location from the existing scene database. This template provides structural information, interaction patterns, and prior information about the environment state in the current scene.

[0022] Next, scene-related sensors, such as laser rangefinders, infrared thermal imagers, and ambient temperature and humidity sensors, are established based on the scene feedback template to capture spatial features and environmental information related to the feedback template. These established scene-related sensors are then used in conjunction with the activated multimodal perception sensors to perform a multimodal perception update. This involves comparing and fusing newly acquired multi-source perception data, including images, sounds, location information, and environmental parameters, with the scene feedback template to correct the target distribution, behavioral relationships, and state parameters in the scene, thereby dynamically updating the current scene information.

[0023] Finally, based on the updated multimodal perception data, hierarchical division is carried out according to spatial distribution, interaction relationship and state characteristics, and the spatial layer (such as object position and configuration), interaction layer (such as focus object and interaction action) and state layer (such as environmental information such as lighting, temperature and humidity) are generated in sequence, thereby establishing a multi-layer scene graph.

[0024] Step S200: After performing dynamic and static channel identification on the multi-layer scene graph, dynamic and static obstacle identifications are established.

[0025] In an embodiment of the present application, during the process of identifying static and dynamic channels in a multi-layer scene graph to establish static and dynamic obstacle identification, the scene feedback template is first used to initialize the static and dynamic channels. This includes updating the channel identification based on the static feature identification in the template and completing the scene-specific adaptation of the static and dynamic channels in combination with the template content. Subsequently, scene feature extraction is performed on the multi-layer scene graph to obtain feature information such as the spatial structure and motion state of the target object. The extracted results are input into the initialized static and dynamic channels, and a dual-layer dynamic and static perception recognition based on time sequence and features is performed to comprehensively determine the static or dynamic properties of the object. Finally, the static and dynamic obstacle identification is established, achieving the classification and identification of fixed and movable obstacles in the scene.

[0026] Furthermore, in the method provided in the embodiment of the application, after the dynamic and static channels are identified on the multi-layer scene graph, dynamic and static obstacle identifications are established, further comprising:

[0027] The scene feedback template is used to initialize the channels of the dynamic and static channels. The dynamic and static channel initialization includes using the static feature identifier of the scene feedback template to update the dynamic and static channel identifier and using the scene feedback template to perform scene-based adaptation of the dynamic and static channels; after extracting the scene features of the multi-layer scene graph, the feature extraction results are input into the initialized dynamic and static channels, and dual-layer dynamic and static perception recognition based on timing and features is performed to establish dynamic and static obstacle identifiers.

[0028] In an embodiment of the present application, a scene feedback template is first used to initialize the channels of the dynamic and static channels. By calling the static feature identifiers contained in the template, such as the ground, wall, and fixed structure boundaries, a region overlap matching method is used to map the static regions in the template to the spatial layers of the current multi-layer scene graph. The positions and shape boundaries are then compared in the image coordinate system to update the identifiers of the static channels in the dynamic and static channels. For example, if the wall region marked in the template has a high degree of overlap with the boundary in the grayscale image of the current scene graph, it is marked as a static obstacle. For example, in a government service hall scenario, permanent structures such as walls, service counters, and seating areas can be pre-configured in the template as static feature identifiers for subsequent static region calibration. Simultaneously, based on the scene type defined in the scene feedback template (e.g., indoor corridor, outdoor plaza, etc.), the corresponding static and dynamic recognition parameter configuration table is called to set settings such as the time window size for background modeling, the frame difference threshold, and the edge detection sensitivity. This automatically adjusts the channel recognition range and judgment conditions, thereby achieving scene-specific adaptation of the dynamic and static channels.

[0029] After completing the channel initialization, the scene feature extraction is performed on the multi-layer scene graph. Specifically, it includes sequentially performing inter-frame difference processing on the image sequence to detect areas in the image where pixels change over time; performing background modeling processing on continuous images to identify areas that remain unchanged for a long time; and combining edge detection methods to identify target areas with obvious contour structures. After integrating the above three types of processing results, annotation information is formed to describe the spatial position of each object in the scene, whether there are signs of motion, and whether it belongs to a static area, that is, the feature extraction result is obtained. For example, when a user enters the business processing window area, it is detected that he appears in a specific area and moves. The movement trajectory is compared with the original static background image, and thus it is identified as a dynamic target, forming an identification mark of the user approaching the event.

[0030] Next, the feature extraction results are fed into the initialized static and dynamic channels, and then a dual-layer dynamic and static perception recognition process based on timing and features is performed. In this process, the feature recognition layer is first activated, using a static target template matching method to identify static features such as the object's shape, area, and boundary stability, generating corresponding feature semantic identifiers and, from these, a first static and dynamic perception identifier. The temporal perception layer is then activated to track the motion trajectory of the same target in the image sequence. A position-based target tracking method is used to analyze its displacement across consecutive frames, generating a second static and dynamic perception identifier. Finally, a perceptual fusion judgment method is used to compare and determine the consistency between the first and second static and dynamic perception identifiers, outputting the final result and completing the dual-layer dynamic and static perception recognition process, establishing a static and dynamic obstacle identifier.

[0031] Furthermore, in the method provided in the embodiment of the application, the execution of dual-layer motion and static perception recognition based on time sequence and features further includes:

[0032] Activate the feature recognition layer, use the feature recognition layer to perform feature recognition authentication of the feature extraction result, establish a feature semantic identifier, and generate a first motion and still perception identifier based on the feature semantic identifier; activate the timing perception layer, perform feature identification on the feature extraction result, perform feature tracking based on the timing relationship, and establish a second motion and still perception identifier based on the feature tracking result; perform perception fusion on the first motion and still perception identifier and the second motion and still perception identifier to complete dual-layer motion and still perception recognition.

[0033] In this embodiment, the feature recognition layer is first activated to analyze the static properties of the feature extraction results. Specifically, a boundary stability determination method is used to extract and align the contours of the target region in consecutive image frames. By calculating the overlap ratio of the contour shapes (for example, using pixel overlap or contour distance error), it is determined whether the target region has significant boundary changes between frames. If the overlap ratio is greater than 90%, the boundary is considered stable. Simultaneously, a background region matching method is used to compare the target region in the current image frame with a stable region generated based on a background modeling method (such as frame averaging). If the region is not identified as a foreground object in consecutive frames and remains within the background region for a long period of time, the target is determined to have stable background properties. For example, a fixed monitoring pole or curbstone, whose shape and position remain unchanged across the image sequence, can be determined to be a static object. Combining the boundary stability and background matching results, a feature semantic identifier is generated, which forms the first motion-static perception identifier, which describes whether the target tends to be a static object in terms of its spatial structure.

[0034] The temporal perception layer is then activated to perform temporal behavior analysis on the feature extraction results. Specifically, an inter-frame position change tracking method is used to record the coordinates of the center position of each target area and calculate its position change within a set time window (e.g., 5 consecutive frames). If the cumulative displacement of the center position in consecutive frames is greater than a threshold (e.g., the total displacement exceeds 15 pixels), the target is judged to have a clear motion trend. In addition, combined with the inter-frame region shape change (e.g., the offset of the boundary contour or morphological deformation), it is determined whether there is discontinuity or local dynamic behavior. For example, a user standing in front of a self-service machine in a government service hall remains standing and does not operate in the first 3 frames. In the last 2 frames, a slight movement displacement of the hand extending toward the touch screen occurs. In this case, it should be recorded as an object with "temporal dynamic behavior." Based on the above judgment results, a second motion-static perception identifier is generated to represent the actual motion state of the target in the time dimension.

[0035] Finally, the first and second motion-stillness perception flags are fused. If the two flags match, for example, both are "static" or "dynamic," the object's motion-stillness status can be directly determined, completing the basic recognition process. If there's a discrepancy, such as if the feature recognition indicates static but the temporal perception indicates dynamic, the state conflict fusion mechanism is activated to comprehensively analyze the strength of the judgment basis for the two results. For example, if the object's boundary is very stable, indicating high confidence static, but small fluctuations are detected in the temporal sequence, it is labeled as a static dynamic object. This indicates that the object, while currently stationary, is inherently movable, such as a person or vehicle. For example, a robot in a service hall may detect a user standing at a self-service terminal for an extended period without operating it. The user's boundary outline is stable and the area overlaps significantly, so the feature recognition layer identifies them as a static object. However, if the user's hand is detected gradually raising and adjusting its posture multiple times, or their head frequently tilts toward the robot, the temporal perception layer identifies continuous small displacements and changes in orientation, identifying potential interaction intent. Finally, after fusion, the object is labeled as a static dynamic object.

[0036] After completing the dual-layer static and dynamic perception recognition through the above process, the recognition results are written into the state layer of the multi-layer scene graph. Using the target's unique identifier as an index and combining it with the corresponding static and dynamic attribute tags, a static and dynamic obstacle identifier is created. This identifier clearly indicates whether each target is a static obstacle, a dynamic obstacle, or a static dynamic obstacle.

[0037] Step S300: After executing interactive attention of the interactive layers in the multi-layer scene graph, an interactive task is established.

[0038] In this embodiment, the robot first uses target detection methods to locate operable objects in the interaction layer. These objects are targets with which the robot can interact, primarily the user's body posture or gestures. Next, it uses multimodal perception sensors (such as RGB cameras, depth sensors, and infrared sensors) to perform interaction perception on these objects, generating interaction perception attention. Based on this interaction perception attention, the robot analyzes the user's interaction intent through multimodal interaction recognition and maps it to specific interaction tasks.

[0039] Furthermore, in the method provided in the embodiment of the application, after executing the interactive attention of the interactive layer in the multi-layer scene graph, establishing the interactive task also includes:

[0040] After positioning the operable objects of the interaction layer, interactive perception is performed through a multimodal perception sensor to establish interactive perception attention; multimodal interaction recognition is performed based on the interactive perception attention to establish the interactive task.

[0041] In an embodiment of the present application, the operable objects in the interaction layer are first located. This process uses target detection methods, such as the YOLO model, to locate the operable objects in the interaction layer, mainly including the user's body posture or gestures. The YOLO model is a deep learning model based on a convolutional neural network (CNN). It recognizes and locates objects by training a large amount of labeled image data. During the training process, the YOLO model learns how to recognize the features of operable objects such as arms and gestures, and can accurately determine the location of these objects in the image. Ultimately, the spatial coordinates and category information of the operable objects are obtained through target detection.

[0042] Next, multimodal sensors are used to perceive interactions and establish interaction awareness. This step first integrates data from multimodal sensors, such as vision, depth, and voice, using a context-aware fusion space. This information is synchronously processed using fusion techniques (such as Kalman filtering) to form a comprehensive perception data space. Semantic extraction methods are then used to extract interaction semantics from the fused data, such as identifying user gestures, movements, or voice commands. Based on this extracted semantic information, the urgency of tasks is analyzed, and task priorities are established using pre-set rules. Finally, interaction awareness attention is generated based on task priorities.

[0043] Finally, multimodal interaction recognition is performed based on interaction perception attention. In this process, a support vector machine (SVM) model is used for multimodal interaction recognition. An SVM is a supervised learning model that uses a training dataset to learn and recognize different interaction modes (such as clicks, touches, and voice). The training process involves a labeled dataset containing different user interaction behaviors (such as gestures, touch locations, and voice commands), as well as labels for each interaction behavior. Using this labeled data, the SVM model learns the characteristics of interaction behaviors (such as gesture speed and touch locations) and classifies them based on the new perception data. After training, the SVM model identifies the user's interaction intent and generates corresponding interaction tasks. For example, in a government service robot scenario, when a user waves, nods, or says something like "I need help" to the robot, image recognition determines the amplitude of their hand movement, and voice recognition captures the request keywords. By integrating this information, the SVM identifies the request guidance intent and generates interaction tasks such as "Guide to window X" or "Start process introduction."

[0044] Furthermore, in the method provided in the embodiment of the application, the interactive perception is performed by the multimodal perception sensor to establish interactive perception attention, and further includes:

[0045] Establish a context-aware fusion space; use the context-aware fusion space to extract semantics from interactive perception data, use the semantic extraction results to extract task urgency features, and construct task priorities; establish interactive perception attention based on the task priorities.

[0046] In the embodiments of this application, multiple sensors (such as visual sensors, voice sensors, depth sensors, and positioning sensors) are first activated to collect multidimensional data from the environment and interactive objects. These sensors provide different types of information, such as images, sounds, location, and status. Through sensor fusion technology, this data from different sensor sources is integrated into a unified data model, forming a context-aware fusion space. This space not only includes the spatial distribution of objects in the physical environment, but also includes behavioral patterns of user interaction and changes in the environment.

[0047] After establishing a context-aware fusion space, semantic extraction is used to understand the sensory data and obtain semantic extraction results. Deep learning algorithms and natural language processing technologies are used to extract semantic information from multimodal sensory data, including visual, acoustic, and location information. For example, image recognition technology can identify user gestures, or speech recognition technology can parse user voice commands. These technologies can identify user intent, such as "move an item" or "open the door," and the identified user intent becomes the semantic extraction result.

[0048] Next, the semantic extraction results are used to extract task urgency features. This process uses a rule engine to analyze the urgency of tasks. For example, when a voice command contains keywords such as "urgent" or "immediately," the task's urgency is automatically assessed as high priority. If the voice command contains expressions such as "slow" or "can wait," the task's priority is assessed as low. In addition to voice commands, task urgency is further assessed based on user interaction behaviors (such as gestures). For example, if a user makes rapid gestures, the task is inferred to be high priority. Conversely, if the user exhibits slow or relaxed movements, the task's urgency may be assessed as low. Through this rule analysis, task urgency features are extracted and quantified into task priorities such as high, medium, and low.

[0049] Finally, interaction-aware attention is established based on task priority. This dynamically adjusts attention based on predefined rules and task priorities. When tasks are of higher priority, these are prioritized, ensuring that urgent tasks are executed as quickly as possible. This intelligently identifies and handles urgent tasks, ensuring that the robot can efficiently execute interactive tasks based on user needs.

[0050] Step S400: performing interaction impact fitting analysis in a multi-layer scene graph according to the interaction task and the dynamic and static obstacle identifiers, and establishing fitting compensation.

[0051] In an embodiment of the present application, when performing an interaction impact fitting analysis within a multi-layer scene graph based on interaction tasks and dynamic and static obstacle identification, first, an interaction path is created based on the interaction task to determine the path required for the robot to perform the task. Then, using the static obstacle identification within the dynamic and static obstacle identification, a static spatial scene is created to identify fixed obstacles in the environment. Next, a distance impact fitting analysis is performed based on the interaction path and the static spatial scene to calculate the impact of static obstacles in the interaction path on task execution and generate a static fitting compensation. Finally, based on the static fitting compensation, a final fitting compensation is established.

[0052] Furthermore, in the method provided in the embodiment of the application, performing interaction impact fitting analysis in a multi-layer scene graph based on the interaction task and the dynamic and static obstacle identifications and establishing fitting compensation further includes:

[0053] An interaction path is created according to the interaction task; a static space scene is created using the static obstacle identifiers in the dynamic and static obstacle identifiers; a distance impact fitting analysis is performed according to the interaction path and the static space scene to establish a static fitting compensation; and the fitting compensation is established based on the static fitting compensation.

[0054] In the embodiment of the present application, an interaction path is first created based on the interaction task. In this process, a path planning algorithm (such as the A* algorithm or the Dijkstra algorithm) is used to calculate the optimal interaction path according to the requirements of the interaction task.

[0055] Next, the static obstacle markers in the dynamic and static obstacle markers are used to create a static spatial scene. During this process, the robot uses the dynamic and static obstacle markers to identify static obstacles in the environment and create a static spatial scene. Static obstacles are typically locations and objects that do not change, such as walls and fixed furniture. The robot uses laser radar (LIDAR) or depth sensors to obtain geometric information about the environment. It then uses environmental modeling techniques (such as point cloud processing or 3D reconstruction) to extract the location, shape, and size of static obstacles, and then constructs a model of the static obstacles in the static spatial scene.

[0056] A distance impact fit analysis is then performed based on the interaction path and static spatial scene. During this process, the robot uses geometric calculation methods (such as Euclidean distance or Manhattan distance) to evaluate the distance between the interaction path and static obstacles and perform a distance impact fit analysis. This analysis calculates the distance between each point in the interaction path and the static obstacle and determines whether there is a potential collision or interference. For example, if a point on the path is too close to an obstacle, potentially causing mission failure, the point is marked as an area that needs to be avoided or adjusted. The output of this process is a static fit compensation, which adjusts the path to ensure safe mission execution.

[0057] Finally, fitting compensation is established based on static fitting compensation. This process first uses the dynamic obstacle identifiers in the static and dynamic obstacle identifiers to create dynamic obstacles in the static spatial scene. Next, the dynamic obstacle mapping prediction model is activated to predict the movement of the dynamic obstacle and generate initial prediction results. Then, based on the interaction path and initial prediction results, an interactive fitting analysis of the dynamic obstacles is performed to generate dynamic fitting compensation. Finally, by combining static and dynamic fitting compensation, the fitting compensation is constructed and the interactive fitting analysis results are obtained.

[0058] Furthermore, in the method provided in the embodiment of the application, establishing the fitting compensation based on the static fitting compensation further includes:

[0059] A dynamic obstacle is created in the static spatial scene using a dynamic obstacle identifier among the dynamic and static obstacle identifiers; a mapping prediction model of the dynamic obstacle is activated, movement prediction of the dynamic obstacle in the static spatial scene is performed, and an initial prediction result is established; an interactive impact fitting analysis of the dynamic obstacle is performed based on the interactive path and the initial prediction result, and a dynamic fitting compensation is established; and the fitting compensation is established using the static fitting compensation and the dynamic fitting compensation.

[0060] In this embodiment of the present application, dynamic obstacles are first created within the static spatial scene using the dynamic obstacle identifiers within the static and dynamic obstacle identifiers. In this step, the robot uses the dynamic obstacle identifiers, already acquired through perception, to insert dynamic obstacles into the static spatial scene. These dynamic obstacle identifiers are calibrated during the environmental perception phase using multimodal sensors (such as RGB cameras, depth sensors, and lidar). Based on the identified dynamic obstacles (such as moving people or objects), they are added to the environmental model of the static spatial scene, forming a real-time, complete scene containing both static and dynamic obstacles.

[0061] Next, the dynamic obstacle mapping prediction model is activated and its movement within the static spatial scene is predicted to generate an initial prediction result. During this process, the dynamic obstacle mapping prediction model uses motion prediction algorithms such as Kalman filtering or particle filtering to predict the future trajectory of the dynamic obstacle based on its historical motion data (such as speed, acceleration, and direction). The Kalman filter model combines historical position and motion information to accurately predict the likely location of a dynamic obstacle within a certain period of time. For example, the Kalman filter can predict the user's position within the next few seconds based on their pace and direction. This process results in an initial prediction result.

[0062] The robot then performs a dynamic obstacle interaction fit analysis based on the interaction path and the initial predictions. This analysis uses geometric calculations, such as Euclidean distance, to evaluate the relationship between the interaction path and the dynamic obstacle. The robot calculates the relative position of each point in the interaction path to the dynamic obstacle to determine if there is a potential collision risk. For example, if a point on the interaction path intersects the predicted trajectory of a dynamic obstacle, the robot adjusts the path to avoid the collision. This analysis results in dynamic fit compensation, which adjusts the path plan to enable the robot to avoid dynamic obstacles and ensure successful mission execution.

[0063] Finally, the final fit compensation is established based on static and dynamic fit compensation. This process combines the effects of static obstacles with predictions of dynamic obstacles, using a path optimization algorithm (such as the A* algorithm) to comprehensively adjust the path. At this stage, static and dynamic fit compensation are combined to ensure that the robot can avoid static obstacles while also coping with interference from dynamic obstacles.

[0064] Step S500: After correcting the interactive task using the fitting compensation, a time-series response action of the robot is established.

[0065] In the present embodiment, the interactive task is first corrected using fitting compensation. Through the aforementioned path optimization and obstacle avoidance analysis, fitting compensation adjusts the robot's path or motion to avoid obstacles and ensure smooth task execution. The corrected interactive task takes into account both static and dynamic obstacles in the environment, ensuring that the robot can accurately perform the task without interference.

[0066] Next, based on the corrected interactive task, the robot establishes a sequential response action using a sequential control algorithm (such as PID control). Sequential response action involves the robot adjusting its movements based on the task's progress and real-time feedback. Based on real-time data such as the robot's current position, speed, and direction, the robot's movement sequence and timing are controlled to ensure the task proceeds as planned.

[0067] Step S600: configuring a tolerant response focus using the interaction impact fitting analysis result, and performing attention association on the temporal response action according to the tolerant response focus to perform scenario adaptive response management.

[0068] In this embodiment of the present application, when configuring tolerant response attention using the results of an interaction impact fitting analysis, the interaction impact fitting analysis results are first analyzed to obtain the time nodes and characteristics of the interaction impact. Based on these analysis results, an offset factor is configured, which is used to tolerantly expand the interaction impact time nodes. Subsequently, based on the tolerant expansion results and the interaction impact characteristics, the configuration of tolerant response attention is completed.

[0069] Next, the robot associates the sequential response actions with the tolerant response attention. This means that as the robot executes a task, the sequential response actions will adjust in real time based on environmental changes. The robot uses tolerant response attention to verify that the current environment meets the expected task execution conditions. If a deviation from the predetermined range is detected and the anomaly threshold is triggered, the robot generates an anomaly warning and immediately takes measures to halt the current task. This process completes scenario-specific adaptive response management.

[0070] Furthermore, in the method provided in the embodiment of the application, the configuration of the tolerant response attention using the interactive impact fitting analysis results also includes:

[0071] Analyze the interaction impact fitting analysis results to obtain interaction impact time nodes and interaction impact characteristics; configure an offset factor according to the fitting compensation, use the offset factor to tolerantly expand the interaction impact time nodes, and complete a tolerant response attention configuration based on the tolerant expansion results and the interaction impact characteristics.

[0072] In the embodiments of the present application, the results of the interaction impact fitting analysis are first analyzed to extract interaction impact time nodes and interaction impact features. Interaction impact time nodes refer to the critical moments in the task execution process where interactions occur. For example, when a robot encounters obstacles or other environmental changes while performing a task, specific time nodes become critical time points for the task. Interaction impact features describe the specific changes that occur at these time nodes, such as the appearance of obstacles or the movement of target objects.

[0073] Next, an offset factor is configured based on the fitting compensation, and this offset factor is used to extend the time window for interaction impact nodes. This process uses a rule engine to calculate an offset factor to extend the time window for interaction impact nodes. The offset factor is dynamically generated based on the intensity of interference and the urgency of the task during execution. During this process, the robot adjusts the task time window by analyzing the impact of obstacles in real time. For example, if a robot detects an obstacle during an object handling task and calculates that it may move into the task path within 3 seconds, the offset factor is used to extend the task execution time to 5 seconds, allowing the robot to temporarily pause and wait for the obstacle to move. This time extension ensures that the robot can continue task execution without being interrupted by brief environmental changes. For example, if the obstacle's position changes, an offset factor of 2 seconds allows the task execution time window to be extended by 2 seconds. If the obstacle moves within these 2 seconds, the robot can continue task execution. If the obstacle is not completely removed, the robot will wait until it is completely removed before continuing the task.

[0074] Finally, based on the tolerance expansion results and interaction impact characteristics, the configuration of tolerant response attention is completed. Based on environmental changes and the adjusted time window during task execution, it is determined which tasks should be executed first and which tasks can be executed later. In this step, the interaction impact characteristics and the tolerance expansion results work together to determine the response priority of tasks. For example, if the robot encounters a temporary obstacle during task execution, but the task itself has a lower priority, the task execution can be delayed until the obstacle is removed and execution can continue. Conversely, if the task priority is higher (such as requiring an immediate response to a user call or handling a frequently repeated user voice request), even if an obstacle appears, the robot will prioritize the task within the allowed time window to ensure timely task completion.

[0075] Furthermore, in the method provided in the embodiment of the application, the method of associating the timing response action with the tolerance response attention to perform scenario-adaptive response management further includes:

[0076] When executing a scene response based on the timing response action, attention verification is performed through the tolerant response attention; if the attention verification triggers an abnormal threshold, an abnormal warning is generated, and the robot is controlled to stop the action according to the abnormal warning, and an early warning is issued.

[0077] In an embodiment of the present application, when a scenario response is performed based on a sequential response action, that is, when the robot adjusts its task execution in real time according to a previously generated sequential response action, attention verification is performed through tolerant response attention. At this stage, the tolerant response attention mechanism verifies whether the current task is affected based on environmental changes during task execution. Tolerant response attention allows the robot to flexibly adjust the order and time window of task execution when encountering short-term interference. For example, if the robot encounters a temporary obstacle during a carrying task, the task is allowed to be paused for a certain period of time until the obstacle is removed, based on a pre-set tolerance range. At this point, through real-time monitoring and verification, it is determined whether the task needs to continue or further adjustments need to be made.

[0078] If the verification triggers an exception threshold, that is, the robot's waiting time exceeds the preset tolerance range while executing the task, an exception warning will be generated. The exception threshold refers to the maximum waiting time allowed during task execution. If the obstacle or interference is not removed within the specified time, exceeding the preset maximum tolerance waiting time, the exception threshold is triggered. For example, if the robot encounters an obstacle while executing a task and detects that the obstacle is still on the path, and the waiting time for the obstacle exceeds the preset maximum waiting time (for example, 10 seconds), the task execution is considered to be affected and the exception threshold is triggered.

[0079] Finally, based on the abnormal warning, the robot stops its action to avoid continuing the mission and potentially causing a safety accident or mission failure. In addition, the robot completes the warning by sounding an alarm.

[0080] In the embodiments of the present application, in summary, the embodiments of the present application have at least the following technical effects:

[0081] The present application activates a multimodal perception sensor, performs multimodal perception of a scene, and establishes a multi-layer scene graph, the multi-layer scene graph including a spatial layer, an interaction layer, and a state layer; after performing dynamic and static channel identification on the multi-layer scene graph, a dynamic and static obstacle identification is established; after performing interactive attention on the interactive layer within the multi-layer scene graph, an interactive task is established; based on the interactive task and the dynamic and static obstacle identification, an interactive impact fitting analysis is performed within the multi-layer scene graph to establish a fitting compensation; after correcting the interactive task using the fitting compensation, a time-series response action of the robot is established; using the interactive impact fitting analysis results to configure a tolerant response attention, the time-series response action is associated with the tolerant response attention to perform scene adaptive response management. The present invention solves the technical problem in the prior art that robots lack multimodal dynamic perception and adaptive response capabilities in complex and changing scenes. By constructing a multi-layer scene graph including a spatial layer, an interaction layer, and a state layer, and integrating dynamic and static identification, interactive fitting analysis, and a tolerant response mechanism, the present invention achieves the technical effect of improving the accuracy of dynamic perception and the adaptability and accuracy of response actions in multiple scenes.

[0082] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0083] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

[0084] This specification and drawings are merely illustrative of the present application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art may make various modifications and variations to this application without departing from the scope of this application. Thus, this application is intended to include such modifications and variations as fall within the scope of this application and its equivalents.

Claims

1. A multimodal perception-driven scenario adaptive response method, characterized in that: The method comprises: Activating a multimodal perception sensor to perform multimodal perception of a scene and establish a multi-layer scene graph, the multi-layer scene graph including a spatial layer, an interaction layer, and a state layer; After performing dynamic and static channel identification on the multi-layer scene graph, dynamic and static obstacle identification is established; After executing the interactive attention of the interactive layers in the multi-layer scene graph, establishing the interactive task; Performing interaction impact fitting analysis within a multi-layer scene graph based on the interaction task and the dynamic and static obstacle identifiers, and establishing fitting compensation; After correcting the interactive task using the fitting compensation, a time-series response action of the robot is established; The results of the interaction impact fitting analysis are used to configure tolerant response attention, and attention association is performed on the temporal response action according to the tolerant response attention to perform scenario adaptive response management.

2. The multimodal perception-driven scene adaptive response method according to claim 1, characterized in that: The activating of the multimodal perception sensor, performing multimodal perception of the scene, and establishing a multi-layer scene graph includes: Activate the positioning sensor in the multimodal perception sensor to obtain real-time location data; After uploading the real-time location data to the cloud, receiving a scene feedback template from the cloud; A scene association sensor is established, and the scene feedback template is used as a basic template. The scene multimodal perception update is performed through the scene association sensor and the multimodal perception sensor to establish the multi-layer scene graph.

3. The multimodal perception-driven scene adaptive response method according to claim 2, characterized in that: After the dynamic and static channels are identified on the multi-layer scene graph, dynamic and static obstacle identifications are established, including: Performing channel initialization of the dynamic and static channels using the scene feedback template, wherein the dynamic and static channel initialization includes updating the dynamic and static channel identifiers using the static feature identifiers of the scene feedback template and performing scene-based adaptation of the dynamic and static channels using the scene feedback template; After extracting scene features from the multi-layer scene graph, the feature extraction results are input into the initialized static and dynamic channels, and dual-layer static and dynamic perception recognition based on timing and features is performed to establish static and dynamic obstacle identification.

4. The multimodal perception-driven scene adaptive response method according to claim 3, characterized in that: The execution of dual-layer motion and stillness perception recognition based on time sequence and features includes: activating a feature recognition layer, performing feature recognition authentication on the feature extraction result using the feature recognition layer, establishing a feature semantic identifier, and generating a first motion-stillness perception identifier based on the feature semantic identifier; activating the temporal perception layer, performing feature identification on the feature extraction result, performing feature tracking based on the temporal relationship, and establishing a second motion and stillness perception identification based on the feature tracking result; The first motion and still perception identifier and the second motion and still perception identifier are subjected to perception fusion to complete double-layer motion and still perception recognition.

5. The multimodal perception-driven scene adaptive response method according to claim 1, characterized in that: After executing the interactive attention of the interactive layer in the multi-layer scene graph, establishing an interactive task includes: After locating the operable objects of the interactive layer, interactive perception is performed through a multimodal perception sensor to establish interactive perception attention; Multimodal interaction recognition is performed according to the interaction perception focus to establish the interaction task.

6. The multimodal perception-driven scene adaptive response method according to claim 5, characterized in that: The interactive perception is performed by the multimodal perception sensor to establish interactive perception attention, including: Establishing context-aware fusion space; Using the context-aware fusion space to perform semantic extraction in the interactive perception data, using the semantic extraction results to extract task urgency features and construct task priorities; An interaction-aware focus is established according to the task priorities.

7. The multimodal perception-driven scene adaptive response method according to claim 1, characterized in that: The performing of interaction impact fitting analysis in a multi-layer scene graph according to the interaction task and the dynamic and static obstacle identifiers, and establishing fitting compensation, includes: creating an interaction path according to the interaction task; Creating a static space scene using the static obstacle identifiers among the dynamic and static obstacle identifiers; Perform distance impact fitting analysis based on the interaction path and the static space scene, and establish static fitting compensation; The fitting compensation is established based on the static fitting compensation.

8. The multimodal perception-driven scene adaptive response method according to claim 7, characterized in that: The establishing the fitting compensation based on the static fitting compensation comprises: Creating a dynamic obstacle on the static space scene using the dynamic obstacle identifier among the dynamic and static obstacle identifiers; Activate the mapping prediction model of the dynamic obstacle, perform movement prediction of the dynamic obstacle in the static space scene, and establish an initial prediction result; Performing interactive impact fitting analysis of dynamic obstacles based on the interactive path and the initial prediction results, and establishing dynamic fitting compensation; The fitting compensation is established by utilizing the static fitting compensation and the dynamic fitting compensation.

9. The multimodal perception-driven scene adaptive response method according to claim 1, characterized in that: The configuration of the tolerance response using the interaction effect fitting analysis results includes: Analyze the interaction impact fitting analysis results to obtain interaction impact time nodes and interaction impact characteristics; An offset factor is configured according to the fitting compensation, the interaction impact time node is tolerantly expanded using the offset factor, and a tolerant response attention configuration is completed according to the tolerant expansion result and the interaction impact feature.

10. The multimodal perception-driven scene adaptive response method according to claim 1, characterized in that: The step of associating the timing response action with attention according to the tolerant response to perform scene adaptive response management includes: When executing a scenario response based on the timing response action, performing attention verification through the tolerant response attention; If the attention verification triggers an abnormal threshold, an abnormal warning is generated, and the robot is controlled to stop moving according to the abnormal warning and an alarm is issued.

Citation Information

Patent Citations

  • Autonomous robot decision-making system based on multi-modal perception fusion and method thereof

    CN119295883A

  • Ai-controlled sensor network for threat mapping and characterization and risk adjusted response

    US20250175456A1