Target navigation method and device based on environmental context, robot and medium

Through multimodal data fusion and quadruple graph-driven spatiotemporal probability analysis, combined with user behavior patterns and environmental context information, the problems of insufficient response to fuzzy instructions and poor coordination between environmental perception and decision-making in existing technologies are solved, achieving a more intelligent and reliable object navigation effect.

CN120609357APending Publication Date: 2025-09-09PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510713215.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

When processing ambiguous or incomplete instructions, existing home assistance robots lack the ability to understand users' historical behavior patterns and cannot effectively combine environmental context information for reasoning and judgment, resulting in inefficient search in complex home environments. In addition, existing systems lack coordination in environmental perception, path planning, and target verification.

Method used

By receiving multimodal input for preprocessing, it generates structured text instructions, environmental semantic segmentation results and action intention analysis results, combines the quadruple map to generate a list of candidate search areas, and constructs a SLAM map for target navigation, and uses visual semantic verification to confirm the target object.

Benefits of technology

It improves the ability to understand ambiguous instructions, enhances environmental context perception and decision-making coordination, improves navigation accuracy and user experience, and lowers the usage threshold for elderly people with cognitive impairment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120609357A_ABST
    Figure CN120609357A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent robots and old-age care, and discloses a target navigation method and device based on environmental context, a robot and a medium, which can be applied to a household auxiliary robot intelligent medicine delivery scene, and the method comprises the following steps: receiving multi-modal input of a user, preprocessing the multi-modal input, and generating preprocessed multi-modal data; reading the current time and the environment state, and performing perceptual analysis based on a pre-constructed tetrad map, the current time and the environment state to generate a candidate search area list; determining a target candidate area from the candidate search area list; generating a target navigation path based on the SLAM map and the target candidate area, and moving the robot to the target candidate area according to the target navigation path; and obtaining a target candidate region image, and performing visual semantic verification based on the target candidate region image and the target object label in the structured text instruction. According to the invention, the robot navigation accuracy and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of intelligent robots and elderly care technology, and in particular to a target navigation method, device, robot, and medium based on environmental context. Background Art

[0002] With the accelerated development of an aging society, the application of indoor assistive robots in elderly care settings is becoming increasingly widespread. Currently, assistive robots are mostly designed to perform tasks such as finding objects, providing reminders, accompanying care, and navigation. They are particularly valuable in helping the elderly find daily necessities. Elderly people with mild cognitive impairment commonly experience short-term memory loss, decreased attention span, and misplaced items. They often find themselves unable to find items such as medicine and glasses in their daily lives.

[0003] Existing home assistance robots primarily rely on explicit voice or visual commands, a form of interaction that presents a significant barrier to adoption for elderly individuals with cognitive impairments. More critically, existing systems lack the ability to understand historical user behavior patterns and effectively integrate environmental context for reasoning and judgment, resulting in poor performance when handling vague, incomplete, or ambiguous commands. Furthermore, existing solutions lack synergy across environmental perception, path planning, and target verification, making it difficult to achieve intelligent decision-making based on multimodal information. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to propose a target navigation method, device, robot and medium based on environmental context, which has the advantages of improving the ability to understand fuzzy instructions, enhancing environmental context perception and decision-making coordination, and achieving more intelligent and reliable object navigation effects.

[0005] In order to solve the above technical problems, an embodiment of the present application provides a target navigation method based on environmental context, including:

[0006] Receiving multimodal input from a user and preprocessing the multimodal input to generate preprocessed multimodal data, wherein the preprocessed multimodal data includes structured text instructions, environment semantic segmentation results, and action intention analysis results, and the structured text instructions include target object labels;

[0007] Read the current time and environmental status, and perform perception analysis based on a pre-built four-tuple map, the current time, and the environmental status to generate a list of candidate search areas;

[0008] Determining a target candidate area from the candidate search area list based on the structured text instruction, the environment semantic segmentation result, and the action intention analysis result;

[0009] Constructing a SLAM map, generating a target navigation path based on the SLAM map and the target candidate area, and moving the robot to the target candidate area according to the target navigation path;

[0010] A target candidate region image is acquired, and visual semantic verification is performed based on the target candidate region image and the target object label in the structured text instruction.

[0011] In order to solve the above technical problems, an embodiment of the present application provides a target navigation device based on environmental context, including:

[0012] A multimodal input module, configured to receive multimodal input from a user and preprocess the multimodal input to generate preprocessed multimodal data, wherein the preprocessed multimodal data includes structured text instructions, environment semantic segmentation results, and action intention analysis results, and the structured text instructions include target object labels;

[0013] A candidate search area generation module is used to read the current time and environmental status, and perform perception analysis based on a pre-built four-tuple map, the current time and the environmental status to generate a candidate search area list;

[0014] a target candidate region determination module, configured to determine a target candidate region from the candidate search region list based on the structured text instruction, the environment semantic segmentation result, and the action intention analysis result;

[0015] A target navigation path generation module is used to construct a SLAM map, generate a target navigation path based on the SLAM map and the target candidate area, and move the robot to the target candidate area according to the target navigation path;

[0016] The visual semantic verification module is used to obtain a target candidate area image and perform visual semantic verification based on the target candidate area image and the target object label in the structured text instruction.

[0017] In order to solve the above technical problems, a technical solution adopted by the present invention is: to provide a robot, including one or more processors; a memory for storing one or more programs, so that the one or more processors can implement any one of the above-mentioned target navigation methods based on environmental context.

[0018] To solve the above technical problems, a technical solution adopted by the present invention is: a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the above-mentioned target navigation methods based on environmental context.

[0019] The embodiment of the present invention provides a target navigation method, device, robot and medium based on environmental context. The method includes: receiving multimodal input from a user, preprocessing the multimodal input, generating preprocessed multimodal data, wherein the preprocessed multimodal data includes structured text instructions, environmental semantic segmentation results and action intention analysis results, and the structured text instructions include target object labels; reading the current time and environmental state, and performing perception analysis based on a pre-constructed four-tuple map, the current time and the environmental state to generate a candidate search area list; determining a target candidate area from the candidate search area list based on the structured text instructions, the environmental semantic segmentation results and the action intention analysis results; constructing a SLAM map, and generating a target navigation path based on the SLAM map and the target candidate area, and moving the robot to the target candidate area according to the target navigation path; obtaining an image of the target candidate area, and performing visual semantic verification based on the image of the target candidate area and the target object label in the structured text instruction. The embodiments of the present invention effectively combine user behavior patterns and environmental context information through multimodal data fusion processing, quadruple graph-driven spatiotemporal probability analysis, multi-dimensional candidate area screening and dynamic verification mechanism, thereby solving the problems of insufficient response to fuzzy instructions and poor coordination between environmental perception and decision-making in existing technologies, and having the advantages of improving navigation accuracy and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 This is a schematic diagram of an application environment of a target navigation method based on environmental context in one embodiment of the present invention;

[0022] Figure 2 This is a flowchart of the implementation of the target navigation method based on environmental context provided by an embodiment of the present application;

[0023] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S1;

[0024] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation method before step S2;

[0025] Figure 5 yes Figure 2A schematic flow chart of a specific implementation of step S2;

[0026] Figure 6 yes Figure 2 A schematic flow chart of a specific implementation of step S3;

[0027] Figure 7 yes Figure 2 A schematic flow chart of a specific implementation of step S4;

[0028] Figure 8 yes Figure 2 A schematic flow chart of a specific implementation of step S5;

[0029] Figure 9 This is a schematic diagram of a target navigation device based on environmental context provided by an embodiment of the present application;

[0030] Figure 10 Schematic diagram of the robot provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0032] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0034] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0035] It should be noted that the target navigation method based on environmental context provided in the embodiments of the present application is generally executed by a robot, and accordingly, the target navigation device based on environmental context is generally configured in the robot.

[0036] The target navigation method based on environmental context provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, the robot can receive user input commands, move to the area corresponding to the commanded object based on the command input, and provide feedback on the object's location to the user, or even grab the object and return it to the user. The robot of this application is a mobile intelligent robot. The present invention is described in detail below through specific embodiments.

[0037] The target navigation method based on environmental context provided by the present invention can be applied to the intelligent medicine delivery scenario of home-assistive robots in the field of elderly care.

[0038] Existing home assistance robots primarily rely on explicit voice or visual commands to perform tasks. This interaction method presents significant limitations for elderly individuals with cognitive impairments. When users provide ambiguous or incomplete instructions—for example, describing an item's function without accurately naming it, or forgetting its specific location—existing systems are unable to effectively interpret the command semantics and reason within the context of the environment. Traditional navigation methods typically rely solely on immediate sensor data and lack the ability to learn from historical user behavior patterns, resulting in inefficient search in complex home environments.

[0039] In order to solve the above problems, it is observed that user behavior has spatiotemporal regularity. For example, medicines are often stored in the bedroom for use at night, and glasses are often placed next to the desk. By constructing a knowledge graph that integrates the four-dimensional features of users, items, time, and space, a probability distribution model of the location of items can be established. At the same time, implicit information such as gesture pointing or gaze direction in multimodal input can help narrow the search scope. Based on this, natural language processing is combined with computer vision technology, candidate areas are generated through spatiotemporal probability calculation, and dynamic path planning and closed-loop verification mechanisms are introduced to form a complete context-aware navigation system. Therefore, the application proposes to receive multimodal input from users and pre-process it to generate data containing structured text instructions, environmental semantic segmentation results, and action intention analysis results, generate a list of candidate search areas through a four-tuple graph based on the current time and environmental status, construct a navigation path after screening the target area based on multi-dimensional information, and finally confirm the target object through visual verification.

[0040] Specifically, when a user issues a fuzzy command, the speech is first converted into structured text containing a target tag. For example, "Find that round thing" is parsed into "Target object tag: round object." The environmental camera captures images in real time and segments them into functional areas such as the kitchen and bedroom, while also capturing the user's pointing gestures to determine location information. The system reads the current time as 8:00 PM and, combined with historical data from the quadruple graph indicating that the user's frequently used nighttime medications are stored on the bedside table, calculates the bedroom as a high-probability candidate. The medication storage area determined by the environmental semantic segmentation is then combined with the gesture's pointing direction to select the coordinates of the bedside table as the target area. The robot constructs a real-time map containing obstacles and plans an obstacle avoidance path. Once in position, the robot captures an image of the cabinet surface with a camera, comparing it to determine whether a round object, such as a medicine box, is present for final verification.

[0041] See also Figure 2 , Figure 2 A specific implementation of a target navigation method based on environmental context is shown.

[0042] It should be noted that the method of the present invention is not limited to the method of Figure 1 The process sequence shown is limited to the following steps:

[0043] S1: Receive multimodal input from a user, and preprocess the multimodal input to generate preprocessed multimodal data.

[0044] Specifically, the robot can receive multimodal user input. It receives natural language input via a voice interface, such as "I put that little medicine box there" or "Help me find my glasses." It also receives environmental images via a visual interface, for example, through an integrated RGBD camera, supplemented by infrared imaging for low-light recognition. The robot can also perform action recognition, supporting user action information such as pointing and gestures. The multimodal input is then preprocessed to generate preprocessed multimodal data.

[0045] Among them, the preprocessed multimodal data includes structured text instructions, environmental semantic segmentation results and action intention analysis results, and the structured text instructions include target object labels.

[0046] See also Figure 3 , Figure 3 A specific implementation of step S1 is shown, which is described in detail as follows:

[0047] S11: Receive the multimodal input of the user, wherein the multimodal input includes user natural language instructions, environment images and user action information.

[0048] S12: Convert the user's natural language instruction into text information, and perform fuzzy semantic expansion on the text information to generate the structured text instruction.

[0049] S13: Constructing an environment point cloud map based on the environment image, and performing room semantic segmentation on the environment point cloud map to generate the environment semantic segmentation result.

[0050] S14: identifying the action intention based on the user action information through a posture estimation model, and performing fusion recognition based on the action intention and the environment image to generate the action intention analysis result.

[0051] Fuzzy semantic expansion refers to the expansion of the vocabulary of natural language instructions through a synonym library and contextual association model. This can be achieved by using a semantic similarity calculation method based on the BERT model to address issues with ambiguous user expressions or limited vocabulary. An environmental point cloud map refers to a three-dimensional spatial structure constructed using depth data collected by an RGB-D camera. This can be achieved by using a SLAM algorithm for real-time mapping, providing a spatial benchmark for room-level semantic segmentation. Room semantic segmentation refers to the semantic annotation of point cloud maps to distinguish different functional areas. This can be achieved using the PointNet++ deep learning model to establish a spatial cognition framework for environmental context. A pose estimation model refers to an algorithm for detecting and tracking key points on the human body. This can be achieved using the OpenPose framework to parse the spatial association between user actions and environmental objects.

[0052] Specifically, when a user issues a vague command like "find glasses," the system first generates a structured command through semantic expansion, including expanded vocabulary such as "reading glasses" and "metal frame." Simultaneously, the system uses a depth camera to collect environmental data and construct a 3D point cloud map. A semantic segmentation model is then used to identify areas such as the bedroom and study. When the user's finger is detected pointing toward the nightstand, the pose estimation model extracts the coordinates of the hand's key points and performs spatial matching based on the location of the nightstand in the environmental image. Ultimately, the system generates an intent analysis result pointing to the nightstand area. The three modal data processing steps are performed in parallel: semantic expansion enhances the robustness of text representation, point cloud segmentation establishes a computable environmental semantic model, and motion fusion recognition compensates for information missing from a single modality. This application effectively addresses the problem of misjudgment of intent caused by insufficient fusion of multimodal input information, improving the accuracy of target object localization in complex indoor environments. The semantic expansion module converts ambiguous natural language into structured commands, point cloud segmentation establishes a semantically annotated environmental model, and motion fusion recognition accurately captures the user's implicit intent. These three collaborative processes provide high-precision multi-source data support for subsequent navigation path planning.

[0053] S2: Read the current time and environmental status, and perform perception analysis based on a pre-built four-tuple map, the current time and the environmental status to generate a list of candidate search areas.

[0054] Specifically, it reads the current time and environmental status (such as kitchen, bedroom, or living room), and uses the spatiotemporal graph to perform probabilistic reasoning (such as searching for drug-related targets with a high probability in the morning); performs room semantic segmentation by integrating structured room perception models such as RoomNet and LayoutGPT, and then performs perception analysis on different areas to generate a list of candidate search areas.

[0055] See also Figure 4 , Figure 4 A specific implementation method before step S2 is shown, which is described in detail as follows:

[0056] S2A: Obtain the historical operation log of the user, and preprocess the historical operation log to generate a preprocessed operation log.

[0057] S2B: Extracting multi-dimensional features from the pre-processed operation log, wherein the multi-dimensional features include user information features, item features, time features, and space features.

[0058] S2C: Constructing the quadruple graph based on the user information feature, the item feature, the time feature, and the space feature.

[0059] S2D: Predict the spatiotemporal distribution probability of items based on the quadruple graph according to a graph neural network to generate a probabilistic behavior model.

[0060] Specifically, raw historical operation logs undergo data cleaning and format conversion to form a standardized set of operation records. A feature extraction module then extracts four feature vectors: user identity tags, item classification codes, time period identifiers, and spatial region codes. User information features can include attributes such as age and usage habits, item features can include dimensions such as item category and frequency of use, temporal features can be broken down into hierarchical levels such as hourly periods, days of the week, and seasons, and spatial features can be decomposed into elements such as room type and regional coordinates. During the construction of the quadruple graph, operation frequency edge weights are established between user and item nodes, and association probability edge weights are established between temporal and spatial nodes, forming a multidimensional interactive network. A graph neural network performs multiple rounds of message passing on the graph, calculating the influence weights of different neighboring nodes through an attention mechanism. Ultimately, it outputs the distribution probability values ​​of each item node at each spatial and temporal location, forming a dynamically updateable probabilistic behavior model. This approach transforms discrete features into structured knowledge representations by constructing a quadruple graph. The topological learning capabilities of graph neural networks are then leveraged to mine nonlinear relationships between multidimensional features, effectively addressing the feature fragmentation problem inherent in traditional approaches. This application can establish a probabilistic association model between user behavior and the spatiotemporal distribution of items, significantly improving the target navigation system's prediction accuracy for item locations. Based on multi-dimensional feature fusion and graph structure relationship mining, the system can accurately infer the possible storage areas of items in different time periods. Especially when processing ambiguous user instructions, it can narrow the search scope based on historical behavior patterns. The embodiments of this application effectively enhance the robot's ability to understand the contextual environment, provide reliable data support for subsequent candidate area generation, and solve the problem of low navigation efficiency caused by the lack of historical behavior analysis in existing systems.

[0061] Among them, historical operation logs refer to data sets that record user actions on items within a specific time period. This can be achieved by storing timestamps, user IDs, item categories, and location coordinates in a database, thereby capturing user behavioral habits. Preprocessed operation logs refer to structured data that has undergone data cleaning and format standardization. This can be achieved through deduplication, outlier filtering, and missing value filling to ensure the effectiveness of subsequent feature extraction. Multidimensional features refer to feature vectors with different semantic attributes extracted from operation logs. This can be achieved through natural language processing techniques to extract user attribute labels, computer vision models to identify item categories, time series analysis to extract periodic patterns, and spatial coordinate clustering to generate regional heat maps, comprehensively covering the spatiotemporal correlations of user behavior. Quadruple graphs refer to heterogeneous graph structures constructed with user nodes, item nodes, time nodes, and spatial nodes as basic elements. This can be achieved by storing node attributes and edge relationships in a graph database, thereby characterizing the complex interactions between multidimensional features. Graph neural networks refer to deep learning models that can process graph-structured data. This can be achieved through a graph attention network that aggregates information about neighboring nodes, thereby mining potential spatiotemporal dependencies in historical behavior patterns.

[0062] See also Figure 5 , Figure 5 A specific implementation of step S2 is shown, which is described in detail as follows:

[0063] S21: Read the current time and the environmental status.

[0064] S22: Generate structured environment description information based on the environment state, and mark object coordinates according to the structured environment description information.

[0065] S23: Using a Bayesian network to perform spatiotemporal probability calculation based on the current time, the object coordinates, and the quadruple map to generate a spatiotemporal probability distribution.

[0066] S24: Calculate the probability of each area according to the spatiotemporal probability distribution and the object coordinates, and generate the candidate search area list.

[0067] Specifically, after the environmental state data is input, the environmental semantic segmentation algorithm identifies and annotates the objects in the scene, forming structured description information including the location and category of the objects. The user behavior pattern data and the current time information in the quadruple map are loaded synchronously, and a joint probability model of time, space and behavior is established through the conditional probability reasoning module of the Bayesian network. The model uses the object coordinate data as the observation variable, the user historical operation data in the quadruple map as the prior probability, and the current time information as the time constraint to calculate the probability of the target object existing in each candidate area. The final generated list of candidate search areas is arranged in descending order according to the probability value to form a priority search sequence. The present application realizes a multi-dimensional fusion analysis of the user's historical behavior pattern and the real-time state of the environment, effectively improving the accuracy of candidate area generation. The search priority is optimized through the spatiotemporal probability model, the number of invalid area traversals is reduced, and the target search efficiency is significantly improved. Dynamic integration of time dimension information ensures that the system can adapt to scene changes in different time periods and enhances the time sensitivity to user behavior patterns.

[0068] Among them, structured environmental description information refers to the conversion of environmental states into structured data containing object categories, location coordinates, and spatial relationships. This can be achieved by using an environmental semantic segmentation algorithm to extract the object bounding box and map it to a three-dimensional coordinate system to establish a digital representation framework for the environmental state. A Bayesian network refers to a conditional probability reasoning tool based on a probabilistic graphical model. It can use the Markov chain Monte Carlo method for parameter learning and probability propagation calculations, and is used to integrate historical behavior data with real-time environmental states for joint probabilistic reasoning. Spatiotemporal probability distribution refers to a probability density function that combines the time dimension and the space dimension. This can be achieved by outputting the probability of occurrence of each candidate area at different time points through a Bayesian network, and is used to quantify the possibility of the target object existing under specific spatiotemporal conditions.

[0069] S3: Determine a target candidate area from the candidate search area list based on the structured text instruction, the environment semantic segmentation result and the action intention analysis result.

[0070] Specifically, large-scale visual language models (such as BLIP2 and OpenFlamingo) are used to align language targets with image regions based on structured text instructions and environmental semantic segmentation results. Prompt is used to construct a description of the target, such as "a small blue box that may contain antihypertensive drugs," to match the room image region. Feature vectors of candidate regions are matched, and the region with the highest similarity is selected as the target candidate region.

[0071] See also Figure 6 , Figure 6 A specific implementation of step S3 is shown, which is described in detail as follows:

[0072] S31: Convert structured text instructions into text feature vectors, and perform region segmentation and visual encoding on the environment semantic segmentation results to generate image region vectors.

[0073] S32: Calculate cosine similarity based on the text feature vector and the image region vector to obtain the similarity value.

[0074] S33: sorting the target candidate search area list into candidate areas according to the similarity value, and determining the target candidate area from the candidate area sorting result.

[0075] Specifically, when the user input contains a description of an ambiguous object, the structured text instruction forms a text vector by extracting key semantic features. The result of the semantic segmentation of the environment is divided into multiple candidate regions, and each region is visually encoded to generate a corresponding image vector. By calculating the cosine similarity between the text vector and the image vector of each region, the candidate region that best matches the text description can be screened out. For example, when the user is looking for a "round white medicine box", the text features are matched with the visual features of each region, and the region containing a white round object is preferentially selected as the target. The present application can effectively combine text semantics and visual features to achieve precise positioning of the target area when the user instruction is incomplete or ambiguous. This method significantly reduces the probability of regional misjudgment due to ambiguous instructions. For example, when the user only describes a "table for putting things", the system can accurately distinguish between different furniture categories such as dining tables and coffee tables. At the same time, through the sorting mechanism of quantitative similarity, the number of times the robot dynamically adjusts the target area is reduced, thereby improving navigation efficiency.

[0076] Among them, structured text instructions refer to standardized instructions containing target object labels formed through semantic analysis. Specifically, this can be achieved by using a natural language processing model to extract entities and relationships from user instructions, in order to eliminate ambiguous information in user input. Environmental semantic segmentation results refer to rasterized data that performs regional division and semantic annotation on scene images. Specifically, this can be achieved by using a deep learning segmentation network to perform pixel-level classification on RGB-D images, in order to establish spatial associations between candidate regions and object categories. Cosine similarity calculation refers to measuring the directional consistency of text features and visual features in vector space. Specifically, this can be achieved by using the inner product operation of the vector after L2 normalization, in order to quantify the degree of matching between multimodal features.

[0077] S4: Constructing a SLAM map, generating a target navigation path based on the SLAM map and the target candidate area, and moving the robot to the target candidate area according to the target navigation path.

[0078] Specifically, the SLAM module is connected to build a map and perform target navigation planning. During the navigation process, obstacles are detected and dynamic obstacle avoidance is performed by combining the YOLO or SAM model.

[0079] See also Figure 7 , Figure 7 A specific implementation of step S4 is shown, which is described in detail as follows:

[0080] S41: Collecting RGB images and depth data, and constructing the SLAM map based on the RGB images and the depth data.

[0081] S42: Acquire the current position, and generate a target navigation path based on the current position, the SLAM map, and the target candidate area.

[0082] S43: Driving the robot toward the target candidate area according to the target path, and using the YOLO target detection algorithm to perform dynamic obstacle avoidance during driving.

[0083] Specifically, RGB images and depth data are fed into the SLAM algorithm for joint optimization, generating a 3D map with real-world scale information. Navigation path planning uses the A* algorithm to search for paths between the current position and the coordinates of candidate target areas, combining the distribution of static obstacles in the map to generate a globally optimal route. During movement, the YOLO algorithm continuously processes camera input at 30 frames per second. Detecting dynamic obstacles triggers the local path replanning module, which adjusts the robot's direction and speed using a dynamic window algorithm.

[0084] Among them, RGB images and depth data refer to the two-dimensional color images and three-dimensional point cloud information obtained synchronously by visual sensors. Specifically, this can be achieved using Kinect or RealSense depth cameras. By fusing the two types of data, the scale drift problem of monocular vision SLAM can be solved. SLAM maps refer to the topological structure containing the three-dimensional coordinates and semantic labels of the environment generated during the simultaneous positioning and map construction process. Specifically, this can be achieved using the ORB-SLAM3 framework, which establishes an accurate environment model through feature point matching and depth data fusion. The YOLO target detection algorithm refers to a target recognition model based on a single-stage convolutional neural network. Specifically, it can be implemented using the YOLOv5 architecture. It outputs the coordinates of the obstacle bounding box and category labels through real-time inference. Depth images provide object distance information to assist in understanding three-dimensional space.

[0085] S5: Acquire a target candidate region image, and perform visual semantic verification based on the target candidate region image and the target object label in the structured text instruction.

[0086] Specifically, after the robot reaches the target candidate area position, it obtains the target candidate area image at the position, and performs visual semantic verification based on the target candidate area image and the target object label in the structured text instruction.

[0087] See also Figure 8 , Figure 8 A specific implementation of step S5 is shown, which is described in detail as follows:

[0088] S51: Acquire the target candidate region image, and calculate an image-text similarity score based on the target candidate region image and the target object label in the structured text instruction.

[0089] S52: If the image-text similarity score exceeds a preset threshold, the visual semantic verification is passed.

[0090] S53: If the image-text similarity score does not exceed the preset threshold, the visual semantic verification fails, and interactive feedback with the user is triggered.

[0091] Specifically, after the robot navigates to the target candidate area, the visual sensor captures a high-resolution image of the current area and extracts image feature vectors using a pre-trained cross-modal model. Simultaneously, the target object label in the structured text instruction is converted into a text feature vector. The cosine similarity between the two vectors is calculated as the image-text similarity score. If this score exceeds a preset threshold, the target object is considered to exist and matches the user instruction. Otherwise, an interactive feedback mechanism is activated. For example, if a user is searching for a "red medicine box" but the medicine box detected in the candidate area image is blue, the system will generate a prompt asking the user whether to expand the search range or confirm the color attribute.

[0092] Among them, the target candidate area image refers to the real-time scene data collected by the visual sensor after the robot reaches the candidate area. It can be implemented by an RGB camera or a depth camera to capture the actual visual features of the target object. The image-text similarity score refers to the degree of matching between the image and the text label calculated by a cross-modal model. It can be implemented by the CLIP model or the visual semantic embedding network to map the visual content and the text description to the same feature space for similarity measurement. The preset threshold refers to the pre-set judgment standard, which can be obtained based on historical verification data statistics or determined through dynamic adjustment algorithm optimization, and is used to balance the strictness and fault tolerance of verification. Interactive feedback refers to the active initiation of information exchange with the user when verification fails. It can be implemented through a voice dialogue interface or a graphical interface interaction to supplement missing information or correct erroneous judgments.

[0093] Furthermore, the robot can proactively clarify and interact with users. For example, when a target is unclear or has multiple candidates, it can initiate natural language feedback: "Are you referring to the box next to the sofa in the living room?" "I can't find it in the kitchen. Should I check in the bedroom?" Sub-modules such as "Selection Confirmation," "History Recall," and "Location Supplement" can be selected to enhance the naturalness of the interaction. The robot also counts the number of failures and automatically optimizes the interaction path and target ranking.

[0094] The embodiments of this application can be applied to intelligent medicine delivery scenarios in the elderly care sector. For example, an elderly person says, "I forgot where I put my medicine." The system analyzes the semantics and identifies "medicine" as the target object. It then uses user history to infer that the time is morning and the target object is likely "blood pressure medication." Using a knowledge graph and environmental perception, it infers that a common location is the kitchen table. A visual model is used to scan the kitchen, and based on language prompts, it searches for "blue box." The robot then navigates to the candidate target and responds, "I found this box. Is this it?" If the user confirms, the robot takes the box. If the user denies, the robot switches to the candidate and continues navigation.

[0095] In an embodiment of the present application, a user's multimodal input is received, and the multimodal input is preprocessed to generate preprocessed multimodal data, wherein the preprocessed multimodal data includes structured text instructions, environmental semantic segmentation results and action intention analysis results, and the structured text instructions include target object labels; the current time and environmental status are read, and a perception analysis is performed based on a pre-constructed four-tuple map, the current time and the environmental status to generate a candidate search area list; a target candidate area is determined from the candidate search area list based on the structured text instructions, the environmental semantic segmentation results and the action intention analysis results; a SLAM map is constructed, and a target navigation path is generated based on the SLAM map and the target candidate area, and the robot is moved to the target candidate area according to the target navigation path; an image of the target candidate area is obtained, and visual semantic verification is performed based on the target candidate area image and the target object label in the structured text instruction. The embodiment of the present invention effectively combines user behavior patterns with environmental context information through multimodal data fusion processing, quad-tuple graph-driven spatiotemporal probability analysis, multi-dimensional candidate area screening, and dynamic verification mechanisms, solving the problems of insufficient response to fuzzy instructions and poor coordination between environmental perception and decision-making in existing technologies. It has the advantages of improving navigation accuracy and user experience. The embodiment of the present application combines the target recognition and navigation method of user historical behavior, environmental context, and language fuzzy instructions, greatly lowering the usage threshold for elderly people with cognitive impairments. The elderly do not need to clearly express the target content; through historical behavior and scene reasoning, intelligent object search and error correction are achieved; compared with traditional rule-based or visual model-based systems, it has stronger generalization and adaptability; and improves the practical value and trust of assistive robots in humanized care scenarios.

[0096] Please refer to Figure 9 , as a response to the above Figure 2 The present application provides an embodiment of a target navigation device based on environmental context, and the device embodiment is similar to Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0097] like Figure 9 As shown, the target navigation device based on environmental context of this embodiment includes: a multimodal input module 61, a candidate search area generation module 62, a target candidate area determination module 63, a target navigation path generation module 64 and a visual semantic verification module 65, wherein:

[0098] A multimodal input module 61 is configured to receive multimodal input from a user and preprocess the multimodal input to generate preprocessed multimodal data, wherein the preprocessed multimodal data includes structured text instructions, environment semantic segmentation results, and action intention analysis results, and the structured text instructions include target object labels;

[0099] A candidate search area generation module 62 is configured to read the current time and environmental status, and perform a perception analysis based on a pre-built four-tuple map, the current time and the environmental status to generate a candidate search area list;

[0100] a target candidate region determining module 63, configured to determine a target candidate region from the candidate search region list based on the structured text instruction, the environment semantic segmentation result, and the action intention analysis result;

[0101] A target navigation path generation module 64 is configured to construct a SLAM map, generate a target navigation path based on the SLAM map and the target candidate area, and move the robot to the target candidate area according to the target navigation path;

[0102] The visual semantic verification module 65 is configured to obtain a target candidate region image and perform visual semantic verification based on the target candidate region image and the target object label in the structured text instruction.

[0103] Furthermore, the multimodal input module 61 includes:

[0104] An input receiving unit, configured to receive the multimodal input from the user, wherein the multimodal input includes a user natural language instruction, an environment image, and user action information;

[0105] a text instruction generating unit, configured to convert the user's natural language instruction into text information, and perform fuzzy semantic expansion on the text information to generate the structured text instruction;

[0106] a semantic segmentation unit, configured to construct an environment point cloud map based on the environment image, and perform room semantic segmentation on the environment point cloud map to generate the environment semantic segmentation result;

[0107] The intention parsing unit is used to identify the action intention based on the user action information through a posture estimation model, and to perform fusion recognition based on the action intention and the environment image to generate the action intention parsing result.

[0108] Furthermore, the candidate search area generating module 62 further includes:

[0109] an operation log obtaining unit, configured to obtain the user's historical operation logs, and pre-process the historical operation logs to generate pre-processed operation logs;

[0110] a feature extraction unit, configured to extract multi-dimensional features from the pre-processed operation log, wherein the multi-dimensional features include user information features, item features, time features, and space features;

[0111] A graph construction unit, configured to construct the four-tuple graph based on the user information feature, the item feature, the time feature, and the space feature;

[0112] A probability prediction unit is used to predict the spatiotemporal distribution probability of items based on the quadruple graph according to a graph neural network to generate a probabilistic behavior model.

[0113] Furthermore, the candidate search area generating module 62 includes:

[0114] An environment status reading unit, configured to read the current time and the environment status;

[0115] an object coordinate generating unit, configured to generate structured environment description information based on the environment state, and mark object coordinates according to the structured environment description information;

[0116] A spatiotemporal probability calculation unit, configured to perform spatiotemporal probability calculation based on the current time, the object coordinates, and the quadruple map using a Bayesian network to generate a spatiotemporal probability distribution;

[0117] The region probability calculation unit is used to calculate the probability of each region according to the spatiotemporal probability distribution and the object coordinates, and generate the candidate search region list.

[0118] Furthermore, the target candidate region determination module 63 includes:

[0119] An image region vector generation unit is used to convert structured text instructions into text feature vectors, and perform region segmentation and visual encoding on the environment semantic segmentation results to generate image region vectors;

[0120] a similarity calculation unit, configured to perform cosine similarity calculation based on the text feature vector and the image region vector, the similarity value;

[0121] The region sorting unit is configured to sort the target candidate search region list into candidate regions according to the similarity value, and determine the target candidate region from the candidate region sorting result.

[0122] Furthermore, the target navigation path generation module 64 includes:

[0123] A data acquisition unit, configured to acquire RGB images and depth data, and construct the SLAM map based on the RGB images and the depth data;

[0124] A path generation unit, configured to obtain a current position and generate a target navigation path based on the current position, the SLAM map, and the target candidate area;

[0125] A dynamic obstacle avoidance unit is used to drive the robot to the target candidate area according to the target path, and to use the YOLO target detection algorithm to perform dynamic obstacle avoidance during driving.

[0126] Furthermore, the visual semantic verification module 65 includes:

[0127] An image-text similarity calculation unit, configured to obtain the target candidate region image and calculate an image-text similarity score based on the target candidate region image and the target object label in the structured text instruction;

[0128] A verification passing unit, configured to pass the visual semantic verification if the image-text similarity score exceeds a preset threshold;

[0129] The verification failure unit is configured to determine that the visual semantic verification has failed if the image-text similarity score does not exceed the preset threshold, and trigger interactive feedback with the user.

[0130] In order to solve the above technical problems, the present application also provides a robot. Figure 10 , Figure 10 This is the basic structural diagram of the robot in this embodiment.

[0131] The robot 7 includes a memory 71, a processor 72, and a network interface 73 that are interconnected via a system bus. Figure 10 Only a robot 7 having three components, memory 71, processor 72, and network interface 73, is shown. However, it should be understood that implementation of all illustrated components is not required, and more or fewer components may be implemented instead. It will be understood by those skilled in the art that a robot herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, and the like.

[0132] The robot can interact with the user through remote controls, touch panels, or voice-controlled devices.

[0133] The memory 71 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, magnetic disk, optical disk, etc. In some embodiments, the memory 71 may be an internal storage unit of the robot 7, such as the robot 7's hard disk or internal memory. In other embodiments, the memory 71 may also be an external storage device of the robot 7, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash memory card, etc. equipped on the robot 7. Of course, the memory 71 may also include both the robot 7's internal storage unit and its external storage device. In this embodiment, the memory 71 is generally used to store the operating system and various application software installed on the robot 7, such as the program code of the target navigation method based on environmental context. In addition, the memory 71 may also be used to temporarily store various types of data that have been output or are about to be output.

[0134] In some embodiments, the processor 72 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 72 is generally used to control the overall operation of the robot 7. In this embodiment, the processor 72 is used to execute program code stored in the memory 71 or process data, such as executing the program code of the aforementioned environment context-based target navigation method to implement various embodiments of the environment context-based target navigation method.

[0135] The network interface 73 may include a wireless network interface or a wired network interface. The network interface 73 is generally used to establish a communication connection between the robot 7 and other electronic devices.

[0136] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores a computer program, and the computer program can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned target navigation method based on environmental context.

[0137] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), including a number of instructions for enabling a terminal device to execute the methods of each embodiment of the present application.

[0138] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.

Claims

1. A target navigation method based on environmental context, characterized in that: include: Receiving multimodal input from a user and preprocessing the multimodal input to generate preprocessed multimodal data, wherein the preprocessed multimodal data includes structured text instructions, environment semantic segmentation results, and action intention analysis results, and the structured text instructions include target object labels; Read the current time and environmental status, and perform perception analysis based on a pre-built four-tuple map, the current time, and the environmental status to generate a list of candidate search areas; Determining a target candidate area from the candidate search area list based on the structured text instruction, the environment semantic segmentation result, and the action intention analysis result; Constructing a SLAM map, generating a target navigation path based on the SLAM map and the target candidate area, and moving the robot to the target candidate area according to the target navigation path; A target candidate region image is acquired, and visual semantic verification is performed based on the target candidate region image and the target object label in the structured text instruction.

2. The target navigation method based on environmental context according to claim 1, characterized in that: The receiving of the multimodal input from the user and preprocessing the multimodal input to generate preprocessed multimodal data includes: receiving the multimodal input from the user, wherein the multimodal input includes a user natural language instruction, an environment image, and user action information; Converting the user's natural language instruction into text information, and performing fuzzy semantic expansion on the text information to generate the structured text instruction; Constructing an environment point cloud map based on the environment image, and performing room semantic segmentation on the environment point cloud map to generate the environment semantic segmentation result; The action intention is identified based on the user action information through a posture estimation model, and fusion identification is performed based on the action intention and the environment image to generate the action intention analysis result.

3. The target navigation method based on environmental context according to claim 2, characterized in that: Before reading the current time and environmental status and performing perception analysis based on a pre-built four-tuple map, the current time and the environmental status to generate a candidate search area list, the method further includes: Obtaining the user's historical operation logs, and preprocessing the historical operation logs to generate preprocessed operation logs; Extracting multi-dimensional features from the pre-processed operation log, wherein the multi-dimensional features include user information features, item features, time features, and space features; Constructing the quadruple graph based on the user information feature, the item feature, the time feature, and the space feature; The spatiotemporal distribution probability of the items is predicted based on the quadruple graph according to the graph neural network to generate a probabilistic behavior model.

4. The target navigation method based on environmental context according to claim 1, characterized in that: The reading of the current time and the environmental state, and the performing of a perception analysis based on a pre-built four-tuple map, the current time and the environmental state to generate a list of candidate search areas include: Read the current time and the environmental status; generating structured environment description information based on the environment state, and marking object coordinates according to the structured environment description information; Using a Bayesian network to perform spatiotemporal probability calculation based on the current time, the object coordinates, and the quadruple map to generate a spatiotemporal probability distribution; The probability of each area is calculated according to the spatiotemporal probability distribution and the object coordinates, and the candidate search area list is generated.

5. The target navigation method based on environmental context according to claim 1, characterized in that: The determining of the target candidate area from the candidate search area list based on the structured text instruction, the environment semantic segmentation result, and the action intention analysis result includes: Convert structured text instructions into text feature vectors, and perform region segmentation and visual encoding on the environment semantic segmentation results to generate image region vectors; Performing cosine similarity calculation based on the text feature vector and the image region vector, the similarity value; The target candidate search area list is sorted into candidate areas according to the similarity value, and the target candidate area is determined from the candidate area sorting result.

6. The target navigation method based on environmental context according to claim 1, characterized in that: The step of constructing a SLAM map, generating a target navigation path based on the SLAM map and the target candidate area, and moving the robot to the target candidate area according to the target navigation path includes: Collecting RGB images and depth data, and constructing the SLAM map based on the RGB images and the depth data; Obtaining a current position, and generating a target navigation path based on the current position, the SLAM map, and the target candidate area; The robot is driven toward the target candidate area according to the target path, and the YOLO target detection algorithm is used for dynamic obstacle avoidance during driving.

7. The target navigation method based on environmental context according to any one of claims 1 to 6, characterized in that: The acquiring of the target candidate region image and performing visual semantic verification based on the target candidate region image and the target object label in the structured text instruction includes: Acquire the target candidate area image, and calculate an image-text similarity score based on the target candidate area image and the target object label in the structured text instruction; If the image-text similarity score exceeds a preset threshold, the visual semantic verification is passed; If the image-text similarity score does not exceed the preset threshold, the visual semantic verification fails, and interactive feedback with the user is triggered.

8. A target navigation device based on environmental context, characterized in that: include: A multimodal input module, configured to receive multimodal input from a user and preprocess the multimodal input to generate preprocessed multimodal data, wherein the preprocessed multimodal data includes structured text instructions, environment semantic segmentation results, and action intention analysis results, and the structured text instructions include target object labels; A candidate search area generation module is used to read the current time and environmental status, and perform perception analysis based on a pre-built four-tuple map, the current time and the environmental status to generate a candidate search area list; a target candidate region determination module, configured to determine a target candidate region from the candidate search region list based on the structured text instruction, the environment semantic segmentation result, and the action intention analysis result; A target navigation path generation module is used to construct a SLAM map, generate a target navigation path based on the SLAM map and the target candidate area, and move the robot to the target candidate area according to the target navigation path; The visual semantic verification module is used to obtain a target candidate area image and perform visual semantic verification based on the target candidate area image and the target object label in the structured text instruction.

9. A robot, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method for target navigation based on environmental context as claimed in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the target navigation method based on environmental context according to any one of claims 1 to 7.

Citation Information

Cited By

  • Robot natural language navigation method, system and equipment

    CN120910245A

  • Robot navigation switching method and device based on working scene, equipment and medium

    CN120991886A

  • Robot control method and device based on visual guidance and medium

    CN121018581A

  • A vision-guided robot control method, device, and medium

    CN121018581B

  • Navigation method, system and equipment based on multi-modal model, medium and product

    CN121230736A