Semantic object searching navigation method, system and equipment based on multi-sensor fusion and storage medium
By integrating RFID and 3D vision information, the problem of poor RFID positioning accuracy is solved, achieving centimeter-level precision positioning and identification of items, ensuring that robots can accurately find target items in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-03-20
AI Technical Summary
Existing RFID positioning technology is susceptible to environmental interference, has poor positioning accuracy, and is difficult to reliably obtain the precise spatial location of items.
By integrating RFID sensing information with 3D visual sensing information, the system obtains the identification information, multiple semantic attribute information, and spatial location information of the object, and marks them as points of interest in the environmental map. Path planning is then used to control the robot to move to the location of the target object.
It achieves centimeter-level precision positioning and recognition of objects in complex environments, improving positioning accuracy and stability, and ensuring that the robot can accurately find the target object.
Smart Images

Figure CN121702397A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of embodied intelligent robots, the Internet of Things and artificial intelligence, and in particular to a semantic object-finding navigation method, system, device and storage medium based on multi-sensor fusion. Background Technology
[0002] In scenarios such as smart warehousing and service robots, the core requirement is for robots to autonomously find and deliver designated items. Successful execution of this task depends on the accurate location and identification of items. Existing technologies generally employ Radio Frequency Identification (RFID) positioning technology, based on Received Signal Strength Indication (RSSI) for item location and identification.
[0003] Existing RFID positioning technology is susceptible to environmental interference and has poor positioning accuracy, making it difficult to reliably obtain the precise spatial location of items. Summary of the Invention
[0004] This application provides a semantic object-finding navigation method, system, device, and storage medium based on multi-sensor fusion to solve the technical problem of difficulty in reliably obtaining the precise spatial location of objects.
[0005] In a first aspect, embodiments of this application provide a semantic object-finding navigation method based on multi-sensor fusion, comprising: In the process of constructing an environmental map, items are located and identified by integrating RFID sensing information and 3D visual sensing information to obtain the item's identification information, multi-semantic attribute information and spatial location information. The item's identification information, multi-semantic attribute information and spatial location information are then marked as points of interest in the environmental map. In response to a received item search command, the item search command is parsed to determine the target item, and the target point of interest corresponding to the target item is matched from the environmental map; Based on the location of the target point of interest, a path is planned, and the robot is controlled to move to the location of the target item.
[0006] In some embodiments, during the construction of the environmental map, the process of locating and identifying objects by fusing RFID sensing information and 3D visual sensing information to obtain object identification information, multi-semantic attribute information, and spatial location information, and marking the object identification information, multi-semantic attribute information, and spatial location information as points of interest in the environmental map, includes: Acquire the RFID sensing information and the 3D visual sensing information; The RFID sensing information and the 3D visual sensing information are synchronized in time and associated with each other. Based on the associated data, a fusion calculation is performed to obtain the identification information, multiple semantic attribute information and spatial location information of the item. The identification information, multiple semantic attribute information, and spatial location information of the item are bound to the coordinates of the environmental map to obtain the point of interest.
[0007] In some embodiments, the RFID sensing information includes RFID signals, and the 3D visual sensing information includes 3D visual point cloud data. The step of synchronizing the RFID sensing information and the 3D visual sensing information in time and associating them with each other, and then performing fusion calculations based on the associative data to obtain the item's identification information, multiple semantic attribute information, and spatial location information includes: The RFID signal is filtered and smoothed to obtain signal strength and phase information. The approximate location of the item is estimated based on the signal strength and phase information. At the same time, the electronic product code (EPC) identifier of the RFID tag is decoded from the RFID signal as the identification information of the item. The 3D visual point cloud data is filtered and preprocessed for target detection to extract the appearance features of the item, and the three-dimensional position of the item is identified based on the appearance features. Based on the identification information of the item, query the item attribute database to obtain multi-semantic attribute information corresponding to the item; The identification information, the multiple semantic attribute information, the appearance features, and the three-dimensional position of the item are associated and fused to obtain the identification information, multiple semantic attribute information, and spatial position information of the item.
[0008] In some embodiments, before binding the identification information, multiple semantic attribute information, and spatial location information of the item with the coordinates of the environmental map to obtain the point of interest, the method further includes: The spatial location information of the items is transformed from their respective sensor coordinate systems to the robot's base coordinate system, and then uniformly transformed to the world coordinate system of the environmental map.
[0009] In some embodiments, binding the item's identification information, multiple semantic attribute information, and spatial location information to the coordinates of the environmental map to obtain the point of interest includes: During the process of constructing a grid map based on the SLAM algorithm, the identification information, multi-semantic attribute information and spatial location information of the identified items are converted into point of interest data in real time. The points of interest data are stored in a map database and associated with the coordinates of the environment map.
[0010] In some embodiments, parsing the item-finding command to determine the target item and matching the target point of interest corresponding to the target item from the environmental map includes: The natural language processing model is used to parse the item search command to obtain multi-semantic attribute information of the target item. The natural language processing model is fine-tuned based on the item attribute database. The target point of interest is determined by matching the multi-semantic attribute information of the target item with the identification information, multi-semantic attribute information, and spatial location information of the items associated with each point of interest in the environmental map.
[0011] In some embodiments, the step of planning a path based on the location of the target point of interest and controlling the robot to move to the location of the target item includes: Based on the robot's current position and the location of the target item's point of interest, the optimal path to reach the target item's location is planned; The robot is controlled to move along the optimal path, and during the movement, it uses onboard sensors to perceive the environment in real time to avoid obstacles.
[0012] Secondly, embodiments of this application provide a semantic object-finding navigation system based on multi-sensor fusion, comprising: The mapping and annotation module is used to locate and identify objects by integrating RFID sensing information and 3D visual sensing information during the construction of an environmental map, thereby obtaining the object's identification information, multi-semantic attribute information and spatial location information, and annotating the object's identification information, multi-semantic attribute information and spatial location information as points of interest in the environmental map. The instruction parsing and matching module is used to respond to the received item-finding instruction, parse the item-finding instruction to determine the target item, and match the target point of interest corresponding to the target item from the environmental map; The path planning and navigation module is used to plan a path based on the location of the target point of interest and control the robot to move to the location of the target item.
[0013] The third invention, according to an embodiment of this application, provides a robot, including: The robot itself; The aforementioned semantic object-finding navigation system based on multi-sensor fusion is integrated into the main control module of the robot body; In addition, the following hardware modules are communicatively connected to the semantic object-finding navigation system based on multi-sensor fusion: wireless radio frequency sensing all-in-one machine, red-green-blue depth RGBD stereo camera, microphone array, lidar and / or inertial measurement unit (IMU).
[0014] Fourthly, embodiments of this application provide a robot, including a memory, a transceiver, and a processor; A memory for storing computer programs; a transceiver for sending and receiving data under the control of the processor; and a processor for reading the computer programs from the memory and performing the following operations: In the process of constructing an environmental map, items are located and identified by integrating RFID sensing information and 3D visual sensing information to obtain the item's identification information, multi-semantic attribute information and spatial location information. The item's identification information, multi-semantic attribute information and spatial location information are then marked as points of interest in the environmental map. In response to a received item search command, the item search command is parsed to determine the target item, and the target point of interest corresponding to the target item is matched from the environmental map; Based on the location of the target point of interest, a path is planned, and the robot is controlled to move to the location of the target item.
[0015] Fifthly, embodiments of this application provide a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the semantic object-finding navigation method based on multi-sensor fusion as described in the first aspect.
[0016] In a sixth aspect, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the semantic object-finding navigation method based on multi-sensor fusion as described in the first aspect.
[0017] The semantic object-finding navigation method, system, device, and storage medium based on multi-sensor fusion provided in this application improves positioning accuracy and stability by integrating RFID and 3D visual information and automatically labeling items during the mapping process, and realizes high-precision autonomous object-finding navigation from object-finding command to accurate arrival at the target object location. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the semantic object-finding navigation method based on multi-sensor fusion provided in an embodiment of this application; Figure 2 This is a flowchart illustrating the method for obtaining points of interest in an environmental map provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the method for wireless radio frequency sensing and 3D vision fusion positioning and recognition provided in an embodiment of this application. Figure 4 This is a schematic diagram of the system architecture of the robot provided in the embodiments of this application; Figure 5 This is a system flowchart of the semantic object-finding navigation method based on multi-sensor fusion provided in the embodiments of this application; Figure 6 This is a schematic diagram of the structure of a semantic object-finding navigation system based on multi-sensor fusion provided in an embodiment of this application; Figure 7 This is a schematic diagram of the physical structure of the robot device provided in the embodiments of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] In scenarios such as smart warehousing and service robots, the core requirement is for robots to autonomously find and deliver designated items. Successful execution of this task depends on the accurate location and identification of items. Existing technologies generally employ Radio Frequency Identification (RFID) positioning technology, based on Received Signal Strength Indication (RSSI) for item location and identification.
[0022] Existing RFID positioning technology is susceptible to environmental interference and has poor positioning accuracy, making it difficult to reliably obtain the precise spatial location of items.
[0023] Figure 1 This is a flowchart illustrating the semantic object-finding navigation method based on multi-sensor fusion provided in an embodiment of this application. Figure 1 As shown, the method may include: Step 101: In the process of constructing the environmental map, the objects are located and identified by integrating RFID sensing information and 3D visual sensing information to obtain the object's identification information, multi-semantic attribute information and spatial location information. The object's identification information, multi-semantic attribute information and spatial location information are marked as points of interest in the environmental map.
[0024] It should be noted that this step overcomes the inherent limitations of single-sensor technologies by synergizing and fusing RFID sensing and 3D vision sensing. Specifically, RFID technology leverages its advantages of non-line-of-sight and batch identification capabilities for coarse target localization and EPC ID acquisition, while 3D vision technology provides precise spatial coordinate measurements. A data fusion algorithm combines the advantages of both, enabling stable output of centimeter-level spatial position accuracy for objects even in complex environments (such as object occlusion or metallic interference). This step is a prerequisite and foundation for achieving subsequent precise navigation.
[0025] Step 102: In response to the received item search instruction, parse the item search instruction to determine the target item, and match the target point of interest corresponding to the target item from the environment map.
[0026] It should be noted that this step, based on the semantic map established in step 101, which is rich in multi-semantic attributes and precise location information of items, transforms the user's natural language commands into specific navigation targets. Since step 101 has ensured the high accuracy and reliability of the semantic and location information of points of interest (POIs) in the map, the parsing and matching in this step is meaningful and can provide the robot planning system with an accurate and reliable navigation endpoint.
[0027] Step 103: Calculate the path based on the location of the target point of interest and control the robot to move to the location of the target item.
[0028] It should be noted that this step is the final execution stage of the task. It utilizes the target location determined in step 102, whose accuracy is guaranteed by step 101, to perform path planning and control, ultimately completing the object retrieval task. The precise spatial location information provided in step 101 is crucial for this step to accurately reach the target.
[0029] In summary, the semantic object-finding navigation method based on multi-sensor fusion provided in this embodiment, through the organic combination of steps 101 to 103, hinges on the fusion positioning scheme proposed in step 101. This scheme effectively solves the technical bottlenecks of poor RFID positioning accuracy and susceptibility to interference, providing a stable and accurate foundation of item location data for the entire object-finding navigation system. This method significantly improves the robot's ability to perceive and locate items in complex environments, thereby ensuring the accuracy and reliability of the final object-finding navigation task.
[0030] Figure 2 This is a flowchart illustrating the method for obtaining points of interest in an environmental map provided in an embodiment of this application. Figure 2As shown, in the process of constructing the environmental map, the identification information, multi-semantic attribute information, and spatial location information of objects are obtained by fusing RFID sensing information and 3D visual sensing information. These information are then marked as points of interest in the environmental map, including: Step 201: Acquire RFID sensing information and 3D visual sensing information.
[0031] It should be noted that this step is the data input stage of fusion sensing, which aims to simultaneously acquire raw data from sensors based on different physical principles, providing an information foundation for subsequent fusion processing.
[0032] For example, an RFID reader periodically transmits radio frequency signals at a set frequency (e.g., 100Hz), collects signals returned by activated RFID tags, and records information such as their EPC identification code, signal strength (RSSI), and phase. Simultaneously, a Red Green Blue Depth (RGBD) stereo camera synchronously acquires RGB images and depth information of the environment at a specific frame rate (e.g., 30Hz), and generates point cloud data containing the three-dimensional coordinates of object surfaces using triangulation principles.
[0033] Step 202: Synchronize and associate the RFID sensing information with the 3D visual sensing information in time, and perform fusion calculations based on the associated data to obtain the item's identification information, multi-semantic attribute information, and spatial location information.
[0034] It should be noted that this step is the core of the integrated positioning and identification process. First, time synchronization ensures the temporal consistency of the two types of data. Then, data association determines whether the RFID target and the 3D visual target correspond to the same physical entity. Finally, the successfully associated data is fused and calculated to obtain accurate and complete item information.
[0035] For example, RFID data streams and 3D vision data streams can be aligned using hardware triggering or software timestamps. Subsequently, association is performed based on a rough spatial consistency principle. For instance, the approximate location of an object estimated by RFID based on Phase Difference of Arrival (PDOA) (e.g., 2-3 meters in front of the robot, 30 degrees to the left) is spatially matched with the 3D bounding boxes of candidate objects identified in the point cloud by 3D vision using object detection algorithms (such as YOLO). This determines which visually detected target corresponds to which RFID tag.
[0036] For successfully associated items, their unique Electronic Product Code (EPC) is obtained from the RFID information, and the item's name, category, and other attributes are retrieved from the database. Simultaneously, the item's precise three-dimensional coordinates and orientation (i.e., spatial location information) are directly extracted from the 3D vision point cloud data. For items with similar appearances but different markings, the RFID identification information provides crucial discrimination criteria for visual recognition, avoiding misidentification.
[0037] Step 203: Bind the item's identification information, multi-semantic attribute information, and spatial location information to the coordinates of the environment map to obtain points of interest.
[0038] It should be noted that this step "anchors" the semantic information of the items obtained in the previous step to the global map, completing the semantic upgrade of the environment map, so that it is upgraded from containing only geometric information to containing both geometric and semantic information.
[0039] For example, the spatial location information of the item obtained through fusion calculation (initially in the camera coordinate system) is first transformed to the robot's base coordinate system using a fixed transformation matrix (TransForm, TF) pre-calibrated by the robot system, and then transformed to the world coordinate system used by the environment map constructed by the Simultaneous Localization and Mapping (SLAM) algorithm. Subsequently, in the world coordinate system, the item's identification information, multi-semantic attribute information, and its precise three-dimensional coordinates (x, y, z) are stored as a structured data entry (i.e., Point of Interest, POI) in the map database and associated with the corresponding grid in the raster map. For example, at the coordinates (x=10.5, y=3.2) of the warehouse map, a POI is recorded with the content "Item E, Category: Router, Attributes: Black, Manufacturer, 4 ports".
[0040] This embodiment integrates the global identification capability of RFID with the local precise ranging capability of 3D vision within a unified spatiotemporal framework through a standardized data flow, ultimately achieving automated and highly reliable binding of item identification with centimeter-level precision location information, providing core data for the construction of semantic maps.
[0041] Furthermore, the RFID sensing information includes RFID signals, and the 3D visual sensing information includes 3D visual point cloud data. The RFID sensing information and 3D visual sensing information are synchronized in time and associated with each other. Based on the associated data, fusion calculations are performed to obtain the item's identification information, multi-semantic attribute information, and spatial location information. This includes: filtering and smoothing the RFID signal to obtain signal strength and phase information, and estimating the approximate location of the item based on the signal strength and phase information to obtain RFID identification information; filtering and target detection preprocessing of the 3D visual point cloud data to obtain item features, identifying the item category and three-dimensional location based on the item features to obtain 3D visual appearance features; associating and matching the RFID identification information with the 3D visual appearance features, and querying the item attribute database based on the RFID identification information to obtain multi-semantic attribute information, thereby comprehensively obtaining the item's identification information, multi-semantic attribute information, and spatial location information.
[0042] Understandably, the fusion of radio frequency identification (RFID) and 3D vision enables precise object positioning and comprehensive identification, including innovative fusion algorithms and multi-source data joint processing methods, which is the core of improving positioning and identification accuracy.
[0043] Figure 3 This is a flowchart illustrating the method for wireless radio frequency sensing and 3D vision fusion positioning and recognition provided in an embodiment of this application. Figure 3 As shown, the steps for obtaining an item's identification information, multiple semantic attribute information, and spatial location information may include: 1) RFID Data Acquisition and Preprocessing: The RFID reader periodically transmits radio frequency signals at a set frequency (e.g., 100Hz) to activate surrounding RFID tags. Upon receiving the tag's reflected signal, it quickly acquires data such as signal strength (RSSI) and phase. To improve data quality, each acquired data is sampled multiple times (e.g., 50 times), and median filtering is used to remove burst noise. A Kalman filter algorithm is then used to smooth the data and reduce signal fluctuations. Signal processing algorithms, such as Phase Difference Positioning Algorithm (PDOA), are used to initially estimate the approximate position of the item relative to the robot. The processed data is then transmitted to the fusion processing unit.
[0044] 2) 3D Vision Data Acquisition and Processing: The 3D camera continuously acquires image information of the environment and objects at a set frame rate. First, the 3D camera projects a specific coded pattern onto the object surface using structured light, and the camera captures images of the pattern deformation from different angles. Using the principle of triangulation, the three-dimensional coordinates of each point on the object surface are calculated to generate 3D point cloud data. The 3D point cloud data is preprocessed by removing outliers through voxelization filtering and smoothing it using bilateral filtering to reduce noise interference. Then, a deep learning-based target detection algorithm (such as the improved YOLO11 algorithm) is used to process the red-green-blue (RGB) image data to identify the target object and obtain its precise three-dimensional position and pose information. Simultaneously, the appearance features of the object, such as shape, color, and texture feature vectors, are extracted, and the processed 3D vision data is transmitted to the fusion processing unit.
[0045] 3) Data Fusion and Identification: In the fusion processing unit, RFID data and 3D visual data are first synchronized in time to ensure consistency. RFID tag identification information (i.e., EPC identifiers) is extracted from the RFID data, and the appearance feature vectors of the objects are extracted from the 3D visual data. A data association model is constructed using deep neural networks (such as Siamese networks). The RFID tag identification information and 3D visual appearance feature vectors are used as inputs to train the model, enabling it to accurately match the relationship between the two, achieving precise object location and comprehensive identification. For example, by learning the correspondence between a large number of object RFID tags and 3D visual features, the model can accurately determine the detailed information of the target object, even distinguishing between objects with similar appearances.
[0046] This application innovatively integrates radio frequency sensing (RFID) with 3D vision to achieve precise positioning. On one hand, RFID technology can quickly detect target item tags and obtain their approximate location range, even when the item is partially obscured. On the other hand, 3D vision technology can accurately acquire the three-dimensional spatial coordinates of the item. By establishing a fusion algorithm between the two, multi-source data is processed jointly. For example, the distance to the item is initially determined by RFID signal strength, and then the point cloud data from 3D vision is used to accurately determine the item's position in space, effectively overcoming the limitations of single technologies and significantly improving positioning accuracy. Compared to the meter-level positioning error of traditional RSSI-based RFID and the instability of 3D vision positioning in complex situations, the fused positioning accuracy can be improved to the centimeter level, accurately determining the location of goods in complex warehouse environments.
[0047] This method not only obtains unique identification information (EPC code) for items from RFID tags and queries the database to obtain multiple semantic attributes such as serial number and category, but also extracts the appearance features of the items using 3D vision. Deep learning algorithms such as PointNet and PointNet++ are used to process the 3D vision data to identify features such as shape, color, and texture of the items. The identification information provided by RFID is correlated and matched with the appearance features extracted by 3D vision, thereby achieving comprehensive and accurate identification of items. For example, for mobile phones of different brands that look very similar, the brand and model information obtained through RFID tags can be accurately distinguished by combining it with subtle appearance differences identified by 3D vision. This fusion identification method greatly improves the accuracy and reliability of identification, avoiding the misjudgments that may occur with single identification methods.
[0048] In some embodiments, before binding the item's identification information, multi-semantic attribute information, and spatial location information to the coordinates of the environmental map to obtain points of interest, the method further includes: The spatial location information of the objects is transformed from their respective sensor coordinate systems to the robot's base coordinate system, and then uniformly transformed to the world coordinate system of the environmental map.
[0049] It should be noted that this embodiment addresses a core challenge in multi-sensor fusion systems—the spatial uniformity of data. RFID antennas and 3D cameras, as independent sensors, each possess their own local coordinate system (sensor coordinate system). The object location information they perceive is initially expressed in their respective coordinate systems and cannot be directly used for the global map. This step introduces the robot's base coordinate system as a unified intermediate reference system and utilizes a fixed transformation matrix (TF) pre-obtained through techniques such as hand-eye calibration to first uniformly transform all the local observation data from the sensors (such as the coarse orientation of RFID signals and the precise coordinates in the 3D camera point cloud) to the robot's body coordinate system. Then, combined with the robot's real-time pose in the global map (world coordinate system) constructed by SLAM, this location information is finally transformed into a unique and stable world coordinate system. This process ensures that the positions of objects perceived from different sources and at different times can be accurately and unambiguously registered on the same global environment map.
[0050] This embodiment eliminates spatial errors caused by different installation positions and observation angles of heterogeneous sensors by establishing a standardized coordinate system transformation chain, providing a stable and unified global spatial reference for the RFID and vision fusion positioning results, thereby ensuring the accuracy and consistency of the coordinates of points of interest (POIs) subsequently marked on the map.
[0051] In some embodiments, the identification information, multiple semantic attribute information, and spatial location information of an item are bound to the coordinates of the environment map to obtain points of interest, including: in the process of constructing a raster map based on the SLAM algorithm, the identification information, multiple semantic attribute information, and spatial location information of the identified items are converted into point of interest data in real time; the point of interest data is stored in a map database and associated with the coordinates of the environment map.
[0052] It should be noted that the method of automatically labeling Points of Interest (POIs) based on object location and recognition information during robot mapping involves the collaborative work of map building algorithms and automatic labeling modules, which is key to improving mapping efficiency and accuracy. Specifically, the methods for autonomous robot mapping and automatic POI labeling can include: 1) Autonomous Map Building: During robot movement, 3D vision cameras or LiDAR acquire real-time 3D point cloud data of the environment, while LiDAR simultaneously obtains precise geometric information. Based on the Simultaneous Localization and Mapping (SLAM) algorithm, the acquired data is processed. Through feature extraction, matching, and optimization, a grid map in the robot's world coordinate system is constructed. During map building, the rich environmental information provided by 3D vision, such as the position and shape of objects like walls, pillars, and shelves, is fully utilized to continuously update and optimize the grid map. Simultaneously, pose and position information provided by the IMU and encoder are combined to improve the accuracy and stability of the grid map construction. During map building, the grid map data is stored in real-time on the robot's local storage device and uploaded to the server for backup and sharing via a wireless communication module.
[0053] 2) Automatic POI Labeling on Grid Maps: After the robot determines the location and identification of items using fusion technology, automatic labeling begins. This module receives item location and identification information from the fusion processing unit, converting the item's name, category, and precise coordinates into POI data format. Then, it searches the map database for the corresponding location and adds the POI information to the map. For example, when mapping a logistics warehouse, once the robot identifies and locates goods on a shelf, it automatically labels the goods' information as the corresponding POI on the map. Simultaneously, a unique identifier is added to each POI for easy subsequent querying and management. After automatic labeling is complete, the navigation grid map update module promptly synchronizes the updated map data to local storage and the server to ensure map information consistency.
[0054] During operation, the mobile robot continuously collects 3D point cloud data of the surrounding environment using 3D vision sensors, and constructs a map in the robot's world coordinate system using Simultaneous Localization and Mapping (SLAM) algorithms. In the map construction process, it fully utilizes the rich environmental geometric information provided by 3D vision (LiDAR), such as the positions of walls and obstacles, to make the map more accurate and detailed. Unlike traditional mapping methods, this application closely integrates the localization and identification information of objects during the mapping process. Once the robot determines the location and identification of an item using fusion technology, the system automatically labels relevant item information, such as name, category, and precise coordinates, as Points of Interest (POIs) on the map. Specifically, an automatic labeling module is embedded in the map-building algorithm. This module receives the item's location and identification results in real time and converts them into POI information, adding it to the map database. For example, when mapping a warehouse, once the robot identifies and locates goods on a shelf, it automatically labels the goods' information as the corresponding POI on the map. This automatic labeling method greatly improves the efficiency and accuracy of map building, reduces the workload and errors of manual labeling, and provides rich and accurate map information for subsequent item navigation.
[0055] In some embodiments, parsing the item search instruction to determine the target item and matching the target point of interest corresponding to the target item from the environmental map includes: parsing the item search instruction using a natural language processing model to obtain multi-semantic attribute information of the target item, wherein the natural language processing model is fine-tuned based on the item attribute database; and matching the multi-semantic attribute information of the target item with the identification information, multi-semantic attribute information and spatial location information of the items associated with each point of interest in the environmental map to determine the target point of interest.
[0056] It should be noted that, to enable the robot to understand voice commands containing object points of interest (POIs), a semantic understanding model based on Natural Language Processing (NLP) was constructed. The model was trained by collecting a large amount of natural language description data related to objects, including their attributes (such as color, shape, and size), location (such as on a table or in a cabinet), and purpose. The model employs a deep learning architecture, such as Recurrent Neural Networks (RNNs) or their variants Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRUs), combined with an attention mechanism to better understand complex semantic relationships. For example, for the command "Find the round white teacup on the third shelf of the blue cabinet," the model can accurately parse the key information such as the object's color, shape, and location.
[0057] The semantic understanding model is trained using a large RFID-based object database that provides natural language descriptions of object attributes, including extended attribute information such as object attributes, location, and purpose in different scenarios. Data preprocessing includes word segmentation, part-of-speech tagging, and named entity recognition. A deep learning architecture, such as a Transformer-based pre-trained model, is employed for fine-tuning on a large-scale corpus of object attributes. Then, for the application scenario described in this application, the pre-trained model is fine-tuned using labeled object semantic data. During fine-tuning, an attention mechanism is used to enable the model to better capture semantic relationships. For example, for the instruction "find the black pen in the red box," the model can accurately parse key information such as the object's color, shape, and location. During training, the cross-entropy loss function is used as the optimization objective, and stochastic gradient descent or its variants are employed for parameter updates to continuously improve the semantic understanding and intent understanding capabilities of the semantic model.
[0058] This embodiment relies on a high-quality item attribute database, which is constructed and enriched by the aforementioned fusion of RFID and visual recognition. RFID provides a unique item identifier (EPC ID) and core attributes, while visual recognition supplements these with appearance attributes such as color and shape. Based on this, a pre-trained language model undergoes domain-adaptive fine-tuning, making it particularly adept at understanding item attributes and spatial relationships within this application scenario. During the matching phase, the system does not perform simple keyword matching but instead executes a multi-attribute weighted similarity calculation, comparing the parsed semantic information with entries in the map POI database to accurately locate the user-specified target among multiple visually similar items.
[0059] This embodiment achieves a reliable conversion from fuzzy and complex natural language instructions to precise and unique target points of interest in the map by constructing a domain-specific semantic understanding model and combining it with a high-precision semantic map, thus bridging the key gap between natural interaction and precise navigation.
[0060] In some embodiments, path planning is performed based on the location of the target point of interest, and the robot is controlled to move to the location of the target item. This includes: planning an optimal path to the location of the target item based on the robot's current location and the location of the target item's point of interest; controlling the robot to move along the optimal path; and using onboard sensors to perceive the environment in real time during the movement to avoid obstacles.
[0061] It's important to note that the constructed semantic understanding model based on natural language processing, along with the path planning and navigation mechanism combined with a map-based POI database, are key to achieving intelligent object-finding navigation. After receiving a user's voice command, the voice interaction module converts the voice signal into text information and transmits it to the semantic understanding model. The semantic understanding model parses the text command, understands the user's intent, and extracts key information such as the target item's location and multiple semantic attributes. Combined with the POI database already marked on the map, it determines the target item's map coordinates. Then, using path planning algorithms such as TED / DWA / A*, it plans the optimal path based on the robot's current position and the target item's location. During path planning, factors such as obstacle information and passage width on the map are considered to ensure the planned path is feasible and efficient. After path planning is complete, the path information is transmitted to the navigation control module.
[0062] The navigation control module controls the robot to move along the planned path based on the path planning results. During movement, the robot uses sensors such as LiDAR and cameras to perceive the surrounding environment in real time, accurately locate its current position, and detect the presence of obstacles. When an obstacle is detected, the robot automatically replans its current direction and speed based on a local path planning algorithm to avoid obstacles. This ensures that the robot can safely and accurately reach the location of the target item. Simultaneously, the robot uses a voice interaction module to calculate the progress of its search and navigation information in real time and provides feedback to the user, such as "Heading to the target item location, estimated arrival time is 1 minute."
[0063] After receiving a voice command, the robot's semantic understanding model parses the command and extracts the Point of Interest (POI) information of the target item. Combined with a database of marked POIs on a map, the robot determines the target item's location. Then, a path planning algorithm, such as A* or Dijkstra's algorithm, is used to plan the optimal path based on the robot's current location and the target item's location. During navigation, the robot uses its onboard sensors, such as LiDAR and cameras, to perceive its surroundings in real time, avoiding obstacles and ensuring accurate arrival at the target item's location. When encountering path obstructions, it can promptly replan its path and continue moving towards the target. This object-based semantic navigation method allows users to interact with the robot in a more natural and flexible way, significantly improving the robot's intelligence in object finding and enhancing the user experience.
[0064] This embodiment combines the target location obtained from semantic parsing with hierarchical path planning (global planning and local replanning) and real-time environmental perception, ensuring that the robot can safely and reliably perform navigation tasks in dynamically changing and complex environments. Ultimately, it transforms accurate semantic understanding into equally accurate physical arrival, completing the "last mile" of intelligent object finding.
[0065] Through the above improvements, this application comprehensively enhances the robot's performance in precise item positioning, identification, map construction and annotation, and item navigation, meeting the intelligent application needs of various complex scenarios. The advantages of this application's embodiments include: 1. Improved positioning and identification accuracy. 1) Precise positioning: By deeply integrating radio frequency identification (RFID) with 3D vision, the limitations of single technologies in positioning are overcome. RFID can quickly detect the approximate location of an item, while 3D vision can accurately determine its spatial coordinates. The combination of the two improves positioning accuracy from the meter-level error of traditional single technologies to the centimeter level. In warehousing and logistics scenarios, this helps the robot accurately locate goods in specific locations on shelves, greatly reducing search time and improving the efficiency of goods entering and leaving the warehouse. For example, in a large multi-layered warehouse, relying on a single technology to locate goods previously might have led the robot to search for the wrong location multiple times, especially for obscured items (not at line-of-sight). However, with the technology of this application, the robot can accurately locate goods in one go, improving the efficiency of item retrieval. 2) Accurate identification: By integrating RFID identification information with appearance features extracted by 3D vision, comprehensive and accurate identification of items is achieved. Compared to traditional single identification methods, this significantly improves the identification accuracy and reduces the false judgment rate. For example, in an electronics warehouse, different models of electronic products may look very similar. Relying solely on 3D vision or RFID technology can easily lead to misjudgment. However, this method combines the two to accurately distinguish between different models of products, thus improving the accuracy of identification.
[0066] 2. Efficient Mapping and Synchronous Accurate Labeling. 1) Efficient Mapping: The robot map building process closely integrates item location and identification information, utilizing rich environmental data acquired through 3D vision and SLAM algorithms to achieve more accurate and detailed map construction. Compared to traditional mapping methods, this reduces the time required for map building and the length of the movement path. Traditional mapping methods may require the robot to traverse the scene environment to build a complete map before manually labeling POIs on the map based on the actual situation. This technology allows the robot to build a map more efficiently based on item information during a single traversal and automatically and synchronously label POIs. 2) Automatic and Accurate POI Labeling: Automatic POI labeling greatly improves the efficiency and accuracy of map building. It avoids the tedious process of manual labeling and potential errors, while also updating map information in real time. In large warehouse environments, manual POI labeling is not only time-consuming and labor-intensive, but also difficult to update promptly after goods locations change. This automatic labeling function can label new items or items with changed locations on the map in real time, ensuring the map remains up-to-date and accurate, providing a reliable foundation for item navigation.
[0067] 3. Enhanced Intelligent Item Finding and Navigation. 1) Semantic Understanding: The constructed semantic understanding model based on natural language processing enables the robot to understand voice commands containing rich semantics. Users can describe various attributes and location information of items using natural language, allowing the robot to accurately find the items, greatly improving the convenience and intelligence of the interaction. Compared with traditional robots that can only recognize simple commands, the robot in this application can handle more complex commands, such as "find the black 10 Gigabit router from a certain manufacturer on the table," meeting the diverse item finding needs of users. 2) Path Planning and Navigation: Combining the target item location information parsed from semantic understanding with the POI database marked on the map for path planning, a better path can be planned. During navigation, the robot perceives the environment in real time and dynamically adjusts the path to ensure accurate arrival at the target item location. This makes item finding and navigation more efficient and reliable, improving the robot's ability to work in complex environments. For example, in office areas with frequent personnel and item traffic, the robot can adjust the path in a timely manner according to environmental changes, quickly finding the target item, increasing the item finding success rate from about 70% to over 90%.
[0068] In some embodiments, a robot includes: a robot body; a semantic object-finding navigation system based on multi-sensor fusion, integrated into the main control module of the robot body; and the following hardware modules communicatively connected to the semantic object-finding navigation system based on multi-sensor fusion: a wireless radio frequency sensing all-in-one machine, a red-green-blue depth RGBD stereo camera, a microphone array, a lidar and / or an inertial measurement unit (IMU).
[0069] Figure 4 This is a schematic diagram of the system architecture of the robot provided in an embodiment of this application. Figure 4 As shown, this embodiment consists of a robot body equipped with a wireless radio frequency sensing integrated machine (RFID / UWB reading device), a microphone array, a depth stereo (RGBD) camera, and a robot main control module. The robot main control module includes software modules such as a wireless radio frequency-vision fusion positioning module, a mapping and positioning module based on laser or vision or laser-vision fusion, a embodied navigation module, a voice interaction module, and item semantic mapping. It can be applied to scenarios such as intelligent warehousing (item sorting) and service robots (item delivery).
[0070] The key hardware modules include: 1) Mobile Robot Body: A wheeled mobile robot with precise mobility, capable of traversing freely in various indoor and outdoor environments. The robot can simultaneously handle complex tasks such as radio frequency sensing, 3D vision perception, object recognition, and semantic understanding. It is equipped with large-capacity memory and high-speed storage devices to ensure rapid data read / write and processing.
[0071] 2) Wireless Radio Frequency Sensing Integrated Machine: Multiple ultra-high frequency (UHF) RFID readers are strategically deployed on the robot to ensure comprehensive signal coverage. This machine is used for identification (EPCID) and coarse positioning of items with RFID tags.
[0072] 3) RGBD Stereo Camera: An RGBD stereo camera is used as a visual sensor to capture real-time 3D information about the environment and objects. The 3D camera is mounted at a corresponding position on the front end of the robot or robotic arm, ensuring a wide and unobstructed field of view. This allows for the identification of object categories within the effective field of view of the stereo camera.
[0073] 4) Microphone Array: Installing microphone arrays (several microphone arrays) on the robot allows for accurate acquisition of user voice commands from any direction. Acquiring voice commands (such as "Please take item E") provides input for voice interaction.
[0074] The robot's main control software acts as its "brain," running various algorithm modules, processing sensor data, and making decisions. The robot's main control software includes: 1) Embodied navigation: Combining SLAM mapping and path planning, it enables robots to autonomously avoid obstacles and move in the environment, such as navigating from the starting point to the location of an object; 2) Voice interaction module: Analyzes the voice collected by the microphone and understands human commands (such as recognizing the semantics of "take item E"); 3) Mapping and localization: Construct an environmental grid map using laser / visual SLAM, and simultaneously determine the robot's own position within the map; 4) Item semantic mapping: Associate the items identified by the sensor (such as “E item”) with POIs (points of interest) in the map, so that “items” become searchable semantic information in the map.
[0075] 5) SLAM grid map (POI): Stores an environment map and marks the location of items (both labeled and unlabeled items are marked as POIs), which is the basis for the robot to "recognize" the environment.
[0076] 6) Passive-vision fusion positioning: It integrates data from RFID (passive) and RGBD camera (vision) to accurately locate objects and make up for the shortcomings of a single sensor (such as RFID positioning being coarse and vision being easily obstructed, while fusion is more stable and accurate).
[0077] Figure 5 This is a system flowchart of the semantic object-finding navigation method based on multi-sensor fusion provided in an embodiment of this application. Figure 5As shown, this describes the robot's journey from initial environmental mapping to precise object localization and automatic point-of-interest (POI) labeling based on radio frequency-vision fusion, enabling the robot to navigate directly to the target object location based on voice commands. The specific steps are as follows: 1. Robot Motion Mapping Initialization The robot starts from the charging station (the origin of the world coordinate system) and scans the scene environment based on laser SLAM or vision vSLAM to build a map and localize itself. During the map scanning process, the robot identifies and locates target objects. 2. RFID-Vision Fusion for Item Location and Identification By enabling the fusion perception of "RFID radio frequency + 3D vision", the robot determines the object identification method. RFID radio frequency identifies and roughly locates the target object (line-of-sight and non-line-of-sight) EPC ID, while 3D vision accurately identifies and classifies the objects within the robot's current field of vision.
[0078] 3. Determine if the item has an RFID tag. For unlabeled items within the visual field, the system detects and identifies the item category (such as cups and snacks) and its position relative to the camera coordinate system based on 3D visual images. For tagged items, the EPC ID of the tag is read by RFID radio frequency, and the item name and category are retrieved from the database. That is, the unique EPC ID of the tag is read by RFID, and the item details are obtained by association with the database (such as the attribute information corresponding to the ID including "black|manufacturer|4 ports|10 Gigabit|router|manufacturing date").
[0079] 4. Transformation of passive sensing device (antenna) / camera coordinate system to world coordinate system to raster map coordinate system.
[0080] The object positions perceived and identified by RFID (passive device coordinate system) and vision (camera coordinate system) are unified to the robot's base coordinate system based on the robot's calibrated passive device fixed transformation matrix TF and the calibrated 3D camera fixed transformation matrix, and then converted to the world coordinate system (robot global reference system). The object positions are then adapted to the grid coordinates of the laser SLAM grid map, so that the object positions are "anchored" to the global map. 5. During the robot's pre-scanning and environmental mapping process (laser SLAM or visual SLAM), POIs are automatically labeled on the grid map simultaneously based on object localization. The robot uses laser SLAM or visual SLAM to build an environmental grid map in real time, and marks the locations of identified objects as "points of interest (POIs)," such as "a black object at (x,y) coordinates | a certain manufacturer | 4 ports | 10 Gigabit | router" on the grid map, in preparation for subsequent voice command matching; 6. Through natural language interaction commands, based on the text command understanding capability of a large language model, the semantic attributes of items are mapped to Points of Interest (POIs). Receive voice / text commands (such as "black|manufacturer|4 ports|10 Gigabit|router|manufacturing date"), parse the semantics (understand that "black router" is the target), match POIs in the raster map, and determine the specific navigation target location.
[0081] 7. The robot's path planning allows it to precisely move to the vicinity of the target item. Based on the coordinates of the target POI in the grid map and combined with the obstacle information of the SLAM map, the robot automatically plans an obstacle avoidance path and controls the robot to move next to the object, completing the task loop of "autonomous object finding".
[0082] In summary, the embodiments of this application achieve EPCID identification and precise positioning of tagged items and query of attribute information through the fusion of "RFID identification + visual recognition". Non-tagged items are classified, identified and precisely positioned through 3D vision. The coordinates of the tagged items and non-tagged items are converted to map coordinates. While the robot is autonomously building a map, the location of the items is automatically marked on the map. Then, combined with a semantic understanding model fine-tuned by a large amount of semantic data of the items, the intent used for speech is understood and the location coordinates of the target object on the map are obtained. Based on the target location, the navigation path is planned. The robot autonomously moves and navigates to the vicinity of the target item and can calculate the item finding progress in real time to complete the semantic item finding task. Typical applications include intelligent warehousing (goods finding) and service robots (item delivery).
[0083] The advantages of this technology make it suitable for various complex scenarios, such as intelligent warehousing, logistics distribution, and indoor service robots. In intelligent warehousing, it improves the automation and intelligence of goods management; in logistics distribution, it helps to quickly locate and sort goods.
[0084] In the field of indoor service robots, they can better meet users' diverse service needs, such as helping users find specific items.
[0085] Compared to traditional technologies, it expands the application scenarios and scope of robots, has strong market potential, and high commercial value.
[0086] The semantic object-finding navigation system based on multi-sensor fusion provided in the embodiments of this application is described below. The semantic object-finding navigation system based on multi-sensor fusion described below can be referred to in correspondence with the semantic object-finding navigation method based on multi-sensor fusion described above.
[0087] Figure 6This is a schematic diagram of the structure of a semantic object-finding navigation system based on multi-sensor fusion provided in an embodiment of this application. Figure 6 As shown, the semantic object-finding navigation system 300 based on multi-sensor fusion includes: The mapping and annotation module 310 is used to locate and identify objects by integrating RFID sensing information and 3D visual sensing information during the construction of an environmental map, thereby obtaining the object's identification information, multi-semantic attribute information and spatial location information, and annotating the object's identification information, multi-semantic attribute information and spatial location information as points of interest in the environmental map. The instruction parsing and matching module 320 is used to respond to the received item-finding instruction, parse the item-finding instruction to determine the target item, and match the target point of interest corresponding to the target item from the environmental map; The path planning and navigation module 330 is used to plan a path based on the location of the target point of interest and control the robot to move to the location of the target item.
[0088] In some embodiments, the mapping and annotation module 310 includes: An information acquisition unit is used to acquire the RFID sensing information and the 3D visual sensing information; The information processing unit is used to synchronize the RFID sensing information and the 3D visual sensing information in time and associate them with data, and perform fusion calculations based on the associated data to obtain the identification information, multiple semantic attribute information and spatial location information of the item. The information binding unit is used to bind the identification information, multiple semantic attribute information and spatial location information of the item with the coordinates of the environmental map to obtain the point of interest.
[0089] In some embodiments, the RFID sensing information includes RFID signals, the 3D visual sensing information includes 3D visual point cloud data, and the information processing unit is specifically used for: The RFID signal is filtered and smoothed to obtain signal strength and phase information. The approximate location of the item is estimated based on the signal strength and phase information. At the same time, the electronic product code (EPC) identifier of the RFID tag is decoded from the RFID signal as the identification information of the item. The 3D visual point cloud data is filtered and preprocessed for target detection to extract the appearance features of the item, and the three-dimensional position of the item is identified based on the appearance features. Based on the identification information of the item, query the item attribute database to obtain multi-semantic attribute information corresponding to the item; The identification information, the multiple semantic attribute information, the appearance features, and the three-dimensional position of the item are associated and fused to obtain the identification information, multiple semantic attribute information, and spatial position information of the item.
[0090] In some embodiments, the mapping and annotation module 310 further includes: The coordinate transformation unit is used to transform the spatial position information of the objects from their respective sensor coordinate systems to the robot's base coordinate system, and then to the world coordinate system of the environmental map.
[0091] In some embodiments, the information binding unit is specifically used for: During the process of constructing a grid map based on the SLAM algorithm, the identification information, multi-semantic attribute information and spatial location information of the identified items are converted into point of interest data in real time. The points of interest data are stored in a map database and associated with the coordinates of the environment map.
[0092] In some embodiments, the instruction parsing and matching module 320 includes: The instruction parsing unit is used to parse the item search instruction using a natural language processing model to obtain multi-semantic attribute information of the target item. The natural language processing model is fine-tuned based on the item attribute database. The point of interest matching unit is used to match the multi-semantic attribute information of the target item with the identification information, multi-semantic attribute information and spatial location information of the items associated with each point of interest in the environmental map to determine the target point of interest.
[0093] In some embodiments, the path planning and navigation module 330 includes: The path planning unit is used to plan the optimal path to the location of the target item based on the robot's current position and the location of the point of interest of the target item. The control and movement unit is used to control the robot to move along the optimal path and to use onboard sensors to perceive the environment in real time during the movement in order to avoid obstacles.
[0094] Figure 7 This is a schematic diagram of the physical structure of the robot device provided in the embodiments of this application. For example... Figure 7As shown, the robot device 400 may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. The processor 410, communication interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call a computer program stored in the memory 430 to execute steps of a semantic object-finding navigation method based on multi-sensor fusion, such as: In the process of constructing an environmental map, items are located and identified by integrating RFID sensing information and 3D visual sensing information to obtain the item's identification information, multi-semantic attribute information and spatial location information. The item's identification information, multi-semantic attribute information and spatial location information are then marked as points of interest in the environmental map. In response to a received item search command, the item search command is parsed to determine the target item, and the target point of interest corresponding to the target item is matched from the environmental map; Based on the location of the target point of interest, a path is planned, and the robot is controlled to move to the location of the target item.
[0095] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0096] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of the semantic object-finding navigation method based on multi-sensor fusion provided in the above embodiments, such as including: In the process of constructing an environmental map, items are located and identified by integrating RFID sensing information and 3D visual sensing information to obtain the item's identification information, multi-semantic attribute information and spatial location information. The item's identification information, multi-semantic attribute information and spatial location information are then marked as points of interest in the environmental map. In response to a received item search command, the item search command is parsed to determine the target item, and the target point of interest corresponding to the target item is matched from the environmental map; Based on the location of the target point of interest, a path is planned, and the robot is controlled to move to the location of the target item.
[0097] On the other hand, embodiments of this application also provide a processor-readable storage medium storing a computer program for causing a processor to execute the steps of the semantic object-finding navigation method based on multi-sensor fusion provided in the above embodiments, such as including: In the process of constructing an environmental map, items are located and identified by integrating RFID sensing information and 3D visual sensing information to obtain the item's identification information, multi-semantic attribute information and spatial location information. The item's identification information, multi-semantic attribute information and spatial location information are then marked as points of interest in the environmental map. In response to a received item search command, the item search command is parsed to determine the target item, and the target point of interest corresponding to the target item is matched from the environmental map; Based on the location of the target point of interest, a path is planned, and the robot is controlled to move to the location of the target item.
[0098] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).
[0099] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0100] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A semantic object-finding navigation method based on multi-sensor fusion, characterized in that, include: In the process of constructing an environmental map, items are located and identified by integrating RFID sensing information and 3D visual sensing information to obtain the item's identification information, multi-semantic attribute information and spatial location information. The item's identification information, multi-semantic attribute information and spatial location information are then marked as points of interest in the environmental map. In response to a received item search command, the item search command is parsed to determine the target item, and the target point of interest corresponding to the target item is matched from the environmental map; Based on the location of the target point of interest, a path is planned, and the robot is controlled to move to the location of the target item.
2. The semantic object-finding navigation method based on multi-sensor fusion according to claim 1, characterized in that, In the process of constructing the environmental map, the identification information, multi-semantic attribute information, and spatial location information of objects are obtained by fusing RFID sensing information and 3D visual sensing information. These object identification information, multi-semantic attribute information, and spatial location information are then marked as points of interest in the environmental map, including: Acquire the RFID sensing information and the 3D visual sensing information; The RFID sensing information and the 3D visual sensing information are synchronized in time and associated with each other. Based on the associated data, a fusion calculation is performed to obtain the identification information, multiple semantic attribute information and spatial location information of the item. The identification information, multiple semantic attribute information, and spatial location information of the item are bound to the coordinates of the environmental map to obtain the point of interest.
3. The semantic object-finding navigation method based on multi-sensor fusion according to claim 2, characterized in that, The RFID sensing information includes RFID signals, and the 3D visual sensing information includes 3D visual point cloud data. The process involves synchronizing the RFID sensing information and the 3D visual sensing information in time and associating them with each other, then performing fusion calculations based on the associative data to obtain the item's identification information, multiple semantic attribute information, and spatial location information, including: The RFID signal is filtered and smoothed to obtain signal strength and phase information. The approximate location of the item is estimated based on the signal strength and phase information. At the same time, the electronic product code (EPC) identifier of the RFID tag is decoded from the RFID signal as the identification information of the item. The 3D visual point cloud data is filtered and preprocessed for target detection to extract the appearance features of the item, and the three-dimensional position of the item is identified based on the appearance features. Based on the identification information of the item, query the item attribute database to obtain multi-semantic attribute information corresponding to the item; Based on the item's identification information, the multi-semantic attribute information, and the appearance features, association matching is performed, and the rough position and the three-dimensional position are fused to obtain the item's identification information, multi-semantic attribute information, and spatial position information.
4. The semantic object-finding navigation method based on multi-sensor fusion according to claim 2, before binding the object's identification information, multiple semantic attribute information, and spatial location information with the coordinates of the environmental map to obtain the point of interest, the method further includes: The spatial location information of the items is transformed from their respective sensor coordinate systems to the robot's base coordinate system, and then uniformly transformed to the world coordinate system of the environmental map.
5. The semantic object-finding navigation method based on multi-sensor fusion according to claim 4, characterized in that, The step of binding the item's identification information, multi-semantic attribute information, and spatial location information with the coordinates of the environmental map to obtain the point of interest includes: During the process of constructing a grid map based on the SLAM algorithm, the identification information, multi-semantic attribute information and spatial location information of the identified items are converted into point of interest data in real time. The points of interest data are stored in a map database and associated with the coordinates of the environment map.
6. The semantic object-finding navigation method based on multi-sensor fusion according to claim 1, characterized in that, The process of parsing the item search command to determine the target item and matching the target point of interest corresponding to the target item from the environmental map includes: The natural language processing model is used to parse the item search command to obtain multi-semantic attribute information of the target item. The natural language processing model is fine-tuned based on the item attribute database. The target point of interest is determined by matching the multi-semantic attribute information of the target item with the identification information, multi-semantic attribute information, and spatial location information of the items associated with each point of interest in the environmental map.
7. The semantic object-finding navigation method based on multi-sensor fusion according to claim 1, characterized in that, The step of planning a path based on the location of the target point of interest and controlling the robot to move to the location of the target item includes: Based on the robot's current position and the location of the target item's point of interest, the optimal path to reach the target item's location is planned; The robot is controlled to move along the optimal path, and during the movement, it uses onboard sensors to perceive the environment in real time to avoid obstacles.
8. A semantic object-finding navigation system based on multi-sensor fusion, characterized in that, include: The mapping and annotation module is used to locate and identify objects by integrating RFID sensing information and 3D visual sensing information during the construction of an environmental map, thereby obtaining the object's identification information, multi-semantic attribute information and spatial location information, and annotating the object's identification information, multi-semantic attribute information and spatial location information as points of interest in the environmental map. The instruction parsing and matching module is used to respond to the received item-finding instruction, parse the item-finding instruction to determine the target item, and match the target point of interest corresponding to the target item from the environmental map; The path planning and navigation module is used to plan a path based on the location of the target point of interest and control the robot to move to the location of the target item.
9. A robot, characterized in that, include: The robot itself; The semantic object-finding navigation system based on multi-sensor fusion as described in claim 8 is integrated into the main control module of the robot body; In addition, the following hardware modules are communicatively connected to the semantic object-finding navigation system based on multi-sensor fusion: wireless radio frequency sensing all-in-one machine, red-green-blue depth RGBD stereo camera, microphone array, lidar and / or inertial measurement unit (IMU).
10. A robot, characterized in that, Includes memory, transceiver, and processor; A memory for storing computer programs; a transceiver for sending and receiving data under the control of the processor; and a processor for reading the computer programs from the memory and performing the following operations: In the process of constructing an environmental map, items are located and identified by integrating RFID sensing information and 3D visual sensing information to obtain the item's identification information, multi-semantic attribute information and spatial location information. The item's identification information, multi-semantic attribute information and spatial location information are then marked as points of interest in the environmental map. In response to a received item search command, the item search command is parsed to determine the target item, and the target point of interest corresponding to the target item is matched from the environmental map; Based on the location of the target point of interest, a path is planned, and the robot is controlled to move to the location of the target item.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the semantic object-finding navigation method based on multi-sensor fusion as described in any one of claims 1 to 7.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the semantic object-finding navigation method based on multi-sensor fusion as described in any one of claims 1 to 7.