Robot object searching method, system and device and computer storage medium
By constructing a scene semantic map and combining it with image and robot data, and using historical video to detect the location of storage containers, the problem of existing item-finding robots being unable to locate accurately has been solved, enabling efficient and accurate item-finding tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-05
- Publication Date
- 2026-04-03
AI Technical Summary
Existing item-finding robot systems cannot accurately indicate the specific location of the target object when dealing with real-world scenarios, resulting in wasted time for users searching for items and affecting user experience.
By constructing a scene semantic map and integrating scene data collected by image acquisition devices and robots, target matching and detection are performed. Combined with historical human detection videos, the location of the storage container is determined, achieving multi-level target localization.
It significantly improves the success rate and accuracy of item finding tasks, saves users time, and optimizes the user experience, especially when the target item is stored away or is in a limited view.
Smart Images

Figure CN121785330A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to a robot object finding method, system, device, and computer storage medium. Background Technology
[0002] With the rapid development of science and technology, robotics and related industries have sprung up, and various robots are widely used in industries, services, and homes to replace or assist humans in completing various tasks. For example, in the home, robots can assist users in daily operations such as retrieving and transporting items. Locating robots can accurately locate target items based on user needs, effectively saving users time.
[0003] Currently, existing item-finding robot systems and methods still have significant shortcomings when dealing with real-world scenarios. In some cases, they cannot accurately inform users of the specific location of the target item, wasting users' search time and affecting user experience.
[0004] Therefore, how to find items more accurately and efficiently is an important problem that urgently needs to be solved. Summary of the Invention
[0005] This application provides a robot object finding method, system, device, and computer storage medium, which can achieve more accurate and efficient object finding.
[0006] In a first aspect, embodiments of this application provide a robot object-finding method, including: In response to a retrieval request, the system performs target matching based on a scene semantic map, which is constructed based on scene data collected by image acquisition devices and robots within the target scene. When a target object is found through target matching, its location information is determined based on the scene semantic map. If no target is found through target matching, real-time scene data is collected for the target scene, and target detection is performed based on the real-time scene data. When a target is detected, its location information is determined based on the scene semantic map and real-time scene data. If no target object is detected by target detection, the location information of the storage container that holds the target object in the target scene is determined based on the historical human detection video. The historical human detection video is obtained by image acquisition equipment as the location information of the target object.
[0007] Secondly, embodiments of this application provide a robot object-finding system, including: The target matching module is used to respond to the item search request and perform target matching based on the scene semantic map. The scene semantic map is constructed based on the scene data collected by the image acquisition device and the robot in the target scene. The first target localization module is used to determine the location information of the target object based on the scene semantic map when the target object is found through target matching. The target detection module is used to collect real-time scene data of the target scene when no target is found through target matching, and to perform target detection based on the real-time scene data. The second target localization module is used to determine the location information of the target object based on the scene semantic map and real-time scene data when the target object is detected by target detection. The storage and positioning module is used to determine the location information of the storage container that holds the target object in the target scene based on historical human detection video when no target object is found through target detection. The historical human detection video is obtained by image acquisition device as the location information of the target object.
[0008] Thirdly, embodiments of this application provide a terminal device, the device including: a processor and a memory storing computer program instructions; When the processor executes computer program instructions, it implements a robot object-finding method as described in the first aspect.
[0009] Fourthly, embodiments of this application provide a computer storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, they implement the robot object finding method as described in the first aspect.
[0010] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the robot object-finding method as described in the first aspect.
[0011] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects: This application provides a robot object-finding method, comprising: First, in response to a user-initiated object-finding request, target matching can be performed on the target object required by the user based on a scene semantic map, which is jointly constructed based on scene data collected by an image acquisition device and a robot. When a target object is successfully matched, the location information of the target object can be output based on the scene semantic map. If a target object cannot be matched, real-time scene data can be collected through an image acquisition device or a robot in the target scene, and target object detection can be performed on the real-time scene data. After a target object is detected, the location information of the target object is output. If no target object is still detected, it is determined that the target object may be stored in a storage container that cannot be directly observed. In this case, the location of the storage container where the target object may be stored can be determined based on pre-stored historical humanoid detection videos, and this location information of the target object is output.
[0012] The technical solution provided in this application can construct a more accurate and blind-spot-free scene semantic map by integrating data collected by image acquisition devices and robots, laying a solid foundation for the smooth and accurate execution of item-finding tasks. Furthermore, the technical solution provided in this application can accurately determine the location of the target object in the scene through a multi-level matching and detection mechanism, significantly improving the success rate and accuracy of the item-finding task execution process. This application can also accurately locate target objects that may be hidden, based on historical video data, avoiding the inability to successfully execute the item-finding task due to the target object being hidden, significantly improving the success rate of item finding, saving user time, and effectively enhancing the user experience during the item-finding process.
[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 A flowchart illustrating a robot object-finding method according to one embodiment of this application; Figure 2 A schematic diagram of the flow framework for a scene semantic map construction process provided in one embodiment of this application; Figure 3 A schematic diagram of the flow framework of a target detection process provided in one embodiment of this application; Figure 4 A schematic flowchart illustrating a process for determining the location information of a storage container, provided as an embodiment of this application; Figure 5 A schematic diagram of the structure of a robot object-finding device provided in another embodiment of this application; Figure 6 This is a schematic diagram of the hardware structure of a terminal device provided in another embodiment of this application. Detailed Implementation
[0016] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0017] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0018] With the rapid development of robotics and related technologies, various robots are widely used in industrial, service, and home settings, playing a vital role in replacing or assisting humans in completing tasks. For example, in the home, robots can assist users in daily operations such as retrieving and transporting items. Among them, item-finding robots can accurately locate target items based on user instructions, significantly improving the efficiency of finding items and saving users a great deal of time.
[0019] However, existing object-finding systems and methods still have significant shortcomings in practical applications. Current mainstream methods primarily rely on data collected by a single sensor for map construction. When the target item is in an occluded area, has limited viewing angles, or is situated within complex furniture layouts, the accuracy and reliability of its location significantly decrease. Furthermore, mainstream object-finding methods typically only identify objects within the sensor's direct visual range, failing to effectively locate target items inside storage containers such as cabinets or drawers. This results in a high failure rate, low efficiency, wasted user time, and negatively impacts practical application effectiveness and user experience.
[0020] To address the aforementioned technical issues, this application provides a robot object-finding method, system, device, and computer storage medium. The method includes: responding to a user-initiated object-finding request, performing target matching based on a scene semantic map, wherein the scene semantic map is jointly constructed based on scene data collected by an image acquisition device and the robot. Upon successful matching of a target object, the location information of the target object can be output based on the scene semantic map.
[0021] If the target object cannot be matched, real-time data collection of the target scene can be performed using image acquisition devices or robots. Target object detection can then be performed on the real-time scene data, and the target object's location information can be output upon detection. If the target object is still not detected, it is determined that the target object may be stored in a container that cannot be directly observed. In this case, the location of the container where the target object might be stored can be determined based on pre-stored historical humanoid detection video, and this location information can be output as the target object's location information.
[0022] The technical solution provided in this application can construct a more accurate and blind-spot-free scene semantic map by integrating data collected from image acquisition devices and robots, laying a solid foundation for the successful execution of object-finding tasks. The technical solution provided in this application can accurately determine the location of the target object in the target scene through multi-level target matching and target detection mechanisms, significantly improving the success rate and accuracy of the object-finding task execution process. This application can also accurately locate target objects that may be stored in containers with storage functions based on historical video data, avoiding the inability to successfully execute the object-finding task due to the target object being stored, significantly improving the success rate of object finding, greatly saving user time, and enhancing the user experience.
[0023] Regarding the execution entity used in the technical solutions provided in this application, it can specifically be a terminal device capable of controlling the image acquisition device and the robot, such as a desktop computer, laptop computer, scene intelligent control center, or the robot itself; or it can be a remote device, such as a server that connects to the image acquisition device and the robot via data. In addition, the execution entity used in this application can also be a software entity, such as a client or software program installed on a terminal device. The specific type of execution entity corresponding to the robot object-finding method, system, device, and storage medium provided in the embodiments of this application is not strictly limited here; it can be flexibly selected and set according to the application scenario and actual needs.
[0024] It should be noted that the embodiments provided in this application do not limit the specific application scenarios corresponding to the robot object finding method, system, device and computer storage medium provided above. The technical solutions provided in the embodiments of this application can be flexibly applied to the actual application scenarios that require finding and locating items according to actual needs.
[0025] For example, in a home setting, when a user needs to find important personal items (such as car keys, house keys, personal identification documents, etc.), the technical solution provided in this application can execute the corresponding item-finding task based on the user's request. Specifically, it can first match the target object (important personal items) based on a scene semantic map pre-built using image acquisition equipment and a robot. If a match is successful, it outputs the corresponding location information in the home setting to the user.
[0026] If a match fails or the user cannot find the target object in the matched location information, it indicates that the target object may have appeared in the target scene after the semantic map was built, or its original location may have changed. In this case, real-time scene data can be collected using image acquisition devices or robots, and target object detection can be performed on the real-time scene data. If the detection is successful, the corresponding location information is output to the user; if the detection fails, it is determined that the target object may not be directly observable and is stored in a storage container.
[0027] Then, based on historical human detection videos collected by image acquisition devices, the system identifies storage containers that the user may have interacted with in the past, and identifies objects that may be stored within them. The location information of the storage containers in the scene semantic map is then used as the location information of the target object and output to the user.
[0028] The technical solution of this application can help users quickly and accurately find important personal items in the home environment, and can also accurately locate target objects that may be in storage containers. This application achieves efficient and accurate execution of the item-finding task through a comprehensively constructed scene semantic map and multi-target localization, fully meeting user needs, greatly saving user time, and significantly optimizing the user's item-finding experience.
[0029] It should be noted that the application scenarios described in the above embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will understand that with the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems. The robot object finding method, system, device, and computer storage medium provided by the embodiments of this application can be applied to any application scenario requiring object finding and positioning.
[0030] Figure 1 This is a flowchart illustrating a robot object-finding method according to one embodiment of this application.
[0031] S101: In response to a lost item request, perform target matching on the target object corresponding to the lost item request based on the scene semantic map.
[0032] In step S101, the technical solution provided by this application can perform target matching on the target object to be found in the object search request based on the user's proposed object search request and based on the pre-built scene semantic map.
[0033] The aforementioned scene semantic map can be constructed by fusing data from the overall scene image captured by the image acquisition device and the scene information obtained by the robot traversing the entire target scene.
[0034] It should be noted that the specific form of the lost item request is not strictly limited in this application. It can be a lost item request sent by the user to the scene intelligent hub or server applying the technical solution of this application through a pre-set software program on a mobile terminal such as a mobile phone or tablet. Alternatively, the user can input the lost item request by voice in the target scene, notifying the terminal device applying the technical solution of this application (such as the scene intelligent hub, lost item robot, etc.) to execute the corresponding lost item task. The choice and setting can be flexibly made according to actual needs and application scenarios.
[0035] Regarding the target detection process for the object corresponding to the item search request in step S101 above, in one embodiment provided in this application, after the scene semantic map is constructed, all item information and corresponding location information contained in the scene semantic map can be recorded and a located list can be constructed and stored in a preset database. Upon receiving an item search request, the type of the target object can be determined based on the item search request, and then the target object can be matched against the pre-saved located list in the database. If a match is successfully found, the location information of the corresponding record is output to the user; if no match is found, the subsequent real-time target detection process is performed.
[0036] Regarding the specific construction process of the scene semantic map, in one embodiment provided in this application, before actually executing the object-finding task, the technical solution provided in this application embodiment can first collect corresponding scene data for the target scene through image acquisition devices and robots respectively, and then fuse the collected scene all-round data to realize the construction of a complete scene semantic map without blind spots.
[0037] Specifically, panoramic images of the target scene can be acquired using image acquisition devices, resulting in at least one panoramic image. All panoramic images contain all the target scene content that the image acquisition devices can directly capture from their current position and angle. Simultaneously, a robot can move along a pre-set route corresponding to the target scene, collecting scene information during its movement using various onboard sensors to obtain the overall scene information for the entire target scene.
[0038] Specifically, the image acquisition device can be a camera, a device with image acquisition and video recording capabilities, such as a camera. The scene information includes, but is not limited to: optical images captured by the robot during movement, collected point cloud data, and inertial measurement data during movement. Optical images can be captured by an optical camera mounted on the robot; point cloud data can be collected by a laser radar mounted on the robot; and inertial measurement data can be collected by an inertial measurement unit (IMU) or module mounted on the robot, specifically including three-axis (horizontal, vertical, and longitudinal) accelerometer data, three-axis gyroscope data, etc.
[0039] Then, a point cloud map corresponding to the target scene can be constructed based on the scene information collected by the robot. The specific method for constructing the point cloud map is not strictly limited in this embodiment. In some embodiments, it can be constructed using Fast LiDAR-Inertial Odometry (Fast-LIO) or Fast LiDAR-Inertial Odometry 2 (Fast-LIO 2) based on scene information. In other embodiments, other methods capable of constructing point cloud maps based on scene information can also be selected, such as tightly-coupled or loosely-coupled methods, which can be flexibly chosen according to actual needs and application scenarios.
[0040] Furthermore, after constructing a point cloud map based on scene information, the panoramic images captured by the image acquisition device can be fused with the point cloud map. During the fusion process, image data from the panoramic images are filled into the corresponding locations in the point cloud map, and the semantic information corresponding to each point cloud location is predicted to obtain a scene semantic map corresponding to the target scene.
[0041] The specific processing methods and approaches for the data fusion and semantic prediction processes described above are not strictly limited in the embodiments of this application. In some embodiments, feature extraction can be performed on point cloud data and panoramic acquired images using the point cloud backbone network and image backbone network in the TransFusion algorithm, respectively. Then, initial semantic prediction is performed on the point cloud features, and the point cloud features and image features are fused through a soft association mechanism to determine the final semantic prediction result.
[0042] In other embodiments, the scene semantic map can also be constructed using algorithms or methods such as Multi-View Point Painting (MVP), 2D-Driven 3D Object Detection, and Cross-Modal Transformer (CMT). The appropriate method can be flexibly selected based on actual needs and application scenarios.
[0043] The above embodiments achieve comprehensive coverage of the target scene through collaborative data acquisition by image acquisition devices and robots, with data complementing and fusing each other to construct a more complete and detailed scene semantic map. Based on the scene semantic map obtained from the above embodiments, the success rate and accuracy of item-finding tasks can be greatly improved in practical applications, minimizing user time spent searching for items and optimizing the user experience.
[0044] It should be noted that this application takes into account that due to the physical location and angle settings of the image acquisition device in the target scene, some scene areas may be occluded during panoramic image acquisition, preventing the acquisition of corresponding image data. The lack of corresponding image data for occluded areas in the point cloud map during the scene semantic map construction process will severely affect the accuracy of the semantic prediction process.
[0045] Based on this, in the embodiments provided in this application, the semantic information and semantic confidence of each point in the point cloud map can be predicted during the fusion process of the point cloud map and the panoramic acquired image, thereby obtaining an initial scene semantic map corresponding to the target scene. The semantic information represents the specific object type that each point in the point cloud map actually corresponds to in the target scene, such as a door, wall, table, wardrobe, etc.
[0046] Semantic confidence represents the accuracy of semantic information. Each point in a point cloud map may output multiple probability vectors during semantic prediction (e.g., calculated using the Softmax function). The semantic type corresponding to the probability vector with the highest probability value is taken as the semantic information of that point, and the probability value is used as the semantic confidence. For example, suppose a point in a point cloud map corresponding to a target scene has the following multiple probability vectors and corresponding semantic types: door 0.02, wall 0.08, wardrobe 0.85, bookshelf 0.05. It can be seen that in this example, the highest probability value for this point is 0.85 for wardrobe, meaning the final predicted semantic information for this point is wardrobe, with a semantic confidence of 0.85 or 85%.
[0047] The above process can determine the semantic information and semantic information confidence level corresponding to each location in the initial scene semantic map. Since the semantic information confidence level of some point clouds lacking image data is low or even 0, the semantic information confidence level corresponding to each location in the initial scene semantic map can be compared with a preset confidence threshold (e.g., 0.6, 0.8, etc.). Locations with semantic information confidence levels less than the preset confidence threshold are considered semantic perception blind spots.
[0048] When the semantic information of some map regions in the initial scene semantic map is the same, the confidence scores of all semantic information in some map regions can be combined and the mean confidence score of semantic information can be calculated. The mean confidence score of semantic information can then be compared with the aforementioned preset confidence threshold. If it is less than the preset confidence threshold, the corresponding region is determined to be a semantic perception blind spot region as a whole.
[0049] In this embodiment, for the identified semantic perception blind spots or areas, a robot can be moved to a location where the blind spot area can be photographed, and images can be collected for the semantic perception blind spot to obtain supplementary images corresponding to the semantic perception blind spot.
[0050] Then, based on the acquired supplementary images, semantic information can be added to the semantic perception blind spots in the initial scene semantic map, so that the confidence level of the semantic information after semantic supplementation can meet the preset confidence threshold. The initial scene semantic map after semantic supplementation can be used as a scene semantic map and as the basis for finding objects when actually performing the object finding task.
[0051] To facilitate understanding of the specific construction process of the scene semantic map, a framework diagram of the scene semantic map construction process is provided for comprehensive explanation. (Refer to...) Figure 2 As shown in the image.
[0052] Figure 2 This is a schematic diagram of the process framework for constructing a scene semantic map according to one embodiment of this application.
[0053] like Figure 2 As shown, after receiving the map construction instruction 201, the image acquisition device can be controlled via 202 to acquire panoramic images, and the robot can be controlled via 203 to acquire scene information and construct a point cloud map. Then, the panoramic images and point cloud map can be fused and semantically predicted via 204 to obtain an initial scene semantic map 205.
[0054] Next, step 206 can be used to determine whether there are semantic perception blind spots based on the initial scene semantic map 205. If so, step 207 can be used to control the robot to plan paths for each semantic perception blind spot, and step 208 can be used to move to each semantic perception blind spot sequentially according to the planned path. Step 209 can be used to acquire supplementary images and input them into step 204 to semantically supplement and update the initial scene semantic map. Figure 2 As shown, multiple rounds of blind spot image acquisition and semantic supplementation can be performed until 206 determines that there are no semantic perception blind spots, thus establishing a scene semantic map 210. Based on the scene semantic map 210, a list of items 211 for the target matching process is determined.
[0055] The above embodiments ensure that the final scene semantic map is complete and intact, without blind spots caused by shooting position and angle, significantly improving the coverage and completeness of the semantic map. Based on the scene semantic map obtained from the above embodiments, the success rate and accuracy of item-finding tasks can be greatly improved in practical applications, maximizing user satisfaction, saving user time, and enhancing user experience.
[0056] S102: When a target object is found through target matching, the location information of the target object is determined based on the scene semantic map.
[0057] In step S102, the technical solution provided in this application embodiment can output the location information of the target object in the scene semantic map as the location information of the target object after the target object is determined based on the pre-built scene semantic map matching, so that the user can locate and find the target object based on the location information.
[0058] S103: If no target object is found through target matching, real-time scene data is collected for the target scene, and target detection is performed on the target object based on the real-time scene data.
[0059] S104: When a target object is detected through object detection, the location information of the target object is determined based on the scene semantic map and real-time scene data.
[0060] In step S103, the technical solution provided in this application embodiment can collect real-time scene data of the target scene through an image acquisition device or a robot when the target object is not successfully matched based on the pre-built scene semantic map. Based on the collected real-time scene data, target detection is performed on the target object corresponding to the item search request to determine the true location of the target object.
[0061] In the embodiments provided in this application, the failure to successfully match a target object based on the pre-constructed scene semantic map can be divided into two cases. First, the target object may be an object that appears in the target scene after the scene semantic map is constructed, but its corresponding item information and location information do not exist in the scene semantic map. In this case, the target matching process in step S101 cannot match the location information corresponding to the target object from the scene semantic map; that is, the target detection process in step S103 is required for real-time target detection.
[0062] The second scenario involves an object that existed in the target scene before the scene semantic map was built, but whose location changed after the scene semantic map was built and before the user submitted the item search request. In this case, although the target object's location information was successfully matched and returned to the user in step S101, the user may not find the target object at the corresponding location in the target scene. The user can then resend the item search request or report a mismatch in the target location. This scenario can also be considered as a failure to successfully match the target object, and the target detection process in step S103 will proceed.
[0063] Regarding the specific processing steps for real-time scene data acquisition and target detection, in one embodiment provided in this application, the target scene can first be acquired in real time using an image acquisition device to obtain the target scene image to be detected.
[0064] Then, based on the image to be detected, the first target detection can be performed. The specific detection method is not strictly limited in this embodiment. In some embodiments, it can be achieved by extracting features from the visual encoder and text encoder in an Open-World Vision Transformer (OwlViT) model, and then performing feature matching on the target object based on the feature data to determine the detection result. In other embodiments, it can also be achieved by using methods such as Grounded Language-Image Pre-training (GLIP) or Grounding DINO to detect the target object in the image. The appropriate method can be flexibly selected based on actual needs and application scenarios.
[0065] After detecting and identifying the target object based on the image to be detected, the corresponding location information of the target object in the scene semantic map can be determined based on the shooting angle and position of the image to be detected and sent to the user, i.e., step S104 above. At the same time, the location information of the target object in the scene semantic map after changes can be updated in real time according to the detection results, ensuring that the location information of the target object can be added or updated in a timely manner, thereby improving the processing efficiency of the next object-finding task.
[0066] The real-time target detection process described in the above embodiments enables accurate real-time positioning of targets, significantly improving the efficiency and accuracy of object-finding tasks and providing users with precise information about the target's location within the target scene. Real-time image acquisition via image acquisition devices significantly enhances the processing efficiency of the target detection process, greatly saving users' time spent searching for objects and improving the user experience.
[0067] When target detection based on an image acquired by an image acquisition device fails to detect the target object, it can be determined that the target object may be located in an area that the image acquisition device cannot directly capture. In this case, a robot with positional mobility can be moved to a blind spot in the scene semantic map to perform real-time image acquisition. Further target detection can then be performed on the real-time acquired images to determine the actual location of the target object.
[0068] Specifically, in the embodiments provided in this application, as can be seen from the above embodiments, during the pre-construction of the scene semantic map, the semantic perception blind spot locations in the target scene where the image acquisition device cannot directly acquire image data can be determined. In this embodiment, when no target object is detected based on the image to be detected acquired in real time by the image acquisition device, the robot can move to a position where image acquisition can be performed for each semantic perception blind spot in the target scene, and image acquisition can be performed for the semantic perception blind spot to obtain the image of the blind spot to be detected corresponding to the semantic perception blind spot.
[0069] Then, based on the acquired images of the blind spots to be detected, a second target detection can be performed on the target object. The specific detection method can be the same as the first target detection process described above. When there are multiple semantic perception blind spots in the target scene, the robot can acquire images one by one according to a preset movement route. After acquiring each image of a blind spot to be detected, the second target detection process is executed, and the robot continues to move towards the next semantic perception blind spot during the detection process, ensuring the detection efficiency of the second target detection process.
[0070] After identifying the target object based on the image of the blind spot to be detected, the location information of the target object in the scene semantic map can be determined and output to the user based on the corresponding location information of the semantic perception blind spot in the image of the blind spot to be detected, as well as the shooting angle and position of the object-finding robot, i.e., step S104 above. Similarly, based on the location information obtained in real time, the scene semantic map is updated in real time to ensure that the location information of the target object can be added or updated in a timely manner, thereby improving the processing efficiency of the next object-finding task.
[0071] To facilitate understanding of the first and second target detection processes described above, a flowchart of a target detection process is provided below for comprehensive explanation. Please refer to the attached diagram for details. Figure 3 As shown.
[0072] Figure 3 This is a schematic diagram of the flow framework of a target detection process provided in one embodiment of this application.
[0073] like Figure 3As shown, if the target object cannot be determined by target matching via 301, the image acquisition device can be controlled via 302 to acquire the image to be detected, and then the target object can be detected on the image to be detected via 303. If the target object is detected in the image to be detected in 303, the location information of the target object in the scene semantic map is output to the user via 304.
[0074] If the target object cannot be detected from the image to be detected, the robot is moved to the semantic perception blind spot during the scene semantic map construction process via 305. Image acquisition is then performed on each of the semantic perception blind spots via 306 to obtain the image of the blind spot to be detected. Then, the target object is detected on the image of the blind spot to be detected via 303. After the target object is determined from the image of the blind spot to be detected, the corresponding location information is output to the user via 304.
[0075] The above embodiments effectively supplement the target detection process using image acquisition devices. Leveraging the robot's flexibility and mobility, image acquisition can be achieved in camera blind spots, providing real-time data for target detection. Based on this embodiment, the detection accuracy and efficiency of targets can be significantly improved, enabling precise localization of newly appearing targets or targets whose positions have changed and are located in camera blind spots after map construction, greatly reducing user search time.
[0076] S105: If no target object is found through target detection, determine the location information of the storage container that holds the target object in the target scene based on historical human detection video, and use it as the location information of the target object.
[0077] In step S105, the technical solution provided in this application embodiment can determine that the target object may be stored inside a storage container that cannot be directly image acquired when the target matching process in step S101 and the target detection process in step S103 have not determined the location information of the target object.
[0078] In this scenario, based on pre-stored historical human detection videos, the interactions between the user and objects with storage functions within a historical timeframe can be determined. Based on this, the location of storage containers that may hold the target object can be located in the scene's semantic map. Then, the location of the storage containers is output to the user as the target object's location information.
[0079] Specifically, in one embodiment provided in this application, a preset interactive behavior detection model can be used to detect the interactive behavior between a person and storage furniture in a pre-stored historical human figure detection video, and determine the interactive behavior detection result.
[0080] Then, based on the interaction behavior detection results, it is possible to determine the storage containers in the target scene that have human interaction within a historical time period, and to determine the location information of the storage containers in the scene semantic map, which is then used as the location information of the target object and output to the user. In this embodiment, the specific model type and structure of the interaction behavior detection model used to perform the aforementioned human-storage container interaction behavior detection are not strictly limited.
[0081] In some embodiments, feature data from historical human detection videos and preset prompt words can be extracted using the video encoder and text encoder in the CogVideo (Cognitive Video) model, respectively. Feature matching is then performed based on the feature data to identify the containers in the scene semantic map that interact with the human as the interaction behavior detection result. The preset prompt words can be flexibly set and can specifically include the human, various containers with storage functions, and text descriptions of the interaction behavior.
[0082] In other embodiments, models or methods with video analysis and feature matching capabilities, such as Contrastive Language-Image Pre-training (CLIP), Bootstrapping Language-Image Pre-training (BLIP), and InternVideo, can also be used to detect containers that interact with people in historical human detection videos. The appropriate model can be flexibly selected based on actual needs and application scenarios.
[0083] Based on the above embodiments, the possible storage location of the target object can be determined by effectively analyzing historical video. Even if the target object is stored in an area that cannot be directly observed, this embodiment can still achieve accurate positioning, breaking through the physical limits of traditional visual object finding, realizing accurate and efficient execution of the object finding task, greatly reducing the time spent by users in finding objects, and improving the user experience during the object finding process.
[0084] It should be noted that, in the process of outputting the determined location information of the storage container as the location information of the target object to the user, in some embodiments of this application, the storage containers that cannot possibly contain the target object can be filtered according to the size of the target object and the actual storage space size of the storage container, so as to prevent the output of erroneous container locations that do not conform to common sense to the user.
[0085] For example, a user inputs a request to find a normal-sized steamer. Based on historical human detection video, the system detects that the user recently interacted with the drawer of a small bedside table. However, since the steamer's size is significantly larger than the drawer's storage space, the furniture location information of the small bedside table can be filtered out, and the location information of other storage containers with a more suitable size for storing the steamer can be output to the user.
[0086] Regarding the acquisition and determination process of the aforementioned historical human detection videos, in one embodiment provided in this application, dynamic video monitoring can be performed using an image acquisition device within a preset time window. When the degree of change in the video content of adjacent video frames in the monitored video exceeds a preset threshold, it is determined that a change has occurred in the monitored video, and the start and end times of the video segment with the change are recorded. The monitored video within the corresponding time period is considered a dynamic video segment. The preset time window can be freely set, such as 7 days or 3 days, and can be flexibly set and adjusted according to actual needs.
[0087] Then, human detection can be performed on the dynamic video clips. When it is confirmed that there are human figures in the dynamic video clips, the dynamic video clips can be marked as "human figures appear" and used as the aforementioned historical human figure detection videos. The specific human detection method is not strictly limited in this application embodiment.
[0088] In some embodiments, the dynamic video clip can be preprocessed and feature extracted using a You Only Look Once Version 8 (YOLOv8) algorithm. A pre-defined detection model then predicts the type of the video features. After post-processing and non-maximum suppression, the prediction result is output, thus determining whether a person exists in the dynamic video clip. In other embodiments, a two-stage detector model or a Transformer-based detector model can be used, and the appropriate model can be flexibly selected based on actual needs and application scenarios.
[0089] To facilitate understanding of the acquisition and determination process of the aforementioned historical humanoid detection videos, as well as the process of determining furniture location information under storage conditions, the following is a comprehensive introduction using a flowchart of the process for determining the location information of a storage container.
[0090] Figure 4 This is a schematic flowchart illustrating a process for determining the location information of a storage container, provided as an embodiment of this application.
[0091] like Figure 4As shown, by performing screen changes and human detection within a preset time window, the historical human detection video 402 is determined. Then, if the target object's location is not determined during the target matching and target detection processes, by performing human-container interaction behavior detection on the historical human detection video 402, at least one container that may contain the target object is identified, and the corresponding container's location information is output to the user by 404.
[0092] The above embodiments can accurately record historical human detection videos through video changes and human detection, providing a valid basis for determining the location information of the storage container for the target object in the above embodiments. This enables accurate positioning of the target object under storage conditions, greatly improving the success rate and accuracy of the item search task, significantly reducing the time spent by users in the process of finding the target object, and enhancing the user experience.
[0093] The above describes a specific implementation of a robot object-finding method provided in this application. The technical solution provided in this application can construct a more accurate and blind-spot-free scene semantic map by fusing data from image acquisition devices and the object-finding robot, laying a solid foundation for the successful execution of the object-finding task. The technical solution provided in this application can accurately determine the location of the target object in the target scene through multi-level target matching and target detection mechanisms, significantly improving the success rate and accuracy of the object-finding task execution process. Furthermore, this application can accurately locate target objects that may be stored in storage containers based on historical video data, avoiding the inability to successfully execute the object-finding task due to the object being stored, significantly improving the success rate of object finding, greatly saving user time, and enhancing the user experience.
[0094] Based on the robot object finding method provided in the above embodiments, this application also provides an embodiment of a robot object finding system.
[0095] Figure 5 This is a schematic diagram of a robot object-finding system according to another embodiment of this application. The robot object-finding system 500 includes: The target matching module 501 is used to respond to the object search request and perform target matching on the target object corresponding to the object search request based on the scene semantic map. The scene semantic map is constructed based on the scene data collected by the image acquisition device and the robot in the target scene. The first target localization module 502 is used to determine the location information of the target object based on the scene semantic map when the target object is found through target matching. The target detection module 503 is used to collect real-time scene data of the target scene when no target object is found through target matching, and to perform target detection on the target object based on the real-time scene data; The second target localization module 504 is used to determine the location information of the target object based on the scene semantic map and real-time scene data when the target object is detected by target detection. The storage and positioning module 505, when no target object is found through target detection, determines the location information of the storage container for the target object in the target scene based on historical human detection video. The historical human detection video is obtained by image acquisition device as the location information of the target object.
[0096] In some embodiments, the robot object finding system 500 described above also includes a map building module 506; Map building module 506 is specifically used for: The target scene is captured by the image acquisition device to obtain at least one panoramic image of the target scene, and the scene information of the target scene is obtained by the robot traversing the target scene to collect information. The scene information includes at least optical images, scene point cloud data and inertial measurement data. Based on the scene information, a point cloud map of the target scene is constructed; The point cloud map and the panoramic acquired image are fused together to predict the semantic information of each point in the point cloud map and determine the scene semantic map.
[0097] In some embodiments, the map building module 506 described above is specifically used for: The point cloud map and the panoramic acquired image are fused together to predict the semantic information and semantic information confidence of each point in the point cloud map, thereby obtaining an initial scene semantic map. Based on the semantic information confidence level and the preset confidence threshold, determine the semantic perception blind spots in the initial scene semantic map; The robot moves to the corresponding position of the semantic perception blind spot in the target scene and performs image acquisition on the semantic perception blind spot to obtain a supplementary image. Based on the supplementary image, the semantic information of the semantic perception blind spot in the initial scene semantic map is semantically supplemented, and the semantically supplemented map is used as the scene semantic map.
[0098] In some embodiments, the target detection module 503 described above is specifically used for: The image acquisition device acquires the image to be detected of the target scene, which is used as the real-time scene data. Based on the image to be detected, a first target detection is performed on the target object; The aforementioned second target positioning module 504 is specifically used for: If the target object is determined through the first target detection, the location information of the target object is determined based on the scene semantic map and the image to be detected.
[0099] In some embodiments, the target detection module 503 described above is specifically used for: If the target object is not identified by the first target detection, the robot moves to the corresponding position of the semantic perception blind spot in the scene semantic map in the target scene, and performs image acquisition on the semantic perception blind spot to obtain the image of the blind spot to be detected, which is used as the real-time scene data. Based on the image of the blind spot to be detected, a second target detection is performed on the target object; If the target object is determined through the second target detection, the location information of the target object is determined based on the scene semantic map and the image of the blind spot to be detected.
[0100] In some embodiments, the storage and positioning module 505 is specifically used for: The interaction behavior detection model is used to detect the interaction behavior between the person and the storage items in the historical human figure detection video, and the interaction behavior detection result is determined. Based on the interaction behavior detection results, the location information of the storage items that interact with the character in the target scene is determined in the scene semantic map, and used as the location information of the storage container.
[0101] In some embodiments, the storage and positioning module 505 is specifically used for: The image acquisition device monitors changes in the target scene within a preset time window to identify dynamic video segments with changing images. The dynamic video clip is subjected to human detection. If a human is found in the dynamic video clip, the dynamic video clip is used as the historical human detection video.
[0102] Figure 6 This is a schematic diagram of the hardware structure of a terminal device provided in another embodiment of this application.
[0103] The terminal device may include a processor 601 and a memory 602 storing computer program instructions.
[0104] Specifically, the processor 601 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0105] Memory 602 may include mass storage for data or instructions. For example, and not limitingly, memory 602 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 602 may include removable or non-removable (or fixed) media. Where appropriate, memory 602 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 602 is non-volatile solid-state memory.
[0106] In a particular embodiment, memory 602 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Thus, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform operations described with reference to a robot object-finding method according to one aspect of this disclosure.
[0107] The processor 601 reads and executes computer program instructions stored in the memory 602 to implement any of the robot object finding methods in the above embodiments.
[0108] In one example, the terminal device may also include a communication interface 603 and a bus 610. Wherein, as... Figure 6 As shown, the processor 601, memory 602, and communication interface 603 are connected through bus 610 and complete communication with each other.
[0109] The communication interface 603 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0110] Bus 610 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 610 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0111] Furthermore, in conjunction with the robot object-finding methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the robot object-finding methods in the above embodiments.
[0112] This application also provides a computer program product, including a computer program, which, when executed, implements any of the robot object finding methods described in the above embodiments.
[0113] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0114] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0115] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0116] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0117] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A robot object finding method, characterized in that, include: In response to a locator request, target matching is performed on the target object corresponding to the locator request based on a scene semantic map, wherein the scene semantic map is constructed based on scene data collected by image acquisition devices and robots within the target scene; If the target object is found through target matching, the location information of the target object is determined based on the scene semantic map; If the target object is not found through the target matching, real-time scene data is collected for the target scene, and target detection is performed on the target object based on the real-time scene data; If the target object is detected by the target detection, the location information of the target object is determined based on the scene semantic map and the real-time scene data; If the target object is not detected by the target detection, the location information of the storage container that holds the target object in the target scene is determined based on the historical human detection video, and is used as the location information of the target object. The historical human detection video is acquired by the image acquisition device.
2. The method according to claim 1, characterized in that, The scene semantic map is constructed based on scene data collected by image acquisition devices and robots within the target scene, including: The target scene is captured by the image acquisition device to obtain at least one panoramic image of the target scene, and the scene information of the target scene is obtained by the robot traversing the target scene to collect information. The scene information includes at least optical images, scene point cloud data and inertial measurement data. Based on the scene information, a point cloud map of the target scene is constructed; The point cloud map and the panoramic acquired image are fused together to predict the semantic information of each point in the point cloud map and determine the scene semantic map.
3. The method according to claim 2, characterized in that, The point cloud map and the panoramic image are fused to predict the semantic information of each point in the point cloud map and determine the scene semantic map, including: The point cloud map and the panoramic acquired image are fused together to predict the semantic information and semantic information confidence of each point in the point cloud map, thereby obtaining an initial scene semantic map. Based on the semantic information confidence level and the preset confidence threshold, determine the semantic perception blind spots in the initial scene semantic map; The robot moves to the corresponding position of the semantic perception blind spot in the target scene and performs image acquisition on the semantic perception blind spot to obtain a supplementary image. Based on the supplementary image, the semantic information of the semantic perception blind spot in the initial scene semantic map is semantically supplemented, and the semantically supplemented map is used as the scene semantic map.
4. The method according to claim 1, characterized in that, Real-time scene data acquisition is performed on the target scene, and target detection is performed on the target object based on the real-time scene data, including: The image acquisition device acquires the image to be detected of the target scene, which is used as the real-time scene data. Based on the image to be detected, a first target detection is performed on the target object; If the target object is detected through the target detection, the location information of the target object is determined based on the scene semantic map and the real-time scene data, including: If the target object is determined through the first target detection, the location information of the target object is determined based on the scene semantic map and the image to be detected.
5. The method according to claim 4, characterized in that, The method further includes: If the target object is not identified by the first target detection, the robot moves to the corresponding position of the semantic perception blind spot in the scene semantic map in the target scene, and performs image acquisition on the semantic perception blind spot to obtain the image of the blind spot to be detected, which is used as the real-time scene data. Based on the image of the blind spot to be detected, a second target detection is performed on the target object; If the target object is determined through the second target detection, the location information of the target object is determined based on the scene semantic map and the image of the blind spot to be detected.
6. The method according to claim 1, characterized in that... Based on historical human detection videos, the location information of the storage container holding the target object in the target scene is determined, including: The interaction behavior detection model is used to detect the interaction behavior between the person and the storage items in the historical human figure detection video, and the interaction behavior detection result is determined. Based on the interaction behavior detection results, the location information of the storage items that interact with the character in the target scene is determined in the scene semantic map, and used as the location information of the storage container.
7. The method according to claim 1, characterized in that, The historical human detection video was acquired by the image acquisition device and includes: The image acquisition device monitors changes in the target scene within a preset time window to identify dynamic video segments with changing images. The dynamic video clip is subjected to human detection. If a human is found in the dynamic video clip, the dynamic video clip is used as the historical human detection video.
8. A robot object-finding system, characterized in that, include: The target matching module is used to respond to a retrieval request and perform target matching on the target object corresponding to the retrieval request based on a scene semantic map. The scene semantic map is constructed based on scene data collected by image acquisition devices and robots within the target scene. The first target localization module is used to determine the location information of the target object based on the scene semantic map when the target object is found through the target matching. The target detection module is used to collect real-time scene data of the target scene when the target object is not found through the target matching, and to perform target detection on the target object based on the real-time scene data; The second target localization module is used to determine the location information of the target object based on the scene semantic map and the real-time scene data when the target object is detected by the target detection. The storage and positioning module is used to determine the location information of the storage container that holds the target object in the target scene based on historical human detection video when the target object is not detected by the target detection. The historical human detection video is acquired by the image acquisition device.
9. A terminal device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the robot object finding method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the robot object-finding method as described in any one of claims 1-7.
Citation Information
Patent Citations
Target search method based on semantic map
CN113505646A
Robot business execution method and device, robot sensing system and robot
CN117873079A
Electronic map updating method, computer device and storage medium
CN119759929A
Semantic map construction method and device, semantic map navigation method and device, electronic equipment and storage medium
CN120141509A