A Method and System for Toilet Recognition Based on Depth Camera and Visual Model
By employing spatiotemporal alignment fusion, visual large model feature extraction, and anti-collision verification, the problems of poor multimodal data fusion quality and low pose estimation accuracy were solved, achieving high-precision and safe operation for toilet recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU WANGCHAO INNOVATION ROBOT CO LTD
- Filing Date
- 2026-07-02
- Publication Date
- 2026-07-31
AI Technical Summary
Existing toilet recognition solutions based on depth cameras and visual models lack rigorous spatiotemporal alignment and time deviation evaluation when processing multimodal data. This results in time delays and spatial projection deviations between color images and depth point clouds, making it impossible to form high-quality fusion features. Furthermore, the detection accuracy is limited when dealing with cluttered backgrounds, the 3D pose estimation accuracy is low, and collisions are prone to occur during robotic arm operations.
By performing spatiotemporal alignment and fusion based on intrinsic and extrinsic coordinate systems, extracting multi-layer convolutional semantic features and detecting anchor boxes using a large visual model, cropping 3D point clouds and registering poses by combining toilet category confidence scores, and performing anti-collision verification by combining hand-eye calibration and safe distance thresholds, high-precision six-DOF pose parameters of toilets are generated.
It achieves high-precision multimodal data fusion for toilet recognition, effectively separating the target from the background and outputting high-precision six-degree-of-freedom pose parameters to ensure the safety and accuracy of the robotic arm operation.
Smart Images

Figure CN122493419A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of toilet recognition, and more specifically, to a toilet recognition method and system based on a depth camera and a large visual model. Background Technology
[0002] With the rapid development of service robots and automation technology, automated cleaning and maintenance operations in complex scenarios such as restrooms have gradually become a focus of industry attention. Accurately identifying and locating various toilets, urinals, and other fixtures is a prerequisite for precise interaction and operation by robots or robotic arms in automated tasks. However, restroom environments often present challenges such as drastic changes in lighting, narrow spaces, and highly reflective ceramic surfaces lacking significant texture features. Traditional two-dimensional vision technology alone cannot acquire the three-dimensional spatial position of the target, making it difficult to guide robotic arms for depth operations; while a single three-dimensional point cloud is insufficient for accurate semantic analysis of complex scenes. Therefore, constructing a toilet fixture recognition solution based on depth cameras and large visual models, integrating the powerful generalization recognition and semantic understanding capabilities of large visual models with the high-precision three-dimensional spatial perception capabilities of depth cameras, has become an inevitable development trend for giving service robots intelligent vision and achieving high-precision six-degree-of-freedom pose estimation and intelligent operation of toilet fixtures.
[0003] However, existing toilet identification schemes based on depth cameras and visual models still have many shortcomings in practical applications. First, in the multimodal data processing stage, existing methods often lack strict spatiotemporal alignment and time deviation evaluation mechanisms, resulting in time delays and spatial projection deviations between color images and depth point clouds, making it impossible to form high-quality fusion features. Second, for toilet target extraction, conventional visual models have limited detection accuracy when dealing with cluttered backgrounds and fail to effectively utilize 2D semantic anchor boxes as masks to guide accurate cropping of 3D point clouds, resulting in extracted 3D point clouds containing a large amount of background noise and adhering objects, greatly increasing the difficulty of subsequent segmentation and computation. In addition, when performing 3D pose estimation, existing algorithms mostly use traditional rigid matching of point clouds, failing to fully utilize the class confidence of the visual model as registration weights, which makes it easy to get trapped in local optima when point clouds are missing or have interference, and unable to obtain high-precision six-DOF pose parameters. Finally, most existing solutions remain at the level of pure visual coordinate output, lacking deep integration with robot hand-eye calibration systems and collision avoidance safety mechanisms. This makes it easy for the robotic arm to collide and damage fragile ceramic toilets due to coordinate calculation errors when performing close-range operations.
[0004] Therefore, the industry urgently needs a comprehensive toilet identification solution that can take into account spatiotemporal fusion, accurate noise reduction, intelligent registration, and anti-collision verification. Summary of the Invention
[0005] To address the aforementioned problems in existing technologies, this application provides a toilet identification method based on a depth camera and a large visual model, comprising: Step 1, performing spatiotemporal alignment and fusion on the acquired raw sensor data stream based on an intrinsic and extrinsic coordinate system to obtain toilet color depth alignment data, wherein the raw sensor data stream includes a toilet color image and a toilet depth point cloud; Step 2, performing multi-layer convolutional semantic feature extraction and anchor box target detection on the color channels in the toilet color depth alignment data based on a large visual model to obtain a two-dimensional anchor box matrix of toilet objects and a toilet category confidence score; Step 3, using the toilet two-dimensional anchor box matrix... The anchor box matrix serves as a pixel space mask. It is used to perform region-of-interest point cloud cropping, statistical filtering for noise reduction, and clustering segmentation on the depth channel of the toilet color depth alignment data to obtain the 3D point cloud for toilet isolation. Step four involves using the toilet category confidence score as a registration weighting factor to perform principal component analysis and iterative nearest-point pose registration on the 3D point cloud for toilet isolation to obtain the six-DOF pose parameters of the toilet. Step five further involves using a hand-eye calibration homogeneous transformation matrix and a preset safety distance threshold to perform world coordinate system transformation and anti-collision differential verification on the six-DOF pose parameters of the toilet to obtain the collision-free operation coordinates of the toilet.
[0006] This application also provides a toilet recognition system based on a depth camera and a large visual model, comprising: a sensor data pre-alignment module, used to perform spatiotemporal alignment and fusion of the acquired raw sensor data stream based on an intrinsic and extrinsic parameter coordinate system to obtain toilet color depth alignment data, wherein the raw sensor data stream includes a toilet color image and a toilet depth point cloud; a semantic extraction and object detection module, used to perform multi-layer convolutional semantic feature extraction and anchor box object detection on the color channels in the toilet color depth alignment data based on a large visual model to obtain a two-dimensional anchor box matrix of toilet objects and a toilet object category confidence score; and a three-dimensional point cloud generation module, used to generate a three-dimensional point cloud matrix of toilet objects using the two-dimensional anchor box matrix of toilet objects. The bounding box matrix serves as a pixel-space mask, performing region-of-interest point cloud cropping, statistical filtering for noise reduction, and clustering segmentation on the depth channel of the toilet color depth alignment data to obtain the 3D point cloud for toilet isolation. The pose parameter generation module uses the toilet category confidence score as a registration weighting factor to perform principal component analysis and iterative nearest-point pose registration on the 3D point cloud for toilet isolation to obtain the six-DOF pose parameters of the toilet. The operation coordinate generation module uses a hand-eye calibration homogeneous transformation matrix and a preset safety distance threshold to perform world coordinate system transformation and anti-collision differential verification on the six-DOF pose parameters of the toilet to obtain the collision-free operation coordinates of the toilet.
[0007] Compared with existing technologies, this application provides a toilet recognition method and system based on a depth camera and a large visual model, aiming to solve problems such as poor multimodal data fusion quality, high noise in point clouds, low pose estimation accuracy, and easy collisions during robotic arm operations. First, it performs spatiotemporal alignment fusion of color images and depth point clouds to obtain high-quality color-depth aligned data, effectively overcoming spatiotemporal biases between multiple sensors. Then, it uses a large visual model to extract features to obtain two-dimensional anchor boxes and category confidence scores, and innovatively uses the two-dimensional anchor boxes as pixel space masks to guide precise cropping and filtering noise reduction of the depth layer, achieving efficient separation of the target from cluttered backgrounds. Next, it uses the category confidence score as a registration weighting factor to perform pose registration on the isolated three-dimensional point cloud, avoiding registration getting trapped in local optima, thereby outputting high-precision six-DOF pose parameters. Finally, it combines the hand-eye calibration matrix and a preset safe distance threshold for anti-collision differential verification to calculate collision-free operating coordinates, completely establishing a technical closed loop from high-precision multimodal recognition to safe and automated robotic arm operations. Attached Figure Description
[0008] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings.
[0009] Figure 1 This is a flowchart of a toilet recognition method based on a depth camera and a large visual model according to an embodiment of this application.
[0010] Figure 2 This is a schematic diagram of the data flow of a toilet recognition method based on a depth camera and a large visual model according to an embodiment of this application.
[0011] Figure 3 This is an actual structural diagram of the intelligent cleaning robot involved in the embodiments of this application.
[0012] Figure 4 This is a flowchart of step three in the toilet recognition method based on a depth camera and a large visual model according to an embodiment of this application.
[0013] Figure 5 This is a block diagram of a toilet recognition system based on a depth camera and a large visual model according to an embodiment of this application.
[0014] The components include: 1. Depth camera; 2. Visual recognition camera; 3. Magnetic quick-release interface structure; 4. Flexible two-finger gripper; 5. Six-degree-of-freedom serial foldable robotic arm; 6. High-pressure water spray actuator; 7. Sewage recycling actuator; 8. Air drying head actuator; 9. Sealed pressure-resistant pipeline; 10. Sealed clean water tank; 11. Sealed sewage tank; 12. Intelligent electronic control scheduling structure; 13. Multimodal sensing data real-time transmission wiring; 14. Side infrared sensor; 15. Multi-line lidar system. Detailed Implementation
[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. It should be understood that the drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0016] To address the shortcomings in the aforementioned technical fields, this application proposes a toilet recognition method based on a depth camera and a large visual model. For example... Figure 1 and Figure 2 As shown, Figure 1 This is a flowchart of a toilet recognition method based on a depth camera and a large visual model according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow of a toilet recognition method based on a depth camera and a large visual model according to an embodiment of this application. An additional illustration showing the physical execution carrier of this method is included herein, such as... Figure 3 As shown in the relevant structural description, Figure 3This is an actual structural diagram of the intelligent cleaning robot involved in the embodiments of this application. The toilet recognition method and system based on depth cameras and large visual models described in this application mainly rely on a multi-functional integrated intelligent cleaning robot hardware platform to achieve overall automated cleaning and safety management in damp and enclosed scenarios such as public toilets. The main body of the robot is equipped with a six-degree-of-freedom serially foldable robotic arm 5 with full spatial attitude adjustment capability. Its end has an integrated magnetic quick-release interface structure 3 and is fixedly connected to a flexible two-finger gripper 4, which can be quickly replaced and control multi-functional extended execution terminals including a high-pressure water spray execution terminal 6, a sewage recycling execution terminal 7, and a drying head execution terminal 8. In order to achieve fluid closed loop, the internal body of the device is designed with independent sealed clean water tank 10 and sealed sewage tank 11, which are directly connected to the execution end through a sealed pressure-resistant pipeline 9 concealed along the inner cavity of the robotic arm. In terms of the multi-dimensional perception system, the robot platform is equipped with heterogeneous distributed sensors that perfectly match the algorithm extraction logic described in this application: a visual recognition camera 2, responsible for collecting high-resolution images in real time, is fixed to the side of the gripper at the end of the robotic arm to follow the robotic arm in performing synchronous approach and movement, while a depth camera 1, responsible for macroscopic depth distance perception and wide field of view scanning, is fixedly mounted on the lower part of the robot body. At the same time, the robot body is also equipped with a side infrared sensor 14 and a multi-line LiDAR system 15. These environmental panoramas and image features captured by the multi-modal perception components are transmitted in real time back to the intelligent electronic control scheduling structure 12 inside the sealed electronic control compartment of the robot body through the built-in multi-modal perception data real-time transmission cable 13. The integrated hardware and software AI algorithm model library and intelligent scheduling engine inside will fully support the spatiotemporal alignment and six-degree-of-freedom pose optimization calculation of the large model proposed in this application, thereby enabling the final calculated anti-collision three-dimensional safety coordinates to be seamlessly converted into physical actions and fluid closed-loop commands driven by servo control, completely overcoming the highly complex problems of automated toilet and blind spot recognition cleaning and full-process operation tracking.
[0017] Step one: Based on the intrinsic and extrinsic coordinate system, the acquired raw sensor data stream is spatiotemporally aligned and fused to obtain toilet color depth aligned data. The raw sensor data stream includes toilet color images and toilet depth point clouds. In this application, the toilet depth point cloud is acquired by a depth camera, and the toilet color image is acquired by a visual recognition camera. It should be understood that in the complex workspace where service robots perform toilet cleaning or maintenance tasks, due to the physical space limitations of hardware deployment, the visual recognition camera responsible for acquiring real-time high-resolution detailed features is usually mounted on the end effector of the robotic arm (e.g., next to the gripper) to move synchronously with the robotic arm, while the depth camera responsible for wide-area distance perception and large field-of-view spatial contour scanning is fixed to the lower end of the robot's base. This heterogeneous multi-sensor distributed physical space layout structure, coupled with the inherent differences in the sampling frequency of the underlying hardware of different sensors, results in significant misalignment and fragmentation of the acquired multimodal data in terms of time and space references. Directly overlaying or fusing unaligned color images with depth point clouds not only causes ghosting and edge tearing of the target object, but also leads to serious discrepancies between the two-dimensional semantic features extracted from the subsequent large visual model and the three-dimensional spatial position. This makes the robotic arm prone to positioning failures or even collision damage during operation. To eliminate this spatiotemporal barrier between multi-source heterogeneous sensors, step one is introduced, which becomes a necessary prerequisite for ensuring the accuracy of subsequent high-precision semantic extraction and six-DOF pose estimation of the toilet.
[0018] In one operable embodiment of this application, step one includes: based on the hardware absolute timestamps carried by the toilet color image and the toilet depth point cloud, performing nearest neighbor time deviation evaluation and super-difference anomalous frame removal on the original sensor data stream to obtain a synchronized color image and a synchronized depth point cloud; based on the camera's internal parameter matrix and external rigid body transformation parameters, performing spatial coordinate mapping from a three-dimensional physical rigid body to a two-dimensional pixel plane on the three-dimensional spatial coordinates of each effective point in the synchronized depth point cloud to obtain a depth pixel projection matrix; based on the addressing topology index recorded by the depth pixel projection matrix, performing channel concatenation stitching and mean filtering gap filling on the three-channel color data of the synchronized color image and the effective depth values of the synchronized depth point cloud to obtain toilet color depth aligned data.
[0019] The implementation process is as follows: The first stage is nearest neighbor time deviation assessment and outlier frame removal. The raw sensor data stream refers to the raw multimedia and depth information stream directly captured by the underlying hardware interface without any secondary processing. After receiving the raw sensor data stream transmitted from the communication bus of the underlying device layer, such as the industrial control motherboard, it is first unpacked according to the underlying network protocol format, removing the transmission packet header and check bits, and discretizing the mixed data stream into two independent data threads: one is a high-frequency camera video frame thread, with the acquisition frequency of the visual recognition camera set to a high value, such as 30 frames per second; the other is a low-frequency depth camera point cloud frame thread, with the depth camera having a relatively low frequency, such as 10 frames per second, due to the involvement of infrared emission and time-of-flight calculation. Subsequently, based on the system global clock crystal oscillator sequence of the intelligent cleaning equipment, the hardware absolute timestamps carried by all unpacked video frames and depth camera point cloud frames are extracted. This hardware absolute timestamp is directly stamped by the underlying microcontroller at the moment the sensor hardware exposure is triggered, and can accurately reflect the absolute physical moment of data acquisition. Next, a comparison loop is initiated using the timeline of the high-frequency video frame thread as a reference point. For each frame in the video frame thread, the absolute value of the time difference between the timestamp of that video frame and the timestamps of all candidate depth camera point cloud frames within the current sliding time window is calculated. Based on this, the corresponding point cloud frame sequence number with the minimum deviation is found through traversal comparisons, thus achieving nearest neighbor matching. To ensure the rigor of the matching and prevent matching failures due to network latency or system thread blocking, a preset threshold is used to determine the minimum deviation value. For example, the preset time deviation threshold is set to 15 milliseconds. When the calculated absolute value of the minimum time difference is less than or equal to 15 milliseconds—for example, if the timestamp of a video frame is calculated to be 1,001,100 milliseconds, and the timestamp of the nearest point cloud frame is 1,001,122 milliseconds, with a deviation of 12 milliseconds—the data pair is considered highly consistent in the time dimension and is retained. If environmental interference or thread blocking causes the minimum deviation to exceed 15 milliseconds, an abnormal frame rejection mechanism is triggered, and all video frames and point cloud frames within that time period are identified as abnormal frames and discarded. This avoids motion blur and spatial misalignment caused by time differences during the robotic arm's movement. After the above rigorous temporal evaluation and filtering, a synchronized color image with strict temporal alignment and matching characteristics and a synchronized depth point cloud with consistent time are finally generated.
[0020] After completing the temporal synchronization, the physical mapping phase proceeds to the spatial dimension. First, the synchronized depth point cloud is traversed and deconstructed to extract the three-dimensional geometric coordinates of each valid point in its cloud array. The point cloud data captured by the depth camera initially represents a discrete set of points in the depth camera's own physical coordinate system, containing Cartesian coordinate values of spatial location. To transform these three-dimensional coordinates to a two-dimensional image perspective, firmware data stored in the device's memory after a pre-implemented joint hand-eye calibration process is required. This firmware data includes the camera's internal parameter matrix and the external rigid body transformation parameters between the depth camera base and the camera's origin center. The camera's internal parameter matrix describes the optical projection properties inside the visual recognition camera, including focal length and principal point coordinates; the external rigid body transformation parameters quantify the relative spatial positional relationship between the depth camera at the bottom of the device and the end effector of the robotic arm, consisting of a rotation matrix and translation vectors.
[0021] For each extracted 3D spatial coordinate point, multi-degree-of-freedom affine transformation and perspective projection operations are performed. Specifically, the 3D coordinates of the point cloud are transformed from the depth camera coordinate system to the visual recognition camera coordinate system using external rigid body transformation parameters. Then, using the camera's internal parameter matrix, the physical geometric points in the sensor's stereo coordinate system are precisely flattened and projected onto an imaginary 2D pixel image plane coordinate system. This spatial coordinate mapping process satisfies the following algebraic geometric relationship:
[0022] In the above formula, , , This represents the true physical coordinates of any valid point extracted from the synchronous depth point cloud in the three-dimensional space of the depth camera's world coordinate system. For example, the coordinates of a point scanned on the edge of a toilet are 450 mm on the horizontal axis, -200 mm on the vertical axis, and 800 mm on the depth axis. This represents the 3×3 orthogonal rotation matrix in the external rigid body transformation parameters. The 3×1 translation vector matrix represents the external rigid body transformation parameters. The two work together to achieve rigid body translation and rotation across the sensor coordinate system. This represents the camera's internal parameter matrix, which consists of the camera's horizontal focal length, vertical focal length, and optical center pixel coordinates. This represents the actual depth value of the point after transformation to the camera coordinate system, that is, the distance along the optical axis. and These represent the ideal sub-pixel coordinates of the horizontal and vertical axes on the image plane of the 2D image after projection calculation. After calculating the coordinates, all obtained 2D discrete target coordinates are normalized, and a small number of sub-pixel bits are discarded to construct a discrete cell-level index, as shown in the calculation. It is 452.7. The value is 805.3, and after discarding the decimal places, the index is fixed at 452 and 805. By constructing this kind of two-dimensional pixel-level coordinate grid, a topological correspondence array between the corresponding three-dimensional spatial point data and the two-dimensional pixel coordinates is established, and the final output is a depth pixel projection matrix containing this strict projection mapping relationship. This matrix is essentially an addressing topology map that records which two-dimensional pixel grid a three-dimensional spatial point should fall into.
[0023] Finally, in the channel fusion stage of the data dimension, an addressing topology index based on the depth pixel projection matrix is executed. First, a high-dimensional data tensor with dimensions completely isomorphic to the synchronized color image is allocated and initialized in the device's graphics processing memory space as a carrier base, for example, a tensor space with a resolution of 1920×1080. Then, all image pixel data from the synchronized color image are arranged according to the original matrix and injected into the carrier base according to the channel index, thus forming the underlying red, green, and blue primary color channels, i.e., the standard color image data layer. Next, based on the precise addressing topology index link recorded by the input depth pixel projection matrix, the effective physical structure depth values provided by the synchronized depth point cloud are matched one by one and pushed into the corresponding pixel index positions of the newly added fourth-dimensional spatial data channel in this carrier base. At this point, an initial four-channel fusion matrix is constructed.
[0024] However, in complex real-world bathroom scenarios, toilet fixtures are often made of smooth ceramic, which is prone to infrared laser speckle absorption or high mirror reflectivity. Combined with the field-of-view occlusion caused by the physical position difference between the two sensors, a large number of depth point clouds are lost after mapping, resulting in many blank, black hole, and other featureless pixel regions in the fourth dimension (depth channel). To eliminate the interference of these gaps on subsequent large-scale visual model convolution calculations, an outward search and filling operation is performed on these empty regions based on the nearest mean-filtered interpolation core. Specifically, a fixed-size sliding neighborhood window, such as a 5×5 pixel matrix window, is set and slid around the empty pixel in the depth channel layer. When a pixel with a depth value of 0 (i.e., a gap) is detected, the statistical algorithm automatically traverses all neighboring pixels with valid depth values within the 5×5 neighborhood window, sums these valid depth values, and calculates the arithmetic mean. Then, the resulting average depth value is used as compensation data to smoothly fill the center of the gap. This process effectively repairs geometric holes caused by high reflectivity or occlusion. The above channel cascading and mean interpolation process satisfies the following fusion mapping formula:
[0025] In the above formula, The final output is located on the x and y coordinates. and The set of fused pixel tensors containing features of four channel dimensions; This represents the red, green, and blue primary color vectors of the underlying layer of the synchronized color image at this coordinate position; This represents an orthogonal concatenation operation performed along the channel dimension of a tensor; Represents a synchronous depth point cloud data array; and Represents the coordinate system index of the spatial topological mapping resolved from the depth pixel projection matrix; This represents the mathematical operator function that performs neighborhood mean filtering and outward smoothing interpolation on the gaps generated after mapping. After this complex matrix channel splicing and filtering dimensionality reduction reconstruction, the dimensionality reduction, compression, and solidification encapsulation of different source modal data were finally completed, generating toilet color depth aligned data with highly integrated color texture information and environmental depth three-dimensional spatial thickness attributes.
[0026] Step two involves extracting multi-layer convolutional semantic features and detecting anchor boxes in the color channels of the toilet color depth alignment data based on a large visual model. This yields a two-dimensional anchor box matrix for the toilet fixtures and a confidence score for the toilet fixture category. Correspondingly, although a unified multimodal benchmark data containing depth spatial thickness has been established after the preceding spatiotemporal synchronization alignment mapping process, the actual toilet cleaning environment is often filled with extremely cluttered visual interference elements such as patterned tiles, reflective mirrors, and intricate water pipes. Furthermore, the simple physical depth point cloud lacks high-level semantic information, making it difficult to directly distinguish between different types of toilet fixtures such as toilets and washbasins. To endow the cleaning robot with advanced visual semantic cognitive capabilities, thereby achieving precise isolation and positioning of specific work targets, this step is used to provide pixel-level logical masking basis and registration confidence prior conditions for the subsequent precise cropping of the 3D point cloud and six-degree-of-freedom pose calculation.
[0027] In one operable embodiment of this application, step two includes: stripping color patches from toilet color depth aligned data and inputting them into a large visual model to perform forward propagation multi-scale semantic feature mapping on the color patches to obtain a high-dimensional multi-scale semantic feature map; performing region response activation and candidate anchor box regression on the high-dimensional multi-scale semantic feature map to obtain an initial generated anchor box matrix and a candidate region feature tensor; performing clustering and filtering to retain overlapping and redundant boxes in the initial generated anchor box matrix to obtain a two-dimensional anchor box matrix for toilets; and performing multi-class conditional probability convergence on the candidate region feature tensor to obtain a toilet category confidence score.
[0028] The implementation process is as follows: First, a forward propagation multi-scale semantic feature mapping stage is performed. The algorithm receives a toilet color depth aligned data tensor with dimensions of 1920×1080 and four channels, output from the previous stage. In the graphics processing unit's memory, channel slicing is performed to decouple these dimensions logically. Specifically, the fourth-dimensional depth channel representing physical distance is stripped away, and a continuous spectral color layer composed of the bottom red, green, and blue primary color channels is extracted, with a tensor shape of 1920×1080×3.
[0029] Subsequently, the extracted color patch layers are directionally input into the onboard AI algorithm model library, activating and loading a large-scale visual model with pre-trained offline generalization weights. This large-scale visual model employs a visual foundation network based on a self-attention transformation architecture. Its specific architecture includes an input image block embedding layer, a feature encoder composed of multiple sets of multi-head self-attention modules stacked alternately, and a bottom-layer multi-layer convolutional backbone network. The network parameters of this large-scale visual model, such as weights and biases, are obtained in a cloud data center using a high-definition image training set covering tens of thousands of real bathroom and indoor environments. This is achieved by calculating the cross-entropy loss between the actual output and the manually labeled ground truth, and then updating it through hundreds of thousands of gradient descent iterations using a backpropagation mechanism. At this stage, the convolutional structure and self-attention layers within the large-scale model are used to perform forward propagation mapping operations on the color patch layers. The bottom-layer convolutional kernels are responsible for extracting high-frequency geometric features such as tile joints and equipment edges, while the high-layer self-attention encoders are responsible for capturing abstract semantic features such as the global structured texture of the toilet and the curved reflective properties of ceramic surfaces. After multi-level downsampling and feature aggregation operations, pyramid response features are finally generated in feature spaces of different resolution scales. For example, the original spatial resolution of 1920×1080 is downsampled and mapped to multidimensional tensors of different scales such as 120×67 and 60×34, and then uniformly output as a high-dimensional multi-scale semantic feature map with a channel depth of 256.
[0030] After extracting rich semantic features, the algorithm proceeds to the region response activation and candidate anchor box regression stage for the high-dimensional multi-scale semantic feature map. Based on the generated high-dimensional multi-scale semantic feature map, a large number of preset reference boxes with different aspect ratios and base scales are densely distributed at the coordinates of each pixel point at different feature levels. For example, the base pixel scales are set to 64, 128, and 256, and three aspect ratios of 1:1, 1:2, and 2:1 are assigned, respectively. Subsequently, the region generation network activation operator is called for processing. The architecture of this operator includes a 3×3 spatial sliding convolutional layer and two parallel 1×1 fully connected convolutional branches, which perform binary classification and spatial regression, respectively. The algorithm scans and compares the pixel gradient peak regions of the high-dimensional multi-scale feature map one by one, and calculates the object score of each preset reference box containing the foreground work target through the classification branch, that is, the preliminary probability of the existence of toilet equipment. By setting a score threshold, such as 0.7, the target pool output of the high response segment is truncated as the candidate region feature tensor. Simultaneously, for these initially identified target object domains, a regression branch is used to further calculate the nonlinear regression offset of the target's true boundary relative to the preset reference box parameters. This calculation process corrects the positional boundaries of the candidate boxes using the following spatial regression compensation formula: ; ; ; ,in, , , , These represent the horizontal axis coordinates, vertical axis coordinates, reference width, and reference height of the geometric center point of the preset reference frame, respectively. , , , This represents the nonlinear regression offset parameter calculated by the region generation network in four degrees of freedom: center bias and scaling. , , , This represents the final center coordinates and pixel size of the predicted bounding box after compensation and correction. Using this computational logic, a set of rectangular pixel coordinates that precisely enclose the potential target can be generated, outputting a generalized initial generated anchor box matrix.
[0031] Next, a clustering and filtering phase is performed on the overlapping redundant boxes in the initially generated anchor box matrix. Due to the dense distribution mechanism, for the same toilet device in the scene, the initially generated anchor box matrix often generates a large number of overlapping redundant boxes clustered in the same physical area. After receiving this matrix, a non-maximum suppression algorithm is applied. First, all bounding boxes in the cluster are sorted in descending order according to the previously calculated gradient response object score, and the one with the highest score is extracted as the benchmark test box. Then, the spatial intersection-union ratio (IUR) parameter between the remaining bounding boxes and the benchmark box is calculated, which is the intersection area divided by the union area. An IUR filtering threshold is set, such as 0.5. If the IUR of a bounding box and the benchmark box is greater than 0.5, it means that both are pointing to the same physical target and the current box confidence is low, so it is removed from the matrix as a redundant interference item. After iterative looping, only the unique bounding box with the highest gradient response confidence is retained in each overlapping cluster. For example, the original six redundant frames generated around the toilet are simplified into a single frame, whose two-dimensional coordinates are precisely locked from the top left corner (x-coordinate 420, y-coordinate 350) to the bottom right corner (x-coordinate 880, y-coordinate 920). Through this clustering and filtering process, a clean and accurate final two-dimensional anchor frame matrix for the toilet is output.
[0032] Finally, a multi-class conditional probability convergence stage is performed on the candidate region feature tensor. The previously acquired candidate region feature tensors are received in parallel, downsampled and flattened using spatial region-of-interest pooling, and then input into a fully connected classifier network attached to the tail of the large visual model. This classifier network consists of multiple layers of linear mapping matrices, used to map the extracted local high-dimensional features to specific categories. In this stage, a normalized exponential function mechanism is used to exponentially map and scale the logistic regression values within the feature tensor. This process aims to evaluate the independent conditional probability distribution of K preset scene objects (e.g., K equals 4, representing categories such as toilets, urinals, sinks, and trash cans), and its convergence calculation satisfies the following formula: ,in, This represents the total number of pre-defined scenario task categories to be identified; This represents the raw, unscaled linear logistic regression score output by the fully connected classifier network for the i-th target class; The exponential operation is based on the natural logarithm constant, which is used to convert all scores into positive numbers and amplify the proportion of high-response features; the denominator is the sum of the exponential scores of all K categories, which serves as a global normalization constant. This represents the final independent conditional probability value of the current candidate region object belonging to the i-th specific physical category, such as a toilet. Using this formula, a classification decision confidence probability sequence consisting of floating-point numbers in the range of 0 to 1 is calculated. For example, the probability of belonging to a toilet is calculated as 0.94, and the probability of belonging to a urinal is 0.03. Finally, this probability sequence is packaged and output to obtain the complete toilet category confidence score.
[0033] Step three involves using the 2D anchor frame matrix of the toilet fixtures as a pixel space mask to perform region-of-interest point cloud cropping, statistical filtering for noise reduction, and clustering segmentation on the depth channel of the toilet color depth alignment data to obtain the 3D point cloud for toilet fixture isolation. It's understandable that after the preceding computational processing, although the visual base network successfully outlines the target on the 2D image plane and obtains high-precision class confidence probabilities, the actual physical space of a bathroom inevitably contains a large number of redundant environmental elements such as background wall tiles, floor structures, and water pipes within the area enclosed by the 2D bounding box. Directly inputting the entire scene's 3D point cloud data into the subsequent pose registration process would not only generate an extremely large computational burden but also cause severe deviations in the pose matching algorithm due to significant background noise, making it impossible to obtain the true spatial coordinates of the equipment. This step is used to utilize the acquired high-confidence 2D anchor frame as a precise pixel space mask, and cross-modal guidance to accurately crop the depth channel containing spatial thickness attributes. Then, through high-frequency statistical filtering and 3D manifold clustering technology, the free flying points caused by reflection and the tightly adhered environmental background in physical space are peeled off layer by layer, and finally, an extremely pure 3D point cloud of toilet furniture with independent geometric topology is extracted.
[0034] Figure 4 This is a flowchart of step three in the toilet recognition method based on a depth camera and a large visual model according to an embodiment of this application. Figure 4 As shown, in one operable embodiment of this application, step three includes: step three-first, constructing the corner pixel boundaries of the toilet's two-dimensional anchor frame matrix into a binary logical mask, and performing pixel-level logical and projection truncation and hard deletion of background points on the depth channel layer stripped from the toilet's color depth alignment data to obtain a coarsely clipped three-dimensional point cloud; step three-second, performing high-frequency outlier noise statistical filtering and noise reduction analysis on the coarsely clipped three-dimensional point cloud to obtain a smooth three-dimensional point cloud; step three-third, performing manifold clustering and lossless segmentation of the adhered background on the smooth three-dimensional point cloud to obtain a toilet-isolated three-dimensional point cloud.
[0035] The implementation process is as follows: First, pixel-level logical and projection truncation and background point hard deletion are performed. The control motherboard receives the toilet color depth alignment data with four-dimensional channels generated in the previous processing stage, and accurately extracts the fourth-dimensional depth channel layer that specifically represents the physical three-dimensional absolute spatial coordinates according to its internal data registry protocol. Next, based on the unique and clean two-dimensional anchor frame matrix of the toilet received in parallel, its precisely locked recognition box, such as the bounding box of the locked toilet, is extracted. The geometric pixel coordinate boundaries of the four corner points of this effective recognition box are extracted, which correspond to the horizontal and vertical extreme values in the two-dimensional image space. For example, the coordinates of the minimum value at the top left corner are 420 on the horizontal axis and 350 on the vertical axis, and the coordinates of the maximum value at the bottom right corner are 880 on the horizontal axis and 920 on the vertical axis. The above corner point boundaries are constructed into a two-dimensional binary logical mask matrix of the same resolution in the device's computing memory. The pixel values inside the mask box are set to 1, and the pixel values outside the mask box are set to 0.
[0036] Subsequently, the logical mask is overlaid on the surface of the stripped depth channel layer, and pixel-level logical AND operations are performed for projection comparison. For invalid background pixels falling outside the calculation boundary mask, since their corresponding logical value is 0, the algorithm directly performs a hard truncation deletion operation on the 3D depth points they index, thereby instantly eliminating a large amount of irrelevant background environment data at the macro level; only the set of valid physical point coordinates within the mask range with a logical value of 1 is aggregated and collected, and the output is encapsulated as a coarsely clipped 3D point cloud for the high-probability toilet target area. This processing process conforms to the following projection truncation logical condition formula:
[0037] in, This represents a coarsely cropped set of 3D point clouds that has been initially isolated and preserved. Represents the horizontal coordinates corresponding to a two-dimensional pixel grid. with vertical coordinate The effective physical lattice vector in three-dimensional space at the location; The depth channel layer represents the entire spatial view extracted from the alignment data; and The horizontal and vertical pixel coordinates of the minimum boundary at the top left corner of the two-dimensional anchor box; and The horizontal and vertical pixel coordinates of the maximum boundary at the lower right corner of the two-dimensional anchor box; This represents the logical AND Boolean operators. The formula strictly limits the retention to only the 3D point cloud located within the decision window boundaries.
[0038] Next, the statistical filtering and noise reduction analysis of high-frequency outliers is performed. The initially extracted, coarsely trimmed 3D point cloud data array is received. Due to multipath reflection of light on the ceramic surface of the bathroom or fluctuations in the underlying hardware of the photosensitive element, a large number of outliers and extremely high-frequency detached points still exist within this point cloud. Therefore, a multidimensional spatial topology search tree structure based on an approximate nearest neighbor search algorithm is established in the device's cache for this spatial point sequence. Each reference coordinate point in the dataset is traversed, and the tree-like index structure is used to retrieve its 's' nearest neighbors within a set 3D physical radius, for example, a preset number of nearest neighbors 's' of 50. The Euclidean physical distance from the reference coordinate point to these 50 nearest neighbors is calculated, and the arithmetic mean of these 50 distances is obtained. Based on the physical spatial distribution laws of nature, it is determined that the average distance distribution of all coordinate points conforms to a Gaussian normal probability model. Therefore, the expected value of the global distance base (i.e., the overall mean) and the standard deviation reflecting the dispersion of the data distribution (i.e., the fluctuation limit) are statistically calculated.
[0039] Subsequently, the mean Euclidean distance calculated for each coordinate point is evaluated one by one. When the mean distance of a reference point exceeds the set upper limit of the Gaussian statistical threshold (this upper limit is composed of the global expectation plus a set multiple of the standard deviation, for example, setting the multiplier factor to twice the standard deviation), the point is determined to be outside the main structure of the target object and belongs to the interference flying point, and is completely removed from memory. After comprehensive cleaning and filtering, a smooth 3D point cloud with highly dense surface details is output. This statistical filtering mechanism follows the following inequality decision formula:
[0040] in, This represents a smooth 3D point cloud output after removing high-frequency flying points; Represents the current reference space coordinates point that is being evaluated; This represents the reference point and its s nearest neighbors. The mean Euclidean distance between them; This represents the calculation of the L2 norm distance between three-dimensional vectors; This represents the expected global average distance calculated from all point cloud data in the current batch. Represents the global distance standard deviation; The scaling multiplier variable represents the pre-defined control over the rigor of the scaling, such as 2.0. This formula ensures that all retained point clouds closely adhere to the physical entity surface.
[0041] Finally, the manifold clustering and lossless segmentation of the adherent background are performed. A smooth 3D point cloud with free-floating points removed is received. A blank cluster record table is initialized in the device's runtime memory, and a 3D coordinate point without attribute labels is randomly extracted from the point cloud set as the first cluster growth starting point, i.e., the seed point. Using the already constructed regional spatial tree topology, nearest neighbor search and comparison are performed radially outward from the seed point. When the 3D physical distance between a found neighboring point and the seed center is less than a preset critical connectivity threshold, the neighboring point is assigned to the same manifold cluster. This critical connectivity threshold is strictly set based on the standard process seam drop in real-world scenarios. For example, there is usually a gap of at least 3 to 5 centimeters between the edge of a toilet and the back wall or the bottom floor tiles; therefore, the connectivity threshold is set to 0.04 meters, or 40 millimeters.
[0042] Subsequently, newly incorporated neighboring points are promoted to new seed centers, and the search continues outward in a relay-like manner, recursively exhibiting a connected, spreading growth state in the spatial dimension. When this spreading continues until no new connected points can be found within the connectivity threshold radius (40 mm), it indicates that a real physical spatial fault gap has been encountered between the toilet entity and the surrounding walls or countertops in three-dimensional space. At this point, the recursive growth of the current cluster is stopped; then, the above process is repeated from the remaining unprocessed scattered points. After a comprehensive scan, the original point cloud is cut and segmented into multiple independent geometric object patch clusters. For the several patch clusters generated by the stripping, the volume element space point count is further cleared and compared. A reasonable point count range is set, such as determining that the point cloud of the target toilet should be between 20,000 and 100,000 points. Support carrier environment patches with too many points, such as background wall sections and floor sections containing hundreds of thousands of points, are removed, while small patches, such as water pipe bracket fragments containing only a few hundred points, are filtered out. Through this series of geometric stripping and clearing mechanisms, the isolated 3D point cloud of toilets, each possessing its own complete geometric topological shape, is accurately extracted and output. The mathematical expression for this clustering process is as follows:
[0043] in, A three-dimensional point cloud representing the final filtered output of the purified toilet fixture. This represents the total number of valid physical patches that meet the target number of points. This represents performing a union operation on all valid cluster patches that meet the criteria to form a complete entity; This represents the central seed node that is in an active, spreading state within the m-th cluster. Represents neighboring candidate physical points within the topological space whose connectivity is being searched; This represents the critical connectivity spatial physical distance threshold pre-set according to the process specifications, namely 40 millimeters as mentioned earlier. Using this clustering screening model, not only is the smooth geometry of the ceramic surface perfectly preserved, but also the physically adhering physical interference is spatially severed without physical damage.
[0044] Step four involves using the toilet category confidence score as a registration weighting factor to perform principal component analysis and iterative nearest-point pose registration on the isolated 3D point cloud of the toilet to obtain the six-DOF pose parameters of the toilet. It's understandable that after the preceding processing steps, although a very pure 3D point cloud of the toilet with independent geometric topology has been successfully extracted from the complex bathroom environment, and a high confidence probability score indicating that the target device belongs to a specific category has been obtained, the 3D point cloud at this point is merely a large set of discrete and disordered coordinate points in space, only representing the geometric contour of the target object's surface. For automated robotic arms requiring precise fitting and deep interaction, the lack of a coordinate system origin and rotation direction representing the spatial pose of the target object makes it impossible to directly generate inverse kinematic control commands. Furthermore, traditional 3D spatial registration algorithms often rely on blind iterative searches, easily getting trapped in local optima without a good initial relative positional relationship, and fail to integrate the semantic confidence output of the visual model into the spatial geometric calculations, resulting in insufficient robustness of the registration. This leads to the use of this step to accurately calculate the six-degree-of-freedom pose parameters that can directly guide the robotic arm to perform interactive operations in three-dimensional space.
[0045] In one operable embodiment of this application, step four includes: extracting the geometric centroid and initial axis of the toilet isolation 3D point cloud based on principal component analysis to obtain geometric centroid and principal axis features; retrieving a standard template point cloud from a pre-set database based on the category label pointed to by the toilet category confidence score, and using the geometric centroid and principal axis features as initial alignment values to perform confidence-weighted iterative nearest point registration optimization on the toilet isolation 3D point cloud and the template point cloud to obtain the optimal mapping transformation matrix; and performing homogeneous matrix decomposition and pose parameterization extraction on the optimal mapping transformation matrix to obtain the toilet's six-degree-of-freedom pose parameters.
[0046] The implementation process is as follows: First, a stage of geometric centroid and initial axis extraction based on principal component analysis is performed. The bottom-level computing unit receives the clean and unadhesive 3D point cloud of the toilet, which contains tens of thousands of effective physical points representing the surface contour of the target object, such as the toilet. The algorithm first iterates through and accumulates the coordinates of this point cloud in 3D Euclidean space, and divides it by the total number of points in the cloud to calculate the mean of the spatial center of the geometric shape. For example, after counting 50,000 effective points, the mean of the geometric center of the toilet point cloud is calculated to be 450 mm on the horizontal axis, -200 mm on the vertical axis, and 800 mm on the depth axis. This mean coordinate is used as the initial centroid of the target in the scene. Subsequently, based on this centroid coordinate, a centering and mean-reduction process is performed on all discrete coordinate points in the point cloud, that is, the centroid coordinate is subtracted from the original coordinate of each point. The significance of this step in algebraic geometry lies in translating the origin of the point cloud's coordinate system to the object's own centroid, completely eliminating its translational bias in global space, and making subsequent attitude analysis completely independent of the object's position.
[0047] After centering, a covariance array matrix reflecting the spatial distribution variance characteristics is constructed using all mean-reduced 3D coordinate vectors through matrix multiplication. The construction of this covariance matrix follows the mathematical formula below: ,in, This represents the calculated 3×3 spatial covariance matrix, which contains information about the degree of dispersion of the point cloud in three orthogonal directions. The total number of valid physical points contained in the 3D point cloud representing toilet isolation is, for example, 50,000. The three-dimensional column vector coordinates of the i-th point in the point cloud set; Represents the coordinates of the three-dimensional column vector of the geometric centroid of the point cloud set obtained from previous calculations; superscript This represents performing a matrix transpose operation on a column vector, transforming it into a row vector, thereby making... A 3×3 local covariance block matrix is generated. After summing and averaging, the covariance array matrix is subjected to deep feature deconstruction using the singular value decomposition (SVD) algorithm to solve for its corresponding eigenvalues and eigenvectors. Among the three orthogonal eigenvectors generated by SVD, the eigenvector with the largest corresponding eigenvalue is extracted. In a physical space sense, this eigenvector represents the direction in which the point cloud data is most dispersed and extends the longest, corresponding to the geometric principal direction of the target, such as the front-to-back symmetrical axis of a toilet or the normal vector of the principal tangent plane. Finally, the calculated 3D centroid coordinates and this principal direction vector are encapsulated and output as a coarse representation of the object's pose state, showing the geometric centroid and principal axis features.
[0048] Next, the process enters the confidence-weighted iterative nearest-point registration optimization stage. In this stage, the device's pre-built database is accessed first. This database is a standard geometry library of Building Information Modeling (BIM) stored in the device's internal solid-state storage. Its data originates from point cloud sets converted from non-destructive mesh models obtained by scanning various standard sanitary ware entities using a high-precision industrial 3D scanner before leaving the factory, or directly exported from computer-aided design 3D software. Each template point cloud within this database has its origin pre-anchored at the geometric center, and its principal axis aligned with the standard coordinate axes. Based on the highest probability category label indicated by the toilet category confidence score passed from the previous core step (e.g., if the label explicitly indicates the current target is a toilet, and the classification confidence score is 0.94), the algorithm precisely retrieves the corresponding toilet model's standard template point cloud from the pre-built database as the registration template set.
[0049] Subsequently, using the geometric centroid and principal axis features calculated in the previous stage as initial values for spatial transformation, a coarse rigid body translation and rotation matrix multiplication are performed to directly move the retrieved standard template point cloud from its default origin position and initially align it to the vicinity of the toilet-isolated 3D point cloud in the scene, ensuring that the spatial poses of the two are roughly consistent. After completing the coarse alignment, the core confidence-weighted iterative nearest-point optimization loop is initiated. In each calculation iteration, the algorithm searches for the closest corresponding point pair between the input target scene point cloud and the standard template point cloud based on the multidimensional spatial tree topology structure. Unlike traditional algorithms, this step innovatively uses the toilet category confidence score output by the large visual model as a global weight compensation factor for the registration error function. The core objective of this weighted iterative registration process is to minimize the sum of weighted Euclidean geometric distances, and its error objective function formula is as follows:
[0050] in, This represents the global registration error scalar that needs to be optimized to a minimum. The confidence score of the toilet category, which serves as a weighting compensation factor, is used as an example. If a specific value of 0.94 is substituted, a high confidence score amplifies the determining power of the registration on the overall pose, while a low confidence score reduces its sensitivity to fine-tuning to prevent divergence. This represents the total number of valid nearest neighbor matching pairs found. The coordinates of the i-th target point in the actual captured 3D point cloud of the toilet enclosure in the scene; The coordinates of the i-th model point in the standard template point cloud that matches the former; The variable represents the 3×3 spatial rotation matrix to be solved; This represents the 3×1 spatial translation vector variable to be solved. The algorithm iteratively solves for the optimal rotation matrix by calculating the cross covariance of point pairs and utilizing singular value decomposition. With translation vector When the overall error change calculated between two consecutive iterations is below an extremely small set threshold, such as 0.001 mm, the point cloud is considered to have achieved physical perfect fitting, and the loop stops. The final converged rotation matrix is then used to calculate this. With translation vector The matrix is reorganized, and the output is an optimal mapping transformation matrix with 4×4 dimensions. This matrix fully records the entire rigid transformation process from the standard attitude to the current real physical attitude.
[0051] In particular, during the standard spatial pose iterative reconstruction process, directly applying the toilet category confidence score output by the large visual model as a static, global linear scalar weight to the Iterative Closest Point (ICP) registration energy function will lead to a series of significant technical bottlenecks. Because a single scalar applies the exact same weight to all registration point pairs, this global uniformity prevents the algorithm from distinguishing the geometric reliability differences between different physical regions in the 3D point cloud. For example, the geometric discriminability and registration constraint of contour feature points such as toilet edges and tank corners are far higher than those of smooth, featureless ceramic sidewall surface points, but in traditional linear registration, the two are incorrectly treated as equals. Furthermore, since the mapping relationship between visual confidence score and spatial registration accuracy is not a simple linear one, direct use can lead to linear response mismatch: when the confidence score is in the intermediate ambiguity range of 0.5 to 0.7, the linear weights cannot generate sufficient decay to effectively suppress registration coordinate drift caused by incorrect template matching; while when the confidence score is extremely high, such as greater than 0.95, the linear weights lack an exponential amplification mechanism and cannot fully utilize this high confidence to accelerate the convergence of the pose matrix. More importantly, if the weight factor remains constant throughout the hundreds of registration iterations and cannot be dynamically and adaptively adjusted as the registration residual gradually decreases, it will inevitably lead to the same constraint strength in the early stage of registration (the coarse alignment stage that requires a tolerant search to avoid local optima) and the late stage of registration (the fine convergence stage that requires strict constraints to obtain extremely high accuracy), thus seriously interfering with the search for the global optimal spatial solution. To completely overcome the limitations of the aforementioned linear and static weights, this step is introduced to construct a nonlinear optimization mechanism that can simultaneously and sensitively perceive semantic confidence, local geometric curvature, and global temporal residual state. By dynamically assigning reasonable constraints to each physical space point during the iteration process, the robustness and absolute accuracy of the final calculated robotic arm pose parameters are significantly improved.
[0052] The specific implementation process of the nonlinear weight iteration mechanism based on point-by-point adaptation and dynamic evolution with iteration covers three consecutive precise calculation stages: local curvature deconstruction, three-factor coupling mapping, and dynamic weighted registration.
[0053] Based on this, in a preferred embodiment of this application, a standard template point cloud is retrieved from a pre-set database based on the category label pointed to by the toilet category confidence score. Using geometric centroid and principal axis features as initial alignment values, a confidence-weighted iterative nearest-point registration optimization is performed between the toilet-isolated 3D point cloud and the template point cloud to obtain the optimal mapping transformation matrix, including: For each point in the 3D point cloud of toilet enclosures, eigenvalue deconstruction of the local neighborhood covariance matrix is performed to obtain the local curvature of each point. It should be understood that only by accurately quantifying the geometric salience of the micro-region where each 3D coordinate point is located can higher alignment priority be given to salient regions such as edges and corners in subsequent registration. In specific implementation, a 3D point cloud of toilet enclosures containing tens of thousands of valid physical coordinates is received, such as the toilet point cloud containing 50,000 valid points extracted previously. For each independent point in this set, the previously constructed micro-principal component analysis framework is reused to extract the set of neighboring points within a preset physical radius centered on that point. Specifically, a pre-built multi-dimensional spatial topology search tree structure in memory is invoked, using the currently evaluated 3D physical point as the center and a preset small spatial retrieval physical radius, such as setting this radius to 15 mm and retrieval outward in a spherical pattern, thereby accurately capturing a subset of locally nearby coordinate points enclosed within this micro-spherical envelope, such as retrieving 35 closely connected neighboring points at the edge of the toilet. Subsequently, the algorithm iterates through and accumulates the 3D coordinate column vectors of all nearest neighbors within the micro-subset, dividing by the number of points to obtain the mean coordinates of the geometric center of this local micro-section region. After establishing this local center, the absolute coordinates of each point in the subset are subtracted from the mean coordinates of the center for local de-meaning, thereby translating the origin of the local coordinate system to the centroid of the micro-section. Finally, using these locally de-meaned relative spatial 3D column vectors, the outer product matrix is multiplied with their own transpose row vectors, and all generated matrices within the subset are accumulated and averaged, thus rigorously constructing the covariance matrix for this local neighborhood at the mathematical and physical level. Through this precise process of point-by-point scanning and micro-dimensional reduction, the algorithm transforms the abstract surface undulations of an object into a micro-local feature distribution pattern that can be analyzed by a quantifiable matrix. Subsequently, singular value decomposition is performed on the covariance matrix to obtain three eigenvalues reflecting the distribution characteristics of the local spatial point cloud. Based on these three eigenvalues, the local curvature of the point is calculated through specific algebraic ratio operations. The calculation process follows the following local curvature characteristic formula:
[0054] In the above formula, This represents the calculated scalar value of the local surface curvature at the i-th three-dimensional point; , and The three eigenvalues, which together represent the local neighborhood covariance matrix of the i-th point, are obtained by deconstructing it; and This is the smallest of the three eigenvalues, and its physical meaning directly corresponds to the spatial dispersion or thickness of the local surface along the normal vector direction. For example, when the scan point is located at the high curvature profile of the toilet rim, the eigenvalues might be calculated as 10.5, 8.2, and 4.1, which can be substituted into the formula to calculate the local curvature. The eigenvalue is approximately 0.18; however, when the scan point is located on the flat surface of the toilet tank, the eigenvalues may be 12.0, 11.5, and 0.2, and the calculated local curvature is only 0.008. This successfully transforms the abstract geometric surface undulations into precise quantitative indicators that can be directly used by the algorithm, laying the data foundation of physical geometric dimensions for breaking the global uniformity weight.
[0055] Based on the local curvature, toilet category confidence score, and global registration residual at each point, a three-factor coupling function is constructed for nonlinear mapping to obtain the adaptive dynamic curvature weight value for each point. Correspondingly, single-dimensional information is insufficient to handle complex on-site reflections and occlusion interference; a nonlinear mathematical model is needed to deeply fuse the macroscopic semantics, microscopic geometric curvature, and macroscopic temporal state of the registration process provided by the large visual model, thereby achieving intelligent weight allocation. In specific implementation, the algorithm receives in parallel the toilet category confidence score output from the visual stage (e.g., 0.94), the local curvature calculated in the previous stage, and the registration residual state recorded in system memory. Based on this, a three-factor coupling function is constructed, containing a Sigmoid family nonlinear mapping term, a geometric enhancement term, and a temporal decay term. The algebraic expression of this function is as follows:
[0056] In the above coupling weight formula, This represents the adaptive dynamic curvature weight value calculated specifically for the i-th registration point pair in the nth registration iteration; the first term of the formula is the Sigmoid nonlinear semantic mapping term, where... The original confidence score for the input toilet category is 0.94. This represents the preset confidence level decision threshold center point, set to 0.6. This represents the kurtosis control parameter used to adjust the width of the nonlinear mapping sensitivity range, such as a preset value of 10.0; the second term in the formula is the geometric enhancement term, where... For the imported local curvature, The preset curvature enhancement coefficient, for example, set to 5.0, is used to control the contribution of geometric features to the weights; the third term in the formula is the temporal residual attenuation term, where... This represents the global registration residual energy value at the end of the (n-1)th iteration. The ratio of the two values represents the initial global residual energy value when the algorithm is first started, and the ratio of the two values constitutes the normalized temporal evolution benchmark. The global temporal decay ratio applies the same scaling factor to all registration point pairs within the same iteration. Its core function is not to change the relative constraint priority between points—the differentiated constraint force between points is entirely determined by the local curvature enhancement terms calculated independently point by point in the first two terms—but rather to dynamically normalize and control the absolute magnitude of the registration energy function at the engineering numerical calculation level: when the global residual is large in the early stage of registration, the ratio is close to 1.0, thus maintaining the energy function in a high numerical magnitude range, ensuring that each matrix element has sufficient effective numerical precision when solving the weighted cross covariance matrix by singular value decomposition; as the iteration progresses and the residual gradually converges to approach zero, the ratio decays synchronously, thereby compressing the absolute value of the energy function to a reasonable range that matches the precision of floating-point operations, effectively preventing stability risks at the engineering implementation level such as covariance matrix condition number degradation, cumulative amplification of floating-point truncation error, and solver numerical oscillation caused by excessively small absolute residual values, thus ensuring that the entire iterative optimization process maintains robust stability of numerical solution in dozens or even hundreds of iterations. With the high curvature point located at the edge of the toilet, Taking 0.18 as an example, since the confidence level of 0.94 is much higher than the threshold of 0.6, the calculated value of the Sigmoid term is close to 0.967, and the geometric enhancement term is amplified to 1.9. If we are currently in the tenth iteration, the residual ratio drops to 0.2, and the final coupling weight of this point is approximately 0.367. Since its local curvature is only 0.008 and the geometric enhancement term is only 1.04, under the same global residual ratio of 0.2, its weight may only be 0.201. It can be seen that the approximately 1.83-fold weight difference between high-curvature edge points and flat surface points is entirely due to the synergistic effect of the point-by-point independent Sigmoid semantic mapping and local curvature geometric enhancement in the first two terms. This differentiated relative weight ratio will not be canceled out by the global proportional scaling of the third term during the singular value decomposition solution process, thus ensuring that the solver prioritizes aligning with key contour feature points with high geometric recognizability in each iteration. In this way, by using the Sigmoid nonlinear mapping to strongly suppress noise in the fuzzy region, fully release weights in the certain region, and take advantage of the greater registration contribution given by high curvature points, combined with the dynamic normalization and stable control of the energy function's numerical magnitude by the global residual decay term, the geometric accuracy of point-by-point differential constraints and the numerical computational stability of the long-range iterative process are perfectly balanced.
[0057] Based on the adaptive dynamic curvature weight value of each point, point-by-point dynamic weight registration optimization is performed on the template point cloud and the toilet isolation 3D point cloud to obtain the optimal mapping transformation matrix. That is, the carefully calculated differential weights from the previous steps need to be truly applied to the physical solution process of spatial alignment to guide the matrix transformation towards the direction with the most significant geometric features and the most reliable semantics. In specific implementation, the algorithm retrieves the standard template point cloud of the corresponding toilet model extracted from a pre-set database, and uses the initial geometric centroid and principal axis as the initial values for spatial transformation. In the iterative solution loop, the aforementioned point-by-point dynamic weights completely replace the global static scalar in the traditional ICP energy function, performing a differential weighted summation of the squared Euclidean distance between the input actual observation point and the template reference point. The energy objective function formula for this weighted registration optimization process is as follows:
[0058] in, This represents the total weighted registration residual energy that needs to be minimized by the algorithm core in the nth iteration; This represents the total number of valid registration point pairs found and established in this round; This represents the adaptive dynamic curvature weight value specific to the i-th point pair in this round. This represents the column vector of the i-th spatial observation point in the 3D point cloud of the toilet enclosure extracted from real-world data within the scene; The column vector representing the i-th reference point in the standard template point cloud that matches the former; This represents the three-row, three-column spatial rotation matrix to be optimized. This represents the three-row, one-column spatial translation vector to be solved; This represents calculating the square of the Euclidean distance in three-dimensional space. The computation engine continuously adjusts... and This causes the overall energy function value to approach its minimum value until the energy difference between two iterations is less than a preset convergence threshold, such as 0.001 mm. Since highly recognizable toilet edge points have a dominant weight in the summation formula, the solver prioritizes and precisely aligns these key geometric frameworks, allowing for reasonable tolerances in flat areas. This significantly improves the robustness of the final output optimal mapping transformation matrix; practical experience shows that this mechanism can reduce small solution errors in the rotation matrix by approximately 30% to 50%. This leap in pose accuracy directly reduces the uncertainty envelope volume in the subsequent collision avoidance safety verification stage, enabling the robotic arm to safely and quickly perform flushing and brushing actions at an absolute spatial position extremely close to the fragile ceramic surface, drastically improving the coverage and efficiency of automated cleaning operations.
[0059] Finally, the homogeneous matrix decomposition and pose parameter extraction stages are performed. The algorithm receives a homogeneous optimal mapping transformation matrix with a dimension of four rows and four columns. In order to make it recognizable and callable by the underlying motion control algorithm of the robotic arm, it is deconstructed into a more physically meaningful parameter form. First, the orthogonal rotation components (i.e., rotation matrix) corresponding to the first three rows and first three columns and the translation component corresponding to the origin of the coordinate system corresponding to the first three rows and fourth column are extracted directly. The three values of the horizontal coordinate, vertical coordinate, and depth coordinate in the translation component are directly extracted as the absolute description of the target's spatial position, such as the extracted spatial position as 452.3 mm on the horizontal axis, -198.6 mm on the vertical axis, and 803.1 mm on the depth axis. For the extracted rotation components, based on the classical geometric trigonometric transformation theory, they are transformed into Euler angle representations that are more in line with the inverse kinematics calculation of the robotic arm, namely yaw angle, pitch angle, and roll angle. The nonlinear mapping solution process from the rotation matrix to Euler angles follows the following consecutive inverse trigonometric function formulas: ; ; Among them, variables , , , , These represent the specific numerical elements of the corresponding rows and columns in the split 3×3 rotation matrix, for example... (Represents the element in the third row and first column of the matrix). This represents the bivariate arctangent mathematical function, which can accurately determine the specific quadrant in which the calculated angle lies based on the signs of the two input parameters, thereby avoiding the ambiguity of positive and negative spans produced by the traditional arctangent function; This represents the calculated pitch angle attitude around the vertical longitudinal axis. The yaw angle attitude represents rotation about the depth optical axis; This represents the roll angle attitude, which is rotated around the horizontal axis. Through rigorous trigonometric calculations, three specific radian values representing the rotational attitude were obtained: a roll angle of 0.05 radians, a pitch angle of -0.12 radians, and a yaw angle of 0.86 radians. Finally, the extracted three-dimensional absolute position coordinate signals and the calculated three-dimensional Euler angle attitude signals are fused and encapsulated at a high level, forming an array containing full-dimensional position and orientation information. The final output of this array is a six-DOF pose parameter for the toilet, possessing the potential for full-space attitude adjustment commands.
[0060] Step five involves using a hand-eye calibration homogeneous transformation matrix and a preset safety distance threshold to perform world coordinate system transformation and anti-collision differential verification on the six-DOF pose parameters of the toilet fixture to obtain collision-free operating coordinates. In other words, although the highly accurate six-DOF pose parameters of the toilet fixture have been precisely calculated after the preceding complex calculations, it's important to note that these pose parameters are only based on the relative position and attitude relationships within the local optical coordinate system of the depth camera. For the automated cleaning robot responsible for executing the final action, the underlying motion planning commands of its robotic arm highly depend on the global world coordinate system established on the physical reference origin of the robot chassis. Without a unified coordinate system transformation, the robotic arm will not be able to recognize the position command. Furthermore, ceramic sanitary ware such as toilets and washbasins are fragile and easily scratched. If, after transformation, the robotic arm's end-effector is directly guided to move in absolute coordinates close to the target physical surface, rigid impacts can easily occur due to the robotic arm's motion inertia or motor execution errors, leading to equipment damage. The purpose of this step is to accurately map local visual coordinates to the robot's global motion coordinate system, and to perform physical-level anti-collision reverse compensation and kinematic envelope verification through fusion process safety distance. This completely opens up the technical closed loop from multimodal visual perception and positioning to safe, accurate, automated physical proximity operation, ensuring that the final output is a set of working coordinates that can directly drive the equipment and are absolutely collision-free.
[0061] In one operable embodiment of this application, step five includes: based on the hand-eye calibration homogeneous transformation matrix, performing robot base coordinate system pose mapping on the camera system homogeneous matrix constructed from the six-degree-of-freedom pose parameters of the toilet to obtain world coordinate system pose data; based on a preset safety distance threshold along the normal vector direction of the target surface, performing reverse incremental compensation and robotic arm reachability envelope verification on the position vector of the world coordinate system pose data to obtain collision-free space operation points.
[0062] The implementation process is as follows: First, the robot's base coordinate system pose mapping stage is performed. The underlying main control computing unit receives the six-DOF pose parameters of the toilet, which include the absolute orientation features of the target, output from the previous stage. These parameters internally contain a three-dimensional Euler angle pose array and a three-dimensional position coordinate vector. Based on the principles of three-dimensional rigid body kinematics, the extracted Euler angle components, such as the roll angle of 0.05 radians, pitch angle of -0.12 radians, and yaw angle of 0.86 radians calculated in the previous step, are first reconstructed into a standard three-row, three-column rotation matrix with orthogonal properties through internal inverse trigonometric function matrix restoration operations. Simultaneously, the extracted spatial position coordinates, such as the horizontal axis of 452.3 mm, the vertical axis of -198.6 mm, and the depth axis of 803.1 mm, are used as spatial translation vectors. Next, the 3×3 rotation matrix is embedded in the top left corner of a 4x4 blank matrix, the translation vector is embedded in the first three rows of the fourth column of the matrix, and three 0s and one 1 are added to the last row of the matrix, thus combining the two to construct a homogeneous matrix of the camera system with four dimensions.
[0063] Subsequently, the pre-burned and stored hand-eye calibration homogeneous transformation matrix is extracted from the underlying parameter configuration solid-state storage module. This hand-eye calibration homogeneous transformation matrix is also a 4x4 high-dimensional rigid body transformation matrix, whose core function is to characterize the fixed geometric relative relationship between the optical axis center inside the vision sensor and the world physical coordinate system (i.e., the base coordinate system) where the robot's reference chassis is located. This matrix is obtained during the robot's factory assembly or initial calibration stage, using a standard black and white checkerboard calibration board, controlling the robotic arm to perform high-frequency visual imaging under multiple known kinematic postures, and then accurately solving for the joint deviations by solving multiple sets of equations based on dual quaternions or nonlinear optimization, and then saving the matrix. After obtaining this underlying calibration matrix, the computation control unit strictly uses the matrix multiplication rule in linear algebra to multiply the camera system homogeneous matrix representing the relative position of the target by the hand-eye calibration homogeneous transformation matrix, thereby mathematically realizing the transformation and projection of the target object from the virtual visual optical space to the real physical working space that the robotic arm can understand. The mathematical mapping expression of this core coordinate system transformation process is as follows: ,in, It represents a homogeneous transformation matrix of a four-by-four camera system, which is completely established in the local optical coordinate system of the depth sensor and is constructed by reconstructing and splicing the pose parameters of the preceding toilet fixtures with six degrees of freedom. This represents a high-precision hand-eye calibration homogeneous transformation matrix retrieved from the underlying persistent configuration file, which contains physical installation offset compensation information for the sensor. This represents the four-by-four homogeneous transformation matrix output after the computation, which is now fully mapped onto the robot's global chassis reference origin. After completing the matrix multiplication mapping, from this... The transformed rotation matrix components and translation position components are extracted from the matrix and then repackaged and organized into world coordinate system pose data. The generated data, such as the transformed global world coordinates of 810.2 mm x, 50.6 mm y, and 605.4 mm d, can now be directly used for the underlying guidance of servo motors, possessing extremely high physical execution value.
[0064] Next, the process proceeds to the reverse incremental compensation and robotic arm reachability envelope verification stage. After acquiring the world coordinate system pose data, the three-dimensional spatial absolute coordinate vector and the rotation matrix representing the target spatial orientation are extracted. In actual automated toilet cleaning operations, if the robotic arm end effector is directly controlled to move to the absolute coordinate point of the physical surface, the actuator itself will inevitably have a rigid collision with the ceramic surface. To avoid this risk, the built-in physical obstacle avoidance and tool parameter library is called in real time to extract the preset safety distance threshold corresponding to different cleaning tool heads such as high-pressure wide-width water spray heads, sewage vacuum suction heads, or soft-bristle rotating brush heads currently mounted on the robotic arm end. This preset safety distance threshold is a scalar length limit parameter with clear physical anti-collision significance. It defines the minimum allowable physical distance between the robotic arm end hardware device and the toilet ceramic surface to maintain normal spraying or brushing operations. This threshold is determined by extensive testing of the optimal working fluid dynamic distance of the cleaning tools before leaving the factory, and by taking into account the maximum braking overshoot error distance of each joint servo motor. For example, when the currently identified mounting tool is a high-pressure wide-angle water jet, to ensure that the water flow fan just covers the inner wall of the toilet and absolutely avoids the metal nozzle from touching the ceramic edge, a preset safety distance threshold is dynamically extracted and set to 150 mm. Subsequently, the algorithm extracts the column vector data corresponding to the depth axis direction from the rotation matrix of the pose data. This unit vector physically represents the normal vector direction of the target object (such as the soiled surface of the toilet), which is the positive spatial direction perpendicular to the target's cross-section. Next, based on the position vector base coordinates of the world coordinate system pose data, the algorithm instructs the target to perform a reverse incremental compensation operation along this normal vector direction. In simpler terms, this forces the planned robotic arm end-effector to retract a safe distance along a trajectory perpendicular to the toilet surface. This spatial reverse collision avoidance compensation operator is equivalent to the following three-dimensional vector compensation formula:
[0065] in, This represents a high-precision column vector of absolute physical three-dimensional coordinates of the target surface extracted from world coordinate system pose data; The preset safe distance threshold in scalar form represents the value of 150 mm set above. The symbol represents the three-dimensional unit normal vector of the target surface stripped from the attitude rotation matrix, representing the positive vertical axis of the interactive interface; the minus sign indicates that a spatial retreat operation is performed along the opposite direction of the normal vector to ensure that the interface is far away from the physical collision interface. This represents the initial coordinates of the safe hovering point in the space, calculated after reverse incremental compensation.
[0066] After calculating the initial safety coordinates, the underlying motion planning layer must execute rigorous differential verification logic. This verification logic inputs the compensated spatial point coordinates into the robotic arm's inverse kinematics solver module for simulation calculations. On one hand, it verifies whether the torsional angles of each mechanical joint required for the point coordinates and corresponding posture exceed the boundaries of the robotic arm's own physically reachable working envelope, preventing singularities or dead angles that could cause jamming. On the other hand, it calls the established static bathroom environment collision box model to verify whether the proposed generated link geometry sweep trajectory will cause spatial interference or conflict with pre-set inherent environmental obstacles in the space (such as glass partitions in the bathroom, side water supply pipes, and other inaccessible areas). If a conflict warning occurs, the algorithm sets a step size around the normal to fine-tune the compensation angle, re-substitutes it into the verification until there is no interference. After this series of anti-collision compensation and inverse kinematics differential verification filtering based on strict physical boundaries, all parameters are finally locked and packaged, outputting a set of toilet collision-free operation coordinates with absolute spatial coordinate positioning information, full posture tilt angle guidance information, and collision immunity certification. At this point, the underlying servo driver generates smooth motion level pulses based on the final work coordinates, controlling the robotic arm to safely, quickly, and accurately approach the target area to carry out automated maintenance work.
[0067] In summary, the toilet recognition method based on depth camera and large visual model according to the embodiments of this application is explained, aiming to solve the problems of poor multimodal data fusion quality, high noise in point clouds, low pose estimation accuracy, and easy collisions during robotic arm operation in the prior art. First, the color image and depth point cloud are spatiotemporally aligned and fused to obtain high-quality color and depth aligned data, effectively overcoming the spatiotemporal bias between multiple sensors. Then, features are extracted using the large visual model to obtain two-dimensional anchor boxes and category confidence scores, and innovatively, the two-dimensional anchor boxes are used as pixel space masks to guide the depth layer for precise cropping and filtering noise reduction, achieving efficient separation of the target from the cluttered background. Next, the category confidence score is used as a registration weighting factor to perform pose registration on the isolated three-dimensional point cloud, avoiding registration getting trapped in local optima, thereby outputting high-precision six-DOF pose parameters. Finally, anti-collision differential verification is performed by combining the hand-eye calibration matrix and the preset safe distance threshold to calculate the collision-free operation coordinates, completely opening up the technical closed loop from high-precision multimodal recognition to safe and automated operation of the robotic arm.
[0068] Figure 5This is a block diagram of a toilet recognition system based on a depth camera and a large visual model, according to an embodiment of this application. Figure 5 As shown, the toilet recognition system 100 based on a depth camera and a large visual model according to an embodiment of this application includes: a sensor data pre-alignment module 110, used to perform spatiotemporal alignment and fusion of the acquired original sensor data stream based on an intrinsic and extrinsic parameter coordinate system to obtain toilet color depth alignment data, wherein the original sensor data stream includes a toilet color image and a toilet depth point cloud; a semantic extraction and target detection module 120, used to perform multi-layer convolutional semantic feature extraction and anchor box target detection on the color channels in the toilet color depth alignment data based on a large visual model to obtain a two-dimensional anchor box matrix of toilets and a toilet category confidence score; and a three-dimensional point cloud generation module 130, used to generate toilets... Using a two-dimensional anchor frame matrix as a pixel space mask, the depth channel in the toilet color depth alignment data is cropped for region of interest point cloud, statistically filtered for noise reduction, and clustered to obtain the toilet-isolated 3D point cloud; the pose parameter generation module 140 is used to perform principal component analysis and iterative nearest point pose registration on the toilet-isolated 3D point cloud using the toilet category confidence score as a registration weighting factor to obtain the toilet's six-DOF pose parameters; the operation coordinate generation module 150 is used to perform world coordinate system transformation and anti-collision differential verification on the toilet's six-DOF pose parameters through hand-eye calibration homogeneous transformation matrix and preset safety distance threshold to obtain the toilet's collision-free operation coordinates.
[0069] Here, those skilled in the art will understand that the specific operations of each step in the toilet recognition system based on depth cameras and large visual models described above have been referenced. Figures 1 to 4 The toilet recognition method based on depth cameras and large visual models has been described in detail, and therefore, its repeated description will be omitted.
Claims
1. A toilet recognition method based on depth camera and large visual model, characterized in that, include: Step 1: Based on the intrinsic and extrinsic coordinate system, perform spatiotemporal alignment and fusion on the acquired raw sensor data stream to obtain toilet color depth aligned data. The raw sensor data stream includes toilet color images and toilet depth point clouds. Step 2: Based on the large visual model, perform multi-layer convolutional semantic feature extraction and anchor box target detection on the color channels in the toilet color depth alignment data to obtain the two-dimensional anchor box matrix of toilet and the confidence score of toilet category. Step 3: Using the two-dimensional anchor frame matrix of the toilet as a pixel space mask, perform region of interest point cloud cropping, statistical filtering noise reduction and clustering segmentation on the depth channel in the toilet color depth alignment data to obtain the toilet isolated three-dimensional point cloud. Step 4: Using the confidence score of toilet category as the registration weighting factor, principal component analysis and iterative nearest point pose registration are performed on the isolated 3D point cloud of toilet to obtain the six-DOF pose parameters of the toilet. Step 5: By calibrating the homogeneous transformation matrix with hand and eye and setting a preset safe distance threshold, the six-degree-of-freedom pose parameters of the toilet are transformed into the world coordinate system and subjected to anti-collision differential verification to obtain the collision-free operation coordinates of the toilet.
2. The toilet recognition method based on depth camera and large visual model according to claim 1, characterized in that, The depth point cloud of the toilet was acquired by a depth camera, while the color image of the toilet was acquired by a visual recognition camera.
3. The toilet recognition method based on depth camera and large visual model according to claim 2, characterized in that, Step one includes: Based on the hardware absolute timestamps carried by the toilet color image and the toilet depth point cloud, the original sensor data stream is evaluated for nearest neighbor time deviation and super-differential abnormal frames are removed to obtain synchronized color image and synchronized depth point cloud. Based on the camera's internal parameter matrix and external rigid body transformation parameters, the three-dimensional spatial coordinates of each effective point in the synchronous depth point cloud are mapped from the three-dimensional physical rigid body to the two-dimensional pixel plane to obtain the depth pixel projection matrix. Based on the addressing topology index recorded by the depth pixel projection matrix, channel concatenation and mean filtering are performed on the three-channel color data of the synchronized color image and the effective depth values of the synchronized depth point cloud to fill gaps and obtain toilet color depth aligned data.
4. The toilet recognition method based on depth camera and large visual model according to claim 1, characterized in that, Step two includes: Color patches are extracted from toilet color depth aligned data and input into a large visual model to perform forward propagation multi-scale semantic feature mapping on the color patches, resulting in a high-dimensional multi-scale semantic feature map. Region response activation and candidate anchor box regression are performed on the high-dimensional multi-scale semantic feature map to obtain the initial generated anchor box matrix and candidate region feature tensor. The overlapping and redundant frames in the initially generated anchor frame matrix are clustered and filtered to retain them, thus obtaining the two-dimensional anchor frame matrix of the toilet. Multi-class conditional probability convergence is performed on the feature tensor of the candidate region to obtain the confidence score of the toilet category.
5. The toilet recognition method based on depth camera and large visual model according to claim 1, characterized in that, Step three includes: The corner pixel boundaries of the two-dimensional anchor frame matrix of the toilet are constructed as a binary logical mask. The depth channel layer stripped from the toilet color depth alignment data is subjected to pixel-level logical and projection truncation and hard deletion of background points to obtain a coarsely clipped three-dimensional point cloud. High-frequency outlier noise statistical filtering denoising analysis is performed on the coarsely cropped 3D point cloud to obtain a smooth 3D point cloud; Manifold clustering and non-destructive segmentation of the adhesive background are performed on the smooth 3D point cloud to obtain the 3D point cloud for toilet isolation.
6. The toilet recognition method based on depth camera and large visual model according to claim 1, characterized in that, Step four includes: Geometric centroid and initial axis features were extracted from the 3D point cloud of toilet enclosure based on principal component analysis. Based on the category label pointed to by the confidence score of the toilet category, the standard template point cloud is retrieved from the pre-set database. The geometric centroid and principal axis features are used as the initial alignment values to perform confidence-weighted iterative nearest point registration optimization on the toilet isolated 3D point cloud and the template point cloud to obtain the optimal mapping transformation matrix. Homogeneous matrix decomposition and pose parameterization are performed on the optimal mapping transformation matrix to obtain the six-DOF pose parameters of the toilet.
7. The toilet recognition method based on depth camera and large visual model according to claim 1, characterized in that, Step five includes: Based on the homogeneous transformation matrix of hand-eye calibration, the camera system homogeneous matrix constructed from the six-degree-of-freedom pose parameters of the toilet is used to perform robot base coordinate system pose mapping to obtain world coordinate system pose data. Based on a preset safe distance threshold along the normal vector direction of the target surface, the position vector of the world coordinate system pose data is subjected to reverse incremental compensation and robotic arm reachability envelope verification to obtain the collision-free space operation point.
8. The toilet recognition method based on depth camera and large visual model according to claim 6, characterized in that, Based on the category label pointed to by the confidence score of the toilet category, a standard template point cloud is retrieved from a pre-set database. Using geometric centroids and principal axis features as initial alignment values, a confidence-weighted iterative nearest-point registration optimization is performed between the isolated 3D point cloud of the toilet and the template point cloud to obtain the optimal mapping transformation matrix, including: For each point in the 3D point cloud of toilet enclosure, the eigenvalues of the local neighborhood covariance matrix are deconstructed to obtain the local curvature of each point; Based on the local curvature of each point, the confidence score of the toilet category, and the global registration residual, a three-factor coupling function is constructed to perform nonlinear mapping to obtain the adaptive dynamic curvature weight value of each point. Based on the adaptive dynamic curvature weight value of each point, the template point cloud and the toilet isolation 3D point cloud are dynamically registered and optimized point by point to obtain the optimal mapping transformation matrix.
9. A toilet recognition system based on a depth camera and a large visual model, characterized in that, include: The sensor data pre-alignment module is used to perform spatiotemporal alignment and fusion of the acquired raw sensor data stream based on the intrinsic and extrinsic parameter coordinate system to obtain toilet color depth alignment data. The raw sensor data stream includes toilet color images and toilet depth point clouds. The semantic extraction and target detection module is used to perform multi-layer convolutional semantic feature extraction and anchor box target detection on the color channels in the toilet color depth alignment data based on a large visual model to obtain a two-dimensional anchor box matrix of toilets and a toilet category confidence score. The 3D point cloud generation module is used to perform region of interest point cloud cropping, statistical filtering noise reduction, and clustering segmentation on the depth channel in the toilet color depth alignment data, using the toilet's 2D anchor frame matrix as a pixel space mask, to obtain the toilet isolated 3D point cloud. The pose parameter generation module is used to perform principal component analysis and iterative nearest point pose registration on the isolated 3D point cloud of the toilet, using the confidence score of the toilet category as the registration weighting factor, to obtain the six-degree-of-freedom pose parameters of the toilet. The operation coordinate generation module is used to perform world coordinate system transformation and anti-collision differential verification on the six-degree-of-freedom pose parameters of the toilet by calibrating the homogeneous transformation matrix by hand and eye and setting a preset safety distance threshold, so as to obtain the collision-free operation coordinates of the toilet.