Intelligent cleaning equipment and method
By adopting deep learning and adaptive feature stitching technology in cleaning equipment, fine-grained and global feature extraction of dirt state images is solved, and the difficulty of identifying and grasping of traditional cleaning equipment in complex environments is improved, and the cleaning efficiency and safety are improved.
Patent Information
- Application Number
- CN202510320779.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional cleaning equipment is difficult to accurately identify and capture dirt in complex or dangerous environments, resulting in inefficient efficiency and insufficient safety, and existing automation equipment is still insufficient to the degree of intelligence.
The deep learning-based neural network model is used to extract the fine-grained and global features of the dirt state image, and the adaptive feature stitching technology of decision anchors is used to perform multi-scale semantic correlation analysis of dirty objects to achieve accurate positioning and capture.
Accurate identification and positioning of different types of dirt in complex environments, improving the accuracy, efficiency and safety of cleaning work, and reducing manpower demand and operational risks.
Smart Images

Figure CN120219697A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent positioning, and more specifically, to an intelligent cleaning and dredging device and method. Background Art
[0002] In the field of traditional cleaning and dredging operations, manual operation still dominates, and there are significant bottlenecks in operation efficiency and safety. Especially when facing cleaning and dredging tasks in complex or dangerous environments, such as deep wells, narrow spaces, or the cleaning of toxic and harmful substances, the limitations of traditional methods are more obvious. In addition, due to the lack of effective real-time monitoring and precise positioning technologies, traditional cleaning and dredging equipment often has difficulty accurately identifying and grasping dirt, resulting in a significant reduction in the effect of cleaning and dredging work and may cause secondary pollution.
[0003] Although some existing automated cleaning and dredging equipment can reduce the human burden to a certain extent, most of the equipment still lacks in intelligence. For example, some equipment can only complete simple moving and suction functions, cannot adjust operation strategies according to actual situations, and do not have the ability to accurately position and effectively process different types of dirt. At the same time, the image recognition systems in the existing technologies usually can only provide basic visual feedback, lack deep learning capabilities, and cannot self-optimize and adjust for complex operation environments, which greatly limits the flexibility and efficiency of cleaning and dredging operations.
[0004] Therefore, an optimized intelligent cleaning and dredging device is expected. Summary of the Invention
[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide an intelligent cleaning and dredging device and method. By using a camera to collect dirt status images, and adopting a neural network model based on deep learning to extract fine-grained features and global image features of the dirt status images, and further using an adaptive feature stitching technology based on decision anchors to adaptively stitch the local features and global semantic features of the dirty object to achieve multi-scale semantic association analysis of the dirty object in a complex scene, and then obtaining the position information of the dirty object based on the semantic analysis result. In this way, the device can accurately identify and locate different types of dirt in a complex environment, not only improving the accuracy, efficiency, and safety of cleaning and dredging work, but also reducing human requirements and operation risks.
[0006] According to one aspect of this application, an intelligent cleaning and dredging device is provided, which includes: An intelligent control device, a camera, and an efficient cleaning and dredging device. The camera is communicably connected to the intelligent control device, and the efficient cleaning and dredging device is communicably connected to the intelligent control device; The camera is used to collect dirt status images and send the dirt status images to the intelligent control device; The intelligent control device is used to identify the soiled object from the soiled state image to obtain the position information of the soiled object, and generate a cleaning action instruction based on the position information of the soiled object; The high-efficiency cleaning device is used to accurately locate and grab the soiled object after receiving the cleaning action instruction; Among them, the intelligent control device is further used to obtain the spatial position information of the soiled object based on the semantic-level search and matching between the soiled state image features and the soiled state global image features of the soiled state image.
[0007] According to another aspect of the present application, an intelligent cleaning method is provided, which includes: Collect the soiled state image and send the soiled state image to the intelligent control device; In the intelligent control device, identify the soiled object from the soiled state image to obtain the position information of the soiled object, and generate a cleaning action instruction based on the position information of the soiled object; After receiving the cleaning action instruction, the high-efficiency cleaning device performs accurate positioning and grabbing of the soiled object.
[0008] Compared with the prior art, the intelligent cleaning device and method provided by the present application collect the soiled state image through a camera, extract the fine-grained features and global image features of the soiled state image by using a neural network model based on deep learning, and further use the adaptive feature stitching technology based on decision-making anchors to adaptively stitch the local features and global semantic features of the soiled object to realize the multi-scale semantic correlation analysis of the soiled object in a complex scene, and then obtain the position information of the soiled object based on the semantic analysis result. In this way, the device can accurately identify and locate different types of soiled objects in a complex environment, which not only improves the accuracy, efficiency and safety of the cleaning work, but also reduces the manpower requirement and operation risk. Description of the Drawings
[0009] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0010] Figure 1 It is a block diagram of the intelligent cleaning device according to the embodiment of the present application; Figure 2 It is a data flow diagram of the intelligent cleaning device according to the embodiment of the present application; Figure 3 It is a block diagram of the intelligent control device in the intelligent cleaning device according to the embodiment of the present application; Figure 4 It is a block diagram of a semantic-level search and matching module for images of dirty objects in an intelligent cleaning device according to an embodiment of the present application; Figure 5 It is a flowchart of an intelligent cleaning method according to an embodiment of the present application. Detailed implementation manners
[0011] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0012] As shown in the present application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0013] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0014] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or below are not necessarily executed precisely in sequence. On the contrary, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0015] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0016] In the technical solution of the present application, an intelligent cleaning device is proposed. By integrating advanced image recognition technology, intelligent algorithms, and efficient mechanical operations, it not only overcomes the problems of low efficiency and insufficient safety of traditional cleaning methods, but also greatly improves the automation degree and intelligent level of the cleaning operation.
[0017] Specifically, in the technical solution of the present application, an intelligent cleaning device is proposed. Figure 1 It is a block diagram of an intelligent cleaning device according to an embodiment of the present application.Figure 2 Schematic diagram of data flow of the intelligent dredging device according to an embodiment of the present application. As Figure 1 and Figure 2 shown, the intelligent dredging device according to an embodiment of the present application includes: a camera 310 for collecting dirt state images and sending the dirt state images to the intelligent control device; an intelligent control device 320 for identifying dirt objects from the dirt state images to obtain the position information of the dirt objects and generating a dredging action instruction based on the position information of the dirt objects; and an efficient dredging device 330 for accurately positioning and grasping the dirt objects after receiving the dredging action instruction.
[0018] Specifically, the camera 310 is used to collect dirt state images and send the dirt state images to the intelligent control device. Among them, the real-time monitoring and data capture of the surrounding environment by the camera are used to initially identify the position and state of the dirt, and accordingly formulate corresponding dredging strategies. It is worth mentioning that in order to ensure that the acquired image quality is high enough for subsequent analysis, the camera may need to have certain technical characteristics, such as high-definition resolution, good low-light performance, etc. In addition, in some special application scenarios, such as dredging work inside pipelines, the camera may also need to have the ability to be waterproof, dustproof, and even corrosion-resistant to meet the requirements of different working environments.
[0019] Specifically, the intelligent control device 320 is used to identify dirt objects from the dirt state images to obtain the position information of the dirt objects and generate a dredging action instruction based on the position information of the dirt objects. It should be understood that in the design and application of the intelligent dredging device, the accuracy of the position information of the dirt objects is the basis for achieving accurate positioning and efficient grasping. Accurately obtaining the position information of the dirt objects can greatly improve the operation efficiency and effect of the device, and at the same time reduce the potential negative impact on the environment. Therefore, in the technical solution of the present application, advanced image recognition technology and intelligent control algorithms are used to accurately extract the position information of the dirt objects from the dirt state images to improve the accuracy and reliability of the dredging work. Therefore, in the embodiment of the present application, the intelligent control device is further used to obtain the spatial position information of the dirt objects based on the semantic-level search and matching between the dirt state image features and the global dirt state image features of the dirt state images. Specifically, in a specific example of the present application, as Figure 3As shown, the intelligent control device 320 includes: a dirt state image feature extraction module 321, configured to extract dirt state image features of a dirt state image to obtain a dirty object image feature coding vector; a dirt state local image feature extraction module 322, configured to extract dirt state local image features of the dirt state image to obtain a set of dirt state local image feature coding vectors; a dirty object image semantic-level search and matching module 323, configured to input the dirty object image feature coding vector and the set of dirt state local image feature coding vectors into a dirty object image semantic-level search coding network to obtain a dirty object image semantic-level search response coding vector; and a position space mapping module 324, configured to obtain position information of the dirty object based on the dirty object image semantic-level search response coding vector.
[0020] Specifically, the dirt state image feature extraction module 321 is used to extract the dirt state image features of the dirt state image to obtain the dirty object image feature coding vector. In the technical solution of the present application, first, the dirt state image is passed through the dirty object target detection network based on the YOLO model to obtain the dirty object ROI image; it should be understood that in deep wells, pipes containing suspended particles or industrial environments with rust textures, dirt often adheres to the pipe wall in irregular forms or forms visual overlap with other impurities, and traditional image processing methods are difficult to distinguish between valid targets and background noise. The target detection network can quickly locate potential dirty areas while retaining the target detail features through a multi-level feature fusion and regional candidate frame generation mechanism, and define accurate semantic analysis boundaries for subsequent processing procedures. Then, the dirty object ROI image is passed through a dirty object image feature extractor based on the ViT model to obtain a dirty object image feature coding vector. Considering the coupling effect of the morphological diversity of dirt and visual interference in complex environments, for example, in scenes such as deep wells and oil pipelines, dirt often presents irregular edges, translucent materials, or visually adheres to the pipe wall. Although traditional convolutional neural networks (CNNs) are effective in extracting local texture features, they are limited by the fixedness of the receptive field and are difficult to capture the non-rigid deformation characteristics of the dirt surface caused by uneven distribution of attachments, light reflection, or fluid dynamics. For example, in the scene of the inner wall of a pipeline containing oily mixed solid residues, the edge of the dirt may present a blurred diffuse form. In contrast, the ViT model can better capture global contextual information. It models the global association between image blocks through a self-attention mechanism, can break through the constraints of the local window, and reconstruct the overall semantic structure of the dirt area from fragmented visual information. Therefore, in the technical solution of the present application, the dirty object ROI image is passed through a dirty object image feature extractor based on the ViT model to convert the dirty object ROI image into a high-dimensional vector with strong semantic representation power, and obtain a dirty object image feature encoding vector. Specifically, the ViT model processes the dirty object ROI image in blocks and uses the self-attention mechanism to evaluate the importance of each dirty object ROI image block, thereby achieving a more detailed and accurate feature representation of the dirty object. This not only improves the recognition accuracy, but also provides parseable deep feature primitives for subsequent semantic-level search.
[0021] Specifically, the dirt state local image feature extraction module 322 is used to extract the dirt state local image features of the dirt state image to obtain a set of dirt state local image feature coding vectors. In the technical solution of the present application, first, the dirt state map is passed through a dirt state image feature extractor based on a dilated convolutional neural network model to obtain a dirt state global image feature coding map; it should be understood that in three-dimensional working spaces such as deep wells and underground pipelines, the distribution of dirt often presents a non-uniform diffusion state. For example, oil stains may adhere to the pipe wall to form a continuous film, while solid particles are randomly accumulated in the concave area. Due to the limitation of the fixed-size convolution kernel, the standard convolutional neural network is difficult to capture the correlation between the microscopic sediment texture and the macroscopic distribution morphology at the same time. By introducing an expandable receptive field, the dilated convolution can achieve the coupling of multi-scale spatial information while maintaining the resolution of the feature map. Therefore, in the technical solution of the present application, the dirt state map is passed through a dirt state image feature extractor based on a hole convolutional neural network model to construct a global semantic field through a hole convolution layer with a layered expansion rate, and a global semantic representation of the dirt state with spatial coherence is constructed to obtain a dirt state global image feature coding map. Among them, the output dirt state global image feature coding map carries the deep semantics of the dirt-environment interaction relationship, which can reflect the microscopic dirt properties and not deviate from the overall environmental constraints. In this way, the device can not only identify specific dirty objects, but also understand the location and possible impact of these objects in the entire environment, providing necessary support for subsequent precise positioning and operation. Further, the image features of the dirt state global image feature coding map are discretized to obtain a set of dirt state local image feature coding vectors. It should be understood that although the global image feature coding map provides the device with a macroscopic view of the entire working area, relying solely on this information is not enough to accurately locate and identify different types of dirt. By discretizing the global feature coding map, a series of coding vectors of local characteristics of the dirt state can be extracted. Each local image feature coding vector of the dirt state can reflect the dirt state characteristics in a specific area. In this way, the system can capture the unique details and specific characteristics of each local area in the environment, thereby enhancing the recognition ability and positioning accuracy of the intelligent cleaning equipment for different types of dirt in complex environments.
[0022] Specifically, the semantic-level search and matching module 323 for soiled object images is configured to input a set of soiled object image feature encoding vectors and local soiled state image feature encoding vectors into a semantic-level search encoding network for soiled object images to obtain a semantic-level search response encoding vector for soiled object images. It should be understood that due to the complex distribution of soiled substances, changing environments, and the presence of risk factors, it is difficult to establish a complete semantic association of the spatial distribution of soiled substances in complex scenarios by simply relying on the local feature encoding of soiled objects or the independent analysis of global image features. Moreover, the split processing of the global background and local details by traditional image processing techniques easily leads to the loss of dynamic relevance at the spatial semantic level of the target soiled substances. For example, when a cleaning device faces a mixed scenario of oil stains and sediment attached to the pipe wall, the high-dimensional features of the soiled object extracted solely by the ViT model may misjudge the boundary of the soiled substances due to the lack of environmental context, while the local set after discretization of the global features extracted only by dilated convolution is difficult to capture the material property differences of tiny soiled substances. To solve the problem of positioning failure caused by the split of soiled substance features in traditional methods, in the technical solution of this application, a set of soiled object image feature encoding vectors and local soiled state image feature encoding vectors are input into the semantic-level search encoding network for soiled object images to establish a dynamic mapping relationship between the ontological attributes and spatial distribution of soiled substances, realize the deep coupling of multi-scale features of soiled objects at the semantic level, and obtain a semantic-level search response encoding vector for soiled object images. Here, the semantic-level search encoding network takes the ontological features of soiled substances extracted by ViT as the anchor points and performs cross-scale semantic alignment with the discretized global environmental features. Its core lies in realizing the dynamic weight fusion of multi-source features through the decision anchor adaptive splicing mechanism, so as to accurately anchor the spatial coordinates of soiled substances under the interference of complex backgrounds. For example, when dealing with the irregular edge formed by the spread of oil stains, the network calculates the feature distribution of the local feature response anchoring matrix and dynamically adjusts the contribution weights of features in different regions in the spatial mapping, enabling the fluid morphological features of the oil stains to form a collaborative expression with the curvature environmental features of the pipe wall, so as to accurately locate the core area of the soiled substances in the complex background. This process essentially constructs a semantic map of the dynamic coupling of soiled substances and the environment, enabling the cleaning device to not only identify the ontological form of soiled substances but also strengthen the feature expression of the key areas of soiled substances. The finally output semantic-level search response encoding vector for soiled object images contains the topological semantic information required for soiled substance positioning, providing a fusion representation that takes into account both microscopic features and macroscopic constraints for subsequent spatial mapping. In this way, the anti-interference ability and positioning accuracy of intelligent cleaning devices in unstructured environments are improved. Specifically, in a specific example of this application, such as Figure 4As shown, the semantic-level search and matching module 323 for the soiled object image includes: a soiled object - dirt state local feature response anchoring unit 3231, which is used to respectively extract the deep implicit features of the soiled object image feature coding vector and the set of dirt state local image feature coding vectors, and construct a local feature response matrix between the two based on a dynamic semantic matching mechanism to obtain a set of soiled object - dirt state local feature response anchoring coding matrices; a decision anchor adaptive splicing weight modulation unit 3232, which is used to calculate the decision anchor adaptive splicing weight factors of each soiled object - dirt state local feature response anchoring coding matrix in the set of soiled object - dirt state local feature response anchoring coding matrices to obtain a set of soiled object - dirt state decision anchor adaptive splicing weight factors; and a fusion unit 3233, which is used to fuse the set of soiled object - dirt state local feature response anchoring coding matrices based on the set of soiled object - dirt state decision anchor adaptive splicing weight factors to obtain a semantic-level search response coding vector for the soiled object image.
[0023] More specifically, the soiled object - dirt state local feature response anchoring unit 3231 is used to respectively extract the deep implicit features of the soiled object image feature coding vector and the set of dirt state local image feature coding vectors, and construct a local feature response matrix between the two based on a dynamic semantic matching mechanism to obtain a set of soiled object - dirt state local feature response anchoring coding matrices. In an embodiment of the present application, first, a deep implicit feature extraction based on fully connected coding is performed on the soiled object image feature coding vector to obtain a soiled object image feature deep implicit coding vector; it should be understood that in the image processing process of traditional cleaning equipment, the original feature coding can often only represent the surface visual attributes of the dirt, and it is difficult to capture the implicit semantic relationship generated by its interaction with the complex environment. Therefore, in the technical solution of the present application, a deep implicit feature extraction based on fully connected coding is performed on the soiled object image feature coding vector to obtain a soiled object image feature deep implicit coding vector. Among them, the fully connected network converts the pixel-level information in the soiled object image feature coding vector into an implicit vector reflecting the physical characteristics of the dirt through the cascaded operation of multiple-layer perceptrons. By performing deep implicit feature extraction on the soiled object image feature coding vector, the system can understand the content of the soiled object image at a higher level and achieve accurate recognition and positioning of different types of dirt. In a specific example of the present application, the following feature extraction formula is used to perform a deep implicit feature extraction based on fully connected coding on the soiled object image feature coding vector to obtain a soiled object image feature deep implicit coding vector; where the feature extraction formula is: ; where is the soiled object image feature coding vector, is matrix multiplication, and They are the weight matrix of the soiled object image features and the bias vector of the soiled object image features, respectively. is the activation function, and is the deep implicit encoding vector of the soiled object image features.
[0024] Next, perform deep implicit feature extraction based on fully connected encoding on each soiled state local image feature encoding vector in the set of soiled state local image feature encoding vectors to obtain a set of soiled state local image feature deep implicit encoding vectors; here, since the soiled objects often show characteristics of non-uniform distribution, variable shapes and high mixing with background interference objects in the operation scenario, the original local image feature encoding vectors can only express shallow visual information at discrete spatial positions and lack the ability to abstractly express the essential attributes of the soiled state. That is, through the multi-layer non-linear transformation of the fully connected neural network, each local feature vector is mapped to a high-dimensional latent semantic space, which is equivalent to constructing a projection field with semantic decoupling ability in the feature space. In this process, through the parameter learning of the neural network weights, the discrete local spatial features of the soiled state are transformed into deep implicit features of the soiled state with global semantic associations, so that the originally isolated spatial feature points are re-encoded into feature nodes carrying high-order semantic information. In this way, not only the representation redundancy caused by environmental noise or light changes in the local features of the soiled state is eliminated, but more importantly, a potential semantic association network between different local features of the soiled state is established, providing a feature basis with semantic interpretability for the subsequent decision-making anchoring mechanism. In a specific example of the present application, the following feature extraction formula is used to perform deep implicit feature extraction based on fully connected encoding on each soiled state local image feature encoding vector in the set of soiled state local image feature encoding vectors to obtain a set of soiled state local image feature deep implicit encoding vectors; where, the feature extraction formula is: ; where, is the set of soiled state local image feature encoding vectors, and are the 1st, 2nd, th, th and and are the weight matrix of the soiled state local image features and the bias vector of the soiled state local image features, respectively, is the th soiled state local image feature deep implicit encoding vector in the set of soiled state local image feature deep implicit encoding vectors.
[0025] Furthermore, each of the depth implicit encoding vectors of the local image features of the dirt states in the set of the depth implicit encoding vector of the dirt object image features and the depth implicit encoding vectors of the local image features of the dirt states is input into the dirt object semantic response decision anchoring component to obtain a set of dirt object-dirt state local feature response anchoring encoding matrices. It should be understood that since the dirt distribution often presents fragmented and unstructured characteristics (such as the mixed form of oil scale and flowing sediment adhering to the inner wall of the pipeline), a single global feature cannot effectively represent the dynamic response relationship of the dirt at the micro local level, and the local features without semantic alignment are prone to falling into the misjudgment trap of isolated spatial positions. Therefore, in the technical solution of this application, each of the depth implicit encoding vectors of the local image features of the dirt states in the set of the depth implicit encoding vector of the dirt object image features and the depth implicit encoding vectors of the local image features of the dirt states is input into the dirt object semantic response decision anchoring component to obtain a set of dirt object-dirt state local feature response anchoring encoding matrices. Here, through the interactive calculation of the semantic response decision anchoring component, a dynamic matching mechanism between the global dirt object semantics and the local environmental features is essentially constructed. In this process, by generating the local feature response anchoring encoding matrix, the overall semantic representation of the dirt object (such as the type and viscosity of the oil stain) is coupled with the dirt states of each local area (such as the sediment thickness and particle distribution density) across dimensions to form a semantic bridge with physical space relevance. In this way, not only the limitation of simple vector superposition in traditional feature stitching is broken through, but also through the matrix-based response relationship expression, the global attributes of the dirt object are mapped to each local feature node, enabling the subsequent decision anchor construction to perform dynamic weight allocation based on the self-consistency between multi-scale features. In a specific example of this application, each of the depth implicit encoding vectors of the local image features of the dirt states in the set of the depth implicit encoding vector of the dirt object image features and the depth implicit encoding vectors of the local image features of the dirt states is input into the dirt object semantic response decision anchoring component according to the following semantic response encoding formula to obtain a set of dirt object-dirt state local feature response anchoring encoding matrices; where the semantic response encoding formula is: ; where is the transposed vector of is the length of denotes vector multiplication, is the th dirt object-dirt state local feature response anchoring encoding matrix in the set of dirt object-dirt state local feature response anchoring encoding matrices.
[0026] More specifically, the decision anchor adaptive splicing weight modulation unit 3232 is configured to calculate the decision anchor adaptive splicing weight factors of each dirty object-dirt state local feature response anchor encoding matrix in the set of dirty object-dirt state local feature response anchor encoding matrices to obtain a set of dirty object-dirt state decision anchor adaptive splicing weight factors. In an embodiment of the present application, first, based on the feature distributions of each dirty object-dirt state local feature response anchor encoding matrix in the set of dirty object-dirt state local feature response anchor encoding matrices, the decision anchor adaptive splicing factors of each dirty object-dirt state local feature response anchor encoding matrix are determined to obtain a set of dirty object-dirt state decision anchor adaptive splicing factors; it should be understood that since there are significant differences in the correlation strength between the dirt states in different local regions and the global semantics of the dirty object, if a feature splicing method with fixed weights is adopted, it will be difficult for the device to dynamically adjust the contribution weights of features in different regions when dealing with complex situations such as blurred boundaries of corrosive dirt and layered coverage of multiphase mixed dirt. In the technical solution of the present application, based on the feature distributions of each dirty object-dirt state local feature response anchor encoding matrix in the set of dirty object-dirt state local feature response anchor encoding matrices, the decision anchor adaptive splicing factors of each dirty object-dirt state local feature response anchor encoding matrix are determined to obtain a set of dirty object-dirt state decision anchor adaptive splicing factors. By analyzing the feature distributions of each dirty object-dirt state local feature response anchor encoding matrix (such as the spatial density of high-activation regions in the matrix, the feature response intensity gradient, etc.), the system can perceive the contribution intensity differences of dirt features at different scales and can uncover the potential physical correlations between the local features of the dirt state and the global semantics of the dirty object. This distribution-driven splicing factor generation mechanism can dynamically adjust the influence proportion of each local feature response anchor encoding matrix in the final semantic-level search response according to the actual physical state of the dirt in the three-dimensional space. In this way, the feature fusion process has the ability to adapt to the environment, so as to achieve accurate semantic-level feature association in scenarios such as dirty object localization, and significantly improve the device's ability to analyze complex dirt forms and localization robustness.
[0027] In particular, it should be understood that analyzing the characteristic distribution of the dirt object - dirt state local feature response anchored coding matrix is essentially a physical property transformation from the "weakly interpretable mean integrity" of the dirt state to the "strongly interpretable maximum locality". Therefore, in the preferred example of this application, based on the eigenvalues of each dirt object - dirt state local feature response anchored coding matrix in the set of dirt object - dirt state local feature response anchored coding matrices, a state transition relationship from the mean to the maximum is constructed through a drift coefficient, and the intermediate transition weight global control is performed with the eigenvalue as the important score to obtain a set of dirt object - dirt state decision anchor adaptive splicing factors. That is, through the interpretable generalization mechanism of the drift coefficient, the global dominance of key local features is enhanced. Among them, the weakly to strongly interpretable generalization process of the drift coefficient constructs a transition bridge from background noise features to target significant features. For example, when dealing with multiphase mixed dirt in a chemical storage tank, some local feature response anchored coding matrices may show a high variance distribution (corresponding to the high-frequency abnormal activation of toxic substance leakage points). At this time, through the regulation of the drift coefficient, the strongly interpretable local features of these matrices (such as the high response extreme value of the leakage point boundary) are promoted to global dominant factors, while the weakly interpretable features in the mean smoothing area (such as background uniform sediment) are down-weighted. This adaptive mechanism based on feature distribution enables the device to establish a physical association between the strongly interpretable local features (high hardness, strong adhesion) in the hardened oil sludge area and the global blockage risk semantics when facing sludge stratification (bottom hardened oil sludge and surface flowing sediment) in the sewage pipe network, so as to accurately distinguish the cleaning priorities of dirt with different textures in the decoding and positioning stage. This dynamic weight distribution mechanism realizes the improvement of the feature representation robustness of the cleaning device under complex working conditions, enabling the device to accurately identify the spatial position relationship between attached dirt and flowing dirt in an environment with unstructured features, providing a reliable pose guidance basis for subsequent precise grasping actions.
[0028] In this example, based on the characteristic distribution of each dirt object - dirt state local feature response anchored coding matrix in the set of dirt object - dirt state local feature response anchored coding matrices, the decision anchor adaptive splicing factor of each dirt object - dirt state local feature response anchored coding matrix is determined by the following decision anchor adaptive splicing factor calculation formula to obtain a set of dirt object - dirt state decision anchor adaptive splicing factors; where, the decision anchor adaptive splicing factor calculation formula is: ; where, represents the variance of represents the mean of is for calculating the number of eigenvalues of represents The number of eigenvalues is used as the difference amplification factor, denotes the maximum value of , denotes the drift coefficient, is the intermediate state transition, is the th eigenvalue of is the th dirt object - dirt state decision anchor adaptive splicing factor in the set of dirt object - dirt state decision anchor adaptive splicing factors.
[0029] Furthermore, the set of dirt object - dirt state decision anchor adaptive splicing factors is weighted based on the Softmax function to obtain the set of dirt object - dirt state decision anchor adaptive splicing weight factors. Since the features of different dirt have distinct spatial distribution characteristics, if equal - weight fusion is adopted, the saliency of key targets will be blurred. Through the exponential transformation property of the Softmax function, the numerical differences of the dirt object - dirt state decision anchor adaptive splicing factors can be mapped to the probability space, enabling the model to autonomously adjust the contribution degrees of different local features according to the current dirt state. It is worth mentioning that through the exponential amplification mechanism of the Softmax function, the weight proportion of high - response local features can be strengthened, while low - response interference features can be suppressed, thereby constructing a physically interpretable weight - assignment logic. In this way, it is ensured that the cleaning device can accurately capture the spatial correlation characteristics of dirt with different densities and forms. In a specific example of this application, the following weight calculation formula is used to weight the set of dirt object - dirt state decision anchor adaptive splicing factors based on the Softmax function to obtain the set of dirt object - dirt state decision anchor adaptive splicing weight factors; where, the weight calculation formula is: ; where is the normalization function, is the th dirt object - dirt state decision anchor adaptive splicing weight factor in the set of dirt object - dirt state decision anchor adaptive splicing weight factors.
[0030] More specifically, the fusion unit 3233 is configured to fuse a set of local feature response anchored coding matrices of the dirty object-dirt state based on a set of decision-making anchors for adaptively splicing weight factors of the dirty object-dirt state to obtain a semantic-level search response coding vector of the dirty object image. Here, by introducing the decision-making anchors for adaptively splicing weight factors of the dirty object-dirt state, a dynamic association bridge between the local features of the dirt state and the global semantics of the dirty object is constructed. This dynamic weight fusion not only solves the problem of unbalanced feature contribution degrees in an unstructured environment in traditional methods, but also realizes the directional enhancement of semantic layer information through the dot product operation of the weight and the feature matrix. In this way, it can help the device more accurately identify the specific location of the pollutant, and further enable the intelligent cleaning device to efficiently complete the cleaning task. In a specific example of the present application, a set of decision-making anchors for adaptively splicing weight factors of the dirty object-dirt state are used to fuse a set of local feature response anchored coding matrices of the dirty object-dirt state according to the following fusion formula to obtain a semantic-level search response coding vector of the dirty object image; wherein, the fusion formula is: ; wherein, is the semantic-level search response coding matrix of the dirty object image, is the shape reshaping operation, is the semantic-level search response coding vector of the dirty object image.
[0031] Specifically, the position space mapping module 324 is configured to obtain the position information of the dirty object based on the semantic-level search response coding vector of the dirty object image. In the technical solution of the present application, the semantic-level search response coding vector of the dirty object image is input into a position space mapper based on a decoder to obtain the position information of the dirty object. It should be understood that although rich feature information about the dirty object and its surrounding environment is extracted and fused in the semantic-level search response coding vector of the dirty object image, this information still needs to be further converted into specific spatial position information in order to efficiently guide actual operations. Here, the position space mapper based on the decoder converts the highly abstract semantic-level search response coding vector of the dirty object image into specific physical coordinates in this process, so as to achieve the purpose of accurately positioning the dirty object.
[0032] In particular, the efficient cleaning device 330 is used to accurately locate and grasp the dirty object after receiving the cleaning action instruction. That is, in this process, according to the received instruction, the efficient cleaning device will start its own positioning system, wherein the efficient cleaning device includes a variety of sensor technologies, such as laser rangefinder, ultrasonic sensor or visual navigation system, so as to accurately locate the target position in a complex working environment. Through these advanced sensing technologies, even in the face of narrow space or different depth working scenes, the efficient cleaning device can effectively overcome obstacles and find the best working path to approach the target. After reaching the predetermined position, the efficient cleaning device performs the grasping operation. It is worth mentioning that considering that different types of dirty objects may have different physical properties, such as hardness, viscosity, etc., the grasping tool is modularized so as to be replaced or adjusted according to the needs of specific tasks. For example, for some harder dirt, a mechanical arm may be used in combination with a claw gripper to grasp; and for more viscous substances, a grasping head with an adsorption function may be used. In addition, in order to ensure the safety and effectiveness of the operation, real-time monitoring of force feedback is required during the grasping process to avoid unnecessary damage to the surrounding environment.
[0033] As described above, the intelligent cleaning device 300 according to the embodiment of the present application can be implemented in various wireless terminals, such as a server with an intelligent cleaning algorithm. In a possible implementation, the intelligent cleaning device 300 according to the embodiment of the present application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the intelligent cleaning device 300 can be a software module in the operating system of the wireless terminal, or can be an application developed for the wireless terminal; of course, the intelligent cleaning device 300 can also be one of the many hardware modules of the wireless terminal.
[0034] Alternatively, in another example, the intelligent cleaning device 300 and the wireless terminal may also be separate devices, and the intelligent cleaning device 300 may be connected to the wireless terminal via a wired and / or wireless network, and transmit interactive information in accordance with an agreed data format.
[0035] Furthermore, an intelligent cleaning method is also provided.
[0036] Figure 5 FIG. 1 is a flow chart of an intelligent cleaning method according to an embodiment of the present application. Figure 5As shown, the intelligent cleaning method according to an embodiment of the present application includes the steps of: S1, collecting an image of the dirt state and sending the image of the dirt state to the intelligent control device; S2, in the intelligent control device, identifying the dirt object from the image of the dirt state to obtain the position information of the dirt object, and generating a cleaning action instruction based on the position information of the dirt object; S3, after receiving the cleaning action instruction, the high-efficiency cleaning device performs precise positioning and grasping of the dirt object.
[0037] In summary, the intelligent cleaning method according to an embodiment of the present application is elucidated. The image of the dirt state is collected by a camera, and the fine-grained features and global image features of the image of the dirt state are extracted by using a neural network model based on deep learning. Further, the adaptive feature stitching technology based on decision anchors is used to adaptively stitch the local features and global semantic features of the dirt object to achieve multi-scale semantic association analysis of the dirt object in a complex scene, and then the position information of the dirt object is obtained based on the semantic analysis result. In this way, the device can achieve precise identification and positioning of different types of dirt in a complex environment, which not only improves the accuracy, efficiency and safety of the cleaning work, but also reduces the manpower requirement and operation risk.
[0038] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications or improvements to the technology in the market, or to enable other ordinary skilled persons in the technical field to understand the embodiments disclosed herein.
Claims
1. An intelligent cleaning device, characterized in that: include: An intelligent control device, a camera and an efficient cleaning device, wherein the camera is communicatively connected to the intelligent control device, and the efficient cleaning device is communicatively connected to the intelligent control device; The camera is used to collect dirt status images and send the dirt status images to the intelligent control device; The intelligent control device is used to identify the dirty object from the dirt state image to obtain the position information of the dirty object, and generate a cleaning action instruction based on the position information of the dirty object; The efficient cleaning device is used to accurately locate and grab the dirty object after receiving the cleaning action instruction; Wherein, the intelligent control device is further used to obtain the spatial position information of the dirty object based on the dirt state semantic level search and matching between the dirt state image features of the dirt state image and the dirt state global image features.
2. The intelligent cleaning device according to claim 1 is characterized in that: The intelligent control device comprises: A dirt state image feature extraction module, used to extract dirt state image features of the dirt state image to obtain a dirt object image feature coding vector; A dirt state local image feature extraction module, used to extract dirt state local image features of the dirt state image to obtain a set of dirt state local image feature encoding vectors; A dirty object image semantic level search matching module is used to input a set of dirty object image feature coding vectors and dirt state local image feature coding vectors into a dirty object image semantic level search coding network to obtain a dirty object image semantic level search response coding vector; The position space mapping module is used to search the response encoding vector based on the semantic level of the dirty object image to obtain the position information of the dirty object.
3. The intelligent cleaning device according to claim 2 is characterized in that: The dirt state image feature extraction module is used for: The dirt state image is passed through a dirt object detection network based on the YOLO model to obtain a dirt object ROI image; The dirty object ROI image is passed through a dirty object image feature extractor based on the ViT model to obtain a dirty object image feature encoding vector.
4. The intelligent cleaning device according to claim 2 is characterized in that: The dirt state local image feature extraction module is used for: The dirt state graph is passed through a dirt state image feature extractor based on a dilated convolutional neural network model to obtain a dirt state global image feature encoding graph; The image features of the global image feature coding map of the dirt state are discretized to obtain a set of local image feature coding vectors of the dirt state.
5. The intelligent cleaning device according to claim 2, characterized in that: The dirty object image semantic level search and matching module comprises: A dirty object-dirty state local feature response anchoring unit, used to extract the deep implicit features of the dirty object image feature coding vector and the dirty state local image feature coding vector respectively, and construct a local feature response matrix between the two based on a dynamic semantic matching mechanism to obtain a set of dirty object-dirty state local feature response anchor coding matrices; A decision anchor adaptive splicing weight modulation unit, used to calculate the decision anchor adaptive splicing weight factor of each dirty object-dirt state local feature response anchor coding matrix in the set of dirty object-dirt state local feature response anchor coding matrices to obtain a set of dirty object-dirt state decision anchor adaptive splicing weight factors; A fusion unit is used to fuse a set of dirty object-dirty state local feature response anchor coding matrices based on a set of dirty object-dirty state decision anchor adaptive splicing weight factors to obtain a dirty object image semantic level search response coding vector.
6. The intelligent cleaning device according to claim 5, characterized in that: The dirty object-dirt state local feature response anchoring unit is used for: Performing deep implicit feature extraction based on fully connected coding on the dirty object image feature coding vector to obtain the dirty object image feature deep implicit coding vector; Performing deep implicit feature extraction based on fully connected coding on each dirt state local image feature coding vector in the set of dirt state local image feature coding vectors to obtain a set of dirt state local image feature deep implicit coding vectors; Each dirty object image feature depth implicit coding vector and each dirty state local image feature depth implicit coding vector in the set of dirty object image feature depth implicit coding vector and dirty state local image feature depth implicit coding vector are respectively input into the dirty object semantic response decision anchor component to obtain a set of dirty object-dirty state local feature response anchor coding matrices.
7. The intelligent cleaning device according to claim 5, characterized in that: The decision anchor adaptive splicing weight modulation unit comprises: A decision anchor adaptive splicing factor calculation subunit is used to determine the decision anchor adaptive splicing factor of each dirty object-dirt state local feature response anchor coding matrix based on the feature distribution of each dirty object-dirt state local feature response anchor coding matrix in the set of dirty object-dirt state local feature response anchor coding matrices to obtain a set of dirty object-dirt state decision anchor adaptive splicing factors; The weighted processing subunit is used to perform weighted processing on the set of dirty object-dirt state decision anchor adaptive splicing factors based on the Softmax function to obtain a set of dirty object-dirt state decision anchor adaptive splicing weight factors.
8. The intelligent cleaning device according to claim 7, characterized in that: The decision anchor adaptive splicing factor calculation subunit is used for: Based on the eigenvalues of each dirty object-dirt state local feature response anchor coding matrix in the set of dirty object-dirt state local feature response anchor coding matrices, a state transition relationship from the mean to the maximum value is constructed through the drift coefficient, and the intermediate transition weights are globally controlled with the eigenvalue as the important score to obtain a set of dirty object-dirt state decision anchor adaptive splicing factors.
9. The intelligent cleaning device according to claim 2, characterized in that: The position space mapping module is used for: The semantic-level search response encoding vector of the dirty object image is input into the decoder-based position space mapper to obtain the position information of the dirty object.
10. An intelligent cleaning method, characterized in that: include: Collecting a dirt state image and sending the dirt state image to an intelligent control device; In the intelligent control device, a dirty object is identified from the dirt state image to obtain position information of the dirty object, and a cleaning action instruction is generated based on the position information of the dirty object; After receiving the cleaning action instruction, the efficient cleaning device accurately locates and grabs the dirty objects.