Cage recognition system and method for mobile robot

The cage data is obtained through cameras and laser scanners, and combined with deep learning technology, it is determined that the cages can be stacked, which solves the problem of manual confirmation of inefficiency in traditional warehousing operations and achieves efficient and accurate cage stacking.

CN119152503BActive Publication Date: 2025-08-29ZHEJIANG KECONG CONTROL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411245633.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-08-29
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

In traditional warehousing operations, drivers need to manually confirm whether the cage can be stacked, resulting in inefficiency and high operating error rates.

Method used

The camera is used to obtain the panoramic image of the cage environment and the laser scanner to obtain the spatial coordinate position data. Combined with deep learning technology, feature vectors are extracted and the classifier is used to determine whether the cage can be stacked.

Benefits of technology

Reduce manual intervention, improve cage stacking efficiency, and reduce operational error rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152503B_ABST
    Figure CN119152503B_ABST
Patent Text Reader

Abstract

The present application relates to the field of cage recognition, and specifically discloses a cage recognition system and method for a mobile robot. The system first obtains a panoramic image of the cage environment captured by a camera and the spatial coordinate position data of the cage captured by a laser scanner, and then uses deep learning technology to perform feature extraction and correlation analysis on the two. Finally, a classifier is used to determine whether the mobile robot can stack the cages, thereby reducing manual intervention, improving efficiency, and reducing the operational error rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cage identification, and more specifically, to a cage identification system and method for a mobile robot. Background Art

[0002] With the development of warehousing and logistics technology, material cages are often used to transport and store goods during warehousing operations, and three-dimensional warehousing can be achieved by stacking material cages, reducing the occupation of storage space.

[0003] In traditional warehouses, every time a forklift driver stacks cages, they need to look outside the forklift to check the cage's status to determine whether it is ready for stacking. However, each time a cage is stacked, the driver must manually confirm the stacking, reducing the efficiency of the cage stacking process.

[0004] Therefore, a cage identification system and method for a mobile robot are desired. Summary of the Invention

[0005] To address the above technical issues, the present application proposes a mobile robot cage recognition system and method. The system first acquires a panoramic image of the cage environment captured by a camera and the cage spatial coordinate position data captured by a laser scanner. Deep learning technology is then used to perform feature extraction and correlation analysis on the two. Finally, a classifier is used to determine whether the mobile robot can stack the cages, thereby reducing manual intervention, improving efficiency, and lowering the error rate of operation.

[0006] According to one aspect of the present application, a cage identification system for a mobile robot is provided, comprising:

[0007] The robot cage recognition data acquisition module is used to obtain the cage environment panoramic image collected by the camera and the cage spatial coordinate position data collected by the laser scanner;

[0008] A robot cage recognition data extraction module is used to extract a cage environment panoramic multimodal association feature vector and a cage space coordinate position semantic feature vector from the cage environment panoramic image collected by the camera and the cage space coordinate position data collected by the laser scanner;

[0009] The mobile robot cage stacking judgment module is used to judge whether the mobile robot can stack the cages based on the panoramic multimodal association feature vector of the cage environment and the semantic feature vector of the cage spatial coordinate position.

[0010] According to another aspect of the present application, a cage identification method for a mobile robot is provided, comprising:

[0011] Acquire the panoramic image of the cage environment captured by the camera and the spatial coordinate position data of the cage captured by the laser scanner;

[0012] Extracting a cage environment panoramic multimodal association feature vector and a cage space coordinate position semantic feature vector from the cage environment panoramic image captured by the camera and the cage space coordinate position data captured by the laser scanner;

[0013] Based on the panoramic multimodal association feature vector of the cage environment and the semantic feature vector of the cage spatial coordinate position, it is determined whether the mobile robot can stack the cages.

[0014] Compared with the existing technology, the present application provides a material cage recognition system and method for a mobile robot, which first obtains a panoramic image of the material cage environment captured by a camera and the material cage spatial coordinate position data captured by a laser scanner, and then uses deep learning technology to perform feature extraction and correlation analysis on the two. Finally, a classifier is used to determine whether the mobile robot can stack the material cages, thereby reducing manual intervention, improving efficiency, and reducing operational error rates. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0016] Figure 1 4 is a block diagram of a cage identification system for a mobile robot according to an embodiment of the present application.

[0017] Figure 2 4 is a block diagram of a robot cage recognition data extraction module in a mobile robot cage recognition system according to an embodiment of the present application.

[0018] Figure 3 This is a block diagram of a cage environment panoramic feature extraction unit in a cage recognition system of a mobile robot according to an embodiment of the present application.

[0019] Figure 4 4 is a block diagram of a mobile robot cage stacking judgment module in a mobile robot cage recognition system according to an embodiment of the present application.

[0020] Figure 5 This is a flowchart of a cage identification method for a mobile robot according to an embodiment of the present application.

[0021] Figure 6is a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0023] Figure 1 : is a block diagram of the cage identification system of the mobile robot in the embodiment of the present application. Figure 1 As shown, the material cage recognition system 100 of the mobile robot according to the embodiment of the present application includes: a robot material cage recognition data acquisition module 110, which is used to obtain a panoramic image of the material cage environment collected by a camera and the material cage spatial coordinate position data collected by a laser scanner; a robot material cage recognition data extraction module 120, which is used to extract the material cage environment panoramic multimodal association feature vector and the material cage spatial coordinate position semantic feature vector from the material cage environment panoramic image collected by the camera and the material cage spatial coordinate position data collected by the laser scanner; a mobile robot material cage stacking judgment module 130, which is used to judge whether the mobile robot can stack the material cages based on the material cage environment panoramic multimodal association feature vector and the material cage spatial coordinate position semantic feature vector.

[0024] In the above-mentioned mobile robot cage recognition system 100, the robot cage recognition data acquisition module 110 is used to obtain a panoramic image of the cage environment captured by a camera and the cage spatial coordinate position data captured by a laser scanner. It should be understood that with the advancement of warehousing and logistics technology, cages are widely used for the transportation and storage of goods. By stacking cages, three-dimensional warehousing can be achieved, thereby effectively saving storage space. In traditional warehousing operations, when stacking cages, forklift drivers usually need to extend their heads out of the forklift to observe the status of the cages to determine whether stacking is possible. This practice is not only time-consuming, but also relies on the driver's manual operation to confirm the feasibility of stacking, thereby reducing the overall efficiency of cage stacking. Therefore, in the technical solution of the present application, by obtaining a panoramic image of the cage environment captured by a camera and the cage spatial coordinate position data captured by a laser scanner, and combining it with deep learning technology, it is determined whether the mobile robot can stack the cages, thereby reducing manual intervention, improving efficiency and reducing operational errors.

[0025] Specifically, acquiring panoramic images of the cage environment captured by a camera and spatial coordinate position data of the cage captured by a laser scanner are key to achieving efficient cage identification and stacking. The panoramic images of the cage environment provided by the camera provide the system with visual information of the cage's surroundings, helping to identify the cage's position, shape, size, and other visual features. This image data enables the system to capture the cage's appearance and the spatial layout of its surroundings—information crucial for understanding the cage's physical presence and environmental conditions. The spatial coordinate position data provided by the laser scanner provides precise spatial coordinate information, depicting the cage's position and geometry in three-dimensional space. Unlike the two-dimensional images provided by a camera, a laser scanner can accurately measure the cage's height, width, and depth, generating a detailed three-dimensional model. This precise spatial data is crucial for assessing the cage's actual dimensions and its spatial relationship to other cages or objects, helping to determine whether it can be stacked safely and efficiently. By combining these two types of data, the system can comprehensively consider the cage's visual characteristics and spatial position, leading to more accurate stacking decisions. Acquiring both types of data is fundamental to implementing an automated cage stacking system, improving identification accuracy and operational safety.

[0026] In the above-mentioned mobile robot cage recognition system 100, the robot cage recognition data extraction module 120 is used to extract the cage environment panoramic multimodal association feature vector and the cage spatial coordinate position semantic feature vector from the cage environment panoramic image collected by the camera and the cage spatial coordinate position data collected by the laser scanner. It should be understood that integrating this information together can form a more comprehensive and accurate cage model, helping the system to consider more factors when making stacking decisions. Doing so not only improves the accuracy of recognition, but also ensures the safety and effectiveness of operations in complex environments. Therefore, extracting and integrating these feature vectors is a key step in realizing an automated cage processing system.

[0027] Figure 2 FIG. 1 is a block diagram of a robot cage recognition data extraction module in a cage recognition system for a mobile robot according to an embodiment of the present application. Figure 2 As shown, in a specific embodiment of the present application, the robot cage recognition data extraction module 120 includes: a cage environment panoramic feature extraction unit 121, used to perform feature extraction on the cage environment panoramic image collected by the camera to obtain the cage environment panoramic multimodal association feature vector; a cage space coordinate position feature extraction unit 122, used to perform feature extraction on the cage space coordinate position data collected by the laser scanner to obtain the cage space coordinate position semantic feature vector.

[0028] It should be understood that the panoramic images captured by the camera provide detailed visual information about the cage and its surroundings. These images include the cage's appearance characteristics, color, texture, and spatial layout. These visual features are very important for identifying the specific type of cage, its location, and its relative relationship to other objects. For example, by analyzing the color and texture in the image, different types of cages can be distinguished, while by identifying the edges and shape of the cage, its spatial layout can be determined. The purpose of feature extraction is to convert this complex visual data into a more concise form that can be used for subsequent processing, namely feature vectors. These feature vectors contain important key information in the image, such as the cage's geometry, relative position, and possible occlusion. Multimodal association feature vectors further integrate this information, taking into account the spatial distribution and contextual relationships in the image, to provide a more comprehensive description of the environment.

[0029] Furthermore, feature extraction is performed on the cage's spatial coordinate position data captured by the laser scanner to accurately describe the cage's position and layout in three-dimensional space. This is crucial for cage identification, processing, and manipulation in automated systems. Laser scanners provide highly accurate three-dimensional spatial data that precisely captures the cage's actual spatial coordinates, shape, and dimensions. This detailed spatial information helps the system better understand the cage's geometry and its specific location within the environment. By performing feature extraction on this spatial data, the system transforms complex three-dimensional coordinate information into a more concise and useful form: semantic feature vectors. These vectors extract the core features related to the cage's spatial location. The feature extraction process involves processing the point cloud data obtained by laser scanning to obtain feature vectors representing the cage's spatial coordinate position. These feature vectors contain information about the cage's geometry, dimensions, and relative position to other objects in the environment. For example, by extracting feature vectors, the system can understand the cage's exact dimensions, position, and possible spatial gaps, thereby determining whether the cage is suitable for the current operation requirements.

[0030] Figure 3 FIG. 1 is a block diagram of a cage environment panoramic feature extraction unit in a cage recognition system for a mobile robot according to an embodiment of the present application. Figure 3 As shown, in a specific embodiment of the present application, the cage environment panoramic feature extraction unit 121 includes: a cage environment panoramic attention feature encoding subunit 1211, which is used to perform attention feature encoding on the cage environment panoramic image captured by the camera to obtain a cage environment panoramic spatial feature map and a cage environment panoramic channel feature map; a cage environment panoramic multimodal feature aggregation subunit 1212, which is used to perform feature aggregation on the cage environment panoramic spatial feature map and the cage environment panoramic channel feature map to obtain the cage environment panoramic multimodal association feature vector.

[0031] It should be understood that the panoramic images captured by the camera contain a large amount of visual information, including not only the specific appearance of the cage, but also the background of the surrounding environment, other objects, and changes in lighting. This information contains features that are crucial for cage recognition and processing. However, image data often contains a large amount of redundant or irrelevant information, so how to extract effective features from it is particularly important. Through attention feature encoding, the system can intelligently focus on key areas in the image, such as the position and shape of the cage, thereby filtering out irrelevant information and improving the accuracy of feature extraction. Attention feature encoding weights important features, allowing the system to quickly extract useful information from massive image data. This method not only improves recognition accuracy, but also significantly improves computational efficiency and reduces unnecessary data processing overhead.

[0032] Furthermore, feature aggregation is performed on the panoramic spatial feature map and the panoramic channel feature map of the cage environment to effectively integrate different types of feature information, thereby gaining a comprehensive understanding of the cage and its environment. The panoramic spatial feature map of the cage environment primarily describes the spatial layout and positional relationships in the image, such as the cage's position, boundaries, and its relative relationship to the surrounding environment. The channel feature map, on the other hand, focuses on the different color and texture features in the image, which are important for identifying the cage's appearance, material, and possible markings. By aggregating these two feature maps, spatial information is combined with visual information to form a comprehensive, multi-dimensional feature representation, which helps to fully understand the actual situation of the cage. Although a single feature map can provide specific information, it often cannot fully capture the complex characteristics of the cage in the environment. The spatial feature map can display the cage's geometric position, while the channel feature map reveals its visual characteristics. Aggregating these two features can provide more accurate cage identification and localization capabilities. For example, by combining spatial position and visual appearance features, the system can more accurately determine the cage's type, state, and relationship with other objects, thereby improving overall recognition accuracy.

[0033] In a specific embodiment of the present application, the cage environment panoramic attention feature encoding subunit 1211 includes: passing the cage environment panoramic image captured by the camera through the cage environment panoramic spatial attention mechanism to obtain the cage environment panoramic spatial feature map; passing the cage environment panoramic image captured by the camera through the cage environment panoramic channel attention mechanism to obtain the cage environment panoramic channel feature map.

[0034] It's understandable that while a panoramic image of a cage environment contains rich visual information, not all areas are critical for identifying the cage or its surroundings. The spatial attention mechanism automatically identifies the most informative regions in the image and focuses attention on these key areas. For example, the attention mechanism allows the system to highlight the cage's edges, position, and relationship to other objects in the environment, while ignoring irrelevant details in the background. This ability to focus on key areas improves the accuracy and relevance of feature extraction. The panoramic spatial feature map of the cage environment primarily involves the spatial layout and geometric features of objects in the image. By weighting the spatial information in the image, the spatial attention mechanism effectively captures and represents the cage's position, size, and relative positional relationships within the environment. Traditional image processing methods often require processing a large amount of pixel information, which can be computationally expensive and inefficient. The spatial attention mechanism uses weighted image processing to emphasize features in important regions, effectively reducing unnecessary data processing. This approach enables the system to extract useful information more quickly, thereby improving processing efficiency. Specifically, the panoramic image of the cage environment captured by the camera is subjected to deep convolution encoding using the convolution encoding part of the cage environment panoramic spatial attention mechanism to obtain an initial convolution feature map; the initial convolution feature map is input into the spatial attention part of the cage environment panoramic spatial attention mechanism to obtain a spatial attention map; the spatial attention map is activated by a Softmax function to obtain a spatial attention feature map; and the spatial attention feature map and the initial convolution feature map are multiplied by the position points to obtain the cage environment panoramic spatial feature map. More specifically, the initial convolution feature map is input into the spatial attention part of the cage environment panoramic spatial attention mechanism to obtain the spatial attention map, including: performing average pooling and maximum pooling along the channel dimension on the initial convolution feature map to obtain an average feature matrix and a maximum feature matrix; cascading and channel-adjusting the average feature matrix and the maximum feature matrix to obtain a channel feature matrix; and using the convolution layer of the spatial attention feature map to perform convolution encoding on the channel feature matrix to obtain the spatial attention map.

[0035] Furthermore, the panoramic image of the cage environment captured by the camera is processed through a channel-based attention mechanism to effectively extract channel-level features, such as color and texture, from the image. This is crucial for accurately identifying and understanding the cage and its environment. In panoramic images captured by the camera, different channels (such as red, green, and blue) carry different types of information, which is crucial for cage recognition and classification. For example, the color channel provides visual features of the cage, while the texture channel reveals surface details. The channel-based attention mechanism weights the features of these different channels, highlighting those that are helpful for recognition, thereby improving feature extraction. The channel-based attention mechanism dynamically adjusts its focus based on the contribution of each channel. This mechanism automatically identifies which channel features are most important for the task at hand and enhances them accordingly. This dynamic weighting improves feature extraction accuracy, enabling the generated channel-based feature maps to more accurately reflect visual characteristics such as the cage's color and texture. Given that panoramic images contain multiple information sources, such as color, brightness, and texture, the channel-based attention mechanism effectively integrates this information. By focusing on the features of different channels, the system can more comprehensively understand the visual features of the cage. Specifically, each layer of the cage environment panoramic channel attention mechanism performs the following operations on the input data in the forward pass of the layer: convolution processing on the input data based on a two-dimensional convolution kernel to generate a convolution feature map; pooling processing on the convolution feature map to generate a pooling feature map; activation processing on the pooling feature map to generate an activation feature map; calculating the quotient of the eigenvalue mean of the feature matrix corresponding to each channel in the activation feature map and the sum of the eigenvalue mean of the feature matrix corresponding to all channels as the weighting coefficient of the feature matrix corresponding to each channel; and weighting the feature matrix of each channel with the weighting coefficient of each channel in the activation feature map to generate a channel attention feature map; wherein the output of the last layer of the cage environment panoramic channel attention mechanism is the cage environment panoramic channel feature map.

[0036] In a specific embodiment of the present application, the cage environment panoramic multimodal feature aggregation subunit 1212 includes: passing the cage environment panoramic spatial feature map and the cage environment panoramic channel feature map through a cage environment panoramic multimodal association convolutional neural network as a feature encoder to obtain a cage environment panoramic multimodal association feature map; and pooling the cage environment panoramic multimodal association feature map to obtain the cage environment panoramic multimodal association feature vector.

[0037] It should be understood that although a single spatial feature map or channel feature map can provide a certain perspective, it is often unable to fully capture all the key information in the environment. By fusing these two feature maps, the multimodal association convolutional neural network can comprehensively consider the characteristics of space and channels, thereby generating a richer and more expressive multimodal association feature map. Among them, the panoramic spatial feature map of the cage environment mainly focuses on the spatial layout and geometric information in the image, while the panoramic channel feature map of the cage environment focuses on the color and texture features of the image. By combining these two feature maps, the information in the image can be fully understood from multiple angles. The spatial feature map provides the position and shape information of the object, while the channel feature map provides the visual details and color information of the object. This combination can more accurately describe the actual environment of the cage, thereby improving recognition accuracy. Specifically, each layer of the cage environment panoramic multimodal association convolutional neural network used as a feature encoder performs convolution processing, mean pooling processing based on the local feature matrix and nonlinear activation processing on the input data in the forward pass of the layer, so that the last layer of the cage environment panoramic multimodal association convolutional neural network used as a feature encoder outputs the cage environment panoramic multimodal association feature map, wherein the input of the cage environment panoramic multimodal association convolutional neural network used as a feature encoder is the cage environment panoramic spatial feature map and the cage environment panoramic channel feature map.

[0038] Furthermore, the panoramic multimodal correlation feature map of the cage environment is pooled in order to achieve dimensionality reduction, extraction and summarization of features, thereby optimizing the performance and efficiency of the model. The panoramic multimodal correlation feature map of the cage environment usually contains a large amount of high-dimensional data, which may lead to large computational workload and high storage requirements. Through pooling, these high-dimensional feature maps can be converted into low-dimensional feature vectors, reducing the computational burden and storage requirements. Pooling effectively compresses data by selecting important areas in the feature map and retaining key feature information. The pooling operation can extract important statistical information from the feature map, such as the maximum value or average value, which can represent the main features of the feature map. In this way, pooling helps summarize and integrate the key information in the image, so that the final feature vector can concisely and comprehensively express the panoramic characteristics of the cage environment.

[0039] In a specific embodiment of the present application, the cage spatial coordinate position feature extraction unit 122 includes: segmenting the cage spatial coordinate position data collected by the laser scanner to obtain a cage spatial coordinate position word sequence; passing the cage spatial coordinate position word sequence through a cage spatial coordinate position semantic encoder to obtain the cage spatial coordinate position semantic feature vector.

[0040] It should be understood that the spatial coordinate data collected by the laser scanner is usually high-dimensional continuous data, which may contain a large number of original coordinate points. These data are difficult to process directly in their original form. Through word segmentation operations, that is, dividing these continuous data points into several meaningful units (i.e., "words"), complex data structures can be converted into structured word sequences. This structured data form is easier to analyze, process and understand. After the spatial coordinate data is segmented, the complexity of the data can be greatly reduced. The word sequence after word segmentation usually has a fixed length and clear boundaries, which makes subsequent calculations and processing more efficient. Compared with processing the original continuous coordinate data, the computing resources and time required to operate and analyze word sequences are greatly reduced.

[0041] Furthermore, the cage spatial coordinate position word sequence is processed through the cage spatial coordinate position semantic encoder in order to convert the discrete word sequence of spatial coordinates into a meaningful vector representation with semantic information. The spatial coordinate position word sequence is only a discrete and basic spatial data representation, which is difficult to directly express complex spatial relationships and features. Through the semantic encoder, the word sequence can be converted into a feature vector with rich semantic information, making the expression of spatial data more comprehensive and in-depth. This vectorized representation can capture the implicit patterns and relationships in spatial data and provide useful information for further analysis. As the output of the semantic encoder, the feature vector can provide a more expressive and structured input for the machine learning model. Compared with the original word sequence, the semantic feature vector has a higher dimension and complexity. Specifically, the embedding layer of the cage spatial coordinate position semantic encoder is used to map each cage spatial coordinate position word in the cage spatial coordinate position word sequence into a cage spatial coordinate position word embedding vector to obtain a sequence of cage spatial coordinate position word embedding vectors; the converter-based Bert model of the cage spatial coordinate position semantic encoder is used to perform global context semantic encoding on the sequence of cage spatial coordinate position word embedding vectors to obtain multiple cage spatial coordinate position feature vectors; and, the multiple cage spatial coordinate position feature vectors are cascaded to obtain the cage spatial coordinate position semantic feature vector.

[0042] In the above-mentioned mobile robot cage recognition system 100, the mobile robot cage stacking judgment module 130 is used to judge whether the mobile robot can stack the cages based on the panoramic multimodal association feature vector of the cage environment and the semantic feature vector of the cage spatial coordinate position. It should be understood that the panoramic multimodal association feature vector of the cage environment provides comprehensive information about the cage environment, including spatial layout, object distribution and other environmental features. These features help the system understand the positional relationship and spatial configuration of the cage in the environment. The semantic feature vector of the cage spatial coordinate position provides detailed position and semantic information of the cage itself, such as size, shape and relative position. Combining these two features can fully understand the environmental conditions and cage status, thereby making more accurate judgments. By analyzing the panoramic feature vector and the semantic feature vector, the system can evaluate whether there is enough space for stacking, whether the cage meets the stacking requirements, and whether obstacles will be encountered during the stacking process. This comprehensive analysis can effectively prevent potential collisions and stacking instability.

[0043] Figure 4 FIG. 1 is a block diagram of a mobile robot cage stacking judgment module in a mobile robot cage recognition system according to an embodiment of the present application. Figure 4 As shown, in a specific embodiment of the present application, the mobile robot cage stacking judgment module 130 includes: a mobile robot cage stacking feature fusion unit 131, which is used to fuse the cage environment panoramic multimodal association feature vector and the cage spatial coordinate position semantic feature vector to obtain a cage stacking judgment classification feature vector; a mobile robot cage stacking feature optimization unit 132, which is used to perform fine-grained feature-class label coupling optimization based on weight adaptation on the cage stacking judgment classification feature vector to obtain an optimized cage stacking judgment classification feature vector; a cage stacking classification judgment unit 133, which is used to pass the optimized cage stacking judgment classification feature vector through a classifier to obtain a classification result, and the classification result is used to judge whether the mobile robot can stack the cage.

[0044] It should be understood that the fusion of the panorama multimodal correlation feature vector of the cage environment and the semantic feature vector of the cage's spatial coordinate position to obtain the cage stacking judgment classification feature vector is intended to comprehensively consider comprehensive information about the environment and the cage, improving the accuracy and efficiency of stacking decisions. This fusion not only optimizes the decision-making process but also enhances the system's robustness and intelligent decision-making capabilities, enabling the robot to more effectively perform stacking tasks in complex environments.

[0045] In particular, in the technical solution of the present application, the weight of the classifier determines how to distribute data points of different categories in the feature space. In the cage stacking judgment problem, the classifier needs to be able to distinguish different stacking states, such as whether the cage can be safely stacked. The decision on cage stacking depends on multiple factors, including environmental characteristics and semantic information of spatial coordinates. Due to the complexity and diversity of the cage environment, the classifier cannot correctly understand all the conditions or environmental restrictions for cage stacking during the training process, and therefore will be overly sensitive or insensitive to specific feature vector patterns, resulting in certain categories of features being mistakenly emphasized or ignored, resulting in category coherence interference. Therefore, in the technical solution of the present application, the cage stacking judgment classification feature vector is optimized based on fine-grained feature-class label coupling optimization based on weight adaptation to obtain an optimized cage stacking judgment classification feature vector.

[0046] Among them, the cage stacking judgment classification feature vector is subjected to fine-grained feature-class label coupling optimization based on weight adaptation to obtain an optimized cage stacking judgment classification feature vector, including: calculating the product between the cage stacking judgment classification feature vector and its transposed vector to obtain a cage stacking judgment feature autocorrelation expression matrix; performing correlation measurement on the category decoding weight matrix of the classifier and the cage stacking judgment feature autocorrelation expression matrix to obtain a cage stacking judgment feature-category weight fine-grained coupling representation matrix; performing probabilistic processing based on the Sigmoid activation function on the cage stacking judgment feature-category weight fine-grained coupling representation matrix to obtain a cage stacking judgment feature-category weight fine-grained coupling probabilistic representation matrix; performing backpropagation expression compensation on the cage stacking judgment classification feature vector based on the cage stacking judgment feature-category weight fine-grained coupling probabilistic representation matrix to obtain a category-related compensation representation vector; and calculating the position-weighted sum between the category-related compensation representation vector and the cage stacking judgment classification feature vector to obtain the optimized cage stacking judgment classification feature vector.

[0047] The following correlation measurement formula is used to measure the correlation between the classifier's category decoding weight matrix and the cage stacking judgment feature autocorrelation expression matrix to obtain a cage stacking judgment feature-category weight fine-grained coupling representation matrix;

[0048] The correlation measurement formula is:

[0049]

[0050] Among them, X 1i Represents the i-th row vector of the class decoding weight matrix of the classifier, X 2jrepresents the jth row vector of the cage stacking judgment feature autocorrelation expression matrix, S represents the covariance matrix of the data set to which the i-th row vector of the category decoding weight matrix of the classifier and the j-th row vector of the cage stacking judgment feature autocorrelation expression matrix belong, S -1 represents the inverse matrix of the covariance matrix, M i,j Represents the eigenvalue of the (i, j)th position of the cage stacking judgment feature-category weight fine-grained coupling probabilistic representation matrix.

[0051] The following compensation formula is used to perform backpropagation expression compensation on the cage stacking judgment classification feature vector based on the cage stacking judgment feature-category weight fine-grained coupling probabilistic representation matrix to obtain a category-related compensation representation vector;

[0052] Wherein, the compensation formula is:

[0053]

[0054] Among them, V r represents the category-related compensation representation vector, M represents the cage stacking judgment feature-category weight fine-grained coupling probabilistic representation matrix, V c represents the cage stacking judgment classification feature vector, Represents matrix multiplication.

[0055] In the technical solution of the present application, when the cage stacking judgment classification feature vector is classified by the classifier, since the weight of the classifier also needs to be adapted to the cage stacking judgment classification feature vector, category coherent interference with the cage stacking judgment classification feature vector may occur. Based on this, in the technical solution of the present application, the cage stacking judgment classification feature vector is optimized based on fine-grained feature-class label coupling based on weight adaptation. First, the autocorrelation matrix of the cage stacking judgment classification feature vector is used to represent the autocorrelation interference feature spectrum, and then the correlation measure between the autocorrelation matrix of the cage stacking judgment classification feature vector and the category decoding weight matrix of the classifier is calculated to simulate the fine-grained feature-class label coherent interference phase between the cage stacking judgment feature and the weight matrix of the classifier. Then, the cage stacking judgment feature-class weight fine-grained coupling representation matrix is ​​used to perform backpropagation expression compensation on the cage stacking judgment classification feature vector, and the equivalent probability intensity representation of the cage stacking judgment classification feature vector in the absence of interference is restored by reverse compensation, and the optimized cage stacking judgment classification feature vector is obtained to improve the accuracy of the classification result.

[0056] Furthermore, the optimized cage stacking judgment classification feature vector incorporates comprehensive data on environmental information and cage characteristics. These features are analyzed by a classifier, which can map complex input data to specific categories, such as "stackable" or "non-stackable." In practical applications, cage stacking tasks may involve complex environmental factors, such as spatial constraints, obstacles, and changes in cage status. The classifier is able to process this complex data and make a comprehensive judgment based on the classification results of the optimized cage stacking judgment classification feature vector. This adaptability enables the robot to maintain efficiency and flexibility in dynamic environments.

[0057] In summary, the embodiment of the present application first obtains a panoramic image of the cage environment captured by a camera and the spatial coordinate position data of the cage captured by a laser scanner, and then uses deep learning technology to perform feature extraction and correlation analysis on the two. Finally, a classifier is used to determine whether the mobile robot can stack the cages, thereby reducing manual intervention, improving efficiency, and reducing operational error rates.

[0058] As described above, the mobile robot cage identification system 100 according to the embodiments of the present application can be implemented in various terminal devices. In one example, the mobile robot cage identification system 100 can be integrated into the terminal device as a software module and / or a hardware module. For example, the mobile robot cage identification system 100 can be a software module in the operating system of the terminal device, or can be an application developed for the terminal device; of course, the mobile robot cage identification system 100 can also be one of the many hardware modules of the terminal device.

[0059] Alternatively, in another example, the mobile robot's cage identification system 100 and the terminal device may also be separate devices, and the mobile robot's cage identification system 100 may be connected to the terminal device via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.

[0060] Figure 5 Flowchart of the cage identification method of the mobile robot according to the embodiment of the present application. Figure 5 As shown, the material cage recognition method of a mobile robot according to an embodiment of the present application includes: S110, acquiring a panoramic image of the material cage environment captured by a camera and material cage spatial coordinate position data captured by a laser scanner; S120, extracting a material cage environment panoramic multimodal association feature vector and a material cage spatial coordinate position semantic feature vector from the material cage environment panoramic image captured by the camera and the material cage spatial coordinate position data captured by the laser scanner; S130, judging whether the mobile robot can stack the material cages based on the material cage environment panoramic multimodal association feature vector and the material cage spatial coordinate position semantic feature vector.

[0061] Here, those skilled in the art will appreciate that the specific operations of each step in the above-mentioned mobile robot cage identification method have been described in detail in the above reference. Figures 1 to 4 The description of the cage recognition system of the mobile robot has been introduced in detail, and therefore, its repeated description will be omitted.

[0062] Below, reference Figure 6 To describe the electronic device according to the embodiment of the present application.

[0063] like Figure 6 As shown, the electronic device 10 includes an input device 11, an input interface 12, a central processing unit 13, a memory 14, an output interface 15, an output device 16, and a bus 17. The input interface 12, the central processing unit 13, the memory 14, and the output interface 15 are interconnected via the bus 17, and the input device 11 and the output device 16 are connected to the bus 17 via the input interface 14 and the output interface 15, respectively, and are further connected to other components of the electronic device 10.

[0064] Specifically, the input device 11 receives input information from the outside and transmits the input information to the central processing unit 13 through the input interface 12; the central processing unit 13 processes the input information based on the computer-executable instructions stored in the memory 14 to generate output information, stores the output information temporarily or permanently in the memory 14, and then transmits the output information to the output device 16 through the output interface 15; the output device 16 outputs the output information to the outside of the electronic device 10 for user use.

[0065] In one embodiment, Figure 6 The electronic device 10 shown can be implemented as a network device, which may include: a memory configured to store programs; a processor configured to run the programs stored in the memory to execute any one of the mobile robot cage identification methods described in the above embodiments.

[0066] According to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program comprising program code for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network and / or installed from a removable storage medium.

[0067] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0068] It is understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present application, and the present application is not limited thereto. Those skilled in the art may make various modifications and improvements without departing from the spirit and substance of the present application, and such modifications and improvements are also considered to be within the scope of protection of the present application.

Claims

1. A cage identification system for a mobile robot, characterized in that: include: The robot cage recognition data acquisition module is used to obtain the cage environment panoramic image collected by the camera and the cage spatial coordinate position data collected by the laser scanner; A robot cage recognition data extraction module is used to extract a cage environment panoramic multimodal association feature vector and a cage space coordinate position semantic feature vector from the cage environment panoramic image collected by the camera and the cage space coordinate position data collected by the laser scanner; A mobile robot cage stacking judgment module is used to judge whether the mobile robot can stack the cages based on the panoramic multimodal association feature vector of the cage environment and the semantic feature vector of the cage spatial coordinate position; Among them, the mobile robot cage stacking judgment module includes: a mobile robot cage stacking feature fusion unit, which is used to fuse the cage environment panoramic multimodal association feature vector and the cage spatial coordinate position semantic feature vector to obtain a cage stacking judgment classification feature vector; a mobile robot cage stacking feature optimization unit, which is used to perform fine-grained feature-class label coupling optimization based on weight adaptation on the cage stacking judgment classification feature vector to obtain an optimized cage stacking judgment classification feature vector; a cage stacking classification judgment unit, which is used to pass the optimized cage stacking judgment classification feature vector through a classifier to obtain a classification result, and the classification result is used to judge whether the mobile robot can stack the cage; Among them, the mobile robot cage stacking feature optimization unit includes: calculating the product between the cage stacking judgment classification feature vector and its transposed vector to obtain a cage stacking judgment feature autocorrelation expression matrix; performing correlation measurement on the category decoding weight matrix of the classifier and the cage stacking judgment feature autocorrelation expression matrix to obtain a cage stacking judgment feature-category weight fine-grained coupling representation matrix; performing probabilistic processing based on the Sigmoid activation function on the cage stacking judgment feature-category weight fine-grained coupling representation matrix to obtain a cage stacking judgment feature-category weight fine-grained coupling probabilistic representation matrix; performing backpropagation expression compensation on the cage stacking judgment classification feature vector based on the cage stacking judgment feature-category weight fine-grained coupling probabilistic representation matrix to obtain a category-related compensation representation vector; calculating the position-weighted sum between the category-related compensation representation vector and the cage stacking judgment classification feature vector to obtain the optimized cage stacking judgment classification feature vector.

2. The mobile robot cage identification system according to claim 1, characterized in that: The robot cage identification data extraction module includes: A cage environment panoramic feature extraction unit is used to extract features from the cage environment panoramic image captured by the camera to obtain a cage environment panoramic multimodal correlation feature vector; The cage space coordinate position feature extraction unit is used to extract features from the cage space coordinate position data collected by the laser scanner to obtain the cage space coordinate position semantic feature vector.

3. The mobile robot cage identification system according to claim 2, characterized in that: The cage environment panoramic feature extraction unit includes: A cage environment panoramic attention feature encoding subunit, configured to perform attention feature encoding on the cage environment panoramic image captured by the camera to obtain a cage environment panoramic space feature map and a cage environment panoramic channel feature map; The cage environment panoramic multimodal feature aggregation subunit is used to perform feature aggregation on the cage environment panoramic spatial feature map and the cage environment panoramic channel feature map to obtain the cage environment panoramic multimodal association feature vector.

4. The cage identification system for a mobile robot according to claim 3, characterized in that: The cage environment panoramic attention feature encoding subunit includes: The cage environment panoramic image captured by the camera is passed through the cage environment panoramic space attention mechanism to obtain the cage environment panoramic space feature map; The cage environment panoramic image captured by the camera is passed through the cage environment panoramic channel attention mechanism to obtain the cage environment panoramic channel feature map.

5. The mobile robot cage identification system according to claim 4, characterized in that: The cage environment panoramic multimodal feature aggregation subunit includes: The cage environment panoramic spatial feature map and the cage environment panoramic channel feature map are passed through a cage environment panoramic multimodal association convolutional neural network as a feature encoder to obtain a cage environment panoramic multimodal association feature map; The cage environment panoramic multimodal association feature map is pooled to obtain the cage environment panoramic multimodal association feature vector.

6. The mobile robot cage identification system according to claim 5, characterized in that: The cage space coordinate position feature extraction unit comprises: Segmenting the cage space coordinate position data collected by the laser scanner to obtain a cage space coordinate position word sequence; The cage space coordinate position word sequence is passed through a cage space coordinate position semantic encoder to obtain the cage space coordinate position semantic feature vector.

7. A method for identifying a cage of a mobile robot, using the cage identification system of the mobile robot according to claim 1, characterized in that: include: Acquire the panoramic image of the cage environment captured by the camera and the spatial coordinate position data of the cage captured by the laser scanner; Extracting a cage environment panoramic multimodal association feature vector and a cage space coordinate position semantic feature vector from the cage environment panoramic image captured by the camera and the cage space coordinate position data captured by the laser scanner; Based on the panoramic multimodal association feature vector of the cage environment and the semantic feature vector of the cage spatial coordinate position, it is determined whether the mobile robot can stack the cages.

8. The method for identifying a cage of a mobile robot according to claim 7, wherein: Extracting a panoramic multimodal association feature vector of the cage environment and a semantic feature vector of the cage space coordinate position from the cage environment panoramic image captured by the camera and the cage space coordinate position data captured by the laser scanner, including: Performing feature extraction on the panoramic image of the cage environment captured by the camera to obtain a panoramic multimodal correlation feature vector of the cage environment; Feature extraction is performed on the cage space coordinate position data collected by the laser scanner to obtain the cage space coordinate position semantic feature vector.

Citation Information

Patent Citations

  • Panorama camera and laser radar fusion method for unmanned boarding bridge obstacle avoidance

    CN117011656A

  • Load state detection method and device of mechanical arm

    CN117754634A

  • Full-automatic screen printing machine for glass bottle body and control method thereof

    CN118493999A