Mobile truck identification system and method for mobile robot
Through cameras, ultrasonic and sound sensors combined with deep learning technology, mobile robots can identify and avoid the visual blind spots and limited recognition range of mobile trucks in industrial environments, and achieve efficient and safe operation of trucks.
Patent Information
- Application Number
- CN202411287005.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-09-13
AI Technical Summary
Existing mobile robots are difficult to effectively identify and avoid mobile trucks in industrial environments, and single camera detection has problems such as blind spots and limited recognition range.
Camera, ultrasonic sensor and sound sensor are used to obtain multi-area images, distance data and sound data, and feature extraction and correlation analysis are combined with deep learning technology to determine whether avoidance operations are needed through the classifier.
Improves mobile robots’ understanding of complex and dynamic industrial environments, ensuring safe and efficient operation and avoiding collisions with trucks.
Smart Images

Figure CN119251788B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of mobile truck identification, and more specifically, to a mobile truck identification system and method for a mobile robot. Background Art
[0002] Mobile robots are intelligent devices capable of autonomously moving and operating in various environments. They are widely used in industries such as industry, logistics, and healthcare. As mobile robots become increasingly popular in the industrial sector, a key challenge is how to enable these robots to effectively identify and avoid moving trucks in dynamic industrial environments.
[0003] In the industrial field, the route, speed, and occasional stops of trucks can affect the normal operation of robots. Currently, most mobile robots rely on single-camera detection to identify moving trucks. Single-camera detection has a single detection angle and is prone to visual blind spots, resulting in limited viewing angles and blind spot issues.
[0004] Therefore, a mobile truck identification system and method for a mobile robot is desired. Summary of the Invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiments of the present application provide a mobile robot mobile truck identification system and method, which first obtains multi-region images of the mobile robot's surrounding environment collected by a camera, ultrasonic distance data between the mobile robot and the surrounding environment collected by an ultrasonic sensor, and environmental sound data at multiple predetermined time points collected by a sound sensor, and then uses deep learning technology to perform feature extraction and correlation analysis on the three. Finally, a classifier is used to determine whether the mobile robot needs to perform an avoidance operation, thereby improving the mobile robot's ability to understand the surrounding environment in a complex and dynamic industrial environment, and ensuring efficient and safe operation of mobile trucks in industrial environments.
[0006] According to one aspect of the present application, a mobile truck identification system for a mobile robot is provided, comprising:
[0007] a mobile truck identification data acquisition module, configured to acquire images of multiple regions of the mobile robot's surroundings captured by the camera, ultrasonic distance data between the mobile robot and the surroundings captured by the ultrasonic sensor, and ambient sound data at multiple predetermined time points captured by the sound sensor;
[0008] a mobile truck identification data extraction module, configured to extract a surrounding truck tracking feature vector and an environmental multimodal association feature vector from the multi-region images of the mobile robot's surrounding environment captured by the camera, the ultrasonic distance data between the mobile robot and the surrounding environment captured by the ultrasonic sensor, and the environmental sound data at multiple predetermined time points captured by the sound sensor;
[0009] The mobile robot avoidance operation judgment module is used to judge whether the mobile robot needs to perform an avoidance operation based on the surrounding environment truck tracking feature vector and the environment multimodal association feature vector.
[0010] According to another aspect of the present application, a mobile truck identification method of a mobile robot is provided, comprising:
[0011] Acquire multi-region images of the mobile robot's surrounding environment captured by a camera, ultrasonic distance data between the mobile robot and the surrounding environment captured by an ultrasonic sensor, and environmental sound data at multiple predetermined time points captured by a sound sensor;
[0012] Extracting a surrounding environment truck tracking feature vector and an environment multimodal association feature vector from the multi-region images of the mobile robot's surrounding environment captured by the camera, the ultrasonic distance data between the mobile robot and the surrounding environment captured by the ultrasonic sensor, and the environmental sound data at multiple predetermined time points captured by the sound sensor;
[0013] Based on the surrounding environment truck tracking feature vector and the environment multimodal association feature vector, it is determined whether the mobile robot needs to perform an avoidance operation.
[0014] Compared with the existing technology, the present application provides a mobile robot mobile truck identification system and method, which first obtains multi-area images of the mobile robot's surrounding environment collected by a camera, ultrasonic distance data between the mobile robot and the surrounding environment collected by an ultrasonic sensor, and environmental sound data at multiple predetermined time points collected by a sound sensor, and then uses deep learning technology to perform feature extraction and correlation analysis on the three. Finally, a classifier is used to determine whether the mobile robot needs to perform an avoidance operation, thereby improving the mobile robot's ability to understand the surrounding environment in a complex and dynamic industrial environment, and ensuring efficient and safe operation of mobile trucks in industrial environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0016] Figure 1 4 is a block diagram of a mobile truck identification system for a mobile robot according to an embodiment of the present application.
[0017] Figure 2 4 is a block diagram of a mobile truck identification data extraction module in a mobile truck identification system of a mobile robot according to an embodiment of the present application.
[0018] Figure 3 4 is a block diagram of a surrounding environment multi-region feature extraction unit in a mobile truck recognition system of a mobile robot according to an embodiment of the present application.
[0019] Figure 4 4 is a block diagram of a mobile robot avoidance operation judgment module in a mobile robot mobile truck recognition system according to an embodiment of the present application.
[0020] Figure 5 Flowchart of a mobile truck identification method for a mobile robot according to an embodiment of the present application.
[0021] Figure 6 is a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0023] Figure 1 This is a block diagram of the mobile truck identification system of the mobile robot according to the embodiment of the present application. Figure 1As shown, the mobile truck recognition system 100 of the mobile robot according to the embodiment of the present application includes: a mobile truck recognition data acquisition module 110, which is used to obtain multi-region images of the mobile robot's surrounding environment collected by a camera, ultrasonic distance data between the mobile robot and the surrounding environment collected by an ultrasonic sensor, and environmental sound data at multiple predetermined time points collected by a sound sensor; a mobile truck recognition data extraction module 120, which is used to extract the surrounding environment truck tracking feature vector and the environmental multimodal association feature vector from the multi-region images of the mobile robot's surrounding environment collected by the camera, the ultrasonic distance data between the mobile robot and the surrounding environment collected by the ultrasonic sensor, and the environmental sound data at multiple predetermined time points collected by the sound sensor; a mobile robot avoidance operation judgment module 130, which is used to judge whether the mobile robot needs to perform an avoidance operation based on the surrounding environment truck tracking feature vector and the environmental multimodal association feature vector.
[0024] In the mobile robot's mobile truck identification system 100, the mobile truck identification data acquisition module 110 is configured to acquire multi-region images of the mobile robot's surroundings captured by a camera, ultrasonic distance data between the mobile robot and its surroundings captured by an ultrasonic sensor, and ambient sound data at multiple predetermined time points captured by an acoustic sensor. It should be understood that mobile robots are intelligent devices capable of autonomously operating and performing tasks in a variety of environments and are widely used in fields such as industry, logistics, and healthcare. With the increasing application of these robots in industrial environments, a major challenge is how to enable robots to effectively identify and avoid moving trucks in complex industrial environments. In industrial environments, a truck's travel path, speed changes, and occasional stops may interfere with the robot's normal operation. Currently, most mobile robots rely on a single camera for truck detection. However, a single camera has a limited field of view and is prone to blind spots, which limits the recognition range and makes it difficult to fully monitor dynamic changes in the environment. Therefore, the technical solution of this application uses deep learning technology to improve the mobile robot's ability to understand its surroundings and determine whether to perform avoidance maneuvers in complex and dynamic industrial environments by acquiring multi-region images of the mobile robot's surroundings captured by a camera, ultrasonic distance data between the mobile robot and its surroundings collected by an ultrasonic sensor, and ambient sound data collected at multiple predetermined time points by an acoustic sensor. This technology is then combined with deep learning technology to improve the mobile robot's ability to understand its surroundings and determine whether to perform avoidance maneuvers in complex and dynamic industrial environments. This ensures that the robot can safely and efficiently interact with moving trucks in industrial environments and avoid potential conflicts.
[0025] Specifically, the data collected by cameras, ultrasonic sensors, and acoustic sensors is used to achieve comprehensive perception of the mobile robot's surroundings, thereby improving the robot's navigation and obstacle avoidance capabilities in complex environments. The camera provides visual information of the robot's surroundings, capturing various objects and obstacles. The use of multi-region images covers different angles and regions around the robot, enabling the system to identify and track trucks and other dynamic objects in the environment. The ultrasonic sensor measures the distance between the object and the sensor by emitting and receiving ultrasonic waves. This distance data helps the system understand the spatial relationship between the robot and its surroundings, including the distance and relative position of obstacles. The ambient sound data recorded by the acoustic sensor provides additional environmental information, such as the presence of sound sources (such as horns and engine noise) that may be associated with moving trucks. The combination of these three types of data creates a multimodal perception system that can more comprehensively identify and track surrounding trucks and other obstacles. This integrated perception capability enables the mobile robot to make more accurate avoidance decisions in complex environments, improving navigation safety and reliability.
[0026] In the mobile robot's mobile truck identification system 100, the mobile truck identification data extraction module 120 is used to extract the surrounding truck tracking feature vector and the environment multimodal association feature vector from the multi-region images of the mobile robot's surrounding environment collected by the camera, the ultrasonic distance data between the mobile robot and the surrounding environment collected by the ultrasonic sensor, and the environmental sound data collected by the sound sensor at multiple predetermined time points. It should be understood that through multimodal feature fusion, a comprehensive description of the environment can be obtained, including the truck's position, motion state, distance to obstacles, and sound sources. The purpose of this is to improve the accuracy and comprehensiveness of environmental perception by comprehensively utilizing data from different sensors.
[0027] Figure 2 FIG. 1 is a block diagram of a mobile truck identification data extraction module in a mobile truck identification system of a mobile robot according to an embodiment of the present application. Figure 2As shown, in a specific embodiment of the present application, the mobile truck identification data extraction module 120 includes: a surrounding environment multi-region feature extraction unit 121, used to perform feature extraction on the multi-region image of the mobile robot's surrounding environment collected by the camera to obtain the surrounding environment truck tracking feature vector; an ultrasonic distance feature extraction unit 122, used to perform feature extraction on the ultrasonic distance data between the mobile robot and the surrounding environment collected by the ultrasonic sensor to obtain an obstacle distance semantic feature vector; an environmental sound feature extraction unit 123, used to perform feature extraction on the environmental sound data at multiple predetermined time points collected by the sound sensor to obtain an environmental sound spatial state feature vector; an environmental multimodal feature fusion unit 124, used to fuse the obstacle distance semantic feature vector and the environmental sound spatial state feature vector to obtain the environmental multimodal association feature vector.
[0028] It's understandable that the environmental images captured by cameras contain a wealth of visual information, including the scene's color, shape, and texture. Feature extraction converts this image data into structured feature vectors, allowing it to extract key truck-related features, such as the truck's outline, color, and logo. In dynamic environments, trucks may be constantly moving. Feature extraction generates truck tracking feature vectors that contain information about the truck's position, speed, and direction. These vectors not only help the system identify the truck in the image but also track its trajectory. This is crucial for achieving continuous monitoring and tracking of targets, especially maintaining continuous tracking of moving trucks.
[0029] Furthermore, feature extraction is performed on the ultrasonic distance data collected by ultrasonic sensors between the mobile robot and its surroundings to improve the accuracy of environmental perception and the system's decision-making capabilities. Ultrasonic sensors provide real-time distance data from obstacles. This data reflects the spatial layout of the robot's surroundings, including the location and distance of obstacles. Feature extraction converts this raw distance data into structured feature vectors, helping the system to more quickly and accurately understand the spatial distribution of the surrounding environment. This real-time perception is crucial for dynamically adjusting the robot's path and avoiding collisions.
[0030] Furthermore, feature extraction of environmental sound data collected by sound sensors at multiple predetermined time points is intended to enhance the perception and decision-making capabilities of mobile robots in complex environments. While cameras and ultrasonic sensors provide visual and distance data, these sensors cannot fully capture the changes in sound in the environment. Feature extraction of sound data can provide complementary information about sound characteristics. For example, the spectral characteristics of sound can help identify specific sound sources or patterns of change, information that may be difficult to detect in visual data. Combining sound features with visual and distance data helps to gain a more comprehensive understanding of environmental conditions.
[0031] Specifically, the goal of fusing the obstacle distance semantic feature vector with the ambient sound spatial state feature vector to generate the multimodal environmental association feature vector is to enhance the mobile robot's comprehensive understanding of its surroundings and its decision-making capabilities. The information provided by different sensors is complementary: visual and ultrasonic sensors primarily detect the spatial position and distance of obstacles, while acoustic sensors provide additional information about ambient sound.
[0032] Figure 3 FIG. 1 is a block diagram of a surrounding environment multi-region feature extraction unit in a mobile truck recognition system of a mobile robot according to an embodiment of the present application. Figure 3 As shown, in a specific embodiment of the present application, the surrounding environment multi-region feature extraction unit 121 includes: a surrounding environment multi-region feature encoding subunit 1211, which is used to feature encode the multi-region image of the surrounding environment of the mobile robot captured by the camera to obtain a surrounding environment truck tracking feature map; a surrounding environment truck feature downsampling subunit 1212, which is used to downsample the surrounding environment truck tracking feature map to obtain the surrounding environment truck tracking feature vector.
[0033] It should be understood that multi-region images provide rich visual information about the surrounding environment. Feature encoding can convert this high-dimensional image data into low-dimensional feature maps. Feature encoding extracts the most representative features in the image, such as the truck's shape, color, and texture, enabling the system to more accurately identify and distinguish the truck from other objects. Feature maps can refine the truck's key visual features, thereby improving recognition accuracy. Raw image data often contains a large amount of redundant information, and directly processing this data may result in computational inefficiency. Feature encoding compresses the image data into smaller feature maps while retaining key visual information. This data compression not only increases processing speed but also reduces computing resource consumption, allowing the system to operate more efficiently in real-time or near-real-time application scenarios.
[0034] Furthermore, downsampling the surrounding truck tracking feature map is performed to improve processing efficiency and system performance while preserving important information. Feature maps typically contain a large number of pixels and complex visual information, which can require significant computing resources to process. Downsampling reduces the data volume by lowering the resolution of the feature map, significantly reducing computational complexity and storage requirements. This enables faster processing and analysis, adapting to the requirements of real-time applications, especially on resource-constrained embedded systems or mobile devices.
[0035] In a specific embodiment of the present application, the surrounding environment multi-region feature encoding subunit 1211 includes: passing the multi-region images of the surrounding environment of the mobile robot captured by the camera through a surrounding environment image denoiser based on deep learning to obtain multiple-region surrounding environment denoised and enhanced images; passing the multiple-region surrounding environment denoised and enhanced images through a surrounding environment truck tracking YOLO target detection network as a feature extractor to obtain multiple-region surrounding environment truck tracking interest regions; passing the multiple-region surrounding environment truck tracking interest regions through a surrounding environment truck tracking feature encoder to obtain the surrounding environment truck tracking feature map.
[0036] It should be understood that images captured by cameras under different environmental conditions (such as low light, strong reflections or high-contrast scenes) may contain various noises. These noises will interfere with the details of the image, making the image blurry or distorted. The deep learning-based denoiser can effectively identify and reduce these noises, thereby improving the clarity and quality of the image. This clear image helps the system more accurately identify and analyze target objects in the surrounding environment, such as trucks, people or other obstacles. The deep learning denoiser learns the difference between noise features and real signals by training the network model, and can extract more real and useful features from the image. Through denoising processing, the system's adaptability to different environmental conditions can be improved, allowing the robot to maintain stable performance when facing various interferences (such as light changes, shadows or reflections). The enhanced image data makes the system more robust in complex environments, and can still achieve accurate target recognition and environmental understanding under non-ideal conditions. Specifically, the multi-region images of the mobile robot's surrounding environment captured by the camera are input into the encoder of the deep learning-based surrounding environment image denoiser, wherein the encoder uses a convolution layer to perform explicit spatial encoding on the multi-region images of the mobile robot's surrounding environment captured by the camera to obtain image features; and the image features are input into the decoder of the deep learning-based surrounding environment image denoiser, wherein the decoder uses a deconvolution layer to perform deconvolution processing on the image features to obtain the denoised and enhanced images of the multiple-region surrounding environments.
[0037] Furthermore, the denoised and enhanced images of the surrounding environment in multiple regions are processed through the YOLO object detection network, primarily to improve the accuracy and efficiency of object detection. The YOLO (You Only Look Once) object detection network can detect and accurately locate multiple objects in an image in real time. Feeding the denoised and enhanced images into the YOLO network leverages its powerful feature extraction and object detection capabilities to accurately identify and locate the truck. This is because the noise in the denoised and enhanced images is significantly reduced after processing, making the features more prominent, helping the YOLO network to more effectively identify and locate the object. The YOLO network was originally designed to achieve efficient real-time object detection. The YOLO network can extract key features from complex images and map them to the object detection task. The denoised and enhanced images provide clearer visual information, which the YOLO network can better extract and utilize, thereby improving the accuracy of truck tracking.
[0038] Furthermore, the surrounding environment of multiple regions of interest (ROI) is processed through a feature encoder, primarily to further extract and enhance key features in the image, improving the accuracy and efficiency of target tracking. The function of the feature encoder is to extract deep, discernible features from the ROI. ROIs often contain important information related to the target, but this information may be obscured by complex background or noise. Using techniques such as convolutional layers and pooling layers, the feature encoder can extract clearer and more useful features from these regions, forming a feature map. In target tracking tasks, the input image typically contains a large amount of pixel information. Directly processing this high-dimensional data can result in high computational complexity and low processing efficiency. By converting the raw image data into a feature map, the feature encoder can map this high-dimensional data into a lower-dimensional space while preserving key feature information. Specifically, each layer of the surrounding environment truck tracking feature encoder is used to perform convolution processing, mean pooling processing based on the local feature matrix, and nonlinear activation processing on the input data in the forward pass of the layer to output the surrounding environment truck tracking feature map by the last layer of the surrounding environment truck tracking feature encoder, wherein the input of the surrounding environment truck tracking feature encoder is the multiple regions of interest of surrounding environment truck tracking.
[0039] In a specific embodiment of the present application, the ultrasonic distance feature extraction unit 122 includes: passing the ultrasonic distance data between the mobile robot and the surrounding environment collected by the ultrasonic sensor through an obstacle distance text convolutional neural network to obtain a plurality of obstacle distance feature vectors; and splicing the plurality of obstacle distance feature vectors into the obstacle distance semantic feature vector.
[0040] It should be understood that ultrasonic sensors provide distance information from obstacles in the environment, and this raw data itself may lack sufficient structured information. Convolutional neural networks (CNNs) can automatically extract recognizable features from this distance data and convert them into feature vectors. These feature vectors contain information such as the location, shape, and size of the obstacles, providing a detailed and accurate data representation for the robot's environmental perception. Ultrasonic sensor data is typically high-dimensional, and direct processing may result in high computational complexity. Through convolution and pooling operations, convolutional neural networks can effectively convert high-dimensional data into low-dimensional feature vectors while retaining key feature information. Specifically, the ultrasonic distance data between the mobile robot and its surrounding environment, collected by the ultrasonic sensor, is segmented to obtain a word sequence; the embedding layer of the obstacle distance text convolutional neural network is used to map each word in the word sequence into a word embedding vector to obtain a sequence of word embedding vectors; and the sequence of word embedding vectors is semantically encoded based on global context using the transformer-based Bert model of the obstacle distance text convolutional neural network to obtain multiple obstacle distance feature vectors.
[0041] Furthermore, multiple obstacle distance feature vectors are concatenated into an obstacle distance semantic feature vector. This is primarily to integrate environmental information from different angles, providing a more comprehensive and richer feature representation, thereby improving environmental understanding and processing capabilities. Each obstacle distance feature vector is typically collected from a specific perspective or sensor position and may only reflect a portion of the environment. By concatenating these multiple feature vectors, distance data from different sources can be fused to form a more comprehensive obstacle distance semantic feature vector. This integration helps form a global understanding of the environment and enhances the system's ability to perceive surrounding obstacles.
[0042] In a specific embodiment of the present application, the ambient sound feature extraction unit 123 includes: extracting an ambient sound spectrum graph from the ambient sound data at multiple predetermined time points collected by the sound sensor; and passing the ambient sound spectrum graph through an ambient sound spectrum space focusing feature encoder to obtain the ambient sound space state feature vector.
[0043] As you can understand, raw sound data is typically time-domain signals, and directly processing these signals can be difficult to extract useful information. Converting sound data into a spectrogram transforms time-domain signals into a frequency-domain representation. A spectrogram displays the intensity distribution of a sound signal at different frequencies, allowing analysts to more clearly understand the frequency characteristics and changing trends of the sound, which is crucial for identifying and classifying it. A spectrogram reveals the frequency composition and intensity variations of a sound, helping to identify different sound sources in an environment. For example, a spectrogram can distinguish different sound patterns in an environment, such as vehicle horns, wind, or human voices. This feature extraction is crucial for applications such as acoustic environment monitoring, sound event detection, and noise control. Using techniques such as Fourier transform, a sound spectrogram can decompose complex sound signals into a series of simple frequency components. This frequency decomposition not only improves sound recognition accuracy but also allows for the extraction of useful information from the spectrum even in the presence of high noise levels or weak sound signals.
[0044] Furthermore, the ambient sound spectrogram is processed through the ambient sound spectrum spatially focused feature encoder. This is primarily to extract representative spatial features from the spectrogram and convert these features into a more informative and structured feature vector. The ambient sound spectrogram provides the distribution of sound in the frequency domain, but in practical applications, relying solely on the spectrogram may not be sufficient to capture the spatial characteristics of the sound source. The ambient sound spectrum spatially focused feature encoder extracts key features related to the spatial location of the sound source by performing feature extraction and spatial focusing on the spectrogram. These features help understand the spatial layout and source distribution of the sound, enabling the system to better analyze and process the ambient sound. The data in the spectrogram is high-dimensional, and direct processing can result in high computational complexity and information redundancy. Through spatial focusing, the feature encoder can highlight the spatial characteristics of the sound, helping the system better understand the source of the sound and its spatial characteristics within the environment. Specifically, the convolution encoding part of the ambient sound spectrum spatial focus feature encoder is used to perform deep convolution encoding on the ambient sound spectrum map to obtain an initial convolution feature map; the initial convolution feature map is input into the spatial attention part of the ambient sound spectrum spatial focus feature encoder to obtain a spatial attention map; the spatial attention map is activated by a Softmax function to obtain a spatial attention feature map; and the spatial attention feature map and the initial convolution feature map are multiplied by the position points to obtain the ambient sound spatial state feature vector. More specifically, the initial convolution feature map is input into the spatial attention part of the ambient sound spectrum spatial focus feature encoder to obtain a spatial attention map, including: performing average pooling and maximum pooling along the channel dimension on the initial convolution feature map to obtain an average feature matrix and a maximum feature matrix; cascading and channel-adjusting the average feature matrix and the maximum feature matrix to obtain a channel feature matrix; and using the convolution layer of the spatial attention feature map to perform convolution encoding on the channel feature matrix to obtain a spatial attention map.
[0045] In the mobile robot mobile truck identification system 100 described above, the mobile robot avoidance operation judgment module 130 is configured to determine whether the mobile robot needs to perform an avoidance operation based on the truck tracking feature vector and the multimodal correlation feature vector of the surrounding environment. It should be understood that the truck tracking feature vector provides information about the truck's specific location, speed, and direction of movement. These feature vectors help the robot understand the truck's current position and trajectory in the environment, thereby predicting the truck's future location. This is crucial for real-time avoidance, as it allows the robot to dynamically adjust its path to avoid collisions with the truck. Based on these features, the robot can calculate the truck's motion trends and determine whether to change its route based on its current and expected positions. The multimodal correlation feature vector integrates data from various sensors, such as vision, lidar, and sound, providing a more comprehensive understanding of the environment. This multimodal information fusion helps improve the robot's ability to recognize environmental complexity, enabling it to more accurately perceive surrounding obstacles and dynamic targets, including trucks. By combining this information, the robot can assess the relationship between the truck's movement and other environmental factors in real time, such as whether there is sufficient space for safe avoidance and whether other obstacles are affecting the avoidance path. The ultimate goal of this comprehensive analysis is to ensure the safety and operational efficiency of the robot.
[0046] Figure 4 FIG. 1 is a block diagram of a mobile robot avoidance operation judgment module in a mobile robot mobile truck recognition system according to an embodiment of the present application. Figure 4 As shown, in a specific embodiment of the present application, the mobile robot avoidance operation judgment module 130 includes: an environmental information feature association unit 131, used to associate the surrounding environment truck tracking feature vector and the environmental multimodal association feature vector to obtain a robot avoidance judgment classification feature vector; an environmental information feature optimization unit 132, used to perform a weight-adapted class coherence interference correction on the robot avoidance judgment classification feature vector to obtain an optimized robot avoidance judgment classification feature vector; an avoidance operation classification judgment unit 133, used to pass the optimized robot avoidance judgment classification feature vector through a classifier to obtain a classification result, and the classification result is used to judge whether the mobile robot needs to perform an avoidance operation.
[0047] It's easy to understand that combining the truck tracking feature vector with the multimodal environmental correlation feature vector generates a classification feature vector for robot avoidance judgment. This feature vector integrates information from both the dynamic and static environments, and after processing, it can clearly identify situations where avoidance is necessary. This approach not only improves decision-making accuracy but also speeds up response, ensuring the robot's safety and operational efficiency in complex environments.
[0048] In particular, in the technical solution of the present application, the robot avoidance judgment classification feature vector is obtained by fusing multiple data sources, namely, multi-region images of the mobile robot's surrounding environment, ultrasonic distance data between the mobile robot and the surrounding environment, and environmental sound data, which contain complex feature information and have different features and semantics. When the features of these multiple data sources are fused, there may be inconsistencies or redundant information between them. When the classifier processes the fused features, it may be affected by this inconsistency, resulting in class coherence interference between the weights and the feature vectors. In particular, in the feature fusion process, the scale, distribution, and importance of the features may be different, which may cause the influence of certain features on the classifier to be overemphasized or ignored. Therefore, in the technical solution of the present application, the robot avoidance judgment classification feature vector is subjected to class coherence interference correction based on weight adaptation to obtain an optimized robot avoidance judgment classification feature vector.
[0049] Among them, the robot avoidance judgment classification feature vector is corrected by quasi-coherence interference based on weight adaptation to obtain an optimized robot avoidance judgment classification feature vector, including: calculating the product between the robot avoidance judgment classification feature vector and its transposed vector to obtain a robot avoidance judgment feature autocorrelation expression matrix; calculating the coherent interference phase between the category decoding weight matrix of the classifier and the robot avoidance judgment feature autocorrelation expression matrix to obtain a robot avoidance judgment feature-weight fine-grained coupling representation matrix; performing probabilistic processing based on the sigmoid activation function on the robot avoidance judgment feature-weight fine-grained coupling representation matrix to obtain a robot avoidance judgment feature-weight fine-grained coupling probabilistic representation matrix; performing backpropagation expression compensation on the robot avoidance judgment classification feature vector based on the robot avoidance judgment feature-weight fine-grained coupling probabilistic representation matrix to obtain a correlation compensation representation vector; and calculating the position-weighted sum between the correlation compensation representation vector and the robot avoidance judgment classification feature vector to obtain the optimized robot avoidance judgment classification feature vector.
[0050] The coherent interference phase between the classifier's category decoding weight matrix and the robot avoidance judgment feature autocorrelation expression matrix is calculated using the following coupling formula to obtain a robot avoidance judgment feature-weight fine-grained coupling representation matrix;
[0051] Wherein, the coupling formula is:
[0052]
[0053] Among them, V 1i Represents the i-th row vector of the class decoding weight matrix of the classifier, V 2jrepresents the j-th row vector of the robot avoidance judgment feature autocorrelation expression matrix, ∑ represents the covariance matrix of the dataset to which the i-th row vector of the class decoding weight matrix of the classifier and the j-th row vector of the robot avoidance judgment feature autocorrelation expression matrix belong, ∑ -1 represents the inverse matrix of the covariance matrix, A(V 1i ,V 2j ) represents the eigenvalue of the (i, j)th position of the robot avoidance judgment feature-weight fine-grained coupling probabilistic representation matrix.
[0054] The following compensation formula is used to perform backpropagation expression compensation on the robot avoidance judgment classification feature vector based on the robot avoidance judgment feature-weight fine-grained coupling probabilistic representation matrix to obtain a related compensation representation vector;
[0055] Wherein, the compensation formula is:
[0056]
[0057] Among them, V r represents the relevant compensation representation vector, A represents the robot avoidance judgment feature-weight fine-grained coupling probabilistic representation matrix, V c represents the robot avoidance judgment classification feature vector, Represents matrix multiplication.
[0058] In the technical solution of the present application, when the robot avoidance judgment classification feature vector is classified by the classifier, since the weight of the classifier also needs to be adapted to the robot avoidance judgment classification feature vector, class coherence interference with the robot avoidance judgment classification feature vector may occur. Based on this, in the technical solution of the present application, the robot avoidance judgment classification feature vector is corrected by class coherence interference based on weight adaptation. First, the autocorrelation matrix of the robot avoidance judgment classification feature vector is used to represent the autocorrelation interference feature spectrum. Then, by calculating the correlation measure between the autocorrelation matrix of the robot avoidance judgment classification feature vector and the class decoding weight matrix of the classifier, the fine-grained feature-class label coherence interference phase between the robot avoidance judgment feature and the classifier weight matrix is simulated. Then, the robot avoidance judgment feature-weight fine-grained coupling representation matrix is used to perform backpropagation expression compensation on the robot avoidance judgment classification feature vector. By reverse compensation, the equivalent probability intensity representation of the robot avoidance judgment classification feature vector in the absence of interference is restored, and the optimized robot avoidance judgment classification feature vector is obtained to improve the accuracy of the classification result.
[0059] Furthermore, the optimized robot avoidance judgment classification feature vector integrates the truck tracking feature vector and the multimodal environmental association feature vector, providing a comprehensive view of the environment. These feature vectors encompass both the dynamic relationship and static information between the robot, the truck, and other environmental factors, covering important factors such as distance, speed, and relative motion. By integrating these features, the classifier can process and analyze this complex data to generate effective classification results. The classifier can be trained to distinguish between the two categories of "avoidance required" and "no avoidance required." This real-time processing capability is crucial for ensuring the robot's safety in dynamic environments. The classifier provides fast, automatic decision support, reducing reliance on human intervention.
[0060] In summary, the embodiment of the present application first obtains multi-area images of the mobile robot's surrounding environment collected by the camera, ultrasonic distance data between the mobile robot and the surrounding environment collected by the ultrasonic sensor, and environmental sound data at multiple predetermined time points collected by the sound sensor, and then uses deep learning technology to perform feature extraction and correlation analysis on the three. Finally, a classifier is used to determine whether the mobile robot needs to perform avoidance operations, thereby improving the mobile robot's ability to understand the surrounding environment in a complex and dynamic industrial environment, and ensuring efficient and safe operation of mobile trucks in industrial environments.
[0061] As described above, the mobile robot mobile truck identification system 100 according to the embodiments of the present application can be implemented in various terminal devices. In one example, the mobile robot mobile truck identification system 100 can be integrated into the terminal device as a software module and / or hardware module. For example, the mobile robot mobile truck identification system 100 can be a software module in the operating system of the terminal device, or can be an application developed for the terminal device; of course, the mobile robot mobile truck identification system 100 can also be one of the many hardware modules of the terminal device.
[0062] Alternatively, in another example, the mobile robot's mobile truck identification system 100 and the terminal device may also be separate devices, and the mobile robot's mobile truck identification system 100 may be connected to the terminal device via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.
[0063] Figure 5 FIG. 1 is a flow chart of a method for identifying a mobile truck of a mobile robot according to an embodiment of the present application. Figure 5As shown, the mobile truck identification method of a mobile robot according to an embodiment of the present application includes: S110, acquiring multi-region images of the mobile robot's surrounding environment collected by a camera, ultrasonic distance data between the mobile robot and the surrounding environment collected by an ultrasonic sensor, and environmental sound data at multiple predetermined time points collected by a sound sensor; S120, extracting a surrounding environment truck tracking feature vector and an environmental multimodal association feature vector from the multi-region images of the mobile robot's surrounding environment collected by the camera, the ultrasonic distance data between the mobile robot and the surrounding environment collected by the ultrasonic sensor, and the environmental sound data at multiple predetermined time points collected by the sound sensor; S130, judging whether the mobile robot needs to perform an avoidance operation based on the surrounding environment truck tracking feature vector and the environmental multimodal association feature vector.
[0064] Here, those skilled in the art will appreciate that the specific operations of each step in the above-mentioned mobile robot mobile truck identification method have been described in detail above. Figures 1 to 4 The mobile truck identification system of the mobile robot has been described in detail, and therefore, its repeated description will be omitted.
[0065] Below, reference Figure 6 To describe the electronic device according to the embodiment of the present application.
[0066] like Figure 6 As shown, the electronic device 10 includes an input device 11, an input interface 12, a central processing unit 13, a memory 14, an output interface 15, an output device 16, and a bus 17. The input interface 12, the central processing unit 13, the memory 14, and the output interface 15 are interconnected via the bus 17, and the input device 11 and the output device 16 are connected to the bus 17 via the input interface 12 and the output interface 15, respectively, and are further connected to other components of the electronic device 10.
[0067] Specifically, the input device 11 receives input information from the outside and transmits the input information to the central processing unit 13 through the input interface 12; the central processing unit 13 processes the input information based on the computer-executable instructions stored in the memory 14 to generate output information, stores the output information temporarily or permanently in the memory 14, and then transmits the output information to the output device 16 through the output interface 15; the output device 16 outputs the output information to the outside of the electronic device 10 for user use.
[0068] In one embodiment, Figure 6The electronic device 10 shown can be implemented as a network device, which may include: a memory configured to store programs; a processor configured to run the programs stored in the memory to execute any one of the mobile robot mobile truck identification methods described in the above embodiments.
[0069] According to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program comprising program code for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network and / or installed from a removable storage medium.
[0070] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0071] It is understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present application, and the present application is not limited thereto. Those skilled in the art may make various modifications and improvements without departing from the spirit and substance of the present application, and such modifications and improvements are also considered to be within the scope of protection of the present application.
Claims
1. A mobile truck identification system for a mobile robot, characterized in that: include: a mobile truck identification data acquisition module, configured to acquire images of multiple regions of the mobile robot's surroundings captured by the camera, ultrasonic distance data between the mobile robot and the surroundings captured by the ultrasonic sensor, and ambient sound data at multiple predetermined time points captured by the sound sensor; a mobile truck identification data extraction module, configured to extract a surrounding truck tracking feature vector and an environmental multimodal association feature vector from the multi-region images of the mobile robot's surrounding environment captured by the camera, the ultrasonic distance data between the mobile robot and the surrounding environment captured by the ultrasonic sensor, and the environmental sound data at multiple predetermined time points captured by the sound sensor; a mobile robot avoidance operation judgment module, configured to judge whether the mobile robot needs to perform an avoidance operation based on the surrounding truck tracking feature vector and the environment multimodal association feature vector; The mobile robot avoidance operation judgment module includes: an environmental information feature association unit, configured to associate the surrounding environment truck tracking feature vector with the environmental multimodal association feature vector to obtain a robot avoidance judgment classification feature vector; an environmental information feature optimization unit, configured to perform weight-adaptive quasi-coherence interference correction on the robot avoidance judgment classification feature vector to obtain an optimized robot avoidance judgment classification feature vector; and an avoidance operation classification judgment unit, configured to pass the optimized robot avoidance judgment classification feature vector through a classifier to obtain a classification result, wherein the classification result is used to determine whether the mobile robot needs to perform an avoidance operation. Among them, the environmental information feature optimization unit includes: calculating the product between the robot avoidance judgment classification feature vector and its transposed vector to obtain a robot avoidance judgment feature autocorrelation expression matrix; calculating the coherent interference phase between the category decoding weight matrix of the classifier and the robot avoidance judgment feature autocorrelation expression matrix to obtain a robot avoidance judgment feature-weight fine-grained coupling representation matrix; performing probabilistic processing on the robot avoidance judgment feature-weight fine-grained coupling representation matrix based on the Sigmoid activation function to obtain a robot avoidance judgment feature-weight fine-grained coupling probabilistic representation matrix; performing backpropagation expression compensation on the robot avoidance judgment classification feature vector based on the robot avoidance judgment feature-weight fine-grained coupling probabilistic representation matrix to obtain a related compensation representation vector; calculating the position-weighted sum between the related compensation representation vector and the robot avoidance judgment classification feature vector to obtain the optimized robot avoidance judgment classification feature vector.
2. The mobile robot mobile truck identification system according to claim 1, characterized in that: The mobile truck identification data extraction module includes: A surrounding environment multi-region feature extraction unit, configured to extract features from the multi-region images of the mobile robot's surrounding environment captured by the camera to obtain a surrounding environment truck tracking feature vector; an ultrasonic distance feature extraction unit, configured to extract features from the ultrasonic distance data between the mobile robot and the surrounding environment collected by the ultrasonic sensor to obtain an obstacle distance semantic feature vector; an ambient sound feature extraction unit, configured to extract features from the ambient sound data collected by the sound sensor at a plurality of predetermined time points to obtain an ambient sound spatial state feature vector; The environmental multimodal feature fusion unit is used to fuse the obstacle distance semantic feature vector and the environmental sound space state feature vector to obtain the environmental multimodal association feature vector.
3. The mobile robot mobile truck identification system according to claim 2, characterized in that: The surrounding environment multi-region feature extraction unit includes: A surrounding environment multi-region feature encoding subunit, configured to perform feature encoding on the multi-region images of the mobile robot's surrounding environment captured by the camera to obtain a surrounding environment truck tracking feature map; The surrounding truck feature downsampling subunit is used to downsample the surrounding truck tracking feature map to obtain the surrounding truck tracking feature vector.
4. The mobile robot mobile truck identification system according to claim 3, characterized in that: The surrounding environment multi-region feature encoding subunit includes: Passing the multi-region images of the mobile robot's surrounding environment captured by the camera through a surrounding environment image denoiser based on deep learning to obtain denoised and enhanced images of the surrounding environment in multiple regions; Passing the plurality of regional surrounding environment denoised and enhanced images through a surrounding environment truck tracking YOLO target detection network as a feature extractor to obtain a plurality of regional surrounding environment truck tracking regions of interest; The plurality of regions of interest of surrounding truck tracking are passed through a surrounding truck tracking feature encoder to obtain the surrounding truck tracking feature map.
5. The mobile robot mobile truck identification system according to claim 4, characterized in that: The ultrasonic distance feature extraction unit includes: Passing the ultrasonic distance data between the mobile robot and the surrounding environment collected by the ultrasonic sensor through an obstacle distance text convolutional neural network to obtain a plurality of obstacle distance feature vectors; The multiple obstacle distance feature vectors are concatenated into the obstacle distance semantic feature vector.
6. The mobile robot mobile truck identification system according to claim 5, characterized in that: The environmental sound feature extraction unit includes: Extracting an ambient sound spectrogram from the ambient sound data at a plurality of predetermined time points collected by the sound sensor; The ambient sound spectrum graph is passed through an ambient sound spectrum spatial focusing feature encoder to obtain the ambient sound spatial state feature vector.
7. A mobile truck identification method for a mobile robot, characterized in that: include: Acquire multi-region images of the mobile robot's surrounding environment captured by a camera, ultrasonic distance data between the mobile robot and the surrounding environment captured by an ultrasonic sensor, and environmental sound data at multiple predetermined time points captured by a sound sensor; Extracting a surrounding environment truck tracking feature vector and an environment multimodal association feature vector from the multi-region images of the mobile robot's surrounding environment captured by the camera, the ultrasonic distance data between the mobile robot and the surrounding environment captured by the ultrasonic sensor, and the environmental sound data at multiple predetermined time points captured by the sound sensor; Determining whether the mobile robot needs to perform an avoidance operation based on the surrounding environment truck tracking feature vector and the environment multimodal association feature vector; Wherein, judging whether the mobile robot needs to perform an avoidance operation based on the surrounding environment truck tracking feature vector and the environment multimodal association feature vector includes: correlating the surrounding environment truck tracking feature vector and the environment multimodal association feature vector to obtain a robot avoidance judgment classification feature vector; performing a weight-adapted quasi-coherence interference correction on the robot avoidance judgment classification feature vector to obtain an optimized robot avoidance judgment classification feature vector; passing the optimized robot avoidance judgment classification feature vector through a classifier to obtain a classification result, and the classification result is used to judge whether the mobile robot needs to perform an avoidance operation; Among them, the robot avoidance judgment classification feature vector is subjected to a quasi-coherent interference correction based on weight adaptation to obtain an optimized robot avoidance judgment classification feature vector, including: calculating the product between the robot avoidance judgment classification feature vector and its transposed vector to obtain a robot avoidance judgment feature autocorrelation expression matrix; calculating the coherent interference phase between the category decoding weight matrix of the classifier and the robot avoidance judgment feature autocorrelation expression matrix to obtain a robot avoidance judgment feature-weight fine-grained coupling representation matrix; performing probabilistic processing based on the Sigmoid activation function on the robot avoidance judgment feature-weight fine-grained coupling representation matrix to obtain a robot avoidance judgment feature-weight fine-grained coupling probabilistic representation matrix; performing backpropagation expression compensation on the robot avoidance judgment classification feature vector based on the robot avoidance judgment feature-weight fine-grained coupling probabilistic representation matrix to obtain a correlation compensation representation vector; and calculating the position-weighted sum between the correlation compensation representation vector and the robot avoidance judgment classification feature vector to obtain the optimized robot avoidance judgment classification feature vector.
8. The mobile truck identification method of a mobile robot according to claim 7, characterized in that: Extracting a surrounding environment truck tracking feature vector and an environment multimodal association feature vector from the multi-region images of the mobile robot's surrounding environment captured by the camera, the ultrasonic distance data between the mobile robot and the surrounding environment captured by the ultrasonic sensor, and the environmental sound data at multiple predetermined time points captured by the sound sensor, including: Performing feature extraction on the multi-region images of the mobile robot's surrounding environment captured by the camera to obtain a truck tracking feature vector of the surrounding environment; Performing feature extraction on ultrasonic distance data between the mobile robot and the surrounding environment collected by the ultrasonic sensor to obtain an obstacle distance semantic feature vector; Performing feature extraction on the ambient sound data collected by the sound sensor at a plurality of predetermined time points to obtain an ambient sound spatial state feature vector; The obstacle distance semantic feature vector and the ambient sound space state feature vector are fused to obtain the ambient multimodal association feature vector.
Citation Information
Patent Citations
Bucket wheel damage judgment method and system based on image recognition
CN115620270A
Production workshop construction operation safety early warning system and method based on Internet of Things
CN118053260A
Environmental perception in autonomous driving using captured audio
US20200209882A1