Autonomous identification and positioning system and method for rescue robots based on industrial vision
By combining industrial vision and LiDAR scanning with inertial measurement and audio sensing information, multimodal image fusion is performed, which solves the problem of inaccurate target positioning in complex environments, achieves efficient target identification and path planning, and improves rescue efficiency.
Patent Information
- Application Number
- CN202510705463.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing rescue robots struggle to accurately locate rescue targets in complex environments, resulting in low rescue efficiency.
By employing industrial vision and lidar collaborative scanning, combined with inertial measurement and audio sensing information, multimodal fusion of thermal imaging and visible light images is performed. Through local human body recognition and attitude reasoning, the complete structural features of the rescue target are supplemented, and path planning is carried out.
Accurately identifying rescue targets in harsh environments improves the precision of target location, ensures that rescue robots can safely and quickly reach the target area, and enhances rescue efficiency.
Smart Images

Figure CN120245004B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and in particular to an autonomous identification and positioning system and method for rescue robots based on industrial vision. Background Art
[0002] Rescue robots are typically equipped with autonomous recognition systems to locate and identify rescue targets in complex environments. Existing target recognition methods mainly rely on RGB image-based or single-sensor technologies, which perform well in open environments. However, in complex environments, they are easily affected by environmental factors such as smoke, dust, insufficient lighting, and heat sources. The rescue target may be partially obscured, making it impossible to accurately identify the location and posture of trapped individuals. This leads to inaccurate target identification in dynamic environments, affecting the timeliness and effectiveness of rescue operations.
[0003] In summary, existing technologies suffer from the problem of low rescue efficiency due to obstructions and interference in complex environments, making it difficult to accurately locate rescue targets. Summary of the Invention
[0004] The purpose of this application is to provide an autonomous identification and positioning system and method for rescue robots based on industrial vision, in order to solve the technical problem in the prior art that it is difficult to accurately locate the rescue target due to occlusion and interference in complex environments, resulting in low rescue efficiency.
[0005] In view of the above problems, this application provides an autonomous identification and positioning system and method for rescue robots based on industrial vision.
[0006] Firstly, this application provides an autonomous identification and positioning system for rescue robots based on industrial vision. The system includes: a collaborative scanning module for initializing an industrial vision device and constructing a robot coordinate map through collaborative scanning with a lidar system; a data fusion module for establishing a unified temporal data fusion frame based on the robot coordinate map, combined with inertial measurement and audio sensing information, to obtain multimodal image data, including thermal imaging images and visible light images; a human body recognition module for mapping the thermal imaging image to the visible light image, performing local human body recognition based on the visible light image, and performing posture reasoning and completion based on local key points to construct human body structural features; a target annotation module for determining rescue targets based on the human body structural features, annotating the target data of the rescue targets on the robot coordinate map to obtain annotation information; and a path planning module for performing path planning based on the annotation information, constructing and updating the path, and then performing rescue operations via the rescue robot.
[0007] Optionally, the device activation unit is used to activate the industrial vision device, verify the connection status and sensor signal validity of the industrial vision device, and obtain a self-test result; the map building unit is used to construct an initial robot coordinate map through collaborative scanning by the industrial vision device and the LiDAR; the map updating unit is used to extract passable areas and obstacle boundaries based on the LiDAR point cloud, set the navigation origin, update the initial robot coordinate map, and generate the robot coordinate map.
[0008] Optionally, a temperature analysis unit is used to perform regional temperature analysis on the thermal imaging image, extract candidate heat source regions, and establish region boundaries; a confidence evaluation unit is used to perform confidence evaluation on the candidate heat source regions to obtain heat source regions, and map them onto the visible light image.
[0009] Optionally, the heat distribution screening channel is used to extract texture information, shape information, and edge information of the heat source candidate region to screen human body heat distribution and obtain heat map morphology; the signal judgment channel is used to judge metabolic activity signals by gas concentration changes obtained by gas concentration sensor and obtain gas concentration; the audio signal acquisition channel is used to acquire audio segments based on voice sensor and detect voice rhythm signals to obtain audio signals; and the heat source region determination channel is used to evaluate the confidence of heat source candidate regions based on the heat map morphology, gas concentration, and audio signal to obtain the heat source region.
[0010] Optionally, the image comparison unit is used to determine whether there is a lack of human body recognition features based on the comparison between the visible light image and the thermal imaging image; the key point extraction unit is used to extract local key points in the visible area if there is a lack of human body recognition features; and the structure reconstruction unit is used to combine the local contour and temperature gradient in the thermal imaging image, use a human pose completion network, and reconstruct human body structural features based on the local key points.
[0011] Optionally, the spatial depth information calculation unit is used to calculate the spatial depth information of the target area where the rescue target is located through binocular vision; the three-dimensional structure correction unit is used to perform three-dimensional structure correction of the spatial depth information by fusing the lidar scanning results to obtain the target spatial information; the target position mapping unit is used to map the target spatial position to the robot coordinate map and record the rescue target position, occlusion level and detection time information of the rescue target; the target annotation unit is used to annotate the target data in the robot coordinate map and upload it to the emergency command platform, and broadcast it to the collaborative rescue robot nodes for sharing.
[0012] Optionally, a path construction channel is used to construct an initial path in the robot coordinate map based on the annotation information; a path update channel is used to collect real-time target space information on the initial path by the rescue robot, and when a new obstacle is detected in the real-time target space information, a path reconstruction mechanism is triggered to update the initial path in the robot coordinate map to obtain an updated path.
[0013] Optionally, an unreachable target marking channel is used to divide reachable paths and unreachable paths based on the updated path, mark the rescue robot as an unreachable target of the rescue target according to the unreachable path, and broadcast it to the collaborative rescue robot for execution.
[0014] Optionally, a data upload channel is used to organize the target data into a standard format data packet and upload it to the emergency command platform via a wireless communication module; a data broadcast channel is used to broadcast the target data in the robot swarm network; and a label update channel is used to set a periodic status update mechanism for the target data, updating the label information and triggering the tracking module when the target data changes.
[0015] Secondly, this application also provides an autonomous identification and positioning method for rescue robots based on industrial vision. The method includes: initializing an industrial vision device; constructing a robot coordinate map through collaborative scanning using the industrial vision device and a lidar; establishing a unified temporal data fusion frame based on the robot coordinate map, combined with inertial measurement and audio sensing information, to obtain multimodal image data, wherein the multimodal image data includes thermal imaging images and visible light images; mapping the thermal imaging image to the visible light image; performing local human body recognition based on the visible light image; performing posture reasoning and completion based on local key points to construct human body structural features; determining the rescue target based on the human body structural features; marking the target data of the rescue target on the robot coordinate map to obtain annotation information; performing path planning based on the annotation information; constructing and updating the path; and performing rescue operations using the rescue robot.
[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0017] The system employs a collaborative scanning module to initialize the industrial vision equipment and construct a robot coordinate map through collaborative scanning with the industrial vision equipment and LiDAR. A data fusion module, based on the robot coordinate map and inertial measurement and audio sensing information, establishes a unified temporal data fusion frame to obtain multimodal image data, including thermal and visible light images. A human body recognition module maps the thermal image to the visible light image, performs local human body recognition based on the visible light image, and performs posture reasoning and completion based on local key points to construct human structural features. A target annotation module determines the rescue target based on the human structural features and annotates the target data of the rescue target on the robot coordinate map to obtain annotation information. A path planning module performs path planning based on the annotation information, constructs and updates the path, and the rescue robot performs the rescue operation. In other words, by using industrial vision and lidar to scan in tandem, combined with inertial measurement and audio sensing information, multimodal fusion of thermal and visible light images is achieved. Even in harsh environments, rescue targets can still be identified. Based on local human body recognition and posture reasoning, the complete structural features of the rescue target are supplemented, the rescue target is determined, and path planning is performed, which improves the accuracy of rescue target positioning and ensures that the rescue robot can safely and quickly reach the target area, thereby improving rescue efficiency.
[0018] The above description is merely an overview of the technical solution of this application. To better understand the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the autonomous identification and positioning system for rescue robots based on industrial vision, as described in this application.
[0021] Figure 2 This is a flowchart illustrating the autonomous identification and positioning method for rescue robots based on industrial vision, as described in this application.
[0022] Figure labeling: Collaborative scanning module 11, data fusion module 12, human body recognition module 13, target annotation module 14, path planning module 15. Detailed Implementation
[0023] This application provides an autonomous identification and positioning system and method for rescue robots based on industrial vision. It addresses the technical problem in existing technologies where accurate target location is difficult due to occlusion and interference in complex environments, leading to low rescue efficiency. By employing collaborative scanning of industrial vision and LiDAR, combined with inertial measurement and audio sensing information, and performing multimodal fusion of thermal and visible light images, the system can identify rescue targets even in harsh environments. Based on local human body recognition and posture reasoning, it completes the structural features of the rescue target, determines the target, and performs path planning, improving the accuracy of target location and ensuring the rescue robot can safely and quickly reach the target area, thereby enhancing rescue efficiency.
[0024] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.
[0025] Example 1, please refer to the appendix. Figure 1 This application provides an autonomous identification and positioning system for rescue robots based on industrial vision. The system is used to implement the steps of an autonomous identification and positioning method for rescue robots based on industrial vision. The autonomous identification and positioning system for rescue robots based on industrial vision includes:
[0026] The collaborative scanning module 11 is used to perform industrial vision device initialization and construct a robot coordinate map through collaborative scanning of the industrial vision device and LiDAR.
[0027] Furthermore, the collaborative scanning module 11 in the industrial vision-based autonomous identification and positioning system for rescue robots is also used for:
[0028] The device activation unit is used to activate the industrial vision device, verify the connection status and sensor signal validity of the industrial vision device, and obtain a self-test result; the map building unit is used to construct an initial robot coordinate map through collaborative scanning by the industrial vision device and LiDAR; the map updating unit is used to extract passable areas and obstacle boundaries based on the laser point cloud, set the navigation origin, update the initial robot coordinate map, and generate the robot coordinate map.
[0029] Specifically, industrial vision devices are used to acquire image or video signals, typically including high-resolution cameras, infrared cameras, depth cameras, etc., capable of capturing visual information in the environment. Initializing industrial vision devices ensures their proper functioning. During initialization, the devices perform self-tests to verify the integrity of their connections and the effectiveness of their sensors. Connection status includes the security of physical connections (such as cables and interfaces), the integrity of logical connections (such as connections established between devices via software or protocols), and communication protocol compatibility (such as whether the same communication protocol is used).
[0030] Simultaneously, the validity of the sensor signals from industrial vision equipment is verified, including signal integrity, accuracy, and stability. Signal integrity checks for damage or loss of image data. For example, this is done by calculating the image checksum and / or using error detection codes to verify data integrity. Signal accuracy verifies whether the image and sensor data accurately reflect environmental conditions, by comparing data from different sensors or using known reference objects. Signal stability ensures that the sensor signal remains stable over a period of time, without abnormal fluctuations or noise, by monitoring historical signal data or using filtering techniques to assess signal stability.
[0031] After the industrial vision equipment completes its self-test, the self-test results are obtained to ensure that all functions are normal, including connection status, whether the sensors are working properly, such as equipment status (e.g., whether it is online or faulty) and the quality of sensor signals (e.g., signal strength and stability).
[0032] After initialization, the industrial vision equipment works simultaneously with LiDAR to collaboratively scan the surrounding environment and construct a preliminary robot coordinate map. LiDAR generates 3D point cloud data of the environment by emitting laser light and receiving echo data, providing geometric information such as obstacle locations and open areas, as well as spatial structure information. The industrial vision equipment provides visual information about the environment through image data, which, combined with the laser point cloud data, enhances the rescue robot's environmental perception capabilities. The core function of collaborative scanning is to improve the accuracy and reliability of environmental perception by combining the information from both. Using either vision equipment or LiDAR alone may be affected by environmental conditions (such as low light, occlusion, and reflection), but through collaborative use, they can complement each other, resulting in a more comprehensive understanding of the environment.
[0033] LiDAR and industrial vision devices generate different data types: LiDAR produces 3D point cloud data, while industrial vision devices produce 2D or depth images. These data are fused to construct a multi-dimensional information system. LiDAR provides spatial structure information (such as distance and position), while vision devices provide object recognition information (such as object shape, color, and texture). After collaborative scanning, the LiDAR point cloud data is processed to filter out noise points and extract valid point clouds; simultaneously, the information in the visual images is used for target detection and image segmentation to identify important objects in the environment. Through coordinate alignment, the data from both systems are fused to generate an initial coordinate map with both spatial and visual information.
[0034] Using data from LiDAR scans and image information from industrial vision devices, an initial robot coordinate map is generated, including obstacles, passable areas, and target objects. The LiDAR point cloud data accurately describes the structure in three-dimensional space, while the vision devices provide morphological information about objects, helping the rescue robot identify specific objects in the environment (such as trees, walls, tables, and people).
[0035] In the point cloud data acquired by LiDAR, the coordinates of each point represent the specific location of that point in the environment. Noise filtering and downsampling are performed on the LiDAR point cloud to remove useless data. Clustering algorithms (such as Euclidean clustering) are then used to divide the point cloud data into different object regions and identify the locations of obstacles. Dense regions in the point cloud data typically represent obstacles; LiDAR can identify these regions based on the density and geometry of the point cloud, separating obstacle boundaries from the surrounding open areas.
[0036] A navigation origin is established as a reference point, typically the coordinates of the rescue robot's current location. For example, using data from LiDAR and other sensors (such as IMU or GPS), the rescue robot can determine its current location and establish a navigation origin. The LiDAR point cloud processing results are then fused with existing map data to dynamically update the initial robot coordinate map, generating a more accurate map that includes a detailed environmental layout, containing information on all obstacles, traversable areas, and the navigation origin. The rescue robot can then perform path planning, obstacle avoidance, and navigation based on this coordinate map. Through the collaborative work of industrial vision equipment and LiDAR, the rescue robot can acquire rich visual and spatial information in complex environments, enabling precise positioning and environmental perception. The point cloud data generated by LiDAR helps the rescue robot identify traversable areas and obstacles, providing a basis for planning the robot's route and preventing collisions in complex environments.
[0037] The data fusion module 12 is used to establish a unified temporal data fusion frame based on the robot coordinate map and inertial measurement and audio sensing information to obtain multimodal image data, wherein the multimodal image data includes thermal imaging images and visible light images.
[0038] Specifically, inertial measurement units (IMUs) and audio sensors are used to acquire inertial measurement and audio sensing information of the target rescue area. An IMU is a sensor used to measure an object's acceleration, angular velocity, and other motion parameters, providing real-time motion information of the rescue target within the target rescue area to assist in positioning and attitude estimation, playing a crucial role, especially in environments with insufficient GPS signals. Audio sensors (such as microphone arrays) are used to collect sound data from the target rescue area environment, enabling the identification and location of sound sources, and detecting shouts or knocking sounds from trapped individuals.
[0039] By fusing robot coordinate maps, inertial measurement, and audio sensor information, information from different sensors is integrated into a unified framework. By synchronizing and adjusting the data from different sensors, temporal consistency is ensured, thereby establishing a multimodal image data fusion frame. Thermal imaging sensors capture thermal images, displaying the temperature distribution of the target area. In rescue missions, thermal images can clearly show the location of trapped personnel, especially in low-light or smoky environments, helping to locate heat sources (trapped personnel) even with poor visibility. Visible light images acquired by cameras provide more detailed information about the target area, confirming whether the target in the thermal imaging image is indeed a trapped person. For example, if a hot spot is detected in the thermal imaging image, visible light images can be used to further confirm the target's external features (such as clothing color, posture, etc.) to ensure accuracy.
[0040] Combining thermal imaging with visible light images further confirms the location and identity of trapped individuals. For example, in a dark environment, thermal imaging might reveal a warmer area, while visible light images can supplement this by confirming the area's shape, color, and other features, ensuring accurate identification. By using audio sensors for localization and motion information provided by the IMU (Integrated Sound Unit), the robot can accurately identify and track the specific location of trapped individuals. When the IMU detects subtle changes in the trapped individual's movement, the audio sensors may pick up distress calls or knocking sounds, helping the robot further confirm the target's location and reducing the possibility of misidentification.
[0041] By fusing data from multiple sensors (such as IMU, audio sensors, thermal imaging, and visible light images), the robot can accurately identify the location of trapped individuals and perform effective localization in complex environments. Even in the presence of obstructions, smoke, or low light, the robot maintains high recognition accuracy.
[0042] The human body recognition module 13 is used to map the thermal imaging image to the visible light image, perform local human body recognition based on the visible light image, perform pose reasoning and completion based on local key points, and construct human body structural features.
[0043] Furthermore, the human body recognition module 13 in the industrial vision-based autonomous identification and positioning system for rescue robots is also used for:
[0044] The temperature analysis unit is used to perform regional temperature analysis on the thermal imaging image, extract candidate heat source regions, and establish region boundaries; the confidence evaluation unit is used to evaluate the confidence of the candidate heat source regions to obtain heat source regions, and map them onto the visible light image.
[0045] The heat distribution screening channel is used to extract texture, shape, and edge information of the heat source candidate region and screen the human body heat distribution to obtain a heat map morphology; the signal judgment channel is used to judge metabolic activity signals by gas concentration changes obtained from a gas concentration sensor and obtain gas concentration; the audio signal acquisition channel is used to acquire audio segments based on a speech sensor and detect speech rhythm signals to obtain an audio signal; and the heat source region determination channel is used to evaluate the confidence of the heat source candidate region based on the heat map morphology, gas concentration, and audio signal to obtain the heat source region.
[0046] Specifically, thermal imaging images undergo regional temperature analysis to determine the temperature distribution across different areas and identify potential heat source regions, or candidate heat source regions. This is typically done by setting a temperature threshold; areas exceeding this threshold are considered candidate heat source regions. The temperature threshold is usually set based on the actual ambient temperature and human body temperature. Candidate heat source regions, identified after regional temperature analysis and filtered according to the set threshold, may include trapped individuals. For example, if the set temperature threshold is 35°C, areas in the thermal imaging image with temperatures exceeding 35°C will be marked as candidate heat source regions, potentially including people (whose body temperature is usually higher than the surrounding environment), machinery, or other heat-generating objects.
[0047] Once candidate heat source regions are identified, their boundaries are established. Edge detection is used to determine the outer contour of these regions. Edge detection algorithms identify object boundaries by finding areas of significant pixel intensity variation in the image. In thermal imaging images, temperature differences between the object and the background cause significant changes in pixel values across different regions. Edges are found by detecting areas with large grayscale variations in the image, typically generating a binary image where edge regions are represented by 1 (white) and non-edge regions by 0 (black). The edge detection results may contain broken edges or small noise points. To ensure the extracted heat source boundaries are complete and accurate, morphological operations (such as dilation and erosion) are typically used to smooth and connect the edges. Dilation connects broken edges, while erosion removes small noise points. Once edge detection is complete, the edge image represents the outer boundary of the candidate heat source region.
[0048] After identifying candidate heat source regions, texture, shape, and edge information are extracted. Texture information refers to the variation pattern of light reflected from the surface in the image, typically reflecting the surface structure of an object, to distinguish the surface types of different objects or regions. By calculating the spatial relationship of pixel grayscale, texture features of the thermal imaging image are extracted, such as contrast, homogeneity, and entropy. Shape information refers to the geometric features of the image region, such as boundary shape, aspect ratio, and roundness, used to distinguish different types of targets (e.g., human bodies, equipment) and further identify whether the target is a trapped person. By calculating the aspect ratio of the candidate heat source region, it is determined whether the region conforms to the conventional proportions of the human body (e.g., the aspect ratio of the human body is generally within a certain range). By analyzing the external contour of the heat source region, the shape information of the region is further extracted, such as whether it presents the common contours of the human body. Edge information refers to parts of the image with drastic changes in brightness or color, helping to determine the contour and shape of objects, reflecting the external boundary of the heat source region. Through a multi-stage edge detection process, the edge information of the candidate heat source region is accurately extracted.
[0049] By analyzing texture, shape, and edge information, regions conforming to the thermal distribution patterns of the human body are selected. The typical thermal distribution of the human body is characterized by a warmer head, a more uniform temperature distribution in the torso and limbs, but lower temperatures in the limbs. The human body usually exhibits certain morphological features, such as a long strip, a near-elliptical shape, or a specific vertical proportion. Based on these characteristics, thermal distribution patterns are selected to generate a thermal map. The thermal map pattern displays the temperature intensity at different locations within the candidate regions, thus helping to confirm the type and location of the target.
[0050] Heat maps of different colors are generated based on the temperature values within a region; higher temperatures are typically represented in red, while lower temperatures are represented in blue. The heat maps correlate temperature values with heat levels, helping to identify whether high-temperature areas match the thermal distribution characteristics of the human body. By analyzing the texture, shape, and edge information of candidate heat source regions, a morphological map showing the thermal distribution of the heat source region is obtained, illustrating the temperature distribution at different locations within the region, thus helping to identify and confirm the nature and location of the target.
[0051] Gas concentration sensors, such as carbon dioxide sensors, carbon monoxide sensors, and oxygen sensors, are used to monitor the gas concentration in a target area. Each sensor detects the concentration of a specific gas using different working principles (such as infrared absorption and semiconductor reactions). During metabolic activities, especially respiration, the human body emits carbon dioxide. When the body is active, metabolic activity increases, and carbon dioxide emissions also increase. By monitoring changes in the concentration of carbon dioxide in the air, it is possible to infer whether a trapped person is active or in an emergency. For example, suppose a trapped person is in a fire; their metabolic activity increases, leading to a significant increase in carbon dioxide emissions. In this case, a gas concentration sensor (such as a carbon dioxide sensor) detects the increased carbon dioxide concentration in the area, thus determining the trapped person's metabolic activity signal.
[0052] Gas concentration sensors continuously record changes in the concentration of specific gases (such as carbon dioxide) in real time. In rescue environments, gas concentration data, combined with thermal imaging images and data from other sensors (such as audio sensors and temperature sensors), can provide additional information for target location. When a gas concentration sensor detects an increase in carbon dioxide concentration, this signal can be combined with data from other sensors (such as heat source candidate areas in thermal imaging images) to improve the accuracy of identifying trapped personnel. Analysis of gas concentration changes can reveal signals of metabolic activity. For example, a sustained increase in carbon dioxide concentration, such as a rapid rise from 500 ppm to 1000 ppm, indicates ongoing metabolic activity in a certain area, usually meaning that the trapped person may still be active. If the carbon dioxide concentration reaches 1500 to 2000 ppm, it may indicate that the trapped person in that area is experiencing drowsiness, dizziness, etc.; if the carbon dioxide concentration reaches 5000 ppm, severe bodily dysfunction may occur, leading to loss of consciousness.
[0053] By analyzing changes in gas concentration, signals of metabolic activity can be identified. Real-time monitoring of gas concentration changes using gas concentration sensors and analysis of the metabolic activity signals of trapped individuals helps in the timely location and identification of them. Simultaneously, it provides gas concentration, i.e., the volume or mass ratio of a certain gas in a specific space, such as the concentration of carbon dioxide or oxygen. For example, sample data on gas concentration changes monitored at a fire scene are shown in Table 1:
[0054] Table 1. Sample data on gas concentration changes
[0055]
[0056] Audio segments are obtained by capturing sound signals from the target rescue area using a voice sensor. These segments consist of sound data collected from the voice sensor over a specific period, including speech, ambient sounds, and shouts. To ensure high-quality audio signal analysis, the audio segments are first preprocessed, including noise reduction and enhancement to remove irrelevant background noise and retain only audio signals relevant to the trapped individuals. Audio signal processing techniques (such as speech signal processing and Fourier transform) are then used to analyze the rhythmic characteristics of the audio signals, detecting pauses, pitch changes, and speech rate. Identifying rapid speech rhythms or prolonged high-frequency sounds helps determine the emotional state and urgency of the speech. For example, when the audio signal exhibits frequent, rapid shouts, a continuous increase in speech rate and volume, it indicates that the trapped individuals are in a state of extreme distress and desperate plea for help.
[0057] Audio clips are input into an existing audio signal recognizer. Based on changes in pitch, speech rate, pauses, and tone, the type of speech (e.g., cries for help, shouts) and the emotions they evoke (e.g., urgency, panic) are determined. Combined with data from other sensors (e.g., thermal imaging, gas concentration sensors), the reliability of the audio signal is further confirmed, ensuring the accuracy of rescue decisions. For example, a faint groan detected at 6 to 10 seconds of the audio clip indicates the presence of a trapped person in a poor physiological state; a loud cry for help detected at 15 to 20 seconds, with rhythm analysis showing high frequency, rapid pace, and emotional agitation, indicates that the trapped person is urgently seeking help and may be seriously injured or drowsy.
[0058] Based on heat map morphology, gas concentration, and audio signal, a confidence assessment of candidate heat source regions is performed. First, the heat map morphology, gas concentration, and audio signal are standardized to eliminate the influence of dimensions. Then, according to the weights corresponding to the heat map morphology, gas concentration, and audio signal, typically with higher weights (heat map morphology 0.4, audio signal 0.4, and gas concentration 0.3), a confidence score is calculated for each candidate heat source region, indicating whether trapped personnel exist in any of these regions. Based on the confidence score, regions with higher scores are selected as heat source regions, indicating that trapped personnel exist in these areas and require urgent rescue. For example, standardizing the heat map morphology, gas concentration, and audio signal of a candidate heat source region yields a standardized value of 0.8 for the heat map morphology, 0.9 for the audio signal, and 0.6 for the gas concentration, resulting in a confidence score of 0.86. This indicates that the candidate heat source region is highly likely to be the location of trapped personnel and therefore requires urgent rescue. Confidence scores are typically between 0 and 1, with scores close to 1 indicating a high probability of trapped individuals in the area.
[0059] Mapping heat source areas onto visible light images helps rescue robots determine whether these areas correspond to trapped individuals. For example, a thermal imaging image might show a danger zone, while a visible light image allows the robot to confirm whether a human body actually exists in that area, further eliminating false positives. By aligning the heat source locations in the thermal imaging image with the visible light image, the location of a trapped person can be marked in the visible light image. By comprehensively using data from thermal imaging images, gas concentration sensors, and voice sensors, and through multimodal information fusion, confidence assessment is performed on candidate heat source areas, ultimately determining the area most likely to contain a trapped person. This not only improves the accuracy of rescue robots but also effectively reduces false positives caused by obstruction and interference.
[0060] Furthermore, the human body recognition module 13 in the industrial vision-based autonomous identification and positioning system for rescue robots is also used for:
[0061] The image comparison unit is used to determine whether there is a lack of human body recognition features based on the comparison between the visible light image and the thermal imaging image; the key point extraction unit is used to extract local key points in the visible area if there is a lack of human body recognition features; the structure reconstruction unit is used to combine the local contour and temperature gradient in the thermal imaging image, use a human pose completion network, and reconstruct human body structural features based on the local key points.
[0062] Specifically, comparing visible light images and thermal images is crucial. Visible light images provide information such as the color and shape of an object, while thermal images provide its temperature. By comparing these two, it's possible to effectively identify any missing human features. For example, due to obstacles or other factors in the environment, certain parts of the human body may not be accurately displayed in a thermal image, or they may be obscured in a visible light image. Determining whether human features are missing refers to situations where the human body outline in the visible light image and the thermal image does not completely overlap or is incompletely displayed. For instance, in a thermal image, certain parts (such as hands or feet) may not be clearly displayed due to distance, angle, or obstruction, resulting in a lack of overall human body outline recognition.
[0063] If a body part or localized heat source is missing in a thermal imaging image, the missing information is usually supplemented by extracting residual features (such as the outlines of hands and feet, the edges of clothing, etc.) from the visible light image. For example, if a person's hands or feet are not clearly shown in a thermal imaging image, but are still visible in a visible light image, they can be completed by extracting local key points. Local key points refer to feature points of certain key parts of the human body, such as hands, feet, head, shoulders, knees, etc. By analyzing the visible light image and extracting these local key points, even if some human features are not visible in the thermal imaging image, they can be completed using information from the visible light image.
[0064] In thermal imaging images, the heat distribution of the human body often exhibits a specific thermal gradient (such as the temperature difference between the core and limbs). By analyzing the contours of local heat sources and the temperature gradients in the thermal image, we can further understand the structural features of the human body, especially when some human features are obscured or blurred in the thermal image. Temperature gradients and local contours can provide important supplementary information. Local contours in thermal imaging images refer to the boundaries of heat sources in different parts of the human body, such as the edges of the torso, arms, and legs. They help identify the overall shape of the human body, especially when parts of the body are obscured; the basic structure of the human body can still be inferred from the temperature gradient. The temperature gradient refers to the rate or intensity of temperature change in a thermal imaging image. Temperature differences between different parts of the human body can be presented through temperature gradients; for example, the torso is usually warmer than the arms and legs.
[0065] Human pose completion networks are deep learning-based models used to infer and complete human poses, reconstructing the complete human structure from local joint information, contour information, and temperature gradients. These networks rely on local keypoint information extracted from visible light or thermal images, typically referring to joints (such as shoulders, elbows, and knees), providing preliminary structural information about the human body. When some information is missing, the network can infer the missing portion based on existing local keypoints and temperature gradient information.
[0066] By combining multimodal data from thermal and visible light images, human pose completion networks can infer the spatial positions of various parts of the human body (such as hands, feet, and head) by integrating local keypoints and temperature gradient data. This helps reconstruct the complete structure of the human body, especially when local data is missing. The completion network can effectively infer and supplement data based on existing information. For example, suppose a thermal image shows that the torso of a trapped person is relatively clear, while the limbs are obscured by obstacles. By combining the temperature information of the torso with visible light image information of the hands and feet, the human pose completion network can infer the posture and position of the hands and feet, ultimately reconstructing the complete body structure.
[0067] Training a human pose completion network typically requires a large dataset of labeled human poses, encompassing poses from various scenes, especially samples with missing or occluded keypoints. Due to the large number of samples needed for training, data augmentation is a common practice. Data augmentation can generate more training samples through rotation, scaling, translation, and cropping, thereby improving the model's generalization ability. For network training, each image in the dataset needs labeled keypoint locations, including those for the head, shoulders, elbows, knees, and ankles. For each keypoint, the dataset records its 2D or 3D coordinates (in different coordinate systems).
[0068] Human pose completion networks generally consist of the following main components: an encoder and a decoder. The encoder extracts features from the input image (thermal imaging or visible light image), and the decoder reconstructs missing keypoints based on the features extracted by the encoder. A graph convolutional network (GCNN) models the relationships between different keypoints as a graph structure, allowing the network to complete the pose based on the spatial constraints between keypoints. During training, a loss function is used to evaluate the difference between the network's output pose and the ground truth pose, such as mean squared error (MSE), which calculates the sum of squared Euclidean distances between the network's output keypoint coordinates and the ground truth coordinates. To ensure the network output conforms to the structure of the human skeleton, a spatial constraint loss is typically added. This loss function restricts the relative positional relationships between keypoints, avoiding outputs that do not conform to conventional human anatomy. During training, input data (such as occluded thermal imaging images or partial keypoints) is forward-propagated through the network. Each layer of the network calculates the predicted coordinates of the human keypoints based on the input image. Based on the calculation results of the loss function, the network performs backpropagation. Through backpropagation, the network parameters (such as convolutional kernel weights) are adjusted according to the error to minimize the loss function, thereby optimizing the network's performance. During training, the dataset is typically divided into multiple batches, each containing a number of samples. Training the network on these batches improves training efficiency and stability. Optimal hyperparameter configurations are selected using methods such as cross-validation or grid search. During training, a validation set is frequently used to evaluate the model and ensure its generalization ability. The validation set is a subset of the training data used to check for overfitting. After training, the network can be applied to real-world environments, taking an image of a human body with occluded portions as input and outputting a complete human pose.
[0069] By combining temperature gradient information from thermal imaging images with local key point data from visible light images, a human pose completion network is used to complete missing human body parts. This solves the problem of human body recognition caused by environmental occlusion or incomplete thermal imaging data. It not only improves the recognition accuracy of rescue robots and enhances their adaptability in complex environments, but also significantly improves rescue efficiency, ensuring that robots can quickly identify and locate trapped personnel under extreme conditions.
[0070] The target annotation module 14 is used to determine the rescue target based on the human body structural features, and to annotate the target data of the rescue target in the robot coordinate map to obtain annotation information.
[0071] Furthermore, the target labeling module 14 in the industrial vision-based autonomous identification and positioning system for rescue robots is also used for:
[0072] The system includes a spatial depth information calculation unit for calculating the spatial depth information of the target area where the rescue target is located using binocular vision; a three-dimensional structure correction unit for fusing lidar scanning results to perform three-dimensional structure correction on the spatial depth information to obtain target spatial information; a target position mapping unit for mapping the target spatial position to the robot coordinate map and recording the rescue target position, occlusion level, and detection time information; and a target annotation unit for annotating the target data on the robot coordinate map and uploading it to the emergency command platform, and broadcasting it to the collaborative rescue robot nodes for sharing.
[0073] Specifically, binocular vision uses two cameras to capture the same scene from different angles. By calculating the differences between the two images, it infers the spatial depth information of a target. The basic principle of binocular vision is to use stereo matching technology to compare two images from different viewpoints, thereby calculating the disparity of each pixel and then deducing the spatial coordinates of that point. Two cameras capture the same scene, and the disparity (the displacement difference of the same point in the two viewpoints) is used to calculate depth information. By matching feature points (such as edges, corners, etc.) on the images, triangulation is used to obtain the depth of each feature point (i.e., its distance from the camera). For example, assuming the baseline of the binocular camera is 10cm, the disparity of a certain feature point in the two images is 2 pixels, and the focal length of the camera is known to be 5cm, then the depth can be calculated using the formula D = ... Where D is the target distance, f is the camera's focal length, B is the baseline distance between the two cameras, and d is the parallax. The target depth information is calculated to be 25cm.
[0074] LiDAR (Light Detection and Ranging) scans the surrounding environment with a laser beam, calculates the time it takes for the laser signal to travel from emission to reflection (or measures the intensity of the reflected light), and thus obtains distance information. LiDAR continuously emits laser beams to acquire environmental reflection information, forming dense 3D point cloud data. Each point represents the coordinate position of a reflected laser signal, typically represented as (x, y, z) coordinates. The LiDAR scan results are used to perform 3D structural correction on the spatial depth information. By combining the point cloud data acquired by LiDAR with the depth information calculated by binocular vision, potential errors or inconsistencies can be corrected. By combining the advantages of both, a more accurate target location can be obtained.
[0075] LiDAR scans the environment, generating a large amount of point cloud data. The coordinates of each point contain the three-dimensional position of the target object. Since the coordinate systems of binocular vision and LiDAR are different, extrinsic parameter calibration is required. Through the calibration process, the rotation and translation relationships between the two are found, ensuring that the data from both can be accurately aligned. Kalman filtering, complementary filtering, and other algorithms are used to weightedly fuse the binocular vision data and the LiDAR data to obtain more accurate spatial information. For example, the LiDAR provides point cloud data for a target as (2.3, 1.4, 0.5), while the binocular vision calculates the distance as 4.0 meters. Through data fusion, combining the accuracy and weights of both, the accurate position of the target is finally obtained as (2.4, 1.3, 0.6).
[0076] Mapping the target's spatial location to the robot's coordinate map transforms it into the robot's perspective, facilitating route planning. The target's position on the robot's coordinate map is recorded, along with its occlusion level and detection time. Occlusion can be determined by the positions of other objects in the environment, while the detection time is obtained through sensor timestamps. Occlusion level indicates the degree to which the target is obscured by obstacles or other objects, such as complete occlusion, partial occlusion, and no occlusion. Recording the specific time the target information was detected helps track the target's state in dynamic environments. For example, assuming the target's spatial location is (3.2, 4.1) and the robot's current position is (1.0, 2.0), the target's position in the robot's coordinate system is calculated as (2.2, 2.1) through coordinate transformation. Simultaneously, the target's occlusion level (e.g., partial occlusion) and detection time (e.g., 1:30 PM) are recorded.
[0077] Target information (location, occlusion level, detection time, etc.) is marked on the robot's coordinate map to help command personnel view the specific location of each target and determine the priority order for rescue. The marked target data is uploaded to the emergency command platform in real time via wireless communication, allowing personnel in the command center to monitor the status of the targets and arrange rescue operations at any time.
[0078] Furthermore, the target labeling module 14 in the industrial vision-based autonomous identification and positioning system for rescue robots is also used for:
[0079] An unreachable target marking channel is used to divide reachable and unreachable paths based on the updated path, mark the rescue robot as an unreachable target of the rescue target according to the unreachable path, and broadcast it to the collaborative rescue robot for execution.
[0080] Specifically, the updated path is divided into reachable and inaccessible paths. Reachable paths are those the robot can traverse smoothly without obstacles or other factors hindering its movement; inaccessible paths are those the robot cannot traverse due to obstacles, confined spaces, environmental changes, or other factors. When a robot detects that a path to a target location is inaccessible, it marks the target as inaccessible and broadcasts this information to other collaborative rescue robots in the same task, including the target's location, obstacle details, reason for inaccessibility, and real-time path data. Upon receiving the inaccessible target information, the collaborative rescue robots, based on their sensor data and map information, may choose different paths, attempting to approach the target area from another angle or via a different route. By dividing paths into reachable and inaccessible paths, robot path planning can be dynamically optimized in real-time in complex environments, preventing robots from entering inaccessible areas and improving rescue efficiency and accuracy.
[0081] Furthermore, the target labeling module 14 in the industrial vision-based autonomous identification and positioning system for rescue robots is also used for:
[0082] The data upload channel is used to organize the target data into a standard format data packet and upload it to the emergency command platform via the wireless communication module; the data broadcast channel is used to broadcast the target data in the robot swarm network; the annotation update channel is used to set a periodic status update mechanism for the target data, and when the target data changes, the annotation information is updated and the tracking module is triggered.
[0083] Specifically, to ensure data consistency and compatibility, target data is organized into standard format data packets. This involves encoding and encapsulating different types of information according to a fixed structure, enabling the data to be recognized and processed by different devices and platforms. The formatted data is then encapsulated into data packets according to a predetermined protocol, with necessary verification information and timestamps added to ensure the integrity and accuracy of the data packets. Target data comprises various information collected during the rescue process regarding targets (such as trapped personnel, obstacles, etc.), including the target's spatial location, status, detection time, and obstruction level.
[0084] The robot transmits target data packets to the emergency command platform via its onboard wireless communication module (such as Wi-Fi, 4G / 5G, or a dedicated wireless communication protocol). On the platform, rescue commanders can view the uploaded target data in real time, enabling task allocation, route planning, and decision support.
[0085] The target data uploaded to the emergency intelligent platform is broadcast within the robot swarm network. Other robots in the swarm can receive this data through the network, preventing duplicate identification of rescue missions. Broadcasting data facilitates collaborative work among multiple robots, allowing them to share key information such as target location and status, thus enabling more effective coordination in rescue missions.
[0086] By setting timers (e.g., every minute or every 5 minutes), the target data is periodically checked and updated to detect changes in the target (such as location and status) and update the data accordingly. When the target's status changes (e.g., changes in target location, changes in the status of trapped personnel), the robot updates the target's annotation information, including the target's latest location and occlusion level. If the target changes, or its data annotation information changes significantly, the tracking module is triggered to continuously track the target, ensuring that the target's dynamic information is updated in a timely manner and accurately tracked. By periodically updating the target data and triggering the tracking module, the accuracy and timeliness of the target information are ensured. If the target's status changes, the rescue strategy can be adjusted in a timely manner to ensure that the target is accurately identified and rescued.
[0087] The path planning module 15 is used to plan a path based on the annotation information, construct an updated path, and carry out rescue operations through the rescue robot.
[0088] The path construction channel is used to construct an initial path in the robot coordinate map based on the annotation information; the path update channel is used to collect real-time target space information on the initial path by the rescue robot. When a new obstacle is detected in the real-time target space information, the path reconstruction mechanism is triggered to update the initial path in the robot coordinate map to obtain the updated path.
[0089] Specifically, based on the annotation information in the robot's coordinate map, a path planning algorithm is used to construct an initial path in the robot's coordinate map. The annotation information provides the target location and obstacle locations, and the robot's coordinate map provides the coordinate system framework. The initial path is a preliminary planned route for the robot from the starting point (the robot's current location) to the target point, taking into account factors such as obstacle locations, robot direction of travel, and speed. Depending on the specific application scenario, a suitable path planning algorithm is selected, such as the A* algorithm, to generate a path based on the starting point, target point, obstacle information, and traversable areas of the environment. For example, suppose the starting point is (1.0, 2.0), the target point is (5.0, 6.0), and the obstacles are located at (3.0, 3.0) and (4.0, 4.0) respectively. The step size is set to 1 unit, indicating that the robot moves 1 unit of horizontal or vertical distance each time. The distance from the current node to the target node is estimated using a heuristic function (such as Manhattan distance). The A* algorithm selects the next optimal node by evaluating the cost of each adjacent node (composed of the heuristic function h(x,y) and the path length g(x,y)). The resulting cost graph is as follows: the cost of (2.0,2.0) is f=g+h=1+7=8; the cost of (3.0,2.0) is f=g+h=2+6=8; and the cost of (5.0,2.0) is f=g+h=3+5=8. The node with the lowest cost is selected as the next moving node, and this process is repeated until the target point is reached. The initial path obtained through path planning is (1.0,2.0)→(2.0,2.0)→(3.0,3.0)→(4.0,4.0)→(5.0,6.0). The rescue robot avoids obstacles and this is the current optimal path.
[0090] As the rescue robot travels along its initial path, it continuously collects real-time environmental data, particularly depth information and obstacle locations within the target area. Using devices such as LiDAR and visual sensors, it monitors for newly added obstacles in the target space. If a new obstacle is detected, the rescue robot triggers a path update through a path reconstruction mechanism. The location of the new obstacle is marked on the robot's coordinate map, and a new path avoiding the obstacle is recalculated using the aforementioned path planning algorithm, updating the path on the robot's coordinate map. The updated path is based on information about the new obstacle and real-time target space data, recalculating and adjusting the robot's travel path to ensure it can avoid new obstacles and proceed safely.
[0091] The robot begins navigation based on the updated path, continuously monitoring the distance between itself and the target to ensure it can successfully avoid obstacles and reach the destination. During path execution, the rescue robot constantly collects environmental data, including the target's and obstacle positions, to ensure it stays on the correct route. If the path changes (e.g., a change in target location or the appearance of dynamic obstacles), the robot will update and replan its path. Upon reaching the target, the robot begins its rescue mission, including clearing obstacles, delivering supplies, locating and rescuing trapped personnel. Throughout the process, the robot adjusts and executes its actions in real time according to mission requirements.
[0092] In summary, the autonomous identification and positioning system for rescue robots based on industrial vision provided in this application has the following technical effects:
[0093] The system employs a collaborative scanning module to initialize the industrial vision equipment and construct a robot coordinate map through collaborative scanning with the industrial vision equipment and LiDAR. A data fusion module, based on the robot coordinate map and inertial measurement and audio sensing information, establishes a unified temporal data fusion frame to obtain multimodal image data, including thermal and visible light images. A human body recognition module maps the thermal image to the visible light image, performs local human body recognition based on the visible light image, and performs posture reasoning and completion based on local key points to construct human structural features. A target annotation module determines the rescue target based on the human structural features and annotates the target data of the rescue target on the robot coordinate map to obtain annotation information. A path planning module performs path planning based on the annotation information, constructs and updates the path, and the rescue robot performs the rescue operation. In other words, by using industrial vision and lidar to scan in tandem, combined with inertial measurement and audio sensing information, multimodal fusion of thermal and visible light images is achieved. Even in harsh environments, rescue targets can still be identified. Based on local human body recognition and posture reasoning, the complete structural features of the rescue target are supplemented, the rescue target is determined, and path planning is performed, which improves the accuracy of rescue target positioning and ensures that the rescue robot can safely and quickly reach the target area, thereby improving rescue efficiency.
[0094] Example 2: Based on the same inventive concept as the industrial vision-based autonomous identification and positioning system for rescue robots in Example 1, this application also provides an industrial vision-based autonomous identification and positioning method for rescue robots. Please refer to the appendix. Figure 2 The industrial vision-based autonomous identification and positioning method for rescue robots includes:
[0095] S100: Initialize the industrial vision equipment and construct a robot coordinate map through collaborative scanning with the industrial vision equipment and LiDAR; S200: Based on the robot coordinate map, combine inertial measurement and audio sensing information to establish a unified temporal data fusion frame to obtain multimodal image data, wherein the multimodal image data includes thermal imaging images and visible light images; S300: Map the thermal imaging image to the visible light image, perform human body local recognition based on the visible light image, perform posture reasoning and completion based on local key points, and construct human body structural features; S400: Determine the rescue target based on the human body structural features, and mark the target data of the rescue target in the robot coordinate map to obtain annotation information; S500: Perform path planning based on the annotation information, construct an updated path, and execute the rescue through the rescue robot.
[0096] Furthermore, the initialization of the industrial vision device, which involves co-scanning with the industrial vision device and LiDAR to construct a robot coordinate map, includes: activating the industrial vision device, verifying the connection status and sensor signal validity of the industrial vision device, and obtaining a self-test result; constructing an initial robot coordinate map through co-scanning with the industrial vision device and LiDAR; extracting passable areas and obstacle boundaries based on the LiDAR point cloud, setting the navigation origin, and updating the initial robot coordinate map to generate the final robot coordinate map.
[0097] Furthermore, mapping the thermal imaging image to the visible light image includes: performing regional temperature analysis on the thermal imaging image, extracting candidate heat source regions, and establishing region boundaries; performing confidence assessment on the candidate heat source regions to obtain heat source regions, and mapping them to the visible light image.
[0098] Furthermore, the step of obtaining the heat source region by performing confidence assessment on the heat source candidate region includes: extracting the texture information, shape information, and edge information of the heat source candidate region, screening human body heat distribution to obtain a heat map morphology; judging metabolic activity signals by gas concentration changes obtained through a gas concentration sensor to obtain gas concentration; acquiring audio segments based on a voice sensor, detecting voice rhythm signals to obtain an audio signal; and performing confidence assessment on the heat source candidate region based on the heat map morphology, gas concentration, and audio signal to obtain the heat source region.
[0099] Furthermore, the step of performing human body local recognition based on the visible light image, performing pose reasoning and completion based on local key points, and constructing human body structural features includes: determining whether there are missing human body recognition features based on the comparison between the visible light image and the thermal imaging image; if there are missing human body recognition features, extracting local key points in the visible area; combining the local contour and temperature gradient in the thermal imaging image, using a human pose completion network, and reconstructing human body structural features based on the local key points.
[0100] Furthermore, the step of determining the rescue target based on the human structural features and marking the target data of the rescue target on the robot coordinate map includes: calculating the spatial depth information of the target area where the rescue target is located through binocular vision; performing three-dimensional structural correction of the spatial depth information by fusing the LiDAR scanning results to obtain the target spatial information; mapping the target spatial position to the robot coordinate map, recording the rescue target position, occlusion level and detection time information of the rescue target; marking the target data on the robot coordinate map and uploading it to the emergency command platform, and broadcasting it to the collaborative rescue robot nodes for sharing.
[0101] Furthermore, the broadcast to the collaborative rescue robot node for sharing includes: constructing an initial path in the robot coordinate map based on the annotation information; collecting real-time target space information on the initial path based on the rescue robot; and triggering a path reconstruction mechanism when a new obstacle is detected in the real-time target space information to update the initial path in the robot coordinate map to obtain an updated path.
[0102] Furthermore, the broadcast to the collaborative rescue robot node sharing includes: dividing the updated path into reachable and unreachable paths, marking the rescue robot as an unreachable target of the rescue target according to the unreachable path, and broadcasting it to the collaborative rescue robot for execution.
[0103] Furthermore, the step of marking the target data on the robot coordinate map and uploading it to the emergency command platform, and broadcasting it to the collaborative rescue robot nodes for sharing, includes: organizing the target data into a standard format data packet and uploading it to the emergency command platform via a wireless communication module; broadcasting the target data in the robot cluster network; and setting a periodic status update mechanism for the target data, updating the marking information and triggering the tracking module when the target data changes.
[0104] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Figure 1The autonomous identification and positioning system and specific examples of the rescue robot based on industrial vision in Embodiment 1 are also applicable to the autonomous identification and positioning method of the rescue robot based on industrial vision in this embodiment. Through the foregoing detailed description of the autonomous identification and positioning system of the rescue robot based on industrial vision, those skilled in the art can clearly understand the autonomous identification and positioning method of the rescue robot based on industrial vision in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here. The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0105] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. An autonomous identification and positioning system for rescue robots based on industrial vision, characterized in that, include: The collaborative scanning module is used to perform industrial vision equipment initialization and construct a robot coordinate map through collaborative scanning between the industrial vision equipment and the LiDAR. The data fusion module is used to establish a unified temporal data fusion frame based on the robot coordinate map and inertial measurement and audio sensing information to obtain multimodal image data, wherein the multimodal image data includes thermal imaging images and visible light images; The human body recognition module is used to map the thermal imaging image to the visible light image, perform local human body recognition based on the visible light image, perform pose reasoning and completion based on local key points, and construct human body structural features. The target annotation module is used to determine the rescue target based on the human structural features, and to annotate the target data of the rescue target in the robot coordinate map to obtain annotation information; The path planning module is used to plan a path based on the annotation information, construct an updated path, and execute the rescue through the rescue robot. The human body recognition module includes: The temperature analysis unit is used to perform regional temperature analysis on the thermal imaging image, extract candidate heat source regions, and establish region boundaries. A confidence assessment unit is used to assess the confidence of the candidate heat source region to obtain the heat source region, and then map it onto the visible light image; The confidence assessment unit includes: The heat distribution filtering channel is used to extract the texture information, shape information and edge information of the heat source candidate region, and filter the human body heat distribution to obtain the heat map shape. The signal judgment channel is used to judge metabolic activity signals by the gas concentration change obtained from the gas concentration sensor, and to obtain the gas concentration. The audio signal acquisition channel is used to acquire audio segments based on the speech sensor and detect speech rhythm signals to obtain audio signals. A heat source region determination channel is used to evaluate the confidence level of candidate heat source regions based on the heat map morphology, gas concentration, and audio signal, and to obtain the heat source region.
2. The autonomous identification and positioning system for rescue robots based on industrial vision as described in claim 1, characterized in that, The human body recognition module includes: An image comparison unit is used to determine whether there are missing human body recognition features based on the comparison between the visible light image and the thermal imaging image; The key point extraction unit is used to extract local key points in the visible area if human recognition features are missing. The structural reconstruction unit is used to combine the local contours and temperature gradients in the thermal imaging image, and use a human pose completion network to reconstruct human structural features based on the local key points.
3. The autonomous identification and positioning system for rescue robots based on industrial vision as described in claim 1, characterized in that, The target annotation module includes: A spatial depth information calculation unit is used to calculate the spatial depth information of the target area where the rescue target is located through binocular vision. A three-dimensional structure correction unit is used to perform three-dimensional structure correction by fusing lidar scanning results to obtain target spatial information; The target location mapping unit is used to map the target spatial location to the robot coordinate map and record the rescue target location, occlusion level and detection time information of the rescue target; The target annotation unit is used to annotate target data on the robot coordinate map and upload it to the emergency command platform, and broadcast it to the collaborative rescue robot nodes for sharing.
4. The autonomous identification and positioning system for rescue robots based on industrial vision as described in claim 3, characterized in that, The target annotation unit includes: The path building channel is used to build an initial path in the robot coordinate map based on the annotation information; The path update channel is used to collect real-time target space information based on the rescue robot on the initial path. When a new obstacle is detected in the real-time target space information, the path reconstruction mechanism is triggered to update the initial path in the robot coordinate map to obtain the updated path.
5. The autonomous identification and positioning system for rescue robots based on industrial vision as described in claim 4, characterized in that, The target annotation unit further includes: An unreachable target marking channel is used to divide reachable and unreachable paths based on the updated path, mark the rescue robot as an unreachable target of the rescue target according to the unreachable path, and broadcast it to the collaborative rescue robot for execution.
6. The autonomous identification and positioning system for rescue robots based on industrial vision as described in claim 3, characterized in that, The target annotation unit further includes: The data upload channel is used to organize the target data into a standard format data packet and upload it to the emergency command platform via a wireless communication module. A data broadcast channel is used to broadcast the target data within the robot swarm network; The annotation update channel is used to set a periodic status update mechanism for the target data. When the target data changes, the annotation information is updated and the tracking module is triggered.
7. The autonomous identification and positioning system for rescue robots based on industrial vision as described in claim 1, characterized in that, The collaborative scanning module includes: The device activation unit is used to activate the industrial vision device, verify the connection status and sensor signal validity of the industrial vision device, and obtain a self-test result. The map building unit is used to construct an initial robot coordinate map through collaborative scanning by the industrial vision equipment and LiDAR; The map update unit is used to extract passable areas and obstacle boundaries based on the laser point cloud, set the navigation origin, update the initial robot coordinate map, and generate the robot coordinate map.
8. A method for autonomous identification and positioning of rescue robots based on industrial vision, characterized in that, The autonomous identification and positioning method for rescue robots based on industrial vision, as described in any one of claims 1 to 7, is executed by the such system. Perform initialization of the industrial vision equipment, and construct a robot coordinate map by coordinating scanning with the industrial vision equipment and LiDAR; Based on the robot coordinate map, combined with inertial measurement and audio sensing information, a unified temporal data fusion frame is established to obtain multimodal image data, wherein the multimodal image data includes thermal imaging images and visible light images; The thermal imaging image is mapped to the visible light image, and human body local recognition is performed based on the visible light image. Pose reasoning and completion are performed based on local key points to construct human body structural features. Based on the aforementioned human structural features, the rescue target is determined, and the target data of the rescue target is marked on the robot coordinate map to obtain the marking information; Based on the labeled information, a path is planned, an updated path is constructed, and a rescue robot is used to carry out the rescue.
Citation Information
Patent Citations
Multi-sensor fused search and rescue robot system
CN112109090A
Control method of mining robot, mining robot and storage medium
CN117140534A