Autonomous recognition and positioning system and method for rescue robot based on industrial vision

Through collaborative scanning of industrial vision and lidar, combined with inertial measurement and audio sensing information, multi-modal image fusion is carried out, which solves the problem of inaccurate positioning of rescue targets in complex environments, and achieves efficient rescue target identification and path planning, improving rescue efficiency.

CN120245004AActive Publication Date: 2025-07-04JIANGSU SANMING ZHIDA TECH CO LTD

Patent Information

Application Number
CN202510705463.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-04
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing rescue robots are difficult to accurately locate rescue targets in complex environments, resulting in low rescue efficiency.

Method used

The coordinated scanning of industrial vision and lidar is adopted, combined with inertial measurement and audio sensing information, and multimodal fusion of thermal imaging images and visible light images is carried out. Through local recognition of human bodies and posture reasoning, the complete structural characteristics of the rescue targets are completed and path planning is carried out.

Benefits of technology

Accurately positioning the rescue targets in harsh environments improves rescue efficiency and ensures that the rescue robots can reach the target area safely and quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120245004A_ABST
    Figure CN120245004A_ABST
Patent Text Reader

Abstract

The invention provides a rescue robot autonomous identification and positioning system and method based on industrial vision, and relates to the technical field of robots, and the system comprises a cooperative scanning module which is used for constructing a robot coordinate map; the data fusion module is used for establishing a data fusion frame with a unified time sequence to obtain multi-modal image data; the human body recognition module is used for mapping the thermal imaging image to a visible light image and executing human body local recognition to obtain human body structure features; the target marking module is used for determining a rescue target and marking the rescue target in the robot coordinate map; and the path planning module is used for performing path planning, constructing an updated path and performing rescue through the rescue robot. According to the method and the device, the technical problem of relatively low rescue efficiency caused by difficulty in accurately positioning the rescue target due to shielding and interference in a complex environment in the prior art can be solved, and the precision of positioning the rescue target is improved by combining industrial vision, so that the rescue efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics technology, and in particular, to an autonomous recognition and positioning system and method for rescue robots based on industrial vision. Background Art

[0002] Rescue robots are usually equipped with an autonomous recognition system for locating and identifying rescue targets in complex environments. Existing target recognition methods mainly rely on technologies based on RGB images or single sensors, which perform well in open environments. However, in complex environments, they are easily interfered by environmental factors such as smoke, dust, insufficient light, heat source interference, etc. The rescue target is partially blocked, and the position and posture of the trapped person cannot be correctly identified, resulting in inaccurate target recognition in a dynamic environment and affecting the timeliness and effectiveness of rescue operations.

[0003] In summary, there are technical problems in the prior art that it is difficult to accurately locate rescue targets due to occlusion and interference in complex environments, resulting in low rescue efficiency. Summary of the Invention

[0004] The purpose of this application is to provide an autonomous recognition and positioning system and method for rescue robots based on industrial vision, so as to solve the technical problem in the prior art that it is difficult to accurately locate rescue targets due to occlusion and interference in complex environments, resulting in low rescue efficiency.

[0005] In view of the above problems, this application provides an autonomous recognition and positioning system and method for rescue robots based on industrial vision.

[0006] In the first aspect, this application provides an autonomous recognition and positioning system for rescue robots based on industrial vision. Among them, the autonomous recognition and positioning system for rescue robots based on industrial vision includes: a collaborative scanning module for performing initialization of industrial vision devices, and constructing a robot coordinate map through collaborative scanning of the industrial vision devices and lidar; a data fusion module for establishing a unified time-series data fusion frame based on the robot coordinate map, combining inertial measurement and audio sensing information, and obtaining multi-modal image data, where the multi-modal image data includes thermal imaging images and visible light images; a human body recognition module for mapping the thermal imaging image to the visible light image, performing local human body recognition based on the visible light image, inferring and completing the posture based on local key points, and constructing a human body structure feature; a target annotation module for determining a rescue target based on the human body structure feature, annotating the target data of the rescue target in the robot coordinate map, and obtaining annotation information; a path planning module for performing path planning based on the annotation information, constructing and updating a path, and performing rescue through a rescue robot.

[0007] Optionally, a device activation unit is configured to activate the industrial vision device, verify the connection status and the sensing signal validity of the industrial vision device, and obtain a self-check result; a map construction unit is configured to construct an initial robot coordinate map by means of collaborative scanning of the industrial vision device and a lidar; a map update unit is configured to extract a passable area and an obstacle boundary according to lidar point clouds, set a navigation origin, and update the initial robot coordinate map to generate the robot coordinate map.

[0008] Optionally, a temperature analysis unit is configured to perform regional temperature analysis on the thermal imaging image, extract a heat source candidate area, and establish an area boundary; a confidence evaluation unit is configured to evaluate the confidence of the heat source candidate area to obtain a heat source area and map it to the visible light image.

[0009] Optionally, a heat distribution screening channel is configured to extract texture information, shape information, and edge information of the heat source candidate area, screen the human heat distribution to obtain a heat map form; a signal judgment channel is configured to judge a metabolic activity signal based on a gas concentration change obtained by a gas concentration sensor to obtain a gas concentration; an audio signal acquisition channel is configured to acquire an audio segment based on a voice sensor and detect a voice rhythm signal to obtain an audio signal; a heat source area determination channel is configured to evaluate the confidence of the heat source candidate area according to the heat map form, the gas concentration, and the audio signal to obtain the heat source area.

[0010] Optionally, an image comparison unit is configured to judge whether there is a lack of human recognition features based on the comparison between the visible light image and the thermal imaging image; a key point extraction unit is configured to extract local key points in the visible area if there is a lack of human recognition features; a structure reconstruction unit is configured to combine the local contour and the temperature gradient in the thermal imaging image, and use a human pose completion network to reconstruct human structure features based on the local key points.

[0011] Optionally, a spatial depth information calculation unit is configured to calculate the spatial depth information of the target area where the rescue target is located by means of binocular vision; a three-dimensional structure correction unit is configured to perform three-dimensional structure correction on the spatial depth information by fusing the lidar scanning result to obtain target spatial information; a target position mapping unit is configured to map the target spatial position to the robot coordinate map, and record the rescue target position, occlusion level, and detection time information of the rescue target; a target annotation unit is configured to annotate the target data on the robot coordinate map and upload it to an emergency command platform, and broadcast it to collaborative rescue robot nodes for sharing.

[0012] Optionally, a path construction channel is used to construct an initial path in the robot coordinate map according to the annotation information; a path update channel is used to collect real-time target space information based on the rescue robot on the initial path. When a new obstacle is detected in the real-time target space information, a path reconstruction mechanism is triggered to update the initial path in the robot coordinate map to obtain an updated path.

[0013] Optionally, an unreachable target marking channel is used to divide the reachable path and the unreachable path based on the updated path, mark the rescue robot as an unreachable target for the rescue target according to the unreachable path, and broadcast it to the cooperative rescue robot for execution.

[0014] Optionally, a data upload channel is used to organize the target data into a standard format data packet and upload it to the emergency command platform through a wireless communication module; a data broadcast channel is used to broadcast the target data in the robot cluster network; an annotation update channel is used to set a periodic status update mechanism for the target data. When the target data changes, the annotation information is updated, and the tracking module is triggered.

[0015] In a second aspect, the present application also provides a method for autonomous recognition and positioning of a rescue robot based on industrial vision. The method for autonomous recognition and positioning of a rescue robot based on industrial vision includes: performing initialization of industrial vision equipment, constructing a robot coordinate map by collaborative scanning of the industrial vision equipment and a lidar; according to the robot coordinate map, combining inertial measurement and audio sensing information to establish a data fusion frame with unified timing to obtain multi-modal image data, where the multi-modal image data includes a thermal imaging image and a visible light image; mapping the thermal imaging image to the visible light image, performing human local recognition based on the visible light image, performing pose inference and completion based on local key points, and constructing a human body structure feature; determining a rescue target based on the human body structure feature, annotating the target data of the rescue target in the robot coordinate map to obtain annotation information; performing path planning based on the annotation information, constructing an updated path, and performing rescue by the rescue robot.

[0016] One or more technical solutions provided in the present application have at least the following technical effects or advantages: Through a collaborative scanning module, it is used to perform the initialization of industrial vision devices, and through the collaborative scanning of the industrial vision devices and lidar, a robot coordinate map is constructed; a data fusion module, which is used to establish a unified time-series data fusion frame according to the robot coordinate map, combined with inertial measurement and audio sensing information, to obtain multimodal image data, where the multimodal image data includes thermal imaging images and visible light images; a human body recognition module, which is used to map the thermal imaging image to the visible light image, perform local human body recognition based on the visible light image, perform pose inference and completion based on local key points, and construct human body structure features; a target annotation module, which is used to determine a rescue target based on the human body structure features, annotate the target data of the rescue target in the robot coordinate map to obtain annotation information; a path planning module, which is used to perform path planning based on the annotation information, construct and update a path, and perform a rescue through a rescue robot. That is to say, through the collaborative scanning of industrial vision and lidar, combined with inertial measurement and audio sensing information, multimodal fusion of thermal imaging images and visible light images is carried out, and the rescue target can still be recognized in a harsh environment. Based on local human body recognition and pose inference, the complete structural features of the rescue target are complemented, the rescue target is determined and path planning is carried out, which improves the accuracy of rescue target positioning, ensures that the rescue robot can safely and quickly reach the target area, and thus improves the rescue efficiency.

[0017] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically gives the specific implementation manners of the present application. It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only exemplary, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0019] Figure 1 It is a schematic structural diagram of a rescue robot autonomous recognition and positioning system based on industrial vision of the present application; Figure 2 It is a schematic flow diagram of a rescue robot autonomous recognition and positioning method based on industrial vision of the present application.

[0020] Description of the attached drawing reference numerals: The collaborative scanning module 11, the data fusion module 12, the human body recognition module 13, the target annotation module 14, and the path planning module 15. Specific implementation manners

[0021] This application provides a rescue robot autonomous recognition and positioning system and method based on industrial vision, which solves the technical problem in the prior art that it is difficult to accurately locate rescue targets due to occlusion and interference in complex environments, resulting in low rescue efficiency. Through the collaborative scanning of industrial vision and lidar, combined with inertial measurement and audio sensing information, multi-modal fusion of thermal imaging images and visible light images is carried out. The rescue target can still be recognized in harsh environments. Based on human body local recognition and pose inference, the complete structural features of the rescue target are complemented, the rescue target is determined and path planning is carried out, improving the accuracy of rescue target positioning, ensuring that the rescue robot can reach the target area safely and quickly, and thus improving the rescue efficiency.

[0022] Next, the technical solutions in this application will be described clearly and completely with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the example embodiments described here. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application. Additionally, it should be noted that for the sake of description, only the parts related to this application are shown in the accompanying drawings rather than all of them.

[0023] Example 1, please refer to the attached Figure 1 , this application provides a rescue robot autonomous recognition and positioning system based on industrial vision. Among them, the rescue robot autonomous recognition and positioning system based on industrial vision is used to implement the steps of the rescue robot autonomous recognition and positioning method based on industrial vision. The rescue robot autonomous recognition and positioning system based on industrial vision includes: The collaborative scanning module 11 is used to perform the initialization of the industrial vision device, and construct a robot coordinate map through the collaborative scanning of the industrial vision device and the lidar.

[0024] Furthermore, the collaborative scanning module 11 in the rescue robot autonomous recognition and positioning system based on industrial vision is further used for: The device activation unit is used to activate the industrial vision device, verify the connection status and the validity of the sensing signal of the industrial vision device, and obtain a self-check result; the map construction unit is used to construct an initial robot coordinate map through the collaborative scanning of the industrial vision device and the lidar; the map update unit is used to extract the passable area and the obstacle boundary according to the laser point cloud, set the navigation origin, and update the initial robot coordinate map to generate the robot coordinate map.

[0025] Specifically, industrial vision devices are devices used to acquire image or video signals, usually including high-resolution cameras, infrared cameras, depth cameras, etc., which can capture visual information in the environment. Initialize the industrial vision device to ensure its normal operation. During the initialization process, the industrial vision device will perform self-checks to confirm whether its connection status is normal and whether the sensors are working effectively. The connection status includes whether the physical connection (such as cables and interfaces) is secure, whether the logical connection (such as the connection established between devices through software or protocols) is normal, and the communication protocol compatibility (such as whether the same communication protocol is used), etc.

[0026] At the same time, verify the validity of the sensing signals of the industrial vision device, including signal integrity, signal accuracy, signal stability, etc. Signal integrity is used to check whether the image data is damaged or lost. For example, verify the integrity of the data by calculating the checksum of the image and / or using error detection codes. Signal accuracy is used to verify whether the image and sensor data accurately reflect the environmental conditions, by comparing the data between different sensors or using known reference objects for verification. Signal stability is used to ensure that the sensing signal remains stable over a period of time without abnormal fluctuations or noise, by monitoring the historical data of the signal or using filtering techniques to evaluate the signal stability.

[0027] After the self-check of the industrial vision device is completed, obtain the self-check results to ensure that all functions are normal, including the connection status and whether the sensors are working properly, such as the device status (such as whether it is online or faulty) and the quality of the sensing signal (such as signal strength, stability).

[0028] After the industrial vision device is initialized, it works simultaneously with the lidar to collaboratively scan the surrounding environment and construct a preliminary robot coordinate map. The lidar emits laser light and receives echo data to generate three-dimensional point cloud data of the environment, providing geometric information such as the positions of obstacles and open areas in space, and providing spatial structure information. The industrial vision device provides visual information of the environment through image data. By combining the lidar point cloud data, it enhances the environmental perception ability of the rescue robot. The core role of collaborative scanning is to improve the accuracy and reliability of environmental perception by combining the information of the two. Using only the vision device or the lidar alone may be affected by environmental conditions (such as low light, occlusion, reflection, etc.), while through collaborative use, they can complement each other to obtain a more comprehensive understanding of the environment.

[0029] LiDAR and industrial vision devices generate different types of data: LiDAR generates three-dimensional point cloud data, while industrial vision devices generate two-dimensional images or depth images. After fusing these data, a multi-dimensional information system is constructed. LiDAR can provide spatial structure information (such as distance, position), while vision devices provide object recognition information (such as the shape, color, texture, etc. of objects). After collaborative scanning, the point cloud data of LiDAR is processed to filter out noise points and extract valid point clouds; at the same time, the information in the visual image is subjected to object detection and image segmentation to identify important objects in the environment. Through coordinate alignment, the data of both are fused to generate an initial coordinate map with spatial and visual information.

[0030] Using the data scanned by LiDAR and the image information provided by industrial vision devices, an initial robot coordinate map including obstacles, passable areas, target objects, etc. is generated. The point cloud data of LiDAR is used to accurately describe the structure of the three-dimensional space, while vision devices provide the shape information of objects to help rescue robots identify specific objects in the environment (such as trees, walls, tables, people, etc.).

[0031] In the point cloud data collected by LiDAR, the coordinates of each point represent the specific position of a certain point in the environment. By performing noise filtering and point cloud downsampling on the LiDAR points, useless data is removed. Using clustering algorithms (such as Euclidean clustering), the point cloud data is divided into different object parts, and the positions of obstacles are identified. The dense areas in the point cloud data usually represent obstacles, and LiDAR can identify these areas according to the density and geometric shape of the point cloud, separating the obstacle boundaries from the surrounding open areas.

[0032] A navigation origin is set as a reference point, usually the coordinate point of the current position of the rescue robot. For example, through the data of LiDAR and other sensors (such as IMU or GPS), the rescue robot can determine its current position and set a navigation origin. The processing results of the LiDAR points are fused with the original map data to dynamically update the initial robot coordinate map, generating a more accurate robot coordinate map that includes the detailed layout of the environment, containing all information such as obstacles, passable areas, and navigation origin. The rescue robot can perform path planning, obstacle avoidance, and navigation based on the rescue robot coordinate map. Through the collaborative work of industrial vision devices and LiDAR, the rescue robot can obtain rich visual and spatial information in a complex environment, thereby achieving precise positioning and environmental perception. The point cloud data generated by LiDAR helps the rescue robot identify passable areas and obstacles, providing a basis for planning the robot's driving route and avoiding collisions of the rescue robot in a complex environment.

[0033] The data fusion module 12 is used to establish a data fusion frame with a unified time sequence according to the robot coordinate map, combining inertial measurement and audio sensor information, and obtain multimodal image data, wherein the multimodal image data includes a thermal imaging image and a visible light image.

[0034] Specifically, inertial measurement and audio sensing information of the target rescue area are obtained through inertial measurement units and audio sensors. IMU is a sensor used to measure the acceleration, angular velocity and other motion parameters of an object. It is used to provide real-time motion information of the rescue target in the target rescue area, assist in positioning and attitude estimation, and is particularly important in environments with insufficient GPS signals. Audio sensors (such as microphone arrays) are used to collect sound data in the target rescue area environment, can identify and locate sound sources, and can detect the shouts or knocks of trapped people.

[0035] The robot coordinate map, inertial measurement and audio sensor information are fused to integrate the information of different sensors into a unified framework. By synchronizing and adjusting the data of different sensors to ensure that they are consistent in time, a multimodal image data fusion frame is established. The thermal imaging image is captured by the thermal imaging sensor to show the distribution of temperature in the target area. In rescue missions, thermal imaging images can clearly show the location of trapped people, especially in low light or smoke environments, which helps to locate the heat source (trapped person) even when the line of sight is poor. Visible light images are collected by the camera to provide more detailed information about the target area and confirm whether the target in the thermal imaging image is a trapped person. For example, a hot spot is detected in the thermal imaging image, and the visible light image is used to further confirm the external features of the target (such as clothing color, posture, etc.) to ensure the accuracy of the target.

[0036] Combine thermal imaging images with visible light images to further confirm the location and identity of the trapped person. For example, in a dark environment, thermal imaging images may reveal an area with higher temperature, while visible light images can be supplemented when confirming the specific shape, color and other features of the area to ensure the accuracy of recognition. Through the positioning of the audio sensor and the motion information provided by the IMU, the robot can accurately identify and track the specific location of the trapped person. When the IMU detects slight changes in the movement of the trapped person, the audio sensor may capture the trapped person's cry for help or knocking sound, helping the robot further confirm the target location and reduce the possibility of misidentification.

[0037] By fusing data from multiple sensors (such as IMU, audio sensors, thermal imaging images, visible light images, etc.), the robot can accurately identify the location of trapped people and effectively locate them in complex environments. The robot can maintain high recognition accuracy even in the presence of occlusion, smoke or low light.

[0038] The human body recognition module 13 is configured to map the thermal imaging image to the visible light image, perform local human body recognition based on the visible light image, perform pose inference and completion based on local key points, and construct human body structure features.

[0039] Furthermore, the human body recognition module 13 in the rescue robot autonomous recognition and positioning system based on industrial vision is further configured to: A temperature analysis unit for performing regional temperature analysis on the thermal imaging image, extracting heat source candidate regions, and establishing regional boundaries; a confidence evaluation unit for evaluating the confidence of the heat source candidate regions to obtain heat source regions and mapping them to the visible light image.

[0040] A heat distribution screening channel for extracting texture information, shape information, and edge information of the heat source candidate regions, screening the human body heat distribution to obtain a heat map form; a signal judgment channel for judging metabolic activity signals based on the gas concentration changes obtained by a gas concentration sensor to obtain the gas concentration; an audio signal acquisition channel for acquiring audio segments based on a voice sensor and detecting voice rhythm signals to obtain audio signals; a heat source region determination channel for evaluating the confidence of the heat source candidate regions according to the heat map form, gas concentration, and audio signals to obtain the heat source regions.

[0041] Specifically, perform regional temperature analysis on the thermal imaging image, analyze the temperature distribution of different regions in the thermal imaging image, and identify possible heat source regions, that is, heat source candidate regions. Usually, it is screened by setting a temperature threshold, and the regions with temperatures exceeding the threshold will be regarded as heat source candidate regions. The temperature threshold is usually set according to the actual environmental temperature and human body temperature. The heat source candidate regions are the possible heat source regions screened out through the set threshold after regional temperature analysis and may include trapped persons. For example, if the set temperature threshold is 35 °C, then the regions in the thermal imaging image with temperatures exceeding 35 °C will be marked as heat source candidate regions, which may include the human body (usually with a body temperature higher than the surrounding environment), mechanical equipment, or other heat-emitting objects.

[0042] After determining the heat source candidate area, establish the boundary of the heat source candidate area, and determine the external contour of the heat source candidate area through edge detection. The edge detection algorithm identifies the object boundary by searching for areas with large pixel intensity changes in the image. In a thermal imaging image, the temperature difference between the object and the background causes significant changes in the pixel values of different regions. By detecting the regions with large gray-scale changes in the image to find the edges, a binary image is usually generated, where the edge part is 1 (white) and the non-edge part is 0 (black). In the edge detection results, there may be broken edges or small noise points. To ensure the integrity and accuracy of the extracted heat source boundary, morphological operations (such as dilation and erosion) are usually used to smooth and connect the edges. The dilation operation can connect broken edges, and the erosion operation removes small noise points. Once the edge detection is completed, the edge image will represent the external boundary of the heat source candidate area.

[0043] After identifying the heat source candidate area, extract the texture information, shape information, and edge information of the heat source candidate area. Texture information refers to the change pattern of the light reflected from the surface in the image, usually manifested as the characteristics of the object surface structure, to distinguish the surface types of different objects or regions. By calculating the spatial relationship of pixel gray levels, texture features of the thermal imaging image are extracted, such as contrast, homogeneity, entropy, etc. Shape information refers to the geometric features of the image area, such as the boundary shape, aspect ratio of the area, roundness, etc., which are used to distinguish different types of targets (e.g., human body, equipment, etc.) and further identify whether the target is a trapped person. By calculating the aspect ratio of the heat source candidate area, it is judged whether the area conforms to the conventional proportion of the human body (e.g., the aspect ratio of the human body is generally within a certain range). By analyzing the external contour of the heat source area, the shape information of the area is further extracted, such as whether it presents the common contour of the human body. Edge information refers to the parts in the image where the brightness or color changes violently, which helps to determine the contour and shape of the object and reflects the external boundary of the heat source area. Through a multi-stage edge detection process, the edge information of the heat source candidate area is accurately extracted.

[0044] Analyze the texture information, shape information, and edge information, and filter out the areas that conform to the human body heat distribution law. The characteristics of human body heat distribution are usually that the head area is hotter, the heat distribution of the torso and limbs is relatively uniform, but the temperature of the limbs is lower. The human body usually presents certain morphological characteristics, such as long strip shape, close to oval shape, or with a specific upper and lower ratio. Based on this, the human body heat distribution is filtered to generate a heat map morphology. The heat map morphology shows the temperature intensity at different positions in the candidate area, thus helping to confirm the type and position of the target.

[0045] Generate a heat map with different colors based on the temperature values within the area. Areas with higher temperatures are usually represented in red, while lower temperatures are represented in blue. In the heat map, the temperature values of the areas are corresponded to the heat levels to help identify whether the high-temperature areas conform to the human heat distribution characteristics. By analyzing the texture, shape, and edge information of the heat source candidate areas, a morphological map showing the heat distribution of the heat source areas is obtained, which demonstrates the temperature distribution at different positions within the area and helps identify and confirm the nature and location of the target.

[0046] Monitor the gas concentration in the target area through gas concentration sensors, such as carbon dioxide sensors, carbon monoxide sensors, oxygen sensors, etc. Each sensor detects the concentration of a specific gas through different working principles (such as infrared absorption, semiconductor reaction, etc.). When the human body is carrying out metabolic activities, especially during the breathing process, carbon dioxide is emitted. When the human body is in an active state, the metabolic activities increase, and the carbon dioxide emission also increases. By monitoring the change in the carbon dioxide concentration in the air, it can be inferred whether the trapped person is in an active state or in an emergency. For example, assume that a trapped person's metabolic activities increase during a fire, resulting in a significant increase in carbon dioxide emissions. At this time, the gas concentration sensor (such as a carbon dioxide sensor) detects an increase in the carbon dioxide concentration in the area, thereby judging the metabolic activity signal of the trapped person.

[0047] When the gas concentration sensor is detecting the gas concentration in real time, it can continuously record the change in the concentration of a specific gas (such as carbon dioxide). In the rescue environment, combined with the data of thermal imaging images and other sensors (such as audio sensors, temperature sensors, etc.), the gas concentration data can provide other information for target positioning. When the gas concentration sensor detects an increase in the carbon dioxide concentration, this signal can be combined with the data of other sensors (such as the heat source candidate areas in the thermal imaging image) to improve the accuracy of identifying trapped persons. By analyzing the change in the gas concentration, the signal of the metabolic activity is judged. For example, when the carbon dioxide concentration continues to increase, such as when the carbon dioxide concentration rapidly increases from 500 ppm to 1000 ppm, it indicates that there is continuous metabolic activity within a certain area, usually meaning that the trapped person may still be in an active state. If the carbon dioxide concentration reaches 1500 ppm to 2000 ppm, it may indicate that the trapped person in this area has felt sleepy, dizzy, etc.; if the carbon dioxide concentration reaches 5000 ppm, the human body functions are severely disordered, resulting in loss of consciousness.

[0048] By analyzing the change in gas concentration, the signal of metabolic activity is determined. Real-time monitoring of the change in gas concentration through a gas concentration sensor and analyzing the metabolic activity signal of the trapped person helps to locate and identify the trapped person in a timely manner, and at the same time obtain the gas concentration, that is, the volume or mass ratio of a certain gas in a specific space, such as the concentration of gases like carbon dioxide, oxygen, etc. For example, in a fire scene, some example data of the monitored gas concentration change is shown in Table 1: Table 1 Partial example data table of gas concentration change

[0049] The voice sensor captures the sound signal in the target rescue area to obtain an audio clip, that is, the sound data collected from the voice sensor within a certain period of time, including voice, ambient sound, shouts, etc. To ensure high-quality audio signal analysis, the audio clip is first preprocessed, including noise removal, noise reduction, and enhancement processing of the audio signal, removing irrelevant background noise and only retaining the audio signal related to the trapped person. Using audio signal processing techniques (such as speech signal processing, Fourier transform, etc.) to analyze the rhythm characteristics of the audio signal, detecting pauses, pitch changes, and speech rate in the speech, and identifying whether there is a rapid speech rhythm or a long-term high audio volume, which helps to judge the emotional state and urgency of the speech. For example, when the audio signal shows frequent rapid shouts, continuous increase in speech rate and volume, it indicates that the trapped person is in an extremely tense state of calling for help.

[0050] The audio clip is input into an existing audio signal recognizer. According to the changes in aspects such as pitch, speech rate, pause, and tone, the type of speech (such as a cry for help, a shout, etc.) and the emotion it generates (such as urgent, panicked, etc.) are judged. Combining the data of other sensors (such as thermal imaging, gas concentration sensor, etc.), the reliability of the audio signal is further confirmed to ensure the accuracy of the rescue decision. For example, a faint groan is recognized in the audio clip from 6 to 10 seconds, indicating the presence of a trapped person with a poor physiological state; from 15 to 20 seconds of the audio, a loud shout for help is recognized, and the rhythm analysis result shows high frequency, rapid, and emotional excitement, indicating that the trapped person is calling for help urgently at this time and may be in a state of serious injury or drowsiness.

[0051] Based on the heat map morphology, gas concentration, and audio signal, confidence assessment is performed on the heat source candidate areas. First, the heat map morphology, gas concentration, and audio signal are standardized to eliminate the influence of dimensions. Then, according to the weights corresponding to the heat map morphology, gas concentration, and audio signal, usually the weights of the audio signal and heat map morphology are relatively large, that is, the weight of the heat map morphology is 0.4, the weight of the audio signal is 0.4, and the weight of the gas concentration is 0.3, the confidence score of the heat source candidate area is calculated, indicating whether there are trapped persons in all the heat source candidate areas. According to the confidence score, the areas with higher scores are selected as the heat source areas, that is, there are trapped persons in these areas and emergency rescue is needed. For example, after standardizing the heat map morphology, gas concentration, and audio signal of a certain heat source candidate area, we get: the value after standardizing the heat map morphology is 0.8, the value after standardizing the audio signal is 0.9, the value after standardizing the gas concentration is 0.6, and the confidence score is 0.86, indicating that this heat source candidate area is very likely to be the location of the trapped person, so emergency rescue is needed. The confidence score is usually between 0 and 1, and a value close to 1 indicates that there are very likely trapped persons in this area.

[0052] Map the heat source areas to the visible light image to help the rescue robot determine whether there are corresponding trapped persons in these heat source areas. For example, the thermal imaging image may show a dangerous area, and through the visible light image, the robot can confirm whether there is really a human body in this area, further eliminating false recognition. By aligning the heat source positions of the thermal imaging image with the visible light image, the positions of the trapped persons can be marked in the visible light image. By comprehensively using the data of the thermal imaging image, gas concentration sensor, and voice sensor, confidence assessment is performed on the heat source candidate areas under multi-modal information fusion, and finally the areas most likely to have trapped persons are determined, which not only improves the accuracy of the rescue robot but also effectively reduces misjudgment caused by occlusion and interference.

[0053] Furthermore, the human body recognition module 13 in the rescue robot autonomous recognition and positioning system based on industrial vision is also used for: The image comparison unit is used to judge whether there is a lack of human body recognition features based on the comparison of the visible light image and the thermal imaging image; the key point extraction unit is used to extract local key points in the visible area if there is a lack of human body recognition features; the structure reconstruction unit is used to combine the local contour and temperature gradient in the thermal imaging image, and use the human body pose completion network to reconstruct the human body structure features based on the local key points.

[0054] Specifically, the visible light image and the thermal imaging image are compared and judged. The visible light image provides information such as the color and shape of an object, while the thermal imaging image provides the temperature information of the object. By comparing these two, it is possible to effectively identify whether there are missing human body features. For example, due to obstacles or other factors in the environment, some parts of the human body may not be accurately displayed in the thermal imaging image, or may be blocked in the visible light image. Judging whether there are missing human body recognition features refers to the situation where the human body contours in the visible light image and the thermal imaging image do not completely overlap or are incompletely displayed. For instance, in the thermal imaging image, some parts (such as hands or feet) may not be clearly shown due to distance, angle, or occlusion, resulting in missing recognition of the overall human body contour.

[0055] If there are missing body parts or local heat sources in the thermal imaging image, the remaining features (such as the contours of hands and feet, the edges of clothing, etc.) are usually extracted from the visible light image to supplement the missing information. For example, if the hands or feet of a person are not clearly shown in the thermal imaging image but are still visible in the visible light image, they can be completed through local key point extraction. Local key points refer to the feature points of certain key parts of the human body, such as hands, feet, head, shoulders, knees, etc. By analyzing the visible light image and extracting these local key points, even if some human body features are not displayed in the thermal imaging image, the information in the visible light image can be used for supplementation.

[0056] In the thermal imaging image, the heat distribution of the human body often presents a specific heat gradient (such as the temperature difference between the core temperature of the human body and the limbs). By analyzing the contours of local heat sources and the temperature gradient in the heat map, the structural features of the human body can be further understood. Especially when some human body features in the thermal imaging image are blocked or blurred, the temperature gradient and local contours can provide important supplementary information. The local contours in the thermal imaging image refer to the heat source boundaries of different parts of the human body, such as the edges of the torso, the contours of the arms and legs, etc., which help to identify the overall shape of the human body. Especially when part of the human body is blocked, the basic structure of the human body can still be inferred through the temperature gradient. The temperature gradient refers to the rate or intensity of temperature change in the thermal imaging image. The temperature differences between different parts of the human body can be presented through the temperature gradient. For example, the torso is usually warmer than the arms and legs.

[0057] The human pose completion network is a deep learning-based model used to infer and complete human poses. It can reconstruct the complete human structure from local joint information, contour information, and temperature gradients. The human pose completion network relies on local key point information extracted from visible light images or thermal imaging images, usually referring to joint parts (such as shoulders, elbows, knees, etc.), which provides the initial structural information of the human body. In the case of partial information loss, the human pose completion network can infer the missing parts based on the existing local key points and temperature gradient information.

[0058] Through multi-modal data of thermal imaging images and visible light images, the human pose completion network can combine local key point and temperature gradient data to infer the spatial positions of various parts of the human body (such as hands, feet, head, etc.), helping to reconstruct the complete human structure. Especially when local data is missing, the completion network can make effective inferences and supplements based on the existing information. For example, assume that the thermal imaging image shows that the torso part of the trapped person is relatively obvious, while the limbs are blocked by obstacles. Through the human pose completion network, by combining the temperature information of the torso with the visible light image information of the hands and feet, the completion network can infer the poses and positions of the hands and feet, and finally reconstruct the complete body structure.

[0059] Training the human pose completion network usually requires a human pose dataset containing a large amount of labeled data, which includes human poses in multiple different scenarios, especially samples with partial key points missing or occluded. Since a large number of samples are required during the training process, data augmentation is a common practice. Data augmentation can generate more training samples through methods such as rotation, scaling, translation, and cropping, thereby improving the generalization ability of the model. To train the network, each image in the dataset needs to have the labeled key point positions, including parts such as the head, shoulders, elbows, knees, and ankles. For each key point, the dataset will record its 2D or 3D coordinates (in different coordinate systems).

[0060] The human body pose completion network generally consists of the following main parts: an encoder-decoder. The encoder is used to extract features from the input image (thermal imaging image or visible light image), and the decoder reconstructs the missing key points based on the features extracted by the encoder. The graph convolutional network is used to model the relationships between different key points as a graph structure, enabling the network to perform pose completion according to the spatial constraints between key points. During training, a loss function is used to evaluate the difference between the pose output by the network and the true annotated pose, such as the mean square error, which calculates the sum of the squared Euclidean distances between the coordinates of the key points output by the network and the true annotated coordinates. To make the network output conform to the structure of the human skeleton, a spatial constraint loss is usually added. This loss function can restrict the relative position relationships between key points and prevent the output of poses that do not conform to the conventional human structure. During training, the input data (such as the thermal imaging image of the occluded part or partial key points) passes through the network for forward propagation. Each layer of the network calculates the predicted coordinates of the human key points based on the input image. According to the calculation result of the loss function, the network performs backpropagation. Through backpropagation, the parameters of the network (such as the weights of the convolutional kernels) are adjusted according to the error to minimize the loss function, thereby optimizing the performance of the network. During training, the dataset is usually divided into multiple batches, each batch containing several samples. The network improves the training efficiency and stability through batch data training. The optimal hyperparameter configuration is selected through methods such as cross-validation or grid search. During the training process, a validation set is often used to evaluate the model to ensure the generalization ability of the model. The validation set is a part of the data separated from the training data and is used to check whether the model is overfitting. After the model training is completed, the network can be applied to the actual environment. By inputting the human image of the occluded part, it outputs the complete human pose.

[0061] By combining the temperature gradient information of the thermal imaging image and the local key point data in the visible light image, the human body pose completion network is used to complete the missing human body parts, solving the human body recognition problem caused by environmental occlusion or incomplete thermal imaging image data. This not only improves the recognition accuracy of the rescue robot, enhances its adaptability in complex environments, but also greatly improves the rescue efficiency, ensuring that the robot can quickly identify and locate the trapped person under extreme conditions.

[0062] The target annotation module 14 is used to determine the rescue target based on the human body structure features and annotate the target data of the rescue target in the robot coordinate map to obtain annotation information.

[0063] Furthermore, the target annotation module 14 in the industrial vision-based rescue robot autonomous recognition and positioning system is also used for: A spatial depth information calculation unit for calculating the spatial depth information of the target area where the rescue target is located through binocular vision; a three-dimensional structure correction unit for fusing the lidar scan results to perform three-dimensional structure correction on the spatial depth information to obtain target spatial information; a target position mapping unit for mapping the target spatial position to the robot coordinate map and recording the rescue target position, occlusion level, and detection time information of the rescue target; a target annotation unit for annotating the target data on the robot coordinate map and uploading it to the emergency command platform and broadcasting it to the collaborative rescue robot nodes for sharing.

[0064] Specifically, binocular vision uses two cameras to capture the same scene from different angles and calculates the spatial depth information of the target by calculating the differences between the two images. The basic principle of binocular vision is to compare two images with different perspectives through stereo matching technology, calculate the disparity of each pixel point, and then deduce the spatial coordinates of this point. The two cameras capture the same scene and use the disparity (the displacement difference of the same point in the two perspectives) to calculate the depth information. By matching the feature points (such as edges, corners, etc.) on the images and using triangulation to obtain the depth of each feature point (i.e., the distance from the camera). For example, assume the baseline of the binocular camera is 10 cm, the disparity of a certain feature point in the two images is 2 pixels, and the focal length of the camera is known to be 5 cm. Then, through the calculation formula D = , where D is the target distance, f is the focal length of the camera, B is the baseline distance between the two cameras, and d is the disparity. Through calculation, the target depth information is obtained as 25 cm.

[0065] The lidar scans the surrounding environment through laser beams, calculates the time from the emission to the reflection of the laser signal (or measures the intensity of the reflected light), and then obtains the distance information. The lidar continuously emits laser beams to obtain the reflection information of the environment and forms dense three-dimensional point cloud data. Each point represents the coordinate position of a reflected laser signal, usually expressed as (x, y, z) coordinates. Through the lidar scan results, three-dimensional structure correction is performed on the spatial depth information, combining the point cloud data obtained by the lidar with the depth information calculated by binocular vision to correct and rectify possible errors or inconsistencies. By combining the advantages of both, a more accurate target position can be obtained.

[0066] The lidar scans the environment to generate a large amount of point cloud data. The coordinates of each point contain the three-dimensional position of the target object. Since the coordinate systems of binocular vision and lidar are different, external parameter calibration is required. Through the calibration process, the rotation and translation relationships between the two are found to ensure that the data of the two can be accurately aligned. Algorithms such as Kalman filtering and complementary filtering are used to perform weighted fusion of the data from binocular vision and lidar, so as to obtain more accurate spatial information. For example, the lidar gives the point cloud data of a certain target as (2.3, 1.4, 0.5), while the distance calculated by binocular vision is 4.0 meters. Through data fusion, combining the accuracy and weight of the two, the accurate position of the target is finally obtained as (2.4, 1.3, 0.6).

[0067] Map the target spatial position to the robot coordinate map, that is, convert it into the target spatial position from the robot's perspective, which is convenient for planning the robot's route. Record the position of the target in the robot coordinate map, and mark the occlusion situation and detection time of the target. The occlusion situation can be judged by the positions of other objects in the environment, and the detection time is obtained through the timestamp of the sensor. The occlusion level indicates the degree to which the target is occluded by obstacles or other objects, such as complete occlusion, partial occlusion, and no occlusion. Recording the specific time when the target information is detected helps to track the target state in a dynamic environment. For example, assuming the target spatial position is (3.2, 4.1) and the current position of the robot is (1.0, 2.0), the position of the target in the robot coordinate system is calculated as (2.2, 2.1) through coordinate transformation. At the same time, record the occlusion level of the target (such as partial occlusion) and the detection time (such as 1:30 PM).

[0068] Mark the target information (position, occlusion level, detection time, etc.) on the robot coordinate map to help the command personnel view the specific positions of each target and determine the priority rescue order. The marked target data is uploaded to the emergency command platform in real time through wireless communication, and the personnel in the command center can monitor the status of the target at any time and arrange the rescue.

[0069] Furthermore, the target annotation module 14 in the rescue robot autonomous recognition and positioning system based on industrial vision is further used for: The unreachable target marking channel is used to divide the reachable path and the unreachable path based on the updated path, mark the rescue robot as an unreachable target for the rescue target according to the unreachable path, and broadcast it to the collaborative rescue robots for execution.

[0070] Specifically, the update path is divided into reachable paths and unreachable paths. A reachable path is a path that the robot can pass through smoothly without obstacles or other factors hindering the robot's movement; an unreachable path is a path that the robot cannot pass through, and the robot may not be able to move smoothly due to factors such as obstacles, narrow spaces, and environmental changes. When the robot detects that the path where the target location is located is unreachable, it marks the target as an unreachable target and broadcasts this information to other cooperative rescue robots in the same task, including the location information of the target, the obstacle situation, the reason for unreachability, and the real-time path data. After receiving the information of the unreachable target, the cooperative rescue robot may choose a different path according to its own sensor data and map information and try to approach the target area from another angle or different route. By dividing the path into reachable paths and unreachable paths, it is possible to optimize the robot path planning in real time and dynamically in a complex environment, avoid the robot entering an impassable area, and improve the rescue efficiency and accuracy.

[0071] Further, the target annotation module 14 in the industrial vision-based rescue robot autonomous recognition and positioning system is further configured to: A data upload channel is used to organize the target data into a standard format data packet and upload it to the emergency command platform through a wireless communication module; a data broadcast channel is used to broadcast the target data in the robot cluster network; an annotation update channel is used to set a periodic status update mechanism for the target data. When the target data changes, update the annotation information and trigger the tracking module.

[0072] Specifically, to ensure the unity and compatibility of data, the target data is organized into a standard format data packet, that is, different types of information are encoded and encapsulated according to a fixed structure, so that the data can be recognized and processed by different devices and platforms. The formatted data is encapsulated into a data packet according to a predetermined protocol, and necessary verification information and timestamps are added to ensure the integrity and accuracy of the data packet. Target data is various information collected during the rescue process about targets (such as trapped persons, obstacles, etc.), including the spatial location, status, detection time, occlusion level, etc. of the target.

[0073] Through the wireless communication module (such as Wi-Fi, 4G / 5G or a dedicated wireless communication protocol) carried by the robot, the target data packet is sent to the emergency command platform. On the platform, rescue commanders can view the uploaded target data in real time for task assignment, path planning, and decision support.

[0074] Broadcast the target data uploaded to the emergency intelligent platform in the robot cluster network. Other robots within the robot cluster can receive this data through the cluster network, avoiding duplicate identification of rescue tasks. Broadcasting data helps to achieve multi-robot collaborative work, sharing key information such as target locations and statuses, so that multiple robots can more effectively collaborate on rescue tasks.

[0075] Periodically check and update the target data by setting a timer (such as every minute or every 5 minutes), regularly detect changes in the target (such as location, status, etc.), and update the data. When the status of the target changes (such as changes in the target location, changes in the status of trapped persons, etc.), the robot will update the annotation information of the target, including the latest location of the target, the occlusion level of the target, etc. If the target changes, or there are significant changes in the data annotation information of the target, trigger the tracking module to continuously track the target to ensure that the dynamic information of the target can be updated in a timely manner and accurately tracked. By regularly updating the target data and triggering the tracking module, ensure the accuracy and timeliness of the target information. If the status of the target changes, be able to adjust the rescue strategy in a timely manner to ensure that the target is accurately identified and rescued.

[0076] The path planning module 15 is used to perform path planning based on the annotation information, construct an updated path, and execute the rescue through the rescue robot.

[0077] The path construction channel is used to construct an initial path in the robot coordinate map according to the annotation information; the path update channel is used to collect real-time target space information on the initial path based on the rescue robot. When a new obstacle is detected in the real-time target space information, trigger the path reconstruction mechanism to update the initial path in the robot coordinate map to obtain an updated path.

[0078] Specifically, according to the annotation information in the robot coordinate map, an initial path is constructed in the robot coordinate map using a path planning algorithm. The annotation information provides the positions of the target location and obstacles, and the robot coordinate map provides the coordinate system framework. The initial path is the preliminary planned route for the robot to move from the starting point (the position where the robot is located) to the target point, taking into account factors such as the positions of obstacles, the moving direction of the robot, and speed. According to the specific application scenario, a suitable path planning algorithm is selected, such as the A* algorithm, which generates a path based on the starting point, target point, obstacle information, and the traversable area of the environment. For example, assume the starting point is (1.0, 2.0), the target point is (5.0, 6.0), and the obstacles are located at (3.0, 3.0) and (4.0, 4.0) respectively. The step size is set to 1 unit, indicating that the robot moves 1 unit of horizontal or vertical distance each time. The distance from the current node to the target node is estimated through a heuristic function (such as the Manhattan distance). The A* algorithm selects the next optimal node by evaluating the cost of each adjacent node (composed of the heuristic function h(x, y) and the path length g(x, y)), and the resulting cost map is as follows: The cost f = g + h = 1 + 7 = 8 for (2.0, 2.0); the cost f = g + h = 2 + 6 = 8 for (3.0, 2.0); the cost f = g + h = 3 + 5 = 8 for (5.0, 2.0). The node with the minimum cost is selected as the next moving node, and this process is repeated until the target point is reached. The initial path obtained through path planning is (1.0, 2.0) → (2.0, 2.0) → (3.0, 3.0) → (4.0, 4.0) → (5.0, 6.0). The rescue robot avoids the obstacles and it is the current optimal path.

[0079] When the rescue robot moves along the initial path, it continuously collects real-time data of the environment, especially the depth information of the target area, the positions of obstacles, etc. Using devices such as lidar and vision sensors, it real-time monitors for new obstacles in the target space. If a new obstacle is detected, the rescue robot will trigger path update through a path reconstruction mechanism. The position of the newly added obstacle is marked in the robot coordinate map, and a path to avoid this obstacle is recalculated through the aforementioned path planning algorithm, updating the path in the robot coordinate map. The updated path is based on the newly added obstacle information and real-time target space data, recalculating and adjusting the moving path of the robot to ensure that the robot can avoid the new obstacle and move forward safely.

[0080] The robot starts to execute navigation according to the updated path. During the path execution, the distance between the robot and the target is detected in real time to ensure that the robot can successfully avoid obstacles and reach the target position. During the path execution, the rescue robot continuously collects environmental data, including the position of the target, the position of obstacles, etc., to ensure that it travels along the correct route. If the path changes (such as a change in the target position, the appearance of dynamic obstacles, etc.), the robot will continue to update the path and re-plan. After reaching the target, the robot starts to perform rescue tasks, including clearing obstacles, delivering supplies, locating and rescuing trapped people, etc. During the execution process, the robot will make real-time adjustments and executions according to the task requirements.

[0081] In summary, the autonomous recognition and positioning system of the rescue robot based on industrial vision provided by this application has the following technical effects: Through the collaborative scanning module, which is used to execute the initialization of the industrial vision device, and through the collaborative scanning of the industrial vision device and the lidar, a robot coordinate map is constructed; the data fusion module is used to establish a unified time-series data fusion frame based on the robot coordinate map, combined with inertial measurement and audio sensing information, to obtain multi-modal image data, where the multi-modal image data includes a thermal imaging image and a visible light image; the human body recognition module is used to map the thermal imaging image to the visible light image, perform local human body recognition based on the visible light image, perform pose inference and completion based on local key points, and construct a human body structure feature; the target annotation module is used to determine the rescue target based on the human body structure feature, annotate the target data of the rescue target in the robot coordinate map to obtain annotation information; the path planning module is used to perform path planning based on the annotation information, construct an updated path, and execute the rescue through the rescue robot. That is to say, through the collaborative scanning of industrial vision and lidar, combined with inertial measurement and audio sensing information, multi-modal fusion of thermal imaging images and visible light images is carried out. The rescue target can still be recognized in harsh environments. Based on local human body recognition and pose inference, the complete structural features of the rescue target are complemented, the rescue target is determined and path planning is carried out, which improves the accuracy of rescue target positioning, ensures that the rescue robot safely and quickly reaches the target area, and thus improves the rescue efficiency.

[0082] Embodiment 2, based on the same inventive concept as the autonomous recognition and positioning system of the rescue robot based on industrial vision in the foregoing Embodiment 1, this application also provides an autonomous recognition and positioning method for the rescue robot based on industrial vision. Please refer to the appendix Figure 2 The autonomous recognition and positioning method for the rescue robot based on industrial vision includes: S100: Initialize the industrial vision device, and construct a robot coordinate map through collaborative scanning of the industrial vision device and the lidar; S200: Based on the robot coordinate map, combine inertial measurement and audio sensing information to establish a data fusion frame with a unified time sequence, and obtain multimodal image data, where the multimodal image data includes a thermal imaging image and a visible light image; S300: Map the thermal imaging image to the visible light image, perform human local recognition based on the visible light image, perform pose inference and completion based on local key points, and construct a human body structure feature; S400: Determine a rescue target based on the human body structure feature, mark the target data of the rescue target on the robot coordinate map to obtain marking information; S500: Perform path planning based on the marking information, construct and update a path, and execute the rescue through a rescue robot.

[0083] Further, the step of initializing the industrial vision device and constructing a robot coordinate map through collaborative scanning of the industrial vision device and the lidar includes: activating the industrial vision device, verifying the connection status and sensing signal validity of the industrial vision device to obtain a self-check result; constructing an initial robot coordinate map through collaborative scanning of the industrial vision device and the lidar; extracting a passable area and an obstacle boundary from the laser point cloud, setting a navigation origin, and updating the initial robot coordinate map to generate the robot coordinate map.

[0084] Further, the step of mapping the thermal imaging image to the visible light image includes: performing regional temperature analysis on the thermal imaging image, extracting a heat source candidate area, and establishing a regional boundary; performing confidence evaluation on the heat source candidate area to obtain a heat source area and mapping it to the visible light image.

[0085] Further, the step of performing confidence evaluation on the heat source candidate area to obtain a heat source area includes: extracting the texture information, shape information, and edge information of the heat source candidate area, screening the human body heat distribution to obtain a heat map form; judging the metabolic activity signal through the gas concentration change obtained by the gas concentration sensor to obtain the gas concentration; acquiring an audio segment based on the voice sensor and detecting the voice rhythm signal to obtain an audio signal; performing confidence evaluation on the heat source candidate area according to the heat map form, gas concentration, and audio signal to obtain the heat source area.

[0086] Further, the method for performing human body local recognition based on the visible light image, inferring and completing the posture based on local key points, and constructing the human body structure features includes: judging whether there is a lack of human body recognition features based on the comparison between the visible light image and the thermal imaging image; if there is a lack of human body recognition features, extracting local key points in the visible area; combining the local contour and temperature gradient in the thermal imaging image, and using a human body posture completion network to reconstruct the human body structure features based on the local key points.

[0087] Further, the method for determining a rescue target based on the human body structure features and marking the target data of the rescue target in the robot coordinate map includes: calculating the spatial depth information of the target area where the rescue target is located through binocular vision; fusing the lidar scanning results to perform three-dimensional structure correction on the spatial depth information to obtain target spatial information; mapping the target spatial position to the robot coordinate map, and recording the rescue target position, occlusion level, and detection time information of the rescue target; marking the target data in the robot coordinate map and uploading it to the emergency command platform, and broadcasting it to the collaborative rescue robot nodes for sharing.

[0088] Further, before broadcasting to the collaborative rescue robot nodes for sharing, it includes: constructing an initial path in the robot coordinate map according to the marking information; collecting real-time target spatial information by the rescue robot on the initial path, and when a new obstacle is detected in the real-time target spatial information, triggering a path reconstruction mechanism to update the initial path in the robot coordinate map to obtain an updated path.

[0089] Further, broadcasting to the collaborative rescue robot nodes for sharing includes: dividing the reachable path and unreachable path based on the updated path, marking the rescue robot as an unreachable target for the rescue target according to the unreachable path, and broadcasting it to the collaborative rescue robot for execution.

[0090] Further, marking the target data in the robot coordinate map and uploading it to the emergency command platform, and broadcasting it to the collaborative rescue robot nodes for sharing includes: organizing the target data into a standard format data packet and uploading it to the emergency command platform through a wireless communication module; broadcasting the target data in the robot cluster network; setting a periodic status update mechanism for the target data, and when the target data changes, updating the marking information and triggering a tracking module.

[0091] Each embodiment in this specification is described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The foregoing Figure 1The autonomous recognition and positioning system of the rescue robot based on industrial vision and the specific examples in Embodiment 1 are equally applicable to the autonomous recognition and positioning method of the rescue robot based on industrial vision in this embodiment. Through the detailed description of the autonomous recognition and positioning system of the rescue robot based on industrial vision above, those skilled in the art can clearly know the autonomous recognition and positioning method of the rescue robot based on industrial vision in this embodiment. Therefore, for the sake of simplicity of the specification, it will not be elaborated here. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

[0092] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. An autonomous recognition and positioning system for a rescue robot based on industrial vision, characterized in that, Including: A collaborative scanning module, which is used to execute the initialization of industrial vision equipment, and collaboratively scan through the industrial vision equipment and lidar to construct a robot coordinate map; A data fusion module, which is used to establish a data fusion frame with unified timing according to the robot coordinate map, combined with inertial measurement and audio sensing information, to obtain multi-modal image data, where the multi-modal image data includes thermal imaging images and visible light images; A human body recognition module, which is used to map the thermal imaging image to the visible light image, perform local human body recognition based on the visible light image, perform pose inference and completion based on local key points, and construct a human body structure feature; A target annotation module, which is used to determine a rescue target based on the human body structure feature, annotate the target data of the rescue target in the robot coordinate map, and obtain annotation information; A path planning module, which is used to perform path planning based on the annotation information, construct and update a path, and execute a rescue through a rescue robot.

2. The autonomous recognition and positioning system of a rescue robot based on industrial vision according to claim 1, wherein, The human body recognition module includes: A temperature analysis unit, which is used to perform regional temperature analysis on the thermal imaging image, extract heat source candidate regions, and establish regional boundaries; A confidence evaluation unit, which is used to evaluate the confidence of the heat source candidate regions to obtain heat source regions and map them to the visible light image.

3. The autonomous recognition and positioning system of a rescue robot based on industrial vision according to claim 2, characterized in that, The confidence evaluation unit includes: A heat distribution screening channel, which is used to extract texture information, shape information, and edge information of the heat source candidate regions, screen the human body heat distribution to obtain a heat map form; A signal judgment channel, which is used to judge the metabolic activity signal through the gas concentration change obtained by a gas concentration sensor to obtain the gas concentration; An audio signal acquisition channel, which is used to acquire an audio segment based on a voice sensor and detect a voice rhythm signal to obtain an audio signal; A heat source region determination channel, which is used to evaluate the confidence of the heat source candidate regions according to the heat map form, gas concentration, and audio signal to obtain the heat source regions.

4. The autonomous recognition and positioning system of a rescue robot based on industrial vision according to claim 1, wherein, The human body recognition module includes: An image comparison unit, which is used to judge whether there is a lack of human body recognition features based on the comparison between the visible light image and the thermal imaging image; A key point extraction unit, which is used to extract local key points in the visible region if there is a lack of human body recognition features; A structure reconstruction unit, which is used to combine the local contour and temperature gradient in the thermal imaging image, and use a human body pose completion network to reconstruct a human body structure feature based on the local key points.

5. The autonomous recognition and positioning system of a rescue robot based on industrial vision according to claim 1, characterized in that, The target annotation module includes: A spatial depth information calculation unit, which is used to calculate the spatial depth information of the target area where the rescue target is located through binocular vision; A three-dimensional structure correction unit, which is used to fuse the lidar scanning result to perform three-dimensional structure correction on the spatial depth information to obtain target spatial information; A target position mapping unit, which is used to map the target spatial position to the robot coordinate map, and record the rescue target position, occlusion level, and detection time information of the rescue target; A target annotation unit, which is used to annotate the target data in the robot coordinate map and upload it to the emergency command platform, and broadcast it to the collaborative rescue robot nodes for sharing.

6. The autonomous recognition and positioning system of the rescue robot based on industrial vision according to claim 5, characterized in that, The target annotation unit includes: A path construction channel for constructing an initial path in the robot coordinate map according to the annotation information; A path update channel for collecting real-time target space information by the rescue robot on the initial path. When a new obstacle is detected in the real-time target space information, a path reconstruction mechanism is triggered to update the initial path in the robot coordinate map to obtain an updated path.

7. The autonomous recognition and positioning system of the rescue robot based on industrial vision according to claim 6, characterized in that, The target annotation unit further includes: An unreachable target marking channel for dividing the reachable path and the unreachable path based on the updated path, marking the rescue robot as an unreachable target of the rescue target according to the unreachable path, and broadcasting it to the cooperative rescue robot for execution.

8. The autonomous recognition and positioning system of a rescue robot based on industrial vision according to claim 5, characterized in that, The target annotation unit further includes: A data upload channel for organizing the target data into a standard format data packet and uploading it to the emergency command platform through a wireless communication module; A data broadcast channel for broadcasting the target data in the robot cluster network; An annotation update channel for setting a periodic status update mechanism for the target data. When the target data changes, the annotation information is updated and the tracking module is triggered.

9. The autonomous recognition and positioning system of a rescue robot based on industrial vision according to claim 1, wherein The cooperative scanning module includes: A device activation unit for activating the industrial vision device, verifying the connection status and the validity of the sensing signal of the industrial vision device, and obtaining a self-check result; A map construction unit for constructing an initial robot coordinate map through the collaborative scanning of the industrial vision device and the lidar; A map update unit for extracting the passable area and the obstacle boundary according to the laser point cloud, setting the navigation origin, and updating the initial robot coordinate map to generate the robot coordinate map.

10. An autonomous recognition and positioning method for a rescue robot based on industrial vision, characterized in that, Executed by the industrial vision-based rescue robot autonomous recognition and positioning system according to any one of claims 1 to 9, the industrial vision-based rescue robot autonomous recognition and positioning method includes: Performing initialization of the industrial vision device, and constructing a robot coordinate map through the collaborative scanning of the industrial vision device and the lidar; According to the robot coordinate map, combining inertial measurement and audio sensing information to establish a data fusion frame with unified time sequence to obtain multimodal image data, wherein the multimodal image data includes a thermal imaging image and a visible light image; Mapping the thermal imaging image to the visible light image, performing human body local recognition based on the visible light image, performing pose inference and completion based on local key points, and constructing a human body structure feature; Determining a rescue target based on the human body structure feature, annotating the target data of the rescue target in the robot coordinate map to obtain annotation information; Performing path planning based on the annotation information, constructing an updated path, and executing rescue by the rescue robot.

Citation Information

Patent Citations

  • Multi-sensor fused search and rescue robot system

    CN112109090A

  • Life characteristic detection and identification method for rescue robot

    CN112248032A

  • 3D human body reconstruction method based on infrared thermal imaging

    CN113112583A

  • First-aid robot

    CN116330256A

  • Control method of mining robot, mining robot and storage medium

    CN117140534A

Cited By

  • Security robot dynamic sensing system integrating space mapping and target recognition

    CN120715961A

  • Dynamic Perception System for Security Robots Integrating Spatial Mapping and Target Recognition

    CN120715961B

  • Robot multi-mode environment feature modeling and recognition positioning method and system

    CN121721943A

  • Emergency rescue robot cooperative control system integrating intelligent perception and multi-mode communication

    CN121912395A

  • Multi-view collaborative operation behavior risk assessment method and system

    CN121999539A