Intelligent unmanned vehicle system for rescue task scene
By designing an intelligent unmanned vehicle system in an automated rescue robot, using radar mapping, path planning, target detection and voice interaction modules, the problems of navigation and decision-making in complex environmental rescue scenarios are solved, and efficient human-computer interaction and task execution are achieved.
Patent Information
- Application Number
- CN202510207318.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
Existing automated rescue robots are unable to meet the navigation requirements of complex environmental rescue scenarios, cannot make decisions in real time, and do not have human-computer interaction functions.
An intelligent unmanned vehicle system for rescue mission scenarios is designed, equipped with a computer motherboard, a depth camera, a three-dimensional lidar and a microphone array, and automatic navigation, target recognition and human-computer interaction are achieved through radar mapping modules, path planning modules, target detection modules and voice interaction modules.
It realizes automatic navigation, real-time decision-making and human-computer interaction in complex environments, and improves the adaptability and task success rate of unmanned vehicles in uncertain environments.
Smart Images

Figure CN120044955A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent unmanned vehicle system for rescue mission scenarios, and belongs to the technical fields of lidar mapping, autonomous driving, target detection and recognition, embedded artificial intelligence, and path planning. Background Art
[0002] In complex natural disaster scenarios, both the weather conditions and the harsh complexity of the scene environment are extremely likely to cause injuries to people. Due to the weak self-protection ability and the weak ability to handle emergencies of first aid personnel in complex scenarios, the probability of injury when facing danger is high. Moreover, their complex professional knowledge leads to huge training costs, and the number of first aid personnel has always been in short supply. In recent years, in response to this problem, the field of fire science and technology in China has continuously focused on the research of automatic robot systems. Currently, the relatively mature solution still requires professional rescue personnel to manually remotely control complex scene rescue robots carrying medical tools to be responsible for rescue tasks. However, this solution can only temporarily alleviate the sharp reduction in the number of first aid personnel. One first aid personnel is still required to operate one automatic robot, so it cannot effectively solve the problem of the shortage of first aid personnel. In addition, since first aid personnel remotely view the affected complex scene through a camera, the rescue efficiency when facing the wounded is not as high as that of on-site rescue. Moreover, the complex scene environment is harsh, and the remote communication signal is easily interfered, and the remotely controlled robot will be disconnected at any time, resulting in mission failure.
[0003] Therefore, it is an urgent need to use unmanned robots equipped with artificial intelligence to replace first aid personnel to complete high-complexity and high-risk rescue tasks in harsh environments. By introducing automatic navigation technology, the robot is given the ability to automatically plan paths and automatically go to designated target points, so that the robot can be set as a semi-automatic or fully automatic robot. When the environmental uncertainty factors change greatly or the task plan needs to be adjusted, it can be flexibly switched to manual control, or when the communication signal is interfered and the manual control function fails, the artificial intelligence automatic control system takes over the robot, improving the robot's ability to cope with risks. For complex scene survey tasks, a camera and visual recognition algorithms are installed to automatically identify the wounded and provide assistance in a chaotic environment, improving the robot's task execution efficiency. A voice system is installed to facilitate the timely feedback of battlefield information by the wounded, and the rescue target is given command authority to affect the decision-making of the automatic robot and improve the rationality of task planning. The existing automatic robot systems have the following problems:
[0004] (1) Most of the existing automatic robot systems are basic modules with basic path planning functions and cannot meet the navigation requirements of complex environment rescue scene tasks.
[0005] (2) After the existing automatic robot systems are equipped with a visual recognition system, they can only perform target recognition and cannot return status information in real time and make decisions.
[0006] (3) In existing automatic robot systems, most command inputs are direct computer commands, and they do not support real-time command adjustment and do not have a human-computer interaction function.
[0007] In summary, in a dangerous environment, in order to successfully complete tasks, an automatic robot must first ensure its own safety. Automatically identifying dangerous roads and avoiding them in a dangerous environment and re-making navigation decisions are also extremely important system functions. The robot evaluates decision-making risks in real time to ensure the success rate of tasks. To effectively solve the shortage problem of fire and emergency rescue personnel, it is imperative to develop an automatic rescue robot with artificial intelligence that is equipped with an embedded software for an automatic robot that can automatically navigate, automatically scan the surrounding environment to identify targets after reaching the target point. Summary of the Invention
[0008] The technical problem to be solved by the present invention is:
[0009] In order to solve the problems that existing automatic rescue robots cannot meet the navigation requirements of rescue scenarios in complex environments, cannot make decisions in real time, and do not have a human-computer interaction function, the present invention provides an intelligent unmanned vehicle system for rescue task scenarios.
[0010] The technical solution adopted by the present invention to solve the above technical problems is:
[0011] An intelligent unmanned vehicle system for rescue task scenarios, the intelligent unmanned vehicle system is mounted on an unmanned vehicle chassis (a small vehicle chassis, including a main body, four motors, and four drive wheels), and it includes a computer main board, a depth camera, a 3D lidar, and a microphone array. The depth camera is used to detect rescue targets and transmit environmental picture information to the computer main board. The lidar is used to sense environmental information and transmit it to the computer main board. The microphone array is used to transmit external voice commands obtained in human-computer interaction to the computer main board. The computer main board controls the driving motors of the unmanned vehicle chassis according to the information fed back by the computer main board, the depth camera, and the lidar, so that the unmanned vehicle can accurately reach the rescue target and return to the starting point after completing the task;
[0012] The processor on the computer main board is configured with a radar mapping module, a path planning module, a target detection module, and a voice interaction module. Each functional module constitutes an embedded software for a rescue unmanned vehicle, which is used for automatic navigation and automatically scanning the surrounding environment to identify targets after reaching the target point;
[0013] The radar mapping module is used to automatically explore the surrounding environment of a complex scene and convert the sensed environmental information into map data;
[0014] The path planning module calculates the optimal cost path between the starting point and the target point based on the acquired map data, and generates motion data for output to the target inspection module; path planning is implemented using 3D lidar mapping and graph theory algorithm path planning;
[0015] The target detection module acquires image data through a depth camera and identifies the target to be detected, generating the coordinate position (3D coordinate data of the target) where the rescue target is located; the information of the coordinate position is passed to the path planning module to drive the unmanned vehicle towards the coordinate position where the target is located;
[0016] The depth camera uses the YOLOv5 image recognition algorithm for target detection and recognition in complex scenarios;
[0017] It is used to calculate the relative distance between the unmanned vehicle and the rescue target in real time. When the set distance is reached, a task completion signal is sent to the voice interaction module through the processor on the computer mainboard;
[0018] The voice interaction module is used to issue a voice prompt according to the task completion signal, prompting the unmanned vehicle to reach the target and wait for subsequent instructions. It is also used to obtain external voice instructions in human-machine interaction, complete the information interaction of complex scenario human-machine instructions, and control the decision-making of the unmanned vehicle in real time (such as returning to the base, going to the next rescue location, or staying in place to continue the search, etc.); the voice interaction module uses an algorithm based on natural language processing (an algorithm within the scope of existing technologies).
[0019] Furthermore, the implementation process of the lidar mapping module is as follows:
[0020] The lidar mapping module includes a data acquisition module, a feature extraction module, and a state estimation module;
[0021] During the mapping process, first, the lidar data collected by the data acquisition module is input into the feature extraction module to obtain planar features and edge features; then, the extracted features and the measurement values of the IMU built into the 3D lidar are input into the state estimation module for state estimation at a frequency of 10Hz - 50Hz; then, the new feature points obtained by the state estimation module are merged with the feature point map constructed so far; the updated map will incorporate more new real-time environmental information; a backpropagation process is added to the state estimation module, and the results obtained from backpropagation and forward propagation are input together into the remaining calculations, and the remaining calculations are within the scope of existing technologies;
[0022] Furthermore, the implementation process of the path planning module is as follows:
[0023] The implementation process of the path planning module is as follows:
[0024] The map data obtained by subscribing to the radar mapping module and the positioning data of the unmanned vehicle calculated by the amcl algorithm are input into the MOVE-BASE framework to plan the global and local paths. MOVE-BASE is the central framework for the robot path planning function under ROS. Then, the path is converted into the speed output instruction of the unmanned vehicle, and finally, the navigation of the unmanned vehicle is realized;
[0025] Furthermore, the MOVE-BASE framework can be implemented based on the Dijkstra algorithm, and the heuristic function and the algorithm for estimating the cost are added. Its main formula is expressed as:
[0026] f(n) = g(n) + h(n)
[0027] Where f(n) is the cost estimation function from the initial point through node n to the target point, g(n) is the actual cost from the initial node to node n in the state space, and h(n) is the estimated cost of the best path from n to the target node; the key to ensuring finding the shortest path (optimal solution) lies in the selection of the evaluation function f(n). The estimated value h(n) is not greater than the actual value of the distance from n to the target node; if h(n) = f(n), that is, the distance estimation h(n) is equal to the shortest distance, then the search will strictly follow the optimal path, and the search efficiency is the highest at this time.
[0028] Furthermore, the MOVE-BASE framework adopts a form of loop calculation, exchanges map data through the ROS communication mechanism. After calling the main function of the MOVE-BASE program, a path planning function task is automatically created, and the ROS message loop is started. When it is judged that there is no second planning node in the system, that is, the ROS server is not preempted by the same type of function instance, the calculation of the global path planning will start; since the planning entity is a user-defined robot model, the system will perform a coordinate system self-check before starting the task. If the randomly specified target point is in the world coordinate system, the global optimal path will be directly calculated. Otherwise, the local path planner will be called, and after coordinate transformation, a temporary path will be generated and an attempt will be made to match the global path.
[0029] If the local path planner successfully performs global planning, the motion vector will be directly calculated. When the local path planner fails, the new target point after coordinate transformation is input into the global path planner to obtain the global path. If both of them fail to plan, it means that the point is unreachable, and the program will end immediately.
[0030] Furthermore, the implementation process of the target detection module (detection and recognition) includes: one is to classify the detection target, and the other is to confirm the coordinates of the detection target;
[0031] The face detection algorithm is implemented using the YOLOv5s model. The network structure of YOLOv5 mainly consists of four parts: Input (input end), Backbone (base network), PANet, and Output (output network).
[0032] First, the Input (input end) is to input the original data image with a size of 608*608. At this stage, the original image is preprocessed using the Mosaic data augmentation method, adaptive anchor boxes, and adaptive image scaling.
[0033] The Backbone (base network) uses the Focus module to achieve fast downsampling, concentrating information on the channels without information loss, making feature extraction more sufficient.
[0034] The network structure PANet adopts the FPN+PAN structure. The FPN layer downsamples, continuously shrinking the image, and improves the target detection effect by fusing high and low-level features. Then, the PAN structure is added for upsampling, making full use of the shallow features of the network for segmentation, and the top layer can receive the position information brought by the bottom layer.
[0035] The Output (output network) is used to output the target detection results and evaluate whether the object localization algorithm is accurate. YOLOv5 uses the loss function GIOU for calculation and adopts the non-maximum suppression algorithm to remove other prediction results whose GIOU_Loss contribution value is not the largest, and outputs the classification result with the highest probability. By analyzing the principle and network structure of the YOLOv5 algorithm model, in the face recognition process, the BiFPN (Bidirectional Weighted Feature Pyramid Network) feature fusion structure is introduced into the YOLOv5 algorithm, and a small target detection layer is added to the feature extraction network to extract the shallow physical information to enhance the detection ability of small targets.
[0036] Furthermore, in the task scenario of the target detection and recognition process, the picture input comes from the D435 binocular depth camera. The image resolution is preliminarily processed, and the picture is scaled according to the input dimension of the YOLOv5 architecture to eliminate the influence of noise and perform range smoothing, making it easier for the base network to extract the main features. After the preprocessing is completed, the picture is input into the trained YOLOv5 model. Due to the subsequent influence on navigation decisions, a continuous frame recognition judgment is added at the output end. Only when the recognition is successful for more than 5 consecutive frames can it be determined that the target detection is successful, and the target position coordinates are calculated and docked with the path planning module.
[0037] Furthermore, the recognition process of the voice interaction module is as follows:
[0038] The overall voice recognition interaction module adopts the iFlytek voice recognition microphone array solution, uses the ASR automatic speech recognition solution and the NLU natural language understanding solution to realize the function of guiding the robot to execute tasks through voice transmission instructions, and finally realizes voice output through the TTS speech synthesis solution;
[0039] ASR automatic speech recognition converts speech into text information, which is the key to speech recognition technology; Using artificial neural networks and training with the backpropagation BP algorithm can improve the accuracy of speech recognition;
[0040] When developing the voice recognition function in a Python environment, it is implemented using third-party library functions. Before recognition, the MFCC features of the test set and the neural network model in the test case have been obtained. First, the trained neural network model is loaded, then the MFCC features of the test set are input into the loaded model for comparison, and finally the result with the highest comparison similarity is output; Mel Frequency Cepstral Coefficients (MFCC) are used as the feature parameters of speech. First, a voice sample in the real-time recording system during task execution is obtained, and the same commands of different commanders are collected and used as the training set and test set for voice recognition; Before feature extraction, the voice is preprocessed, and the preprocessed voice sample will be used for feature extraction to obtain MFCC features.
[0041] A computer-readable storage medium stores a computer program, and the computer program is configured to implement the execution steps of the rescue unmanned vehicle embedded software when called by a processor.
[0042] An intelligent unmanned vehicle for a rescue task scenario, characterized in that: the intelligent unmanned vehicle for a rescue task scenario includes the intelligent unmanned vehicle system for a rescue task scenario, and a memory communicatively connected to a processor on a computer motherboard stores instructions executable by the processor on the computer motherboard. When the instructions are executed by the processor, the processor can execute the rescue unmanned vehicle embedded software to enable the automated robot to automatically navigate and automatically scan the surrounding environment to identify targets after reaching the target point.
[0043] The present invention has the following beneficial effects:
[0044] The present invention has developed a set of embedded software for automated rescue unmanned vehicles (robots) that can automatically navigate and automatically scan the surrounding environment to identify targets after reaching the target point. The automated design equipped with artificial intelligence enables the rescue unmanned vehicle to flexibly adjust its decision when facing complex environments, thereby improving the adaptability of the unmanned rescue unmanned vehicle system to uncertain environmental factors. The embedded software for the rescue unmanned vehicle provided by the present invention includes: a laser radar mapping module: used to perceive environmental information and convert it into map data; a path planning module: based on the acquired map data, the optimal cost path between the starting point and the target point is calculated, and motion data is generated and output to the system; a target detection module: the image data is acquired through a depth camera and the target to be detected is identified, and the three-dimensional coordinate data of the target is generated and output to the system; a voice interaction module: used to obtain external voice commands in human-computer interaction, and to control the decision-making of the rescue unmanned vehicle system in real time. The radar mapping module and the path planning module constitute the navigation function.
[0045] Aiming at complex mission scenarios, the present invention develops an embedded system for a rescue unmanned vehicle with visual recognition function, voice interaction function and automatic navigation function. Facing complex scenarios, a laser radar is used to obtain environmental point cloud information to generate an environmental map, and the optimal cost path between the starting point and the target point is calculated based on the acquired map data, and motion data is generated and output to the system, which has high accuracy and data real-time performance. The image data is acquired through a depth camera and the target to be detected is identified, and the three-dimensional coordinate data of the target is generated and output to the system. The YOLOv5 algorithm is used to identify a variety of objects, and the decision-making system is used to complete complex functions such as danger avoidance and target tracking, which has high intelligence. The voice interaction module is used to realize the real-time voice command input function, and a set of intelligent human-computer interaction subsystems are constructed, so that the automatic robot decision-making system has strong adaptability and stability to complex environments.
[0046] The present invention has developed a set of embedded software for automated rescue unmanned vehicles that can automatically navigate and automatically scan the surrounding environment to identify targets after reaching the target point. The automated rescue unmanned vehicle equipped with artificial intelligence can effectively solve the shortage of fire emergency personnel, enhance the ability of the rescue unmanned vehicle to cope with environmental uncertainties, and improve the robustness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is the overall design diagram of the intelligent unmanned vehicle system;
[0048] Figure 2 This is the FAST-LIO mileage calculation method flow (specific process of radar mapping) used in the mapping technology of the present invention;
[0049] Figure 3 It is the program topology flow of the automatic navigation function (path planning and control output);
[0050] Figure 4 It is a schematic diagram of the object detection and recognition function of the present invention;
[0051] Figure 5 It is a schematic diagram of the speech recognition and interaction function of the present invention;
[0052] Figure 6 It is a physical photo of an intelligent unmanned vehicle (showing the component modules of the unmanned vehicle) for rescue mission scenarios based on the technical solution of the present invention;
[0053] Figure 7 It is a photo of the experimental scenario of the unmanned vehicle developed based on the present invention. Specific implementation manners
[0054] Combined with the attached Figures 1-6 , the implementation of the intelligent unmanned vehicle system for rescue mission scenarios described in the present invention is elaborated as follows:
[0055] To solve the above problems, an intelligent unmanned vehicle system for rescue mission scenarios described in the present invention mainly focuses on the following four aspects: lidar three-dimensional mapping technology, path planning and obstacle avoidance decision-making technology, target inspection technology (visual target recognition technology), and voice interaction technology (speech recognition and interaction technology). In response to the need for the rescue unmanned vehicle (robot) to automatically explore the surrounding environment of complex scenarios, a design plan for the automatic navigation sub-module is proposed, and it is developed using the design scheme of three-dimensional lidar mapping and graph theory algorithm path planning; in response to the need for target detection and recognition in complex scenarios, a design plan for the visual recognition sub-module is proposed, and it is developed using the scheme based on the YOLOv5 image recognition algorithm; in response to the need for human-machine command information interaction in complex scenarios, a design plan for the speech recognition sub-module is proposed, and it is developed using the speech recognition scheme based on natural language processing algorithms. After the successful development of each functional sub-module, they will be combined into a complete embedded software system for the rescue unmanned vehicle, and simulation and verification work for the hardware environment scheme will be carried out. The overall technical solution is as Figure 1 shown. The hardware of the rescue unmanned vehicle system includes a microphone voice array equipped with a speaker, a depth camera, a lidar (integrated with IMU), a UWB positioning module, and an ORIN edge core processor. The processor board receives the voice information from the microphone array, the visual images and RGBD maps from the binocular camera, the three-dimensional point cloud map from the lidar, and the data of each sensor. The original input data is distributed to each subsystem for data processing and calculation. USB, HDMI, and network ports are reserved externally as debugging interfaces, and the intelligent brain control system communicates with the torso control system using the network port to issue decision instructions.
[0056] (1) Lidar mapping technology
[0057] During the mapping process, the lidar input data is first fed into the feature extraction module to obtain plane features and edge features. The extracted features and IMU measurements are then fed into the state estimation module to perform state estimation at 10Hz-50Hz. After that, the estimated pose registers the feature points into the global frame and merges them with the feature point map built so far. The updated map will add more new points in the next step. In a significant difference from past algorithms, a backpropagation process is added to the IMU state estimation module, and the results obtained from the backpropagation are input into the remaining calculations together with the results obtained from the forward propagation.
[0058] The FAST-LIO algorithm belongs to laser SLAM, and its algorithm flow is as follows Figure 2 As shown. This algorithm can choose to directly register the raw point cloud to the map (and then update the map) without extracting features, which can take advantage of subtle features in the environment to improve accuracy. No artificially designed feature extraction module is used here, making it naturally adaptable to emerging lidars with different scanning modes; it can also choose to maintain the map through an incremental kd-tree data structure kd-tree, which supports incremental updates (i.e. point insertion, deletion) and dynamic balancing. Compared with existing dynamic data structures (octrees, R-trees, nanoflann kd-trees), kd-trees achieve excellent overall performance while naturally supporting downsampling. We conduct an exhaustive benchmark comparison on 19 sequences from various open lidar datasets. Due to the improved computational efficiency of the iKD tree, we directly register the raw point cloud to the map, which can achieve more accurate and reliable point cloud registration even in violent motion and very cluttered environments. We call this raw point-based registration a direct method, eliminating artificially designed feature extraction and making the system naturally applicable to different lidar sensors.
[0059] MOVE-BASE is the central hub of the path planning function for robots (in this case, unmanned vehicles) under ROS. It subscribes to data such as lidar, map, and amcl positioning, and then plans global and local paths, and then converts the paths into speed information of unmanned vehicles, ultimately achieving unmanned vehicle navigation.
[0060] The MOVE-BASE package provided by ROS allows us to specify the target position and direction in the established map, and then MOVE-BASE controls the unmanned vehicle to reach the desired target position based on the sensor information of the unmanned vehicle. Its main functions include: combining the odometry information calculated by the unmanned vehicle code disk for path planning, and outputting the forward speed and turning speed. These two speeds are automatically derived based on the maximum speed and minimum speed set in the configuration file.
[0061] like Figure 3It is the program topology process of the automatic navigation function of the driverless vehicle. On the left side are the functional modules of the above-mentioned radar mapping part. The final output format of the mapping module is encapsulated 3D point cloud map data, 2D grid map data, IMU odometer data, and adaptive Monte Carlo localization data. The program outside the module only has the read permission to prevent overwriting for data security and stability. It directly inputs the MOVE-BASE navigation framework for path calculation, plans the global and local paths, and then converts the paths into the speed information of the driverless vehicle, finally realizing the navigation of the driverless vehicle.
[0062] The path planning algorithm of the present invention uses the Dijkstra algorithm and incorporates the heuristic function and the algorithm for estimating the cost. Its main formula is expressed as:
[0063] f(n) = g(n) + h(n)
[0064] Among them, f(n) is the cost estimation function from the initial point through node n to the target point, g(n) is the actual cost from the initial node to node n in the state space, and h(n) is the estimated cost of the best path from n to the target node. The key to ensuring finding the shortest path (optimal solution) lies in the selection of the evaluation function f(n): the estimated value h(n) is not greater than the actual value of the distance from n to the target node. In this case, the number of searched points is large, the search range is large, and the efficiency is low, but the optimal solution can be found. And if h(n) = f(n), that is, the distance estimation h(n) is equal to the shortest distance, then the search will strictly follow the optimal path, and the search efficiency is the highest at this time. If the estimated value is greater than the actual value, the number of searched points is small, the search range is large, and the efficiency is high. Therefore, the latter is adopted. Generally, the actual cost of the map grid points is reduced and the evaluation coefficient is increased. Even if there is a risk of falling into a local optimum, the planned path can approach the global optimal path through cyclic planning. Essentially, it is to increase the confidence coefficient of the newly planned path, so it can effectively avoid obstacles that suddenly appear on the local map.
[0065] The MOVE-BASE framework adopts a form of cyclic calculation. Figure 3Both the global planner and the local planner in (the planning algorithm) need to be specified by the user, and data exchange is carried out through the ROS communication mechanism. It supports using the path planning package encapsulated by the official website or a custom path planning algorithm. After calling the main function of the MOVE-BASE program, the system will automatically create an instance of the path planning function and start the ROS message loop. When it is judged that there is no second planning node in the system, that is, the ROS server is not preempted by the same type of function instance, the calculation of the global path planning will start. Since the planning entity is a user-defined unmanned vehicle model, the system will perform a coordinate system self-check before starting the task. If the randomly specified target point is in the world coordinate system, the global optimal path will be directly calculated; otherwise, the local path planner will be called, and after coordinate transformation, a temporary path will be generated and an attempt will be made to match the global path. If the local path planner successfully conducts global planning, the motion vector will be directly calculated. When the local path planner fails, the new target point after coordinate transformation will be input into the global path planner to obtain the global path. If both fail, it means that the point is unreachable, and the program will end immediately. After the speed vector calculation is successful, it will be reported to the server through the ROS communication system, and the motion control module will read the latest frame of motion instructions from the server and execute them.
[0066] (2) Target detection and recognition technology
[0067] In the present invention, the target detection and recognition technology is divided into two task parts: one is to classify the detection target, and the other is to confirm the coordinates of the detection target.
[0068] The YOLO series of algorithms are deep neural network models applied to target detection. The high performance and easy-to-use characteristics of YOLOv5 make it more widely used in the field of target detection. In this paper, the face detection algorithm is studied based on the YOLOv5s model. The network structure of YOLOv5 is mainly composed of four parts, namely Input (input end), Backbone (benchmark network), PANet, and Output (output network).
[0069] First, the Input (input end) is to input the original data image, and the image size is 608*608. At this stage, the original image is preprocessed using the Mosaic data augmentation method, adaptive anchor boxes, and adaptive image scaling. This method enriches the data and improves the detection ability for small targets.
[0070] The Backbone uses the Focus module to achieve fast downsampling. This module can concentrate information onto channels without information loss, making the subsequent feature extraction more thorough. The use of the SPP module improves the scale invariance of the image, effectively increasing the receptive field of the backbone features, making it easier for the network to converge and improving the accuracy. The PANet network structure adopts the FPN+PAN structure. The FPN layer downsamples, continuously shrinking the image, and improves the object detection effect by fusing high and low level features. The subsequent PAN structure is added for upsampling, making full use of the shallow features of the network for segmentation, and the top layer can also receive the rich position information brought by the bottom layer.
[0071] The Output is used to complete the output of the object detection results. To evaluate whether the object localization algorithm is accurate, YOLOv5 uses the loss function GIOU for calculation, and uses the non-maximum suppression algorithm to remove other prediction results whose GIOU_Loss contribution value is not the largest, and outputs the classification result with the highest probability. By analyzing the principle and network structure of the YOLOv5 algorithm model, it is found that the data detection results of small models and ultra-small models are inaccurate during the face recognition process. Therefore, we introduce the BiFPN (Bidirectional Weighted Feature Pyramid Network) feature fusion structure into the YOLOv5 algorithm, add a small object detection layer to the feature extraction network, and extract the physical information of the shallow layer to enhance the detection ability of small objects.
[0072] As Figure 4 shown, in the task scenario targeted by the present invention, the picture input comes from the D435 binocular depth camera. The original picture resolution collected at a certain frame rate is not 608*608. It is necessary to perform preliminary image resolution processing, scale the picture according to the input dimension of the YOLOv5 architecture, eliminate the influence of noise, perform range smoothing so that the backbone network is more likely to extract the main features, and will not be affected by unnecessary interference, and the edge detection effect is more obvious. After the preprocessing is completed, the picture is input into the trained YOLOv5 model. Due to the subsequent relationship affecting navigation decisions, it is necessary to ensure the accuracy of the object detection module and minimize the large amount of computing power wasted caused by incorrect feedback of navigation target points due to misrecognition. Therefore, continuous frame recognition judgment is added at the output end. Only when the recognition is successful for more than 5 consecutive frames can it be determined that the object detection is successful, and the target position coordinates are calculated and docked with the path planning module.
[0073] (3) Voice interaction recognition technology
[0074] The voice recognition interaction module as a whole adopts the iFlytek voice recognition microphone array solution, uses the ASR automatic speech recognition solution and the NLU natural language understanding solution to realize the function of guiding the unmanned vehicle to execute tasks through voice transmission instructions, and finally realizes voice output by the TTS speech synthesis solution. The overall architecture is as Figure 5shown.
[0075] ASR automatic speech recognition is about converting speech into text information, which is the key to speech recognition technology. The use of artificial neural networks and back-propagation BP algorithm for training can improve the accuracy of speech recognition. With the continuous development of edge terminal computing power, it is possible to choose to use a hidden Markov model combined with a Gaussian mixture model to form a Gaussian mixture model for speech recognition, which has significant effects but consumes a lot of computing power. Both solutions are mainstream speech recognition solutions, but because the present invention (this topic) needs to take into account other modules with high computing power requirements, the former is finally chosen.
[0076] When using Python environment to develop speech recognition function, some steps can be implemented using third-party library functions, so that better results can be achieved without cumbersome operations. The main processes include feature extraction, establishment of neural network model, recognition, etc. Before recognition, the test set MFCC features and neural network model in the test case have been obtained. First, the trained neural network model is loaded, and then the MFCC features of the test set are input into the loading model for comparison, and finally the result of maximum comparison similarity is output. The present invention adopts Mel frequency cepstral coefficient (MFCC) as the feature parameter of speech. First, a speech sample in a real-time recording system of a mission flight sequence or a speech sample in a real-time recording system of a mission execution sequence is obtained, and the same commands of different commanders are collected, which are used as training sets and test sets for speech recognition. The speech is preprocessed before feature extraction, and the preprocessed speech samples will be used for feature extraction and MFCC features. Among them, Whisper is an open source speech recognition engine based on DeepSpeech and Kaldi projects, aiming to provide a simple and easy-to-use speech recognition function. Whisper happens to focus on end-to-end automatic speech recognition, which can convert speech into text, and then through subsequent development, the text can be mapped to driverless car instructions to achieve voice command effects.
[0077] It has been verified that the intelligent unmanned vehicle system for rescue mission scenarios proposed by the present invention solves the technical problems proposed by the present invention. The product development based on the technical solution of the present invention has been completed. After practical application, the technical effects and practicality claimed by the present invention have been verified. The present invention has been verified through practical application. Figure 6 and Figure 7 shown.
[0078] The present invention proposes an intelligent unmanned vehicle system for rescue mission scenarios, in which the processor of the computer mainboard is configured with a radar mapping module, a path planning module, a target detection module, and a voice interaction module. Each functional module constitutes the embedded software of the rescue unmanned vehicle, which is used for automatic navigation and automatic scanning of the surrounding environment to identify the target after reaching the target point; this is the inventive point of the present invention.
[0079] The embedded software or algorithm of the rescue unmanned vehicle proposed by the present invention is the underlying technical core of the present invention, and products in the upstream and downstream of the industry can be derived based on the software or algorithm.
[0080] Based on the software algorithm (method) proposed by the present invention, an embedded software system for a rescue unmanned vehicle is developed using a programming language. The system has program modules corresponding to the steps of the above technical solution, and executes the steps in the above-mentioned embedded software of the rescue unmanned vehicle when running.
[0081] The computer program of the developed system (software) is stored on a computer-readable storage medium. The computer program is configured to implement the steps of the above-mentioned embedded software of the rescue unmanned vehicle when called by a processor. That is, the present invention is materialized on a carrier to become a computer program product.
[0082] An intelligent unmanned vehicle for a rescue mission scenario, the intelligent unmanned vehicle for a rescue mission scenario includes the intelligent unmanned vehicle system for a rescue mission scenario. A memory communicatively connected to a processor on a computer motherboard stores instructions executable by the processor on the computer motherboard. The instructions are executed by the processor so that the processor can execute the embedded software of the rescue unmanned vehicle, enabling the autonomous robot to automatically navigate and automatically scan the surrounding environment to identify targets after reaching the target point.
[0083] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, application specific ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0084] The computing programs (also referred to as programs, software, software applications, or code) in the present invention include machine instructions for a programmable processor, and these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used in the present invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., magnetic disks, optical disks, memories, programmable logic device PLD) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0085] It should be understood that various forms of the flow shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and all are within the protection scope of the present invention.
Claims
1. An intelligent unmanned vehicle system for rescue mission scenarios, characterized in that: The intelligent unmanned vehicle system is mounted on the chassis of the unmanned vehicle, and includes a computer motherboard, a depth camera, a three-dimensional laser radar, and a microphone array. The depth camera is used to detect the rescue target and transmit the environmental image information to the computer motherboard, the laser radar is used to sense the environmental information and transmit it to the computer motherboard, and the microphone array is used to transmit the external voice command obtained in the human-computer interaction to the computer motherboard. The computer motherboard controls the driving motor action of the unmanned vehicle chassis according to the information fed back by the computer motherboard, the depth camera, and the laser radar, so that the unmanned vehicle accurately reaches the rescue target and returns to the starting point after completing the task; The processor on the computer motherboard is equipped with a radar mapping module, a path planning module, a target detection module, and a voice interaction module. Each functional module constitutes the embedded software of the rescue unmanned vehicle, which is used for automatic navigation and automatic scanning of the surrounding environment to identify the target after reaching the target point; Radar mapping module, which is used to automatically explore the surrounding environment of complex scenes and convert perceived environmental information into map data; The path planning module calculates the optimal cost path between the starting point and the target point based on the acquired map data, and generates motion data to output to the target inspection module; Path planning is achieved by using 3D laser radar mapping and graph theory algorithm path planning; The target detection module acquires image data through the depth camera and identifies the target to be detected, generating the coordinate position of the rescue target; The coordinate position information is transmitted to the path planning module, driving the unmanned vehicle to move towards the target coordinate position; The depth camera uses the YOLOv5 image recognition algorithm to detect and identify targets in complex scenes; It is used to calculate the relative distance between the unmanned vehicle and the rescue target in real time. When the set distance is reached, the processor on the computer motherboard sends a task completion signal to the voice interaction module; The voice interaction module is used to issue voice prompts based on the task completion signal, prompting the unmanned vehicle to reach the target and wait for subsequent instructions. It is also used to obtain external voice commands in human-computer interaction, complete human-computer command information interaction in complex scenarios, and control the decision-making of the unmanned vehicle in real time. The voice interaction module adopts a natural language processing algorithm.
2. The intelligent unmanned vehicle system for rescue mission scenarios according to claim 1 is characterized in that: The implementation process of the radar mapping module is as follows: The radar mapping module includes a data acquisition module, a feature extraction module, and a state estimation module; In the process of mapping, the lidar data collected by the data acquisition module is first input into the feature extraction module to obtain plane features and edge features; The extracted features and the measurements of the 3D LiDAR’s built-in IMU are then fed into the state estimation module, which performs state estimation at a frequency of 10Hz-50Hz. The state estimation module then merges the new feature points obtained with the feature point map built so far; the updated map incorporates more new real-time environmental information. A back propagation process is added to the state estimation module, and the results obtained by back propagation are input into the remaining calculations together with the results obtained by forward propagation.
3. The intelligent unmanned vehicle system for rescue mission scenarios according to claim 2 is characterized in that: The implementation process of the path planning module is: The implementation process of the path planning module is: The map data obtained by subscribing to the radar mapping module and the unmanned vehicle positioning data calculated by the amcl algorithm are input into the MOVE-BASE framework to plan the global and local paths. MOVE-BASE is the central framework for the robot path planning function under ROS. The path is then converted into the speed output command of the unmanned vehicle, and finally the navigation of the unmanned vehicle is realized.
4. The intelligent unmanned vehicle system for rescue mission scenarios according to claim 3 is characterized in that: The MOVE-BASE framework can be implemented based on the Dijkstra algorithm, and adds heuristic functions and cost estimation algorithms. Its main formula is expressed as: f(n)=g()n+(h) Where f(n) is the cost estimation function from the initial point via node n to the target point, g(n) is the actual cost from the initial node to node n in the state space, and h(n) is the estimated cost of the best path from n to the target node; The key to ensuring the shortest path, i.e. the optimal solution, is the selection of the evaluation function f(n), where the evaluation value h(n) is not greater than the actual distance from n to the target node. If h(n)=f(n), that is, the distance estimate h(n) is equal to the shortest distance, then the search will be strictly carried out along the optimal path, and the search efficiency is the highest at this time.
5. An intelligent unmanned vehicle system for rescue mission scenarios according to claim 3 or 4, characterized in that: The MOVE-BASE framework uses a cyclic calculation method to exchange map data through the ROS communication mechanism. After calling the main function of the MOVE-BASE program, a path planning function task is automatically created and the ROS message loop is started. When it is determined that there is no second planning node in the system, that is, the ROS server is not preempted by a similar function instance, the calculation of the global path planning will begin. Since the planning subject is a user-defined robot model, the system will perform a coordinate system self-check before starting the task. If the randomly specified target point is located in the world coordinate system, the global optimal path is directly calculated. Otherwise, the local path planner is called to generate a temporary path after coordinate transformation and try to match the global path. If the local path planner successfully performs global planning, the motion vector is directly calculated. When the local path planner fails, the new target point after coordinate transformation is input into the global path planner to obtain the global path. If both planning fail at the same time, it means that the point is unreachable and the program ends immediately.
6. The intelligent unmanned vehicle system for rescue mission scenarios according to claim 5, characterized in that: The implementation process of the target object detection module includes: first, classifying the detection target, and second, confirming the coordinates of the detection target; The face detection algorithm is developed using the YOLOv5s model. The network structure of YOLOv5 is mainly composed of four parts: input, backbone, PANet, and output. First, the input end Input is the input raw data image, the image size is 608*608. In this stage, the original image is preprocessed using Mosaic data enhancement, adaptive anchor frame, and adaptive image scaling; The baseline network Backbone uses the Focus module to achieve fast downsampling, focusing information on channels without information loss, making feature extraction more sufficient; The network structure PANet adopts FPN+PAN structure. The FPN layer downsamples and continuously shrinks the image. The target detection effect is improved by fusing high-level and low-level features. The PAN structure is added later for upsampling, making full use of the shallow features of the network for segmentation. The top layer can receive the location information brought by the bottom layer. The output network Output is used to complete the output of target detection results and evaluate whether the object positioning algorithm is accurate. YOLOv5 uses the loss function GIOU calculation, and uses the non-maximum suppression algorithm to remove other prediction results whose GIOU_Loss contribution value is not the largest, and outputs the classification result with the highest probability. By analyzing the principles and network structure of the YOLOv5 algorithm model, the bidirectional weighted feature pyramid network BiFPN feature fusion structure is introduced into the YOLOv5 algorithm in the face recognition process, and a small target detection layer is added to the feature extraction network to extract shallow physical information to enhance the detection ability of small targets.
7. The intelligent unmanned vehicle system for rescue mission scenarios according to claim 6, characterized in that: In the task scenario of the target detection and recognition process, the image input comes from the D435 binocular depth camera. The image resolution is preliminarily processed, and the image is scaled according to the input dimension of the YOLOv5 architecture to eliminate the influence of noise and perform range smoothing, making it easier for the baseline network to extract the main features. After the pre-processing is completed, the image is input into the trained YOLOv5 model. Due to the subsequent impact on navigation decisions, continuous frame recognition judgment is added at the output end. Only when more than 5 consecutive frames of recognition are successful can the target detection be determined to be successful, and the target position coordinates are calculated and connected to the path planning module.
8. The intelligent unmanned vehicle system for rescue mission scenarios according to claim 7, characterized in that: The recognition process of the voice interaction module is: The speech recognition interaction module adopts the speech recognition microphone array solution as a whole, uses the ASR automatic speech recognition solution and the NLU natural language understanding solution to realize the function of voice transmission instructions to guide the robot to perform tasks, and finally uses the TTS speech synthesis solution to realize speech output; ASR automatic speech recognition is the key to speech recognition technology, which converts speech into text information. Using artificial neural networks and back propagation BP algorithm for training can improve the accuracy of speech recognition. When developing speech recognition functions in Python environment, third-party library functions are used to implement it. Before recognition, the test set MFCC features and neural network models in the test case have been obtained. First, the trained neural network model is loaded, and then the MFCC features of the test set are input into the loaded model for comparison. Finally, the maximum comparison similarity result is output; Mel frequency cepstral coefficient MFCC is used as the feature parameter of speech. First, a speech sample in the real-time recording system during mission execution is obtained, and the same commands of different commanders are collected as the training set and test set of speech recognition; The speech is preprocessed before feature extraction, and the preprocessed speech samples will be used for feature extraction and obtain MFCC features.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the execution steps of the rescue unmanned vehicle embedded software described in any one of claims 1 to 8 when called by a processor.
10. An intelligent unmanned vehicle for rescue mission scenarios, characterized by: The intelligent unmanned vehicle for rescue mission scenarios includes the intelligent unmanned vehicle system for rescue mission scenarios as described in any one of claims 1-8, and the memory on the computer motherboard that is communicatively connected to the processor stores instructions that can be executed by the processor on the computer motherboard. The instructions are executed by the processor so that the processor can execute the rescue unmanned vehicle embedded software as described in any one of claims 1-8, so as to realize that the automated robot can automatically navigate and automatically scan the surrounding environment to identify the target after reaching the target point.
Citation Information
Patent Citations
On-site investigation and material supply method, system and equipment based on rescue robot
CN113532440A
Multi-sensor fused mine inspection rescue robot and control method thereof
CN116352722A
Mine emergency rescue vehicle capable of realizing autonomous construction of map
CN116476718A
Multi-mode slam method suitable for vision-assisted laser fusion IMU in indoor environment
CN117367427A
Cited By
Autonomous recognition and positioning system and method for rescue robot based on industrial vision
CN120245004A
Air-ground adaptive fusion sensing method
CN120747703A