Multi-sensory fusion bionic search robot with human-computer two-way interaction function

By integrating vision, hearing, touch and smell modules into a multi-sensory fusion bionic search robot, and combining it with a spiking neural network for information fusion decision-making, the problem of autonomous exploration and target recognition of search robots in complex environments has been solved, and more efficient target object recognition and operation has been achieved.

CN117021112BActive Publication Date: 2025-10-28ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311201895.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2025-10-28
Estimated Expiration
2043-09-18

AI Technical Summary

Technical Problem

Existing search robots are limited by their single perception modality in complex environments, making it impossible for them to explore autonomously and perform tasks effectively. In particular, they struggle to accurately identify and manipulate targets when vision is limited or obstructed by obstacles.

Method used

A multi-sensory fusion bionic search robot is adopted, integrating vision, hearing, touch and olfaction modules, and combining spiking neural networks for information fusion decision-making, so as to realize the autonomous movement of the robot body and the comprehensive judgment of target objects.

Benefits of technology

The ability to move autonomously and operate flexibly in complex environments improves the accuracy and efficiency of target identification, reduces perception bias caused by a single sensor, provides comprehensive environmental awareness, and facilitates the execution of search and rescue and bomb disposal missions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117021112B_ABST
    Figure CN117021112B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-sensory fusion bionic search robot with bidirectional human-machine interaction. The robot includes a vision module, an auditory module, a tactile module, an olfactory module, a motor drive module, a bionic motion control module, and an information fusion module. The motor drive module is connected to the bionic motion control module for motion control of the search robot. The information fusion module connects to the vision, auditory, tactile, and olfactory modules to acquire multimodal information. Furthermore, this robot can be used in various application fields such as disaster relief, explosive detection, and live object detection. The robot can be used in various application fields such as hazard source search and explosive detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomimetic search robots, specifically relating to a multi-sensory fusion biomimetic search robot with two-way human-machine interaction capabilities. Background Technology

[0002] Search robots play a crucial role in disaster relief and explosive ordnance detection, areas where the rapid and efficient discovery and processing of critical information is essential. Mobile robots, equipped with sensors, perceive the world remotely and can operate in environments that pose a danger or constraint to humans. However, autonomously searching for targets in complex environments remains a challenge.

[0003] Current search robots primarily employ vision-based methods, including optical / thermal imagers, cameras, and infrared sensors. However, these methods have inherent limitations, such as limited field of view, the presence of obstacles, and generally poor visibility under rubble, thus affecting their effectiveness. To effectively detect targets under cover, sound and odor signal analysis are viable alternatives. Acoustic life detectors, which utilize audio signal processing to identify the sound information of targets, have been widely used in commercial applications. Installing microphone arrays on mobile robots enables sound source localization and human-robot interaction, allowing for listening and recording without entering specific locations, providing a more flexible approach to searches.

[0004] Furthermore, gases, as markers of targets, can be sensitively detected by electronic nose systems in a non-contact manner. Gas sensors can detect changes in the composition and concentration of gases in the environment, thereby enabling target detection. Odors in the air disregard physical boundaries and the absence of light, providing a method for identifying and tracking remote sources, opening new possibilities for disaster search and rescue and explosives detection. However, the low gas concentration places high demands on the sensitivity and reliability of the sensors, which is an important factor to consider when applying odor signal analysis.

[0005] Currently, visual recognition is the most mature and widely used technology, primarily in visual navigation and object recognition. Vision contains rich information; cameras can capture environmental information and extract features. This effectively enables object recognition, and the vision module can return real-time scene information to remote operators, providing search and command personnel with an intuitive understanding of the scene. However, vision is susceptible to ambient light. Strong light can cause image overexposure, resulting in the loss of object features, while weak light can prevent the comprehensive extraction of object feature information, affecting object identification. Furthermore, in disaster relief and search and rescue scenarios, obstacles often obstruct the view, also impacting the effectiveness of visual judgment.

[0006] In search mission scenarios, a standalone mobile robot can only provide environmental information to the operator but cannot interact with the environment independently, still requiring human intervention to complete the task. A robotic arm, limited by its fixed workspace, can only operate flexibly within a confined area and cannot escape the constraints of its base. Therefore, a vehicle-arm system combining these two technologies allows for free platform movement and flexible manipulation of the target, making it more suitable for replacing human operators in complex environments. This vehicle-arm system, combined with the aforementioned multi-sensory information system (audio, auditory, tactile, and olfactory), enables autonomous movement and exploration of the environment, allowing it to perform search and detection tasks in complex environments.

[0007] In conclusion, when faced with complex environments, single-modal perception often has limitations. Summary of the Invention

[0008] To address the shortcomings of the existing technology, this invention provides a multi-sensory fusion bionic search robot with two-way human-machine interaction capabilities.

[0009] The objective of this invention is achieved through the following technical solution:

[0010] A multi-sensory fusion bionic search robot with human-machine two-way interaction is proposed, comprising: a robot body and a multi-modal perception fusion decision module set on the robot body;

[0011] The robot body includes a robot mechanical part, a motor drive module connected to the robot mechanical part, and a bionic motion control module communicatively connected to the motor drive module.

[0012] The multimodal perception fusion decision module includes: an olfactory module, a visual module, an auditory module, a tactile module, and an information fusion module that is communicatively connected to the bionic modules of the visual module, auditory module, and tactile module.

[0013] The vision module is used to capture visual information of the surrounding environment of the bionic search robot, and the vision module uses a spiking neural network to identify and classify the visual information of the surrounding environment and generate a pulse signal containing the path planning of the bionic search robot moving to the target object;

[0014] The auditory module is used to receive voice commands with target object tags; and the auditory module uses a spiking neural network to identify the direction of the target object relative to the bionic search robot and generate a pulse signal containing the direction information;

[0015] The tactile module is used to acquire at least a portion of the contour information and contour friction coefficient information of the target object; and the tactile module uses a spiking neural network to process the at least a portion of the contour information and contour friction coefficient information of the target object to generate a pulse signal containing the contour information and contour friction coefficient information of the target object.

[0016] The olfactory module is used to acquire odor information emitted by the target object; and the olfactory module uses a spiking neural network to determine the type of odor emitted by the target object and generates a pulse signal containing the type of odor emitted by the target object.

[0017] The information fusion module receives pulse signals generated by one or more of the vision, hearing, touch, and olfaction modules, integrates the pulse signals by aligning them in the time dimension, and processes them through a spiking neural network to obtain a decision result. It then sends the information containing the decision result to the bionic motion control module, which in turn sends motor motion parameter information to the motor drive module, thereby controlling the movement of the robot's mechanical parts.

[0018] Furthermore, the olfactory module, the visual module, the auditory module, and the tactile module, as well as the bionic motion control module, are all bidirectionally connected to the operator via a terminal. The terminal can be operated by the operator, who receives information from at least one of the olfactory module, the visual module, the auditory module, and the tactile module through the terminal and issues control commands to the bionic motion control module.

[0019] Furthermore, for different search targets, multiple modules generate pulse signals containing path planning for the multi-sensory fusion bionic search robot to move towards the suspected explosive device and the rescuer based on the captured location information of the suspected explosive device and the rescuer;

[0020] In the explosives search scenario, the information fusion module receives the pulse signal from the vision module and sends the pulse signal from the vision module directly to the bionic motion control module. The bionic motion control module sends motor motion parameter information to the motor drive module, and the motor drive module controls the mechanical part of the robot to move to the suspected explosives.

[0021] The auditory module receives the sound information emitted by the suspected explosive and uses a spiking neural network to generate a pulse signal of the target object relative to the direction of the bionic search robot. The olfactory module obtains the odor information emitted by the suspected explosive and uses a spiking neural network to generate a pulse signal containing the type of odor emitted by the suspected explosive.

[0022] When the suspected explosive device is visible, the vision module uses a spiking neural network to generate pulse signals representing the appearance features of the suspected explosive device. The information fusion module receives the pulse signals generated by the hearing and olfactory modules, as well as the pulse signals representing the appearance features of the suspected explosive device generated by the vision module. It integrates the various pulse signals through time-dimensional alignment and processes them through a spiking neural network to obtain a decision result. The decision result involves the location and type of the suspected explosive device. The information containing the decision result is sent to the terminal. If the suspected explosive device is an explosive device, the operator decides to dismantle it. The operator sends control information for dismantling the explosive device to the bionic motion control module through the terminal. This causes the bionic motion control module to send motor motion parameter information to the motor drive module, which then controls the robotic arm in the robot's mechanical part to dismantle the explosive device.

[0023] When the suspected explosive device is invisible, the information fusion module receives pulse signals generated by the auditory and olfactory modules, integrates the pulse signals through time-dimensional alignment, and processes them through a spiking neural network to obtain a decision result. The decision result involves the location and type of the suspected explosive device, and the information containing the decision result is sent to the terminal. If the suspected explosive device is an explosive device, the operator decides to dismantle it. The operator sends control information for dismantling the explosive device to the bionic motion control module through the terminal, causing the bionic motion control module to send motor motion parameter information to the motor drive module. Then, the motor drive module controls the robotic arm in the robot's mechanical part to extend to the location of the explosive device and dismantle it.

[0024] In search and rescue scenarios, the vision module generates pulse signals containing the path planning of the bionic search robot moving to the target object based on the captured position information of the target object, and combines them with the information fusion module to generate response control strategies for different scenarios;

[0025] The information fusion module receives the pulse signal from the vision module and sends the pulse signal from the vision module directly to the bionic motion control module. The bionic motion control module sends motor motion parameter information to the motor drive module, and the motor drive module controls the mechanical parts of the robot to move to the object to be searched and rescued.

[0026] The auditory module receives the sound information emitted by the object to be searched and rescued and uses a pulse neural network to generate pulse signals of the object to be searched and rescued relative to the direction of the bionic search robot;

[0027] The tactile module acquires at least a portion of the contour information and contour friction coefficient information of the object to be searched and rescued, and uses a pulse neural network to generate a pulse signal containing the contour information and contour friction coefficient information of the object to be searched and rescued.

[0028] The olfactory module acquires the odor information emitted by the object to be searched and rescued and uses a spiking neural network to generate a pulse signal containing the type of odor emitted by the object to be searched and rescued;

[0029] When the object to be rescued is visible, the vision module uses a spiking neural network to generate pulse signals of the object's appearance features. The information fusion module receives the pulse signals generated by the hearing and olfactory modules, as well as the pulse signals of the object's appearance features generated by the vision module. It integrates the pulse signals by aligning them in the time dimension and processes them through the spiking neural network to obtain a decision result. The decision result involves the location and type of the object to be rescued. The information containing the decision result is sent to the terminal. If the object to be rescued is a target object, the operator decides to rescue the target object. The operator sends control information for rescuing the target object to the bionic motion control module through the terminal. This causes the bionic motion control module to send motor motion parameter information to the motor drive module, which then controls the robotic arm in the robot's mechanical part to rescue the target object.

[0030] When the object to be rescued is invisible, the auditory module receives the voice information emitted by the object and uses a spiking neural network to generate a pulse signal containing the voice information. The information fusion module receives the pulse signals generated by the auditory module, the olfactory module, and the tactile module, and integrates the pulse signals containing the voice information by aligning them in a time dimension and then processes them through the spiking neural network to obtain a decision result. The decision result involves the location and type of the object to be rescued, and the information containing the decision result is sent to the terminal. If the object to be rescued is a target object, the operator decides to rescue the target object. The operator sends control information for rescuing the target object to the bionic motion control module through the terminal, causing the bionic motion control module to send motor motion parameter information to the motor drive module. Then, the motor drive module controls the robotic arm in the robot's mechanical part to extend to the location of the target object and rescue it.

[0031] Furthermore, the vision module includes a scene visual data capture unit, an event stream spatiotemporal feature extraction and encoding unit, an event stream pulse pooling unit, an event stream temporal resolution adjustment and integration unit, and a deep spiking neural network (SNN) learning spatiotemporal features and recognition unit;

[0032] The scene visual data capture unit uses a dynamic visual sensor to read dynamic information in the scene in real time and converts it into an event stream in address event expression format;

[0033] The event stream spatiotemporal feature extraction and encoding unit extracts and encodes spatiotemporal features from the event stream obtained after processing by the captured scene visual data unit, so as to transform visual information into feature codes of interest.

[0034] The event stream pulse pooling unit performs pulse pooling on the event stream obtained after processing by the event stream spatiotemporal feature extraction and encoding unit.

[0035] The event stream temporal resolution adjustment and integration unit adjusts and integrates the event stream temporal resolution obtained after processing by the event stream pulse pooling unit;

[0036] The deep spiking neural network (SNN) learning spatiotemporal features and recognition unit uses the event stream obtained after processing by the event stream time resolution adjustment and integration unit to learn spatiotemporal features in order to identify the target object.

[0037] Furthermore, the auditory module has a ring-shaped microphone array. It estimates the angle information of the sound source by calculating the time difference of sound wave arrival and records the raw data of the microphone array as auditory information. The auditory information is encoded into a pulse signal through biomimetic pulse coding and input into a biomimetic spiking neural network. The biomimetic spiking neural network consists of four layers of biomimetic neurons to recognize the input speech, thereby controlling the movement of the mobile robot. At the same time, the encoded auditory information is input into the information fusion module to realize the multimodal recognition task.

[0038] Furthermore, the tactile module adopts a hand-like structure, integrating a force sensor at the top of the finger to acquire the three-dimensional force intensity at the contact point. The acquired force intensity can provide basic information for generating the contact surface contour information. When the tactile module explores the environment, it can acquire the contact force between the end of the tactile module and the environment at that point. Based on the force feedback information, the contour on a certain time series is reconstructed, providing contour information for the robotic arm in the bionic search robot.

[0039] Furthermore, the olfactory module has an electronic nose array with multiple MOS sensors. Based on the recorded data from the electronic nose array, the olfactory encoding part adopts a three-layer spiking neural network architecture, in which the third layer provides negative feedback to the second layer, and the pulses of the second layer are output as the encoded result, converting the odor signal into a pulse signal. The pulse signal is input into the information fusion module to realize the multimodal recognition task.

[0040] Furthermore, the robot's mechanical parts include a vehicle body and a robotic arm. The vehicle body is an autonomous mobile vehicle with a two-wheel differential chassis, capable of autonomous navigation and precise tracking of a predetermined trajectory. The robotic arm is a six-degree-of-freedom collaborative robotic arm with torque feedback.

[0041] Furthermore, the dynamic vision sensor in the scene visual data capture unit reads dynamic information in the scene in real time and converts it into an event stream S in address event expression format. tr ={e i |e i =[x i ,y i ,t i ,p i ] T ,i∈[1,m]};

[0042] Where m is the total number of events, and e is each event in the event stream. i It has four pieces of information, namely the spatial coordinate x i ∈[1,...,H],y i ∈[1,...,W], where H and W are the resolutions of the event camera, and the timestamp t i and polarity p i ∈{-1,1}, where -1 and 1 represent "OFF" and "ON" events respectively. An "ON" event means that the light intensity increases, and an "OFF" event means that the light intensity decreases.

[0043] The event stream spatiotemporal feature extraction and encoding unit extracts and encodes the event stream obtained after processing by the captured scene visual data unit. It can be divided into a spatiotemporal event plane calculation stage and an encoding stage.

[0044] The spatiotemporal event plane computation phase involves using a spatiotemporal event plane descriptor to calculate the correlation between events within a spatiotemporal neighborhood. This descriptor includes a spatial correlation kernel function. A time-related kernel function For any event e i The mathematical expression for the spacetime event plane is as follows:

[0045]

[0046]

[0047]

[0048] Wherein, spatial coordinate x j ∈[1,…,H],y j ∈[1,…,W], u is the proportionality coefficient controlling the temporal and spatial correlation, σl and τ l These are the spatial standard deviation and temporal decay constant of the spatiotemporal plane, with the subscript 'l' indicating the l-th feature coding layer; A i It is event e i Spatiotemporal event plane description operator, spatial correlation kernel function S i (·) Focus on e i Centered on R l This function represents the spatial relationships of events within a rectangular neighborhood of a radius; it assigns greater weight to events that are spatially closer. (Temporal correlation kernel function) Focus on the latest events that occur within the time window τ. l The temporal relationship of internal events, i.e., satisfying t j =max(t(x) j ,y j )),t j ∈[t i -τ l ,t i This function assigns smaller weights to historical events that occurred earlier, assuming that those events have a smaller temporal correlation with the current event. Therefore, the greater the spatiotemporal correlation between the current event and events in its spatiotemporal neighborhood, the greater the weight it is assigned. This is beneficial for more accurately describing the spatiotemporal dynamic changes of the event flow.

[0049] Encoding phase: First N l The spatiotemporal surface of each event is used to initialize N. l Feature template:

[0050]

[0051] After initialization, these templates are updated using an online clustering algorithm for each event e. i Its spatiotemporal surface is obtained through the spatiotemporal event plane calculation stage, and is calculated using the following formula. Euclidean distance between each template:

[0052]

[0053] and The closest k-th template will be event e i Encoded as e i =[x i ,y i ,t i ,k] T Meanwhile, the k-th template will be updated using an online clustering algorithm:

[0054]

[0055]

[0056] Where, n k The number of events encoded as the k-th feature is α, which is the sum of n. k The control update magnitude, i.e., the smaller the number of events encoded as the k-th feature, the larger the corresponding template update magnitude, β represents. and Cosine correlation between them;

[0057] Two-layer feature encoding is used to extract features at different spatial and temporal scales in the event stream. The first layer uses small-scale spatiotemporal parameters τ1 and R1 to encode fine-grained features, while the second layer uses larger-scale spatiotemporal parameters τ2 and R2 to encode more complex features across spatiotemporal distances.

[0058] The event stream pulse pooling unit performs pulse pooling on the event stream obtained after processing by the event stream spatiotemporal feature extraction and encoding unit. This operation includes a spatial pooling stage and a refractory period stage.

[0059] The spatial pooling stage involves preserving the local key features of events through spatial pooling. Specifically, it first encodes events e within adjacent, non-overlapping r×r pixel regions. i =[x i ,y i ,t i ,k i ] T Perform space pooling, that is:

[0060] e i =[x i / r,y i / r,t i ,k i ] T

[0061] Where r is the pooling radius;

[0062] Refractory period: A spiking neuron receives spatially pooled events. This neuron operates with a low threshold and refractory period to ensure that a single input is sufficient to trigger an output pulse and to remain silent for a period of time. Therefore, any input event will trigger the neuron outside the refractory period to pulse, and then the neuron will remain silent until the refractory period t. ref Finally, only the events that cause neurons to fire are transmitted to the spiking neural network for classification;

[0063] The event stream temporal resolution adjustment and integration unit adjusts and integrates the event stream temporal resolution obtained after processing by the event stream pulse pooling unit. Currently, deep spiking neural networks (SNNs) cannot directly process microsecond-level event stream data. Therefore, a time collapse strategy is used to adjust the event stream temporal resolution and aggregate events with the same time. This strategy first divides the event stream into multiple time slices, then assumes that events within a time slice have the same timestamp, setting the position with an event input to 1 and otherwise to 0. Thus, the event stream within a slice is transformed into a binary input S for one time step. t The formula is expressed as follows:

[0064] S t (y i ,x i ,k i ) = 1

[0065] Among them, y i ,x i ,k i For the event flow within the t-th time slice {e i |t i The coordinates and number of feature channels of ∈[α×t,α×(t+1)-1]};

[0066] Therefore, we can obtain t (H,W,N) l A binary pulse vector of size 1 is used as the input to the spiking neural network at each time step;

[0067] The deep spiking neural network (SNN) learning spatiotemporal features and recognition unit inputs the event stream obtained after processing by the event stream temporal resolution adjustment and integration unit into the SNN, and trains the deep SNN using the spatiotemporal backpropagation (STBP) algorithm. This step is divided into the SNN feedforward calculation stage and the backpropagation training stage.

[0068] SNN feedforward computation stage: The discrete-time iterative LIF neuron model used in SNN is expressed by the following formulas for membrane potential update and pulse firing dynamics:

[0069]

[0070]

[0071]

[0072] Where λ is the membrane potential decay coefficient of the neuron, V thr It is the membrane potential threshold. This represents the pulse firing state controlled by the step function g(x). When the membrane potential exceeds the threshold, the neuron fires a pulse to pass to the next layer and resets the membrane potential at the next moment. The specific formula for calculating postsynaptic potential updates is as follows:

[0073]

[0074] in, This represents the synaptic weight between the i-th postsynaptic neuron in the l-th layer and the j-th presynaptic neuron in the previous layer; It is the pulse output of the j-th presynaptic neuron in layer l-1 at time t+1;

[0075] Backpropagation training phase: Due to the non-differentiable nature of pulse firing, STBP uses a rectangular function to replace the derivative of the neuron's output pulse with respect to the membrane potential, thereby supporting the backpropagation of errors.

[0076]

[0077] Where a1 is a hyperparameter that determines the gradient width;

[0078] Finally, the loss function of STBP is set to minimize the mean squared error of all samples under a given time window T:

[0079]

[0080] Where y s and Let represent the label vector and the output impulse vector of the s-th training sample in the last layer, respectively.

[0081] The beneficial effects of this invention are as follows: This invention primarily addresses the problems of existing search robots, such as their single perception component and inability to explore autonomously. It develops a biomimetic robot that combines visual, auditory, tactile, and olfactory perception methods. In complex environments, it can comprehensively assess these four senses for judgment. Furthermore, its perception and control methods, based on a spiking neural network, more closely resemble human judgment, achieving a comprehensive perception and judgment effect similar to that of a human in real-world scenarios. The robot first uses a visual module to gain an overall perception of the environment, used to determine the search area. Then, the auditory module determines the location of the personnel to be searched, determining the robot's direction of travel. Simultaneously, it recognizes and controls voice commands. Upon reaching a general location, it combines visual information with the tactile module to explore the environment and determine the approximate outline of contact surfaces. Afterward, the olfactory module assesses vital signs and collects and analyzes ambient gases to detect hazardous materials. This comprehensive environmental assessment avoids the perceptual biases caused by single sensors and provides a more complete understanding of the task environment, facilitating subsequent search and rescue operations. Attached Figure Description

[0082] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0083] Figure 1 This is a schematic diagram of a biomimetic search robot module and system based on multiple senses (sight, hearing, touch, smell) proposed in one embodiment.

[0084] Figure 2 This is a schematic diagram of the operation of the robot system proposed in one embodiment;

[0085] Figure 3 Here is a flowchart of the vision module proposed in one embodiment;

[0086] Figure 4 This is a block diagram of the auditory module proposed in one embodiment;

[0087] Figure 5 This is a block diagram of a tactile module proposed in one embodiment;

[0088] Figure 6 Here is a block diagram of the olfactory module proposed in one embodiment;

[0089] Figure 7 This is a block diagram of an information fusion module proposed in one embodiment;

[0090] Figure 8 Here is a block diagram of a robotic bomb disposal task proposed in one embodiment;

[0091] Figure 9 This is a block diagram of a robot search and rescue mission proposed in one embodiment. Detailed Implementation

[0092] To better understand the technical solution of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0093] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0094] This invention aims to overcome the limitations of existing target-searching robots in terms of multiple senses and intelligent discrimination. This invention provides a search robot with multi-sensory biomimetic capabilities encompassing sight, hearing, touch, and smell.

[0095] like Figure 1 As shown, in one embodiment, a multi-sensory fusion bionic search robot with human-machine two-way interaction function is provided. This robot can be used in various application fields such as hazard source search and explosive detection. Because it has multiple sensing methods and information fusion modules, it can make comprehensive judgments based on multiple information such as shape, smell, and sound in complex environments, avoiding ignoring some environmental information or misjudging, and facilitating the implementation of search and rescue and bomb disposal tasks.

[0096] The multi-sensory fusion bionic search robot with human-machine two-way interaction function includes: a robot body and a multi-modal perception fusion decision module set on the robot body;

[0097] The robot body includes a robot mechanical part, a motor drive module connected to the robot mechanical part, and a bionic motion control module that is communicatively connected to the motor drive module.

[0098] In one embodiment, the robot's mechanical part is a robot platform, which is a vehicle-arm system. The vehicle-arm system includes a vehicle body (mobile robot chassis) and a robotic arm, as shown in the schematic diagram of its structure and operation process. Figure 2 As shown, the vehicle body is an autonomous mobile vehicle with a two-wheel differential chassis, capable of autonomous navigation and precise tracking of predetermined trajectories. The robotic arm is a six-degree-of-freedom collaborative robotic arm with torque feedback, possessing precise position control to meet the precise control requirements of search and rescue missions. The robotic arm also features torque feedback, which can detect collisions upon contact with people during rescue missions, preventing secondary injuries to the target. This vehicle-arm system achieves collaborative operation through ROS node communication. Combined with an information fusion module, it can autonomously explore the environment.

[0099] The vehicle uses an auditory module to determine the approximate direction of the target object, and combines this with a visual module to identify the object, thus defining a search area. Upon reaching the search location, a tactile and olfactory module is mounted at the end of the robotic arm. For areas indistinguishable by vision, the robot uses touch to explore the environment, acquiring its outline and providing information for the search path. Simultaneously, the olfactory module at the end of the robotic arm can be brought close to collect odor information. This odor detection allows for the early detection of human vital signs, explosives, and hazardous gases in the environment, providing prior information for subsequent search and rescue operations.

[0100] The multimodal perception fusion decision module includes: a vision module, an auditory module, a tactile module, an olfactory module, and an information fusion module that is communicatively connected to the vision module, auditory module, tactile module, olfactory module, and bionic motion control module;

[0101] The vision module is used to capture visual information of the surrounding environment of the bionic search robot, and the vision module uses a spiking neural network to identify and classify the visual information of the surrounding environment and generate pulse signals containing the path planning of the bionic search robot to move to the target object.

[0102] In one embodiment, the vision module includes a scene visual data capture unit, an event stream spatiotemporal feature extraction and encoding unit, an event stream pulse pooling unit, an event stream temporal resolution adjustment and integration unit, and a deep spiking neural network (SNN) learning spatiotemporal features and recognition unit. For details, please refer to... Figure 3 ;

[0103] Among them, the scene visual data capture unit uses a dynamic visual sensor to read dynamic information in the scene in real time and converts it into an event stream in address event expression format;

[0104] The event stream spatiotemporal feature extraction and encoding unit extracts and encodes spatiotemporal features from the event stream obtained after processing by the captured scene visual data unit, so as to transform visual information into feature codes of interest.

[0105] The event stream pulse pooling unit performs pulse pooling on the event stream obtained after processing by the event stream spatiotemporal feature extraction and encoding unit.

[0106] The event stream temporal resolution adjustment and integration unit adjusts and integrates the event stream temporal resolution obtained after processing by the event stream pulse pooling unit;

[0107] The Deep Spike Neural Network (SNN) learns spatiotemporal features and the recognition unit uses the event stream, which has been processed by the event stream temporal resolution adjustment and integration unit, to learn spatiotemporal features in order to identify the target object.

[0108] The method for using the vision module in a multi-sensory fusion bionic search robot with human-computer two-way interaction capabilities includes the following steps:

[0109] Step 1: Use the Dynamic Vision Sensor (DVS) in the scene visual data capture unit to read dynamic information in the scene in real time and convert it into an event stream S in Address Event Representation (AER) format. tr ={e i |e i =[x i ,y i ,t i ,p i ] T ,i∈[1,m]};

[0110] Where m is the total number of events, and e is each event in the event stream. i It has four pieces of information, namely the spatial coordinate x i ∈[1,...,H],y i ∈[1,...,W], where H and W are the resolutions of the event camera, and the timestamp t i and polarity p i ∈{-1,1}, where -1 and 1 represent "OFF" and "ON" events respectively. An "ON" event means that the light intensity increases, and an "OFF" event means that the light intensity decreases.

[0111] Step 2: The event stream obtained after being processed by the scene visual data unit in Step 1 is subjected to spatiotemporal feature extraction and encoding using the event stream spatiotemporal feature extraction and encoding. This can be divided into a spatiotemporal event plane calculation stage and an encoding stage.

[0112] The spatiotemporal event plane computation phase involves using a spatiotemporal event plane descriptor to calculate the correlation between events within a spatiotemporal neighborhood. This descriptor includes a spatial correlation kernel function. A time-related kernel function For any event e i The mathematical expression for the spacetime event plane is as follows:

[0113]

[0114]

[0115]

[0116] Wherein, spatial coordinate x j ∈[1,...,H],y j ∈[1,...,W], u is the proportionality coefficient controlling the temporal and spatial correlation, σ l and τ l These are the spatial standard deviation and temporal decay constant of the spatiotemporal plane, with the subscript 'l' indicating the l-th feature coding layer; A i It is event e i Spatiotemporal event plane description operator, spatial correlation kernel function Focus on e i Centered on R l This function represents the spatial relationships of events within a rectangular neighborhood of a radius; it assigns greater weight to events that are spatially closer. (Temporal correlation kernel function) Focus on the latest events that occur within the time window τ. l The temporal relationship of internal events, i.e., satisfying t j =max(t(x) j ,y j )),t j∈[t i -τ l ,t i This function assigns smaller weights to historical events that occurred earlier, assuming that those events have a smaller temporal correlation with the current event. Therefore, the greater the spatiotemporal correlation between the current event and events in its spatiotemporal neighborhood, the greater the weight it is assigned. This is beneficial for more accurately describing the spatiotemporal dynamic changes of the event flow.

[0117] Encoding phase: First N l The spatiotemporal surface of each event is used to initialize N. l Feature template:

[0118]

[0119] After initialization, these templates are updated using an online clustering algorithm for each event e. i Its spatiotemporal surface is obtained through the spatiotemporal event plane calculation stage, and is calculated using the following formula. Euclidean distance between each template:

[0120]

[0121] and The closest k-th template will be event e i Encoded as e i =[x i ,y i ,t i ,k] T Meanwhile, the k-th template will be updated using an online clustering algorithm:

[0122]

[0123]

[0124] Where, n k The number of events encoded as the k-th feature is α, which is the sum of n. k The control update magnitude, i.e., the smaller the number of events encoded as the k-th feature, the larger the corresponding template update magnitude, β represents. and Cosine correlation between them;

[0125] Two-layer feature encoding is used to extract features at different spatial and temporal scales in the event stream. The first layer uses small-scale spatiotemporal parameters τ1 and R1 to encode fine-grained features, while the second layer uses larger-scale spatiotemporal parameters τ2 and R2 to encode more complex features across spatiotemporal distances.

[0126] Step 3: The event stream obtained after being processed by the event stream spatiotemporal feature extraction and encoding unit in step 2 is pulse-pooled using the event stream pulse pooling unit. This operation includes a spatial pooling stage and a refractory period stage.

[0127] The spatial pooling stage involves preserving the local key features of events through spatial pooling. Specifically, it first encodes events e within adjacent, non-overlapping r×r pixel regions. i =[x i ,y i ,t i ,k i ] T Perform space pooling, that is:

[0128] e i =[x i / r,y i / r,t i ,k i ] T

[0129] Where r is the pooling radius;

[0130] Refractory period: A spiking neuron receives spatially pooled events. This neuron operates with a low threshold and refractory period to ensure that a single input is sufficient to trigger an output pulse and to remain silent for a period of time. Therefore, any input event will trigger the neuron outside the refractory period to pulse, and then the neuron will remain silent until the refractory period t. ref Finally, only events that cause neurons to fire are transmitted to the spiking neural network for classification.

[0131] Step 4: The event stream temporal resolution adjustment and integration unit adjusts and integrates the event stream temporal resolution obtained after processing by the event stream pulse pooling unit in Step 3. Currently, deep spiking neural networks (SNNs) cannot directly process microsecond-level event stream data. A time collapse strategy is adopted to adjust the temporal resolution of the event stream and aggregate events with the same time. This strategy first divides the event stream into multiple time slices, then assumes that events within a time slice have the same timestamp, sets the position with an event input to 1, and otherwise sets it to 0. Therefore, the event stream within a slice is transformed into a binary input S of one time step. t The formula is expressed as follows:

[0132] S t (y i ,x i ,k i ) = 1

[0133] Among them, y i ,x i ,k iFor the event flow within the t-th time slice {e i |t i The coordinates and number of feature channels of ∈[α×t,α×(t+1)-1]};

[0134] Therefore, we can obtain t (H,W,N) l A binary pulse vector of size 1 is used as the input to the spiking neural network at each time step;

[0135] Step 5: The deep spiking neural network (SNN) is used to learn spatiotemporal features and the recognition unit inputs the event stream obtained after the event stream time resolution adjustment and integration unit in step 5 into the SNN, and the deep SNN is trained using the spatiotemporal backpropagation (STBP) algorithm. This step is divided into the SNN feedforward calculation stage and the backpropagation training stage.

[0136] SNN feedforward computation stage: The discrete-time iterative LIF neuron model used in SNN is expressed by the following formulas for membrane potential update and pulse firing dynamics:

[0137]

[0138]

[0139]

[0140] Where λ is the membrane potential decay coefficient of the neuron, V thr It is the membrane potential threshold. This represents the pulse firing state controlled by the step function g(x). When the membrane potential exceeds the threshold, the neuron fires a pulse to pass to the next layer and resets the membrane potential at the next moment. The specific formula for calculating postsynaptic potential updates is as follows:

[0141]

[0142] in, This represents the synaptic weight between the i-th postsynaptic neuron in the l-th layer and the j-th presynaptic neuron in the previous layer; It is the pulse output of the j-th presynaptic neuron in layer l-1 at time t+1;

[0143] Backpropagation training phase: Due to the non-differentiable nature of pulse firing, STBP uses a rectangular function to replace the derivative of the neuron's output pulse with respect to the membrane potential, thereby supporting the backpropagation of errors.

[0144]

[0145] Where a1 is a hyperparameter that determines the gradient width;

[0146] Finally, the loss function of STBP is set to minimize the mean squared error of all samples under a given time window T:

[0147]

[0148] where y s and Let represent the label vector and the output impulse vector of the s-th training sample in the last layer, respectively.

[0149] The auditory module is used to receive voice commands with target tags; and the auditory module uses a spiking neural network to identify the direction of the target relative to the bionic search robot and generate a pulse signal containing directional information.

[0150] like Figure 4 As shown, in one embodiment, the auditory module is used to receive a voice command with a target object tag; and the auditory module uses a spiking neural network to identify the direction of the target object relative to the bionic search robot and generates a pulse signal containing the direction information;

[0151] The specific auditory module features a ring-shaped microphone array that transmits data via UART serial communication. The microphone array estimates the angle of the sound source by calculating the time difference of sound wave arrival. Simultaneously, the raw data from the microphone array is recorded as auditory information. This auditory information is encoded into pulse signals using a biomimetic pulse coder and input into a biomimetic spiking neural network. This network consists of four layers of biomimetic neurons, modeled as a LIF neuron model, enabling the recognition of input speech and thus controlling the movement of the mobile robot (including forward and backward movements). The encoded auditory information is also input into the information fusion module to perform multimodal recognition tasks.

[0152] like Figure 5 As shown, the tactile module is used to acquire at least a portion of the contour information and contour friction coefficient information of the target object; and the tactile module uses a pulse neural network to process the joint errors caused by the inverse kinematics of the robotic arm, so that the robotic arm can acquire surface contour information more accurately to generate a pulse signal containing the contour information and contour friction coefficient information of the target object.

[0153] In one embodiment, the tactile module adopts a hand-like structure, integrating a force sensor at the fingertip to acquire the three-dimensional force intensity at the contact point. This acquired force intensity provides basic information for generating the contact surface contour. When the tactile module explores the environment, it can acquire the contact force between the end of the tactile module and the environment at that point, and based on the force feedback information F... ext And action X generated by the basal ganglia dThis method reconstructs contours over a specific time series, providing Cartesian spatial position information for the admittance control of the robotic arm in a biomimetic search robot. Furthermore, for robotic arms equipped with tactile modules, a spiking neural network can be used to correct the inverse kinematics model of the robotic arm, thus improving the Cartesian spatial pose information provided by the robotic arm. d It is necessary to convert the inverse kinematics model of the robotic arm to the joint space q. d However, there will be certain deviations between the actual structure of the robotic arm and the inverse kinematics model, resulting in discrepancies in the inverse kinematics calculations of the robotic arm. The joint spatial position may not be accurate, thus affecting control precision. However, the contact force information acquired by the haptic module can correct for this error and compensate for its impact. This is achieved by using the contact force information acquired through haptic feedback. d And from the actual position x reached by the robotic arm after its movement, the position error e in Cartesian space can be calculated. x By inputting this value into the SNN-like cerebellar controller, a Δq can be trained to compensate for the error of the inverse kinematics model. d , will Δq d By inputting the results into the inverse kinematics model solution, the model error can be compensated.

[0154] Specifically, the haptic module integrates a pressure sensor array at the fingertip, with each module containing 16 sensors. These sensors can independently measure the force values ​​along the x, y, and z axes in space. A built-in temperature sensor compensates for the impact of temperature drift on force measurement. The haptic module employs a flexible design, allowing it to contact fragile or soft objects without damaging their surfaces. Its flexibility also ensures high elasticity of the sensor surfaces, making them less prone to damage in complex environments. When the haptic module comes into contact with the external environment, the sensed contact force is fed back to the robotic arm in real time. This external force feedback ensures that the robotic arm maintains a constant force when in contact with the environment, and simultaneously, the environmental contour is obtained based on the location information of the haptic module.

[0155] like Figure 6 As shown, the olfactory module is used to acquire odor information emitted by the target object; and the olfactory module uses a spiking neural network to determine the type of odor emitted by the target object and generates a pulse signal containing the type of odor emitted by the target object.

[0156] In one embodiment, the olfactory module has an electronic nose array with multiple MOS sensors. The olfactory encoding part adopts a three-layer spiking neural network architecture based on the recorded electronic nose array data. The third layer provides negative feedback to the second layer, and the pulse of the second layer is output as the encoded result. The odor signal is converted into a pulse signal, and the pulse signal is input into the information fusion module to realize the multimodal recognition task.

[0157] Specifically, the olfactory module features an electronic nose array of twelve olfactory sensors arranged in four groups of three types. It connects to the main control unit via UART to collect odor information from the air in real time. The gas information is pulse-coded using a designed biomimetic olfactory bulb network. The signal amplitude is used as input, employing a three-layer spiking neural network architecture. The first and second layers are excitatoryly connected, the second and third layers are excitatoryly connected, and the third layer is inhibitoryly connected to the second layer, with fixed weights. The signal from the second layer is used as the output pulse signal. The neurons utilize the Izhikevich neuron model. The encoded pulse signal is then input to the information fusion module to perform subsequent downstream tasks.

[0158] like Figure 7 As shown, a multimodal information fusion algorithm is introduced into the information fusion module to better handle the information input from the four modules: vision, hearing, touch, and smell. First, image information from the vision module and sound source direction information and voice control information from the hearing module are input to the motor drive module. This integrated information is connected to the bionic motion control module for the search robot's motion control strategy. The touch module provides external object information, and the olfactory module provides odor information. Image and sound information are integrated with touch and olfactory information. After processing, the signals from all four modalities are encoded into pulse signals.

[0159] Specifically, the information fusion module receives pulse signals generated by one or more modules from the vision, hearing, touch, and olfaction modules. It integrates these pulse signals through time-dimensional alignment and processes them using a spiking neural network (involving object recognition and analysis) to obtain a decision result. This decision result is then sent to the bionic motion control module, which in turn sends motor motion parameters to the motor drive module. The motor drive module then controls the robot's mechanical movements. Through this process, the robot can gain a more comprehensive understanding of objects in its environment, enabling more accurate decisions and responses. Furthermore, by introducing a multimodal information fusion algorithm, the robot can more intelligently perceive and interpret its environment, improving system performance and adaptability, thus better enabling it to complete various tasks.

[0160] In one embodiment, the vision module, hearing module, tactile module, olfactory module, and bionic motion control module are each communicatively connected to a terminal. The terminal can be operated by an operator, who receives information sent by at least one of the vision, hearing, tactile, and olfactory modules through the terminal and issues control commands to the bionic motion control module.

[0161] like Figure 8As shown, in one embodiment, during bomb disposal, the vision module generates a pulse signal containing a path plan for a bionic search robot to move toward the suspected explosive based on the captured location information of the suspected explosive.

[0162] The information fusion module receives the pulse signal from the vision module and sends the pulse signal directly to the bionic motion control module. The bionic motion control module sends motor motion parameter information to the motor drive module, and the motor drive module controls the mechanical parts of the robot to move toward the suspected explosive.

[0163] The auditory module receives sound information emitted by the suspected explosive and uses a spiking neural network to generate a pulse signal relative to the direction of the target object in the bionic search robot. The olfactory module obtains odor information emitted by the suspected explosive and uses a spiking neural network to generate a pulse signal containing the type of odor emitted by the suspected explosive.

[0164] When a suspected explosive device is visible, the vision module uses a spiking neural network to generate pulse signals representing the appearance features of the suspected explosive device. The information fusion module receives the pulse signals generated by the hearing and olfactory modules, as well as the pulse signals representing the appearance features of the suspected explosive device generated by the vision module. It integrates these pulse signals by aligning them in the time dimension and processes them through the spiking neural network to obtain a decision result. The decision result involves the location and type of the suspected explosive device. The information containing the decision result is then sent to the terminal. If the suspected explosive device is indeed an explosive device, the operator decides to dismantle it. The operator sends control information for dismantling the explosive device to the bionic motion control module through the terminal. This causes the bionic motion control module to send motor motion parameter information to the motor drive module, which in turn controls the robotic arm in the robot's mechanical part to dismantle the explosive device.

[0165] When the suspected explosive device is invisible, the information fusion module receives pulse signals generated by the auditory and olfactory modules, integrates the pulse signals by aligning them in the time dimension, and processes them through a spiking neural network to obtain a decision result. The decision result involves the location and type of the suspected explosive device, and sends the information containing the decision result to the terminal. If the suspected explosive device is an explosive device, the operator decides to dismantle it. The operator sends control information for dismantling the explosive device to the bionic motion control module through the terminal, which in turn sends motor motion parameter information to the motor drive module. The motor drive module then controls the robotic arm in the robot's mechanical part to extend to the location of the explosive device and dismantle it.

[0166] like Figure 9 As shown, in one embodiment, during search and rescue, the vision module generates a pulse signal containing a path plan for the bionic search robot to move to the object to be searched based on the captured position information of the object to be searched;

[0167] The information fusion module receives the pulse signal from the vision module and sends the pulse signal directly to the bionic motion control module. The bionic motion control module sends motor motion parameter information to the motor drive module, and the motor drive module controls the mechanical parts of the robot to move to the object to be searched and rescued.

[0168] The auditory module receives sound information emitted by the object to be searched and uses a spiking neural network to generate pulse signals of the object relative to the direction of the bionic search robot.

[0169] The tactile module acquires at least a portion of the contour information and contour friction coefficient information of the object to be searched and rescued, and uses a pulse neural network to generate a pulse signal containing the contour information and contour friction coefficient information of the object to be searched and rescued.

[0170] The olfactory module acquires odor information emitted by the object to be searched and uses a spiking neural network to generate a pulse signal containing the type of odor emitted by the object to be searched;

[0171] When the object to be rescued is visible, the vision module uses a spiking neural network to generate pulse signals of the object's appearance features. The information fusion module receives the pulse signals generated by the hearing and olfactory modules, as well as the pulse signals of the object's appearance features generated by the vision module. It integrates the pulse signals by aligning them in the time dimension and processes them through the spiking neural network to obtain a decision result. The decision result involves the location and type of the object to be rescued, and sends the information containing the decision result to the terminal. If the object to be rescued is a target object, the operator decides to rescue the target object. The operator sends control information for rescuing the target object to the bionic motion control module through the terminal, which in turn sends motor motion parameter information to the motor drive module. The motor drive module then controls the robotic arm in the robot's mechanical part to rescue the target object.

[0172] When the object to be rescued is invisible, the auditory module receives the voice information emitted by the object and uses a spiking neural network to generate pulse signals containing the voice information. The information fusion module receives the pulse signals generated by the auditory and olfactory modules, as well as the pulse signals containing the voice information generated by the tactile module. It integrates the various pulse signals by aligning them in the time dimension and processes them through the spiking neural network to obtain a decision result. The decision result involves the location and type of the object to be rescued, and the information containing the decision result is sent to the terminal. If the object to be rescued is the target object, the operator decides to rescue the target object. The operator sends control information for rescuing the target object to the bionic motion control module through the terminal, so that the bionic motion control module sends motor motion parameter information to the motor drive module. Then, the motor drive module controls the robotic arm in the robot's mechanical part to extend to the position of the target object and rescue it.

[0173] This invention addresses the problems of existing search robots, such as relying on a single sensory component and lacking autonomous exploration capabilities. It develops a biomimetic robot that combines visual, auditory, tactile, and olfactory perception methods. In complex environments, it can comprehensively assess these four senses for judgment, and its perception and control methods, based on spiking neural networks, more closely resemble human judgment. In real-world scenarios, it can achieve a comprehensive perception and judgment similar to that of a human. The robot first uses a visual module to gain an overall perception of the environment, used to determine the search area. Then, the auditory module determines the location of the personnel being searched, determining the robot's direction of travel. Simultaneously, it recognizes and controls voice commands. Upon reaching a general location, it combines visual information with its tactile module to explore the environment and determine the approximate outline of contact surfaces. Afterward, the olfactory module assesses vital signs and collects and analyzes ambient gases to detect hazardous materials. This comprehensive environmental assessment avoids the perceptual biases of single sensors and provides a more complete understanding of the mission environment, facilitating subsequent search and rescue operations.

[0174] The above are merely preferred embodiments of one or more embodiments of this specification and are not intended to limit the scope of one or more embodiments of this specification. Any modifications made within the spirit and principles of one or more embodiments of this specification are permitted.

Claims

1. A multi-sensory fusion bionic search robot with two-way human-machine interaction capabilities, characterized in that: include: The robot body and the multimodal perception fusion decision module mounted on the robot body; The robot body includes a robot mechanical part, a motor drive module connected to the robot mechanical part, and a bionic motion control module communicatively connected to the motor drive module. The multimodal perception fusion decision module includes: an olfactory module, a visual module, an auditory module, a tactile module, and an information fusion module that is communicatively connected to the visual module, the auditory module, and the tactile module; The vision module is used to capture visual information of the surrounding environment of the bionic search robot, and the vision module uses a spiking neural network to identify and classify the visual information of the surrounding environment and generate a pulse signal containing the path planning of the bionic search robot to move to the target object; The auditory module is used to receive voice commands with target object tags; and the auditory module uses a spiking neural network to identify the direction of the target object relative to the bionic search robot and generates a pulse signal containing the direction information; The tactile module is used to acquire at least a portion of the contour information and contour friction coefficient information of the target object; and the tactile module uses a spiking neural network to process the at least a portion of the contour information and contour friction coefficient information of the target object to generate a pulse signal containing the contour information and contour friction coefficient information of the target object. The olfactory module is used to acquire odor information emitted by the target object; and the olfactory module uses a spiking neural network to determine the type of odor emitted by the target object and generates a pulse signal containing the type of odor emitted by the target object. The information fusion module receives pulse signals generated by one or more of the vision module, hearing module, tactile module, and olfactory module, integrates the pulse signals by aligning them in a time dimension, and processes them through a spiking neural network to obtain a decision result. The information containing the decision result is then sent to the bionic motion control module, which sends motor motion parameter information to the motor drive module, thereby controlling the movement of the robot's mechanical parts. The vision module includes a scene visual data capture unit, an event stream spatiotemporal feature extraction and encoding unit, an event stream pulse pooling unit, an event stream temporal resolution adjustment and integration unit, and a deep spiking neural network (SNN) learning spatiotemporal features and recognition unit. The scene visual data capture unit uses a dynamic visual sensor to read dynamic information in the scene in real time and converts it into an event stream in address event expression format; The event stream spatiotemporal feature extraction and encoding unit extracts and encodes spatiotemporal features from the event stream obtained after processing by the captured scene visual data unit, so as to transform visual information into feature codes of interest. The event stream pulse pooling unit performs pulse pooling on the event stream obtained after processing by the event stream spatiotemporal feature extraction and encoding unit. The event stream temporal resolution adjustment and integration unit adjusts and integrates the event stream temporal resolution obtained after processing by the event stream pulse pooling unit; The deep spiking neural network (SNN) learning spatiotemporal features and recognition unit uses the event stream obtained after processing by the event stream temporal resolution adjustment and integration unit to learn spatiotemporal features in order to identify the target object; The dynamic vision sensor in the scene visual data capture unit reads dynamic information in the scene in real time and converts it into an event stream S in address event expression format. tr ={e i |e i =[x i ,y i ,t i ,p i ] T ,i∈[1,m]}; Where m is the total number of events, and e is each event in the event stream. i It has four pieces of information, namely the spatial coordinates x i ∈[1,…,H],y i ∈[1,…,W], where H and W are the resolutions of the event camera, and the timestamp t i and polarity p i ∈{-1,1}, where -1 and 1 represent "OFF" and "ON" events respectively. An "ON" event means that the light intensity increases, and an "OFF" event means that the light intensity decreases. The event stream spatiotemporal feature extraction and encoding unit extracts and encodes the event stream obtained after processing by the captured scene visual data unit. It can be divided into a spatiotemporal event plane calculation stage and an encoding stage. The spatiotemporal event plane computation phase involves using a spatiotemporal event plane descriptor to calculate the correlation between events within a spatiotemporal neighborhood. This descriptor includes a spatial correlation kernel function. A time-related kernel function For any event e i The mathematical expression for the spacetime event plane is as follows: Wherein, spatial coordinate x j ∈[1,…,H],y j ∈[1,...,W], u is the proportionality coefficient controlling the temporal and spatial correlation, σ l and τ l These are the spatial standard deviation and temporal decay constant of the spatiotemporal plane, with the subscript 'l' indicating the l-th feature coding layer; A i It is event e i Spatiotemporal event plane description operator, spatial correlation kernel function Focus on e i Centered on R l This function represents the spatial relationships of events within a rectangular neighborhood of a radius; it assigns greater weight to events that are spatially closer. (Temporal correlation kernel function) Focus on the latest events that occur within the time window τ. l The temporal relationship of internal events, i.e., satisfying t j =max(t(x) j ,y j )),t j ∈[t i -τ l ,t i This function assigns smaller weights to historical events that occurred earlier, assuming that those events have a smaller temporal correlation with the current event. Therefore, the greater the spatiotemporal correlation between the current event and events in its spatiotemporal neighborhood, the greater the weight it is assigned. This is beneficial for more accurately describing the spatiotemporal dynamic changes of the event flow. Encoding phase: First N l The spatiotemporal surface of each event is used to initialize N. l Feature template: After initialization, these templates are updated using an online clustering algorithm for each event e. i Its spatiotemporal surface is obtained through the spatiotemporal event plane calculation stage, and is calculated using the following formula. Euclidean distance between each template: and The closest k-th template will be event e i Encoded as e i =[x i ,y i ,t i ,k] T Meanwhile, the k-th template will be updated using an online clustering algorithm: Where, n k The number of events encoded as the k-th feature is α, which is the sum of n. k The control update magnitude, i.e., the smaller the number of events encoded as the k-th feature, the larger the corresponding template update magnitude, β represents. and Cosine correlation between them; Two-layer feature encoding is used to extract features at different spatial and temporal scales in the event stream. The first layer uses small-scale spatiotemporal parameters τ1 and R1 to encode fine-grained features, while the second layer uses larger-scale spatiotemporal parameters τ2 and R2 to encode more complex features across spatiotemporal distances. The event stream pulse pooling unit performs pulse pooling on the event stream obtained after processing by the event stream spatiotemporal feature extraction and encoding unit. This operation includes a spatial pooling stage and a refractory period stage. The spatial pooling stage involves preserving the local key features of events through spatial pooling. Specifically, it first encodes events e within adjacent, non-overlapping r×r pixel regions. i =[x i ,y i ,t i ,k i ] T Perform space pooling, that is: e i =[x i / r,y i / r,t i ,k i ] T Where r is the pooling radius; Refractory period: A spiking neuron receives spatially pooled events. This neuron operates with a low threshold and refractory period to ensure that a single input is sufficient to trigger an output pulse and to remain silent for a period of time. Therefore, any input event will trigger the neuron outside the refractory period to pulse, and then the neuron will remain silent until the refractory period t. ref Finally, only the events that cause neurons to fire are transmitted to the spiking neural network for classification; The event stream temporal resolution adjustment and integration unit adjusts and integrates the event stream temporal resolution obtained after processing by the event stream pulse pooling unit. Currently, deep spiking neural networks (SNNs) cannot directly process microsecond-level event stream data. Therefore, a time collapse strategy is used to adjust the event stream temporal resolution and aggregate events with the same time. This strategy first divides the event stream into multiple time slices, then assumes that events within a time slice have the same timestamp, setting the position with an event input to 1 and otherwise to 0. Thus, the event stream within a slice is transformed into a binary input S for one time step. t The formula is expressed as follows: S t (y i ,x i ,k i )=1 Among them, y i ,x i ,k i For the event flow within the t-th time slice {e i |t i The coordinates and number of feature channels of ∈[α×t,α×(t+1)-1]}; Therefore, we can obtain t (H,W,N) l A binary pulse vector of size 1 is used as the input to the spiking neural network at each time step; The deep spiking neural network (SNN) learning spatiotemporal features and recognition unit inputs the event stream obtained after processing by the event stream temporal resolution adjustment and integration unit into the SNN, and trains the deep SNN using the spatiotemporal backpropagation (STBP) algorithm. This step is divided into the SNN feedforward calculation stage and the backpropagation training stage. SNN feedforward computation stage: The discrete-time iterative LIF neuron model used in SNN is expressed by the following formulas for membrane potential update and pulse firing dynamics: Where λ is the membrane potential decay coefficient of the neuron, V thr It is the membrane potential threshold. This represents the pulse firing state controlled by the step function g(x). When the membrane potential exceeds the threshold, the neuron fires a pulse to pass to the next layer and resets the membrane potential at the next moment. The specific formula for calculating postsynaptic potential updates is as follows: in, This represents the synaptic weight between the i-th postsynaptic neuron in the l-th layer and the j-th presynaptic neuron in the previous layer; It is the pulse output of the j-th presynaptic neuron in layer l-1 at time t+1; Backpropagation training phase: Due to the non-differentiable nature of pulse firing, STBP uses a rectangular function to replace the derivative of the neuron's output pulse with respect to the membrane potential, thereby supporting the backpropagation of errors. Where a1 is a hyperparameter that determines the gradient width; Finally, the loss function of STBP is set to minimize the mean squared error of all samples under a given time window T: Where y s and Let represent the label vector and the output impulse vector of the s-th training sample in the last layer, respectively.

2. The multi-sensory fusion bionic search robot with human-machine two-way interaction function according to claim 1, characterized in that: The olfactory module, visual module, auditory module, and tactile module, as well as the bionic motion control module, are all bidirectionally connected to the operator via a terminal. The terminal can be operated by the operator, who receives information from at least one of the olfactory module, visual module, auditory module, and tactile module through the terminal and issues control commands to the bionic motion control module.

3. The multi-sensory fusion bionic search robot with human-machine two-way interaction function according to claim 1, characterized in that: For different search targets, multiple modules generate pulse signals containing the path planning of the multi-sensory fusion bionic search robot to move to the suspected explosive or rescuer based on the captured location information of the suspected explosive or rescuer, and combine them with the information fusion module to generate response control strategies for different scenarios; In the explosives search scenario, the information fusion module receives the pulse signal from the vision module and sends the pulse signal from the vision module directly to the bionic motion control module. The bionic motion control module sends motor motion parameter information to the motor drive module, and the motor drive module controls the mechanical part of the robot to move to the suspected explosives. The auditory module receives the sound information emitted by the suspected explosive and uses a spiking neural network to generate a pulse signal of the target object relative to the direction of the bionic search robot. The olfactory module obtains the odor information emitted by the suspected explosive and uses a spiking neural network to generate a pulse signal containing the type of odor emitted by the suspected explosive. When the suspected explosive device is in a visible state, the vision module uses a spiking neural network to generate pulse signals of the appearance features of the suspected explosive device; The information fusion module receives pulse signals generated by the auditory and olfactory modules, as well as pulse signals representing the appearance features of the suspected explosive device generated by the visual module. It integrates these pulse signals through time-dimensional alignment and processes them using a spiking neural network to obtain a decision result. This decision result relates to the location and type of the suspected explosive device. The information containing the decision result is then sent to the terminal. If the suspected explosive device is indeed an explosive device, the operator decides to dismantle it. The operator sends control information for dismantling the explosive device to the bionic motion control module via the terminal. This causes the bionic motion control module to send motor motion parameter information to the motor drive module, which in turn controls the robotic arm in the robot's mechanical part to dismantle the explosive device. When the suspected explosive device is invisible, the information fusion module receives pulse signals generated by the auditory and olfactory modules, integrates the pulse signals through time-dimensional alignment, and processes them through a spiking neural network to obtain a decision result. The decision result involves the location and type of the suspected explosive device, and the information containing the decision result is sent to the terminal. If the suspected explosive device is an explosive device, the operator decides to dismantle it. The operator sends control information for dismantling the explosive device to the bionic motion control module through the terminal, causing the bionic motion control module to send motor motion parameter information to the motor drive module. Then, the motor drive module controls the robotic arm in the robot's mechanical part to extend to the location of the explosive device and dismantle it. In a search and rescue scenario, the vision module generates a pulse signal containing the path planning of the bionic search robot moving to the object to be searched, based on the captured position information of the object to be searched. The information fusion module receives the pulse signal from the vision module and sends the pulse signal from the vision module directly to the bionic motion control module. The bionic motion control module sends motor motion parameter information to the motor drive module, and the motor drive module controls the mechanical parts of the robot to move to the object to be searched and rescued. The auditory module receives the sound information emitted by the object to be searched and rescued and uses a pulse neural network to generate a pulse signal of the object to be searched and rescued relative to the direction of the bionic search robot; The tactile module acquires at least a portion of the contour information and contour friction coefficient information of the object to be searched and rescued, and uses a pulse neural network to generate a pulse signal containing the contour information and contour friction coefficient information of the object to be searched and rescued. The olfactory module acquires the odor information emitted by the object to be searched and rescued and uses a spiking neural network to generate a pulse signal containing the type of odor emitted by the object to be searched and rescued; When the object to be rescued is visible, the vision module uses a spiking neural network to generate pulse signals of the object's appearance features. The information fusion module receives the pulse signals generated by the hearing and olfactory modules, as well as the pulse signals of the object's appearance features generated by the vision module. It integrates the pulse signals by aligning them in the time dimension and processes them through the spiking neural network to obtain a decision result. The decision result involves the location and type of the object to be rescued. The information containing the decision result is sent to the terminal. If the object to be rescued is a target object, the operator decides to rescue the target object. The operator sends control information for rescuing the target object to the bionic motion control module through the terminal. This causes the bionic motion control module to send motor motion parameter information to the motor drive module, which then controls the robotic arm in the robot's mechanical part to rescue the target object. When the object to be rescued is invisible, the auditory module receives the voice information emitted by the object and uses a spiking neural network to generate a pulse signal containing the voice information. The information fusion module receives the pulse signals generated by the auditory module, the olfactory module, and the tactile module, and integrates the pulse signals containing the voice information by aligning them in a time dimension and then processes them through the spiking neural network to obtain a decision result. The decision result involves the location and type of the object to be rescued, and the information containing the decision result is sent to the terminal. If the object to be rescued is a target object, the operator decides to rescue the target object. The operator sends control information for rescuing the target object to the bionic motion control module through the terminal, causing the bionic motion control module to send motor motion parameter information to the motor drive module. Then, the motor drive module controls the robotic arm in the robot's mechanical part to extend to the location of the target object and rescue it.

4. The multi-sensory fusion bionic search robot with human-machine two-way interaction function according to claim 1, characterized in that: The auditory module has a ring-shaped microphone array. It estimates the angle information of the sound source by calculating the time difference of sound wave arrival and records the raw data of the microphone array as auditory information. The auditory information is encoded into pulse signals through biomimetic pulse coding and input into a biomimetic spiking neural network. The biomimetic spiking neural network consists of four layers of biomimetic neurons to recognize the input speech, thereby controlling the movement of the mobile robot. At the same time, the encoded auditory information is input into the information fusion module to realize the multimodal recognition task.

5. The multi-sensory fusion bionic search robot with human-machine two-way interaction function according to claim 1, characterized in that: The tactile module adopts a hand-like structure, integrating a force sensor at the top of the finger to acquire the three-dimensional force intensity at the contact point. The acquired force intensity can provide basic information for generating the contact surface contour information. When the tactile module explores the environment, it can acquire the contact force between the end of the tactile module and the environment at that point. Based on the force feedback information, the contour on a certain time series is reconstructed, providing contour information for the robotic arm in the bionic search robot.

6. The multi-sensory fusion bionic search robot with human-machine two-way interaction function according to claim 1, characterized in that: The olfactory module has an electronic nose array with multiple MOS sensors. The olfactory encoding part adopts a three-layer spiking neural network architecture based on the recorded electronic nose array data. The third layer provides negative feedback to the second layer, and the pulse of the second layer is output as the encoded result, converting the odor signal into a pulse signal. The pulse signal is input into the information fusion module to realize the multimodal recognition task.

7. The multi-sensory fusion bionic search robot with human-machine two-way interaction function according to claim 1, characterized in that: The robot's mechanical components include a vehicle body and a robotic arm. The vehicle body is an autonomous mobile vehicle with a two-wheel differential chassis, capable of autonomous navigation and precise tracking of a predetermined trajectory. The robotic arm is a six-degree-of-freedom collaborative robotic arm with torque feedback.

Citation Information

Patent Citations

  • Positioning robot for self-recognition and positioning system thereof

    CN112720448A

  • Bionic vision fusion severe environment imaging device and method

    CN115631123A