Intelligent ridable machine blind guiding device and method
Through the intelligent rideable robot guide device, high-precision navigation and path planning are carried out by combining voice, image and lidar data, which solves the navigation accuracy and safety issues of existing guide equipment in complex environments and realizes a stable riding experience and intelligent interaction.
Patent Information
- Application Number
- CN202510711433.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing guide technology has difficulty achieving high-precision navigation, path planning, stable interaction, and safe riding in complex environments. Existing equipment also performs poorly in dynamic obstacles and complex terrain, and cannot provide a stable riding experience.
An intelligent rideable robot guide device is used, including a load-bearing module, a data acquisition module, a control module, a wireless communication module, a monitoring and warning module, and a battery module. It collects and processes data through voice input and output, cameras, lidar, GPS positioning, and inertial measurement units, and combines artificial intelligence algorithms for path planning and posture control to achieve high-precision perception and dynamic stability.
It improves navigation accuracy, enhances multimodal interaction capabilities, improves system safety and user experience, ensures stable control in dynamic environments, and supports efficient energy management and wireless charging.
Smart Images

Figure CN120617009A_ABST
Abstract
Description
Technical Field
[0001] The present embodiment relates to the field of robotics and artificial intelligence technology, and more particularly to an intelligent rideable robotic blind-guiding device and method. Background Art
[0002] With the development of society and advancements in technology, travel assistance tools for the visually impaired have evolved from traditional guide dogs to intelligent assistive devices. However, existing guide technology still has many limitations. Traditional methods struggle to accurately capture the perceptual characteristics of speech and cannot fully meet the travel needs of the visually impaired.
[0003] In the existing technology, first, traditional guide dogs can help visually impaired people with basic navigation and obstacle avoidance after professional training, but they have the following problems: high training costs, usage restrictions and single functions. They can only provide basic guidance functions and are unable to perform complex environmental perception, path planning and remote interaction; second, the shortcomings of existing electronic guide devices. Electronic guide devices (such as smart canes and wearable navigation devices) have gradually developed, but still have the following defects: limited navigation accuracy: reliance on GPS or ultrasonic sensors, insufficient accuracy in complex environments (such as indoors, areas with dense obstacles); weak interactive capabilities: most devices only provide simple voice prompts and lack intelligent natural interaction capabilities (such as voiceprint recognition and face verification); insufficient stability: existing devices have difficulty coping with dynamic obstacles and complex terrain (such as steps and slopes), and cannot provide a stable riding experience.
[0004] There are still huge challenges in current robot guide technology, and the visually impaired urgently need to provide more intelligent, stable and safe travel assistance solutions. Summary of the Invention
[0005] The main purpose of this embodiment is to provide an intelligent rideable robot guide device and method to solve at least one of the above technical problems.
[0006] To solve the above technical problems, the technical solution adopted in this embodiment is: an intelligent rideable robot blind guide device and method, including: The carrying module includes a robot dog and a seat provided on the robot dog, serving as a carrying body for a user to sit on; The data acquisition module is used to receive user information and transmit it to the control module. It specifically includes: a voice input and output module for voice interaction; a camera for collecting image data; a lidar for environmental perception; a GPS positioning module for positioning; and an inertial measurement unit for posture detection. The control module is used to obtain data from each module and send corresponding control instructions. Specifically, it includes: the robot guide dog chip, which uses artificial intelligence algorithms based on user information to generate corresponding control instructions to control the movement of the robot dog; Wireless communication module, used for remote interactive communication and sending and receiving instructions between the control module and other modules; Monitoring and early warning module, including guardian remote monitoring module, for remote monitoring; The battery module, including batteries and charging modules, is used for power supply and power management.
[0007] In the preferred embodiment, the robot dog is a high-power quadruped robot dog with a standing size of 1000*715*470 (mm), a total weight of 60kg, a maximum speed of ≥4m / s, a maximum climbing angle of ≤45°, a step / obstacle height of ≥20CM, a protection grade of IP67, an operating temperature of -20℃~55℃, a battery life of 2.5-4h, and a range of ≥10km; The voice input and output module includes: The microphone is an externally polarized condenser microphone with a built-in 22mm gold-plated large diaphragm pickup capsule and an equivalent noise level of 15dB. The speaker is 4.5W, the sound source response time is less than 250ms, and the maximum volume is 95dB; The robot guide dog chip is an intelligent acceleration card equipped with an ultra-thin AE7100 chip, with a power consumption of 10W, LPDDR4 / 4x 128-bit memory, 8GB / 16GB capacity, compatibility with CUDA, ONNX, and Runtime API, and an INT8 computing performance of 25.6TOPS. The camera is an infrared high-definition camera with a sensor type of 1 / 2.7" Progressive Scan CMOS, a minimum illumination of 0.002 Lux @ (F1.2, AGC ON) for color, 0 Lux for infrared light on, an infrared distance of ≥50 meters, and a wide dynamic range of 120dB. The laser radar is a two-dimensional laser radar with a laser wavelength of 905nm, a scanning angle of 270 degrees, a scanning frequency of 15Hz / 30Hz, an angular resolution of 0.1 degrees / 0.3 degrees, and a working area of 0.05 meters to 5 meters.
[0008] An intelligent rideable robot blind-guiding method, applicable to the intelligent rideable robot blind-guiding device, comprises the following steps: Step 1: The user authorizes the first use of the system, extracts the user's voiceprint feature data and facial feature data when speaking, and authorizes the user; Step 2: Authorize the user to issue a riding command, execute the riding command after confirming the user's identity, and control the posture stability during riding; Step 3: Authorize the user to issue a charging command. After confirming the user's identity, the charging command is executed. The user is first placed in a safe location, and then the charging is carried out by the user. Step 4: During driving, when an impassable road surface is detected, an early warning is initiated and relevant operations are performed, specifically: Step 4-1: Collect road condition information during driving; Step 4-2: Use artificial intelligence to perform cluster analysis on road conditions, that is, to divide them into passable roads, semi-passable roads, and impassable roads; passable roads refer to roads that meet the following requirements, that is, artificial intelligence uses cameras and lidar to determine that the road width is greater than the width of the intelligent rideable robot guide dog, the uphill and downhill angles are less than the maximum angle of the intelligent rideable robot guide dog's ability to pass, the road slope height or obstacle height or pothole depth is less than the maximum height of the intelligent rideable robot guide dog's ability to pass, and the traffic light is green; semi-passable roads refer to roads that do not meet the requirements of passable roads but meet the following requirements: Artificial intelligence uses cameras and lidar to determine whether the road width is temporarily occupied by motor vehicles or pedestrians, whether there is congestion between vehicles and pedestrians, whether obstacles will move or are moving, whether large animals are passing, and whether the traffic light is red or yellow. Impassable roads refer to roads that are neither passable nor semi-passable. Artificial intelligence uses cameras and lidar to determine conditions such as: large areas of deep water accumulation, roadblocks or obstacles that are too high, broken roads, bridge collapses, fires, serious traffic accidents, and permanent or temporary no-entry signs. Preferably, the artificial intelligence uses the YOLOv7 model, inputs camera and lidar data, and performs joint learning and cluster analysis. Step 4-3: Prioritize drivable roads. If there are none, wait for the semi-drivable roads to become drivable. If there are only impassable roads, issue an alarm. Step 4-4: After the alarm, search for a new route to the destination or charging station, negotiate and confirm with the user, and then drive according to the confirmed route.
[0009] In a preferred embodiment, the step 1 comprises: Step 1-1: Collect voice and face data when the user uses it for the first time; Step 1-2: Extract voiceprint feature data; Step 1-3: Extract facial feature data; Step 1-4: Extract the correlation feature data between voiceprint and facial features; Step 1-5: Save authorized user feature data.
[0010] In the preferred solution, steps 1-2 extract the user's voiceprint feature data and facial feature data when speaking. Voiceprint recognition uses a full-process process of Gaussian filtering, framing, windowing, FFT transformation, Mel spectrum and DCT transformation to ultimately generate a Mel frequency cepstral coefficient feature matrix, specifically including: Obtain the user's voiceprint feature data, use a Gaussian filter to remove noise, then use a framing function to divide it into frames and then use a windowing function to smooth it; The smoothed signal is converted into a spectrum signal by fast Fourier transform, and the user voice signal with linear spectrum is mapped to the Mel scale using a Mel filter bank. The Mel spectrum of the user voice is then logarithmically operated and the Mel frequency cepstral coefficient features of the user voice S8 are then discrete cosine transformed. Based on the user voice with Mel-frequency cepstral coefficient characteristics, the timbre, pitch, intensity, rhythm and intonation vectors are calculated to obtain the voiceprint feature matrix of the user voice, i.e., the voiceprint feature data S9 of the user voice. The formula is: .
[0011] In the preferred embodiment, the robot guide dog chip in steps 1-3 extracts facial feature data of the user while speaking. Facial feature extraction includes grayscale conversion, mean filtering, and histogram equalization preprocessing, combined with geometric features, texture features, heat map analysis, and lip shape dynamic recognition, specifically including: Collect the user's face image in real time while they are speaking, convert the image into grayscale, then use mean filtering to remove noise, and then perform histogram equalization; Obtain the user's face image M4 after histogram equalization, calculate the vectors of geometry, texture, heat map and mouth shape, and obtain the facial feature matrix when the user speaks, that is, the facial feature data M5 when the user speaks. The formula is: ; A confidence threshold is set for the detection and matching of M5, and is adjusted according to the different characteristics of daytime, nighttime, rainy and snowy environments.
[0012] In the preferred solution, steps 1-4 extract correlation feature data between the user's voiceprint feature data and the facial feature data when speaking. The correlation feature analysis fuses the voiceprint and facial data through the YOLOv7 model and introduces the SE attention mechanism to achieve channel weighting, specifically including: Use the YOLOv7 model and initialize it: set the confidence threshold to 0.25, set the non-maximum suppression threshold to 0.45, and load the pre-trained YOLOv7 model weight file; The YOLOv7 model inputs the data of S9 and S5, aligns them in time, and inputs the original input correlation feature map U of the Head, which is divided into nine channels; Introducing the SE attention mechanism, applying the global average pooling operation to the original input correlation feature map input to the Head Compression is performed on each channel; the original input is associated with the feature map from Compress to The eigenvector Z of , with the number of channels unchanged, is: ; in, is the global average pooling result of channel c; this is a scalar value representing the average value of all pixel values of channel c; represents the global average pooling operation; is the original input associated feature map, H is the height of the original input associated feature map, that is, the number of pixels of the original input associated feature map in the vertical direction; W is the width of the original input associated feature map, that is, the number of pixels of the original input associated feature map in the horizontal direction; is the number of channels of the original input associated feature map, that is, the dimension of the original input associated feature map in the depth direction; Input original input associated feature map In the channel ,Location The value at represents the pixel value at a specific position in the original input associated feature map; is the normalization factor; Adaptive recalibration Operation, obtain the weight of [[1,1,C], a total of 9 channels, use the compressed global importance information to generate the excitation weight of each channel, the formula is: ; Among them, s represents the output channel weight vector, which represents the importance weight of each channel and has a shape of [1, 1, C]; Represents the activation function, which is used to generate channel weights; is the global descriptor of the input, with a shape of [1,1,C], obtained by global average pooling; represents the weight parameter of the model; σ is the activation function; Calculate the output correlation feature map U2 of the attention mechanism, and multiply the generated channel weight with the original input correlation feature map U to achieve channel-level feature weighting, that is, obtain the correlation feature data U2 between the user's voiceprint feature data and the facial feature data when speaking: ; in, Represents the value of the output correlation feature map U2 on channel c, which is the result after channel weighting; represents the scaling function used to apply the channel weights to the original input associated feature map U; Represents the value of the original input associated feature map U on channel c; Represents the weight of channel c, which is determined by the activation function Generate, indicating the importance of channel c; The output correlation feature map U2 between the user's voiceprint feature data and the facial feature data when speaking is matched according to the preset confidence threshold.
[0013] In a preferred embodiment, the step 2 comprises: Step 2-1: Receive riding instructions and collect current user characteristics; Step 2-2: Match and verify with the authorization feature; Step 2-3: Plan and confirm the optimal cycling route. The path planning algorithm combines the global A* method with the local dynamic window method to avoid local optimal solutions through adaptive weight adjustment, supporting real-time avoidance of low obstacles and dynamic targets. Steps 2-4: Real-time control of posture stability during riding.
[0014] In a preferred embodiment, step 3 includes: Step 3-1: The battery module monitors the power level in real time and issues a voice alarm when the power level is below the threshold. Step 3-2: Receive charging instructions and perform chip verification triple feature matching; Step 3-3: Based on the segmented path principle of "current location → safe unloading point → charging pile", use the A* algorithm to optimize the turning points and plan the optimal charging route; Step 3-4: Place the user in a safe location according to the optimal charging route and proceed to charge. After charging is completed, the device returns to the starting point, allowing the guardian to remotely monitor the charging status.
[0015] In the preferred embodiment, in step 3, during the charging process, the selection of a safe placement point is based on real-time environmental scanning by a laser radar (104), requiring that there are no dynamic obstacles within 5 meters and that the ground flatness is ≤5°, and the selection is confirmed with the user through voice; Battery status prediction uses a Kalman filter algorithm, combined with real-time power consumption (positively correlated with speed, u=kv²) to predict remaining battery life with an error of ≤5%.
[0016] This embodiment provides an intelligent rideable robot guide device and method, including a user riding on a carrying module and transmitting necessary information to a control module through a data acquisition module. The control module generates corresponding control instructions through an artificial intelligence algorithm, and each module executes the corresponding instructions through a wireless communication module to control the driving of the robot dog; and the monitoring and early warning module can provide early warning through remote monitoring. During the driving process, the battery module detects the power level of the battery and charging module 108 and flexibly adjusts them; thereby achieving high-precision perception, intelligent authentication, dynamic stability control, wireless charging and multimodal interaction.
[0017] The technical effects are as follows: 1. Improved the accuracy of voiceprint feature extraction. By simulating the human ear's nonlinear perception of frequency (sensitive to low frequencies and coarse at high frequencies), the extracted voiceprint features (such as timbre, pitch, and rhythm) are more consistent with actual auditory perception, improving feature robustness and enhancing noise suppression capabilities.
[0018] Second, the reliability of multimodal interaction has been improved. Triple authentication of voiceprint + face + associated features (based on YOLOv7 + SE attention mechanism) is adopted to prevent misoperation or malicious interference, ensuring that only authorized users can issue commands, realizing dynamic environment analysis and decision-making, and significantly improving anti-interference capabilities.
[0019] 3. It improves system safety, active obstacle avoidance capabilities, and user experience. The A* algorithm can quickly find the target path through heuristic functions. The voiceprint features extracted by the Mel spectrum are linked with the inertial measurement unit (IMU) data to trigger the IMU's rapid posture adjustment. Combined with the artificial potential field method, it can predict the risk of user body imbalance and adjust the robot dog's center of gravity in advance, reducing the risk of falling.
[0020] 4. It improves the efficiency of energy management and wireless charging, supports Kalman filtering to predict remaining battery life, actively reminds and plans charging routes when the battery is low, automatically unloads users to a safe location during charging, and returns to pick up after charging is complete, without the need for human intervention throughout the process. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The present embodiment will be further described below with reference to the accompanying drawings and examples: Figure 1 1 is a structural diagram of the blind guide device according to this embodiment; Figure 2 4 is a flow chart of the blind guiding method in this embodiment. DETAILED DESCRIPTION
[0022] Example 1 like Figure 1-2 As shown, an intelligent rideable machine guide device for the blind comprises: The carrying module includes a robot dog 100 and a seat 109 provided on the robot dog 100 , serving as a carrying body for a user to sit on.
[0023] The data acquisition module is used to receive user information and transmit it to the control module. It specifically includes: a voice input and output module 102 for voice interaction; a camera 103 for collecting image data; a lidar 104 for environmental perception; a GPS positioning module 105 for positioning; and an inertial measurement unit 106 for posture detection.
[0024] The control module is used to obtain data from each module and send corresponding control instructions. Specifically, it includes: the robot guide dog chip 101 uses artificial intelligence algorithms based on user information to generate corresponding control instructions to control the driving of the robot dog 100.
[0025] The wireless communication module 107 is used for remote interactive communication and sending and receiving instructions between the control module and other modules.
[0026] The monitoring and early warning module includes a guardian remote monitoring module 110 for remote monitoring.
[0027] The battery module, including batteries and a charging module 108, is used for power supply and power management.
[0028] In this embodiment, Figure 1 As shown, the user sits on the carrying module and transmits necessary information to the control module through the data acquisition module. The control module generates corresponding control instructions through the artificial intelligence algorithm. Through the wireless communication module 107, each module executes the corresponding instructions to control the driving of the robot dog; and the monitoring and warning module can provide early warnings through remote monitoring. During the driving process, the battery module detects the power of the battery and the charging module 108 and adjusts them flexibly; thus, through high-precision perception, intelligent authentication, dynamic stability control, wireless charging and multimodal interaction, the shortcomings of traditional guide tools in navigation accuracy, safety, battery life, interactive experience, etc. are solved, and the intelligence, stability and safety of visually impaired people's travel are greatly improved.
[0029] In the preferred embodiment, the robot dog 100 provides a carrying body for this embodiment, which is used for visually impaired people to ride. It is equipped with a robot guide dog chip 101, a voice input and output module 102, a camera 103, a laser radar 104, a GPS positioning module 105, an inertial measurement unit 106, a wireless communication module 107, a battery and charging module 106, and a seat 109; preferably, it is a high-power four-legged robot dog with a standing size of 1000*715*470 (mm), a total weight of 60kg, a maximum movement speed ≥4m / s, a maximum climbing angle ≤45°, a step / obstacle height ≥20CM, a protection level IP67, an operating temperature of -20℃~55℃, a battery life of 2.5-4h, a cruising range ≥10km, an external communication interface USB2.0, USB3.0 or Ethernet WiFi, an external power supply of 5V, 12V or 24V, and a load capacity of not less than 60kg. Furthermore, the computer included in the robot dog 100 can serve as a slave computer and be controlled by the robot guide dog chip 101 .
[0030] The robotic guide dog chip 101 provides the core computing power for this embodiment and runs artificial intelligence algorithms. Preferably, it is a smart accelerator card equipped with an ultra-thin AE7100 chip, featuring a fully programmable design. It consumes 10W of power, has 8GB / 16GB of LPDDR4 / 4x 128-bit memory, and is compatible with CUDA, ONNX, and the Runtime API. It supports models such as Llama2, Stable Diffusion, Yolov5, and ResNet. It boasts an INT8 computing performance of 25.6TOPS. Measuring 80mm long by 22mm wide, it offers high flexibility for various AI applications. Its computing power reaches 25.6TOPS, and its memory bandwidth reaches 60GBs, ensuring efficient and stable processing and data transmission. Furthermore, the robotic guide dog chip 101, acting as a host computer, can control the robot dog 100's built-in computer, thereby driving the robot dog 100 and controlling its speed and direction.
[0031] The voice input and output module 102 provides voice input and output functions for this embodiment and includes a microphone and speaker. Preferably, the microphone is an externally polarized condenser microphone with a built-in 22mm gold-plated large-diaphragm pickup capsule and an equivalent noise level of 15dB, ensuring high sound reproduction and fineness. Preferably, the speaker is a 4.5W speaker with a sound source response time of less than 250ms, a maximum volume of 95dB, software and hardware adjustable, and an operating temperature range of -40°C to 60°C.
[0032] Camera 103 provides depth imaging and target detection for this embodiment during the day and at night. Preferably, it is an infrared high-definition camera with a sensor type of 1 / 2.7" Progressive Scan CMOS, a minimum illumination of 0.002 Lux @F1.2 for color, AGC ON, 0 Lux infrared light on, infrared distance ≥ 50 meters, and a shutter speed of 1 / 3 s to 1 / 100,000s, wide dynamic range of 120dB, day / night mode, infrared filter, focal length and field of view of 2.7 to 12mm, horizontal field of view of 110° to 35°, vertical field of view of 58° to 20°, diagonal field of view of 132° to 41°, maximum aperture of F1.2, and maximum image size of 1920×1080. Furthermore, it features intelligent alerting, including area intrusion detection, boundary crossing detection, area entry detection, area exit detection, object left behind detection, object removal detection, loitering detection, parking detection, crowd gathering detection, and rapid movement detection.
[0033] The laser radar 104 provides stereo imaging of regional objects for this embodiment. Preferably, it is a two-dimensional laser radar with a laser wavelength of 905nm, a scanning angle of 270 degrees, a scanning frequency of 15Hz / 30Hz, an angular resolution of 0.1 degrees / 0.3 degrees, a working area of 0.05 meters to 5 meters, a self-learning function, automatic scanning and generation of the environment, 16 area groups, each area group containing 3 areas, a power supply voltage of DC9V~28V, a power consumption of 2W, a protection level of IP65, can recognize almost any shape, a measurement error of ±30mm, vibration resistance, and an operating temperature of -20℃~55℃.
[0034] The GPS positioning module 105 provides precise positioning for this embodiment. Preferably, it is a high-precision, low-power quad-mode satellite positioning module, namely, single-mode, dual-mode, and multi-mode operation of Beidou + GPS + Galileo + Glonass, and can be switched between each other through commands. The operating voltage is 3.0V~3.5V, supports SBAS, QZSS, and A-GNSS assisted positioning, with an acquisition sensitivity of -147dB, a tracking sensitivity of -163dB, a positioning accuracy of less than 3 meters, supports powering active antennas, and an operating temperature of -40℃~85℃.
[0035] The Inertial Measurement Unit 106 is a compact, three-axis fiber-optic gyroscope (FOG) inertial measurement unit (IMU). It features a built-in high-precision three-axis fiber-optic gyroscope (FMG) and accelerometer. It outputs acceleration and angular velocity information without relying on external signal input, allowing users to calculate the azimuth, roll, and pitch angles of the measured vehicle. It is suitable for inertial measurement in various states, including motion, vibration, and static conditions. The unit offers a supply voltage of 12-24V, a ripple of ≤50mV, a measurement range of ±300° / s, and an operating temperature range of -40°C to +85°C. Dimensions: L97.3×W90×H70mm. Gyroscope performance parameters include bias stability of 10s, 1σ≤0.08° / h, and bias repeatability of 1σ≤0.08° / h, with a measurement range of ±300deg / s. Accelerometer performance parameters include bias stability of 10s, 1σ≤0.5mg, and bias repeatability of 1σ≤0.5mg at room temperature, with a measurement range of ±30g. Furthermore, by adopting high-reliability MEMS accelerometers and three-axis fiber optic gyroscopes, the original data deviation is estimated accordingly through a 6-state Kalman filter with appropriate gain, and the measurement accuracy is guaranteed by the algorithm. The parameters are compensated through various means such as nonlinear compensation, orthogonal compensation, temperature compensation and drift compensation, which can greatly eliminate errors and improve the product accuracy level.
[0036] The wireless communication module 107 provides wireless remote communication for this embodiment. Preferably, it is a multi-protocol communication module that supports Bluetooth, Zigbee, Thread, Proprietary, and Wi-Fi. It has a maximum flash memory of 3200kB / RAM of 512kB and an output power range of -20~19.5dBm.
[0037] The battery and charging module 106 provides energy supply for this embodiment; preferably, it is a high-power wireless power supply module with a rated power of 1300W, an input voltage of 220V±15%, 50Hz, an input current of 6A, a charging voltage of 54.6V±0.2, an output current of 20A±5%, an output type constant current-constant voltage charging curve, a charging efficiency of 87%, a transmitting heat dissipation method of natural heat dissipation of the aluminum shell, a receiving heat dissipation method of natural heat dissipation of the aluminum shell, an operating environment of -20-55℃, RH<90%, a storage temperature of -20-55℃, RH<90%, a protection temperature of 85℃, a hysteresis of 10℃, output overvoltage protection, and output overcurrent protection.
[0038] The seat 109 provides stable support for the visually impaired person to ride, and is preferably a seat with a front armrest. Furthermore, the seat may be equipped with a gravity sensor to monitor whether the rider is riding stably.
[0039] The guardian remote monitoring module 110 provides users with communication and control of the robot dog 100 and can communicate with the wireless communication module 107; preferably, it is a smartphone with a matching APP installed, a 5300mAh large-capacity battery, 50W wireless super fast charging, a rear camera with a 50-megapixel ultra-wide-angle camera with an aperture of F1.4~F4.0, OIS optical image stabilization + a 40-megapixel ultra-wide-angle camera with an aperture of F2.2 + a 12-megapixel periscope telephoto camera with an aperture of F3.4, OIS optical image stabilization + a 1.5-megapixel multi-spectral channel red maple primary color camera, a front camera with a 13-megapixel ultra-wide-angle camera with an aperture of F2.4, a screen size of 6.7 inches, a body memory ROM of 512GB, a screen type OLED, supports a 1-120Hz LTPO adaptive refresh rate, 1440Hz high-frequency PWM dimming, a 300Hz touch sampling rate, an operating system HarmonyOS 4.3, and a WLAN protocol that supports 802.11 a / b / g / n / ac / ax, 2x2 MIMO, HE160, 1024QAM, 8 Spatial-stream Sounding MU-MIMO, WLAN frequency 2.4GHz and 5GHz, support WLAN hotspot, WLAN direct connection, Bluetooth Bluetooth 5.2, support low-power Bluetooth, support SBC, AAC, support LDAC and L2HC high-definition audio, OTG supports maximum output current 1A / 5V when reverse power supply, infrared remote control, Skylink, positioning supports GPSL1+L5 dual-band / AGPS / GLONASS / Beidou B1I+B1C+B2a+B2b quad-band / GALILEOE1+E5a+E5b tri-band / QZSSL1+L5 dual-band / NavIC, product operating temperature 0℃~35℃.
[0040] This embodiment uses the devices of the above modules to provide an intelligent rideable robot guide device, providing a more intelligent, stable and safe travel assistance solution for the visually impaired.
[0041] Example 2 Further illustrate with reference to Example 1, Figure 2 As shown, an intelligent rideable blind-guiding machine method is applicable to an intelligent rideable blind-guiding machine device of Example 1, comprising the following steps: Step 1: The user uses the authorization for the first time, extracts the user's voiceprint feature data and facial feature data when speaking and authorizes it.
[0042] Step 1-1, when a user uses the device for the first time, the user sits on the seat 109 and speaks a few words to the voice input and output module 102 , and the camera 103 takes a picture of the user's face while the user is speaking.
[0043] In step 1-2, the voice input and output module 102 sends the voice signal of the user speaking to the robot guide dog chip 101, and the camera 103 sends the captured face of the user speaking to the robot guide dog chip 101.
[0044] Step 1-3: The robot guide dog chip 101 extracts the user's voiceprint feature data and facial feature data when speaking and authorizes them, while analyzing the correlation feature data between the user's voiceprint feature data and facial feature data when speaking; Step 1-4: Extract the correlation feature data between the voiceprint and facial features.
[0045] Step 1-5: Save authorized user feature data.
[0046] Step 2: Authorize the user to issue riding instructions, execute the riding instructions after confirming the user's identity, and control the posture stability during riding.
[0047] Step 3: Authorize the user to issue a charging command, execute the charging command after confirming the user's identity, place the user in a safe place first, and then go to charge by yourself.
[0048] Step 4: During driving, when an impassable road surface is detected, an early warning is initiated and relevant operations are performed, specifically: Step 4-1: Collect road condition information during driving.
[0049] Step 4-2: Perform cluster analysis on the road conditions through artificial intelligence, that is, classify them into passable roads, semi-passable roads, and impassable roads; passable roads refer to roads that meet the following requirements, that is, artificial intelligence determines through the camera 103 and the laser radar 104 that the width of the road surface is greater than the width of the intelligent rideable robot guide dog, the uphill and downhill angles are less than the maximum angle of the intelligent rideable robot guide dog's ability to pass, the road slope height or obstacle height or pothole depth is less than the maximum height of the intelligent rideable robot guide dog's ability to pass, and the traffic light is green; semi-passable roads refer to roads that do not meet the requirements of passable roads but meet the following requirements, and artificial intelligence determines through the camera 103 and the laser radar 104 that the road surface can pass through the road surface, the uphill and downhill angles are less than the maximum angle of the intelligent rideable robot guide dog's ability to pass, and the traffic light is green. The camera 103 and the lidar 104 determine whether the width of the road is temporarily occupied by motor vehicles or pedestrians, there is congestion of vehicles and pedestrians, obstacles will move or are moving, large animals are passing, and the traffic light is red or yellow. An impassable road surface refers to a road surface that does not meet the requirements of a traversable road surface or a semi-traversable road surface. Artificial intelligence uses the camera 103 and the lidar 104 to determine the following: large-scale and deep water accumulation, roadblocks or obstacles that are too high, broken roads, collapsed bridges, fires, serious traffic accidents, and permanent or temporary no-entry signs. Preferably, the artificial intelligence uses the YOLOv7 model, inputs the camera 103 and the lidar 104 data, and performs joint learning and cluster analysis.
[0050] Step 4-3: Prioritize passable roads. If there are none, wait for the semi-passable roads to be converted into passable roads. If there are only impassable roads, issue an alarm.
[0051] Step 4-4: After the alarm, search for a new route to the destination or charging station, negotiate and confirm with the user, and then drive according to the confirmed route.
[0052] In this embodiment, steps 1-4 ensure that only authorized users can operate through triple biometric features (voiceprint + face + associated features), achieving accurate extraction and association analysis of voiceprint and facial features. When a riding instruction is issued, the triple matching degree is calculated, and if the match is successful, the instruction is executed. In path planning and riding control, the start / end point coordinates are obtained through GPS + lidar, and the improved A* algorithm (adaptive obstacle weighting) is used to plan the optimal path. The posture is monitored in real time, and the motion stability is dynamically adjusted in combination with the artificial potential field method. Intelligent charging management is also performed. This improves the safety, intelligence, and reliability of travel and enhances the user experience.
[0053] In the preferred solution, steps 1-2 extract the user's voiceprint feature data and facial feature data when speaking. Voiceprint recognition uses a full process of Gaussian filtering, framing, windowing, FFT transformation, Mel spectrum and DCT transformation to ultimately generate a Mel frequency cepstral coefficient (MFCC) feature matrix, specifically including: In sub-step 1-2-1, the robot guide dog chip 101 reads the user's speech signal S1 collected by the speech input and output module 102, and uses a Gaussian filter (Gauss) to remove noise, thereby obtaining the user's speech signal S2 after removing background noise: .
[0054] In sub-step 1-2-2, the robot guide dog chip 101 uses the framing function Frame() to frame the user voice signal S2 after removing the background noise, that is, to divide it into a series of short frames, each frame containing a signal of a certain time, preferably, the length of each frame is 20-40 milliseconds; further, the framed user voice signal S3 is obtained: .
[0055] In sub-step 1-2-3, the robot guide dog chip 101 uses the windowing function Window() to window the framed user voice signal S3, that is, smoothing each frame signal and smoothing the edge of each frame so that the edge signal of the frame gradually decreases to zero; further, the windowed user voice signal S4 is obtained: .
[0056] In sub-step 1-2-4, the robot guide dog chip 101 uses fast Fourier transform (FFT) to convert the windowed user voice signal S4 into a spectrum signal to obtain the frequency information of the sound signal, i.e., the user voice signal S5 after fast Fourier transform: .
[0057] In sub-step 1-2-5, the robot guide dog chip 101 uses the Mel filter bank Mel() to map the user voice signal S5 of the linear spectrum to the Mel scale to obtain the Mel spectrum S6 of the user voice: .
[0058] In sub-step 1-2-6, the robot guide dog chip 101 uses the logarithmic operation Log() to perform a logarithmic operation on the Mel spectrum S6 of the user's voice to make the feature distribution closer to the perceptual characteristics of human hearing, and obtain the logarithmic Mel spectrum of the user's voice S7: .
[0059] In sub-step 1-2-7, the robot guide dog chip 101 uses discrete cosine transform (DCT) to transform the logarithmic Mel frequency spectrum of the user voice S7 to obtain the Mel frequency cepstral coefficient (MFCC) feature of the user voice S8: .
[0060] In sub-step 1-2-8, the robot guide dog chip 101 analyzes the user voice S8 based on the Mel-Frequency Cepstral Coefficient (MFCC) features, calculates timbre, pitch, intensity, rhythm, and intonation vectors, and obtains the voiceprint feature matrix of the user voice, i.e., the voiceprint feature data S9 of the user voice: .
[0061] In sub-step 1-2-9, the robot guide dog chip 101 sets a confidence threshold for detecting and matching the user's voiceprint feature data S9. The robot guide dog chip 101 matches the user's voiceprint feature data based on the confidence threshold. Preferably, the user sets a confidence threshold, and only user voiceprint feature data with a confidence level above this threshold is detected. The default value is typically 0.2, but can be adjusted based on the characteristics of the surrounding human voice, quiet, or noisy environment.
[0062] In the preferred embodiment, in steps 1-3, the robot guide dog chip (101) extracts facial feature data of the user when speaking. The facial feature extraction includes grayscale conversion, mean filtering, and histogram equalization preprocessing, combined with geometric features, texture features, heat map analysis, and lip shape dynamic recognition, specifically including: In sub-step 1-3-1, the robot guide dog chip 101 obtains the facial image M1 of the user when speaking captured by the camera 103 in real time, and grayscales the captured facial image of the user when speaking, that is, converts the color image into a grayscale image M2, using the following formula: ; in, is the pixel value of the grayscale image, 0.299, 0.587, and 0.114 respectively represent the contribution of red, green, and blue to human eye perception, and R, G, and B are the pixel values of the red, green, and blue channels in the image respectively.
[0063] Sub-step 1-3-2: the robot guide dog chip 101 uses the mean filtering method Remove the noise from the face image M2 and obtain the denoised face image M3 of the user speaking. The formula is: ; in, is the pixel value of the filtered image, is the original image pixel value, is the filter size, is the coordinate position of the current pixel in the image, is the offset within the filter template used to traverse the pixels within the filter range, Normalization factor, used to divide the summation result by the total number of pixels in the filter to ensure that the pixel values of the output image are within a reasonable range.
[0064] In sub-step 1-3-3, the robot guide dog chip 101 performs histogram equalization on the denoised face image M3 to enhance the image contrast and obtain a histogram-equalized user face image M4. The formula is: ; ; ; ; in, Grayscale The number of occurrences, Grayscale The number of pixels, is the grayscale level, Indicates the Gray levels, Used to calculate the cumulative distribution function, which represents the probability of the occurrence of pixels less than or equal to a certain gray level. is the total number of pixels in the graphic, Grayscale The cumulative distribution function of .
[0065] In sub-steps 1-3-4, the robot guide dog chip 101 analyzes the histogram-equalized user face image M4, calculates geometry, texture, heat map, and mouth shape vectors, and obtains a facial feature matrix when the user is speaking, i.e., facial feature data M5 when the user is speaking. .
[0066] Furthermore, geometric features refer to the shape, outline, and positional relationship of the eyes, nose, and mouth of the face. By calculating parameters such as the distance between the eyes of two faces and the relative position of the mouth and nose, different faces can be distinguished. Texture features refer to texture features such as skin color, wrinkles, and spots on the face, which are different on each person's face. By analyzing the changes in texture on the face image when speaking, a specific face can be identified. Heat map technology refers to the detection of hot spot distribution on the face, and based on the characteristics of thermal energy radiated by the human body, different faces can be identified. Mouth shape refers to the changes in the user's mouth shape when speaking, including the deformation of the lips, the exposure of teeth, and the shape of the tongue, which are all related to a person's physiological characteristics or long-term habits.
[0067] In sub-step 1-3-5, the robotic guide dog chip 101 sets a confidence threshold for detecting and matching the user's facial feature data M5. The robotic guide dog chip 101 matches the user's facial feature data M5 according to the confidence threshold. Preferably, the user sets a confidence threshold, and only facial feature data from the user's speech with a confidence level above this threshold is detected. The default value is typically 0.3, but can be adjusted based on different characteristics of daytime, nighttime, rainy, and snowy environments.
[0068] In the preferred solution, steps 1-4 extract the correlation feature data between the user's voiceprint feature data and the facial feature data when speaking. The correlation feature analysis uses the YOLOv7 model to fuse the voiceprint and facial data, and introduces the SE attention mechanism to achieve channel weighting, specifically including: In sub-step 1-4-1, the robot guide dog chip 101 initializes the YOLOv7 model, sets the confidence threshold of the detection result to 0.25, sets the threshold of non-maximum suppression to 0.45, and loads the pre-trained YOLOv7 model weight file.
[0069] In sub-step 1-4-2, the robot guide dog chip 101 inputs the voiceprint feature data S9 of the user's voice obtained in sub-step 1-3-1 and the facial feature data M5 of the user when speaking obtained in sub-step 1-3-2 to YOLOv7, and aligns them in time, that is, inputs the original input association feature map U to the Head; and divides them into nine channels, namely timbre, pitch, intensity, rhythm, and intonation in sub-step 1-3-1-8, and geometry, texture, heat map, and lip shape in sub-step 1-3-2-4.
[0070] In sub-step 1-4-3, the robot guide dog chip 101 introduces the Squeeze-and-Excitation Network (SE) attention mechanism to help the YOLOv7 network adaptively select and weight features from different channels, thereby improving network efficiency and performance, enhancing network generalization capabilities, and adapting to different tasks and application requirements. For example, feature enhancement of user voice in noisy environments and feature enhancement of user faces in daytime or nighttime environments can be performed.
[0071] In sub-step 1-4-4, the robot guide dog chip 101 applies a global average pooling operation to the original input correlation feature map input to the Head through SE Compression is performed on each channel, which is divided into 9 channels: timbre, pitch, intensity, rhythm, intonation, geometry, texture, heat map and mouth shape. The original input can be associated with the feature map from Compress to The feature vector Z of is obtained, and the number of channels remains unchanged, thereby converting the original input associated feature map of each channel into a single value, and obtaining the most representative value of each channel of the original input associated feature map. This value represents the global importance of the channel. The calculation formula is as follows: ; in, is the global average pooling result of channel c. This is a scalar value representing the average of all pixel values of channel c. Represents a global average pooling operation. is the original input associated feature map, H is the height of the original input associated feature map, that is, the number of pixels in the vertical direction of the original input associated feature map. W is the width of the original input associated feature map, that is, the number of pixels in the horizontal direction of the original input associated feature map. is the number of channels of the original input associated feature map, that is, the dimension of the original input associated feature map in the depth direction. Input original input associated feature map In the channel ,Location The value at represents the pixel value at a specific position in the original input associated feature map. is a normalization factor used to divide the summation result by the total dimension of the original input associated feature map , ensuring that the range of output values is consistent with the input values.
[0072] Sub-steps 1-4-5: Adaptive recalibration of the robot guide dog chip 101 Operation. The robot guide dog chip 101 generates weights [[1,1,C]], which are divided into 9 channels: timbre, pitch, intensity, rhythm, intonation, geometry, texture, heat map, and lip shape. It uses the previously compressed global importance information to generate the excitation weight for each channel. It inputs the global importance information into a small multi-layer perceptron (MLP), passes through some fully connected layers and nonlinear activation functions, and generates a weight vector with the same number of channels. These weight vectors represent the importance weight of each channel: ; Among them, s represents the output channel weight vector, which represents the importance weight of each channel and has a shape of [1, 1, C]. Represents the activation function, which is used to generate channel weights. is the global descriptor of the input, with a shape of [1,1,C], obtained by global average pooling. represents the weight parameter of the model. σ is the activation function, which is used to limit the output value to between 0 and 1.
[0073] ; Among them, g represents a multi-layer perceptron used to process the global descriptor z. is the weight matrix of the first fully connected layer, used for channel compression. δ is the ReLU activation function, used to introduce nonlinearity.
[0074] In sub-steps 1-4-6, the robot guide dog chip 101 calculates the output correlation feature map U2 of the attention mechanism and multiplies the generated channel weights with the original input correlation feature map U to achieve channel-level feature weighting, that is, to obtain the correlation feature data U2 between the user's voiceprint feature data and the facial feature data when speaking: ; in, It represents the value of the output correlation feature map U2 on channel c, which is the result after channel weighting. Represents a scaling function that applies channel weights to the original input associated feature map U. Represents the value of the original input associated feature map U on channel c. Represents the weight of channel c, which is determined by the activation function Generate, indicating the importance of channel c.
[0075] The SE attention mechanism used in this embodiment determines the proportion of the channel in training according to its own importance, increases the proportion of important channels, and reduces the proportion of unimportant channels, which can effectively improve the effect of user voice and face detection.
[0076] In sub-step 1-4-7, the robot guide dog chip 101 sets a confidence threshold for the output correlation feature map U2. Based on the confidence threshold, the robot guide dog chip 101 matches the correlation feature data between the user's voiceprint feature data and the facial feature data during speech. Preferably, the user sets a confidence threshold, and only correlation feature data with a confidence level above this threshold is detected. The default value is typically 0.25, but can be adjusted based on different characteristics of daytime, nighttime, and noisy environments.
[0077] In a preferred embodiment, step 2 includes: Step 2-1: Receive riding instructions and collect current user characteristics, specifically: The authorized user issues a riding instruction to the voice input and output module 102. When the robot guide dog chip 101 detects the riding instruction, it triggers the camera 103 to capture the user's face while speaking. The robot guide dog chip 101 extracts the voiceprint feature data of the user issuing the riding instruction and the facial feature data of the user speaking, and simultaneously analyzes the correlation feature data between the voiceprint feature data of the user issuing the riding instruction and the facial feature data of the user speaking.
[0078] Step 2-2: Verify the match with the authorization feature, specifically: The robot guide dog chip 101 matches the voiceprint feature data of the user issuing the riding instruction with the saved voiceprint feature data of the authorized user; further, the robot guide dog chip 101 matches the facial feature data of the user when issuing the riding instruction with the saved facial feature data of the authorized user when speaking; further, the robot guide dog chip 101 matches the correlation feature data between the voiceprint feature data and the facial feature data when the user issues the riding instruction with the saved correlation feature data between the voiceprint feature data and the facial feature data when speaking of the authorized user.
[0079] If the matching degree calculated by the robot guide dog chip 101 is relatively high, the riding instruction is confirmed; otherwise, if the matching degree calculated by the robot guide dog chip 101 is relatively low, the riding instruction is directly ignored, effectively preventing the surrounding human voices from interfering with the work of the intelligent rideable robot guide dog.
[0080] Steps 2-3: Plan and confirm the optimal cycling route. The path planning algorithm combines global A* with the local dynamic window algorithm (DWA). It avoids local optimal solutions through adaptive weight adjustment and supports real-time avoidance of low obstacles (such as steps) and dynamic targets (such as pedestrians). Specifically: The robot guide dog chip 101 plans the optimal cycling route based on the current cycling instruction and reports the optimal cycling route to the user through the voice input and output module 102. The user confirms or modifies the optimal cycling route through the voice input and output module 102. The robot guide dog chip 101 confirms the user's confirmation or modification of the cycling route by matching the voiceprint feature data, matching the facial feature data when speaking, and matching the associated feature data between the two. If the matching degree calculated by the robot guide dog chip 101 is relatively high, the optimal or modified cycling route is confirmed.
[0081] Steps 2-4: Real-time control of posture stability during riding, specifically: The robot guide dog chip 101 controls the intelligent rideable robot guide dog to ride according to the confirmed optimal or modified riding route, and detects and controls the posture stability of the intelligent rideable robot guide dog in real time during riding to prevent the user from falling.
[0082] In this embodiment, the matching verification of the authorization features in steps 2-1 to 2-2 is the same as that in step 1 and will not be repeated here.
[0083] Furthermore, confirming the optimal cycling route also includes: Sub-step 2-3-1: The robot guide dog chip 101 analyzes the riding end position from the user's riding instruction. , is the end point coordinate.
[0084] Sub-step 2-3-2: The robot guide dog chip 101 obtains the current location of the intelligent rideable robot guide dog through the GPS positioning module 105. , The current location coordinates of the intelligent rideable robotic guide dog.
[0085] In sub-step 2-3-3, the robot guide dog chip 101 uses the A* algorithm to perform global path planning, initializes the open list and the closed list, and sets the starting point Add to the open list, and leave the closed list empty. Define a cost function. Based on the evaluation function, judge the distance between the robot dog 100 and the surrounding environment, i.e., the target point, during each movement. This algorithm makes the robot dog 100's path planning more purposeful and efficient. The cost estimation function is as follows: ; in, is the total cost from the current node a to the target point; is the moving cost from the starting point to the current point a; It is the estimated cost from the current point a to the end point, also known as the heuristic function.
[0086] In sub-step 2-3-4, the robot guide dog chip 101 introduces an adaptive cost weight function. To prevent the A* algorithm from falling into a local optimum and improve its path search efficiency, the robot guide dog chip 101 incorporates the obstacle ratio in the local map into the heuristic function. This allows the robot dog 100 to adaptively adjust the weight of the heuristic function during movement. When there are many obstacles or charging stations, the heuristic function weight is increased; when there are few obstacles or charging stations, the heuristic function weight is decreased.
[0087] ; Where, yes The algorithm expands the obstacle ratio in the local map from the parent node of the node to the end point. It can be expressed as follows: .
[0088] in, is the number of obstacles; The parent node coordinates of the current node in the algorithm planning process; is the end point coordinate.
[0089] In sub-steps 2-3-5, the robot guide dog chip 101 runs the A* algorithm, iterates through the nodes, and searches for the node from the current position. To the end of the ride The optimal path L*( , ). If the open list is empty, it means there is no path to reach, and the algorithm ends. If the open list is not empty, take The smallest node a. If a is the end point , the algorithm ends, and the path is returned by backtracking the parent node to build the path. Otherwise, add a to the closed list and check all adjacent nodes of a.
[0090] In substep 2-3-6, the robot guide dog chip 101 performs inflection point optimization. Starting from the endpoint, the robot traverses the next two nodes in the order in which they are stored, determining whether the lines connecting them are aligned. If so, the traversal continues; otherwise, the robot traverses the node starting point, stopping at the last traversed point and completing the node optimization. The robot guide dog chip 101 determines whether the distance from the obstacle in the local map to the line connecting the two points is less than a threshold. If so, the robot guide dog chip optimizes the inflection point to prevent the user from riding at excessive angles, causing instability and even the risk of falling.
[0091] Sub-step 2-3-7, the robot guide dog chip 101 processes the adjacent nodes. For each adjacent node m: if m is in the closed list, skip; if m is not in the open list, calculate 、 and , and add it to the open list; if m is already in the open list, check whether the cost of reaching m through the current path is smaller. If so, update 、 and , and update the parent node of m to n.
[0092] In sub-steps 2-3-8, the robot guide dog chip 101 repeats sub-steps 2-5-6 until it finds the end point. Or the open list is empty.
[0093] Sub-step 2-3-9, the robot guide dog chip 101 builds a path, when the end point When it is added to the open list and processed, the robot guide dog chip 101 constructs a path L ( , ).
[0094] In sub-step 2-3-10, the robot guide dog chip 101 determines whether the algorithm is finished. If so, the robot guide dog chip 101 returns the optimal path L*( , ), otherwise return to substep 2-5-3 to continue the calculation. If the open list is empty at the end of the algorithm and no path is found, then return "no path".
[0095] In sub-step 2-3-11, the robot guide dog chip 101 reports the optimal riding route L*( , ), or No Path.
[0096] Furthermore, the user confirms or modifies the optimal cycling route through the voice input and output module 102, and the robot guide dog chip 101 confirms the user's confirmation or modification of the cycling route by matching the voiceprint feature data, matching the facial feature data when speaking, and matching the associated feature data between the two. If the matching degree calculated by the robot guide dog chip 101 is relatively high, the optimal or modified cycling route is confirmed.
[0097] In sub-step 2-4, the robot guide dog chip 101 controls the intelligent rideable robot guide dog to ride according to the confirmed optimal or modified riding route, and monitors and controls the posture stability of the intelligent rideable robot guide dog in real time during riding to prevent the user from falling.
[0098] In sub-step 2-4-1, the inertial measurement unit 106 detects the acceleration and angular velocity information of the intelligent rideable robotic guide dog in real time and transmits the information to the robotic guide dog chip 101.
[0099] In sub-step 2-4-2, the robot guide dog chip 101 calculates the azimuth, roll, and pitch angles of the intelligent rideable robot guide dog based on the acceleration and angular velocity information detected in real time by the inertial measurement unit 106, and measures the inertial parameters of the intelligent rideable robot guide dog in various postures, including motion, vibration, and stillness, while riding.
[0100] In sub-step 2-4-3, the robot guide dog chip 101 initializes the parameters related to the artificial potential field method and defines the potential field parameters, including the strength and attenuation coefficient of the gravitational field and the strength and range of the repulsive field.
[0101] In sub-step 2-4-4, the robot guide dog chip 101 constructs a potential field for the motion posture of the robot dog 100, and constructs a gravitational field and a repulsive field respectively. The distance to the stable equilibrium point Increase and decrease, usually defined as: ; in, represents the gravitational field strength, represents the distance between the robot dog 100 and the stable equilibrium point, and η is the gravitational coefficient.
[0102] Strength of the repulsive field The distance from the easy falling point Decrease and increase, usually defined as: ; in, represents the strength of the repulsive field, It indicates the distance of the robot dog 100 from the falling point. is the repulsion coefficient, It's a safe distance.
[0103] In sub-steps 2-4-5, the robot guide dog chip 101 calculates the magnitude of the resultant force. During the movement of the robot dog 100, whether its posture is stable or prone to falling depends on whether the resultant potential field it is subjected to is the superposition of the repulsive potential field and the gravitational potential field. Through the superposition of potential fields, the robot dog 100 can move towards a stable equilibrium point while avoiding points prone to falling. The formation of this composite field is based on the principle of potential field superposition, and the magnitude of the resultant potential field is the vector sum of the gravitational potential field and the repulsive potential field. Its expression is: ; in, represent the strength of the gravitational and repulsive fields respectively.
[0104] The direction of the resultant force determines the posture adjustment direction of the robot dog 100, and the magnitude of the resultant force determines the posture adjustment speed of the robot dog 100. The direction of the repulsive force deviates from the easy-to-fall point.
[0105] In sub-step 2-4-6, the robot guide dog chip 101 detects the local minimum of the posture adjustment. When the net force is zero or very small, but has not yet reached a stable equilibrium point, the posture adjustment of the robot dog 100 is stuck in a local minimum. Random perturbations can be introduced to cause the robot dog 100 to escape the local minimum, or the parameters of the attractive and repulsive forces can be adjusted to change the distribution of the potential field.
[0106] In sub-step 2-4-7, the robot guide dog chip 101 detects the oscillation of the posture adjustment. If the posture adjustment of the robot dog 100 repeatedly jumps sideways in a certain area, the step length parameter can be reduced or the range of the repulsive field can be increased to reduce the oscillation of the posture adjustment.
[0107] In sub-step 2-4-8, the robot guide dog chip 101 determines whether the robot dog 100 has reached a stable equilibrium point. If the distance between the robot dog 100's current position and the stable equilibrium point is less than a certain threshold, the robot dog 100 is considered to have reached a stable equilibrium point, and the algorithm terminates. Otherwise, the algorithm returns to step 2-4-4 to continue calculating the posture resultant force and updating the position.
[0108] In sub-steps 2-4-9, the robot guide dog chip 101 executes the above steps in a loop until the robot dog 100 reaches a stable balance point, that is, maintains posture stability during riding to prevent the user from falling.
[0109] In a preferred embodiment, step 3 includes: Step 3-1: The battery and charging module 106 monitors the battery level in real time. When the battery level is low, the robotic guide dog chip 101 issues a charging alarm via the voice input and output module 102. The robotic guide dog chip 101 calculates a battery status prediction model and uses Kalman filtering to predict the remaining battery life: ; in, Indicates the current time k The estimated value of electrical energy, is the battery attenuation coefficient, Indicates the last moment k The estimated value of electric energy is −1, represents the influence coefficient of control usage on electric energy estimation, is the current power consumption, which is related to the speed v: , represents process noise and represents random error in the model.
[0110] The authorized user issues a charging instruction to the voice input and output module 102. When the robot guide dog chip 101 detects the charging instruction, it triggers the camera 103 to capture the user's face when speaking. The robot guide dog chip 101 extracts the voiceprint feature data of the user issuing the charging instruction and the facial feature data of the user when speaking, and simultaneously analyzes the correlation feature data between the voiceprint feature data when the user issues the charging instruction and the facial feature data when speaking.
[0111] Step 3-2: Receive charging instructions and perform triple feature matching for chip verification; further, the authorized user can also speak to the intelligent rideable robot guide dog through the wireless communication module 107 and the guardian remote monitoring module 110 to issue charging instructions and the location of the charging station.
[0112] The robot guide dog chip 101 matches the voiceprint feature data of the user issuing the charging instruction with the saved voiceprint feature data of the authorized user; further, the robot guide dog chip 101 matches the facial feature data of the user when issuing the charging instruction with the saved facial feature data of the authorized user when speaking; further, the robot guide dog chip 101 matches the correlation feature data between the voiceprint feature data and the facial feature data when the user issues the charging instruction with the saved correlation feature data between the voiceprint feature data and the facial feature data when speaking of the authorized user.
[0113] If the matching degree calculated by the robot guide dog chip 101 is relatively high, the charging instruction is confirmed; otherwise, if the matching degree calculated by the robot guide dog chip 101 is relatively low, the charging instruction is directly ignored, and the robot guide dog chip 101 controls the intelligent rideable robot guide dog to enter a power saving state.
[0114] Step 3-3: Based on the segmented path principle of "current location → safe unloading point → charging pile", use the A* algorithm to optimize the turning points and plan the optimal charging route.
[0115] The robot guide dog chip 101 plans the optimal charging route based on the current charging instruction, and reports the optimal charging route to the user through the voice input and output module 102.
[0116] The user confirms or modifies the optimal charging route through the voice input and output module 102. The robot guide dog chip 101 confirms the user's confirmation or modification of the charging route by matching the voiceprint feature data, matching the facial feature data when speaking, and matching the associated feature data between the two. If the matching degree calculated by the robot guide dog chip 101 is relatively high, the optimal or modified charging route is confirmed.
[0117] Step 3-4: Place the user in a safe location according to the optimal charging route and proceed to charging. Return to the starting point after charging is complete. The guardian is supported to remotely monitor the charging status. Specifically: The robot guide dog chip 101 controls the intelligent rideable robot guide dog to unload the user to a safe location, and then proceeds to the charging pile for charging according to the confirmed optimal or modified charging route. After charging is completed, the robot returns to the safe location where the user unloaded.
[0118] The triple feature matching in step 3-2 of this embodiment will not be described in detail.
[0119] Step 3-3 plans the optimal charging route, which is similar to steps 2-4-1 to 2-4-11 and will not be repeated here.
[0120] In this embodiment, in steps 3-4, the robot guide dog chip 101 controls the intelligent rideable robot guide dog to unload the user to a safe location, then proceed to the charging station for charging according to the confirmed optimal or modified charging route, and return to the safe location where the user unloaded after charging is completed.
[0121] Preferably, the battery and charging module 106 use wireless charging to facilitate use by visually impaired people.
[0122] Furthermore, after charging is completed, the robot guide dog chip 101 follows the optimal charging route L*( , )Return to a safe location Load the user and let the user ride again.
[0123] In this embodiment, if the matching degree calculated by the robot guide dog chip 101 is relatively high, the charging instruction is confirmed; otherwise, if the matching degree calculated by the robot guide dog chip 101 is relatively low, the charging instruction is directly ignored, and the robot guide dog chip 101 controls the intelligent rideable robot guide dog to enter a power saving state.
[0124] The robot guide dog chip 101 plans the optimal charging route according to the current charging instruction, and reports the optimal charging route to the user through the voice input and output module 102.
[0125] In the preferred solution, in step 3 of the charging process, the selection of a safe placement point is based on real-time environmental scanning by the laser radar (104), requiring that there are no dynamic obstacles within 5 meters and the ground flatness is ≤5°, and confirmation is made with the user through voice.
[0126] Battery status prediction uses a Kalman filter algorithm, combined with real-time power consumption (positively correlated with speed, u=kv²) to predict remaining battery life with an error of ≤5%.
[0127] The above embodiments are merely preferred technical solutions of this embodiment and should not be construed as limiting this embodiment. The scope of protection of this embodiment shall be the technical solutions described in the claims, including equivalent alternatives to the technical features of the technical solutions described in the claims. In other words, equivalent alternatives and improvements within this scope are also within the scope of protection of this embodiment.
Claims
1. An intelligent rideable machine guide device, characterized in that: include: A carrying module comprises a robot dog (100) and a seat (109) arranged on the robot dog (100), serving as a carrying body for a user to sit on; The data acquisition module is used to receive user information and transmit it to the control module, specifically including: a voice input and output module (102) for voice interaction; a camera (103) for collecting image data; a laser radar (104) for environmental perception; a GPS positioning module (105) for positioning; and an inertial measurement unit (106) for posture detection. A control module, used for acquiring data from each module and sending corresponding control instructions, specifically comprising: a robot guide dog chip (101), which uses an artificial intelligence algorithm based on user information to generate corresponding control instructions to control the movement of the robot dog (100); A wireless communication module (107) for remote interactive communication and sending and receiving instructions between the control module and other modules; A monitoring and early warning module, including a guardian remote monitoring module (110), is used for remote monitoring; The battery module includes a battery and a charging module (108) for power supply and power management.
2. The intelligent rideable robot blind guiding device according to claim 1, characterized in that: The robot dog (100) is a high-power quadruped robot dog with a standing size of 1000*715*470 (mm), a total weight of 60 kg, a maximum movement speed ≥4 m / s, a maximum climbing angle ≤45°, a step / obstacle height ≥20 cm, a protection grade of IP67, an operating temperature of -20°C to 55°C, a battery life of 2.5-4 hours, and a battery life of ≥10 km; The speech input and output module (102) includes: The microphone is an externally polarized condenser microphone with a built-in 22mm gold-plated large diaphragm pickup capsule and an equivalent noise level of 15dB. The speaker is 4.5W, the sound source response time is less than 250ms, and the maximum volume is 95dB; The robot guide dog chip (101) is an intelligent acceleration card equipped with an ultra-thin AE7100 chip, with a power consumption of 10W, a memory of LPDDR4 / 4x 128bit, a capacity of 8GB / 16GB, compatibility with CUDA, ONNX, and Runtime API, and a computing performance of INT825.6TOPS; The camera (103) is an infrared high-definition camera, with a sensor type of 1 / 2.7" Progressive Scan CMOS, a minimum illumination of 0.002 Lux @ (F1.2, AGC ON) for color, 0 Lux for infrared light on, an infrared distance of ≥50 meters, and a wide dynamic range of 120 dB; The laser radar (104) is a two-dimensional laser radar with a laser wavelength of 905 nm, a scanning angle of 270 degrees, a scanning frequency of 15 Hz / 30 Hz, an angular resolution of 0.1 degrees / 0.3 degrees, and a working area of 0.05 meters to 5 meters.
3. An intelligent rideable machine guiding the blind method, characterized in that: An intelligent rideable machine blind guiding device applicable to the present invention comprises the following steps: Step 1: The user authorizes the first use of the system, extracts the user's voiceprint feature data and facial feature data when speaking, and authorizes the user; Step 2: Authorize the user to issue a riding command, execute the riding command after confirming the user's identity, and control the posture stability during riding; Step 3: Authorize the user to issue a charging command. After confirming the user's identity, the charging command is executed. The user is first placed in a safe location, and then the charging is carried out by the user. Step 4: During driving, when an impassable road surface is detected, an early warning is initiated and relevant operations are performed, specifically: Step 4-1: Collect road condition information during driving; Step 4-2: Perform cluster analysis on the road conditions through artificial intelligence, that is, divide them into passable road surfaces, semi-passable road surfaces, and impassable road surfaces; a passable road surface refers to a road surface that meets the following requirements, that is, artificial intelligence determines through cameras (103) and laser radar (104) that the road surface can pass through a width greater than the width of the intelligent rideable robot guide dog, the uphill and downhill angles are less than the maximum climbing angle of the intelligent rideable robot guide dog, the road surface slope height or obstacle height or pit depth is less than the maximum height of the intelligent rideable robot guide dog's passing ability, and the traffic light is green; a semi-passable road surface refers to a road surface that does not meet the requirements of a passable road surface but meets the following requirements, artificial intelligence determines through cameras (103) and laser radar (104) that the road surface can pass through a width greater than the width of the intelligent rideable robot guide dog, the uphill and downhill angles are less than the maximum climbing angle of the intelligent rideable robot guide dog, the road surface slope height or obstacle height or pit depth is less than the maximum height of the intelligent rideable robot guide dog's passing ability, and the traffic light is green; The optical radar (104) judges whether the width of the road is temporarily occupied by motor vehicles or pedestrians, whether there is congestion of vehicles and pedestrians, whether obstacles will move or are moving, whether large animals are passing, and whether the traffic light is red or yellow. An impassable road surface refers to a road that does not meet the requirements of a traversable road surface or a semi-traversable road surface. The artificial intelligence judges the following through the camera (103) and the laser radar (104): road closure or construction, large-scale and deep water accumulation, roadblocks or obstacles that are too high, road fractures, bridge collapses, fires, serious traffic accidents, and permanent or temporary no-entry signs. Preferably, the artificial intelligence uses a YOLOv7 model, inputs the camera (103) and the laser radar (104) data, and performs joint learning and cluster analysis. Step 4-3: Prioritize drivable roads. If there are none, wait for the semi-drivable roads to become drivable. If there are only impassable roads, issue an alarm. Step 4-4: After the alarm, search for a new route to the destination or charging station, negotiate and confirm with the user, and then drive according to the confirmed route.
4. The intelligent rideable robot blind guiding method according to claim 3, characterized in that: The step 1 comprises: Step 1-1: Collect voice and face data when the user uses it for the first time; Step 1-2: Extract voiceprint feature data; Step 1-3: Extract facial feature data; Step 1-4: Extract the correlation feature data between voiceprint and facial features; Step 1-5: Save authorized user feature data.
5. The intelligent rideable robot blind guiding method according to claim 4, characterized in that: Steps 1-2 extract the user's voiceprint feature data and facial feature data when speaking. Voiceprint recognition uses a full-process process of Gaussian filtering, framing, windowing, FFT transformation, Mel spectrum and DCT transformation to ultimately generate a Mel frequency cepstral coefficient feature matrix, specifically including: Obtain the user's voiceprint feature data, use a Gaussian filter to remove noise, then use a framing function to divide it into frames and then use a windowing function to smooth it; The smoothed signal is converted into a spectrum signal by fast Fourier transform, and the user voice signal with linear spectrum is mapped to the Mel scale using a Mel filter bank. The Mel spectrum of the user voice is then logarithmically operated and the Mel frequency cepstral coefficient features of the user voice S8 are then discrete cosine transformed. Based on the user voice with Mel-frequency cepstral coefficient characteristics, the timbre, pitch, intensity, rhythm and intonation vectors are calculated to obtain the voiceprint feature matrix of the user voice, i.e., the voiceprint feature data S9 of the user voice. The formula is: 。 6. The intelligent rideable robot blind guiding method according to claim 5, characterized in that: In steps 1-3, the robot guide dog chip (101) extracts facial feature data of the user when speaking. The facial feature extraction includes grayscale conversion, mean filtering, and histogram equalization preprocessing, combined with geometric features, texture features, heat map analysis, and lip shape dynamic recognition, specifically including: Collect the user's face image in real time while they are speaking, convert the image to grayscale, then use mean filtering to remove noise, and then perform histogram equalization; Obtain the user's face image M4 after histogram equalization, calculate the vectors of geometry, texture, heat map and mouth shape, and obtain the facial feature matrix when the user speaks, that is, the facial feature data M5 when the user speaks. The formula is: ; A confidence threshold is set for the detection and matching of M5, and it is adjusted according to the different characteristics of daytime, nighttime, rainy and snowy environments.
7. The intelligent rideable robot blind guiding method according to claim 6, characterized in that: In steps 1-4, correlation feature data between the user's voiceprint feature data and the facial feature data when speaking is extracted. The correlation feature analysis fuses the voiceprint and facial data through the YOLOv7 model and introduces the SE attention mechanism to achieve channel weighting. Specifically, it includes: Use the YOLOv7 model and initialize it: set the confidence threshold to 0.25, set the non-maximum suppression threshold to 0.45, and load the pre-trained YOLOv7 model weight file; The YOLOv7 model inputs the data of S9 and S5, aligns them in time, and inputs the original input correlation feature map U of the Head, which is divided into nine channels; Introducing the SE attention mechanism, applying the global average pooling operation to the original input correlation feature map input to the Head Compression is performed on each channel; the original input is associated with the feature map from Compress to The eigenvector Z of , with the number of channels unchanged, is: ; in, For channel c The global average pooling result of ; This is a scalar value indicating the channel c The average value of all pixel values; represents the global average pooling operation; is the original input associated feature map, H is the height of the original input associated feature map, that is, the number of pixels of the original input associated feature map in the vertical direction; W is the width of the original input associated feature map, that is, the number of pixels of the original input associated feature map in the horizontal direction; is the number of channels of the original input associated feature map, that is, the dimension of the original input associated feature map in the depth direction; Input original input associated feature map In the channel ,Location The value at represents the pixel value at a specific position in the original input associated feature map; is the normalization factor; Adaptive recalibration Operation, obtain the weight of [[1,1,C], a total of 9 channels, use the compressed global importance information to generate the excitation weight of each channel, the formula is: in ,s Represents the output channel weight vector, which represents the importance weight of each channel and has a shape of [1,1, C ]; Represents the activation function, which is used to generate channel weights; The global descriptor of the input has a shape of [1,1, C ], obtained by global average pooling; Represents the weight parameters of the model; σ is the activation function; Calculate the output correlation feature map U2 of the attention mechanism, and multiply the generated channel weight with the original input correlation feature map U to achieve channel-level feature weighting, that is, obtain the correlation feature data U2 between the user's voiceprint feature data and the facial feature data when speaking: ; in, Indicates that the output correlation feature map U2 is in channel c The value above is the result after channel weighting; represents the scaling function used to apply the channel weights to the original input associated feature map U; Indicates the original input associated feature map U in channel c The value on Indicates channel c The weight of the activation function Generate, representing the channel c the importance of The output correlation feature map U2 between the user's voiceprint feature data and the facial feature data when speaking is matched according to the preset confidence threshold.
8. The intelligent rideable robot blind guiding method according to claim 3, characterized in that: The step 2 includes: Step 2-1: Receive riding instructions and collect current user characteristics; Step 2-2: Match and verify with the authorization feature; Step 2-3: Plan and confirm the optimal cycling route. The path planning algorithm combines the global A* method with the local dynamic window method to avoid local optimal solutions through adaptive weight adjustment, supporting real-time avoidance of low obstacles and dynamic targets. Steps 2-4: Real-time control of posture stability during riding.
9. The intelligent rideable robot blind guiding method according to claim 3, characterized in that: The step 3 comprises: Step 3-1: The battery module monitors the power level in real time and issues a voice alarm when the power level is below the threshold. Step 3-2: Receive charging instructions and perform chip verification triple feature matching; Step 3-3: Based on the segmented path principle of "current location → safe unloading point → charging station", use the A* algorithm to optimize the turning points and plan the optimal charging route; Step 3-4: Place the user in a safe location according to the optimal charging route and proceed to charge. After charging is completed, the device returns to the starting point, allowing the guardian to remotely monitor the charging status.
10. The intelligent rideable robot blind guiding method according to claim 9, characterized in that: In the step 3, during the charging process, the selection of a safe placement point is based on real-time environmental scanning by the laser radar (104), requiring that there are no dynamic obstacles within 5 meters and that the ground flatness is ≤5°, and the selection is confirmed with the user through voice; Battery status prediction uses a Kalman filter algorithm, combined with real-time power consumption (positively correlated with speed, u=kv²) to predict remaining battery life with an error of ≤5%.