A mobile chassis augmented positioning system and method
Patent Information
- Application Number
- CN202611008212.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-08
AI Technical Summary
[0018]本发明的有益效果包括:本发明能够在标准环境下实现转向定位平均偏差≤1.5°,相比传统开环控制方案精度提升4倍以上;系统待机平均功耗≤18mW,相比视觉持续工作方案续航时间提升300%以上;在高噪声、低光照等极端环境下识别成功率≥88%,可广泛应用于教育机器人、家用清洁机器人、陪伴机器人等量产化产品。
Smart Images

Figure CN122506492B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of positioning control technology, and in particular to a mobile chassis enhanced positioning system and a mobile chassis enhanced positioning method. Background Technology
[0002] As smart mobile devices penetrate the consumer market, the demand for sound-guided positioning technology for low-cost, mass-producible mobile vehicles (such as educational robots, home cleaning robots, and companion robots) is growing.
[0003] In existing technologies, sound source localization-guided robot steering and localization solutions typically rely on high-performance computing platforms such as Raspberry Pi and Jetson, coupled with expensive sensors such as gimbal servo motors and high-precision IMUs (Inertial Measurement Units) to achieve multimodal data fusion and real-time control. However, existing solutions generally suffer from problems such as high hardware costs, poor portability, insufficient rotation control precision, high system power consumption, and weak robustness in complex environments. In particular, they are not suitable for mass production scenarios using low-cost MCUs (Microcontroller Units) with limited resources, such as the ESP32-S3.
[0004] Existing technologies for achieving sound source localization that integrates auditory and visual perception rely on gimbal structures, increasing hardware costs and complexity; or on high-performance computing platforms, which cannot be implemented on low-cost MCUs; or on methods such as visual teaching boards and nonlinear learning machines for sound source angle training, but these methods require complex hardware setups and training processes; or on radar as an auxiliary sensor, which is costly. Furthermore, existing technologies do not optimize chassis rotation control, failing to address the accuracy issues in low-cost chassis rotation control. Therefore, existing technologies have not effectively solved the core problems of simultaneously achieving high-precision closed-loop rotation control, audio-visual fusion localization, and low-power optimization for a mobile chassis on a low-cost MCU platform.
[0005] In order to overcome the above-mentioned defects of the existing technology, there is an urgent need in this field for a mobile chassis enhanced positioning technology that can solve the core problems of existing technologies such as reliance on high-performance hardware, high cost, insufficient rotation control precision, high system power consumption, and weak adaptability to complex environments. Through multi-module collaborative innovation, a high-precision, low-power, and highly robust sound source guided positioning function can be achieved on a low-cost MCU platform. Summary of the Invention
[0006] The following provides a brief overview of one or more aspects to offer a basic understanding of them. This overview is not an exhaustive summary of all conceived aspects, nor is it intended to identify key or decisive elements of all aspects, nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed descriptions that follow.
[0007] To overcome the aforementioned deficiencies in the existing technology, this invention provides a mobile chassis enhanced positioning system and a mobile chassis enhanced positioning method, which can solve the core problems of existing technologies such as reliance on high-performance hardware, high cost, insufficient rotation control precision, high system power consumption, and weak adaptability to complex environments. Through multi-module collaborative innovation, it achieves high-precision, low-power, and highly robust sound source guided positioning function on a low-cost MCU platform.
[0008] Specifically, the mobile chassis enhanced positioning system provided by the first aspect of the present invention includes: a voice wake-up and sound source localization unit configured to acquire a sound source angle; an IMU attitude calculation unit configured to use a built-in digital motion processor to feed back the real-time yaw angle of the mobile chassis; an adaptive PID closed-loop motion control unit configured to calculate a real-time angle deviation value based on the sound source angle and the real-time yaw angle, and dynamically adjust PID parameters and motor PWM signals based on the real-time angle deviation value to drive the mobile chassis to rotate; a visual dynamic wake-up and detection unit configured to perform visual detection when the real-time angle deviation value is less than a threshold to determine the visual angle and its confidence level corresponding to the target; and an audio-visual multimodal decision unit configured to use a three-level confidence multimodal optimal orientation decision method to integrate the sound source angle and the visual angle and their confidence levels to make a final orientation selection.
[0009] Furthermore, in some embodiments of the present invention, the adaptive PID closed-loop motion control unit is also configured to adjust the PID parameters in conjunction with the real-time load of the motor current and / or the vibration data fed back in real time by the IMU attitude calculation unit.
[0010] Furthermore, in some embodiments of the present invention, the three-level confidence multimodal selection orientation decision method includes: determining the sound source angle as the target angle in response to the confidence level being in the low confidence level range; determining the visual angle as the target angle in response to the confidence level being in the high confidence level range; and fusing the sound source angle and the visual angle using a weighted fusion strategy in response to the confidence level being in the medium confidence level range to determine the target angle.
[0011] Furthermore, in some embodiments of the present invention, the audiovisual multimodal decision unit is further configured to reduce the weight of the sound source angle during the weighted fusion process and trigger visual multi-frame cumulative detection to determine the visual angle when the signal-to-noise ratio is less than a preset threshold, and / or reduce the default threshold of the confidence of the visual angle and combine it with the sound source angle to realize the visual detection of the visual dynamic wake-up and detection unit when the illuminance is less than a preset threshold.
[0012] Furthermore, in some embodiments of the present invention, a dynamic power consumption management and reliability assurance module is also included. The dynamic power consumption management and reliability assurance module is configured to use an event-driven mechanism to switch frequency modes and / or to set the task priority of the adaptive PID closed-loop motion control unit to the highest level through preemptive task scheduling.
[0013] Furthermore, according to the second aspect of the present invention, the above-described mobile chassis enhanced positioning method is executed by the above-described mobile chassis enhanced positioning system provided in the first aspect of the present invention. The method includes the steps of: acquiring the sound source angle and the real-time yaw angle of the mobile chassis; calculating a real-time angle deviation value based on the sound source angle and the real-time yaw angle, and dynamically adjusting PID parameters and motor PWM signals based on the real-time angle deviation value to drive the mobile chassis to rotate; performing visual detection when the real-time angle deviation value is less than a threshold to determine the visual angle corresponding to the target and its confidence level; and using a three-level confidence multimodal optimal orientation decision method to integrate the sound source angle and the visual angle and their confidence level to make a final orientation selection.
[0014] Furthermore, in some embodiments of the present invention, the step of calculating the real-time angle deviation value based on the sound source angle and the real-time yaw angle, and dynamically adjusting the PID parameters and the motor PWM signal based on the real-time angle deviation value to drive the mobile chassis to rotate includes: adjusting the PID parameters in combination with the real-time load of the motor current and / or the vibration data fed back in real time by the IMU attitude calculation unit.
[0015] Furthermore, in some embodiments of the present invention, the three-level confidence multimodal selection orientation decision method includes: determining the sound source angle as the target angle in response to the confidence level being in the low confidence level range; determining the visual angle as the target angle in response to the confidence level being in the high confidence level range; and fusing the sound source angle and the visual angle using a weighted fusion strategy in response to the confidence level being in the medium confidence level range to determine the target angle.
[0016] Furthermore, in some embodiments of the present invention, the step of using a three-level confidence multimodal orientation selection method to integrate the sound source angle and the visual angle and their confidence to make a final orientation selection includes: reducing the weight of the sound source angle in the weighted fusion process and triggering visual multi-frame cumulative detection to determine the visual angle when the signal-to-noise ratio is less than a preset threshold; and / or reducing the default threshold of the confidence of the visual angle and combining it with the sound source angle to achieve the visual detection when the illuminance is less than a preset threshold.
[0017] Furthermore, in some embodiments of the present invention, the steps include: using an event-driven mechanism to implement frequency mode switching; and / or setting the task priority of the adaptive PID closed-loop motion control unit to the highest level through preemptive task scheduling.
[0018] The beneficial effects of this invention include: the invention can achieve an average steering positioning deviation of ≤1.5° in a standard environment, which is more than 4 times more accurate than the traditional open-loop control scheme; the average standby power consumption of the system is ≤18mW, which is more than 300% more endurance than the vision continuous working scheme; the recognition success rate is ≥88% in extreme environments such as high noise and low light, and it can be widely used in mass-produced products such as educational robots, household cleaning robots, and companion robots. Attached Figure Description
[0019] The above-described features and advantages of the present invention will be better understood after reading the following detailed description of embodiments of the present disclosure in conjunction with the accompanying drawings. In the drawings, components are not necessarily drawn to scale, and components having similar related properties or features may have the same or similar reference numerals.
[0020] Figure 1 A schematic diagram of a mobile chassis enhanced positioning system provided according to some embodiments of the present invention is shown; Figure 2 An architecture diagram of a mobile chassis augmented positioning system based on connectivity and data interaction, according to some embodiments of the present invention, is shown. Figure 3 A flowchart of a mobile chassis enhanced positioning method according to some embodiments of the present invention is shown; Figure 4 A flowchart of a mobile chassis enhanced positioning method according to some embodiments of the present invention is shown; Figure 5 A flowchart of adaptive PID closed-loop motion control provided according to some embodiments of the present invention is shown; Figure 6 A flowchart of a multimodal selection orientation decision method provided according to some embodiments of the present invention is shown; Figure 7A hardware circuit block diagram of a mobile chassis enhanced positioning system according to some embodiments of the present invention is shown.
[0021] Figure label: 100: Mobile chassis enhanced positioning system; 110: Voice wake-up and sound source localization unit; 111: Voice wake-up module; 112: Voice activity detection module; 113: Sound source localization module; 120: IMU attitude calculation unit; 130: Adaptive PID closed-loop motion control unit; 140: Visual dynamic wake-up and detection unit; 141: Visual inspection module; 142: Visual Reasoning Module; 150: Audio-visual multimodal decision unit; 160: Dynamic power consumption management and reliability assurance module; S310~S340: Steps. Detailed Implementation
[0022] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention is presented in conjunction with preferred embodiments, this does not mean that the features of the invention are limited to these embodiments. On the contrary, the purpose of describing the invention in conjunction with embodiments is to cover other options or modifications that may be derived based on the claims of the present invention. To provide a thorough understanding of the invention, many specific details will be included in the following description. The invention may also be implemented without using these details. Furthermore, to avoid confusion or obscuring the focus of the invention, some specific details will be omitted in the description.
[0023] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0024] Furthermore, the terms "upper," "lower," "left," "right," "top," "bottom," "horizontal," and "vertical" used in the following description should be understood as the orientations shown in the relevant paragraphs and accompanying drawings. These relative terms are for illustrative purposes only and do not imply that the described apparatus must be manufactured or operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0025] It is understood that although terms such as "first," "second," and "third" may be used herein to describe various components, regions, layers, and / or parts, these components, regions, layers, and / or parts should not be limited by these terms, and these terms are only used to distinguish different components, regions, layers, and / or parts. Therefore, the first components, regions, layers, and / or parts discussed below may be referred to as second components, regions, layers, and / or parts without departing from some embodiments of the present invention.
[0026] As mentioned above, existing technologies for guiding robot steering based on sound source localization typically rely on high-performance computing platforms such as Raspberry Pi and Jetson, coupled with expensive sensors such as gimbal servo motors and high-precision IMUs (Inertial Measurement Units) to achieve multimodal data fusion and real-time control. However, existing solutions generally suffer from problems such as high hardware costs, poor portability, insufficient rotation control precision, high system power consumption, and weak robustness in complex environments. In particular, they are not suitable for mass production scenarios using low-cost MCUs (Microcontroller Units) with limited resources, such as the ESP32-S3.
[0027] Existing technologies for achieving sound source localization that integrates auditory and visual perception rely on gimbal structures, increasing hardware costs and complexity; or on high-performance computing platforms, which cannot be implemented on low-cost MCUs; or on methods such as visual teaching boards and nonlinear learning machines for sound source angle training, but these methods require complex hardware setups and training processes; or on radar as an auxiliary sensor, which is costly. Furthermore, existing technologies do not optimize chassis rotation control, failing to address the accuracy issues in low-cost chassis rotation control. Therefore, existing technologies have not effectively solved the core problems of simultaneously achieving high-precision closed-loop rotation control, audio-visual fusion localization, and low-power optimization for a mobile chassis on a low-cost MCU platform.
[0028] To overcome the aforementioned deficiencies in the existing technology, this invention provides a mobile chassis enhanced positioning system and a mobile chassis enhanced positioning method, which can solve the core problems of existing technologies such as reliance on high-performance hardware, high cost, insufficient rotation control precision, high system power consumption, and weak adaptability to complex environments. Through multi-module collaborative innovation, it achieves high-precision, low-power, and highly robust sound source guided positioning function on a low-cost MCU platform.
[0029] The following will refer to Figures 1 to 7 The principles and implementation of this invention are described in detail.
[0030] Please refer to Figure 1 and Figure 2 , Figure 1 A schematic diagram of a mobile chassis enhanced positioning system according to some embodiments of the present invention is shown. Figure 2 An architecture diagram of a mobile chassis augmented positioning system based on connectivity and data interaction, according to some embodiments of the present invention, is shown.
[0031] like Figure 1 As shown, the mobile chassis enhanced positioning system 100 may include a voice wake-up and sound source localization unit 110, an IMU attitude calculation unit 120, an adaptive PID closed-loop motion control unit 130, a visual dynamic wake-up and detection unit 140, an audio-visual multimodal decision unit 150, and a dynamic power consumption management and reliability assurance module 160.
[0032] In some non-limiting embodiments, the mobile chassis enhanced positioning method provided by the present invention can be implemented via the mobile chassis enhanced positioning system 100 provided by the present invention.
[0033] The architecture diagram of the mobile chassis augmented positioning system 100 based on connectivity and data interaction can be shown as follows: Figure 2 As shown, the mobile chassis augmented positioning system 100 can be divided into a control and execution layer, a decision and management layer, a processing and algorithm layer, and a perception layer.
[0034] The working principle of the aforementioned mobile chassis enhanced positioning system 100 will be described below with reference to some embodiments of mobile chassis enhanced positioning methods. Those skilled in the art will understand that these embodiments are merely non-limiting implementations provided by the present invention, intended to clearly demonstrate the main concepts of the invention and provide specific solutions convenient for public implementation, rather than limiting all functions or all operating methods of the system. Similarly, this system is also only one non-limiting implementation provided by the present invention, and does not constitute a limitation on the executing entities and execution order of the steps in these methods.
[0035] Please refer to Figure 3 and Figure 4 , Figure 3and Figure 4 A flowchart of a mobile chassis enhanced positioning method according to some embodiments of the present invention is shown.
[0036] like Figure 3 As shown, the mobile chassis augmented positioning system 100 can perform step S310: acquiring the sound source angle and the real-time yaw angle of the mobile chassis.
[0037] The mobile chassis augmented positioning system 100 can first use the voice wake-up and sound source localization unit 110 to obtain the sound source angle and use the IMU attitude calculation unit 120 to use the built-in digital motion processor to provide feedback on the real-time yaw angle.
[0038] Please refer to the reference. Figure 1 and Figure 4 The voice wake-up and sound source localization unit 110 may include a voice wake-up module 111, a voice activity detection module 112 and a sound source localization module 113. The mobile chassis enhanced positioning system 100 can first be continuously monitored and woke up by the voice wake-up module 111 of the voice wake-up and sound source localization unit 110.
[0039] The voice wake-up module 111 can be located in Figure 2 The processing and algorithm layer in the architecture diagram is connected to the microphone ring array in the perception layer. Preferably, the element spacing of the microphone ring array is 6 cm. The voice wake-up module 111 continuously monitors the external ambient audio through the configured lightweight neural network architecture and wakes up the monitoring to trigger the subsequent sound source localization process when a trigger condition is met (e.g., a preset keyword is detected).
[0040] Preferably, the voice wake-up module 111 can adopt a lightweight keyword detection architecture based on deep separable convolution (CNN) and recurrent neural network (RNN), and reduce the false trigger rate of the mobile chassis enhanced positioning system 100 through a triple verification mechanism of audio preprocessing, Mel frequency cepstral coefficients (MFCC) feature extraction, and confidence evaluation.
[0041] Preferably, the voice wake-up module 111 can enter a duty cycle monitoring mode in the idle state to achieve low-power wake-up triggering. For example, the voice wake-up module 111 can activate 10ms for audio sampling every 200ms to keep the average power consumption below 12mW.
[0042] More preferably, the mobile chassis enhanced positioning system 100 can use the voice activity detection module 112 to detect audio in the environment and eliminate noise. The voice activity detection module 112 can be configured with a dual threshold decision method of spectral entropy and energy, which can ensure a voice detection rate of ≥90% in an environment with a signal-to-noise ratio ≥15dB, and still maintain a detection rate of more than 85% when the signal-to-noise ratio is lower than 10dB, ensuring reliable triggering in noisy environments. The voice activity detection module 112 can adapt well to low signal-to-noise ratio environments, and can transmit the detected voice audio to the voice wake-up module 111. The voice wake-up module 111 analyzes the acquired voice audio to determine whether to wake up the monitoring system.
[0043] Then, as Figure 4 As shown, after the mobile chassis enhanced positioning system 100 is awakened and monitored, the sound source positioning module 113 performs sound source positioning and angle calculation to obtain the sound source angle.
[0044] The sound source localization module 113 can be located in the processing and algorithm layer and connected to the microphone ring array in the perception layer. It is a processing unit that calculates the relative angle of the sound source through the microphone ring array.
[0045] The sound source localization module 113 can use the generalized cross-correlation (GCC-PHAT) algorithm to calculate the signal arrival time difference, and combine the array geometry parameters of the microphone ring array to calculate the horizontal direction angle with the front of the device as the reference, and use this horizontal direction angle as the sound source angle. Here, the sound source horizontal direction angle α is based on the front of the device as 0°, with a range of ±180° and an angular resolution of ±5°.
[0046] Thus, by utilizing the voice wake-up and sound source localization unit 110, the mobile chassis enhanced positioning system 100 can significantly reduce the computational resource consumption of audio processing while acquiring the sound source angle. It can operate stably on a low-cost MCU with an average power consumption of less than 12mW. It still maintains a detection rate of more than 85% in environments with a signal-to-noise ratio of less than 10dB, and has strong triggering reliability in noisy environments.
[0047] The IMU attitude calculation unit 120 can be configured to use a built-in digital motion processor to provide feedback on the real-time yaw angle of the moving chassis.
[0048] The IMU attitude calculation unit 120 can be a sensor component capable of providing real-time attitude data. Preferably, the IMU attitude calculation unit 120 can be implemented using an MPU6050 sensor with a built-in digital motion processor (DMP). The IMU attitude calculation unit 120 can communicate with the main control MCU via an I2C (inter-integrated circuit bus) interface.
[0049] The IMU attitude calculation unit 120 directly outputs attitude information (including pitch, roll, and yaw angles) in quaternion format based on the pre-configured digital motion processor firmware. The main control MCU can read the quaternions at a certain frequency (e.g., 100Hz) and convert them into Euler angles, and further extract the yaw angle of the horizontal rotation of the mobile chassis as the real-time accumulated rotation angle. The real-time cumulative rotation angle This refers to the real-time yaw angle of the mobile chassis.
[0050] The IMU attitude resolution unit 120 utilizes a built-in digital motion processor to achieve... Figure 2 Pose estimation in the processing and algorithm layer.
[0051] Preferably, attitude calculation can be implemented using a complementary filtering algorithm, a sensor data fusion algorithm that can suppress attitude drift and improve angle calculation accuracy. Specifically, gyroscope data from the sensors is used for short-term angle integration, while accelerometer data provides a gravity vector reference to correct for long-term drift-induced deviations in the gyroscope data. Furthermore, the filtering coefficients of the gyroscope data and the accelerometer data can be dynamically adjusted according to the motion state of the moving chassis. When the moving chassis is stationary or moving slowly, the accelerometer is prioritized, with its weight significantly increased; when the moving chassis is rotating, the gyroscope is prioritized, with its weight significantly increased. In some embodiments, the dynamic adjustment of the filtering coefficients can be achieved by setting the accelerometer weight to 0.8 in static conditions and the gyroscope weight to 0.9 in dynamic conditions, thereby controlling the dynamic attitude error within 0.5°. Preferably, the aforementioned sensor data can be mean-filtered using a sliding window of a preset length (e.g., 10) to effectively suppress high-frequency vibration noise while maintaining a low system response delay (e.g., ≤30ms).
[0052] Please continue to refer to this. Figure 3 The mobile chassis enhanced positioning system 100 can perform step S320: calculate the real-time angle deviation value based on the sound source angle and the real-time yaw angle, and dynamically adjust the PID parameters and motor PWM signal based on the real-time angle deviation value to drive the mobile chassis to rotate.
[0053] exist Figure 2 In the embodiment shown, the mobile chassis can be configured with a 12V DC geared motor and a TB6612 motor driver chip to drive the movement of the mobile chassis.
[0054] The mobile chassis enhanced positioning system 100 can use the adaptive PID closed-loop motion control unit 130 to realize the calculation of real-time angle deviation value and the dynamic adjustment of PID parameters and motor PWM signal.
[0055] For the dual-wheel differential chassis design of the mobile chassis, the core of the adaptive PID closed-loop motion control unit 130 is PID (proportional / integral / derivative) closed-loop control based on the real-time yaw angle fed back by the IMU attitude calculation unit 120. Preferably, the control cycle can be strictly 10ms.
[0056] The adaptive PID closed-loop motion control unit 130 uses the target sound source horizontal direction angle α output by the voice wake-up and sound source localization unit 110 as the set value, and the real-time yaw angle output by the IMU attitude calculation unit 120. As feedback value, calculate the real-time angle deviation value. Then based on the real-time angle deviation value By dynamically adjusting the PID parameters and dynamically adjusting the motor differential PWM duty cycle through the PID discretization formula, closed-loop steering control of the chassis is achieved.
[0057] Please refer to Figure 2 and Figure 5 , Figure 5 A flowchart of adaptive PID closed-loop motion control provided according to some embodiments of the present invention is shown.
[0058] like Figure 2 As shown, the adaptive PID closed-loop motion control unit 130 may include a location located at Figure 2 The adaptive PID controller and overshoot prevention mechanism in the control and execution layer.
[0059] The adaptive PID closed-loop motion control unit 130 can use an adaptive PID parameter adjustment algorithm to achieve dynamic adjustment of PID parameters.
[0060] like Figure 5 As shown, the adaptive PID closed-loop motion control unit 130 can be based on real-time angle deviation values. To dynamically set the basic segment ratio coefficient in the adaptive PID parameter adjustment algorithm This is to ensure that the motor output force is greater when the real-time yaw angle differs significantly from the horizontal angle of the sound source, and less when the difference is small. For example, the adaptive PID closed-loop motion control unit 130 can base its output force on the real-time angle deviation value. To set the basic segmented proportional coefficient in the adaptive PID parameter adjustment algorithm in segments In real-time angle deviation value At that time, a scaling factor can be set. It is 0.8; in real-time angle deviation value At that time, a scaling factor can be set. It is 0.5; in real-time angle deviation value At that time, a scaling factor can be set. The value is 0.3. Based on the calculated adaptive PID parameters, a PWM signal is output to drive the motor and rotate the mobile chassis.
[0061] The adaptive PID closed-loop motion control unit 130 can also be set with an anti-overshoot braking mechanism, which performs reverse braking when the real-time angle deviation value is less than a preset threshold. When the adaptive PID closed-loop motion control unit 130 is in use, it can immediately trigger the braking command of the mobile chassis, perform reverse braking, and output a reverse PWM pulse (e.g., a reverse PWM pulse with a 10% duty cycle) for a period of time (e.g., 20ms) to eliminate the overshoot caused by the rotational inertia of the mobile chassis.
[0062] like Figure 2 As shown, the adaptive PID closed-loop motion control unit 130 can also use extended Kalman filtering to fuse the data fed back by the IMU attitude calculation unit 120 with the incremental data of the motor encoder to suppress the deviation caused by wheel slippage and control the final positioning error within ±1°.
[0063] Preferably, the adaptive PID closed-loop motion control unit 130 can further adjust the PID parameters in combination with the real-time load of motor current and / or the vibration data fed back in real time by the IMU attitude calculation unit 120 to adapt to different load and road surface scenarios.
[0064] In some embodiments, the adaptive PID closed-loop motion control unit 130 can adjust the proportional coefficient in conjunction with the real-time load of the motor current. and differential coefficients When the real-time load indicator of the motor current increases (current increases), the adaptive PID closed-loop motion control unit 130 can reduce the proportional coefficient. and integral coefficient This is to reduce the driving force of the mobile chassis under heavy loads. For example, the proportional coefficient increases by 100mA for every 100mA increase in current. Decrease by 0.1, integral coefficient Decrease by 0.02.
[0065] In some embodiments, the adaptive PID closed-loop motion control unit 130 can adjust the proportional coefficient by combining the vibration data fed back in real time by the IMU attitude calculation unit 120. When the vibration data fed back in real time by the IMU attitude calculation unit 120 exceeds the threshold, it indicates that the road surface is bumpy. This will cause the mobile chassis to be in a state of easy slippage or bouncing, reducing the rotation efficiency of the mobile chassis and weakening the control response. At this time, the adaptive PID closed-loop motion control unit 130 can increase the derivative coefficient. This is to compensate for the loss of control force caused by uneven road surfaces. For example, when the vibration amplitude exceeds a threshold, the differential coefficient... Increased by 30%.
[0066] The adaptive PID closed-loop motion control unit 130 can output a motor PWM signal based on the adaptive PID parameters calculated in the above steps to drive the motor to rotate the mobile chassis.
[0067] Thus, addressing the problems of insufficient rotation control precision and poor robustness in existing technologies, the ability of sound source localization to output only a rough angle, large rotational inertia of low-cost dual-wheel differential chassis, easy overshoot in open-loop control, and positioning deviations generally exceeding 8°, as well as the lack of an effective real-time feedback closed-loop control mechanism, the mobile chassis enhanced positioning system 100 provided by this invention utilizes a voice wake-up and sound source localization unit 110, an IMU attitude calculation unit 120, and an adaptive PID closed-loop motion control unit 130 to introduce an adaptive PID closed-loop control strategy based on real-time angle feedback in the rotation control of the mobile chassis. By dynamically adjusting the motor PWM duty cycle through dynamically set PID parameters, smooth acceleration and precise braking are achieved, significantly suppressing rotation overshoot and chattering, and controlling the steering positioning deviation within ±2°. Compared with traditional open-loop control schemes, the accuracy is improved by more than 4 times. Furthermore, compared with existing technologies that cannot dynamically adjust control parameters according to chassis load and ground conditions, and are prone to control performance degradation in different application scenarios, this invention can adapt to different loads and road surface scenarios and has strong anti-interference capabilities.
[0068] Please continue to refer to this. Figure 3 The mobile chassis augmented positioning system 100 can perform step S330: when the real-time angle deviation value is less than a threshold, perform visual detection to determine the visual angle corresponding to the target and its confidence level.
[0069] like Figure 4 As shown, during the process of the adaptive PID closed-loop motion control unit 130 controlling the rotation of the mobile chassis, the mobile chassis enhanced positioning system 100 can continuously monitor the real-time yaw angle and calculate the real-time angle deviation value. When the real-time angle deviation value meets the threshold, the mobile chassis enhanced positioning system 100 can perform visual human detection.
[0070] The mobile chassis augmented positioning system 100 can achieve visual detection using the visual dynamic wake-up and detection unit 140.
[0071] The visual dynamic wake-up and detection unit 140 can be configured with a dynamic wake-up mechanism linked to rotation control. The visual dynamic wake-up and detection unit 140 is in a power-off sleep state by default. When the real-time yaw angle of the moving chassis... Real-time angular deviation value of the horizontal direction angle α of the target sound source The visual dynamic wake-up and detection unit 140 will only be woken up and switched to the working state when the time is right. Specifically, the mobile chassis enhanced positioning system 100 can power the visual dynamic wake-up and detection unit 140 through a MOSFET module to wake up the visual dynamic wake-up and detection unit 140. The switching time from sleep to ready can be less than or equal to 100ms, and the static power consumption is reduced to less than 0.5mW. After the mobile chassis completes the subsequent orientation alignment, the visual dynamic wake-up and detection unit 140 immediately returns to the power-off sleep state.
[0072] Thus, in view of the problems of high power consumption, poor battery life and inability to operate for long periods of time in low-cost lithium battery-powered products caused by the significant increase in computational burden of vision modules in the prior art, the mobile chassis enhanced positioning system 100 provided by the present invention utilizes the configured dynamic wake-up mechanism, the visual dynamic wake-up and detection unit 140 can only activate visual detection when the mobile chassis approaches the direction of the sound source target, thus avoiding high power consumption caused by continuous operation.
[0073] like Figure 1 As shown, the visual dynamic wake-up and detection unit 140 may include a visual detection module 141 and a visual reasoning module 142.
[0074] The visual inspection module 141 can be located as follows: Figure 2 The perception layer in Figure 2 In the embodiment shown, the visual inspection module can be implemented using an OV2640 camera module.
[0075] After the visual dynamic wake-up and detection unit 140 is woken up and switched to the working state, the visual detection module 141 can acquire a single frame image and transmit it to the visual inference module 142. Preferably, when no human body is detected in the single frame image acquired by the visual detection module 141, the visual dynamic wake-up and detection unit 140 can be switched back to the power-off sleep state to further reduce unnecessary power consumption.
[0076] The visual reasoning module 142 can be located in Figure 2 The processing and algorithm layer is described. The visual inference module 142 can be implemented using the MobileNet-SSD lightweight network architecture. The MobileNet-SSD lightweight network architecture is a lightweight neural network model specifically designed for hardware with limited computing power for object detection. When a target (human body) is detected, the target region is marked with a bounding box. This lightweight network architecture can be used for real-time human detection and target localization in the mobile chassis augmented positioning system 100 and achieve efficient inference on a low-cost MCU platform. Through the lightweight neural network model, the visual inference module 142 can guarantee detection efficiency and accuracy.
[0077] The resolution of the images input to the visual inference module 142 is constrained to 320×240. On a resource-constrained platform (such as ESP32-S3) of the mobile chassis augmented positioning system 100, the inference speed of this lightweight network architecture is ≥12fps, meaning the lightweight network architecture can process at least 12 images per second and complete human detection in each image. The default human detection confidence threshold is set to 0.5. After the visual inference module 142 detects a human, it can calculate the pixel offset between the bounding box center and the image center. Combined with the camera's horizontal field of view Calculate the angle of deviation of the target human body relative to the front of the moving chassis, that is, the visual angle β corresponding to the target human body.
[0078] In some embodiments, the visual angle β can be determined based on the following formula: , in, The horizontal width of the input image. In the above embodiment, It can be 320.
[0079] Preferably, in order to accelerate the conversion between pixel offset and visual angle β, the mobile chassis augmented positioning system can also use the checkerboard method to pre-calibrate the camera's intrinsic parameter matrix, so that the visual inference module 142 can quickly convert the angle by looking up a table during online calculation, thereby making the single calculation time less than or equal to 3ms.
[0080] In this way, the visual dynamic wake-up and detection unit 140 completely avoids the high power consumption problem caused by the continuous operation of the visual module in the prior art. The static power consumption of the visual dynamic wake-up and detection unit 140 is reduced to below 0.5mW, and the average standby power consumption of the mobile chassis enhanced positioning system 100 is ≤18mW. Compared with the existing visual continuous operation scheme, the battery life of the present invention is increased by more than 300% while balancing power consumption and response speed.
[0081] Please continue to refer to this. Figure 3 and Figure 4 The mobile chassis enhanced positioning system 100 can perform step S340: using a three-level confidence multimodal optimal orientation decision method to integrate the sound source angle and visual angle and their confidence to make the final orientation selection.
[0082] The mobile chassis augmented positioning system 100 can utilize the audio-visual multimodal decision unit 150 to achieve multimodal angle fusion decision-making, thereby selecting the final orientation of the mobile chassis.
[0083] The audio-visual multimodal decision unit 150 can be located in Figure 2The decision-making and management system of the audio-visual multimodal decision-making unit 150 can adopt a three-level confidence multimodal selection orientation decision-making method to integrate the sound source angle obtained by the voice wake-up and sound source localization unit 110 and the visual angle and their confidence levels obtained by the visual dynamic wake-up and detection unit 140 to make the final orientation selection. Specifically, the three-level confidence multimodal selection orientation decision-making method can include determining the sound source angle as the target angle when the confidence level is in the low confidence level range, determining the visual angle as the target angle when the confidence level is in the high confidence level range, and using a weighted fusion strategy to fuse the sound source angle and visual angle to determine the target angle when the confidence level is in the medium confidence level range.
[0084] Please refer to Figure 6 , Figure 6 A flowchart of a multimodal selection orientation decision method provided according to some embodiments of the present invention is shown.
[0085] like Figure 6 As shown, the audiovisual multimodal decision unit 150 can acquire the target sound source horizontal direction angle α output by the voice wake-up and sound source localization unit 110 and the visual angle β and confidence level of the visual angle β output by the visual dynamic wake-up and detection unit 140. .
[0086] Then, the audio-visual multimodal decision unit 150 can determine the confidence level. Whether it is within the high confidence interval, Figure 6 In the illustrated embodiment, the high confidence interval is the confidence level. When confidence level When the confidence level is within the high confidence range, the audiovisual multimodal decision unit 150 can preferentially determine the visual angle β as the target angle; conversely, when the confidence level is low, the visual angle β is lower than the target angle β. When not within the high confidence interval, the audiovisual multimodal decision unit 150 can determine the confidence level. Further judgment is made based on the interval in which it is located.
[0087] exist Figure 6 In the illustrated embodiment, the medium confidence interval can be The low confidence interval can be The audio-visual multimodal decision unit 150 can determine the confidence level. Whether it is located in the low confidence interval, when the confidence level is... When the target angle is in the low confidence range, the audio-visual multimodal decision unit 150 can automatically switch the target angle to the horizontal direction angle α of the target sound source, that is, use the sound source angle α as the final target angle; conversely, when the confidence level is high... When it is not within the low confidence interval, it indicates the confidence level. To meet the requirements of the medium confidence interval, the audio-visual multimodal decision unit 150 can employ a weighted fusion strategy to fuse the sound source angle and visual angle to determine the target angle. .
[0088] In some embodiments, target angle It can be determined by the following formula: , , in, For weights.
[0089] Preferably, the additional computational load resulting from using the weighted fusion strategy to determine the target angle can be controlled within 5% to ensure that it does not crowd out other computing resources, thereby enabling the mobile chassis augmented positioning system 100 to run on a resource-constrained platform.
[0090] Thus, the audio-visual multimodal decision unit 150 uses a three-level confidence multimodal selection orientation decision method to make the final orientation selection by combining the sound source angle, visual angle and their confidence. When visual detection fails, it automatically switches to sound source localization, thereby improving the robustness of the mobile chassis positioning in complex environments.
[0091] In addition, the audio-visual multimodal decision unit 150 can be designed with an extreme environment audio-visual deep coupling strategy to solve the problems of weak adaptability of existing technologies when facing complex environments and easy failure of detection in extreme environments such as high noise and low light.
[0092] When the signal-to-noise ratio is less than a preset threshold, the audio-visual multimodal decision unit 150 can reduce the weight of the sound source angle in the weighted fusion process and trigger visual multi-frame cumulative detection to determine the visual angle, and / or when the illuminance is less than a preset threshold, the audio-visual multimodal decision unit 150 can reduce the default threshold of the confidence of the visual angle and combine it with the sound source angle to realize the visual detection process of the visual dynamic wake-up and detection unit 140.
[0093] In high-noise environments (such as signal-to-noise ratio ≤ 5dB), the audio-visual multimodal decision unit 150 can automatically reduce the weight of the sound source angle α during the weighted fusion process, and after triggering multi-frame cumulative visual detection (such as 3-frame cumulative detection), the visual angles obtained from the multi-frame visual detection are averaged to determine the final visual angle β, so as to improve the positioning stability in high-noise environments.
[0094] In low-light environments (e.g., illuminance ≤ 50 lux), the confidence level of the visual angle β is automatically reduced. The default threshold (e.g., reduced from 0.5 to 0.3) is set, and the region of interest (ROI) of the visual dynamic wake-up and detection unit 140 during the visual detection process is set as the region corresponding to the sound source based on the sound source angle α. The visual dynamic wake-up and detection unit 140 only performs inference on the region corresponding to the sound source, thereby further improving the detection speed and accuracy in low light environment. In some embodiments, the inference speed can be increased to more than 25fps while reducing the false detection rate.
[0095] In addition, the audio-visual multimodal decision unit 150 can be configured with an anti-oscillation state machine. When the visual dynamic wake-up and detection unit 140 fails visual detection three times in a row, it will automatically disable the visual dynamic wake-up and detection unit 140 for several control cycles (such as 50 control cycles, 500ms) to prevent system instability caused by frequent mode switching of the visual dynamic wake-up and detection unit 140.
[0096] In this way, the audio-visual multimodal decision unit 150 effectively reduces the computational load while ensuring positioning accuracy, with the additional computational load controlled within 5%; it can still maintain stable positioning performance in extreme environments such as high noise and low light, and the system robustness and environmental adaptability are greatly improved.
[0097] After that, as Figure 4 As shown, based on the final orientation determined by the audio-visual multimodal decision unit 150, the mobile chassis augmented positioning system 100 can control the mobile chassis to accurately align with the target.
[0098] In addition, such as Figure 1 and Figure 2 As shown, the mobile chassis enhanced positioning system 100 may further include a dynamic power management and reliability assurance module 160. This dynamic power management and reliability assurance module 160 may be located in... Figure 2 The decision-making and management layers are designed to enable dynamic power consumption management and preemptive task scheduling.
[0099] The dynamic power consumption management and reliability assurance module 160 can use an event-driven mechanism to achieve frequency mode switching of the mobile chassis enhanced positioning system 100.
[0100] The mobile chassis augmented positioning system 100 can dynamically adjust its clock frequency according to the task load, so that it can operate in a high-frequency mode (e.g., 240MHz) during the data processing phase and switch to a low-frequency mode (e.g., 80MHz) in standby mode. After the sound source localization is completed, the dynamic power consumption management and reliability assurance module 160 can immediately put the mobile chassis augmented positioning system 100 into sleep mode, retaining only the timing reading function of the IMU attitude calculation unit 120 and the motor control loop function of the adaptive PID closed-loop motion control unit 130, thereby controlling the overall power consumption fluctuation range of the mobile chassis augmented positioning system 100 within ±15%.
[0101] In this way, by utilizing dynamic power management, the dynamic power management and reliability assurance module 160 can coordinate the energy efficiency optimization mechanism of the working state of each module, and realize the hierarchical sleep and on-demand wake-up of modules through event-driven, so as to balance the system power consumption and response speed.
[0102] The dynamic power consumption management and reliability assurance module 160 can ensure real-time performance through preemptive task scheduling, setting the task priority of the adaptive PID closed-loop motion control unit 130 to the highest level. In some embodiments, the angle control loop task of the adaptive PID closed-loop motion control unit 130 can have the highest priority, followed by the sound source localization task of the voice wake-up and sound source localization unit 110, and the visual processing task of the visual dynamic wake-up and detection unit 140 has the lowest priority, thereby ensuring strict synchronization of the control cycle of the adaptive PID closed-loop motion control unit 130. The data reading of the IMU attitude calculation unit 120 and the motor control output of the adaptive PID closed-loop motion control unit 130 can be triggered by the same timer to avoid timing deviations.
[0103] In this way, through preemptive task scheduling, the dynamic power consumption management and reliability assurance module 160 can prioritize the execution timing of the angle control loop, ensuring the real-time performance and stability of steering control.
[0104] In addition, the dynamic power management and reliability assurance module 160 can be designed with anomaly handling and degradation operation mechanisms. The dynamic power management and reliability assurance module 160 can continuously monitor the working status of each module. When the data of the IMU attitude calculation unit 120 is continuously abnormal for more than a preset number of cycles (such as 10 cycles), it automatically switches the mobile chassis augmented positioning system 100 to pure encoder control mode. When the visual dynamic wake-up and detection unit 140 fails continuously, it automatically switches the mobile chassis augmented positioning system 100 to pure sound source positioning mode to ensure that the mobile chassis augmented positioning system 100 can still maintain basic positioning function when the sensor fails.
[0105] Through the aforementioned dynamic power consumption management and reliability assurance module 160, the mobile chassis enhanced positioning system 100 provided by the present invention can control the overall power consumption fluctuation range within ±15%, resulting in significant energy efficiency optimization. The preemptive scheduling ensures the real-time performance and stability of steering control, and the anomaly handling mechanism ensures that the system can still maintain basic functions when the sensor fails, thus greatly improving reliability.
[0106] The aforementioned mobile chassis enhanced positioning system 100 can use ESP32-S3-WROOM-1-N8R8 as the main control MCU, and the modules work together through standardized data interfaces to achieve accurate sound source guidance and positioning of the low-cost mobile chassis in complex environments.
[0107] The aforementioned mobile chassis enhanced positioning system 100 adopts a layered modular architecture, with all algorithms designed for lightweight operation. All functions can be implemented entirely on a single ESP32-S3 chip, eliminating the need for a high-performance processor and gimbal structure, resulting in a minimalist hardware solution. Compared to existing solutions that rely on high-performance processors and gimbal structures, which suffer from high hardware costs, poor portability, and inability to be deployed on low-cost MCUs, this invention reduces hardware BOM (Bill of Materials) costs by more than 70% compared to existing technologies. It also boasts strong portability, perfectly adapting to consumer-grade mass production applications, and solving the core pain points of existing solutions—high costs and inability to be deployed on low-cost MCUs.
[0108] Please refer to Figure 7 , Figure 7 A hardware circuit block diagram of a mobile chassis enhanced positioning system according to some embodiments of the present invention is shown.
[0109] exist Figure 7 In the embodiment shown, the microphone ring array and inertial measurement unit (IMU) of the sensing layer can be connected to the main control MCU and used as inputs to the main control MCU. The main control MCU can use the MOSFET module to control the camera of the sensing layer and use the H-bridge motor drive to realize closed-loop control of the chassis differential motor.
[0110] During operation, the voice wake-up and sound source localization unit 110 of the mobile chassis augmented positioning system 100 continuously monitors the ambient audio. Upon detecting preset keywords, it triggers the sound source localization process, calculates the target angle of the sound source, and sends it to the adaptive PID closed-loop motion control unit 130. The adaptive PID closed-loop motion control unit activates the PID closed-loop control loop, driving the chassis to rotate based on the attitude data fed back in real time by the IMU attitude calculation unit 120. When the chassis approaches the target direction, the visual dynamic wake-up and detection unit 140 is activated to perform human detection and angle correction. Then, the audiovisual multimodal decision unit 150 determines the final target angle based on the audiovisual data, controlling the chassis to accurately align with the target. After positioning is completed, each module enters its corresponding sleep state.
[0111] In some embodiments, the aforementioned mobile chassis enhanced positioning system 100 is repeatedly tested 100 times in a standard indoor environment (temperature 25°C, signal-to-noise ratio 20dB, illuminance 500 lux). The core test results are as follows: average steering positioning deviation 1.2°, maximum deviation 1.8°, overshoot rate 2%; average standby power consumption 17.8mW, peak operating power consumption 118mW, 1000mAh lithium battery can last up to 42 hours; positioning success rate ≥88% in extreme environments such as high noise and low light.
[0112] This invention can achieve an average steering and positioning deviation of ≤1.5° under standard conditions, which is more than 4 times more accurate than the traditional open-loop control scheme; the average standby power consumption of the system is ≤18mW, which is more than 300% more endurance than the vision continuous working scheme; the recognition success rate is ≥88% in extreme environments such as high noise and low light, and it can be widely used in mass-produced products such as educational robots, home cleaning robots, and companion robots.
[0113] In summary, the mobile chassis enhanced positioning system and method provided by this invention can solve the core problems of existing technologies, such as reliance on high-performance hardware, high cost, insufficient rotation control precision, high system power consumption, and weak adaptability to complex environments. Through multi-module collaborative innovation, it can achieve high-precision, low-power, and highly robust sound source guided positioning function on a low-cost MCU platform.
[0114] Although the methods described above are illustrated and depicted as a series of actions for the sake of simplicity, it should be understood and appreciated that these methods are not limited by the order of the actions, as some actions may occur in a different order and / or concurrently with other actions from the illustrations and descriptions herein or not illustrated and described herein but which may be understood by those skilled in the art, according to one or more embodiments.
[0115] Those skilled in the art will understand that information, signals, and data can be represented using any of a variety of different techniques and skills. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof.
[0116] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in a generalized manner in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention.
[0117] The various illustrative logic modules and circuits described in conjunction with the embodiments disclosed herein may be implemented or performed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but in alternatives, it may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.
[0118] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read and write information to / from the storage medium. In an alternative, the storage medium may be integrated into the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and storage medium may reside as discrete components in the user terminal.
[0119] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functionality may be stored or transmitted as one or more instructions or code on or through a computer-readable medium. A computer-readable medium includes both computer storage media and communication media, encompassing any medium that facilitates the transfer of a computer program from one location to another. A storage medium may be any available medium accessible to a computer. By way of example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible to a computer. Any connection is also legitimately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of a medium. As used in this article, disk and disc include compact discs (CDs), laser discs, optical discs, digital multi-purpose discs (DVDs), floppy disks, and Blu-ray discs. Disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of these should also be included within the scope of computer-readable media.
[0120] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A mobile chassis enhanced positioning system, characterized in that, include: A voice wake-up and sound source localization unit, wherein the voice wake-up and sound source localization unit is configured to acquire the sound source angle; The IMU attitude calculation unit is configured to use a built-in digital motion processor to provide feedback on the real-time yaw angle of the moving chassis. An adaptive PID closed-loop motion control unit is configured to calculate a real-time angle deviation value based on the sound source angle and the real-time yaw angle, and dynamically adjust the PID parameters and motor PWM signal based on the real-time angle deviation value to drive the mobile chassis to rotate. A visual dynamic wake-up and detection unit is configured to perform visual detection when the real-time angle deviation value is less than a threshold to determine the visual angle corresponding to the target and its confidence level. as well as An audio-visual multimodal decision unit is configured to use a three-level confidence multimodal orientation selection method to integrate the sound source angle and the visual angle and their confidence levels to make a final orientation selection. The three-level confidence multimodal orientation selection method includes: determining the sound source angle as the target angle in response to the confidence level being in a low confidence range; determining the visual angle as the target angle in response to the confidence level being in a high confidence range; and fusing the sound source angle and the visual angle using a weighted fusion strategy in response to the confidence level being in a medium confidence range to determine the target angle.
2. The mobile chassis enhanced positioning system as described in claim 1, characterized in that, The adaptive PID closed-loop motion control unit is also configured to adjust the PID parameters in conjunction with the real-time load of the motor current and / or the vibration data fed back in real time by the IMU attitude calculation unit.
3. The mobile chassis enhanced positioning system as described in claim 1, characterized in that, The audiovisual multimodal decision unit is further configured to reduce the weight of the sound source angle during the weighted fusion process and trigger visual multi-frame cumulative detection to determine the visual angle when the signal-to-noise ratio is less than a preset threshold, and / or reduce the default threshold of the confidence of the visual angle and combine it with the sound source angle to realize the visual detection of the visual dynamic wake-up and detection unit when the illuminance is less than a preset threshold.
4. The mobile chassis enhanced positioning system as described in claim 1, characterized in that, It also includes a dynamic power consumption management and reliability assurance module, which is configured to use an event-driven mechanism to switch frequency modes and / or set the task priority of the adaptive PID closed-loop motion control unit to the highest through preemptive task scheduling.
5. A mobile chassis enhanced positioning method, characterized in that, The method is performed by the mobile chassis augmentation positioning system as described in any one of claims 1 to 4, and the method includes the following steps: Obtain the sound source angle and the real-time yaw angle of the moving chassis; Calculate the real-time angle deviation value based on the sound source angle and the real-time yaw angle, and dynamically adjust the PID parameters and motor PWM signal based on the real-time angle deviation value to drive the mobile chassis to rotate. When the real-time angle deviation value is less than the threshold, visual detection is performed to determine the visual angle corresponding to the target and its confidence level; as well as A three-level confidence multimodal orientation selection method is used to integrate the sound source angle, the visual angle, and their confidence levels to make the final orientation selection. The three-level confidence multimodal orientation selection method includes: In response to the confidence level being in the low confidence range, the sound source angle is determined as the target angle; In response to the confidence level being in the high confidence interval, the visual angle is determined to be the target angle; and In response to the confidence level being in the middle confidence range, a weighted fusion strategy is used to fuse the sound source angle and the visual angle to determine the target angle.
6. The mobile chassis enhanced positioning method as described in claim 5, characterized in that, The step of calculating the real-time angle deviation value based on the sound source angle and the real-time yaw angle, and dynamically adjusting the PID parameters and motor PWM signal based on the real-time angle deviation value to drive the mobile chassis to rotate includes: The PID parameters are adjusted by combining the real-time load of the motor current and / or the vibration data fed back in real time by the IMU attitude calculation unit.
7. The mobile chassis enhanced positioning method as described in claim 5, characterized in that, The steps of using a three-level confidence multimodal orientation selection method to combine the sound source angle, the visual angle, and their confidence levels to make the final orientation selection include: If the signal-to-noise ratio is less than a preset threshold, reduce the weight of the sound source angle during the weighted fusion process and trigger visual multi-frame cumulative detection to determine the visual angle; and / or When the illuminance is less than a preset threshold, the default threshold of the confidence level of the visual angle is reduced, and the visual detection is achieved by combining the sound source angle.
8. The mobile chassis enhanced positioning method as described in claim 5, characterized in that, It also includes the following steps: Frequency mode switching is implemented using an event-driven mechanism; and / or The task priority of the adaptive PID closed-loop motion control unit is set to the highest through preemptive task scheduling.
Citation Information
Patent Citations
Heading control device and method based on electronic differential chassis
CN112141210A
Positioning control strategy for automatic charging process of Mecanum wheel mobile chassis based on IMU and camera
CN119806130A