Multi-scene intelligent control system based on intelligent voice recognition technology
By using a multi-scenario intelligent control system based on intelligent voice recognition technology, combined with drones and ground vehicles, air-land coordination and multi-mode motion switching are achieved, solving the perception and mobility problems of unmanned systems in complex environments and improving the flexibility and efficiency of mission execution.
Patent Information
- Application Number
- CN202511597874.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2026-03-03
AI Technical Summary
Existing unmanned systems have limited perception range and insufficient mobility in complex environments. Remote control operation requires high precision and is easily affected by line-of-sight obstruction and signal interference. They also lack the flexibility to respond to emergencies and the convenience of human-machine interaction.
The system employs a multi-scenario intelligent control system based on intelligent voice recognition technology. It combines an air-land collaborative platform, a hierarchical hardware control architecture, a multimodal interaction module, and a collaborative task management module. It integrates drones and ground vehicles, supports remote voice and web-based control and visual autonomous decision-making, and designs a flexible priority arbitration mechanism to achieve air-land collaboration and multi-mode motion switching.
It improves the environmental adaptability and task completion rate of unmanned systems in complex environments, lowers the operational threshold, enhances the flexibility and efficiency of task execution, and strengthens the robustness and scalability of the system.
Smart Images

Figure CN121596791A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent device technology, specifically to a multi-scenario intelligent control system based on intelligent voice recognition technology. Background Technology
[0002] With the continuous development of robotics and artificial intelligence, unmanned systems are increasingly being used in complex scenarios such as disaster relief, environmental exploration, and security patrols. These scenarios typically involve harsh environments and complex terrains, placing extremely high demands on the adaptability and intelligence of unmanned platforms. Currently, mainstream unmanned systems are mainly divided into two categories: one is unmanned aerial vehicles (UAVs) focused on aerial operations, whose advantages lie in obstacle crossing and wide field of vision; the other is all-terrain vehicles focused on ground transportation, whose advantages lie in strong load capacity and relatively long endurance.
[0003] However, single aerial or ground platforms often struggle to complete tasks efficiently in complex environments due to limited perception range and insufficient maneuverability. Existing unmanned systems mostly rely on traditional remote control or autonomous operation based on preset programs. Remote control requires highly skilled operators and is susceptible to line-of-sight obstruction and signal interference in complex environments, while fully autonomous operation lacks the flexibility to handle emergencies and the convenience of human-machine interaction. Therefore, we propose a multi-scenario intelligent control system based on intelligent voice recognition technology. Summary of the Invention
[0004] The purpose of this invention is to provide a multi-scenario intelligent control system based on intelligent voice recognition technology to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, a multi-scenario intelligent control system based on intelligent voice recognition technology is provided, including an air-land collaborative platform, a hierarchical hardware control architecture, a multimodal interaction module, and a collaborative task management module. The air-land collaborative platform includes a drone platform with flight mode and a vehicle platform with ground-based capabilities, and the drone platform and the vehicle platform are connected by a detachable connection structure. The hierarchical hardware control architecture includes a first processor as an upper-level decision-making unit, a second processor as a lower-level motion control unit, and a power module that supplies power to the air-land collaborative platform. The multimodal interaction module is integrated into the hierarchical hardware control architecture and is used to receive and process voice control commands from users, web page control commands from remote clients, and visual data commands from environmental perception sensors. The collaborative task management module runs on the first processor and is configured to generate and assign task instructions to the UAV platform and the vehicle platform based on the instructions of the multimodal interaction module and the data of the environmental perception sensor.
[0006] Preferably, the hierarchical hardware control architecture specifically includes: The first processor is a Raspberry Pi running the Linux operating system, which carries image processing, advanced decision-making algorithms, and collaborative task management modules; The second processor is an STM32 microcontroller, which is directly connected to and controls the motor, servo motor and inertial measurement unit, executes motion commands from the first processor, and realizes attitude stabilization control of the aircraft or vehicle body; The unmanned aerial vehicle platform and the vehicle platform exchange commands and data via a wireless communication module.
[0007] Preferably, the multimodal interaction module includes a voice control submodule, which specifically comprises: By integrating a mature speech recognition chip and software library, it receives users' voice commands; The recognized voice commands are validated and semantically parsed, and then converted into preset control commands. By designing control algorithms that include error handling and fault tolerance mechanisms, control commands are transformed into executable actions for drone or vehicle platforms, and real-time voice feedback is provided.
[0008] Preferably, the system supports a priority arbitration mechanism for multimodal commands. When voice control commands, web page control commands, and self-generated visual tracking commands are received simultaneously, the system executes the command with the highest priority according to preset priority rules, which are configured by the user through a web page.
[0009] Preferably, a multi-mode motion switching module is also included, wherein the multi-mode motion switching module specifically comprises: Based on the data collected by the environmental perception sensor, the current terrain complexity or obstacle information is determined in real time; When it is determined that the ground obstacle is insurmountable, a mode switching command is generated to control the air-land cooperative platform to switch from the ground tire driving mode to the short-distance, low-altitude flight mode in order to overcome the obstacle; After overcoming the obstacle, the platform is controlled to smoothly switch back from flight mode to ground driving mode.
[0010] Preferably, the environmental perception sensor includes a camera, an ultrasonic sensor, and an inertial measurement unit; the system further includes a sensor data fusion and processing module, which specifically comprises: Time synchronization and coordinate system unification of data from different sensors; A filtering algorithm is used to fuse multi-source data to generate comprehensive perception information on the surrounding environment, obstacle positions, and the user's own posture. Based on the comprehensive sensing information, a decision-making basis is provided for autonomous tracking and multi-mode motion switching.
[0011] Preferably, the collaborative task management module supports collaborative strategies including split strategies and collaborative strategies. The specific separation strategy is as follows: the UAV platform and the vehicle platform are physically separated and independently execute different tasks, and the UAV platform and the vehicle platform share environmental information or task status through wireless communication. The specific collaborative strategy is as follows: the UAV platform conducts wide-area reconnaissance and navigation in the air, the vehicle platform performs precision operations or material transportation on the ground, and the UAV platform and the vehicle platform collaborate to complete the same task based on shared information.
[0012] Preferably, a remote interaction and data display platform is also included. This platform is a web application built using the Vue and Spring Boot frameworks, which allows users to remotely monitor the real-time status, sensor data, and task execution screens of the air-land collaborative platform, and to issue remote commands.
[0013] A task execution method for a multi-scenario intelligent control system based on intelligent speech recognition technology as described in any one of the above-mentioned methods, comprising the following steps: S1. The system receives task triggering instructions through the multimodal interaction module; S2. The collaborative task management module plans the task execution strategy based on the instruction type and current environment information, and determines the collaborative mode between the UAV platform and the vehicle platform. S3. The environmental perception sensor continuously collects environmental data, which is then analyzed by the sensor data fusion and processing module and fed back to the collaborative task management module for dynamic decision-making and adjustment. S4, the collaborative task management module sends specific motion commands to the corresponding processors in the hierarchical hardware control architecture, driving the air-land collaborative platform to execute tasks; S5. Status data during task execution is uploaded in real time to the remote interaction and data display platform for user monitoring and intervention.
[0014] Compared with the prior art, the beneficial effects of the present invention are: This invention creatively combines unmanned aerial vehicles (UAVs) and ground vehicles into an air-ground collaborative platform, overcoming the functional limitations of a single platform. The UAVs endow the system with the ability to traverse vertical obstacles and acquire high-altitude visibility, while the ground vehicles provide a stable foundation for ground movement and operations. Combined with a multi-mode motion switching strategy, the system can intelligently select between driving and flying based on the terrain, thereby achieving unprecedented environmental adaptability and mission completion rates in complex scenarios such as disaster relief and field exploration.
[0015] This invention integrates three interaction modes: voice control, remote web-based operation, and visual autonomous decision-making. It also features a flexible priority arbitration mechanism, significantly lowering the operational threshold and enabling even non-professional users to issue commands via natural voice. Furthermore, it ensures timely response to higher-level instructions in emergency situations. Simultaneously, the collaborative task management module intelligently schedules air and ground units to execute separate or collaborative strategies, such as a "air reconnaissance-ground response" collaborative mode. This deeply integrates human decision-making wisdom with the machine's autonomous execution capabilities, significantly improving the flexibility and efficiency of task execution.
[0016] This invention employs a hierarchical hardware control architecture and modular software design, decoupling complex perception, decision-making, and control tasks, thus balancing the performance of the system in handling complex algorithms with the real-time performance and stability of the underlying control. Through sensor data fusion and processing, the accuracy of environmental perception is improved. Furthermore, the remote interaction platform based on Vue and Spring Boot facilitates functional expansion and secondary development, ensuring not only the system's robustness but also providing a low-cost, high-performance, and reusable platform solution for technological innovation in fields such as intelligent robots and unmanned systems. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the invention; Figure 2 This is a flowchart of the method of the present invention; Figure 3 This is a speech recognition control logic diagram of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Please see Figure 1-3 The present invention provides a technical solution: a multi-scenario intelligent control system based on intelligent voice recognition technology, including an air-land collaborative platform, a hierarchical hardware control architecture, a multimodal interaction module, and a collaborative task management module.
[0020] The air-land collaborative platform includes a drone platform with flight mode and a vehicle platform with ground-based capabilities, and the drone platform and the vehicle platform are connected by a detachable connection structure.
[0021] It should be noted that the air-land collaborative platform is composed of a quadcopter drone platform and a four-wheeled ground vehicle platform connected by a detachable mechanical snap-fit structure. The drone platform adopts a lightweight carbon fiber frame, which is mounted on a brushless motor, electronic speed controller and propeller; the vehicle platform adopts a high-strength plastic body and is equipped with shock-absorbing Mecanum wheels and its drive motor.
[0022] The hierarchical hardware control architecture includes a first processor as the upper-level decision-making unit, a second processor as the lower-level motion control unit, and a power supply module to power the air-land collaborative platform. Specifically, the hierarchical hardware control architecture includes: the first processor is a Raspberry Pi running the Linux operating system, which carries image processing, advanced decision-making algorithms, and collaborative task management modules; the second processor is an STM32 microcontroller, which directly connects to and controls motors, servos, and inertial measurement units, executes motion commands from the first processor, and realizes attitude stabilization control of the aircraft or vehicle; the UAV platform and the vehicle platform interact with each other through a wireless communication module.
[0023] It should be noted that the upper-level decision-making unit uses a Raspberry Pi 4B as the primary processor, running the Ubuntu Linux operating system. The Raspberry Pi 4B is responsible for running complex software modules, including: a web server backend, a visual processing algorithm (OpenCV), a lightweight neural network model (YOLOv5-lite), a collaborative task management module, and a communication interface with the web frontend. It connects to the camera via USB and communicates with the lower-level controller via GPIO ports. The lower-level motion control unit uses an STM32F407 microcontroller as the secondary processor. The STM32 is responsible for tasks with high real-time requirements. It reads data from the inertial measurement unit via the I2C bus, generates PWM signals to precisely control the drone's motors and the vehicle's servos, and runs a PID control algorithm to achieve attitude stabilization control of the aircraft or vehicle. The drone and the vehicle platform exchange commands and status data through an NRF24L01 wireless radio frequency module. The power module uses a high-energy-density lithium polymer battery and is equipped with corresponding voltage conversion circuits to provide stable power at different voltage levels to the Raspberry Pi, STM32, sensors, and actuators.
[0024] The multimodal interaction module is integrated into a hierarchical hardware control architecture. It receives and processes voice control commands from users, web control commands from remote clients, and visual data commands from environmental perception sensors. The multimodal interaction module includes a voice control submodule, which specifically: receives user voice commands by integrating a mature voice recognition chip and software library; verifies the validity and semantics of the recognized voice commands, converting them into preset control commands; and transforms the control commands into executable actions for the UAV or vehicle platform by designing a control algorithm that includes error handling and fault tolerance mechanisms, providing real-time voice feedback. The system supports a priority arbitration mechanism for multimodal commands. When voice control commands, web control commands, and self-generated visual tracking commands are received simultaneously, the system executes the highest priority command according to preset priority rules, which are configured by the user through the web interface.
[0025] It should be noted that the voice control submodule uses a mature offline speech recognition chip, which is connected to the STM32. During implementation, a set of key command words (such as "take off," "forward," "turn left," "separate," and "reconnaissance") are pre-programmed into the chip. When the user issues a voice command, the LD3320 recognizes it and sends the recognition result to the Raspberry Pi via serial port. The voice control service on the Raspberry Pi performs validity verification and semantic parsing of the command, converting it into a specific control command (such as {"action":"take_off"}). Then, using an algorithm that includes error handling and fault tolerance mechanisms, the command is issued to the collaborative task management module. Simultaneously, real-time voice feedback such as "command received" can be provided through the connected speech synthesis module. A web application is built using the Vue.js front-end framework and the Spring Boot back-end framework. Users access the webpage through a browser and can see real-time video streams, sensor data, and virtual control buttons. Control commands generated by clicking buttons or dragging a joystick are sent in real-time to the Spring Boot service on the Raspberry Pi via the WebSocket protocol. An OpenCV program running on the Raspberry Pi continuously captures camera footage and calls the deployed YOLOv5-lite model for real-time target detection (such as identifying "injured persons" and "fire sources"). Upon target detection, visual tracking instructions are automatically generated. Set up a priority arbiter within the collaborative task management module. The default priority rules are: emergency visual commands (such as obstacle avoidance) > web-based manual commands > voice commands > regular autonomous visual commands. If the system is automatically cruising and the user suddenly commands "emergency return" via the web interface, the arbiter will immediately interrupt the cruising and execute the return. The user can customize this priority order in the "Settings" page on the web interface.
[0026] The collaborative task management module runs on the first processor and is configured to generate and distribute task instructions to the UAV platform and the vehicle platform based on the instructions of the multimodal interaction module and the data of the environmental perception sensor. The collaborative strategies supported by the collaborative task management module include split strategy and collaborative strategy. The separate strategy is as follows: the UAV platform and the vehicle platform are physically separated and independently perform different tasks, and the UAV platform and the vehicle platform share environmental information or task status through wireless communication; the collaborative strategy is as follows: the UAV platform conducts wide-area reconnaissance and navigation in the air, the vehicle platform conducts precision operations or material transportation on the ground, and the UAV platform and the vehicle platform cooperate to complete the same task based on shared information.
[0027] It should be noted that the collaborative task management module dynamically invokes different collaborative strategies based on received instructions and sensing data. When the split strategy is implemented, the user gives the voice command "UAV reconnaissance, vehicle standby", the collaborative task management module controls the connection mechanism to unlock, the UAV takes off and hovers in the air to reconnoiter the designated area, and transmits the real-time image back to the Raspberry Pi on the vehicle via NRF24L01, and then the Raspberry Pi transmits it back to the web page, while the vehicle remains on standby. When the collaborative strategy is implemented, the UAV takes off first when the system performs a "search and rescue" mission. It uses its aerial vision advantage to conduct a wide-area search. After finding the target, it sends the GPS coordinates or visual features to the vehicle via wireless communication. After receiving the information, the vehicle autonomously plans a path to drive to the target point for close-range confirmation or material delivery. During this period, the UAV continuously provides aerial relay and environmental monitoring for the vehicle.
[0028] The environmental perception sensors include a camera, an ultrasonic sensor, and an inertial measurement unit; the system also includes a sensor data fusion and processing module, which specifically comprises: The system synchronizes data from different sensors in time and coordinates them; it uses filtering algorithms to fuse multi-source data and generate comprehensive perception information on the surrounding environment, obstacle positions, and its own attitude; based on the comprehensive perception information, it provides decision-making basis for autonomous tracking and multi-mode motion switching.
[0029] It should be noted that the sensor data fusion and processing module is implemented in software on the Raspberry Pi. First, the acceleration / angular velocity data from the IMU, the distance data from the ultrasonic sensor, and the visual odometry data from the camera are timestamped and their coordinates are transformed to the same vehicle coordinate system. Then, the Kalman filter algorithm is used to fuse these multi-source data to generate more accurate and stable comprehensive environmental information.
[0030] It also includes a multi-mode motion switching module, which specifically includes: Based on data collected by environmental perception sensors, the system can determine the current terrain complexity or obstacle information in real time. When it is determined that a ground obstacle is insurmountable, a mode switching command is generated to control the air-land cooperative platform to switch from ground tire driving mode to short-distance, low-altitude flight mode to overcome the obstacle. After overcoming the obstacle, the control platform smoothly switches back from flight mode to ground driving mode.
[0031] It should be noted that the multi-mode motion switching module makes decisions based on the sensor fusion results. When the vehicle is in driving mode, if the fused data determines that there is a ditch ahead with a width exceeding the length of the vehicle, the multi-mode motion switching module will generate a mode switching command. The STM32 first controls the vehicle to stop, then unlocks the drone motors, performs a vertical takeoff to a predetermined low altitude (e.g., 2 meters), flies a short distance over the ditch, and then performs a landing. After landing, the drone motors stop, the vehicle motors are restarted, and the system switches back to ground mode.
[0032] It also includes a remote interaction and data display platform, which is a web application built using the Vue and Spring Boot frameworks. This platform allows users to remotely monitor the real-time status, sensor data, and task execution screens of the air-land collaborative platform, and to issue remote commands.
[0033] A task execution method for a multi-scenario intelligent control system based on intelligent speech recognition technology for any of the above-mentioned applications includes the following steps: S1. The system receives task trigger commands through the multimodal interaction module; S2. The collaborative task management module plans the task execution strategy based on the command type and current environmental information, and determines the collaborative mode between the UAV platform and the vehicle platform; S3. Environmental perception sensors continuously collect environmental data, which is analyzed by the sensor data fusion and processing module and then fed back to the collaborative task management module for dynamic decision-making and adjustment; S4. The collaborative task management module issues specific motion commands to the corresponding processors in the hierarchical hardware control architecture to drive the air-land collaborative platform to execute the task; S5. Status data during task execution is uploaded in real time to the remote interaction and data display platform for user monitoring and intervention.
[0034] It should be noted that this invention takes a search and rescue mission at a disaster site as an example. During operation, rescuers issue a "start search and rescue" command via web page or voice. The collaborative task management module parses the command, determines the "collaborative strategy," and instructs the drone to take off for global reconnaissance, while the vehicle follows on the ground and provides relay support. The camera on the drone transmits images in real time, and the YOLOv5-lite model identifies targets with suspected vital signs. At the same time, ultrasonic and IMU data are fused to sense the distance between the drone and obstacles, and this information is fed back to the management module in real time. The management module dynamically adjusts the strategy: it commands the drone to hover and mark the target location, while simultaneously commanding the vehicle to move towards the target point. During the vehicle's movement, if the sensors detect a collapsed obstacle ahead, the multi-mode motion switching module is triggered, and the control platform briefly flies over the obstacle. Throughout the process, information such as platform status, video stream, and target location is displayed in real time on the rescuers' web page interface, and rescuers can issue new commands at any time via web page or voice as needed.
[0035] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multi-scenario intelligent control system based on intelligent voice recognition technology, characterized in that, It includes an air-land collaborative platform, a layered hardware control architecture, a multimodal interaction module, and a collaborative task management module; The air-land collaborative platform includes a drone platform with flight mode and a vehicle platform with ground-based capabilities, and the drone platform and the vehicle platform are connected by a detachable connection structure. The hierarchical hardware control architecture includes a first processor as an upper-level decision-making unit, a second processor as a lower-level motion control unit, and a power module that supplies power to the air-land collaborative platform. The multimodal interaction module is integrated into the hierarchical hardware control architecture and is used to receive and process voice control commands from users, web page control commands from remote clients, and visual data commands from environmental perception sensors. The collaborative task management module runs on the first processor and is configured to generate and assign task instructions to the UAV platform and the vehicle platform based on the instructions of the multimodal interaction module and the data of the environmental perception sensor.
2. The multi-scenario intelligent control system based on intelligent voice recognition technology according to claim 1, characterized in that, The hierarchical hardware control architecture specifically includes: The first processor is a Raspberry Pi running the Linux operating system, which carries image processing, advanced decision-making algorithms, and collaborative task management modules; The second processor is an STM32 microcontroller, which is directly connected to and controls the motor, servo motor and inertial measurement unit, executes motion commands from the first processor, and realizes attitude stabilization control of the aircraft or vehicle body; The unmanned aerial vehicle platform and the vehicle platform exchange commands and data via a wireless communication module.
3. The multi-scenario intelligent control system based on intelligent voice recognition technology according to claim 1, characterized in that, The multimodal interaction module includes a voice control submodule, which specifically comprises: By integrating a mature speech recognition chip and software library, it receives users' voice commands; The recognized voice commands are validated and semantically parsed, and then converted into preset control commands. By designing control algorithms that include error handling and fault tolerance mechanisms, control commands are transformed into executable actions for drone or vehicle platforms, and real-time voice feedback is provided.
4. The multi-scenario intelligent control system based on intelligent voice recognition technology according to claim 1, characterized in that, The system supports a priority arbitration mechanism for multimodal commands. When voice control commands, web page control commands, and self-generated visual tracking commands are received simultaneously, the system executes the command with the highest priority according to preset priority rules, which are configured by the user through a web page.
5. The multi-scenario intelligent control system based on intelligent voice recognition technology according to claim 1, characterized in that, It also includes a multi-mode motion switching module, which specifically comprises: Based on the data collected by the environmental perception sensor, the current terrain complexity or obstacle information is determined in real time; When it is determined that the ground obstacle is insurmountable, a mode switching command is generated to control the air-land cooperative platform to switch from the ground tire driving mode to the short-distance, low-altitude flight mode in order to overcome the obstacle; After overcoming the obstacle, the platform is controlled to smoothly switch back from flight mode to ground driving mode.
6. The multi-scenario intelligent control system based on intelligent voice recognition technology according to claim 1, characterized in that, The environmental perception sensor includes a camera, an ultrasonic sensor, and an inertial measurement unit; the system also includes a sensor data fusion and processing module, which specifically comprises: Time synchronization and coordinate system unification of data from different sensors; A filtering algorithm is used to fuse multi-source data to generate comprehensive perception information on the surrounding environment, obstacle positions, and the user's own posture. Based on the comprehensive sensing information, a decision-making basis is provided for autonomous tracking and multi-mode motion switching.
7. The multi-scenario intelligent control system based on intelligent voice recognition technology according to claim 1, characterized in that, The collaborative task management module supports two collaborative strategies: split strategy and collaboration strategy. The specific separation strategy is as follows: the UAV platform and the vehicle platform are physically separated and independently execute different tasks, and the UAV platform and the vehicle platform share environmental information or task status through wireless communication. The specific collaborative strategy is as follows: the UAV platform conducts wide-area reconnaissance and navigation in the air, the vehicle platform performs precision operations or material transportation on the ground, and the UAV platform and the vehicle platform collaborate to complete the same task based on shared information.
8. The multi-scenario intelligent control system based on intelligent voice recognition technology according to claim 1, characterized in that, It also includes a remote interaction and data display platform, which is a web application built using the Vue and Spring Boot frameworks. This platform allows users to remotely monitor the real-time status, sensor data, and task execution screens of the air-land collaborative platform, and to issue remote commands.
9. A task execution method for a multi-scenario intelligent control system based on intelligent speech recognition technology as described in any one of claims 1-8, characterized in that, Includes the following steps: S1. The system receives task triggering instructions through the multimodal interaction module; S2. The collaborative task management module plans the task execution strategy based on the instruction type and current environment information, and determines the collaborative mode between the UAV platform and the vehicle platform. S3. The environmental perception sensor continuously collects environmental data, which is then analyzed by the sensor data fusion and processing module and fed back to the collaborative task management module for dynamic decision-making and adjustment. S4, the collaborative task management module sends specific motion commands to the corresponding processors in the hierarchical hardware control architecture, driving the air-land collaborative platform to execute tasks; S5. Status data during task execution is uploaded in real time to the remote interaction and data display platform for user monitoring and intervention.