Intelligent voice control sweeping music robot integrated with visual interaction and control method thereof
Through the intelligent voice-controlled sweeping music robot integrating visual interaction and AI interaction technology, the single function problem of sweeping robot and smart audio is solved, the whole-house music playback and personalized cleaning services are realized, and the user interaction experience and equipment adaptability are improved.
Patent Information
- Application Number
- CN202510803238.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-08-01
AI Technical Summary
The existing sweeping robots and smart audio equipment have single functions, lacking deep fusion and intelligent interaction capabilities, and cannot meet users' music needs and personalized interactive experiences in different areas of the whole house.
Design an intelligent voice-controlled sweeping music robot that integrates visual interaction and AI interaction, including high-definition cameras, microphone arrays, multi-core microprocessors, etc., to achieve multi-module collaborative work through visual and voice recognition technology, providing whole-house music playback and personalized cleaning services.
It has achieved the deep integration of sweeping robots and smart audio, providing whole-house mobile music playback and personalized and intelligent interactive experience, improving users' quality of life and the adaptability and stability of equipment.
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart home devices, and particularly relates to an intelligent voice-controlled sweeping music robot integrated with visual interaction and its control method, which is particularly suitable for devices and methods that combine the floor sweeping and cleaning function with the smart speaker function, and realize diverse home scene applications through visual interaction and voice control with AI interaction capabilities. Background Art
[0002] With the continuous development of smart home technology, sweeping robots and smart speakers have become common smart devices in households. Traditional sweeping robots mainly focus on floor cleaning tasks and achieve automatic cleaning through built-in sensors and navigation algorithms, but their functions are relatively single and lack diverse interaction methods with users. While smart speakers can realize functions such as music playback and information query through voice recognition, they are usually fixedly placed in a certain position, and the range of their music playback is limited, unable to meet the user's need to enjoy music at any time in different areas of the whole house.
[0003] In recent years, although there have been some technical solutions that attempt to integrate different functions into a single device, there are still many deficiencies in the integration of sweeping robots and smart speakers. For example, when integrating functions in existing devices, they often simply stack functions without achieving deep integration and collaborative work of the two functions. At the same time, in terms of interaction methods, most rely on single voice control, and the voice control lacks intelligent AI interaction capabilities, can only execute simple instruction operations, is difficult to understand complex semantic and emotional needs of users, and cannot provide a more personalized and natural interaction experience. Therefore, there is an urgent need for a smart home device and control method that can deeply combine a sweeping robot and a smart speaker and introduce innovative interaction technologies, especially voice control with powerful AI interaction capabilities. Summary of the Invention
[0004] Object of the Invention The object of the present invention is to provide an intelligent voice-controlled sweeping music robot integrated with visual interaction and its control method. By deeply integrating the floor sweeping and cleaning function with the smart speaker function, and introducing visual interaction technology and intelligent voice control technology with AI interaction capabilities, the functional boundaries of home devices are expanded, the intelligent interaction experience of users in cleaning and entertainment scenarios is improved, and the whole-house mobility of music playback and diverse, personalized, and intelligent human-computer interaction are realized.
[0005] Technical Solution
[0006] The intelligent voice-controlled sweeping music robot integrated with visual interaction described in the present invention includes a robot body, a visual interaction module, a voice recognition and AI interaction module, a music playback module, a floor sweeping and cleaning module, a control module, and a moving module.
[0007] The robot body is the supporting structure of the entire device, used to mount and secure the other modules. Its outer shell is constructed of a high-strength, lightweight composite material, ensuring structural strength while reducing the robot's overall weight and facilitating mobility. The outer shell is specially treated to be wear-resistant, scratch-resistant, and easy to clean, making it suitable for complex home environments.
[0008] The visual interaction module includes a high-definition camera and an image processing unit. The high-definition camera, mounted on the top or front of the robot, is equipped with a wide-angle lens and infrared night vision. It can capture human movements and the surrounding environment in real time under various lighting conditions, and transmits the captured image information to the image processing unit in a high-frame-rate, high-resolution format. The image processing unit, based on a deep learning algorithm and equipped with multiple human motion recognition models and environmental scene analysis models, performs in-depth analysis and processing of image information. It can not only recognize user body movements such as waving, nodding, and hand gestures, but also identify environmental conditions, including changes in furniture layout, the location of obstacles, and the degree of floor stains, and transmit the processed results to the control module. For example, when a user performs a specific combination of gestures, the image processing unit can accurately recognize the gesture combination and convert it into a corresponding complex control command signal. If significant stains are detected on the floor, the image processing unit transmits the location and severity of the stains to the control module, enabling the sweeping and cleaning module to perform targeted cleaning.
[0009] The voice recognition and AI interaction module includes a microphone array, a voice recognition processor, a natural language processing unit, and a dialogue management system. The microphone array adopts a multi-channel design, with beamforming technology and noise suppression function, capable of receiving voice signals from different directions omnidirectionally, effectively filtering out ambient noise, and accurately collecting the user's voice commands. The voice recognition processor uses advanced voice recognition algorithms to convert the received voice signals into text information. The natural language processing unit, based on a deep learning natural language processing model, performs semantic understanding, sentiment analysis, and intent recognition on the text information, can understand the user's complex semantic expressions, and identify the user's emotional tendencies and real needs. The dialogue management system generates reasonable response strategies and control instructions according to the analysis results of the natural language processing unit, combined with the current dialogue context and device status. For example, when the user says "I'm a bit tired and want to listen to some relaxing music", the voice recognition processor converts the voice into text, the natural language processing unit analyzes that the user feels fatigued and needs relaxing music, and the dialogue management system, on the one hand, controls the music playback module to play soothing music, and on the other hand, responds to the user through speech synthesis technology, such as "Okay, playing some soothing music for you, hope it can help you relax"; when the user says "This music doesn't sound good", the dialogue management system will understand that the user is not satisfied with the current music, and then interact with the user to ask about the user's specific preferences, such as "What style of music do you like? Is it classical, pop, or other styles?", and then adjust the music playback according to the user's answer. This module also supports multi-round conversations, can continuously understand the changing needs and intents of the user during the conversation process, and achieve natural and smooth AI interactive communication.
[0010] The music playback module includes an audio decoding chip, a digital signal processor, a power amplifier, and speakers. The audio decoding chip supports the decoding of multiple audio formats, including common ones such as MP3, WAV, FLAC, etc., and can perform efficient decoding processing on music files stored locally or music data received through the network. The digital signal processor further optimizes the decoded audio signals, such as equalization adjustment, sound effect enhancement, noise reduction, etc., and adjusts the parameters of the audio signals according to different music styles and user settings to provide better sound quality. The power amplifier amplifies the optimized audio signals to drive the speakers to play music. The speakers adopt high-quality audio units and are installed at appropriate positions on the robot body through reasonable layout design to ensure that the music can be clearly and evenly played throughout the house. At the same time, the speakers also have waterproof and dustproof functions to adapt to the complex environment during the robot's cleaning process.
[0011] The floor cleaning module includes a vacuum motor, side brushes, a roller brush, a dust collection box, and a cleaning mode adjustment device. The vacuum motor uses a high-performance brushless motor, which features strong suction, low noise, and long lifespan, and can generate stable and powerful suction to suck dust, garbage, etc. on the ground into the dust collection box through the air duct. The side brushes and the roller brush adopt special bristle materials and structural designs. The side brushes can rotate flexibly to sweep the garbage at the corners and edges into the suction range of the vacuum motor; the roller brush effectively rolls up the garbage on the ground and sends it into the air duct by rotating at high speed. The two work together to improve the cleaning efficiency. The dust collection box adopts a large-capacity and easily detachable design, which is convenient for users to clean regularly. The cleaning mode adjustment device can automatically adjust the suction of the vacuum motor, the rotation speeds of the side brushes and the roller brush according to the instructions of the control module, and implement different cleaning modes, such as standard mode, strong mode, silent mode, etc., to adapt to different cleaning scenarios and requirements. For example, in the standard mode, the vacuum motor works with medium suction, and the side brushes and the roller brush rotate at normal speeds, which is suitable for daily floor cleaning; in the strong mode, the vacuum motor works with maximum suction, and the rotation speeds of the side brushes and the roller brush increase, which can effectively clean stubborn stains and a large amount of garbage; in the silent mode, the vacuum motor works with low suction, and the rotation speeds of the side brushes and the roller brush decrease to reduce noise generation, which is suitable for use when the user is resting.
[0012] The control module is the core control unit of the entire robot, which uses a high-performance multi-core microprocessor and is equipped with an embedded operating system and a real-time control system. The control module is used to receive the information transmitted by the visual interaction module and the voice recognition and AI interaction module, and issue control instructions to the music playback module, the floor cleaning module, and the movement module according to the preset programs and algorithms to achieve the coordinated work of each module. At the same time, the control module also has the functions of data storage and management, can store data such as the user's usage habits, personalized settings, and cleaning history records, and analyzes these data through data analysis algorithms to provide more personalized services for users. For example, according to the music styles and time periods that the user often plays, automatically recommend music that meets the user's preferences at the corresponding time; according to the user's cleaning habits and the characteristics of the home environment, optimize the cleaning path planning and cleaning mode selection. In addition, the control module supports connecting to the mobile APP or the smart home central control system through the wireless network. Users can remotely control the robot through the mobile APP, view the device status, set tasks, adjust parameters, etc.; they can also be linked with other smart home devices to achieve richer smart home scenario applications. For example, when it detects that the user comes home, it automatically starts the floor cleaning music robot, plays welcome music and starts cleaning.
[0013] The mobile module includes a driving motor, wheels, a suspension system, and navigation sensors. The driving motor uses a high-precision servo motor, which can accurately control the rotation speed and steering of the wheels, enabling the robot to move flexibly on the ground. The wheels are made of anti-slip and wear-resistant rubber material and are equipped with an independent suspension system, which can effectively adapt to different ground materials and terrains, such as carpets, floors, tiles, etc., ensuring the stability and passability of the robot during movement. The navigation sensors include multiple sensors such as lidar, infrared sensors, ultrasonic sensors, and visual positioning cameras. The lidar is used to scan the surrounding environment in real time to construct a high-precision three-dimensional map; the infrared sensors and ultrasonic sensors are used to detect obstacles at close range to avoid collisions; the visual positioning camera combines image processing algorithms to achieve visual positioning and navigation of the robot. Through the collaborative work of multiple sensors, the mobile module can real-time sense the environmental information around the robot, and use advanced path planning algorithms, such as improved AI algorithms, Dijkstra algorithms, etc., to plan the optimal movement path, achieve automatic navigation and full-house coverage cleaning, and can also accurately move in different areas according to the user's needs when playing music.
[0014] Based on the above intelligent voice-controlled sweeping music robot, the present invention also provides a control method, and the specific steps are as follows: Initialization: After the robot is powered on, the control module first performs initialization settings on the visual interaction module, voice recognition and AI interaction module, music playback module, sweeping and cleaning module, and mobile module. It sequentially detects whether the hardware connections of each module are normal and whether the software programs are successfully loaded. If an abnormality is detected, the control module uses voice synthesis technology to send a detailed fault prompt to the user, such as "The visual interaction module connection is abnormal. Please check whether the camera is loose", and displays the fault information and solution suggestions in a graphic form on the display screen (if any) equipped on the robot or the mobile APP, facilitating the user to conduct troubleshooting and repair.
[0015] Information collection: The high-definition camera of the visual interaction module captures human actions and environmental images in real time at a set frame rate and resolution, and transmits the image information to the image processing unit. The image processing unit preprocesses the image information, including operations such as denoising, enhancement, and cropping, and then uses a deep learning model for human action recognition and environmental scene analysis, and transmits the recognized and analyzed results to the control module in a specific data format.
[0016] The microphone array of the speech recognition and AI interaction module continuously collects the user's speech signals, optimizes the speech signals through beamforming technology and noise suppression functions, and then transmits the optimized speech signals to the speech recognition processor. After converting the speech signals into text information, the speech recognition processor transmits it to the natural language processing unit. The natural language processing unit performs semantic understanding, sentiment analysis, and intent recognition on the text information, and transmits the analysis results to the dialogue management system.
[0017] Instruction judgment and processing: The control module judges and analyzes the received visual interaction information and speech instruction information.
[0018] If it is a speech instruction, the dialogue management system first determines whether the instruction is a clear control instruction according to the analysis results of the natural language processing unit, such as music playback control instructions (including play, pause, previous song, next song, adjust volume, switch music style, etc.), floor cleaning control instructions (including start cleaning, stop cleaning, schedule cleaning, select cleaning mode, etc.), movement control instructions (including go to a specified location, return to the charging dock, etc.). If it is a clear control instruction, according to the preset instruction mapping relationship, it sends a control instruction to the corresponding module. If the instruction is not a clear control instruction, but content such as the user's question, suggestion, or expression of emotion, the dialogue management system combines the current dialogue context and device status, uses dialogue strategies to generate response content, interacts with the user through speech synthesis technology, and at the same time sends control instructions to the corresponding module at an appropriate time according to the user's intent. For example, when the user says "What's the weather like today", after the dialogue management system obtains the weather information by querying the weather interface, it responds to the user "It's sunny and the temperature is suitable today". If the user then says "Then play some relaxing outdoor music", the dialogue management system sends an instruction to play relaxing outdoor music to the music playback module.
[0019] If it is visual interaction information, the control module converts the recognized body movements into corresponding control instructions according to the preset correspondence between body movements and instructions, and sends them to the corresponding module. At the same time, for some complex combinations of body movements or continuous movements, the control module generates a more complex control instruction sequence by analyzing the timing and logical relationships of the movements to achieve richer interaction functions.
[0020] Module execution: After the music playback module receives the music playback control instruction from the control module, the audio decoding chip decodes the corresponding music data. The digital signal processor optimizes the decoded audio signal according to the instruction. The power amplifier amplifies the audio signal, and the speaker plays the corresponding music. At the same time, parameters such as the volume, sound effect, and playback mode of the music are adjusted according to the instruction. For example, when receiving an instruction to switch the music style, the digital signal processor adjusts the equalization parameters of the audio signal to make the music present the sound effect characteristics of different styles.
[0021] After the floor cleaning module receives the floor cleaning control instruction, the cleaning mode adjustment device adjusts the suction power of the dust suction motor, the rotation speeds of the side brush and the roller brush according to the instruction, and starts the corresponding cleaning mode. The dust suction motor, the side brush, and the roller brush work together to clean the floor garbage and suck it into the dust collection box. During the cleaning process, the capacity of the dust collection box and the cleaning progress are monitored in real time, and the information is fed back to the control module.
[0022] After the movement module receives the movement control instruction, the drive motor accurately controls the rotation speed and steering of the wheels according to the instruction. The navigation sensor senses the environmental information in real time, and uses the path planning algorithm to plan and adjust the movement path, so that the robot moves to the specified position or executes the corresponding movement task according to the requirements. During the movement process, the position and attitude information of the robot itself are continuously monitored and fed back to the control module.
[0023] Feedback and adjustment: During the operation of the robot, each module feeds back the working state information to the control module in real time. For example, the music playback module feeds back information such as the currently playing music track, volume, playback progress, and audio quality; the floor cleaning module feeds back information such as the cleaning progress, remaining capacity of the dust collection box, working state of the dust suction motor, and rotation speeds of the side brush and the roller brush; the movement module feeds back information such as the current position, movement speed, attitude information, and navigation path planning situation. The control module adjusts and optimizes the working states of each module according to the feedback information, using the preset adjustment strategies and algorithms. When the dust collection box is close to being full, the control module controls the robot to return to the charging dock according to the planned path and prompts the user to clean the dust collection box through voice; when it is detected that the environmental noise is large during the music playback, the control module automatically adjusts the volume and sound effect parameters to ensure a good auditory experience; when the movement module encounters an obstacle that cannot be avoided, the control module adjusts the movement strategy, such as trying to detour from other directions or pausing to wait for the user to handle. At the same time, the control module stores the feedback information and adjustment records for subsequent data analysis and optimization. Description of the Drawings
[0025] Figure 1 : Flowchart of the system architecture of the intelligent voice-controlled floor cleaning music robot This flow chart shows the system architecture of the intelligent voice-controlled sweeping music robot, clarifying the connection relationships and information transmission directions among various modules.
[0026] The visual interaction (s101) module transmits the processed visual interaction information to the control module (s102).
[0027] The voice recognition (s103) and AI interaction module (s104) transmit the voice commands and the results after AI processing to the control module.
[0028] The control module (s102) sends control commands related to music playback to the music playback module (s105).
[0029] The control module sends control commands related to floor sweeping and cleaning to the floor sweeping and cleaning module (s106).
[0030] The control module sends control commands related to movement to the movement module (s107).
[0031] Figure 2 : Flow chart of the control method for the intelligent voice-controlled sweeping music robot This flow chart details the control method process of the intelligent voice-controlled sweeping music robot.
[0032] s201: After the robot is powered on, it enters the initialization stage, and the control module performs initialization settings and hardware and software detections on each module.
[0033] s202: Enter the information collection stage, and the visual interaction module and the voice recognition and AI interaction module respectively collect information, process it, and transmit it.
[0034] s203: The control module judges and processes the received information, distinguishes voice commands and visual interaction information, and performs corresponding processing.
[0035] s204: Each functional module performs corresponding operations according to the instructions of the control module, such as music playback, floor sweeping and cleaning, movement, etc.
[0036] s205: During the operation of the robot, each module feeds back the working status information to the control module, and the control module makes adjustments and optimizations.
[0037] Beneficial effects
[0038] Deep integration of functions: The present invention deeply integrates the cleaning function of the sweeping robot and the music playback function of the intelligent speaker. It not only realizes the integration of functions, but also through the collaborative work between the control module and each module, enables the two functions to be organically combined and cooperate with each other, expanding the functional boundaries of household appliances, providing users with more comprehensive and convenient household services, reducing the number of devices in the home, and saving space.
[0039] Interactive experience upgrade: Introduce visual interaction technology and voice control technology with AI interaction capabilities. Users can interact with the robot in multiple ways. Visual interaction provides an intuitive and natural interaction means to meet the control needs of users when voice communication is inconvenient; the voice control with AI interaction capabilities can understand the complex semantic and emotional needs of users, achieve natural and smooth dialogue communication, provide personalized services, greatly improve the convenience, interestingness and intelligent level of human-computer interaction, and meet the usage habits and needs of different users.
[0040] Music playback optimization: Through the precise navigation and flexible movement of the mobile module, the music playback breaks through the fixed area limit and realizes whole-house mobile playback. At the same time, the audio processing technology and speaker design of the music playback module ensure high-quality music playback and good auditory effects. Users can enjoy personalized music services at any corner of the home, and can also accompany music that suits their mood and scene when doing housework such as cleaning, improving the quality of life.
[0041] High degree of intelligence: The control module adopts an advanced hardware and software architecture, combined with data analysis and learning algorithms, and can automatically optimize the operating parameters and working modes of the device according to the user's usage habits and the device's working status, realizing intelligent operation. At the same time, it supports linkage with mobile phone APPs and other smart home devices, further expanding the application scenarios of the device and creating a more intelligent and convenient home living environment for users.
[0042] Strong adaptability: Each module of the robot is designed with full consideration of the complex home use environment, such as the night vision function of the visual interaction module, the waterproof and dustproof design of the music playback module, the multiple cleaning modes of the floor cleaning module, the suspension system and multi-sensor navigation of the mobile module, etc., enabling the robot to adapt to different lighting conditions, floor materials and home scenes, ensuring the stability and reliability of the device.
Claims
1. An intelligent voice-controlled sweeping music robot integrated with visual interaction, characterized in that, Including: A robot body, serving as the load-bearing structure for each functional module; A visual interaction module, including a high-definition camera and an image processing unit. The high-definition camera is installed on the top or front of the robot body and is used to capture human actions and environmental images in real time. The image processing unit analyzes and processes the image information based on deep learning algorithms to identify user limb actions and environmental state features; A voice recognition and AI interaction module, including a microphone array, a voice recognition processor, a natural language processing unit, and a dialogue management system. The microphone array adopts a multi-channel design and has beamforming technology. The natural language processing unit realizes semantic understanding, sentiment analysis, and intent recognition based on a deep learning model. The dialogue management system supports multi-round dialogue interactions; A music playback module, including an audio decoding chip, a digital signal processor, a power amplifier, and a speaker. The digital signal processor performs optimization processing such as equalization adjustment and sound effect enhancement on the audio signal; A floor cleaning module, including a suction motor, side brushes, a roller brush, a dust collection box, and a cleaning mode adjustment device. The cleaning mode adjustment device automatically adjusts the suction power of the suction motor and the rotation speed of the side brushes and roller brush according to control instructions; A movement module, including a drive motor, wheels, a suspension system, and navigation sensors. The navigation sensors include a lidar, an infrared sensor, an ultrasonic sensor, and a visual positioning camera; A control module, which connects and controls the above-mentioned modules to work together, receives information from the visual interaction module and the voice recognition and AI interaction module, and sends control instructions to the music playback module, the floor cleaning module, and the movement module.
2. The intelligent voice-controlled sweeping music robot with integrated visual interaction according to claim 1, characterized in that: The natural language processing unit of the voice recognition and AI interaction module includes: A semantic understanding sub-unit, which uses a Transformer architecture model to parse the semantic meaning of the user's voice text; A sentiment analysis sub-unit, which identifies the sentiment tendency in the user's voice based on the BERT model; An intent recognition sub-unit, which judges the category of the user's instruction intent through a multi-classifier.
3. The intelligent voice-controlled sweeping music robot with integrated visual interaction according to claim 1, characterized in that: The image processing unit of the visual interaction module includes: A human body action recognition unit, which uses the OpenPose algorithm to identify user limb action features; An environmental analysis unit, which identifies the degree of ground stains and the distribution of obstacles based on the MaskR-CNN model; A feature extraction unit, which extracts feature vectors and classifies the recognition results.
4. The intelligent voice-controlled sweeping music robot with integrated visual interaction according to claim 1, characterized in that: The control module includes: A data storage unit, which stores user usage habits, personalized settings, and device operation data; A data analysis unit, which analyzes and mines the stored data based on machine learning algorithms; A decision-making and control unit, which generates an optimal control strategy according to the data analysis results and the current instructions.
5. The intelligent voice-controlled sweeping music robot with integrated visual interaction according to claim 1, characterized in that: The digital signal processor of the music playback module includes: An audio equalizer, which automatically adjusts the frequency response curve according to the music style; A 3D sound effect processor, which simulates spatial audio effects through the HRTF algorithm; A dynamic range compressor, which adaptively adjusts the audio dynamic range.
6. A control method for a robot according to any one of claims 1-5, characterized in that, Including the following steps: S201 Initialization: The control module performs initialization settings and status detection on each functional module; S202 Information collection: The visual interaction module collects human actions and environmental images, and the speech recognition and AI interaction module collects speech commands and performs semantic analysis; S203 Command judgment and processing: The control module differentiates the command types and converts the speech commands and visual interaction information into control signals; S204 Module execution: Each functional module performs corresponding operations according to the control signals; S205 Feedback and adjustment: Each module feeds back its working status to the control module, and the control module adjusts the control strategy according to the feedback information.
7. The control method according to claim 6, wherein: In the S203 command judgment and processing step, it includes: Classify the intent of the speech command, and identify it as a music control type, a cleaning control type, or a movement control type; Perform action semantic mapping on the visual interaction information, and convert specific limb actions into corresponding control commands; Combine historical interaction data and current scene information for command optimization processing.
8. The control method according to claim 6, wherein: In the S205 feedback and adjustment step, it includes: The music playback module feeds back audio quality parameters, and the control module dynamically adjusts the sound effect parameters; The floor cleaning module feeds back the dust collection box capacity and the cleaning coverage rate, and the control module plans the cleaning path; The movement module feeds back the real-time position and obstacle information, and the control module optimizes the navigation strategy.
9. The control method according to claim 6, characterized in that: It also includes a learning and optimization step. The control module optimizes the control strategy based on the user's historical interaction data using a reinforcement learning algorithm, including: Establish a user preference model to predict the possible command intents of the user; Optimize the multi-task scheduling algorithm to balance the priorities of cleaning tasks and music services; Dynamically adjust the sensor parameters to improve the environmental perception accuracy.
10. The control method according to claim 6, wherein: In the S202 information collection step, the speech recognition and AI interaction module uses an attention mechanism to process multi-round conversations, including: Construct a dialogue context vector to retain historical dialogue information; Calculate the attention weights of the current input and the context; Generate a coherent response strategy according to the attention weights.
Citation Information
Cited By
Intelligent interactive door control system
CN120997935A