Intelligent robot language recognition interaction method
Through the combination of microphone array and speech recognition processor, the voice communication barrier between the elderly and intelligent robots is solved, efficient voice interaction and daily care assistance are achieved, and the quality of life of the elderly is improved.
Patent Information
- Application Number
- CN202510816560.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies make it difficult to achieve efficient and barrier-free voice communication between the elderly and intelligent robots, and are unable to meet the multi-dimensional emotional needs and daily care needs of the elderly group.
A microphone array is used to capture sound signals, which are pre-processed, recognized, and analyzed by a speech recognition processor to generate voice signals for interaction. It supports multi-language recognition and plays responses through speakers.
It enables barrier-free voice communication between the robot and the elderly, enhances social connections, assists with daily care, such as reminding them to drink water and take medicine, and reduces the loneliness of the elderly.
Smart Images

Figure CN120612943A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent robots, and in particular relates to a language recognition and interaction method for intelligent robots. Background Art
[0002] Against the backdrop of an accelerating global aging population, effectively improving the well-being of the elderly and addressing their multifaceted emotional needs has become a pressing social issue. Leveraging the rapid development of intelligent and information technology, intelligent companion robots are integrating cutting-edge technologies such as artificial intelligence, the Internet of Things, and robotics to create a comprehensive, 24 / 7, multimodal care system. As a fusion of artificial intelligence and robotics, intelligent companion robots have become a crucial tool for addressing elderly care needs. Their research and application encompass multiple dimensions, including emotional companionship, health monitoring, and daily living assistance, and are experiencing a complex development in terms of technological breakthroughs, market expansion, and ethical considerations. Intelligent robotic companion systems can not only significantly reduce labor costs but also improve care effectiveness through precise and personalized service models. Currently, in response to the dual challenges of an aging population and digital transformation, intelligent companion robots have become a key area of global academic research and industrial application, with significant cross-national differences in their technological evolution. Summary of the Invention
[0003] In view of this, the present invention proposes an intelligent robot language recognition interaction method, comprising the following steps: S10 captures sound signals from the surrounding environment through a microphone array; S20, input processor for speech recognition; S30, recognizing the content of the voice command, the processor parses it and controls the robot to perform corresponding actions or responses; S40: Generate a voice signal and play it through a speaker.
[0004] Preferably, before the S20 speech recognition, the input sound signal is preprocessed, and the preprocessing includes filtering and noise reduction.
[0005] Preferably, the speech recognition in S20 is to convert the sound signal into text or instructions, and the speech recognition includes feature extraction and pattern matching.
[0006] Preferably, the feature extraction is to extract feature information related to speech from the sound signal.
[0007] Preferably, the pattern matching compares the extracted feature information with a speech template or model stored in a memory.
[0008] Preferably, the recognition of voice instructions in S30 includes text analysis and voice coding.
[0009] Preferably, the text analysis converts the text into speech units, which include syllables and phonemes.
[0010] Preferably, the speech encoding is to combine speech units into a coherent speech signal.
[0011] Compared with the prior art, the intelligent robot language recognition interaction method disclosed in the present invention has at least the following beneficial effects: The robot enables seamless voice communication with users, enabling intelligent voice recognition and human-machine interaction through API calls. The robot can engage in conversations with seniors at any time, integrating voice communication into the intelligent care process to help alleviate loneliness and enhance social connections through daily interactions. The robot can also remind seniors to drink water, take medicine, and other daily tasks, helping them maintain healthy habits and schedules. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to make the purpose, technical solutions and beneficial effects of the present invention more clear, the present invention provides the following drawings for illustration: Figure 1 This is a flowchart of the steps of the intelligent robot language recognition interaction method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a voice interaction command box for the intelligent robot language recognition interaction method according to an embodiment of the present invention; Figure 3 A schematic diagram of setting the recognized language type for the intelligent robot language recognition interaction method according to an embodiment of the present invention; Figure 4 A schematic diagram of a language recognition instruction box for setting up a language recognition interaction method for an intelligent robot according to an embodiment of the present invention; Figure 5 This is a schematic diagram of setting the recognition language instruction box parameters for the intelligent robot language recognition interaction method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0013] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0014] See also Figure 1 The flow chart of the intelligent robot language recognition interaction method of the present invention includes the following steps: S10 captures sound signals from the surrounding environment through a microphone array; S20, input processor for speech recognition; S30, recognizing the content of the voice command, the processor parses it and controls the robot to perform corresponding actions or responses; S40: Generate a voice signal and play it through a speaker.
[0015] The microphone array includes four directional microphones with a sensitivity of 20mV / Pa±3dB at 1kHz and a frequency range of 150Hz to 12kHz. These four microphones are MicroFront, MicroLeft, MicroRear, and MicroRight. They are located on the robot's left and right ears, forehead, and back of the head. This allows the robot to detect external sounds while also identifying the direction of the sound source. Each microphone corresponds to a channel and has independent sampling and encoding circuitry. The robot's microphones can perform functions such as audio playback, reading user input, and simple recording. The external volume level received is determined by the output variable energy (variable). A higher energy value indicates a higher volume level received by the robot.
[0016] The robot uses a microphone array to capture sound signals from its surroundings and converts them into electrical signals. Before speech recognition, the robot first preprocesses the input speech signal. This step aims to remove noise, enhance the quality of the speech signal, and make it more suitable for subsequent processing. Through filtering and noise reduction techniques, the robot ensures the clarity and accuracy of the speech signal, laying a solid foundation for subsequent speech recognition.
[0017] The preprocessed sound signal is fed into the processor for speech recognition. The processor runs a speech recognition algorithm, converting the sound signal into text or commands. Speech recognition algorithms typically include steps such as feature extraction and pattern matching. Feature extraction extracts feature information from the sound signal that is useful for speech recognition. Pattern matching compares the extracted feature information with speech templates or models stored in memory to determine the content of the user's voice command.
[0018] Once the user's voice command content is recognized, the processor will parse the command and control the robot to perform the corresponding action or response. For situations where a voice reply is required, the processor will run a speech synthesis algorithm to convert the text or command into a voice signal. The speech synthesis algorithm usually includes steps such as text analysis and speech coding. Text analysis is responsible for converting text into speech units (such as syllables, phonemes, etc.): speech coding combines these speech units into a coherent speech signal. Finally, the generated voice signal is played through the speaker to achieve voice interaction with the user.
[0019] See also Figure 2During voice interaction, I used the command box, which includes "Speak," "Set Recognition Language," "Speech Recognition," and "Switch Case." The "Speak" command box allows the robot to read out text or sentences preset in the "Text" field. Clicking the parameter setting icon in the lower left corner of the "Speak" command box allows you to adjust the voice shaping and speed parameters.
[0020] See also Figure 3 The Set Recognition Language command box is used to set the language type for speech recognition. The robot supports 32 languages, which can be selected and set in the options below the command box. In the specific embodiment, the type of speech recognition is set to Mandarin.
[0021] See also Figure 4 The speech recognition command box is used to determine whether the speech generated by the microphone matches the set text. The command box includes two input terminals and three output terminals. The first output terminal is activated when speech recognition starts. The second output terminal (wordRecognized output terminal) will be triggered when the set word is recognized and output the recognized text. The third output terminal (onNothing output terminal) is triggered when the set text is not recognized.
[0022] See also Figure 5 , click the wrench in the lower left corner of the command box to open the settings interface. The three settings are: Word list: Fill in the words you want the robot to recognize, and separate each word with an English semicolon; Confidence threshold: Set the threshold for robot recognition; Enable word spotting: Set whether to fuzzy recognize the set words.
[0023] The Switch Case command box is equivalent to a conditional branch statement. You can preset keywords in the Switch Case command box so that the robot can trigger corresponding voice interaction or joint movements after recognizing the keywords. You can add more keyword options to the Switch Case command box to enrich the robot's voice recognition and interaction functions. In this case, the output port behind the Switch Case command box will also increase accordingly.
[0024] In the specific embodiment, the robot uses the NAO robot as an intelligent carrier. As a typical representative of bipedal humanoid service robots, the NAO robot has demonstrated unique technical advantages in the field of intelligent elderly care since it was released by the French Aldebaran Robotics company in 2006.
[0025] In addition to the above embodiments, the present invention may also have other implementations. Any technical solution formed by equivalent replacement or equivalent transformation is within the scope of protection required by the present invention.
[0026] The present invention has been described in detail above, but the specific implementation of the present invention is not limited thereto. Those skilled in the art may make various modifications or alterations without departing from the spirit and scope of the claims of this application.
Claims
1. An intelligent robot language recognition interaction method, characterized in that: The following steps are involved: S10 captures sound signals from the surrounding environment through a microphone array; S20, input processor for speech recognition; S30, recognizing the content of the voice command, the processor parses it and controls the robot to perform corresponding actions or responses; S40: Generate a voice signal and play it through a speaker.
2. The intelligent robot language recognition interaction method according to claim 1, characterized in that: Before the speech recognition in S20, the input sound signal is preprocessed, and the preprocessing includes filtering and noise reduction.
3. The intelligent robot language recognition interaction method according to claim 1, characterized in that: The speech recognition in S20 is to convert the sound signal into text or instructions, and the speech recognition includes feature extraction and pattern matching.
4. The intelligent robot language recognition interaction method according to claim 3, characterized in that: The feature extraction is to extract feature information related to speech from the sound signal.
5. The intelligent robot language recognition interaction method according to claim 3, characterized in that: The pattern matching compares the extracted feature information with the speech template or model stored in the memory.
6. The intelligent robot language recognition interaction method according to claim 1, characterized in that: The recognition of the voice instruction in S30 includes text analysis and voice coding.
7. The intelligent robot language recognition interaction method according to claim 6, characterized in that: The text analysis converts text into phonetic units, which include syllables and phonemes.
8. The intelligent robot language recognition interaction method according to claim 7, characterized in that: The speech coding is to combine speech units into a coherent speech signal.