Robotic multi-chamber sound production device

By simulating the coordinated vocalization of multiple cavities in humans through a multi-cavity sound-producing device, the problem of monotonous vocalization in traditional robots has been solved, achieving more realistic and richer sound output, and enhancing the immersiveness and anthropomorphic effect of human-computer interaction.

CN224544620UActive Publication Date: 2026-07-24SHANGHAI TODAY XINDONG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
SHANGHAI TODAY XINDONG TECHNOLOGY CO LTD
Filing Date
2025-09-05
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing robot voice-generating devices rely on traditional electronic synthesis technology, resulting in monotonous and mechanical sounds that lack the natural listening experience and emotional tension of human voices, thus hindering the anthropomorphic upgrade of human-computer interaction.

Method used

The device employs a multi-cavity sound-generating mechanism, including a control module and multiple sound-generating modules, which are respectively placed in the robot's head, neck, or torso. The initial sound waves are processed through acoustic cavities to simulate the coordinated sound generation of multiple cavities in humans, achieving spatial superposition and coupling of sound waves to enhance realism and immersion.

Benefits of technology

It improves the realism and layering of the robot's voice, enhances the immersiveness and simulation of human-computer interaction, and makes the output sound closer to the natural human voice, adapting to the interaction needs of different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN224544620U_ABST
    Figure CN224544620U_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to the technical field of robot control, and particularly relates to a robot multi-cavity sound production device.The device comprises a control module and at least one sound production module; the control module comprises a processor and a signal driver; the sound production module comprises a vibration module and an acoustic cavity in communication with the vibration module; the sound production module is in communication connection with the signal driver of the control module through a cable or a wireless communication module, the vibration module is arranged to be driven by the control module to generate an initial sound wave, the acoustic cavity is arranged to acoustically process the initial sound wave, and when the sound production module is multiple, the multiple sound production modules are arranged at different positions of a head, a neck or a trunk of a robot body respectively.The device can simulate the spatial listening of human sound production and the sound characteristics of the common sound production of multiple positions, enhances the authenticity of robot sound production, makes the robot output have natural human voice characteristics, enhances the immersion and naturalness of human-computer interaction, and improves the interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application relate to the field of robot control technology, specifically to a multi-cavity sound-generating device for robots. Background Technology

[0002] The statements herein are provided merely as background information in connection with this application and do not necessarily constitute prior art.

[0003] With the development of artificial intelligence and automation technologies, the demand for service and companion robots has surged, and their core competitiveness lies in highly human-like interaction. "Voice," as a key carrier of emotion and information transmission in human-computer interaction, directly impacts user experience—a natural human voice can reduce psychological barriers and enhance immersion. However, most robots currently rely on traditional electronically synthesized speech technology, either splicing together preset audio segments or adjusting basic parameters such as pitch and speech rate. This results in monotonous, mechanical output voices that lack the natural listening experience and emotional tension of a human voice, making interactions feel stiff and hindering their upgrade from "tool attributes" to "emotional interaction attributes." Therefore, breaking through traditional technologies and enabling robots to emit voices with realistic human characteristics is crucial to improving their simulation accuracy. Utility Model Content

[0004] A brief overview of this application is provided below to offer a basic understanding of certain aspects thereof. It should be understood that this overview is not an exhaustive summary of the application. It is not intended to identify key or essential parts of the application, nor is it intended to limit its scope. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description that follows.

[0005] Embodiments of this application provide a multi-cavity sound-generating device for a robot. The device includes: a control module and at least one sound-generating module; the control module includes a processor and a signal driver; the sound-generating module includes a vibration module and an acoustic cavity communicating with the vibration module; the sound-generating module is communicatively connected to the signal driver of the control module via a cable or a wireless communication module; the vibration module is configured to generate an initial sound wave driven by the control module; the acoustic cavity performs acoustic processing on the initial sound wave; when there are multiple sound-generating modules, the multiple sound-generating modules are respectively located at different positions of the robot's head, neck, or torso.

[0006] The method provided in the embodiments of this application receives external information through the processor of the control module and generates drive signals to control the sound-generating module. The signal driver controls the vibration module of the sound-generating module to generate initial sound waves. This facilitates intelligent and context-specific sound generation by the robot based on input information. Furthermore, it allows the robot to emit sounds suitable for the current scene from appropriate spatial locations on its body, simulating the spatial auditory experience of different positions during human conversation. Simultaneously, the acoustic cavity performs acoustic processing on the sound waves, enhancing the layering of the robot's voice. Furthermore, when multiple drive signals are present, multiple sound-generating modules can be controlled to emit sound simultaneously, forming spatially superimposed coupled sound waves. This more realistically simulates the sound characteristics of multiple parts of the human voice emitting sound together, enhancing the realism of the robot's voice and improving the immersive experience during human-computer interaction.

[0007] These and other advantages of this application will become more apparent from the following detailed description of preferred embodiments in conjunction with the accompanying drawings. Attached Figure Description

[0008] To further illustrate the above and other advantages and features of this application, the specific embodiments of this application will be described in more detail below with reference to the accompanying drawings. The drawings, together with the following detailed description, are included in and form a part of this specification. Elements having the same function and structure are indicated by the same reference numerals. It should be understood that these drawings only depict typical examples of this application and should not be considered as limiting the scope of this application.

[0009] Figure 1 This is a schematic diagram of the structure of a robot multi-cavity sound-generating device according to an embodiment of this application.

[0010] It should be noted that the accompanying drawings are not necessarily drawn to scale, but are shown only in a schematic manner without affecting the reader's understanding.

[0011] Explanation of reference numerals in the attached figures:

[0012] 10. Control module; 20. Sound generation module; 30. Vibration module; 40. Acoustic cavity. Detailed Implementation

[0013] Exemplary embodiments of this application will be described below with reference to the accompanying drawings. For clarity and brevity, not all features of actual implementations are described in the specification. However, it should be understood that many implementation-specific decisions must be made in the development of any such actual embodiment to achieve the developer's specific goals, such as complying with constraints related to the system and business, and these constraints may vary depending on the implementation. Furthermore, it should be understood that while development work can be very complex and time-consuming, such development work is merely a routine task for those skilled in the art who benefit from the content of this application.

[0014] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the equipment structure and / or processing steps closely related to the solution according to this application are shown in the accompanying drawings, while other details that are not closely related to this application are omitted.

[0015] It should be noted that, unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning as understood by a person with ordinary skills in the field to which this application pertains.

[0016] In the description of the embodiments of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0017] In related technologies, traditional methods for robot voice production often use a single speaker as the output, combined with electronic synthesis technology to generate speech signals. Voice production is achieved through pre-set parameter splicing or tone adjustment. This differs fundamentally from the complex physiological vocal mechanism of humans, which relies on the coordination of multiple cavities such as the larynx, nose, and pharynx—it cannot simulate the spatial and temporal sound coupling effect of these multiple cavities during human vocalization. This directly results in a monotonous and flat timbre in the robot's output voice, a mechanical and stiff rhythm, and frequent issues such as formant shift and rhythm misalignment when repeating the same phrase. In contrast, the natural human voice possesses a vivid auditory experience precisely because its multi-cavity resonance creates rich layers. The limitations of these technologies make robots appear stiff and rigid in human-computer interaction, making it difficult to stably reproduce anthropomorphic voice characteristics and hindering the realization of highly realistic interactive experiences.

[0018] To address the aforementioned technical problems, embodiments of this application provide a multi-cavity sound-generating device for robots. Figure 1 This is a schematic diagram of the structure of a robot multi-cavity sound-generating device according to an embodiment of this application, as shown below. Figure 1As shown, the device includes: a control module 10 and at least one sound-generating module 20; the control module 10 includes a processor and a signal driver; the sound-generating module 20 includes a vibration module 30 and an acoustic cavity 40 communicating with the vibration module 30; the sound-generating module 20 is communicatively connected to the signal driver of the control module 10 via a cable or a wireless communication module, the vibration module 30 is configured to generate an initial sound wave driven by the control module 10, and the acoustic cavity 40 performs acoustic processing on the initial sound wave; when there are multiple sound-generating modules 20, the multiple sound-generating modules 20 are respectively located at different positions of the robot's head, neck, or torso.

[0019] The method provided in the embodiments of this application receives external information through the processor of the control module 10 and generates a drive signal to control the sound-emitting module 20. The signal driver controls the vibration module 30 of the sound-emitting module 20 to generate initial sound waves. This facilitates intelligent and scenario-based sound processing of the robot based on input information. Furthermore, it can emit sounds suitable for the current scene from appropriate spatial positions on the robot body, simulating the spatial auditory experience of different positions during human conversation. Simultaneously, the acoustic cavity 40 performs acoustic processing on the sound waves, enhancing the layering of the robot's voice. Furthermore, when there are multiple drive signals, multiple sound-emitting modules 20 at different positions on the head, neck, or torso can be controlled to emit sounds simultaneously, forming spatially superimposed coupled sound waves. This more realistically simulates the sound characteristics of multiple parts of the human body emitting sounds together, enhancing the realism of the robot's voice and improving the immersive experience during human-computer interaction.

[0020] In some embodiments, the processor may include: a data receiving module configured to receive input information; a feature extraction module configured to determine feature information of the input information based on the input information; a response analysis module configured to determine the robot's response content based on the determined feature information, the response content including the sound characteristics and specific sound information of the robot's response; a parameter translation module configured to determine the sound-emitting modules 20 that need to emit sound based on the response content, and determine the real-time sound parameters of each sound-emitting module 20; and a signal generation module configured to determine the corresponding control drive signal based on the determined sound-emitting modules 20 that need to emit sound and the real-time acoustic parameters of the sound-emitting modules 20, and transmit it to the signal driver. In these embodiments, through the analysis and processing of input information, precise control of the robot's sound emission can be achieved.

[0021] In some embodiments, the feature information of the input information includes emotional tags (such as emotional tendencies such as joy and sadness), phoneme sequences (such as basic articulation units of speech), etc., and the response content may include elements such as speech text content, pitch, and rhythm. In these embodiments, by analyzing the input information, the robot can analyze the emotions and content in the current interaction scenario, and then generate a human-like voice response that matches the scenario. This is beneficial to the intelligentization of the robot's voice, making the robot's voice more in line with the interaction needs in terms of semantics and expression.

[0022] In some embodiments, the driving signal may include control instructions for parameters such as frequency, amplitude, and phase of the sound-generating module 20, thereby driving one or more sound-generating modules 20 to emit sound according to predetermined parameters. In these embodiments, by controlling the acoustic parameters such as frequency, amplitude, and phase of the sound-generating module 20, the sound waves emitted by the sound-generating modules 20 at different locations are superimposed and coupled in the time and space dimensions, thereby achieving the purpose of mimicking the coordinated sound production of multiple parts of the human body.

[0023] In some embodiments, the vibration module 30 is a miniature piezoelectric ceramic sheet or a miniature loudspeaker.

[0024] The embodiments provided in this application use a signal driver to drive a micro piezoelectric ceramic sheet or a micro loudspeaker to emit initial sound waves, which makes it easy to set the sound-emitting module 20 in small parts such as the mouth and nose of the robot body.

[0025] In some embodiments, the acoustic cavity 40 is configured to simulate the acoustic characteristics of a human body part corresponding to the location on the robot body.

[0026] The embodiments provided in this application process the initial sound wave through an acoustic cavity 40 that simulates the acoustic characteristics of the human body parts corresponding to the robot body. This process can adjust the resonance frequency and optimize the frequency distribution of the initial sound wave, so that the output target sound wave matches the resonance frequency and resonance peak distribution of the corresponding human voice parts. This makes the resonance peak of the output target sound wave more in line with the characteristics of human voice, and gives the target sound wave a resonance effect that is closer to human voice, thus achieving the goal of making the target sound wave closer to natural human voice.

[0027] In some embodiments, the structure of the acoustic cavity 40 can be modeled in finite element software, and an acoustic propagation simulation system can be constructed by combining transfer functions, physical acoustic models and other tools: based on preset sound parameters (such as target frequency band, resonance intensity, etc.), the propagation loss, resonance gain and phase change law of the sound wave in the acoustic cavity 40 are calculated by the model, and the driving signal parameters that enable the initial sound wave to be accurately converted into the target sound wave containing the corresponding response content after resonance of the cavity are derived in reverse, so as to realize the accurate mapping from sound parameters to driving signal and ensure the matching degree between the acoustic characteristics of the target sound wave and the response content.

[0028] In some embodiments, one or more sound-generating modules include one or more combinations of the following: a first sound-generating module disposed at the head of the robot body, wherein the cavity space of its acoustic cavity is constructed as a biomimetic structure that simulates the acoustic characteristics of the human oral cavity and / or nasal cavity; a second sound-generating module disposed at the neck of the robot body, wherein the cavity space of its acoustic cavity is constructed as a biomimetic structure that simulates the acoustic characteristics of the human pharynx; and a third sound-generating module disposed at the torso of the robot body, wherein the cavity space of its acoustic cavity is constructed as a biomimetic structure that simulates the acoustic characteristics of the human chest.

[0029] In the embodiments provided in this application, by setting the first, second, and third sound-generating modules on the robot body at positions corresponding to common human vocalization sites (mouth, nasal cavity, pharynx, torso, etc.), the spatial auditory perception of human vocalization can be simulated when one or more sound-generating modules emit sound, enhancing the realism and immersion of the sound. Simultaneously, the first sound-generating module processes sound waves by simulating the acoustic cavities of the human mouth and / or nasal cavity, enabling the generation of enhanced resonance in specific high-frequency bands of the sound waves. This results in high-frequency resonance peaks that match the oral and nasal articulation characteristics of natural human voices, such as dentition and labial consonants, making the output sound waves closer to the oral cavity's resonance characteristics in human speech. The first module enhances the expressive quality of the voice, improving its detail and emotional nuances (such as subtle changes in tone). The second module processes sound waves by simulating the acoustic cavity of the human throat, reinforcing specific mid-frequency bands to generate a fundamental frequency and a stable fundamental frequency range. This ensures a smooth transition between high and low frequencies, avoiding the excessive fundamental frequency fluctuations common in traditional electronic synthesized voices. The third module processes sound waves by simulating the acoustic cavity of the human chest, reinforcing specific low-frequency bands to simulate the low-frequency response and sound quality of the human chest during resonance. This functional differentiation allows the robot to adjust its unit coordination to output sounds more suited to different scenarios (such as whispered conversations or the transmission of commands requiring penetrating power), expanding the applicability of human-computer interaction. The robot's voice is made more realistic. At the same time, by having sound modules in different positions emit sound simultaneously, the spatial listening experience of the robot's voice is enhanced, and the layering and fullness of the robot's voice are improved.

[0030] In some embodiments, the second voice module can emit speech content in the mid-frequency band of 200Hz to 2.5kHz.

[0031] In some embodiments, the first sound-generating module is primarily used to generate high-frequency resonant peaks; the third sound-generating module is primarily used to enhance low-frequency resonance and fundamental tone energy.

[0032] In the embodiments provided in this application, a sound wave with a high-frequency resonant peak is generated by a first sound-generating module, and a sound wave with enhanced low-frequency resonance and fundamental tone energy is generated by a third sound-generating module. This allows the output sound waves to couple naturally in space, forming a precisely coupled sound wave that covers the high and low frequency bands of human voice. This makes the robot's voice closer to the characteristics of a natural human voice. The enhancement of the high-frequency resonant peak can enhance the expressive detail of the speech (such as subtle changes in tone in emotional expression), while the enhancement of low-frequency resonance and fundamental tone energy can improve the penetration and appeal of the voice.

[0033] Regarding the embodiments of this application, it should also be noted that, without conflict, the embodiments of this application and the features in the embodiments can be combined with each other to obtain new embodiments.

[0034] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. The scope of protection of this application shall be determined by the scope of the claims.

Claims

1. A multi-cavity sound-generating device for robots, characterized in that, The device includes: Control module and at least one sound-generating module; The control module includes a processor and a signal driver; The sound-generating module includes a vibration module and an acoustic cavity connected to the vibration module; The sound-generating module is connected to the signal driver of the control module via a cable or wireless communication module. The vibration module is configured to generate an initial sound wave driven by the control module. The acoustic cavity performs acoustic processing on the initial sound wave. When there are multiple sound-generating modules, the multiple sound-generating modules are respectively set at different positions on the head, neck or torso of the robot body.

2. The apparatus according to claim 1, characterized in that, The vibration module is a miniature piezoelectric ceramic sheet or a miniature loudspeaker.

3. The apparatus according to claim 1, characterized in that, The acoustic cavity is configured to simulate the acoustic characteristics of a human body part corresponding to the location on the robot body.

4. The apparatus according to claim 1, characterized in that, One or more of the sound-generating modules comprise one or more combinations selected from: The first sound-generating module is located in the head of the robot body, and the cavity space of its acoustic cavity is constructed as a biomimetic structure that simulates the acoustic characteristics of the human oral cavity and / or nasal cavity. The second sound-generating module is located in the neck of the robot body, and the cavity space of its acoustic cavity is constructed as a biomimetic structure that simulates the acoustic characteristics of the human throat. The third sound-generating module is located in the torso of the robot body, and the cavity space structure of its acoustic cavity is a biomimetic structure that simulates the acoustic characteristics of the human chest.

5. The apparatus according to claim 4, characterized in that, The first sound-generating module is mainly used to generate high-frequency resonant peaks; the third sound-generating module is mainly used to enhance low-frequency resonance and fundamental tone energy.