Head movement device of multi-modal sensing robot

By integrating visual and auditory modules into the humanoid robot's head motion device and combining it with motion control, the problem of the low installation position of the AI ​​interaction system was solved, efficient human-computer interaction and sound recognition in complex environments were achieved, and the simulation performance and usability of the robot head were improved.

CN120680481APending Publication Date: 2025-09-23HANGZHOU LANXIN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510947203.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing humanoid robot head vision system is installed at a relatively low position, resulting in poor sound reception efficiency and complex environment sound recognition capabilities of the AI ​​interaction system, making it difficult to meet the needs of efficient interaction in dynamic and complex environments.

Method used

A head movement device for a multimodal perception robot is designed, integrating vision, hearing, and motion control. By rotating the connected head and neck mechanism around the X-axis, combined with the vision module, voice module, and motion module, the installation height and integration of the AI ​​interaction system are improved.

Benefits of technology

The human simulation performance and usability of the head movement device are improved, the sound collection efficiency and complex environment sound recognition ability of the voice module are enhanced, stable nodding/shaking movements of the head are achieved, and the reliability of human-computer interaction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120680481A_ABST
    Figure CN120680481A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent robots, in particular to a head movement device of a multi-modal sensing robot, which comprises a head mechanism and a neck mechanism which are rotatably connected around an X axis, the head mechanism comprises a front mask, a rear mask, a framework, a main board, a visual module, a voice module and a motion module; the front face shield and the rear face shield are correspondingly and detachably connected with the front end face and the rear end face of the framework; the main plate is connected with the framework; the framework is arranged in a cavity defined by the front face cover and the rear face cover. The visual module, the voice module and the motion module are all arranged in the cavity and are all connected with the framework and the main board; wherein the moving die comprises an adapter and a swivel motor; the adapter is connected with the neck mechanism; a machine shell of the rotating head motor is arranged in a clamping groove of the framework. A rotating shaft of the rotating head motor is connected with the adapter so that a machine shell of the rotating head motor can rotate around the axis of the rotating shaft relative to the adapter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent robots, and in particular to a head movement device of a multimodal perception robot. Background Art

[0002] Current humanoid robots' multimodal fusion navigation and obstacle avoidance capabilities meet the basic requirements of structured scenarios, but still face reliability challenges in dynamic and complex environments. Vision and laser obstacle avoidance have established a clear division of labor and complementarity. Vision systems, with their semantic understanding capabilities, dominate complex interactive scenarios such as home and healthcare, while laser systems maintain their precision advantage in industrial and emergency response areas. Multimodal fusion and biomimetic perception will become the technological high ground in the next phase, pushing the humanoid robot obstacle avoidance success rate from the current 90% to 99%.

[0003] Among them, the vision system solutions for humanoid robot heads vary. Most humanoid products still use a single visual camera solution, which is only placed in front of the head for visual obstacle avoidance. Functions such as rear obstacle avoidance are compensated by cameras or lidars installed on the back. Secondly, the robot's AI interaction system is generally placed on its torso, including microphones, speakers, etc. However, the installation position of the AI ​​interaction system is relatively low, resulting in poor sound reception efficiency and ability to recognize and locate sounds in complex environments during normal interactions with humans (1.7 meters tall), reducing the overall performance of the product. Summary of the Invention

[0004] (1) Technical issues to be resolved

[0005] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a head movement device for a multimodal perception robot, a collaborative device for the head and neck of a humanoid robot integrating vision, hearing and motion control, which solves the technical problem of the humanoid robot head being compatible with the AI ​​voice interaction system.

[0006] (2) Technical solution

[0007] In order to achieve the above-mentioned object, the head movement device of the multimodal perception robot of the present invention includes a head mechanism and a neck mechanism connected to rotate around the X-axis;

[0008] The head mechanism includes a front mask, a rear mask, a frame, a main board, a visual module, a voice module, and a motion module; the front mask and the rear mask are detachably connected to the front and rear faces of the frame respectively; the main board is connected to the frame; the frame is built into a cavity enclosed by the front mask and the rear mask; the visual module, the voice module, and the motion module are all built into the cavity and are connected to the frame and the main board;

[0009] Among them, the motion module includes an adapter and a head turning motor; the adapter is connected to the neck mechanism; the casing of the head turning motor is arranged in the slot of the skeleton; the rotating shaft of the head turning motor is connected to the adapter so that the casing of the head turning motor can rotate around the axis of the rotating shaft relative to the adapter.

[0010] Optionally, the front mask includes a front shell and a front transparent cover that are snap-connected; a plurality of first protrusions are provided on the front transparent cover; a plurality of first positioning holes are opened on the frame; the plurality of first protrusions are correspondingly inserted into the plurality of first positioning holes and then connected by bolts;

[0011] The rear cover comprises a rear shell and a rear transparent cover that are snap-fitted together; a plurality of second protrusions are provided on the rear transparent cover; a plurality of second positioning holes are formed on the frame; the plurality of second protrusions are correspondingly inserted into the plurality of second positioning holes and then connected by bolts;

[0012] The front shell and the rear shell are correspondingly connected to the frame.

[0013] Optionally, the visual module includes a light-distributing member and a light-emitting member;

[0014] The light-distributing members made of light-transmitting material are correspondingly provided on both sides of the interior of the front shell;

[0015] The light emitting element is built into the light uniforming element.

[0016] Optionally, the voice module includes an annular silicone pad, an annular microphone, and a sound card PCB board arranged inside the front shell, and both end surfaces of the annular silicone pad are correspondingly attached to the inner wall of the front shell and the annular microphone;

[0017] The front shell is provided with a plurality of first through holes; the annular silicone pad is provided with a plurality of second through holes; the annular microphone is provided with a plurality of sound pickup holes; the first through holes, the second through holes and the sound pickup holes are aligned along the axial direction;

[0018] The sound card PCB is arranged inside the frame; the ring microphone can transmit the sound signal to the sound card PCB and then to the main board.

[0019] Optionally, the vision module includes a binocular camera, a monocular RGB camera, an auxiliary RGB camera, an interactive screen, a telefocus monocular camera and an acquisition board;

[0020] The three groups of binocular cameras are arranged in an equilateral triangle in the front, left rear and right rear of the skeleton; the two monocular RGB cameras are arranged on the front side of the skeleton; the auxiliary RGB cameras are respectively arranged on the left rear and right rear of the skeleton to assist the binocular cameras in imaging; the interactive screen is arranged in the front lower part of the skeleton; the telefocus monocular camera is arranged in the front middle part of the skeleton; and the acquisition board is arranged on the top of the skeleton;

[0021] Among them, the telefocus monocular camera can transmit the image of the portrait head to the main board, and the main board locates the position of the speaking portrait through face recognition technology and / or lip recognition technology; the three groups of binocular cameras correspondingly aggregate visual imaging data to the acquisition board and merge it into 360° panoramic perception; the two monocular RGB cameras correspondingly aggregate visual imaging data to the acquisition board for visual obstacle avoidance.

[0022] Optionally, the main board processes the portrait position information and sound information, gives an accurate voice answer through the voice model, and then outputs audio through the voice module.

[0023] Optionally, the binocular camera and / or the monocular RGB camera are installed on the skeleton in an oblique downward manner.

[0024] Optionally, the voice module further includes a square silicone pad and a first speaker arranged inside the rear shell, and both end surfaces of the square silicone pad are correspondingly attached to the inner wall of the rear shell and the first speaker.

[0025] Optionally, the neck mechanism comprises a base, a nodding motor and a housing;

[0026] The output shaft of the nodding motor is connected to the adapter to drive the head mechanism to rotate around the X-axis;

[0027] The bottom end of the shell is connected to the base, and the top end is provided with an air-avoiding groove for avoiding the air-avoiding of the adapter.

[0028] Optionally, the front transparent cover and the rear transparent cover are both acrylic masks with light transmittance and curvature.

[0029] (3) Beneficial effects

[0030] The beneficial effects of the present invention are:

[0031] The combination of the visual, voice, and motion modules significantly enhances the human head simulation performance and integration of the head motion device. Furthermore, the integration increases the installation height of the AI ​​interactive system, improves the voice module's sound pickup efficiency and its ability to identify and locate sounds in complex environments, and enhances the performance of the head motion device.

[0032] The head mechanism and the neck mechanism are connected by rotating around the X-axis, thereby realizing the nodding action of the head mechanism. The shaking action of the head mechanism is then realized through the motion module, so as to achieve the purpose of the head motion device to simulate the nodding / shaking of the human body. Compared with the method in which the housing and the neck mechanism are fixed, and the rotating shaft is connected to the head mechanism, the present invention increases the contact force area between the motion mechanism and the head mechanism by using the housing as the driving end and connecting it to the card slot; and then by adding an adapter, the contact force area between the rotating shaft and the neck mechanism is increased, and the device is more stable when driven, and the adapter enlarges the distance between the head mechanism and the neck mechanism. During the nodding / shaking action, it is not easy for the shells of the head mechanism and the neck mechanism to produce abnormal noise due to friction. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a schematic diagram of the structure of the head movement device of the multimodal perception robot of the present invention from a front perspective;

[0034] Figure 2 This is a schematic diagram of the structure of the head movement device of the multimodal perception robot of the present invention from a rear perspective;

[0035] Figure 3 This is an exploded schematic diagram of the head mechanism of the present invention from a front perspective;

[0036] Figure 4 is an exploded schematic diagram of the head mechanism of the present invention from a rear perspective;

[0037] Figure 5 Schematic diagram of the head mechanism and neck mechanism of the present invention;

[0038] Figure 6 A schematic diagram of the posture of the head movement device of the present invention during nodding action;

[0039] Figure 7 A schematic diagram of the posture of the head movement device of the present invention during head turning action;

[0040] Figure 8 An exploded view of the head mechanism of the present invention;

[0041] Figure 9 is a structural schematic diagram of the first positioning hole of the present invention;

[0042] Figure 10 is a schematic structural diagram of the first protruding member of the present invention;

[0043] Figure 11 It is a schematic structural diagram of the buckle edge and lip edge of the present invention;

[0044] Figure 12 Schematic diagram of the exploded front shell, rear shell and head mechanism of the present invention;

[0045] Figure 13 A front view of the head movement device of the multimodal perception robot of the present invention;

[0046] Figure 14 for Figure 13 Cross-sectional view along the middle edge FF;

[0047] Figure 15 This is a schematic diagram of the layout of the monocular RGB camera and binocular camera of the present invention arranged tilted downward;

[0048] Figure 16 Schematic diagram of three sets of binocular cameras arranged in an equilateral triangle according to the present invention;

[0049] Figure 17 is an exploded schematic diagram of the neck mechanism of the present invention from a front perspective;

[0050] Figure 18 It is an exploded schematic diagram of the neck mechanism of the present invention from a rear perspective.

[0051] [Description of Reference Numerals]

[0052] 110: Head mechanism; 110-1: Front shell; 110-2: Rear shell; 110-3: Front transparent cover; 110-3-1: First protrusion; 110-4: Rear transparent cover; 110-5: Light diffuser; 110-6: Light emitting element; 110-7: Frame; 110-7-1: First positioning hole; 110-8: Acquisition board; 110-9: Ring-shaped silicone pad; 110-10: Square silicone pad; 110-11: Adapter; 110-12: Rotating head motor; 110-13: Monocular RGB camera; 110-14: Binocular camera; 110-15: Auxiliary RGB camera; 110-16: Interactive screen; 110-17: Telefocus monocular camera; 110-18: Sound card PCB; 110-19: Main board; 110-20: Heat dissipation holes;

[0053] 120: Neck mechanism; 120-1: Base; 120-2: Housing; 120-3: Nodding motor;

[0054] 130 - 1 : ring microphone; 130 - 2 : first speaker; 130 - 3 : second speaker. DETAILED DESCRIPTION

[0055] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.

[0056] It should be noted that all directional indications in the embodiments of the present invention (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0057] In addition, the terms "first," "second," and so on, used in this disclosure are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referenced. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this disclosure, "plurality" means at least two, such as two or three, unless otherwise specifically defined.

[0058] In the present invention, unless otherwise specified or limited, the terms "connect," "fix," etc. should be understood in a broad sense. For example, "fix" can mean fixed connection, detachable connection, or integration; "connection" can mean mechanical connection or electrical connection; it can mean direct connection or indirect connection through an intermediate medium; it can mean internal communication between two elements or interaction between two elements, unless otherwise specified. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0059] See also Figures 1 to 5 as well as Figure 13 The present invention provides a head movement device of a multimodal perception robot, which includes a head mechanism 110 and a neck mechanism 120 that are connected and rotated around the X-axis, where the X-axis is the horizontal direction of the line connecting the left and right ears; the head mechanism 110 includes a front mask, a rear mask, a skeleton 110-7, a main board 110-19, a visual module, a voice module, and a motion module; the front mask and the rear mask are detachably connected to the front and rear faces of the skeleton 110-7, and screw connection is optional; the main board 110-19 is connected to the skeleton 110-7; the skeleton 110-7 is built into the front mask and the rear mask. The housing of the rotating head motor 110-12 is arranged in the slot of the skeleton 110-7; the shaft of the rotating head motor 110-12 is connected to the adapter 110-11 so that the housing of the rotating head motor 110-12 can rotate around the axis of the shaft relative to the adapter 110-11.

[0060] Because the head motion device contains numerous electrical components (control boards, acquisition boards, cameras, etc.), its power consumption is relatively high. Therefore, heat dissipation holes 110-20 are provided on both sides of the front shell 110-1. This is achieved by rationally positioning the components on the frame 110-7. The head rotation motor 110-12 and the nodding motor 120-3 can be articulated motors (servo motors) or other rotating mechanisms.

[0061] The front and rear covers form the outline of the head, the skeleton serves as the mounting bracket for the various components, and the main board serves as the control center. The visual module is used to locate the human body, the voice module is used to receive and output audio, and the motion module is used to achieve nodding / shaking movements. The combination of the visual module, voice module, and motion module significantly improves the simulation performance and integration level of the human head of the head motion device. Furthermore, the installation height of the AI ​​interactive system is increased after integration, which improves the voice module's sound reception efficiency and its ability to identify and locate sounds in complex environments, thereby enhancing the performance of the head motion device.

[0062] like Figure 6 As shown, the head mechanism 110 and the neck mechanism 120 are connected to rotate around the X axis, thereby realizing the nodding action of the head mechanism 110; Figure 7 As shown, the head mechanism 110 is then driven by the motion module to achieve the purpose of the head motion device to simulate human nodding / shaking. The adapter 110-11 is secured by the neck mechanism 120, the shaft of the head-turning motor 110-12 is connected to the adapter 110-11, and the housing of the head-turning motor 110-12 is engaged with the slot of the frame 110-7, enabling the housing to rotate relative to the neck mechanism 120 about the shaft to achieve the head-turning motion. Compared with the method in which the housing and the neck mechanism 120 are fixed and the rotating shaft is connected to the head mechanism 110, the present invention increases the contact force area between the motion mechanism and the head mechanism 110 by using the housing as the driving end and connecting it to the card slot; and then by adding an adapter 110-11, the contact force area between the rotating shaft and the neck mechanism 120 is increased, and the device is more stable when driven, and the adapter 110-11 increases the distance between the head mechanism 110 and the neck mechanism 120. During nodding / shaking movements, it is less likely for abnormal noise to be generated between the shells of the head mechanism 110 and the neck mechanism 120 due to friction.

[0063] like Figures 8 to 11As shown, the front mask includes a front shell 110-1 and a front transparent cover 110-3 that are snap-fitted together; a plurality of first protrusions 110-3-1 are provided on the front transparent cover 110-3; a plurality of first positioning holes 110-7-1 are opened on the skeleton 110-7; the plurality of first protrusions 110-3-1 are correspondingly engaged with the plurality of first positioning holes 110-7-1 and are connected by bolts; the rear mask includes a rear shell 110-2 and a rear transparent cover 110-4 that are snap-fitted together; a plurality of second protrusions are provided on the rear transparent cover 110-4; a plurality of second positioning holes are opened on the skeleton 110-7; the plurality of second protrusions are correspondingly engaged with the plurality of second positioning holes and are connected by bolts; the front shell 110-1 and the rear shell 110-2 are correspondingly connected to the skeleton 110-7, and screw connection is optional. In this embodiment, a buckle edge is formed on the front housing 110-1, and a corresponding lip edge is formed on the front transparent cover 110-3. The buckle edge and the lip edge engage and connect, enabling quick assembly and disassembly of the front housing 110-1 and the front transparent cover 110-3. The first protrusion 110-3-1 is a circular boss formed by connecting a large cylinder and a small cylinder arranged concentrically. The free end of the circular boss is stepped, and the shape of the first positioning hole 110-7-1 is correspondingly configured, which can effectively improve positioning accuracy and connection strength.

[0064] The traditional connection method of the front and rear shells is to fix them directly with screws. When the number of times they are disassembled and assembled is large, large repeated positioning errors will occur at the threaded connection, causing the curvature of the transparent mask in front of the internal camera (front transparent mask 110-3 and rear transparent mask 110-4) to change, thereby requiring recalibration of the internal camera components, which prolongs maintenance time. The present invention, based on the screw connection, adds a positioning structure of a first protrusion 110-3-1 and a first positioning hole 110-7-1, which effectively improves the installation accuracy of the transparent mask. Even if it is disassembled and assembled multiple times, the installation accuracy can be effectively guaranteed. The position error of the transparent mask during the head assembly and maintenance process can be controlled within 0.1mm, eliminating the need to recalibrate the internal camera components, shortening maintenance time and solving the maintainability defects of the head mask. Similarly, the rear shell 110-2 and the rear transparent mask 110-4 can be set up in an analogous manner to the front shell 110-1 and the front transparent mask 110-3, and will not be repeated here.

[0065] Furthermore, the visual module includes a light-distributing member 110-5 and a light-emitting member 110-6; light-distributing members 110-5 made of a translucent material are provided on both sides of the interior of the front shell 110-1; and the light-emitting member 110-6 is built into the light-distributing member 110-5. Since the front transparent cover 110-3 and the rear transparent cover 110-4 are curved, their curvature will affect the imaging effect of each camera inside. The present invention arranges the light-distributing member 110-5 and the light-emitting member 110-6 on both sides of the interior of the front shell 110-1, freeing up space for the installation of the camera, allowing the camera to be installed close to the inner surface of the translucent mask, thereby reducing imaging differences. In addition, the light-distributing member 110-5 and the light-emitting member 110-6 are installed independently of the frame 110-7 and close to the translucent mask, so that the light-distributing member 110-5 and the light-emitting member 110-6 have less impact on the layout of the components on the frame 110-7, and the high integration performance is more reliable.

[0066] See also Figure 12 The voice module includes a ring-shaped silicone pad 110-9, a ring-shaped microphone 130-1, and a sound card PCB 110-18, which are disposed within the front housing 110-1. The two ends of the ring-shaped silicone pad 110-9 correspond to the inner wall of the front housing 110-1 and the ring-shaped microphone 130-1. The front housing 110-1 is provided with a plurality of first through-holes. The ring-shaped silicone pad 110-9 is provided with a plurality of second through-holes. The ring-shaped microphone 130-1 is provided with a plurality of sound pickup holes. The first through-holes, second through-holes, and sound pickup holes are aligned axially. The sound card PCB 110-18 is disposed within the frame 110-7. The ring-shaped microphone 130-1 can transmit sound signals to the sound card PCB 110-18, and then to the mainboard 110-19. Specifically, the ring-shaped ring-shaped silicone pad 110-9 and the ring-shaped microphone 130-1 provide a better sound pickup effect. Secondly, a ring-shaped silicone pad 110-9 is arranged between the ring microphone 130-1 and the inner wall of the front shell 110-1. The ring-shaped silicone pad 110-9 allows the ring microphone 130-1 to be set tightly against the inner wall of the front shell 110-1, further improving the sound reception effect. In addition, the coaxial arrangement of the first through hole, the second through hole and the pickup hole allows external sound to be accurately transmitted to the ring microphone 130-1, further improving the sound reception effect. The sound card PCB board further processes the sound signal and transmits it to the main board 110-19 for processing. The main board 110-19 compares the sound signal with the voice model, gives an accurate answer, and finally outputs the corresponding audio to achieve the purpose of human-computer interaction and communication.

[0067] Optionally, the voice module also includes a square silicone pad 110-10 and a first speaker 130-2 arranged inside the rear shell 110-2, and the two end surfaces of the square silicone pad 110-10 correspond to the inner wall of the rear shell 110-2 and the first speaker 130-2. Similarly, the square silicone pad 110-10 enables the first speaker 130-2 to be set close to the inner wall of the rear shell 110-2, thereby improving the sound effect of the head movement device and effectively avoiding the echo inside the cavity. Secondly, howling will be generated when the microphone and the speaker are acoustically coupled. The present invention adds a square silicone pad 110-10, which can play a shock-absorbing effect, effectively preventing the vibration generated by the first speaker 130-2 from causing the shell of the head mechanism 110 to vibrate, thereby avoiding acoustic coupling between the microphone and the speaker, and ensuring the normal use of the ring microphone 130-1 and the first speaker 130-2. In addition, a sound cavity is formed between the first speaker 130-2 and the inner side of the rear shell 110-2, and a sound hole is opened on the rear shell 110-2. Sound is transmitted to the outside through the sound cavity and the sound hole, ensuring that the first speaker 130-2 can work normally.

[0068] like Figure 3 、 Figure 4 、 Figure 14 、 Figure 15 and Figure 16As shown, the visual module includes a binocular camera 110-14, a monocular RGB camera 110-13, an auxiliary RGB camera 110-15, an interactive screen 110-16, a telefocus monocular camera 110-17 and an acquisition board 110-8; three groups of binocular cameras 110-14 are arranged in an equilateral triangle in front of the skeleton 110-7 (at the bridge of the nose), the left rear and the right rear; two monocular RGB cameras 110-13 are arranged on the front side of the skeleton 110-7, at the position of the eyes; auxiliary RGB cameras 110-15 are correspondingly arranged on the left rear and the right rear of the skeleton 110-7 to assist the binocular camera 110-14 in imaging; the interactive screen 110-16 is arranged in the front lower part of the skeleton 110-7, at the position of the mouth, and can be seen through the front transparent cover 11 0-3 sees the picture displayed on the interactive screen 110-16; the telephoto monocular camera 110-17 is set in the front middle of the skeleton 110-7, at the tip of the nose; the acquisition board 110-8 is set on the top of the skeleton 110-7; among them, the telephoto monocular camera 110-17 can transmit the picture of the portrait's head to the main board 110-19, and the main board 110-19 locates the position of the speaking portrait through face recognition technology and / or lip recognition technology; the three groups of binocular cameras 110-14 respectively aggregate the visual imaging data to the acquisition board 110-8, and merge it into 360° full-domain / panoramic perception, so as to realize the visual navigation function without blind spots; the two monocular RGB cameras 110-13 respectively aggregate the visual imaging data to the acquisition board 110-8 for visual obstacle avoidance.

[0069] Specifically, the monocular camera 110-17 transmits the image it sees to the main board 110-19. The main board 110-19 can recognize the human figure (head) appearing in front of it, and recognize that the human figure is speaking through face recognition / lip recognition technology, and locate the position of the human figure. The three groups of binocular cameras 110-14 are installed with the central axis of the head as the starting point, and the three groups of binocular cameras 110-14 are equally divided into three groups and formed into an equilateral triangle structure to achieve visual full-domain / panoramic perception (360°). Optionally, depending on the field of view of the camera, the number of binocular cameras 110-14 that achieve full-domain perception can be adjusted to four groups. It should be noted that the cameras, circuits and signal control of the present invention are prior art. The present invention aims to achieve the integration of each camera component, visual signal acquisition module, voice module, visual interaction module, and head joint motion module into the same head motion device through a reasonable layout of each component, and achieve visual full-domain / panoramic perception (360°), improve the maintainability of the head, and solve the acoustic interference problem of the speaker and microphone.

[0070] By integrating multiple binocular camera arrays 110-14 inside the head of the humanoid robot, full-area / panoramic perception (360°) is achieved while avoiding the need to install and arrange cameras or lidars inside the torso of the humanoid robot, thereby increasing the layout space inside the torso and facilitating the installation and arrangement of the required components inside the torso.

[0071] Furthermore, the main board 110-19 processes the portrait position information and sound information, gives an accurate voice answer through the voice model, and then outputs the audio through the voice module. The voice model is a voice database recorded in the main board 110-19, which is used to give answers to the received external sounds to achieve the purpose of human-computer interaction and communication. Specifically, the head telephoto monocular camera (110-17) first identifies the speaking portrait through the portrait + lip shape recognition technology, and locates the position of the speaking portrait relative to the robot; the array sound signal is received by the recognition ring microphone 130-1, and the sound signal of the designated person is accurately screened and transmitted to the sound card PCB board 110-18, and then the data is transmitted to the main board 110-19. The main board 110-19 locates the specific direction of the picked-up sound through the face recognition + lip shape recognition technology, and uses the received sound data to identify the words spoken by the portrait in the direction in the external sound source, and then gives an accurate answer through the voice model, and then outputs the audio through the speaker. This AI interaction solution first locates the speaker's position, then the specific direction of the sound. Combining these two, it identifies the speaker's utterance and finally outputs the audio response after processing it through a speech model. This AI interaction solution effectively prevents audio interaction errors, improves human-computer interaction performance, and effectively avoids the difficulty of identifying the voice of a specific speaker in noisy environments.

[0072] In addition, the binocular camera 110-14 and / or the monocular RGB camera 110-13 are installed on the skeleton 110-7 at an angle downward. The mounting surfaces of the three groups of binocular cameras 110-14 on the skeleton 110-7 are at a certain angle to the vertical plane to reduce the ground blind spot after imaging. Similarly, the two monocular RGB cameras 110-13 used for visual obstacle avoidance in the front are also installed at an angle to improve obstacle avoidance performance. The mounting surface of the telephoto monocular camera 110-17 on the skeleton 110-7 is parallel to the vertical plane. Its function is to collect portrait information directly in front and does not require tilted installation. The correct angle layout of the visual navigation camera (binocular camera 110-14), the visual obstacle avoidance camera (monocular RGB camera 110-13), and the camera for face recognition / lip shape recognition (telephoto monocular camera 110-17) achieves the purpose of voice interaction within a distance of 0.5-1.5 meters and close-range obstacle recognition. According to the height of the final head motion device used (such as a 1.6-meter robot or a 1.7-meter robot), the installation angle and installation height of each camera can be adjusted appropriately.

[0073] See also Figure 17and Figure 18 The neck mechanism 120 includes a base 120-1, a nodding motor 120-3 and a shell 120-2; the output shaft of the nodding motor 120-3 is connected to the adapter 110-11 to drive the head mechanism 110 to rotate around the X-axis; the bottom end of the shell 120-2 is connected to the base 120-1, and an air avoidance groove is provided at the top to avoid the air adapter 110-11, further reducing the abnormal noise caused by the friction of the shell. Specifically, the nodding motor 120-3 is connected to the adapter 110-11 at the bottom of the head mechanism 110. When the output shaft of the nodding motor 120-3 rotates, it will drive the adapter 110-11 to rotate together, thereby realizing the nodding action. The shell 120-2 is fixed to the base 120-1 through two mounting holes on the back. Optionally, a second speaker 130-3 is provided on the inclined surface in front of the base 120-1, and the housing 120-2 has a sound-transmitting hole at the installation position of the second speaker 130-3 to achieve sound transmission. The positions of the first speaker 130-2 and the second speaker 130-3 can be adjusted appropriately, such as being installed at the head, ears or neck. The present invention uses the shock-absorbing design of the square silicone pad 110-10 to ensure that the sound reception of the ring microphone 130-1 and the sound playback of the dual-channel speakers (the first speaker 130-2 and the second speaker 130-3) do not interfere with each other.

[0074] Secondly, both the front and rear transparent covers 110-3 and 110-4 are acrylic masks with both light transmittance and curvature. The transmittance and curvature of the front and rear acrylic masks allow the internal cameras to function properly, making the internal structure virtually invisible from the outside and ensuring that the internal cameras are not exposed to the outside world. Because the front and rear transparent covers 110-3 and 110-4 are curved, their curvature affects the imaging quality of the internal cameras. Therefore, during installation, the lenses must be close to the inner surfaces of the front and rear transparent covers to minimize image differences.

[0075] In summary, the present invention integrates various camera components, visual signal acquisition modules, voice modules, visual interaction modules, and head joint motion modules into the same head motion device, and optimizes the spatial layout to achieve full-domain / panoramic visual perception (360°), improve the maintainability of the head, and solve the acoustic interference problem between speakers and microphones.

[0076] It should be understood that the above description of the specific embodiments of the present invention is merely for the purpose of illustrating the technical approach and features of the present invention. Its purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. However, the present invention is not limited to the above-described specific embodiments. Any changes or modifications made within the scope of the claims of the present invention are intended to be covered by the scope of protection of the present invention.

Claims

1. A head movement device for a multimodal perception robot, characterized in that: The head movement device comprises a head mechanism (110) and a neck mechanism (120) connected to rotate around an X-axis; The head mechanism (110) comprises a front mask, a rear mask, a skeleton (110-7), a main board (110-19), a visual module, a voice module and a motion module; the front mask and the rear mask are detachably connected to the front end face and the rear end face of the skeleton (110-7) respectively; the main board (110-19) is connected to the skeleton (110-7); the skeleton (110-7) is built into a cavity enclosed by the front mask and the rear mask; the visual module, the voice module and the motion module are all built into the cavity and are all connected to the skeleton (110-7) and the main board (110-19); The motion module comprises an adapter (110-11) and a rotating head motor (110-12); the adapter (110-11) is connected to the neck mechanism (120); the housing of the rotating head motor (110-12) is arranged in a slot of the skeleton (110-7); the rotating shaft of the rotating head motor (110-12) is connected to the adapter (110-11), so that the housing of the rotating head motor (110-12) can rotate relative to the adapter (110-11) around the axis of the rotating shaft.

2. The head movement device of the multimodal perception robot according to claim 1, characterized in that: The front mask comprises a front shell (110-1) and a front transparent cover (110-3) that are snap-connected; a plurality of first protrusions (110-3-1) are provided on the front transparent cover (110-3); a plurality of first positioning holes (110-7-1) are provided on the frame (110-7); the plurality of first protrusions (110-3-1) are correspondingly inserted into the plurality of first positioning holes (110-7-1) and then connected via bolts; The rear cover comprises a rear shell (110-2) and a rear transparent cover (110-4) that are snap-connected; a plurality of second protrusions are provided on the rear transparent cover (110-4); a plurality of second positioning holes are provided on the frame (110-7); the plurality of second protrusions are correspondingly inserted into the plurality of second positioning holes and then connected via bolts; The front shell (110-1) and the rear shell (110-2) are correspondingly connected to the skeleton (110-7).

3. The head movement device of the multimodal perception robot according to claim 2, characterized in that: The visual module includes a light-distributing component (110-5) and a light-emitting component (110-6); The light-distributing members (110-5) made of a light-transmitting material are correspondingly provided on both sides of the interior of the front shell (110-1); The light-emitting component (110-6) is built into the light-distributing component (110-5).

4. The head movement device of the multimodal perception robot according to claim 2, characterized in that: The voice module comprises an annular silicone pad (110-9), an annular microphone (130-1) and a sound card PCB (110-18) arranged inside the front shell (110-1), and two end surfaces of the annular silicone pad (110-9) are correspondingly attached to the inner wall of the front shell (110-1) and the annular microphone (130-1); The front shell (110-1) is provided with a plurality of first through holes; the annular silicone pad (110-9) is provided with a plurality of second through holes; the annular microphone (130-1) is provided with a plurality of sound pickup holes; the first through holes, the second through holes and the sound pickup holes are aligned along the axial direction; The sound card PCB board (110-18) is arranged inside the skeleton (110-7); the ring microphone (130-1) can transmit sound signals to the sound card PCB board (110-18) and then to the main board (110-19).

5. The head movement device of the multimodal perception robot according to claim 4, characterized in that: The vision module includes a binocular camera (110-14), a monocular RGB camera (110-13), an auxiliary RGB camera (110-15), an interactive screen (110-16), a telefocus monocular camera (110-17) and an acquisition board (110-8); Three groups of binocular cameras (110-14) are arranged in an equilateral triangle position in front, on the left rear, and on the right rear of the skeleton (110-7); two monocular RGB cameras (110-13) are arranged on the front side of the skeleton (110-7); auxiliary RGB cameras (110-15) are respectively arranged on the left rear and right rear of the skeleton (110-7) to assist the binocular cameras (110-14) in imaging; the interactive screen (110-16) is arranged at the front lower part of the skeleton (110-7); the telefocus monocular camera (110-17) is arranged at the front middle part of the skeleton (110-7); and the acquisition board (110-8) is arranged on the top of the skeleton (110-7); The telefocus monocular camera (110-17) can transmit an image of a person's head to the main board (110-19), and the main board (110-19) locates the position of the speaking person through face recognition technology and / or lip recognition technology; the three groups of binocular cameras (110-14) correspondingly aggregate visual imaging data to the acquisition board (110-8) to merge into a 360-degree panoramic perception; and the two monocular RGB cameras (110-13) correspondingly aggregate visual imaging data to the acquisition board (110-8) for visual obstacle avoidance.

6. The head movement device of the multimodal perception robot according to claim 5, characterized in that: The main board (110-19) processes the portrait position information and the sound information, gives an accurate voice answer through the voice model, and then outputs the audio through the voice module.

7. The head movement device of the multimodal perception robot according to claim 5, characterized in that: The binocular camera (110-14) and / or the monocular RGB camera (110-13) are installed on the skeleton (110-7) in a tilted downward direction.

8. The head movement device of the multimodal perception robot according to claim 2, characterized in that: The voice module further includes a square silicone pad (110-10) and a first speaker (130-2) arranged inside the rear shell (110-2), and two end surfaces of the square silicone pad (110-10) are correspondingly attached to the inner wall of the rear shell (110-2) and the first speaker (130-2).

9. The head movement device of the multimodal perception robot according to claim 8, characterized in that: The neck mechanism (120) comprises a base (120-1), a nodding motor (120-3) and a housing (120-2); The output shaft of the nodding motor (120-3) is connected to the adapter (110-11) so as to be able to drive the head mechanism (110) to rotate around the X-axis; The bottom end of the housing (120-2) is connected to the base (120-1), and a top end is provided with an air-avoiding groove for avoiding the adapter (110-11).

10. The head movement device of the multimodal perception robot according to claim 2, characterized in that: The front transparent cover (110-3) and the rear transparent cover (110-4) are both acrylic masks with light transmittance and curvature.