Robot and head assembly thereof, control method, data processing method
Patent Information
- Application Number
- CN202610846669.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-08-18
AI Technical Summary
但是相关技术中,机器人多传感器融合数据的精度不佳,导致机器人控制效果较差
[0026] The head assembly described in this specification includes a housing and a sensing component disposed within the housing. The sensing component includes a camera, a microphone array, and a positioning sensor. The microphone array comprises multiple microphones spatially distributed and non-centrally symmetrically arranged. The main control unit performs fusion processing based on data from each sensing component and executes robot control based on the fused data. The head assembly utilizes multi-sensor fusion to improve the robot's environmental perception capabilities. Furthermore, the spatially distributed microphone array allows for the perception of sound wave propagation path differences in the vertical direction using microphones at different heights. This enables accurate localization and identification of sound sources at different heights. Simultaneously, the microphones located at the lower level can accurately pick up sound source data from below. Combined with beamforming algorithms, the received sound beam can be directionally focused below the robot, effectively improving voice wake-up and recognition success rates in scenarios where sound sources are located at lower levels. In addition, the non-centrally symmetrical arrangement of microphones on the same horizontal plane effectively eliminates or mitigates problems such as positioning ambiguity caused by mirrored sound sources, improving sound source localization accuracy and thus enhancing the robot's perception accuracy.
Smart Images

Figure CN122584318A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of embodied robot technology, specifically to a robot and its head assembly, control method, and data processing method. Background Technology
[0002] As a product of the integration of artificial intelligence and mechanical engineering, embodied robots have a head system similar to the human brain, undertaking core functions such as visual perception, voice interaction, and data processing. The perception system of the robot's head often includes multi-source heterogeneous sensors, which drive the robot's interactive behavior through the fusion and processing of multimodal heterogeneous data.
[0003] For example, a robot's head perception system typically integrates multiple sensors such as cameras, microphones, LiDAR, and positioning modules. The robot uses multi-sensor fusion algorithms to fuse data from multiple sources and then uses the fused data to locate the target object. However, in related technologies, the accuracy of multi-sensor fusion data for robots is not good, resulting in poor robot control performance. Summary of the Invention
[0004] To address at least one of the aforementioned technical problems, this specification provides a robot head assembly, a robot, a data processing method, and a robot control method.
[0005] In a first aspect, embodiments of this specification provide a robot head assembly, including: The housing has at least one receiving cavity; The motherboard and various types of sensing components are disposed in the cavity, including a microphone array for acquiring audio data; The microphone array includes multiple microphones, which are spaced apart in the horizontal and vertical directions of the receiving cavity, and at least some of the microphones are non-centrally symmetrically distributed in the horizontal plane; The main control unit on the motherboard is configured to perform multi-sensor fusion of the sensing data collected by the various types of sensing components to obtain fused sensing data, and to perform robot control based on the fused sensing data.
[0006] In some possible implementations, the microphone array includes a front microphone, a rear microphone, and side microphones; The front microphone includes at least three microphones, which are spaced apart on the front side of the receiving cavity; The side-mounted microphone includes at least two left-side microphones and at least two right-side microphones. The at least two left-side microphones are vertically spaced on the left side of the receiving cavity, and the at least two right-side microphones are vertically spaced on the right side of the receiving cavity. The rear microphone includes at least one microphone, which is located on the rear side of the receiving cavity.
[0007] In some possible implementations, the front microphone, the rear microphone, at least one of the left microphones and at least one of the right microphones are positioned close to the top of the head assembly, while the remaining left microphones and right microphones are positioned close to the neck of the head assembly.
[0008] In some possible implementations, the head assembly also includes a display unit comprising a curved display screen and an ambient light sensor; The display screen is located on the front side of the receiving cavity; The ambient light sensor is mounted on a flexible circuit board that carries at least part of the front microphone, and is detachably electrically connected to the motherboard through the flexible circuit board. The main control unit is configured to adjust the display parameters of the display screen based on the ambient light intensity data collected by the ambient light sensor.
[0009] In some possible implementations, the sensing component is disposed on a flexible circuit board and is detachably electrically connected to the motherboard via the flexible circuit board.
[0010] In some possible implementations, the housing is provided with at least one expansion interface for pluggable connection to at least one external sensing component; The main control unit is configured to identify the type of the external sensing component and load the corresponding driver configuration when the external sensing component is detected to be connected.
[0011] In some possible implementations, the main control unit is configured to receive time-stamped data collected by the various types of sensing components, align the data based on the system clock to obtain synchronization frame data, and perform multi-sensor fusion based on the synchronization frame data to obtain the fused sensing data, wherein the synchronization frame data represents the data collected by each sensing component at the same physical moment.
[0012] In some possible implementations, the sensing component is configured to configure a first timestamp for the collected data based on its own local clock; The main control unit is configured to map the first timestamp of the sensing component to the system timestamp based on a pre-set clock mapping relationship, wherein the clock mapping relationship represents the correspondence between the system clock and the local clock of the sensing component, and the main control unit performs alignment processing on the data according to the system timestamp to obtain synchronization frame data.
[0013] In some possible implementations, the main control unit is configured to determine the position information of the target object based on the fused sensing data, and generate control commands based on the position information, the control commands being used to control the pose of the head assembly and / or the robot body.
[0014] In some possible implementations, the various types of sensing components also include cameras and positioning sensors; The camera is used to acquire image data, and the camera includes at least one of the following: an RGB camera and a depth camera; The positioning sensor is used to collect positioning data of the target object relative to the head component, and the positioning sensor includes at least one of the following: an ultra-wideband positioning module, a WiFi positioning module, and a Bluetooth positioning module.
[0015] Secondly, some embodiments of this specification provide a data processing method applied to the main control unit of the robot head assembly in any of the above embodiments, the method comprising: Data collected by various types of sensing components is acquired. The data collected by the sensing components includes a first timestamp, which is configured by the sensing component based on its own local clock. Based on a pre-set clock mapping relationship, the first timestamp of the sensing component is mapped to the system timestamp, wherein the clock mapping relationship represents the correspondence between the system clock and the local clock of the sensing component; The data is aligned according to the system timestamp to obtain synchronization frame data, which represents the data collected by each sensing component at the same physical moment. Multi-sensor fusion is performed based on the synchronous frame data to obtain fused sensing data.
[0016] In some possible implementations, the process of pre-setting the clock mapping relationship includes: For any sensing component, obtain the offset between the local clock and the system clock of the sensing component, and collect multiple sets of data pairs between the local clock and the system clock based on the offset; By performing linear fitting or constructing a data lookup table based on multiple sets of data pairs, the clock mapping relationship between the local clock and the system clock of the sensing component can be obtained.
[0017] In some possible implementations, the data is aligned according to the system timestamp to obtain synchronization frame data, including: Based on the system timestamp, the data is timestamped and interpolated to determine the data collected by each sensing component at the same physical moment as synchronization frame data.
[0018] Thirdly, some embodiments of this specification provide a robot control method, including: Acquire fused perception data from multiple sensing components of the robot, wherein the fused perception data is generated by the method described in any of the above embodiments; The location information of the target object is determined based on the fused sensing data, and control commands are generated based on the location information. Robot control is executed based on the control commands.
[0019] In some possible implementations, human-machine interaction control is performed based on the control commands, including at least one of the following: The azimuth and / or pitch angle of the head assembly are controlled based on the control command, so that the front side of the head assembly faces the target object; Based on the control command, the microphone array is controlled to perform beamforming, so that the acquisition beam of the microphone array targets the target object; The robot's body is moved according to the control commands, causing the robot to move toward the location of the target object.
[0020] Fourthly, some embodiments of this specification provide a robot, including: a body; and a head assembly as described in any of the above embodiments, the head assembly being disposed on the body.
[0021] Fifthly, some embodiments of this specification provide a data processing apparatus applied to the robot head assembly of any of the above embodiments, comprising: The clock mapping unit is configured to acquire data collected by the sensing component. The data is configured with a first timestamp, which is configured by the sensing component based on its own local clock. Based on a pre-set clock mapping relationship, the first timestamp of the sensing component is mapped to the system timestamp. The clock mapping relationship represents the correspondence between the system clock and the local clock of the sensing component. The frame synchronization unit is configured to align the data according to the system timestamp to obtain synchronization frame data, wherein the synchronization frame data represents the data collected by each sensing component at the same physical moment. The data fusion unit is configured to perform multi-sensor fusion based on the synchronization frame data to obtain fused sensing data.
[0022] In some implementations, the clock mapping unit is configured as follows: For any sensing component, obtain the offset between the local clock and the system clock of the sensing component, and collect multiple sets of data pairs between the local clock and the system clock based on the offset; By performing linear fitting or constructing a data lookup table based on multiple sets of data pairs, the clock mapping relationship between the local clock and the system clock of the sensing component can be obtained.
[0023] In some implementations, the frame synchronization unit is configured as follows: Based on the system timestamp, the data is timestamped and interpolated to determine the data collected by each sensing component at the same physical moment as synchronization frame data.
[0024] Sixthly, some embodiments of this specification provide a robot control device, which includes a control unit configured to: The fused perception data of multiple sensing components of the robot is acquired, and the fused perception data is generated by the data processing method of any of the aforementioned embodiments; The location information of the target object is determined based on the fused sensing data, and control commands are generated based on the location information. Human-machine interaction control is performed based on the control commands.
[0025] In some implementations, the control unit is configured to: The azimuth and / or pitch angle of the head assembly are controlled based on the control command, so that the front side of the head assembly faces the target object; Based on the control command, the microphone array is controlled to perform beamforming, so that the acquisition beam of the microphone array targets the target object; The robot's body is moved according to the control commands, causing the robot to move toward the location of the target object.
[0026] The head assembly described in this specification includes a housing and a sensing component disposed within the housing. The sensing component includes a camera, a microphone array, and a positioning sensor. The microphone array comprises multiple microphones spatially distributed and non-centrally symmetrically arranged. The main control unit performs fusion processing based on data from each sensing component and executes robot control based on the fused data. The head assembly utilizes multi-sensor fusion to improve the robot's environmental perception capabilities. Furthermore, the spatially distributed microphone array allows for the perception of sound wave propagation path differences in the vertical direction using microphones at different heights. This enables accurate localization and identification of sound sources at different heights. Simultaneously, the microphones located at the lower level can accurately pick up sound source data from below. Combined with beamforming algorithms, the received sound beam can be directionally focused below the robot, effectively improving voice wake-up and recognition success rates in scenarios where sound sources are located at lower levels. In addition, the non-centrally symmetrical arrangement of microphones on the same horizontal plane effectively eliminates or mitigates problems such as positioning ambiguity caused by mirrored sound sources, improving sound source localization accuracy and thus enhancing the robot's perception accuracy. Attached Figure Description
[0027] Figure 1 This is a schematic diagram of the robot's structure in some embodiments of this specification.
[0028] Figure 2 This is a schematic diagram of the head assembly in some embodiments of this specification.
[0029] Figure 3 This is a structural block diagram of the head component in some embodiments of this specification.
[0030] Figure 4 This is a schematic diagram illustrating the principle of sound source localization in some embodiments of this specification.
[0031] Figure 5 This is a schematic diagram of the microphone array arrangement in some embodiments of this specification.
[0032] Figure 6 This is an internal structural diagram of the head assembly in some embodiments of this specification.
[0033] Figure 7 This is an exploded view of the head assembly in some embodiments of this specification.
[0034] Figure 8 This is a schematic diagram of the head assembly in some embodiments of this specification.
[0035] Figure 9 This is a flowchart of a data processing method in some embodiments of this specification.
[0036] Figure 10 This is a flowchart of a data processing method in some embodiments of this specification.
[0037] Figure 11 This is a flowchart of a data processing method in some embodiments of this specification.
[0038] Figure 12 This is a flowchart of a data processing method in some embodiments of this specification. Detailed Implementation
[0039] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0040] Embossed robots are intelligent robots with physical entities that can perceive and interact with their environment in real time. Their forms can include humanoid robots, wheeled robots, etc. Currently, embossed robots are widely used in service, industrial, and other fields. The perception system is the core of an embossed robot's intelligent interaction. Taking a humanoid robot as an example, the robot's head system is similar to the human brain, undertaking core functions such as visual perception, voice interaction, and data processing. The accuracy of the perception system directly determines the effectiveness of subsequent task execution.
[0041] Robot head perception systems often incorporate multiple heterogeneous sensors, such as cameras, microphone arrays, LiDAR, positioning modules, and tactile sensors. Multi-sensor fusion algorithms are used to fuse multimodal heterogeneous data, driving the robot's interactive behavior. Taking multi-sensor fusion for target object localization as an example, the robot can collect environmental data through modules such as cameras, microphone arrays, and positioning sensors, and achieve high-precision object localization based on this data. For instance, images captured by the camera determine the visual orientation of the target object, audio data collected by the microphone array determines the sound source orientation, and the positioning sensor actively detects the relative distance to the target object. These heterogeneous data are then fused and verified simultaneously to construct a high-precision spatial position perception system.
[0042] Microphone systems are primarily used to perceive user speech and ambient sound, locating the direction and angle of sound sources. In related technologies, robot microphone systems suffer from weak beamforming capabilities, resulting in poor speech recognition performance in far-field scenarios (e.g., sound source greater than 2 meters from the robot) or high-noise scenarios (e.g., ambient noise greater than 65dB). For example, voice wake-up and recognition rates are below 70%, and the sound source localization direction is ambiguous, with positioning angle errors exceeding ±15°. For instance, if a user speaks while standing 30° to the left front of the robot, the robot's head will be facing directly forward in response, leading to a poor human-computer interaction experience.
[0043] Based on this, the embodiments of this specification provide a robot head assembly, a robot, a data processing method, and a robot control method. The robot provided in this specification can be any suitable type of robot, including but not limited to: humanoid robots, wheeled robots, bipedal robots, quadrupedal robots, etc., and this specification does not impose any limitations on this.
[0044] For example Figure 1 Image (a) shows the form of a bipedal humanoid robot. Figure 1 Figure (b) illustrates the form of a humanoid wheeled robot. The robot includes a head assembly and a body, which includes the robot's torso and other parts, such as legs, arms, hands, and feet, which will not be described in detail in this specification.
[0045] Figure 2This specification shows schematic diagrams illustrating the external structure of the head assembly in some embodiments, such as... Figure 2 As shown, the head assembly 100 is connected to the robot's body via the neck assembly 200, which provides the head assembly with degrees of freedom in three dimensions. For example, the head assembly 100 can generate pitch freedom (forward and backward tilting around the x-axis), lateral freedom (left and right swaying around the y-axis), and rotational freedom (horizontal rotation around the z-axis).
[0046] In this specification, for ease of description and understanding below, the following is used: Figure 2 Taking the oxyz coordinate system shown as an example, and combining it with the definition of the orientation of the human head, the following directions are defined for the head assembly 100: the direction parallel to the y-axis is the front-back direction of the head assembly 100, where +y is front and -y is back; the direction parallel to the x-axis is the left-right direction of the head assembly 100, where +x is left and -x is right; and the direction parallel to the z-axis is the up-down direction of the head assembly 100, where +z is up and -z is down. Unless otherwise specified, these defined directions will be used in the following descriptions of this specification.
[0047] In some embodiments, the head assembly 100 exemplified in this specification includes a housing, which serves as the outer shell structure of the head assembly. At least one receiving cavity can be formed inside the housing, which is a cavity structure that can accommodate the relevant circuit structure and electrical components of the head assembly 100.
[0048] In one example, see Figure 2 As shown, the head assembly 100's housing includes a front shell 111 and a rear shell 112. The front shell 111 covers the front of the head assembly 100, and the rear shell 112 covers the top, sides, and back of the head, thus the front shell 111 and the rear shell 112 combine to form an outer shell structure that covers the entire robot head. Of course, those skilled in the art will understand that... Figure 2 The housing structure shown is only an example of the solution in this specification. In other embodiments, the housing of the head assembly 100 may take any other shape suitable for implementation, which will not be described in detail here.
[0049] The head assembly 100 houses various types of sensing components and a mainboard. Sensing components refer to the front-end sensors in the head assembly 100 used to sense the external environment and collect raw data; for example, in some embodiments, sensing components include cameras, microphone arrays, and positioning sensors. The mainboard refers to the main control circuit board in the head assembly 100 used to carry and integrate various circuits and electronic devices; for example, in one example, the mainboard may include a PCB (Printed Circuit Board) and / or an FPC (Flexible Printed Circuit).
[0050] In some implementations, the motherboard is equipped with a main control unit, which refers to the computing module in the head component 100 that is responsible for receiving and processing data from the multi-sensing components. Figure 3 The following diagram illustrates the structural block diagram of the head assembly 100 in some embodiments of this specification. Figure 3 Please provide an explanation.
[0051] like Figure 3 As shown, in some examples, the main control unit includes a processor 101 and a memory 102. The processor 101 can be any type of processor with one or more processing cores, which can perform single-threaded or multi-threaded operations, used to parse instructions to perform operations such as acquiring data, performing logical operations, and issuing processing results.
[0052] The memory 102 may include a non-volatile computer-readable storage medium, such as at least one disk storage device, flash memory device, etc., which has a program storage area for storing non-volatile software programs, non-volatile computer-executable programs, and modules, which can be called by the processor to cause the processor to execute one or more method steps. The memory 102 may also include a volatile random access storage medium, or a storage portion such as a hard disk, as a data storage area for storing the calculation results and data output by the processor 101.
[0053] The processor 101 and memory 102 establish a communicative connection with other electronic devices in the head assembly 100 via bus 103. Bus 103 can be any protocol bus suitable for realizing electronic communication, including but not limited to: I2C (Inter-Integrated Circuit), SPI (Serial Peripheral Interface), USB (Universal Serial Bus), MIPI (Mobile Industry Processor Interface), etc., which are not limited in this specification.
[0054] Camera 120 is used to acquire environmental image data. It is worth noting that camera 120 in this specification can be any type of camera, or any combination of multiple types of cameras. For example, in one example, camera 120 may include any one or more combinations of a monochrome camera, an RGB camera, a depth (RGB-D) camera, an infrared camera, a laser array, and a ToF (Time of Flight) sensor.
[0055] Camera 120 is used to collect visual information from the robot. Similar to human eyes, camera 120 is typically positioned in front of the robot's head assembly 100, for example... Figure 2 In the example, camera 120 is mounted on the front shell 111 and positioned above the front shell 111, thereby facilitating the acquisition of environmental images in front of the robot's movement.
[0056] Microphone array 150 is used to acquire audio data. In this embodiment, microphone array 150 includes multiple microphones, which are distributed on the head assembly 100. This causes differences in the arrival times of audio signals from the same sound source at each microphone. Combined with the TDOA (Time Difference of Arrival) algorithm, the sound source can be located. Furthermore, the audio signals acquired by the multiple microphones are subjected to beamforming and signal differential processing to enhance and reduce noise in the sound source signal, thereby improving the signal-to-noise ratio of the user's voice signal.
[0057] The positioning sensor 130 is used to collect positioning data between the target object and the head assembly. This positioning data includes, for example, orientation and / or distance. Its core function is to provide spatial awareness for the robot. In the embodiments described in this specification, the positioning sensor 130 can employ any functional module suitable for achieving positioning, including but not limited to UWB (UltraWideband) positioning modules, WiFi positioning modules, Bluetooth positioning modules, ultrasonic positioning modules, infrared positioning modules, etc., without limitation.
[0058] In some embodiments, the head assembly 100 may also include other sensing components 180, such as laser dot matrix, infrared sensors, etc., which are not limited in this specification.
[0059] It is worth noting that the human-computer interaction input between the user and the robot mainly relies on voice commands. Therefore, the accuracy of the audio data collected by the microphone array 150 directly affects the accuracy of the multi-sensor fusion data.
[0060] In related technologies, the microphone array of the robot head assembly 100 is typically arranged symmetrically at the same horizontal height, for example, combined with... Figure 2As shown, a microphone can be placed on each side of the robot's head, near the ears, to mimic the structure of human binaural ears and form a symmetrical dual-microphone array. However, the sound source localization of the microphone array depends on the time difference between the arrival of the sound source at different microphones. This symmetrically distributed microphone array has difficulty distinguishing sound sources in mirror-symmetrical directions. For example, when the sound source is in a mirror image position directly in front of and behind the robot, the robot may easily misjudge it as the same location.
[0061] In some related technical solutions, in order to solve the problem of confusion in the positioning of the front and rear mirror images, a microphone is usually added in front of the head component 100 to form a three-array layout in front and on the left and right sides. Although this layout can alleviate the front and rear mirror image problem, if the sound source is located in the left and right mirror image position, the positioning problem is still easy to occur.
[0062] More importantly, the microphone arrays in related technologies only consider the horizontal (oxy-plane) arrangement, ignoring the differences in audio signal propagation in the robot's vertical (z-axis) direction. For example... Figure 4 In the example, the height of sound source S1 is comparable to that of head assembly 100; for example, sound source S1 could be at the height of an adult standing. Conversely, the height of sound source S2 is much lower than that of the robot head assembly 100; for example, sound source S2 could be at the height of a child or adult squatting. Figure 4 In the example, the voice command 2 emitted by the sound source S2 has a large height difference with the horizontal plane where the microphone array in the head assembly 100 is located, resulting in poor directivity of the microphone array's sound beam and a lower recognition success rate for voice command 2 compared to voice command 1.
[0063] In the embodiments described in this specification, the microphone array includes multiple microphones that are spatially distributed not only horizontally but also vertically, forming a three-dimensional microphone array layout. This arrangement allows for sound propagation with differences in both horizontal distance and vertical path length, enabling better acquisition of audio data at different spatial heights and improving positioning accuracy. Furthermore, the multiple microphones on the same horizontal plane are arranged non-centrally symmetrically, ensuring that the paths of the sound source signal reaching each microphone are not identical, effectively mitigating the positioning ambiguity problem caused by mirrored sound sources.
[0064] In some embodiments, the microphone array of the head assembly 100 includes a front microphone, a rear microphone, and side microphones. The front microphone is primarily used to collect audio data from the front of the robot, and includes at least three microphones spaced apart on the front side of the internal receiving cavity of the head assembly 100. The rear microphone is primarily used to collect audio data from the rear of the robot, and includes at least one microphone located on the rear side of the internal receiving cavity of the head assembly 100. The side microphones are used to collect audio data from the left and right sides of the robot, and include at least two left-side microphones and at least two right-side microphones. The at least two left-side microphones are spaced apart vertically on the left side of the receiving cavity, and the at least two right-side microphones are spaced apart vertically on the right side of the receiving cavity.
[0065] Figure 5 This specification illustrates a schematic diagram of the spatial arrangement of the microphone array in the head assembly 100 according to an exemplary embodiment. The following is a description in conjunction with... Figure 5 Examples will be provided to illustrate this.
[0066] exist Figure 5 In the example, the oy direction is the front of the head assembly 100, and the ox direction is the left side of the head assembly 100. In this example, the microphone mic0 is positioned directly in front of the head assembly 100. The line connecting mic0 and the center point o' of its plane is defined as the initial position (angle 0°), and the counterclockwise rotation direction of the line connecting mic0 and the center point o' around the center point o' is defined as the positive direction.
[0067] exist Figure 5 In the example, the front-facing microphones include three microphones: mic0, mic2, and mic1. The angle between the line connecting mic2 and the center point o' and the initial position is θ, where θ ranges from 5° to 45°. For example, in one example, θ = 30°. The angle between the line connecting mic1 and the center point o' and the initial position is α, where θ ranges from -5° to -45°. For example, in one example, α = -30°.
[0068] The rear microphone includes mic7, wherein the line connecting mic7 and the center point o' forms an angle of 180° with the initial position, that is, mic7 is located directly behind the head assembly 100.
[0069] The left microphone consists of two microphones, mic3 and mic4. The line connecting mic3 and the center point o' forms a 90° angle with the initial position, meaning mic3 is positioned directly to the left of the head assembly 100. Microphone 4 is vertically aligned with mic3 and positioned below mic3. The right microphone consists of two microphones, mic5 and mic6. The line connecting mic5 and the center point o' forms a -90° angle with the initial position, meaning mic5 is positioned directly to the right of the head assembly 100. Microphone 6 is vertically aligned with mic5 and positioned below mic5.
[0070] Figure 6 It shows Figure 5 The arrangement of the microphone array in the head assembly 100 is shown in the image below for clarity. Figure 6 The housing of the head assembly 100 is concealed. Combined with... Figure 5 and Figure 6 As can be seen in this example, mic0~mic2, mic3, mic5, and mic7 are located on the same horizontal plane. That is, these microphones are at the same or nearly the same height in the vertical direction and are closer to the top of the robot head assembly 100. Meanwhile, mic4 and mic6 are located below this horizontal plane and are closer to the neck of the robot head assembly 100.
[0071] In some embodiments of this specification, each microphone of the microphone array is disposed inside the receiving cavity of the head assembly 100. For example, the microphones can be fitted inside the housing of the head assembly 100, and the fitting method includes, but is not limited to, clips, adhesives, etc. To improve the sound pickup effect of each microphone, a pickup hole can be provided on the housing corresponding to the position of the microphone.
[0072] The basic principle of the Time Difference of Arrival (TDOA) algorithm is to use the time difference between the arrival times of sound waves from the same sound source at different microphones to calculate the azimuth and / or distance of the sound source relative to the microphone array, thereby achieving sound source localization. After locating the sound source, the delay and gain of each microphone signal can be adjusted according to the sound source's orientation to form a directional pickup beam, enhancing the audio signal in the direction of the sound source and weakening the audio signal in other directions, thus achieving directional sound pickup.
[0073] In this embodiment, considering that user voice commands mostly originate from the front of the robot, a large number of microphones are placed on the front of the head assembly to improve the audio data acquisition effect from the front of the robot. Simultaneously, at least one microphone is placed directly behind the robot to increase the audio data gain from the rear, ensuring that user commands from the rear can be properly picked up. Furthermore, microphones spaced apart along the vertical direction are arranged on the left and right sides. This three-dimensional microphone array allows for the perception of audio data propagation path differences in the vertical direction using microphones at different heights, enabling accurate localization and identification of sound sources at different heights.
[0074] Furthermore, the microphone located on the neck can accurately pick up audio data from below the robot's head assembly, and combined with beamforming algorithms, it can focus the sound pickup direction below the head assembly, effectively improving voice wake-up and recognition success rates in scenarios where the sound source is low, such as with children. In addition, the microphones on the same horizontal plane are arranged in a non-centrally symmetrical manner, effectively eliminating or mitigating problems such as positioning ambiguity caused by mirrored sound sources, and improving the accuracy of sound source localization.
[0075] See Figure 6 As shown, the microphone array includes microphones that are distributed within the housing cavity, and each microphone can establish an electrical connection with the motherboard 300 via a flexible printed circuit board (FPC). For the front-facing microphone, which includes multiple microphones such as mic0~mic2, these microphones can be mounted on the same FPC, thereby improving circuit integration. Furthermore, in... Figure 6 In the example, the camera 120 is located below the front microphone FPC, and the camera 120 can also be connected to the motherboard 300 via the FPC.
[0076] It is worth noting that the above Figure 5 and Figure 6 The microphone array arrangement shown in the example is merely an exemplary approach to the solution described in this specification. In other implementations, the microphone array may also adopt other implementations and is not limited to the above example. This specification will not elaborate further on these aspects.
[0077] Combination Figure 3As shown in the example in this specification, the input data for the main control unit to perform multi-sensor fusion includes at least: image data acquired by camera 120, audio data acquired by microphone array 150, and positioning data acquired by positioning sensor 130. After the processor of the main control unit acquires this multi-source sensor data, it executes a corresponding algorithm based on each sensor data, such as determining the visual positioning of the target object based on image data, determining the sound source positioning of the target object based on audio data, and determining the spatial positioning of the target object based on positioning data. Then, the multi-sensor data is fused to obtain fused perception data, which is the precise positioning data of the target object obtained from multiple types of sensing components. Afterwards, the main control unit can perform human-computer interaction control based on the fused perception data. This will be explained later in this specification and will not be detailed here.
[0078] As described above, in this embodiment, the robot head assembly utilizes multi-sensor fusion to enhance the robot's environmental perception capabilities. Furthermore, through a three-dimensional spatially arranged microphone array, microphones at different heights can sense the difference in sound wave propagation path in the vertical direction, enabling accurate localization and identification of sound sources at different heights. Simultaneously, the microphones located at the bottom can accurately pick up sound source data from below, and combined with beamforming algorithms, the received sound beam can be directionally focused below the robot, effectively improving voice wake-up and recognition success rates in scenarios where sound sources are low. Additionally, the microphones on the same horizontal plane are arranged in a non-centrally symmetrical manner, effectively eliminating or mitigating problems such as positioning ambiguity caused by mirrored sound sources, thus improving sound source localization accuracy.
[0079] Through comparative testing, it was found that in the relevant technical solutions, the horizontally symmetrical microphone array used in the robot's head assembly had a speech recognition rate of less than 70% and a sound source positioning angle error of 15°~20° when the sound source distance was greater than 2 meters or the ambient noise was greater than 65dB. This manual, however, does not meet these requirements. Figure 5 and Figure 6 In the implementation method, within a 360° horizontal omnidirectional range and a distance of 5 meters from the sound source, the accuracy of voice wake-up and recognition reaches over 95%, and the false wake-up rate is less than 0.1%. When the distance from the sound source is 2 meters, the accuracy of sound source localization is less than 5°.
[0080] In some embodiments, the robot's head assembly 100 also includes a display unit, which includes a display screen 140 and an ambient light sensor 160. The display screen 140 is used to display a user interface and is located on the front side of the receiving cavity of the head assembly 100. The ambient light sensor 160 is used to detect ambient light intensity data and send the light intensity data to the main control unit. The main control unit adjusts the display parameters of the display screen according to the ambient light intensity data.
[0081] For example, in some implementations, combined with Figure 2 and Figure 6 As shown, the display screen 140 can be disposed inside the front shell 111 of the head assembly 100, and the front shell 111 is a light-transmitting portion at least at the location corresponding to the display screen 140. For example, in one example, the front shell 111 is made of a transparent or translucent material. In another example, the front shell 111 is made of a transparent or translucent material only at the location corresponding to the display screen 140, and the remaining locations are made of a non-transparent material. The material of the light-transmitting portion can be, for example, glass, acrylic, sapphire, transparent ceramic, etc., and this specification does not limit this.
[0082] It is worth noting that in the relevant technical solutions, the display screen 140 usually adopts a rigid display screen, which cannot automatically adjust the brightness according to the ambient light (poor visibility under strong light and glaring under dim light), and also lacks flexible display capabilities, and cannot simulate human eye focus, side gaze and other micro-expressions, resulting in rigid interaction, lack of user emotional resonance, and a significant reduction in long-term interaction willingness.
[0083] Therefore, in some embodiments of this specification, the display screen 140 adopts a flexible curved display screen, for example... Figure 6 In the example, the curvature of the display screen 140 is adapted to the front shell 111 of the head assembly 100. For example, in one example, the flexible display screen 140 can be a flexible OLED (Organic Light-Emitting Diode) screen, which has advantages over traditional LCD (Liquid Crystal Display) screens such as being ultra-thin and flexible, having fast response times, and a wide color gamut. In the embodiments described in this specification, the curvature of the display screen 140 can be used to mimic micro-expressions such as focusing and glancing in the human eye, thereby making human-computer interaction more dynamic.
[0084] In addition, in some embodiments, the ambient light sensor 160 can be integrated with the aforementioned front microphones (mic0~mic2) on the same flexible circuit board (FPC). This improves the integration of electrical components and, being close to the top of the robot's head, is less likely to be blocked, thus enabling accurate acquisition of changes in ambient light intensity.
[0085] In this specification, the ambient light sensor 160 collects ambient light intensity data, which represents the brightness parameter of the ambient light. The unit of the brightness parameter can be expressed in lux. In some embodiments, the mapping relationship between the display parameters of the display screen 140 and the ambient light intensity data can be pre-configured. For example, in one example, the mapping relationship between the display parameters and the ambient light intensity data is shown in Table 1 below: Table 1 Mapping Relationship Table In the example in Table 1, the ambient light intensity data is divided into three data intervals. Each data interval is pre-configured with corresponding screen display parameters, which are display brightness in nits. The ambient light sensor 160 collects the ambient light intensity data and sends it to the main control unit. The main control unit determines the corresponding display brightness by looking up the mapping relationship shown in Table 1 above, and then drives the screen display based on the display brightness, so that the display screen 140 adjusts its brightness according to the ambient light intensity.
[0086] It is worth noting that the display parameters of the display screen 140 are not limited to display brightness; for example, they may also include display color temperature, etc. This manual does not impose any restrictions on this.
[0087] As described above, in this embodiment, by detecting ambient light intensity and adaptively adjusting the display parameters, the brightness of the display screen can automatically adapt to the current ambient light intensity, avoiding problems such as poor visibility in strong light and glare in low light. Furthermore, the flexible curved display screen can mimic micro-expressions such as human eye focus and sideways glances, making human-computer interaction more dynamic.
[0088] Figure 7 An exploded view of the head assembly 100 in some embodiments of this specification is shown. See also Figure 7 As shown, in this example, the housing of the head assembly 100 includes a front shell 111, a rear shell 112, and a neck shell 113. The front shell 111 and the rear shell 112 are assembled and enclosed to form a cavity for carrying various electrical components, and the neck shell 113 can serve as the outer shell structure of the robot's neck.
[0089] A motherboard 300 is housed within the receiving cavity of the head assembly 100. The motherboard 300 houses the main control unit and its peripheral circuits, circuit function modules, etc. Various sensing components and electrical elements included in the head assembly 100 can be connected to the motherboard 300 via an FPC to achieve communication with the main control unit.
[0090] It is worth noting that in related technologies, most of the various sensing components and sensors in the robot head are directly soldered to the motherboard and do not support hot-swapping replacement. If a component fails, the entire head assembly needs to be replaced, resulting in high maintenance costs. Furthermore, it does not support the upgrading or replacement of individual components or the expansion of functions, making it difficult to flexibly adapt to different application scenarios.
[0091] In some embodiments of this specification, the various sensing components and peripheral sensors included in the head assembly 100 can all be detachably electrically connected to the motherboard 300 via an FPC. For example, in one example, the FPC of the sensing component can be hot-swapped via a BTB (Board-to-board) connector. A male BTB connector can be provided on the motherboard 300, and the FPC of the sensing component presses against the female BTB connector, thereby achieving a detachable connection through the mating of the male and female connectors.
[0092] Combination Figure 7 For example, the small board 301 integrates microphone arrays mic0~mic2 and an ambient light sensor 160, with the small board 301 connected to a BTB connector female via an FPC cable. Similarly, the camera 120, positioning sensor 130, display screen 140, and other sensing components 180 can all be detachably electrically connected to the main board 300 via FPC and BTB connectors. This way, when a component fails or needs to be replaced, only that component needs to be removed and replaced individually, without replacing the entire head assembly 100, greatly reducing maintenance costs and increasing the flexibility of the head assembly 100, allowing users to freely replace the required parts.
[0093] In some embodiments, the head component 100 further includes at least one expansion interface 303 for connecting to external sensing components. For example... Figure 7 and Figure 8 In the example, an expansion interface 303 is provided on the housing in the top region of the head assembly 100. The expansion interface 303 can be any interface type suitable for implementation, including but not limited to: USB protocol interface, Type-C protocol interface, MIPI CSI protocol interface, MicroUSB protocol interface, or any type of protocol interface used for communication of electronic devices in the future.
[0094] For example, in one exemplary scenario, a new LiDAR sensor is to be added to the head assembly 100 to sense a radar dot matrix image in front of the robot. In this case, an electrical connection can be established with an external LiDAR sensor through the expansion interface 303, and the raw data collected by the LiDAR sensor can then communicate with the main control unit on the motherboard 300 through the expansion interface 303.
[0095] As described above, in this embodiment, all sensing components and related sensors in the head assembly are pluggable to the motherboard via standardized interfaces. The modular design enables hot-swapping and rapid replacement of components, reducing maintenance costs and complexity. Furthermore, the modular equipment can be manufactured and tested independently, resulting in higher overall assembly efficiency. Additionally, the expansion interfaces enhance the hardware expandability of the robot head assembly, allowing for the on-site installation of various types of sensors and improving the flexibility of the equipment's functionality.
[0096] In some implementations, see Figure 7 As shown, the head assembly 100 also includes a wireless communication unit 302, which can establish wireless communication with the peer device based on a near-field communication protocol. For example, in some embodiments, the wireless communication unit 302 may be a Bluetooth communication unit, a WiFi communication unit, a LoRA module, etc., and this specification does not limit this.
[0097] The head assembly 100 may further include ear assemblies 310, which are disposed on the left and right sides of the head assembly 100, forming an ear structure similar to that of a human ear. In some embodiments, the ear assemblies 310 may include indicator lights, which indicate the robot's operating status through lighting effects and / or color changes.
[0098] For example, in one example, the indicator light of the ear assembly 310 may include multiple colors. When the robot is operating normally, the indicator light of the ear assembly 310 displays blue; when the robot malfunctions, the indicator light of the ear assembly 310 displays red; and when the robot's battery is low, the indicator light of the ear assembly 310 displays yellow. In another example, the indicator light of the ear assembly 310 may include multiple lighting effects. When the robot is operating normally, the indicator light of the ear assembly 310 remains constantly lit; when the robot malfunctions, the indicator light of the ear assembly 310 flashes at a high frequency; and when the robot's battery is low, the indicator light of the ear assembly 310 flashes at a low frequency. Of course, those skilled in the art will understand that the indicator light of the ear assembly 310 can indicate even more types of operating states through combinations of colors and lighting effects, which will not be elaborated upon in this specification.
[0099] It is worth noting that in a multi-sensor fusion system, clock synchronization is the benchmark for improving the accuracy of multi-sensor data fusion. Clock synchronization means that the data collected by multiple sensors should correspond to the same physical time in the real world.
[0100] It is understandable that in a robot's head system, different types of sensing components are typically driven by independent controllers. For example, each sensor is equipped with an independent crystal oscillator (clock), resulting in a lack of a unified time reference between different types of sensing components. To achieve clock synchronization of multi-source data, relevant technical solutions usually employ software synchronization. That is, the system time when the data arrives at the main control unit is used as the data timestamp, thereby unifying multiple data streams into a single system timestamp and aligning the data based on this unified system timestamp.
[0101] However, real-world testing revealed that this data fusion solution frequently suffers from a disconnect between the robot's visual, auditory, and spatial positioning data across time and space, resulting in a significant fragmentation of perception. For example, when the robot is dynamically tracking a user's location, there are substantial discrepancies between the target locations obtained from visual positioning, sound source localization, and spatial positioning. This can cause the robot's body to move towards the user, but its face to be turned in a different direction when speaking, and even lead to tracking failures and error messages.
[0102] Further research revealed that this issue arises because, in the multi-sensor fusion system architecture of robots, data transmission from multiple sensing components is affected by factors such as transmission paths, bus protocols, and data jitter, resulting in significant differences in data latency. Even data collected at the same physical moment from multiple sources arrives at the main control unit at different times. Consequently, the software timestamp configured by the main control unit for the multi-sensor data based on the system clock has an error compared to the actual data acquisition time. This leads to time errors in the multi-sensor data included in a set of synchronization frames during data alignment processing, resulting in poor accuracy of the fused data obtained based on the synchronization frame data, thus causing data distortion.
[0103] To address this issue, some embodiments of this specification collect the local hardware timestamps of each sensing component and combine them with timestamp mapping to map the hardware timestamps of multiple sensors to a unified system time, thereby eliminating or mitigating signal time errors and improving the accuracy of synchronization frame data.
[0104] In some implementations, each sensing component has a local clock (CLKn), which refers to an independent crystal oscillator integrated into each sensing component and serves as the timing source within the sensor. For example, taking a camera as an example, the camera's internal local clock records a timestamp at the start of exposure for each frame of image data; this timestamp is the camera's local hardware timestamp, representing the acquisition time of one frame of image data. As another example, taking a positioning module as an example, the positioning module's local clock records a timestamp at the acquisition time of each frame of positioning data; this timestamp is the positioning module's local hardware timestamp, representing the acquisition time of one frame of positioning data.
[0105] It's understandable that the hardware timestamps marked by the sensing components correspond to the actual physical time. However, because each sensing component timestamps independently, there are differences in clock speed, start time, and sampling frequency between different sensing components. Therefore, multi-sensor data cannot be directly timestamped. For example, in one scenario, suppose that at the same physical time T0, a camera, microphone array, and positioning sensor each collect one frame of data. The hardware timestamp corresponding to the image data collected by the camera might be 4000ms, the hardware timestamp corresponding to a frame of audio data collected by the microphone array might be 15450ms, and the hardware timestamp corresponding to a frame of positioning data collected by the positioning sensor might be 850ms. As you can see, although the actual collection time of the three sensors is the same physical time, the hardware timestamps marked are completely different due to the differences in their respective clock standards.
[0106] Therefore, in the embodiments of this specification, the clock mapping relationship between the local clock (CLKn) of each sensing component and the system clock (CLKs) can be configured in advance according to the correspondence between the system clock (CLKs) and the local clock of each sensing component. The system clock (CLKs) refers to the host standard clock specified by the main control unit. In the embodiments of this specification, the clock mapping relationship represents the correspondence between the local clock of a certain sensing component and the system clock.
[0107] In some implementations, the clock mapping relationship between the local clock of sensing component n and the system clock can be represented by a function relationship. For example, the clock mapping relationship can be represented as: CLKs=Yn(CLKn), where Yn() represents the mapping function of the nth sensing component.
[0108] In some exemplary embodiments, the mapping function between the local clock of the sensing component and the system clock can be obtained through linear fitting. For example, a set of local clock and system clock data, denoted as (CLKn, CLKs), can be collected at the same physical moment through data acquisition. Multiple sets of data are collected continuously, and then linear fitting is performed based on the multiple sets of data to obtain the mapping function relationship corresponding to the sensing component. This function mapping relationship is the clock mapping relationship described in this specification. For example, in one example, the clock mapping relationship of sensing component 1 is represented as: CLKs_1 = 0.8 × CLK1; the clock mapping relationship of sensing component 2 is represented as: CLKs_2 = 0.355 × CLK2; and so on.
[0109] In other implementations, the clock mapping relationship between the local clock of sensing component n and the system clock can be represented using a lookup table. For example, the clock mapping relationship can be shown in Table 2 below: Table 2 Clock Mapping Relationship Table In some exemplary embodiments, the clock mapping table of the sensing components can be obtained by continuous data acquisition, as those skilled in the art will understand, and will not be described in detail here.
[0110] The above describes the configuration process for the clock mapping relationship of only one sensing component. The clock mapping relationship for each sensing component can be configured through the above process and stored in the main control unit.
[0111] During the multi-sensor data fusion process, each sensing component continuously collects frame data at a fixed sampling frequency and marks each frame with a hardware timestamp according to its local clock. Next, the main control unit can use a pre-configured clock mapping relationship to uniformly map the hardware timestamps corresponding to the frame data of each sensing component to a system timestamp. The system timestamp is the time identifier corresponding to the system clock CLKs.
[0112] For example Figure 9 This specification illustrates a flowchart of the main control unit's data flow processing for the multi-channel sensing component in some embodiments. The following section, in conjunction with... Figure 9 Examples will be provided to illustrate this.
[0113] In some implementations, the main control unit includes a clock mapping unit, a frame synchronization unit, a data fusion unit, and a control unit. The clock mapping unit performs clock mapping on the raw data collected by the multi-channel sensing components, uniformly mapping the local clocks of the sensing components to the system clock. The frame synchronization unit synchronizes and aligns the frame data from the multi-channel sensing components, filtering out a set of synchronized frame data belonging to the same physical moment. The data fusion unit, based on a fusion algorithm, fuses the multi-channel sensing component data included in a set of synchronized frame data to obtain fused sensing data. The control unit generates control commands for the robot's head assembly and / or body based on the fused sensing data.
[0114] exist Figure 9 In this example, the multi-channel sensing components include the aforementioned camera, microphone array, and positioning sensor. Each sensing component continuously collects data at a predetermined sampling frequency and uses its own local clock to mark each frame of data with a first timestamp, which is the hardware timestamp marked by the local clock.
[0115] For example, taking a frame of image data C1 captured by the camera as an example, its corresponding first timestamp Tc1 is: the hardware timestamp Tc1 based on the camera's local clock. Similarly, taking a frame of positioning data U2 captured by the positioning sensor as an example, its corresponding first timestamp Tu2 is: the hardware timestamp Tu2 based on the positioning sensor's local clock.
[0116] As described above, the clock mapping unit can invoke the pre-configured clock mapping relationship of each sensing component to perform clock mapping on the first timestamp of each frame of data, thereby converting the timestamp information from the local clock of the sensing component to the system clock. For example, taking camera data as an example, the clock mapping unit can invoke the clock mapping relationship corresponding to the camera to convert the first timestamp corresponding to each frame of image data into a system timestamp. For example, taking image data C1 as an example, the first timestamp Tc1 of image data C1 is converted into the system timestamp Ts1 through clock mapping.
[0117] For multi-channel sensing components, the clock mapping unit sequentially executes the aforementioned clock mapping process, thereby uniformly converting the first timestamp of the data stream from each sensing component into a system timestamp identified by the system clock. It can be understood that after clock mapping, the timestamp information of each frame of data from the sensing component is based on a unified system clock standard. Therefore, data with the same system timestamp contained within a multi-channel sensing component are data collected at the same physical moment.
[0118] Based on this principle, the frame synchronization unit can filter out data with the same system timestamp from the multi-sensor component data and bind them into a group of synchronization frame data. For example, Figure 9 In the example, the data C1, M2, and U1 included in the synchronization frame data Vsync1 correspond to the data at the same physical moment. Similarly, the data C3, M3, and U2 included in the synchronization frame data Vsync2 correspond to the data at the same physical moment.
[0119] It is understandable that for multi-channel sensing components, the time information of each frame of data they collect is based on the time marked by the local clock, rather than the time of arrival at the main control unit. Therefore, it can effectively avoid the influence of factors such as signal transmission delay and data jitter. The mapped system timestamp is the time information corresponding to the real physical time, which eliminates or alleviates the delay error and improves the accuracy of data synchronization.
[0120] Continue to refer to Figure 9 As shown, after obtaining the synchronization frame data, the data fusion unit can perform fusion processing on the data from the multi-sensing components included in the synchronization frame data based on the fusion algorithm to obtain fused sensing data. It is worth noting that the data fusion algorithm can be any suitable algorithm, such as weighted fusion algorithm, Kalman filter algorithm, Bayesian estimation algorithm, etc., and this specification does not impose any restrictions on it. For example, in one example, for the synchronization frame data Vsync1, the fused sensing data V1 obtained after the data fusion unit performs fusion processing represents: the orientation and / or position information of the target object calculated based on the multi-sensing components (camera, microphone array, and positioning data) at the same physical moment.
[0121] The control unit can then generate continuous control commands for controlling the robot's behavior based on the fused sensing data. These control commands can be used to control any one or more components of the robot; for example, in some implementations, they can control the pose of the robot's head assembly and / or body. In an example scenario, taking automatic user tracking as an example, the control commands generated by the control unit can control the robot's head assembly to face the user's location, while simultaneously controlling the robot's body to move in the same direction.
[0122] As described above, in this embodiment, the local clock of the sensing component is used as the time reference for data acquisition. Furthermore, clock mapping is used to uniformly map the data from multiple sensing components to the system clock standard. This effectively eliminates or mitigates time synchronization errors caused by data transmission delays and jitter, improves synchronization data accuracy, and effectively correlates multiple data streams with the same system time to the same physical moment, achieving high-precision time alignment. Experimental verification shows that the synchronization scheme in this specification can control the synchronization error of multiple sensing components to within 2ms, meeting the accuracy requirements of various multi-sensor fusion algorithms.
[0123] In some implementations, considering that different sensing components have different sampling frequencies, the frame synchronization unit may not be able to find a set of synchronization frame data with the same system timestamp when performing time alignment processing.
[0124] For example, in one example, the system timestamp corresponding to camera data C1 is 800ms, the system timestamp corresponding to microphone data M2 is 799.5ms, and the system timestamp corresponding to positioning sensor data U1 is 800.5ms. For such a set of data, which are very close in timestamp but not completely consistent, in order to achieve synchronization alignment, some embodiments of this specification can combine timestamp matching and data interpolation algorithms to interpolate a set of data with close timestamps into a set of synchronization frame data.
[0125] For example, the frame synchronization unit can use a timestamp matching algorithm to identify data with a timestamp difference less than or equal to 1ms as candidate synchronization frame data. The set of data in the example above is a candidate synchronization frame data. Next, using 800ms as the time base, microphone data at timestamp 800ms can be obtained by interpolation based on several frames before and after microphone data M2. For example, combining the current 799.5ms microphone data M2 and the next frame's 800.5ms microphone data M3, a weighted average can be used to estimate the microphone data at 800ms. This estimation process is called data interpolation. Similarly, for positioning sensor data U1, the positioning sensor data at 800ms can be estimated by data interpolation based on the preceding 799.5ms positioning sensor data U0. After aligning each data stream to 800ms, a set of synchronization frame data corresponding to the system timestamp 800ms can be obtained.
[0126] As can be seen from the above, in the implementation of this specification, during the data synchronization process of the multi-channel sensing components, timestamp matching and data interpolation are used to align the multi-channel sensing data to the same physical moment, further improving the alignment accuracy of the synchronization frame data and increasing the number of synchronization frame data, thereby ensuring the accuracy of the fused data.
[0127] In some implementations, the main control unit is configured to receive time-stamped data collected by the sensing components, align the data based on the system clock to obtain synchronization frame data, and perform multi-sensor fusion based on the synchronization frame data to obtain the fused sensing data, wherein the synchronization frame data represents the data collected by each sensing component at the same physical moment.
[0128] In combination with the above Figure 9 For example, the time stamp of the data collected by the sensing components is the hardware timestamp marked by the local clock of the sensing components, which is also the aforementioned first timestamp. The clock mapping unit of the main control unit maps the first timestamp of the data collected by the multiple sensing components to a unified system timestamp. The frame synchronization unit performs time alignment processing on the multiple sensing data according to the unified timestamp standard to obtain synchronized frame data. The data fusion unit performs data fusion on the synchronized frame data based on the fusion algorithm to obtain fused sensing data. The control unit generates control commands for controlling at least one device of the robot based on the fused sensing data.
[0129] As described above, in this embodiment, the local clock of the sensing component is used as the time reference for data acquisition. Furthermore, clock mapping is used to uniformly map the data from multiple sensing components to the system clock standard. This effectively eliminates or mitigates time synchronization errors caused by data transmission delays and jitter, improves the accuracy of synchronized data, and effectively associates multiple data streams with the same system time to the same physical moment, achieving high-precision time alignment. During the data synchronization process of multiple sensing components, timestamp matching and data interpolation are used to align the multiple sensing data streams to the same physical moment, further improving the alignment accuracy of synchronization frame data and increasing the number of synchronization frame data streams, thereby ensuring the accuracy of the fused data.
[0130] In some embodiments, this specification provides a data processing method that can be applied to the header component 100 of any of the foregoing embodiments and executed by the main control unit of the header component 100, for example, by the processor 101 of the main control unit.
[0131] like Figure 10 As shown, in some embodiments, the data processing method exemplified in this specification includes: S101. Obtain the data collected by the sensing component.
[0132] Combination Figure 9 As shown, the multi-channel sensing component takes the aforementioned camera, microphone array, and positioning sensor as examples. Each sensing component continuously collects data at a predetermined sampling frequency and uses its own local clock to mark each frame of data with a first timestamp, which is the hardware timestamp marked by the local clock.
[0133] S102, The main control unit maps the first timestamp of the sensing component to the system timestamp based on the pre-set clock mapping relationship.
[0134] In some implementations, the main control unit may be configured with a clock mapping relationship for each sensing component in advance. The clock mapping relationship represents the correspondence between the local clock of the sensing component and the system clock. Figure 11 This specification illustrates the process by which the main control unit configures the clock mapping relationship in some embodiments, including: S111. For any sensing component, obtain the offset between the local clock and the system clock of the sensing component, and collect multiple sets of data pairs between the local clock and the system clock based on the offset.
[0135] S112. Based on multiple sets of data pairs, perform linear fitting or data interpolation to obtain the clock mapping relationship between the local clock system clock of the sensing component.
[0136] As can be seen from the foregoing, the local clock of the sensing component refers to the crystal oscillator inside the sensing component itself, while the system clock refers to the host standard clock specified by the main control unit. The two are not at the same frequency, but have a certain time offset.
[0137] In some implementations, the offset between the local clock of the sensing component and the system clock can be determined by pre-calibration or data acquisition. This offset is generally a fixed value. By acquiring the local clock of the sensing component and adding the offset, the corresponding system clock can be obtained, thus acquiring a set of data pairs. Repeating this acquisition process can yield multiple sets of data pairs.
[0138] Then, by combining the linear fitting or data lookup table construction methods described in the foregoing implementation examples, the clock mapping relationship between the local clock and the system clock can be obtained. Those skilled in the art will understand this from the foregoing, and it will not be elaborated further in this specification. After configuring the clock mapping relationship corresponding to each sensing component, these clock mapping relationships can be stored in the main control unit.
[0139] Combination Figure 9 As shown, for multi-channel sensing component data, the main control unit can call the corresponding clock mapping relationship to uniformly map the first timestamp of multiple data channels to the system timestamp. For example, taking camera data, the clock mapping unit can call the clock mapping relationship corresponding to the camera to convert the first timestamp of each frame of image data into the system timestamp. For example, taking image data C1 as an example, the first timestamp Tc1 of image data C1 is converted into the system timestamp Ts1 through clock mapping, and so on.
[0140] S103. Align the data according to the system timestamp to obtain synchronization frame data.
[0141] It is understandable that after clock mapping, the timestamp information of each frame of data from the sensing component is the time under a unified system clock standard. Therefore, data with the same system timestamp contained in multiple sensing components are data collected at the same physical moment.
[0142] Then, the main control unit can filter out data with the same system timestamp from the multi-sensor component data and bind them into a group of synchronization frame data. For example Figure 9 In the example, the data C1, M2, and U1 included in the synchronization frame data Vsync1 correspond to the data at the same physical moment. Similarly, the data C3, M3, and U2 included in the synchronization frame data Vsync2 correspond to the data at the same physical moment.
[0143] S104. Perform multi-sensor fusion based on the synchronization frame data to obtain fused perception data.
[0144] After obtaining the synchronization frame data, the data fusion unit can perform fusion processing on the data from the multi-sensing components included in the synchronization frame data based on the fusion algorithm to obtain fused sensing data. It is worth noting that the data fusion algorithm can be any suitable algorithm, such as weighted fusion algorithm, Kalman filter algorithm, Bayesian estimation algorithm, etc., and this specification does not impose any restrictions on it. In the examples in this specification, fused sensing data represents the orientation and / or position information of the target object calculated based on the multi-sensing components at the same physical moment.
[0145] For example, in some implementations, the multi-channel sensing component takes the aforementioned camera, microphone array, and positioning sensor as examples. First, the main control unit can preprocess the multi-channel data included in the synchronization frame data. For example, it can filter and denoise the image data acquired by the camera, perform bandpass filtering on the audio data acquired by the microphone array, and perform moving average filtering on the positioning data acquired by the positioning sensor to improve data accuracy.
[0146] Next, the main control unit can use a normalization algorithm to normalize the image data, audio data, and positioning data, quantizing the multiple data streams into data within the same value range. The normalization algorithm can be, for example, min-max normalization or Z-score, and there are no restrictions on which one to use.
[0147] Then, weight values can be pre-configured for the data of each sensing component. The weight value represents the credibility of the sensing component; the higher the weight value, the higher the credibility of the corresponding sensing component. The sum of the weight values of the multi-sensing components used for fusion is 1. Finally, the data result obtained by weighted summation of the multi-sensing component data and their corresponding weight values is the fused sensing data.
[0148] Of course, those skilled in the art will understand that data fusion algorithms are not limited to the weighted fusion algorithms in the examples above, and this specification will not elaborate further on this.
[0149] As can be seen from the above, in the embodiments of this specification, the local clock of the sensing component is used as the time reference for data acquisition, and the data of multiple sensing components are uniformly mapped to the system clock standard by combining clock mapping. This can effectively eliminate or alleviate the time synchronization error caused by data transmission delay and data jitter, improve the accuracy of synchronized data, effectively associate multiple data of the same system time with the same physical time, and achieve high-precision time alignment.
[0150] In some embodiments, this specification provides a robot control method that can be applied to a robot. The robot exemplified in this specification can be any suitable robot type, such as... Figure 1 The bipedal humanoid robot or wheeled robot shown in this specification are not limited to these types of robots.
[0151] like Figure 12 As shown, in some embodiments, the control method exemplified in this specification includes: S121. Obtain fused perception data from multiple sensing components of the robot.
[0152] S122. Determine the location information of the target object based on the fused sensing data, and generate control commands based on the location information.
[0153] S123. Human-machine interaction control is executed based on control commands.
[0154] As can be seen from the foregoing, the fused perception data is the fused perception data calculated by the main control unit of the robot head assembly 100 through the methods and processes of any of the aforementioned implementation methods, and this specification will not elaborate further on this.
[0155] It is understandable that fused sensing data reflects the orientation and / or position information of a target object calculated based on the multi-path sensing components at the same physical moment. Therefore, after determining the fused sensing data of the multi-path sensing components, the position information of the target object can be determined based on this fused sensing data. The position information may include orientation, location, and other information.
[0156] Then, control commands for controlling the robot can be generated based on the position information of the target object. In some implementations, the control commands are used to control the pose of the head assembly and / or the robot body.
[0157] For example, in one instance, human-machine interaction control based on control commands includes controlling the azimuth and / or pitch angle of the robot's head assembly 100 based on control commands, so that the front of the head assembly faces the target object.
[0158] In this example, control commands are used to control the movement of the robot's head assembly, in conjunction with the aforementioned... Figure 2 As shown, the head assembly's motion includes multiple degrees of freedom, such as pitch freedom around the x-axis, yaw freedom around the y-axis, and horizontal rotation freedom around the z-axis. The horizontal rotation angle around the z-axis is the azimuth angle, the forward / backward rotation angle around the x-axis is the pitch angle, and the left / right yaw angle around the y-axis is the roll angle.
[0159] In some embodiments of this specification, control commands can control the movement of one or more degrees of freedom of the head assembly, thereby causing the front of the head assembly to face the target object, thus allowing the user to experience a human-computer interaction similar to face-to-face conversation.
[0160] For example, in another example, human-computer interaction control based on control commands includes: controlling the microphone array to perform beamforming based on control commands, so that the acquisition beam of the microphone array executes the target object.
[0161] In this example, the microphone array can focus the pickup beam of the microphone array on the direction of the target object based on the location information of the target object, combined with a beamforming algorithm, thereby improving the gain of the user's voice signal, reducing environmental noise, and improving the signal-to-noise ratio and speech recognition accuracy.
[0162] It is worth noting that those skilled in the art can understand and fully implement the specific process of the beamforming algorithm by referring to the principles described above in this specification, and this specification will not elaborate further on it.
[0163] For example, in another instance, human-machine interaction control based on control commands includes: controlling the robot's body movement based on control commands, so that the robot moves toward the location of the target object.
[0164] For example, in an automatic following scenario, control commands are used to control the robot to move to follow the position of the target object. After determining the position information of the target object, the robot's motion system (such as a wheeled drive system, a bipedal motion system, etc.) can be controlled to move toward the position of the target object, thereby achieving automatic following.
[0165] Of course, those skilled in the art will understand that the human-computer interaction control of a robot is not limited to the above-mentioned example scenarios, and may include many other forms of human-computer interaction control, which this specification cannot exhaustively list.
[0166] As can be seen from the above, in the embodiments of this specification, by acquiring the fused perception data obtained by the fusion of multiple sensing components, the position information of the target object is accurately determined and corresponding control commands are generated, which can realize multi-dimensional human-computer interaction control, such as controlling the robot's head or body to follow the user, controlling the microphone array to pick up sound in a directional manner, etc., effectively improving the intelligence of human-computer interaction and the user interaction experience.
[0167] In some embodiments, this specification provides a robot including a body and a head assembly as described in any of the above embodiments. The robot provided in this specification can be any suitable type of robot, including but not limited to: humanoid robots, wheeled robots, bipedal robots, quadrupedal robots, etc., and this specification does not impose any limitations on this.
[0168] As described above, in this embodiment, the local clock of the sensing component is used as the time reference for data acquisition. Furthermore, clock mapping is used to uniformly map the data from multiple sensing components to the system clock standard. This effectively eliminates or mitigates time synchronization errors caused by data transmission delays and jitter, improves the accuracy of synchronized data, and effectively associates multiple data streams with the same system time to the same physical moment, achieving high-precision time alignment. During the data synchronization process of multiple sensing components, timestamp matching and data interpolation are used to align the multiple sensing data streams to the same physical moment, further improving the alignment accuracy of synchronization frame data and increasing the number of synchronization frame data streams, thereby ensuring the accuracy of the fused data.
[0169] In some embodiments, this specification provides a data processing apparatus applied to the robot head assembly of any of the above embodiments. For example... Figure 9 As shown, the data processing device includes: The clock mapping unit is configured to acquire data collected by the sensing component. The data is configured with a first timestamp, which is configured by the sensing component based on its own local clock. Based on a pre-set clock mapping relationship, the first timestamp of the sensing component is mapped to the system timestamp. The clock mapping relationship represents the correspondence between the system clock and the local clock of the sensing component. The frame synchronization unit is configured to align the data according to the system timestamp to obtain synchronization frame data, wherein the synchronization frame data represents the data collected by each sensing component at the same physical moment. The data fusion unit is configured to perform multi-sensor fusion based on the synchronization frame data to obtain fused sensing data.
[0170] In some implementations, the clock mapping unit is configured as follows: For any sensing component, obtain the offset between the local clock and the system clock of the sensing component, and collect multiple sets of data pairs between the local clock and the system clock based on the offset; By performing linear fitting or constructing a data lookup table based on multiple sets of data pairs, the clock mapping relationship between the local clock and the system clock of the sensing component can be obtained.
[0171] In some implementations, the frame synchronization unit is configured as follows: Based on the system timestamp, the data is timestamped and interpolated to determine the data collected by each sensing component at the same physical moment as synchronization frame data.
[0172] In some embodiments, this specification provides a robot control device, with reference to Figure 9 As shown, the control device includes a control unit, which is configured to: The fused perception data of multiple sensing components of the robot is acquired, and the fused perception data is generated by the data processing method of any of the aforementioned embodiments; The location information of the target object is determined based on the fused sensing data, and control commands are generated based on the location information. Human-machine interaction control is performed based on the control commands.
[0173] In some implementations, the control unit is configured to: The azimuth and / or pitch angle of the head assembly are controlled based on the control command, so that the front side of the head assembly faces the target object; Based on the control command, the microphone array is controlled to perform beamforming, so that the acquisition beam of the microphone array targets the target object; The robot's body is moved according to the control commands, causing the robot to move toward the location of the target object.
[0174] Obviously, the above embodiments are merely examples for clear illustration and are not intended to limit the embodiments. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all embodiments here. However, obvious variations or modifications derived therefrom remain within the scope of protection created by this specification.
Claims
1. A robot head assembly, comprising: The housing has at least one receiving cavity; The motherboard and various types of sensing components are disposed in the cavity, including a microphone array for acquiring audio data; The microphone array includes multiple microphones, which are spaced apart in the horizontal and vertical directions of the receiving cavity, and at least some of the microphones are non-centrally symmetrically distributed in the horizontal plane; The main control unit on the motherboard is configured to perform multi-sensor fusion of the sensing data collected by the various types of sensing components to obtain fused sensing data, and to perform robot control based on the fused sensing data.
2. The head assembly according to claim 1, wherein the microphone array comprises a front microphone, a rear microphone, and side microphones; The front microphone includes at least three microphones, which are spaced apart on the front side of the receiving cavity; The side-mounted microphone includes at least two left-side microphones and at least two right-side microphones. The at least two left-side microphones are vertically spaced on the left side of the receiving cavity, and the at least two right-side microphones are vertically spaced on the right side of the receiving cavity. The rear microphone includes at least one microphone, which is located on the rear side of the receiving cavity.
3. The head assembly according to claim 2, characterized in that, The front microphone, the rear microphone, at least one left microphone, and at least one right microphone are positioned close to the top of the head assembly, while the remaining left microphone and right microphone are positioned close to the neck of the head assembly.
4. The head assembly according to claim 1 further includes a display unit, the display unit including a curved display screen and an ambient light sensor; The display screen is located on the front side of the receiving cavity; The ambient light sensor is mounted on a flexible circuit board that carries at least part of the front microphone, and is detachably electrically connected to the motherboard through the flexible circuit board. The main control unit is configured to adjust the display parameters of the display screen based on the ambient light intensity data collected by the ambient light sensor.
5. The head assembly according to claim 1, wherein the sensing component is disposed on a flexible circuit board and is detachably electrically connected to the motherboard through the flexible circuit board.
6. The head assembly according to claim 1, wherein the housing is provided with at least one expansion interface for pluggable connection with at least one external sensing component; The main control unit is configured to identify the type of the external sensing component and load the corresponding driver configuration when the external sensing component is detected to be connected.
7. The head assembly according to any one of claims 1 to 6, wherein the main control unit is configured to receive time-stamped data collected by the various types of sensing components, align the data based on the system clock to obtain synchronization frame data, and perform multi-sensor fusion based on the synchronization frame data to obtain the fused sensing data, wherein, The synchronization frame data represents the data collected by each sensing component at the same physical moment.
8. The head component according to claim 7, wherein the sensing component is configured to configure a first timestamp for the collected data based on its own local clock; The main control unit is configured to map the first timestamp of the sensing component to the system timestamp based on a pre-set clock mapping relationship, wherein... The clock mapping relationship represents the correspondence between the system clock and the local clock of the sensing component. The main control unit performs alignment processing on the data according to the system timestamp to obtain synchronization frame data.
9. The head assembly according to claim 1, wherein the main control unit is configured to determine the position information of the target object based on the fused perception data, and generate control instructions based on the position information, the control instructions being used to control the pose of the head assembly and / or the robot body.
10. The head assembly of claim 1, wherein the plurality of sensing components further comprises a camera and a positioning sensor; in, The camera is used to acquire image data, and the camera includes at least one of the following: an RGB camera and a depth camera; The positioning sensor is used to collect positioning data of the target object relative to the head component, and the positioning sensor includes at least one of the following: an ultra-wideband positioning module, a WiFi positioning module, and a Bluetooth positioning module.
11. A data processing method, applied to the main control unit of a robot head assembly as described in any one of claims 1 to 10, the method comprising: Data collected by various types of sensing components is acquired. The data collected by the sensing components includes a first timestamp, which is configured by the sensing component based on its own local clock. Based on a pre-set clock mapping relationship, the first timestamp of the sensing component is mapped to the system timestamp, wherein the clock mapping relationship represents the correspondence between the system clock and the local clock of the sensing component; The data is aligned according to the system timestamp to obtain synchronization frame data, which represents the data collected by each sensing component at the same physical moment. Multi-sensor fusion is performed based on the synchronous frame data to obtain fused sensing data.
12. The method according to claim 11, wherein the process of pre-setting the clock mapping relationship includes: For any sensing component, obtain the offset between the local clock and the system clock of the sensing component, and collect multiple sets of data pairs between the local clock and the system clock based on the offset; By performing linear fitting or constructing a data lookup table based on multiple sets of data pairs, the clock mapping relationship between the local clock and the system clock of the sensing component can be obtained.
13. The method according to claim 11, wherein aligning the data according to the system timestamp to obtain synchronization frame data, comprising: Based on the system timestamp, the data is timestamped and interpolated to determine the data collected by each sensing component at the same physical moment as synchronization frame data.
14. A robot control method, comprising: Acquire fused perception data from multiple sensing components of a robot, wherein the fused perception data is generated by the method described in any one of claims 11 to 13; The location information of the target object is determined based on the fused sensing data, and control commands are generated based on the location information. Robot control is executed based on the control commands.
15. The control method according to claim 14, characterized in that, Executing robot control based on the control commands includes at least one of the following: The azimuth and / or pitch angle of the head assembly are controlled based on the control commands, so that the front side of the head assembly faces the target object. Based on the control command, the microphone array is controlled to perform beamforming, so that the acquisition beam of the microphone array is pointed at the target object; The robot's body is moved according to the control commands, causing the robot to move toward the location of the target object.
16. A robot comprising: body; as well as The head assembly according to any one of claims 1 to 10 is disposed on the body.