Bionic service type humanoid robot head structure

By integrating multimodal sensors and high-performance processors into the head structure of a humanoid robot, the problem of humanoid robots being unable to provide timely feedback on changes in the external environment and recognize the taste of food in existing technologies has been solved, achieving more efficient environmental perception and interaction capabilities.

CN118952165BActive Publication Date: 2025-12-09SANMEN TONGSHUN RIVET
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411005219.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2025-12-09
Estimated Expiration
2044-07-25

AI Technical Summary

Technical Problem

Existing humanoid robots cannot respond promptly to changes in the external environment, nor can they automatically identify the texture of food, and their interactive capabilities are limited.

Method used

A biomimetic service humanoid robot head structure was designed, which includes a sensory system, a vision system, a nasal system, and an ear system. It adopts multimodal sensors and high-performance processors (such as NVIDIA's Project GR00T and Isaac robot platform), combined with image processing algorithms and machine learning models, to achieve real-time perception and feedback of the external environment.

Benefits of technology

It enhances the robot's ability to interact with the outside world, enabling it to test and evaluate the taste of food without human testers, and achieves timely and effective feedback to the external environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118952165B_ABST
    Figure CN118952165B_ABST
Patent Text Reader

Abstract

The application discloses a bionic service type humanoid robot head structure, which comprises a head structure, a sensory system, a mouth system connected with the sensory system, a visual system, a nose system and an ear system, the sensory system comprises a perception center, a sensor network, a signal processing unit and an execution unit, the sensor network is connected with the signal processing unit, the signal processing unit is connected with the perception center, and the perception center is connected with the execution unit. The sensory system formed by the perception center, the sensor network, the signal processing unit and the execution unit can make timely and effective feedback to the changes of the external environment, and the interaction ability with the external environment is greatly improved. Meanwhile, the humanoid robot can test and evaluate food without human testers according to the taste sensor. The taste perception system can simulate the function of the human taste perception system, so that the taste of food can be automatically evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of robot bionics, and relates to a head structure of a bionic service type humanoid robot. BACKGROUND

[0002] A "humanoid robot" can be defined as a robot that has certain attributes of the appearance and functions of a human (such as a torso, a head, arms, legs), the ability to communicate verbally with humans using speech recognition and voice synthesis, etc. This kind of robot aims to reduce the cognitive distance between humans and machines.

[0003] The existing "humanoid robot" has limited interaction ability with the outside world and cannot make timely feedback according to the changes in the outside environment. In addition, the humanoid robot cannot complete the corresponding work in places where food texture needs to be automatically identified. SUMMARY

[0004] The present application is to overcome at least one deficiency of the prior art, and provides a head structure of a bionic service type humanoid robot.

[0005] In order to achieve the above-mentioned purpose, the present application adopts the following technical solution: a head structure of a bionic service type humanoid robot, comprising a skull structure, a sensory system, a mouth system connected with the sensory system, a vision system, a nose system and an ear system, the sensory system comprising a perception center, a sensor network, a signal processing unit and an execution unit, the sensor network being connected with the signal processing unit, the signal processing unit being connected with the perception center, and the perception center being connected with the execution unit.

[0006] The sensor network comprises a plurality of sensors for capturing information of the external environment.

[0007] The signal processing unit is used for receiving raw signals from the sensor network, performing signal processing and preprocessing, and transmitting the processed signals to the perception center.

[0008] The perception center receives the processed signals, analyzes and makes decisions, and obtains the state of the external environment and the required response.

[0009] The execution unit receives the instructions of the perception center and executes corresponding actions or tasks.

[0010] Further, the sensor network comprises a vision sensor, the signal processing unit comprises a central processing unit and a graphics processing unit, the graphics processing unit and the vision sensor constitute a visual tracking system, the vision sensor captures visual information of the surrounding environment and transmits it to the graphics processing unit for analysis, identification and tracking, the graphics processing unit has an image processing algorithm built-in, the vision sensor captures continuous image frames, and the image processing algorithm is used to identify and extract the required images or objects, specifically comprising the following steps:

[0011] Step S1: The visual sensor captures real-time images or videos in the environment and transmits them to the graphics processing unit;

[0012] Step S2: The image processing algorithm pre-processes the captured images;

[0013] Step S3: Feature extraction: specific information is extracted from the pre-processed images;

[0014] The feature extraction method includes edge detection based on edge detection algorithm, corner detection based on corner detection algorithm, and texture analysis based on texture analysis algorithm;

[0015] Step S4: Object recognition;

[0016] Identify and classify objects in the image using a set of machine learning models;

[0017] Step S5: Object localization and tracking;

[0018] After object recognition, determine the specific location of the object in the image based on the bounding box, which frames each recognized object in the image to locate it;

[0019] Step S6: Output results;

[0020] The graphics processing unit outputs the results of recognition and localization to the perception center, which analyzes the results and determines whether to generate corresponding decisions. If so, adjust the parameters and execute step S2 to process the newly captured images. If not, adjust the model, call the deep learning model, and execute step S4.

[0021] Further, the sensory system also includes a sensory central anti-electromagnetic interference system and a sensory central storage chip family. The sensory central anti-electromagnetic interference system realizes electromagnetic interference protection. The sensory central anti-electromagnetic interference system covers sensitive components with conductive or magnetic materials. The sensory central storage chip family is used for storing the memory and storage devices of the data collected from the sensor network, including random access memory, fast access memory, and solid state disk. Fast access memory is used to temporarily store data in processing, and solid state disk is used for long-term data storage.

[0022] Further, the head structure is provided with an inspection hatch, a storage chip card slot and a hatch. The hatch is detachably installed in the storage chip card slot, the sensory central storage chip family is installed in the storage chip card slot, and the inspection hatch is hinged to the head structure.

[0023] Further, the head structure is provided with an oral space, and the oral system is installed in the oral space. The oral system includes a jaw structure, a tooth structure, a lip structure, and a tongue structure. The oral system simulates the functions and perception processes of the human mouth.

[0024] Further, the tongue structure also includes a taste perception system that simulates human taste perception, which includes several high-temperature-resistant taste sensors distributed on the tongue surface for sensing five tastes, the high-temperature-resistant taste sensors are integrated into a taste sensor module and converted into electrical signals, the tongue structure is provided with taste signal lines, the electrical signals are transmitted to the signal processing unit through the taste signal lines for processing, the signal processing unit converts the original signals into digital signals and performs filtering and processing, the processed signals are sent to the perception center, the perception center analyzes, decodes and discriminates the signals, identifies the characteristic patterns of the five basic tastes, and compares them with the pre-stored patterns to determine the types and degrees of the tastes.

[0025] Further, the visual system is installed on the head structure, including an eye shell, and eyebrow structure, eyelid structure and eyeball installed in the eye shell, the eyebrow structure includes an eyebrow moving support and two groups of eyebrow moving assemblies, the two groups of eyebrow moving assemblies are connected with the eyebrow moving support respectively, the eyebrow moving assemblies control the up and down movement of the eyebrow moving support, the eyelid structure includes an upper eyelid support, an upper eyelid, a lower eyelid, a lower eyelid support and an eyelid moving assembly, the eyelid moving assembly controls the opening and closing movement of the upper eyelid support and the lower eyelid support, and the eyeball includes a body assembly and an eye moving assembly, the body assembly is movably installed in the middle of the eyelid structure, and the eye moving assembly controls the up and down movement and rotation of the body assembly.

[0026] Further, when the robot detects that a person or object enters its surrounding environment, the visual video tracking system starts capturing images and transmits them to the perception center for analysis, if the perception center determines that the visual direction needs to be adjusted to track the target, sends corresponding control instructions to the eye moving assembly 332 to trigger the eyeball movement; The triggering of the eye movement mainly includes target detection, target importance evaluation, target movement analysis, and the priority and requirements of the current task, which specifically includes the following steps:

[0027] Step A1. Target detection and confirmation:

[0028] 1.1. Detection: The visual video tracking system detects a person or object in the image;

[0029] 1.2. Confirmation: Confirm whether the detected object meets the conditions and characteristics of tracking;

[0030] Step A2. Target importance and priority evaluation:

[0031] 2.1. Importance: Evaluate the importance of the target, which depends on the specific requirements of the task;

[0032] 2.2 Priority: In a multi-target environment, based on the dynamic behavior of the target, the relevance to the task, or pre-set rules, determine which target has a higher tracking priority;

[0033] Step A3. Target dynamic analysis:

[0034] 3.1 Motion tracking: Analyze the motion trajectory of the target to determine whether it is moving, the speed and direction of movement;

[0035] 3.2 Predict future position: Use motion estimation model to predict the future position of the target to determine how to adjust the line of sight most effectively;

[0036] Step A4. Task and environmental requirements:

[0037] 4.1 Task requirements: Determine whether to adjust the line of sight according to the current task requirements of the robot or changes in the environment;

[0038] Step A5. Line of sight adjustment strategy:

[0039] 5.1 Instant adjustment: If the position or speed of the target does not match the predetermined tracking strategy, adjust the direction of the line of sight to keep the target in the center of the field of view or in a suitable position;

[0040] 5.2 Predictive adjustment: Adjust the direction of the line of sight in advance based on the prediction model of the target motion;

[0041] When the robot needs to change the direction of the line of sight, the perception center analyzes the current environment according to the position of the target and the current position of the eye, and determines the target position that needs to be adjusted. The execution unit sends a control signal to the eye movement component 332 to adjust the direction and position of the eye.

[0042] Further, the nose system includes a nose bridge and an olfactory sensor, the nose bridge is provided with a left nostril and a right nostril, and the olfactory sensor is installed in the left nostril and the right nostril respectively. The olfactory sensor receives and analyzes odor information in the surrounding environment, which simulates the olfactory perception mechanism in the human nasal cavity, perceives different odor molecules, and converts them into information that can be understood and responded by the system. The ear system includes ears and auditory sensors, the ears are provided with two groups, which are symmetrically located on both sides of the head structure, and the auditory sensors are fixedly arranged in the ears for receiving sound in the surrounding environment. The signal processing unit receives and processes the sound to obtain a sound signal that can be recognized by the system, and the sound signal is transmitted to the sensory center.

[0043] Further, the step of receiving and processing the sound by the signal processing unit is:

[0044] Step 1: Sound capture: The auditory sensor 52 captures the sound signal in the surrounding environment and converts the sound wave into an electrical signal;

[0045] Step 2: Preprocessing: denoising and filtering the electrical signals;

[0046] Step 3: Extracting sound features: extracting sound features from the preprocessed signals;

[0047] Step 4: Sound recognition: using machine learning algorithms or pattern recognition techniques to analyze and compare sound features to identify the type and meaning of the sound, training the model to recognize different sound categories and processing them;

[0048] Step 5: Action triggering: the sensory center sends instructions to the execution unit according to the identified sound type and meaning, and the execution unit triggers the corresponding action or response.

[0049] In summary, the benefits of the present application are:

[0050] The present application can make timely and effective feedback to the changes in the external environment through the sensory center, sensor network, signal processing unit and execution unit, greatly improving the interaction ability with the outside world; at the same time, the taste sensor enables the humanoid robot to test and evaluate food without human testers, and the taste perception system can simulate the function of human taste perception system, thereby automatically evaluating the taste of food. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 It is a schematic diagram of the head structure of the present application.

[0052] Figure 2 It is a top view of the head structure of the present application.

[0053] Figure 3 It is a schematic diagram of the jaw structure and tooth structure of the present application Figure 1 .

[0054] Figure 4 It is a schematic diagram of the jaw structure and tooth structure of the present application Figure 2 .

[0055] Figure 5 It is a schematic diagram of the tooth structure and lip structure of the present application Figure 1 .

[0056] Figure 6 It is a schematic diagram of the tooth structure and lip structure of the present application Figure 2 .

[0057] Figure 7 It is a schematic diagram of the tooth structure and lip structure of the present application Figure 3 .

[0058] Figure 8 It is a schematic diagram of the visual system of the present application Figure 1 .

[0059] Figure 9 schematic diagram of the visual system of the present application Figure 2 .

[0060] Figure 10 schematic diagram of the visual system of the present application Figure 3 .

[0061] Figure 11 schematic diagram of the eyeball of the present application.

[0062] Figure 12 schematic diagram of the eyeball of the present application.

[0063] Figure 13 schematic diagram of the nose system and ear system of the present application. DETAILED DESCRIPTION

[0064] The present application can be implemented or applied in other different embodiments, and the details in the specification can be modified or changed based on different views and applications without departing from the spirit of the present application. It should be noted that the following examples and features in the examples can be combined with each other without conflict.

[0065] It should be noted that the diagrams provided in the following examples only illustrate the basic concept of the present application in a schematic manner, and only the components related to the present application are shown in the diagrams, not the number, shape and size of the components when actually implemented. The actual implementation of each component can be a random change in type, number and proportion, and the component layout pattern can also be more complex.

[0066] All directional indications (such as up, down, left, right, front, back, transverse, longitudinal, etc.) in the embodiments of the present application are only used to explain the relative position relationship, movement condition, etc. between the components in a certain specific posture, and if the specific posture changes, the directional indications will also change accordingly.

[0067] Due to installation errors and other reasons, the parallel relationship referred to in the embodiments of the present application can actually be an approximate parallel relationship, and the vertical relationship can actually be an approximate vertical relationship.

[0068] Embodiment one:

[0069] As shown in Figures 1-13 , the bionic service type humanoid robot head structure includes a skull structure 1 and a sensory system, a mouth system 2, a visual system 3, a nose system 4 and an ear system 5 integrated and installed on the skull structure.

[0070] The sensory system includes a sensing center, a sensor network, a signal processing unit, and an execution unit. The sensor network signal is connected to the processing unit, the signal processing unit is connected to the sensing center, and the sensing center is connected to the execution unit.

[0071] A sensor network comprises several sensors, including visual sensors, auditory sensors, olfactory sensors, gustatory sensors, and other types of sensors, to capture various information about the external environment, such as sound, images, and smells. The sensor network transmits the captured information to the signal processing unit. The sensor network provides various information about the external environment and serves as the input source for the perception center.

[0072] The signal processing unit receives raw signals from the sensor network, performs signal processing and preprocessing to enhance the accuracy and stability of the signals, and then transmits the processed signals to the sensing center to provide the sensing center with more reliable information.

[0073] As the intelligent core of the humanoid robot, the perception center receives and processes information, analyzes and makes decisions to determine the state of the external environment and the required response, so as to realize the perception of the surrounding environment and the response to external stimuli. The perception center is interconnected with the oral system 2, visual system 3, nasal system 4 and ear system 5 to realize overall perception and decision-making.

[0074] The perception center utilizes NVIDIA's Project GR00T and the high-performance processors of the Isaac robotic platform. Leveraging NVIDIA's proprietary SoC (System-on-a-Chip) technology, it is optimized for the needs of complex robotic systems to achieve highly automated and intelligent operation, specifically including the following functions:

[0075] 1. High-performance computing support:

[0076] NVIDIA's SoC design is primarily aimed at providing sufficient processing power to support the operation of AI algorithms and models in robots. This chip integrates a high-efficiency GPU (Graphics Processing Unit), enabling robots to rapidly process visual and perceptual data, performing complex image recognition, object detection, and environmental understanding in real time. This is because GPUs are designed for parallel processing of large amounts of data, making them ideal for performing deep learning tasks.

[0077] 2. Multimodal sensor fusion:

[0078] In humanoid robot applications, data from multiple sensors, including vision, hearing, touch, and position awareness, need to be integrated and processed. NVIDIA's SoC supports this advanced sensor fusion, enabling robots to better understand and adapt to their environment. For example, a robot can simultaneously process visual data from a camera and force feedback from sensors to achieve more precise object manipulation and navigation.

[0079] 3. Low latency real-time response:

[0080] Humanoid robots require extremely low response times when performing tasks such as delivery, rescue, or collaborative work. NVIDIA's SoC ensures low latency and high-speed data processing capabilities by optimizing computing paths and improving data transfer efficiency. This allows robots to react quickly in dynamic and unpredictable environments.

[0081] 4. Energy efficiency:

[0082] Considering the need for humanoid robots to operate for long periods on battery power, energy efficiency is a key factor in designing SoCs. NVIDIA's chips optimize energy consumption while maintaining high performance through advanced manufacturing processes and power management techniques, extending the operating time of robots.

[0083] 5. Software and hardware synergy:

[0084] The Isaac robot platform not only includes the hardware SoC but also a complete software development kit (SDK) that includes simulators, development tools, and pre-trained AI models. These software tools are tightly integrated with the SoC hardware, allowing developers to customize and optimize robot behavior for specific application scenarios, improving development efficiency and robot performance.

[0085] The execution unit receives instructions from the perception center and performs corresponding actions or tasks, such as movement, operation of external devices, etc.

[0086] The execution unit can use servo motors, electromagnetic brakes, linear actuators, inductors, and brakes.

[0087] The signal processing unit, as the core processing unit for all sensory data, specifically includes a central processing unit (CPU) and a graphics processing unit (GPU).

[0088] The central processing unit is responsible for processing program instructions, managing software operations, and other computing tasks, while the graphics processing unit (GPU) is used for image analysis and machine vision, such as NVIDIA's Jetson series.

[0089] The central processing unit (CPU) can use conventional NVIDIA Jetson AGX Xavier, Intel Core i7-1185G7, AMD Ryzen 9 5900HX, Qualcomm Snapdragon 888, Apple M1, etc.

[0090] The graphics processing unit (GPU) can use conventional NVIDIA GeForce RTX 3080, AMD Radeon RX 6800XT, Intel Iris Xe Graphics G7, Qualcomm Adreno 660, Apple M1 GPU, etc.

[0091] The graphics processing unit and the visual sensor constitute a visual recording tracking system, and the visual sensor captures visual information of the surrounding environment and transmits it to the graphics processing unit for analysis, recognition and tracking.

[0092] The graphics processing unit is built-in image processing algorithm, the visual sensor captures continuous image frames, and then identifies and extracts the required images or objects through the image processing algorithm, which includes the following steps:

[0093] Step S1: The visual sensor captures real-time images or videos in the environment and transmits them to the graphics processing unit;

[0094] Step S2: The image processing algorithm pre-processes the captured images;

[0095] The pre-processing includes image enhancement such as resizing, normalization, denoising, etc., to improve the effect and efficiency of subsequent processing steps;

[0096] Step S3: Feature extraction: extract specific information from pre-processed images, which can help identify objects in the image;

[0097] The feature extraction method includes edge detection based on edge detection algorithm, corner detection based on corner detection algorithm, texture analysis based on texture analysis algorithm, etc.

[0098] Step S4: Object recognition;

[0099] Use the set machine learning model to identify and classify objects in the image;

[0100] The machine learning model has been trained on a large amount of labeled data to recognize different types of objects;

[0101] Step S5: Object localization and tracking;

[0102] After object recognition, the specific location of the object in the image is determined based on bounding boxes, which frame each recognized object in the image to locate it;

[0103] Step S6: output the result;

[0104] The graphics processing unit outputs the results of recognition and localization to the perception center, which analyzes the results and determines whether to generate corresponding decisions. If so, adjust the parameters and execute step S2 to process the newly captured image. If not, adjust the model and call the deep learning model (such as convolutional neural network) to execute step S4.

[0105] Step S4 object recognition is achieved through machine learning and computer vision techniques. Common methods include:

[0106] Object classification: classifying objects in an image into predefined categories.

[0107] Object detection: detecting the location and bounding box of objects in an image.

[0108] Object tracking: tracking the motion trajectory of a specific object in consecutive image frames.

[0109] Depending on the different actions that can be taken based on the image, the program or algorithm is predefined. For example, if the image recognition system detects a human face, the robot's head may turn and face the face; if a specific object is detected, the robot may perform tasks or actions related to the object, such as grabbing, moving, etc.

[0110] This embodiment proposes a mathematical model of an image recognition algorithm based on neural network multi-model feature fusion, which integrates image features from multiple sensors to improve the accuracy of image recognition. Specifically:

[0111] Assume the number of sensors is n, each sensor can capture images and extract image features, let the image features captured by the i-th sensor be represented as where i ∈ {1, 2,..., n}. These image features are integrated into a feature vector, represented as x ∈ R m where

[0112] A deep neural network is established for image recognition, with the following structure:

[0113] h = σ(W1x + b1)

[0114] y = softmax(W2h + b2)

[0115] where y is the model's predicted class probability distribution, h represents the hidden layer bias vector, and W1 ∈ Rh×m and W2 ∈ R k×h denote the weight matrices from input to hidden and hidden to output layers, respectively, R h×m denotes the weight matrix from input to hidden layer, R k×h denotes the weight matrix from hidden to output layer, b1 ∈ R h and b2 ∈ R k denote the bias vectors for hidden and output layers, respectively, k denotes the output layer bias vector, R h denotes the weight matrix for input layer, R k denotes the weight matrix for hidden layer; σ() denotes the activation function, softmax() denotes the softmax function.

[0116] To train the neural network, a cross-entropy loss function is used:

[0117]

[0118] where g is the model's predicted class probability distribution, is the true class label, and j denotes the class index.

[0119] The loss function is minimized using gradient descent, i.e., updating the weights and biases:

[0120]

[0121] where α is the learning rate, are the gradients of the loss function with respect to the weight matrices W1, W2 and bias vectors b1, b2, respectively.

[0122] A deep convolutional neural network (CNN) is used for image recognition. The CNN contains multiple convolutional, pooling, and fully connected layers for feature extraction and classification. Let the input to the CNN be x, and the output be y ∈ R k , which denotes the predicted class result for the image.

[0123] The output y of the CNN can be represented as:

[0124] h1 = ReLU(W1x + b1)

[0125] h2 = MaxPooling(h1)

[0126] h3 = ReLU(W2h2 + b2)

[0127] h4 = MaxPooling(h3)

[0128] h5 = ReLU(W3h4 + b3)

[0129] h6 = Flatten(h5)

[0130] y = Softmax(W4h6 + b4)

[0131] where b3, b4 are bias vectors, W3, W4 are weight matrices, ReLU() represents the rectified linear unit activation function, MaxPooling() represents the max pooling operation, Flatten() represents the flattening of multi-dimensional data into a one-dimensional vector, and Softmax() represents the softmax function.

[0132] The cross-entropy loss function is used to measure the difference between the model's predicted results and the true labels:

[0133]

[0134] where g is the model's predicted class probability distribution, is the true class label, k represents the number of classes, and j is the model's initial parameter value.

[0135] The loss function is minimized by gradient descent, i.e., updating the weights and biases:

[0136]

[0137] where α is the learning rate, and represent the gradients of the loss function with respect to the weight matrix W1 and the bias vector b1, respectively.

[0138] Calculation example:

[0139] Suppose there are 3 sensors that capture color features, texture features, and shape features of an image, respectively, and the output feature vectors of each sensor are x1∈R 10 x1∈R10, x2∈R 15 x2∈R15, and x3∈R 20 x3∈R20.

[0140] The goal is to use a deep convolutional neural network to identify the class of the image, assuming there are 5 classes.

[0141] First, initialize the parameters of the neural network. Assume that the hidden layer contains 20 neurons and the learning rate is 0.01.

[0142] [W1∈R 20×(10+15+20) , W2∈R 5×20 ][b1∈R 20 , b2∈R 5 ]

[0143] Next, prepare the training data. Assume there are 1000 samples, and each sample's label is represented by one-hot encoding. The random gradient descent method will be used to update the parameters, and each update uses one sample.

[0144] Then, the training and prediction of the model can be performed. Assume that 100 iterations of training have been completed, and a new image needs to be classified and predicted.

[0145] Finally, the prediction result will be output, and the accuracy of the model will be calculated.

[0146] Operation process:

[0147] Suppose the hidden layer contains 20 neurons, and the learning rate is 0.01. The initialization parameters are as follows:

[0148] \[ W1 \in R 20×45 , W2 \in R 5×20 \]\[ b1 \in R 20 , b2 \in R 5 \]

[0149] Prepare the training data: assume there are 1000 samples, and each sample's label is represented by one-hot encoding.

[0150] Use the random gradient descent method to update the parameters, and each update uses one sample for training. Complete 100 iterations of training.

[0151] Based on the above premise, assume that there is a new image that needs to be classified and predicted. The output result is class 3, and the accuracy is 80%.

[0152] The visual sensor can use a conventional Sony IMX477 CMOS image sensor, OmniVision OV5670 CMOS image sensor, ON Semiconductor AR0144 CMOS image sensor, Samsung S5K4H7YX CMOS image sensor, Canon 5D Mark IV CMOS image sensor, etc.

[0153] The sensory system also includes a sensory central electromagnetic interference prevention system and a sensory central storage chip family. The sensory central electromagnetic interference prevention system realizes electromagnetic interference (EMI) protection. The sensory central electromagnetic interference prevention system uses EMI shielding materials, which use conductive or magnetic materials to cover sensitive components to reduce the impact of external electromagnetic waves. The sensory central storage chip family is used to store the memory and storage devices of the data collected from various sensor networks, including random access memory (RAM): fast access memory used for temporary storage of data in processing, and solid state drive (SSD): used for long-term data storage, with faster read and write speed and high durability.

[0154] The sensory central storage chip family can adopt conventional Samsung PM9A3 E1.S SSD, Western Digital WD Black SN850 NVMe SSD, SK Hynix Gold P31 NVMe SSD, Crucial P5 Plus NVMe SSD, Intel Optane SSD 905P U.2 SSD, etc.

[0155] The rear side of the head structure 1 is provided with an inspection hatch 11, and the inspection hatch 11 is hinged to the head structure 1. This hinged structure allows the inspection hatch 11 to be opened and stopped at a certain angle, facilitating access to the internal structure of the head structure 1. The inspection hatch 11 facilitates the replacement and maintenance of the sensory system of the robot. The design of the inspection hatch 11 allows technicians to easily replace or upgrade these components, as well as perform routine maintenance and fault diagnosis.

[0156] The rear side of the head structure 1 is also provided with a storage chip card slot and a hatch 12, and the hatch 12 is detachably installed in the storage chip card slot. The sensory central storage chip family is installed in the storage chip card slot.

[0157] The head structure 1 is provided with an oral cavity space, and the oral system 2 is installed in the oral cavity space. The oral system 2 includes a jaw structure 20, a tooth structure, a lip structure 21, and a tongue structure 23. The oral system 2 is connected to the sensory system to simulate the functions and sensory processes of the human oral cavity.

[0158] The jaw structure 20 includes an upper jaw shell 201, a jaw movable assembly, and a lower jaw shell 202. The left and right ends of the upper jaw shell 201 and the lower jaw shell 202 are rotationally connected by an upper and lower jaw fixed shaft 203. The upper jaw shell 201 and the lower jaw shell 202 can rotate in the up-down direction relative to the upper and lower jaw fixed shaft 203. The upper and lower jaw fixed shaft 203 is installed in the head structure 1, so that the upper jaw shell 201 and the lower jaw shell 202 can move up and down relative to the head structure 1.

[0159] The upper jaw shell 201 and the lower jaw shell 202 are made of abs material.

[0160] The jaw movable assembly is provided with two groups, and the two groups of jaw movable assemblies are symmetrically arranged on the left and right sides. The jaw movable assembly connects the upper jaw shell 201 and the lower jaw shell 202. The jaw movable assembly receives control signals from the execution unit to control the upper jaw shell 201 and the lower jaw shell 202 to perform corresponding movements.

[0161] The jaw moving assembly includes a brushless servo 204, a jaw rear connecting rod structure 205 and a jaw front connecting rod structure 206, the jaw rear connecting rod structure 205 and the jaw front connecting rod structure 206 are connected with the brushless servo 204 respectively, the jaw rear connecting rod structure 205 and the jaw front connecting rod structure 206 respectively include an upper moving joint 2051, a middle connecting rod 2052 and a lower moving joint 2053, the upper moving joint 2051 and the lower moving joint 2053 are connected at two ends of the middle connecting rod 2052, the upper moving joint 2051 is connected with a fixed pin of the upper jaw shell 201, the lower moving joint 2053 is connected with a fixed pin of the lower jaw shell 202, the middle connecting rod 2052 is provided as an extension rod, the brushless servo 204 is connected with the upper moving joint 2051 to realize the movement control of the upper jaw shell 201 and the lower jaw shell 202, realize the front and rear movement speed control of the upper jaw shell 201 and the lower jaw shell 202 relative to the skull structure 1, and local micro-motion can also be realized.

[0162] The brushless servo 204 adopts a 15kg level to provide sufficient power for the movement of the jaw.

[0163] In the embodiment, the upper jaw shell 201 and the lower jaw shell 202 are controlled to perform corresponding movements by the jaw moving assembly, the jaw moving assembly is connected with an execution unit, the execution unit sends a jaw control signal to the jaw moving assembly after processing a jaw moving trigger instruction issued by a perception center, and then controls the movement of the upper jaw shell 201 and the lower jaw shell 202.

[0164] The jaw moving trigger instruction includes instructions for triggering the up-down movement and the left-right micro-motion of the jaw, which are respectively a jaw up-down movement trigger instruction and a jaw left-right micro-motion trigger instruction.

[0165] The jaw up-down movement trigger instruction includes

[0166] The chewing instruction: when the perception center detects food particles or other objects in the oral cavity space, the upper jaw shell 201 is triggered to move up and down to simulate the chewing action and send the food into the oral cavity.

[0167] The voice instruction: when the humanoid robot receives a specific oral instruction, the up-down movement of the jaw is triggered, such as "mouth closing", "mouth opening" and the like.

[0168] The voice interaction instruction: when the humanoid robot has a voice conversation with the user, the up-down movement of the upper jaw is triggered, such as the start of sound, the end of sound and the like.

[0169] The visual recognition instruction: when a specific gesture or object movement is detected by a visual sensor, the up-down movement of the upper jaw is triggered, such as detecting the movement of the hand or the object approaching the mouth.

[0170] Emotion expression instruction: when the humanoid robot needs to simulate a smile or cry, etc. Expression, according to the emotion recognition algorithm to trigger the up and down movement of the upper jaw;

[0171] Jaw left and right micro-motion trigger instruction includes

[0172] Visual tracking instruction: when a specific object or person is detected to move left and right through a visual sensor, trigger the small left and right movement of the upper jaw to simulate the attention or observation action of the human;

[0173] Voice interaction instruction: when the humanoid robot has a voice conversation with the user, according to the directionality of the voice instruction, trigger the small left and right movement of the upper jaw to express the intention of listening or responding;

[0174] Emotion simulation instruction: according to the emotion recognition algorithm to analyze the emotional signals in the environment, for example, when the user's smile or nervousness is detected, trigger the micro-motion of the upper jaw to establish a closer emotional connection with the user.

[0175] The tooth structure includes an upper tooth bed 241, an upper tooth 242, a lower tooth 243, and a lower tooth bed 244, the upper tooth bed 241 and the lower tooth bed 244 are fixedly arranged in the skull structure 1, the upper tooth 242 is fixedly arranged in the upper tooth bed 241, the lower tooth 243 is fixedly arranged in the lower tooth bed 244, the upper tooth bed 241 is fixedly arranged in the upper jaw shell 201, the lower tooth bed 244 is fixedly arranged in the lower jaw shell 202, the upper tooth 242 and the lower tooth 243 move synchronously with the upper jaw shell 201 and the lower jaw shell 202, and the 15kg level brushless servo 204 transmits torque to provide sufficient biting force for the tooth structure.

[0176] In this embodiment, the upper tooth bed 241, the upper tooth 242, the lower tooth 243, and the lower tooth bed 244 are all made of high molecular polymer, and the number of teeth of the upper tooth 242 and the lower tooth 243 is self-set according to needs.

[0177] The lip structure 21 includes an upper lip 211, a lip movable assembly, and a lower lip 212, the upper lip 211 is movably connected to the upper jaw shell 201, the lower lip 212 is movably connected to the lower jaw shell 202, the lip movable assembly is connected with the upper lip 211 and the lower lip 212, and the up-down movement and the left-right movement of the upper lip 211 and the lower lip 212 are controlled through the lip movable assembly.

[0178] There are eight groups of lip movable assemblies in total, four groups of lip movable assemblies are used for the connection between the upper lip 211 and the upper jaw shell 201, the four groups of lip movable assemblies are evenly distributed along the length direction of the upper lip 211, and are specifically distributed at both ends of the upper lip 211 and the middle part of the upper lip 211, and the other four groups of lip movable assemblies are used for the connection between the lower lip 212 and the lower jaw shell 202, and the distribution manner is the same as above, which will not be repeated here.

[0179] The lip moving assembly includes a servo steering engine 213, a moving connecting rod 214 and a moving telescopic rod 215. The servo steering engine 213 is installed on the upper jaw shell 201, the moving connecting rod 214 is connected to the servo steering engine 213, one end of the moving telescopic rod 215 is connected to the moving connecting rod 214, and the other end of the moving telescopic rod 215 is rotatably connected to the upper lip 211. The lower lip 212 is movably connected to the lower jaw shell 202 through the moving assembly, which is the same as the above and will not be repeated here.

[0180] The servo steering engine 213 provides power for the movement of the upper lip 211 and the lower lip 212. The servo steering engine 213 and the moving connecting rod 214 are connected through a universal ball, and the moving connecting rod 214 and the moving telescopic rod 215 are connected through a universal ball, realizing 360° rotation. When the two groups of lip moving assemblies located in the middle of the upper lip 211 and / or the lower lip 212 are activated, the up-and-down movement of the upper lip 211 and / or the lower lip 212 can be realized. When the lip moving assemblies located at both ends of the upper lip 211 and / or the lower lip 212 are activated, the left-and-right movement of the upper lip 211 and / or the lower lip 212 can be realized.

[0181] In this embodiment, the upper lip 211 and the lower lip 212 are made of high-temperature-resistant silicone polymer.

[0182] In this embodiment, the upper lip 211 and the lower lip 212 are controlled to perform corresponding movements by the lip moving assembly. The lip moving assembly is connected to an execution unit. The execution unit receives a lip moving trigger instruction, processes it, and sends a lip control signal to the lip moving assembly, thereby controlling the movement of the upper lip 211 and the lower lip 212.

[0183] The lip moving trigger instruction includes instructions for triggering the up-and-down movement and the left-and-right movement of the lips, namely, a lip up-and-down movement trigger instruction and a lip left-and-right movement trigger instruction.

[0184] The lip up-and-down movement trigger instruction includes

[0185] Voice instruction: When the humanoid robot receives a specific verbal instruction, the up-and-down movement of the lips is triggered, such as "mouth closing" and "mouth opening".

[0186] Visual recognition instruction: When a specific gesture or object movement is detected by a visual sensor, the up-and-down movement of the lips is triggered, such as detecting hand movements or objects approaching the mouth.

[0187] Emotion expression instruction: When the humanoid robot needs to simulate a smile, expression change or sound, the up-and-down movement of the lips is triggered according to the emotion recognition algorithm.

[0188] The lip left-and-right micro-motion trigger instruction includes

[0189] Visual tracking instructions: When a specific object or person is detected moving left and right by the visual sensor, trigger the slight left and right movement of the lips to simulate the human attention or observation action;

[0190] Voice interaction instructions: When the humanoid robot has a voice conversation with the user, according to the directionality of the voice instructions, trigger the slight left and right movement of the lips to express the intention of listening or responding;

[0191] Emotion simulation instructions: According to the emotion recognition algorithm to analyze the emotional signals in the environment, for example, when the user's smile or nervousness is detected, trigger the slight movement of the lips to establish a closer emotional connection with the user.

[0192] The oral cavity space is provided with a mounting bracket 230, and the tongue structure 23 is installed in the mounting bracket 230. The tongue structure 23 includes a controllable telescopic servo, a tongue fetus 231, and a tongue joint structure 232. The tongue joint structure 232 is composed of a plurality of tongue joints 233, which are sequentially installed and connected through connecting pipelines, forming a chain structure similar to a snake, which can realize telescopic and curling movement. One end of the tongue joint structure 232 is connected with the mounting bracket 230 and the controllable telescopic servo, which ensures the movement of the tongue joint structure 232 according to the instructions. The other end of the tongue joint structure 232 is a free end. The tongue fetus 231 is nested outside the tongue joint structure 232. The tongue fetus 231 is made of APS+PC material, and a layer of sheath can be nested outside the tongue fetus 231.

[0193] The sheath of the tongue fetus 231 is made of special high-molecular polymer, specifically engineering plastic polyphenylene sulfide (PPS). PPS has very good thermal stability (can be used at 200℃ for a long time, and can withstand up to 260℃ for a short time), and is also very resistant to chemical corrosion and oxidation. PPS has high mechanical strength and certain elasticity.

[0194] The tongue structure 23 also includes a taste perception system that simulates the human taste perception ability. The taste perception system includes a plurality of high-temperature-resistant taste sensors uniformly distributed on the surface of the tongue fetus 231 for perceiving five basic tastes: sour, sweet, bitter, salty, and umami. Each type of taste has 2 sensors distributed on the left and right sides to ensure comprehensive coverage and accurate perception of taste stimulation.

[0195] The taste perception system of the humanoid robot in this embodiment enables testing and evaluation of food without human testers. The taste perception system can simulate the function of the human taste perception system, thereby automatically evaluating the taste of food.

[0196] The high-temperature-resistant taste sensor integrates a taste sensor module, which can detect the taste in the oral cavity and convert it into an electrical signal. The taste signal line is distributed inside the tongue section structure 232, and the electrical signal is transmitted to the signal processing unit through the taste signal line for processing. The signal processing unit converts the original signal into a digital signal and performs filtering and processing to enhance the accuracy of the signal. The processed signal is sent to the perception center, which analyzes, decodes, and discriminates the signal. It can identify the characteristic patterns of the five basic tastes and compare them with the pre-stored patterns to determine the type and degree of the taste.

[0197] The process of signal analysis by the perception center is as follows:

[0198] The perception center receives digital signals processed by the signal processing unit. These signals are first analyzed to determine their intensity and pattern, which is the first step in identifying the taste. By amplifying and filtering these signals, the perception center can distinguish which signals are relevant and which are background noise or irrelevant information.

[0199] The process of signal decoding by the perception center is as follows:

[0200] The decoding step involves converting the electrical signals obtained in the analysis phase into specific taste information. This process requires the use of known neural encoding patterns, which are pre-set through extensive sensory tests and data analysis. During the decoding process, the nervous system matches the received signals with these pre-set patterns, such as sweet, sour, bitter, salty, and umami.

[0201] The process of signal discrimination by the perception center is as follows:

[0202] In the discrimination phase, the perception center uses the previous decoding results to determine the specific taste type and intensity. This step involves comparing the decoded taste patterns with the standard taste patterns stored in the database. Through comparison, the perception center can accurately identify the taste being experienced and assess its intensity.

[0203] The analysis, decoding, and discrimination of signals by the perception center is a highly integrated and automated process that relies on advanced neural networks and machine learning techniques to improve the accuracy and efficiency of identification. In artificial systems, algorithms and neural network models that simulate these biological processes are used to perform similar tasks.

[0204] This embodiment divides the taste level into 0-10 levels to represent the robot's acceptance of different tastes. The threshold can also be adjusted according to specific circumstances.

[0205] The perception center discriminates the type and degree of the taste and sends corresponding control signals to the execution unit to perform corresponding actions.

[0206] The robot expresses the robot's feeling for different tastes through the movement of the mouth, sound or expression, etc. according to the detected taste and taste level, that is, the execution unit makes corresponding reactions and decisions according to the preset mode and task requirements, so that the tongue structure performs specific actions such as opening and closing, chewing, swallowing, etc., so that the humanoid robot can more realistically simulate the human taste perception process. Specifically:

[0207] Sour: The tongue slightly curls or makes a sour expression, and the voice has a certain sharpness.

[0208] Sweet: The tongue licks the lips or smiles, and the voice has a pleasant tone.

[0209] Bitter: The tongue slightly contracts or makes an unpleasant expression, and the voice may have a depressed or uncomfortable tone.

[0210] Salt: The tongue stretches out or licks the lips, making a motion of thirst for water, and the voice has a feeling of thirst.

[0211] Fresh: Slightly open the mouth, accompanied by the action of licking the lips, and the voice has a fresh or energetic tone.

[0212] The voice with emotional tone or feeling in the above is established by using a conventional method to build an acid / sweet / bitter / salt / fresh model, the model establishes the association between taste and voice, and the corresponding voice is called by recognizing the taste.

[0213] In the embodiment, the tongue structure 23 is controlled by a controllable telescopic servo to perform corresponding movements. The controllable telescopic servo is connected to the execution unit, and the execution unit sends a tongue control signal to the controllable telescopic servo after processing the tongue activity trigger instruction, thereby controlling the movement of the tongue structure 23.

[0214] The tongue activity trigger instruction includes instructions for triggering the extension and contraction movement and the curling movement of the tongue, which are respectively the tongue extension and contraction movement trigger instruction and the tongue curling movement trigger instruction;

[0215] The tongue extension and contraction movement trigger instruction includes

[0216] Voice instruction: When the humanoid robot receives a specific oral instruction, the extension and contraction movement of the tongue is triggered, such as "tongue out" and "tongue back".

[0217] Visual recognition instruction: When a specific gesture or object movement is detected by the visual sensor, the extension and contraction movement of the tongue is triggered, such as detecting the movement of the hand or the object approaching the mouth.

[0218] Emotion expression instruction: When the humanoid robot needs to simulate swallowing, spitting out objects or expression changes, the extension and contraction movement of the tongue is triggered according to the emotion recognition algorithm.

[0219] The tongue curling motion triggering instruction includes

[0220] Voice interaction instruction: when the humanoid robot has a voice conversation with the user, the tongue curling motion is performed according to the content and emotion of the voice instruction, such as simulating spitting, licking, etc.

[0221] Emotion simulation instruction: according to the emotion recognition algorithm, analyze the emotional signals in the environment, for example, when detecting the user's smile, surprise or anger, trigger the micro-motion or curling motion of the tongue to establish a closer emotional connection with the user.

[0222] The oral cavity space is also provided with a perception organ, which is a device simulating the function of the human oral cavity. The perception organ perceives the stimulation in the oral cavity by capturing various stimulation signals inside the oral cavity, such as taste, temperature, texture, etc. The perception organ integrates a variety of sensors to simulate and evaluate the sensory experience of humans during food intake, including the above-mentioned taste sensors, touch sensors, temperature sensors, pressure sensors, etc. Several types of sensors work together to fully simulate the taste, temperature, texture, etc. of food.

[0223] The installation position of the perception organ simulates the structural layout of the human oral cavity. The taste sensors and temperature sensors are distributed on the surface of the tongue fetus, the touch sensors and pressure sensors are distributed on the tooth structure, and the temperature sensors are distributed on different parts of the simulated oral cavity to fully capture the experience during food intake.

[0224] The taste of food can be captured by the taste sensor, and then processed by the signal processing unit. The perception center analyzes the signal to obtain the taste. The temperature of the food can be captured by the temperature sensor. The texture of the food refers to the physical composition and feeling of the food, such as hardness, viscosity, humidity, and granularity. The texture is captured by the touch sensor. For example, the tooth structure can evaluate the hardness of the food through the pressure sensor, and the touch sensor can evaluate the viscosity or roughness of the surface of the food.

[0225] The head structure is provided with two groups of symmetrical visual mounting holes, and the visual system 3 is provided with two groups, and the two groups of visual systems 3 are respectively arranged at the positions of the two groups of visual mounting holes, and the two groups of visual systems 3 are the same in structure, and the structure of one group of visual systems 3 in the embodiment is described. The visual system 3 comprises an eye shell 30, and a brow structure 31, an eyelid structure 32 and an eyeball 33 arranged in the eye shell 3. The brow structure 31 is arranged above the visual mounting hole, the eyelid structure 32 and the eyeball 33 are arranged in the visual mounting hole, the brow structure 31 comprises a brow movable support 311 and two groups of brow movable assemblies 312, the two groups of brow movable assemblies 312 are respectively connected with the brow movable support 311, and the brow movable assemblies 312 are used for realizing the up-down movement of the brow movable support 311. The eyelid structure 32 comprises an upper eyelid support 321, an upper eyelid 322, a lower eyelid 323, a lower eyelid support 324 and an eyelid movable assembly. The eyelid movable assembly controls the opening and closing movement of the upper eyelid support 321 and the lower eyelid support 324. The eyeball 33 comprises a body assembly 331 and an eye movable assembly 332. The body assembly 33 is movably arranged in the middle part of the eyelid structure 32, and the eye movable assembly 332 controls the up-down movement and rotation of the body assembly 331.

[0226] The eye shell 30 is an external structure for protecting and supporting the internal components of the eye.

[0227] The brow movable support 311, the upper eyelid support 321 and the lower eyelid support 324 are made of abs material. The upper eyelid 322 and the lower eyelid 323 are made of high-temperature-resistant silicone polymer, and the eyelashes can be arranged on the upper eyelid 322 and the lower eyelid 323.

[0228] The brow movable assembly 312 comprises a servo steering wheel 3121 and a brow connecting rod 3122. The upper end of the brow connecting rod 3122 is connected with an upper movable joint 3123, the lower end of the brow connecting rod 3122 is connected with a lower movable joint 3124, the servo steering wheel 3121 is arranged on the head structure, the upper movable joint 3123 is connected with the servo steering wheel 3121 through a brow movable pin 3125, and the brow movable pin 3125 can realize the rotation similar to the universal ball. The brow movable support 311 is provided with a brow fixing pin, and the lower movable joint 3124 is connected with the brow fixing pin.

[0229] The servo steering wheel 3121 is started to control the rotation of the brow movable pin 3125. When the brow movable pin 3125 rotates from the lower part to the upper part, the brow connecting rod 3122 is driven to move upwards, and then the brow movable support 311 is driven to move upwards. When the brow movable pin 3125 rotates from the upper part to the lower part, the brow connecting rod 3122 is driven to move downwards, and then the brow movable support 311 is driven to move downwards, so as to realize the up-down movement of the brow movable support 311.

[0230] The two groups of brow movable supports 311 of the two groups of visual systems 3 can be driven by the brow movable assemblies 312 to move synchronously or asynchronously, so as to realize the expression of various emotions of the robot.

[0231] In this embodiment, the eyebrow movable support 311 is controlled to perform corresponding movements by two groups of eyebrow movable assemblies 312. The eyebrow movable assemblies 312 are connected to an execution unit. The execution unit sends eyebrow movable support control signals to the eyebrow movable assemblies 312 after processing eyebrow movable support activity trigger instructions, thereby controlling the movement of the eyebrow movable support structure 23.

[0232] The eyebrow movable support activity trigger instructions are up-and-down movement trigger instructions for the eyebrow movable support and asynchronous movement instructions for the two groups of eyebrow movable supports in left-and-right positions, including:

[0233] Visual recognition instructions: when the user's eye expressions or gestures are detected by the camera or other visual sensors, the up-and-down movement of the right eyebrow is triggered according to the user's eye direction or gesture action. For example, when the user raises the eyebrows upward or frowns downward, the corresponding action of the right eyebrow is triggered;

[0234] Voice interaction instructions: when the humanoid robot has a voice conversation with the user, the up-and-down movement of the right eyebrow is triggered according to the content and directionality of the voice instructions to express the attention or response action of the robot.

[0235] The asynchronous movement trigger instructions for the two groups of eyebrow movable supports include

[0236] Visual tracking instructions: the head movement or facial expressions of the user are tracked by the camera or other visual sensors. When the user's head or eyes turn left and right, the left-and-right micro-movement of the right eyebrow is triggered to simulate the eye direction or expression change of a person.

[0237] Voice interaction instructions: when the humanoid robot has a voice conversation with the user, the left-and-right micro-movement of the right eyebrow is triggered according to the content and directionality of the voice instructions to express the attention or response action of the robot.

[0238] The upper eyelid 322 is fixedly arranged on the upper eyelid support 321, and the lower eyelid 323 is fixedly arranged on the lower eyelid support 324. The upper eyelid support 321 and the lower eyelid support 324 are respectively provided with tooth ends at left and right ends. The two tooth ends are coaxially arranged inside and outside and meshed. When one of the tooth ends is driven to rotate, the upper eyelid support 321 and the lower eyelid support 324 rotate in opposite directions due to the meshing movement, so that the upper eyelid support 321 and the lower eyelid support 324 can perform opening and closing movements to realize blinking operation. The opening and closing frequency can be set as needed.

[0239] The eyelid movement assembly provides power for the eyelid movement assembly. Specifically, the eyelid movement assembly includes an eyelid opening and closing rotating shaft 325 and a servo steering engine. The servo steering engine is installed on the skull structure, and the output shaft of the servo steering engine is fixedly connected with the eyelid opening and closing rotating shaft 325. The eyelid opening and closing rotating shaft 325 is fixedly connected with the inner tooth end. The servo steering engine starts to transmit power to the inner tooth end through the eyelid opening and closing rotating shaft 325, so as to realize the opening and closing movement of the upper eyelid support 321 and the lower eyelid support 324.

[0240] The two groups of eyelid structures 32 of the two groups of visual systems 3 can be driven by the eyelid movement assembly to perform synchronous opening and closing movement or asynchronous opening and closing movement.

[0241] In this embodiment, the upper eyelid support 321 and the lower eyelid support 324 are controlled to perform corresponding movements by the eyelid movement assembly. The servo steering engine is connected with an execution unit. The execution unit receives eyelid support trigger instruction processing and sends eyelid support control signal to the servo steering engine, so as to control the movement of the eyelid support.

[0242] The eyelid support trigger instruction includes an instruction for triggering the opening and closing movement of the upper eyelid support 321 and the lower eyelid support 324.

[0243] The opening and closing movement trigger instruction of the upper eyelid support 321 and the lower eyelid support 324 includes

[0244] Visual recognition instruction: when the user's eye expression or gesture is detected by the camera or other visual sensor, the opening and closing movement of the right upper eyelid support is triggered according to the user's eye state. For example, when the user closes his eyes or opens his eyes, the corresponding action of the right upper eyelid support is triggered.

[0245] Emotion simulation instruction: according to the emotion recognition algorithm, the emotional signal in the environment is analyzed. For example, when the user's tired, surprised or happy emotion is detected, the opening and closing movement of the right upper eyelid support is triggered to simulate the corresponding eye expression change.

[0246] The body assembly 331 includes an eyeball profiling shell 3315, an eyeball profiling shell rear cover 3316 and a rear cover central groove locking nut 3317. The visual sensor 3311, the variable focus movable motor seat 3312, the integrated circuit board 3313 and the fixed connecting rod seat rubber ring 3314 are sequentially and fixedly arranged in the eyeball profiling shell 3315. The eyeball profiling shell rear cover 3316 is fixedly arranged at the end of the eyeball profiling shell 3315 and is locked by the rear cover central groove locking nut 3317.

[0247] The visual sensor 3311 can adopt a miniature high-definition camera, which has the functions of automatic exposure, automatic focusing, automatic white balance and automatic light compensation, and is used to capture visual information of the surrounding environment.

[0248] The visual sensor 3311 is part of the visual video tracking system, which captures high-definition video and images and transmits them to the graphics processing unit for processing.

[0249] The eye movement assembly 332 is used to control the movement of the eyeball, enabling it to rotate in horizontal and vertical directions, thereby changing the direction of the line of sight. The eye movement assembly 332 includes an eyeball support rod 3320, an eyeball connecting rod 3322, and a servo steering engine 3325. The eyeball support rod 3320 is connected to the eyeball 33 and an eyeball connecting block 3324 in the middle, which is installed on the head structure. The servo steering engine 3325 is installed on the eyeball connecting block 3324. The eyeball connecting rod 3322 has a movable joint at both ends. One end is connected to the eyeball profiling shell 3315 through a fixed pin, and the other end is connected to the servo steering engine 3325 through a movable pin. A group of servo steering engines 3325 controls the movement of a group of eyeball connecting rods 3322.

[0250] When several servo steering engines 3325 control the forward movement of the body assembly 331, and the remaining several servo steering engines 3325 control the backward movement of the body assembly 331, the body assembly 331 can achieve upward or downward turning. In this embodiment, by controlling different servo steering engines 3325 to drive different eyeball connecting rods 3322, the eyeball can rotate in horizontal and vertical directions, thereby changing the direction of the line of sight.

[0251] In this embodiment, when the robot detects that a person or object enters its surrounding environment, the visual video tracking system will start capturing images and transmitting them to the perception center for analysis. If the perception center determines that the direction of the line of sight needs to be adjusted to track the target, it sends corresponding control instructions to the eye movement assembly 332 to trigger the movement of the eyeball;

[0252] The trigger of eye movement mainly includes target detection, target importance evaluation, target movement analysis, and the priority and requirements of the current task, which specifically includes the following steps:

[0253] Step A1. Target detection and confirmation:

[0254] 1.1. Detection: The visual video tracking system detects a person or object in the image;

[0255] 1.2. Confirmation: Confirm whether the detected object meets the conditions and characteristics of tracking, such as specific shapes, colors, or known markers.

[0256] Step A2. Target importance and priority evaluation:

[0257] 2.1. Importance: Evaluate the importance of the target, depending on the specific requirements of the task (for example, safety monitoring, interactive tasks, or important objects in specific scenarios).

[0258] 2.2 Priority: In a multi-target environment, based on the dynamic behavior of the target, relevance to the task, or pre-set rules, determine which target has a higher tracking priority.

[0259] Step A3. Target Dynamic Analysis:

[0260] 3.1 Motion Tracking: Analyze the motion trajectory of the target to determine whether it is moving, the speed and direction of movement.

[0261] 3.2 Predict Future Position: Use motion estimation models to predict the future position of the target to determine how to adjust the line of sight most effectively.

[0262] Step A4. Task and Environmental Requirements:

[0263] 4.1 Task Requirements: Based on the current task requirements of the robot (e.g., whether it needs to continuously monitor a certain area or object) or environmental changes (e.g., other moving objects, changes in light, etc.), determine whether to adjust the line of sight.

[0264] Step A5. Line of Sight Adjustment Strategy:

[0265] 5.1 Immediate Adjustment: If the position or speed of the target does not match the predetermined tracking strategy, immediately adjust the direction of the line of sight to keep the target in the center of the field of view or in a suitable position.

[0266] 5.2 Predictive Adjustment: Based on the prediction model of target motion, adjust the direction of the line of sight in advance to reduce reaction delay, improve the continuity and accuracy of tracking.

[0267] The triggering instructions can specifically include

[0268] Visual tracking instructions: Real-time capture of user head movement or facial expressions through cameras or other visual sensors, and trigger synchronous rotation of left and right eyeballs according to user head up, down, left or right movement. When the user's head is raised, lowered, turned left or right, the corresponding movement of the left and right eyeballs is triggered to simulate the change of human line of sight.

[0269] Gesture recognition instructions: The humanoid robot is equipped with gesture recognition technology, which triggers the movement of the eyeballs according to the user's hand movements. For example, when the user's fingers point in a certain direction, the eyeballs move in the corresponding direction.

[0270] Context-aware instructions: Through environmental perception technology such as sound sensors or depth cameras, the sound or object position in the surrounding environment is perceived, triggering asynchronous movement of the left and right eyeballs. For example, when the humanoid robot detects that the sound comes from the left, the left eyeball may turn left, and the right eyeball remains stationary to simulate the direction of human attention.

[0271] Emotion simulation instructions: based on emotion recognition algorithm to analyze the user's emotional changes, and according to the emotional changes to trigger the asynchronous movement of the eyeball. For example, when detecting that the user shows anxious or surprised emotion, the left and right eyeballs may present asynchronous up-down and left-right movement to express the robot's reaction to the user's emotion.

[0272] When the robot needs to change the direction of the line of sight, the perception center will analyze the current environment and determine the target position that needs to be adjusted. Specifically, according to the target position and the current position of the eyeball, the execution unit sends corresponding control signals to the eye movement component 332, so that it adjusts the direction and position of the eyeball, thereby aiming at the target.

[0273] The robot needs to change the direction of the line of sight is related to target tracking, task requirements, and environmental changes,

[0274] Target tracking: deviation of the target position detected by the visual sensor from the expected or planned path;

[0275] Task requirements: for example, when navigating, avoiding obstacles, or interacting with people, the line of sight needs to be changed to adapt to changes in the environment or focus.

[0276] Environmental changes: changes that occur in the environment may require the robot to adjust the line of sight to obtain more information or better understand the surrounding environment.

[0277] Specifically includes the following steps:

[0278] Step B1: determination of the target position;

[0279] The specific position or coordinates of the target are usually determined through the following steps:

[0280] 1.1. Image capture: first, the visual sensor captures the image of the current field of view;

[0281] 1.2. Image processing: identify objects in the image through the signal processing unit;

[0282] 1.3. Object localization: localize the position of the object in three-dimensional space based on a deep learning model;

[0283] 1.4. Coordinate conversion: convert the coordinates of the target object into the position in the global coordinate system according to the coordinates of the robot's own position and orientation;

[0284] Step B2: adjust the eyeball position and line of sight direction;

[0285] The adjustment process includes:

[0286] 2.1. Error calculation: calculate the deviation between the target position and the current line of sight center;

[0287] 2.2. Motion planning: Calculate the angle and direction of eye rotation based on the deviation;

[0288] 2.3. Adjustment execution: Adjust the line of sight direction by controlling the eye movement assembly 332 to make the target located in the center of the field of view or other predetermined position.

[0289] The nose system 4 includes a nose bridge 41 provided with left and right nostrils and an olfactory sensor 42 mounted in the left and right nostrils respectively. The olfactory sensor 42 is an important component for receiving and analyzing odor information in the surrounding environment, which simulates the olfactory perception mechanism in the human nasal cavity, can perceive different odor molecules, and convert them into information that can be understood and responded by the system.

[0290] The olfactory sensor 42 can achieve the following functions:

[0291] Environmental perception: The olfactory sensor 42 can perceive various odors in the surrounding environment, such as flower fragrance, food aroma, smoke smell, etc., which enables the humanoid robot to respond to changes in the surrounding environment in a timely manner.

[0292] Safety protection: The olfactory sensor 42 can also detect dangerous odors such as gas leaks, gas smells or the presence of toxic gases. Once these dangerous odors are detected, the system can trigger appropriate safety protection measures such as stopping robot operation, alarming or shifting to a safe area.

[0293] Health monitoring: By detecting odors in the surrounding environment through the olfactory sensor 42, the humanoid robot can monitor some health-related information such as air quality and environmental pollution level, which helps to protect the health of the robot and human users.

[0294] Task execution: In some specific tasks, odor information may be one of the important information necessary for task execution. For example, in search and rescue tasks, the robot may need to find the location of the trapped person through odor; in the field of agriculture, the robot may need to detect the maturity of crops or the presence of pests and diseases through odor.

[0295] Emotional expression: In some cases, odor can also be used as a way of emotional expression. For example, the robot may be designed to "smell" the fragrance of flowers and express a happy emotion, or "smell" the smoke and express a worried or nervous emotion.

[0296] The olfactory sensor 42 not only provides environmental odor information, but also converts these information into meaningful data for the robot, helping the robot better understand and adapt to its surrounding environment. Through the olfactory sensor, the bionic humanoid robot can achieve more intelligent and humanized interaction and task execution, enhancing its applicability and practicality in various application scenarios.

[0297] The ear system 5 includes ears 51 and auditory sensors 52, the ears 51 are provided in two groups, symmetrically located on both sides of the head structure, and the auditory sensors 52 are fixedly arranged on the ears 51 to receive the sound in the surrounding environment. The signal processing unit receives and processes the sound to obtain a sound signal that can be recognized by the system, and the sound signal is transmitted to the sensory center.

[0298] The step of receiving and processing the sound by the signal processing unit is:

[0299] Step 1: Sound capture: The auditory sensor 52 captures the sound signal in the surrounding environment and converts the sound wave into an electrical signal;

[0300] Step 2: Preprocessing: Denoising and filtering processing are performed on the electrical signal to eliminate environmental noise and enhance the clarity and distinguishability of the sound;

[0301] Step 3: Extracting sound features: Extracting sound features from the preprocessed signal;

[0302] The sound features include frequency, amplitude, sound duration, etc., which are helpful for analyzing and recognizing the sound;

[0303] Step 4: Sound recognition: Using machine learning algorithms or pattern recognition techniques to analyze and compare the sound features to identify the type and meaning of the sound. The trained model can recognize different sound categories, such as voice instructions, environmental noise, alarm sounds, etc., and classify and process them.

[0304] Step 5: Action triggering: The sensory center sends instructions to the execution unit according to the identified sound type and meaning, and the execution unit triggers the corresponding action or response.

[0305] For example, if the user's voice instruction is recognized, the execution unit may perform the corresponding task or action; if an emergency alarm sound is recognized, the execution unit may take appropriate safety measures or issue a warning signal.

[0306] Different sounds can trigger different actions or responses, and the specific triggering conditions depend on the design of the system and the preset task requirements. For example: the user speaks a specific voice instruction, which can trigger the robot to perform corresponding operations, such as moving to a specific location, performing a specific task, etc.

[0307] When an alarm sound or abnormal sound is captured, the system may trigger the robot to take emergency measures, such as stopping the current task and moving to a safe location.

[0308] After the ear sensor 52 of the humanoid robot receives the sound, the sound is converted into understandable information through the above processing steps, and the corresponding action or response is triggered according to the recognition result to realize the interaction with the environment and the task execution.

[0309] The voice instruction and the voice interaction instruction in the embodiment are captured by the hearing sensor, processed by the signal processing unit, and sent to the execution unit by the perception center. The visual tracking / identification instruction and the emotion simulation instruction are captured by the visual sensor, processed by the signal processing unit, and sent to the execution unit by the perception center. The emotion expression instruction is sent to the execution unit by the perception center.

[0310] Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor shall belong to the scope of protection of the present application.

Claims

1. A biomimetic service humanoid robot head structure, characterized in that it comprises a skull structure, a sensory system, a mouth system connected to the sensory system, a vision system, a nose system, and an ear system, the sensory system comprising a perception center, a sensor network, a signal processing unit, and an execution unit, the sensor network being connected to the signal processing unit, the signal processing unit being connected to the perception center, and the perception center being connected to the execution unit; The sensor network comprises a plurality of sensors that capture information about the external environment. The signal processing unit receives raw signals from the sensor network, processes and pre-processes them, and transmits the processed signals to the perception center. The perception center receives the processed signals, analyzes and makes decisions, and obtains the state of the external environment and the required response. The execution unit receives instructions from the perception center and performs corresponding actions or tasks. The sensor network includes a vision sensor, and the signal processing unit includes a central processing unit and a graphics processing unit. The graphics processing unit and the vision sensor form a visual tracking system. The vision sensor captures visual information about the surrounding environment and transmits it to the graphics processing unit for analysis, recognition, and tracking. The graphics processing unit has built-in image processing algorithms. The vision sensor captures continuous image frames, and the image processing algorithms identify and extract the required images or objects. The specific steps are as follows: Step S1: The vision sensor captures real-time images or videos in the environment and transmits them to the graphics processing unit. Step S2: The image processing algorithm pre-processes the captured images. Step S3: Feature extraction: extract specific information from the pre-processed images. The feature extraction methods include edge detection based on edge detection algorithms, corner detection based on corner detection algorithms, and texture analysis based on texture analysis algorithms. Step S4: Object recognition. Use a set of machine learning models to identify and classify objects in the image. The mathematical model of the image recognition algorithm based on neural network multi-model feature fusion integrates image features from multiple sensors. Specifically: Assume there are n sensors, each of which can capture images and extract image features. Let the image features captured by the i-th sensor be represented as where i∈{1,2,...,n}; integrate these image features into a feature vector, denoted as x∈R m where A deep neural network is established for image recognition, and its structure is as follows: h = σ(W1x + b1) y = softmax(W2h + b2) where y is the model's predicted class probability distribution, h represents the hidden layer bias vector, W1 e R h×m and W2 e R k×h denote the weight matrices from the input layer to the hidden layer and from the hidden layer to the output layer, respectively, R h×m denotes the weight matrix from the input layer to the hidden layer, R k ×h denotes the weight matrix from the hidden layer to the output layer, b1 e R h and b2 e R k denote the bias vectors of the hidden layer and the output layer, respectively, k denotes the output layer bias vector, R h denotes the weight matrix of the input layer, R k denotes the weight matrix of the hidden layer; σ() denotes the activation function, and softmax() denotes the softmax function. To train the neural network, a cross-entropy loss function is used: where y is the model's predicted class probability distribution, is the true class label, and j denotes the class index. The loss function is minimized by gradient descent, i.e., updating the weights and biases: where a is the learning rate, are the gradients of the loss function with respect to the weight matrices W1, W2 and the bias vectors b1, b2, respectively. A deep convolutional neural network (CNN) is used for image recognition; the CNN contains multiple convolutional layers, pooling layers, and fully connected layers for feature extraction and classification; let the input of the CNN be x and the output be y e R k , which represents the predicted class result of the image. The output y of the CNN can be represented as: h1 = ReLU(W1x + b1) h2 = MaxPooling(h1) h3 = ReLU(W2h2 + + b2) h4 = MaxPooling(h3) h5 = ReLU(W3h4 + b3) h6 = Flatten(h5) y = Softmax(W4h6 + b4) Wherein, b3, b4 are bias vectors, W3, W4 are weight matrices, ReLU() represents a rectified linear unit activation function, MaxPooling() represents a maximum pooling operation, Flatten() represents flattening multi-dimensional data into a one-dimensional vector, and Softmax() represents a softmax function. A cross-entropy loss function is used to measure the difference between the model prediction result and the true label: Wherein, y is the class probability distribution predicted by the model, is the true class label, k represents the number of classes, and j is the initial parameter value of the model. The loss function is minimized by gradient descent, that is, the weights and biases are updated: Wherein, a is the learning rate, and respectively represent the gradient of the loss function with respect to the weight matrix W1 and the bias vector b1. Step S5: object positioning and tracking; After object recognition, the specific position of the object in the image is determined based on the bounding box, and the bounding box frames each recognized object in the image to position it; Step S6: output result; The graphics processing unit outputs the recognition and positioning results to the perception center, which analyzes the results and determines whether to generate a corresponding decision. If so, adjust the parameters and execute step S2 to process the newly captured image. If not, adjust the model, call the deep learning model, and execute step S4. When the robot detects that a person or object enters its surrounding environment, the visual video tracking system starts capturing images and transmits them to the perception center for analysis. If the perception center determines that the line of sight needs to be adjusted to track the target, it sends corresponding control instructions to the eye movement component to trigger eye movement. The triggering of eye movement mainly includes target detection, target importance assessment, target movement analysis, and the priority and requirements of the current task, which specifically includes the following steps. Step A1. Target detection and confirmation: 1.

1. Detection: The visual video tracking system detects people or objects in the image. 1.

2. Confirmation: Confirm whether the detected object meets the conditions and characteristics of tracking; Step A2. Target importance and priority assessment: 2.

1. Importance: Assess the importance of the target, depending on the specific requirements of the task. 2.2 Priority: In a multi-target environment, based on the dynamic behavior of the target, relevance to the task, or preset rules, determine which target has a higher tracking priority; Step A3. Target dynamic analysis: 3.1 Motion tracking: Analyze the motion trajectory of the target to determine whether it is moving, the speed and direction of movement; 3.2 Predict future position: Use motion estimation models to predict the future position of the target to determine how to adjust the line of sight most effectively; Step A4. Task and environmental requirements: 4.1 Task requirements: Determine whether to adjust the line of sight based on the current task requirements of the robot or changes in the environment; Step A5. Line of sight adjustment strategy: 5.1 Instant adjustment: If the position or speed of the target does not match the predetermined tracking strategy, adjust the direction of the line of sight to keep the target in the center of the field of view or in a suitable position; 5.2 Predictive adjustment: Adjust the direction of the line of sight in advance based on the prediction model of the target motion; When the robot needs to change the direction of the line of sight, the perception center analyzes the current environment, determines the target position that needs to be adjusted based on the target position and the current eye position, and the execution unit sends control signals to the eye movement component 332 to adjust the direction and position of the eye; The need for the robot to change the direction of the line of sight is related to target tracking, task requirements, and environmental changes. Target tracking: deviation of the target position detected by the visual sensor from the expected or planned path; Task requirements: Change the line of sight to adapt to changes in the environment or focus when navigating, avoiding obstacles, or interacting with people; Environmental changes: Changes in the environment may require the robot to adjust the line of sight to obtain more information, Specifically includes the following steps: Step B1: Determination of target position; The specific position or coordinates of the target are determined by the following steps: 1.

1. Image capture: First, the visual sensor captures the image of the current field of view; 1.

2. Image processing: Identify objects in the image through the signal processing unit; 1.

3. Object localization: Localize the position of the object in three-dimensional space based on a deep learning model; 1.

4. Coordinate conversion: convert the coordinates of the target object into a global coordinate system according to the coordinates of the robot's own position and orientation; Step B2: adjust the eye position and line of sight direction; The adjustment process includes: 2.

1. Error calculation: calculate the deviation between the target position and the current line of sight center; 2.

2. Motion planning: calculate the angle and direction of the eye rotation according to the deviation; 2.

3. Execute adjustment: adjust the line of sight direction by controlling the eye movement assembly 332 to make the target located in the field of view center or other predetermined position.

2. The bionic service-type humanoid robot head structure according to claim 1, characterized in that: The sensory system also includes a sensory central anti-electromagnetic interference system and a sensory central storage chip family. The sensory central anti-electromagnetic interference system realizes electromagnetic interference protection, and the sensory central anti-electromagnetic interference system covers the sensitive components with conductive or magnetic materials. The sensory central storage chip family is used for storing the memory and storage devices of the data collected from the sensor network, including random access memory, fast access memory and solid state disk. The fast access memory is used for temporarily storing data in processing, The solid state disk is used for long-term data storage.

3. The bionic service-type humanoid robot head structure according to claim 1, characterized in that: The head structure is provided with an inspection hatch, a storage chip card slot and a hatch. The hatch is detachably installed in the storage chip card slot, the sensory central storage chip family is installed in the storage chip card slot, and the inspection hatch is hinged to the head structure.

4. The bionic service-type humanoid robot head structure according to claim 1, characterized in that: The head structure is provided with an oral space, and the oral system is installed in the oral space. The oral system includes a jaw structure, a tooth structure, a lip structure and a tongue structure, which simulates the function and perception process of human oral cavity.

5. The bionic service-type humanoid robot head structure according to claim 4, characterized in that: The tongue structure also includes a taste perception system that simulates human taste perception ability. The taste perception system includes a plurality of high-temperature-resistant taste sensors distributed on the surface of the tongue fetus, which are used to perceive five tastes. The high-temperature-resistant taste sensors integrate the taste sensor module and convert it into an electrical signal. The tongue section structure is distributed with a taste signal line, and the electrical signal is transmitted to the signal processing unit through the taste signal line for processing. The signal processing unit converts the original signal into a digital signal and performs filtering and processing. The processed signal is sent to the perception center, which analyzes, decodes and discriminates the signal, identifies the characteristic pattern of the five basic tastes, and compares it with the pre-stored pattern to determine the type and degree of the taste.

6. The bionic service-type humanoid robot head structure according to claim 1, characterized in that: The visual system is installed in the head structure, including an eye shell and a brow structure, an eyelid structure and an eyeball installed in the eye shell. The brow structure includes a brow movement support and two groups of brow movement assemblies, which are respectively connected with the brow movement support. The brow movement assemblies control the up and down movement of the brow movement support. The eyelid structure includes an upper eyelid support, an upper eyelid, a lower eyelid, a lower eyelid support and an eyelid movement assembly. The eyelid movement assembly controls the opening and closing movement of the upper eyelid support and the lower eyelid support. The eyeball includes a body assembly and an eye movement assembly. The body assembly is movably installed in the middle of the eyelid structure, and the eye movement assembly controls the up and down movement and rotation of the body assembly.

7. The bionic service-type humanoid robot head structure according to claim 1, characterized in that: The nose system includes a nose bridge and olfactory sensors, the nose bridge is provided with left and right nostrils, and the olfactory sensors are respectively installed in the left and right nostrils. The olfactory sensors receive and analyze odor information in the surrounding environment, simulate the olfactory perception mechanism in the human nasal cavity, perceive different odor molecules, and convert them into information that can be understood and responded to by the system. The ear system includes ears and auditory sensors, the ears are provided with two groups, symmetrically located on both sides of the head structure, and the auditory sensors are fixedly arranged on the ears for receiving sound in the surrounding environment. The signal processing unit receives and processes the sound to obtain a sound signal that can be recognized by the system, and the sound signal is transmitted to the sensory center.

8. The bionic service-type humanoid robot head structure according to claim 7, characterized in that: The signal processing unit receives and processes the sound in the following steps: Step 1: sound capture: the auditory sensor captures the sound signal in the surrounding environment and converts the sound wave into an electrical signal; Step 2: Preprocessing: denoising and filtering processing of the electrical signal; Step 3: extracting sound features: extracting sound features from the preprocessed signal; Step 4: sound recognition: using machine learning algorithms or pattern recognition techniques to analyze and compare the sound features to identify the type and meaning of the sound, training the model to recognize different sound categories and classify and process them; Step 5: action trigger: the sensory center sends instructions to the execution unit according to the identified sound type and meaning, and the execution unit triggers the corresponding action or response.

Citation Information

Patent Citations

  • Intelligent cognitive robot and cognitive system thereof

    CN104493827A

  • Image sensor control method, device and system

    CN111756990A

  • Bionic robot head and neck structure and bionic robot

    CN117124343A

  • Intelligent robot control system for accompanying old people

    CN118034152A

  • Bionic service type humanoid robot head structure

    CN222857996U