Bionic service type humanoid robot head structure
By designing a humanoid robot head structure that includes multiple sensory systems and simulates human oral functions, the problem of limited feedback capabilities of humanoid robots in the prior art and inability to automatically recognize the taste of food is solved, and more efficient external interaction and automated evaluation is achieved.
Patent Information
- Application Number
- CN202421774591.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2034-07-25
AI Technical Summary
Existing humanoid robots have limited feedback capabilities when the external environment changes and cannot automatically recognize the taste of food.
A bionic service humanoid robot head structure is designed, including the head structure, sensory system, oral system, visual system, nose system and ear system. The sensory system realizes perception and feedback to the external environment through the perception center, sensor network, signal processing unit and execution unit. The oral system includes a jaw structure, a tooth structure, a lip structure and a tongue structure, and has the ability to simulate human oral functions.
By enhancing the perception of the external environment, humanoid robots can promptly and effectively feedback environmental changes, and automatically evaluate the taste of food by simulating the human taste perception system, improving the ability to interact with the outside world and the efficiency of automated evaluation.
Smart Images

Figure CN222857996U_ABST
Abstract
Description
Technical Field
[0001] The utility model belongs to the field of robot bionics and relates to a head structure of a bionic service-type humanoid robot. Background Art
[0002] A "humanoid robot" can be defined as a robot that has certain attributes of human appearance and functionality (e.g., torso, head, arms, legs), the ability to communicate verbally with humans using speech recognition and sound synthesis, etc. This type of robot aims to reduce the cognitive distance between humans and machines.
[0003] Existing "humanoid robots" have limited ability to interact with the outside world and are unable to provide timely feedback based on changes in the external environment. In addition, in places where automatic identification of the taste of food is required, humanoid robots are unable to complete the corresponding tasks. Summary of the invention
[0004] In order to overcome at least one deficiency of the prior art, the utility model provides a bionic service-type humanoid robot head structure.
[0005] In order to achieve the above-mentioned purpose, the utility model adopts the following technical scheme: a bionic service humanoid robot head structure, including a skull structure, a sensory system, an oral system connected to the sensory system, a visual system, a nasal system and an ear system, the oral system, the visual system, the nasal system and the ear system are installed on the skull structure, the oral system includes a jaw structure, a tooth structure, a lip structure and a tongue structure, the visual system includes an eyebrow structure, an eyelid structure and an eyeball, the nasal system includes a nose bridge and an olfactory sensor installed on the nose bridge, and the ear system includes an ear and an auditory sensor installed on the ear.
[0006] Furthermore, the jaw structure includes an upper jaw shell, a jaw movable component and a lower jaw shell. The left and right ends of the upper jaw shell and the lower jaw shell are rotatably connected through upper and lower jaw fixed shafts, and the upper and lower jaw fixed shafts are installed on the skull structure.
[0007] Furthermore, the jaw movable assembly includes a brushless servo, a jaw rear connecting rod structure and a jaw front connecting rod structure, the jaw rear connecting rod structure and the jaw front connecting rod structure are respectively connected to the brushless servo, the jaw rear connecting rod structure and the jaw front connecting rod structure respectively include an upper movable joint, a middle connecting rod and a lower movable joint, the upper movable joint and the lower movable joint are connected to the two ends of the middle connecting rod, the upper movable joint is connected to the fixing pin of the upper jaw shell, the lower movable joint is connected to the fixing pin of the lower jaw shell, and the brushless servo is connected to the upper movable joint.
[0008] Furthermore, the tooth structure includes an upper gum, upper teeth, lower teeth and a lower gum, the upper gum and the lower gum are fixed to the skull structure, the upper teeth are fixed to the upper gum, the lower teeth are fixed to the lower gum, the upper gum is fixed to the upper jaw shell, and the lower gum is fixed to the lower jaw shell.
[0009] Furthermore, the lip structure includes an upper lip, a lip movable assembly and a lower lip. The upper lip is movably connected to the upper jaw shell, the lower lip is movably connected to the lower jaw shell, the lip movable assembly is connected to the upper lip and the lower lip, and the up and down movement and left and right movement of the upper lip and the lower lip are controlled by the lip movable assembly.
[0010] Furthermore, the lip movable assembly includes a servo servo, a movable connecting rod and a movable telescopic rod. The servo servo is installed on the upper jaw shell, the movable connecting rod is connected to the servo servo, one end of the movable telescopic rod is connected to the movable connecting rod, and the other end of the movable telescopic rod is rotatably connected to the upper lip.
[0011] Furthermore, the oral system is provided with an oral space, a mounting frame is fixedly provided in the oral space, the tongue structure is installed on the mounting frame, the tongue structure includes a controllable telescopic servo, a tongue tire and a tongue segment structure, the tongue segment structure is composed of a plurality of tongue segments, the plurality of tongue segments are installed in sequence and connected through connecting pipelines, one end of the tongue segment structure is connected to the mounting frame and to the controllable telescopic servo, the other end of the tongue segment structure is a free end, and the tongue tire is nested on the outside of the tongue segment structure.
[0012] Furthermore, the tongue structure also includes a taste perception system, which includes a plurality of high-temperature resistant taste sensors distributed on the surface of the tongue.
[0013] Furthermore, the sensory system includes a perception center, a sensor network, a signal processing unit and an execution unit. The sensor network is connected to the signal processing unit, the signal processing unit is connected to the perception center, and the perception center is connected to the execution unit. The sensory system also includes a sensory central electromagnetic interference protection system and a sensory central storage chip family. The sensory central electromagnetic interference protection system realizes electromagnetic interference protection. The sensory central electromagnetic interference protection system uses conductive or magnetic materials to cover sensitive components. The sensory central storage chip family is a memory and storage device for storing data collected from the sensor network, including random access memory, fast access memory and solid state hard disk. The fast access memory is used to temporarily store data in processing, and the solid state hard disk is used for long-term data storage.
[0014] Furthermore, the head structure is provided with an inspection hatch, a memory chip card slot and a hatch, the hatch can be detachably installed in the memory chip card slot, the sensory central memory chip cluster is installed in the memory chip card slot, and the inspection hatch is hinged to the head structure.
[0015] In summary, the utility model is beneficial in that:
[0016] The utility model can make timely and effective feedback to the changes of the external environment by forming a sensory system through a perception center, a sensor network, a signal processing unit and an execution unit, thereby greatly improving the ability to interact with the outside world; at the same time, according to the taste sensor, the humanoid robot can test and evaluate food without a human experimenter, and the taste perception system can simulate the function of the human taste perception system, thereby automatically evaluating the taste of the food. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the head structure of the utility model.
[0018] Figure 2 It is a top view of the head knot of the utility model.
[0019] Figure 3 Schematic diagram of the jaw structure and tooth structure of the utility model Figure 1 .
[0020] Figure 4 Schematic diagram of the jaw structure and tooth structure of the utility model Figure 2 .
[0021] Figure 5 Schematic diagram of the tooth structure and lip structure of the utility model Figure 1 .
[0022] Figure 6 Schematic diagram of the tooth structure and lip structure of the utility model Figure 2 .
[0023] Figure 7 Schematic diagram of the tooth structure and lip structure of the utility model Figure 3 .
[0024] Figure 8 Schematic diagram of the visual system of the utility model Figure 1 .
[0025] Fig. 9 Schematic diagram of the visual system of the utility model Figure 2 .
[0026] Fig.10 Schematic diagram of the visual system of the utility model Figure 3 .
[0027] Fig.11 It is a schematic diagram of an eyeball of the utility model.
[0028] Fig.12 It is a schematic diagram of the eyeball decomposition of the utility model.
[0029] Fig.13 It is a schematic diagram of the nasal system and the ear system of the utility model. DETAILED DESCRIPTION
[0030] The following describes the implementation of the present invention through specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific implementations, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0031] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed at will, and the component layout may also be more complicated.
[0032] All directional indications in the embodiments of the present invention (such as up, down, left, right, front, back, horizontal, vertical...) are only used to explain the relative position relationship, movement status, etc. between the components in a certain specific posture. If the specific posture changes, the directional indication will also change accordingly.
[0033] Due to installation errors and other reasons, the parallel relationship referred to in the embodiments of the present invention may actually be an approximately parallel relationship, and the vertical relationship may actually be an approximately vertical relationship.
[0034] Embodiment 1:
[0035] like Figure 1-Figure 13 As shown, the head structure of the bionic service humanoid robot includes a skull structure 1 and a sensory system, an oral system 2, a visual system 3, a nasal system 4 and an ear system 5 which are integrated and installed in the skull structure.
[0036] The sensory system includes a perception center, a sensor network, a signal processing unit and an execution unit. The sensor network signal is connected to the processing unit, the signal processing unit is connected to the perception center, and the perception center is connected to the execution unit.
[0037] The sensor network includes several sensors, including visual sensors, auditory sensors, olfactory sensors, taste sensors and other types of sensors, so as to capture various information of the external environment, such as sound, image, smell, etc. The sensor network transmits the captured information to the signal processing unit. The sensor network provides various information of the external environment and is the input source of the perception center.
[0038] The signal processing unit is used to receive the original signal from the sensor network, perform signal processing and preprocessing, enhance the accuracy and stability of the signal, and then transmit the processed signal to the perception center to provide more reliable information for the perception center.
[0039] As the intelligent core of the humanoid robot, the perception center receives processed information, performs analysis and decision-making, and determines the state of the external environment and the required response to achieve perception of the surrounding environment and response to external stimuli. The perception center is interconnected with the oral system 2, the visual system 3, the nasal system 4 and the ear system 5 to achieve overall perception and decision-making.
[0040] The Perception Center uses the high-performance processors of NVIDIA's Project GR00T and Isaac robotics platforms, and uses NVIDIA's unique SoC (system-on-chip) technology to optimize the design for the needs of complex robotic systems to achieve highly automated and intelligent operations. Specifically, it has the following functions:
[0041] 1. High performance computing support:
[0042] The main purpose of NVID IA's SoC design is to provide sufficient processing power to support the operation of AI algorithms and models in robots. This chip integrates an efficient GPU (graphics processing unit), which enables robots to quickly process visual and perception data and perform complex image recognition, object detection and environmental understanding in real time. This is because GPUs are designed for parallel processing of large amounts of data and are very suitable for deep learning tasks.
[0043] 2. Multimodal sensor fusion:
[0044] In humanoid robot applications, data from multiple sensors needs to be integrated and processed, including vision, hearing, touch, and position perception. NVIDIA's SoC supports this advanced sensor fusion, enabling robots to more accurately understand and adapt to their environment. For example, a robot can simultaneously process visual data from a camera and force feedback from a sensor to achieve more precise object manipulation and navigation.
[0045] 3. Low latency and real-time response:
[0046] Humanoid robots require extremely low response times when performing tasks such as delivery, rescue, or collaborative work. NVIDIA SoCs ensure low latency and high-speed data processing capabilities by optimizing computing paths and improving data transmission efficiency. This allows robots to respond quickly in dynamic and unpredictable environments.
[0047] 4. Energy efficiency:
[0048] Considering that humanoid robots need to run for a long time under battery power, energy efficiency becomes a key factor when designing SoCs. While maintaining high performance, NVID IA's chips also optimize energy consumption through advanced manufacturing processes and power management technologies, extending the robot's operating time.
[0049] 5. Collaborative optimization of software and hardware:
[0050] The Isaac robot platform includes not only hardware SoC, but also a complete software development kit (SDK), including simulators, development tools, and pre-trained AI models. These software tools are tightly integrated with the SoC hardware, enabling developers to customize and optimize robot behavior for specific application scenarios, thereby improving development efficiency and robot performance.
[0051] The execution unit receives instructions from the perception center and performs corresponding actions or tasks, such as moving, operating external devices, etc.
[0052] The execution unit can adopt servo motor, electromagnetic brake, linear actuator, sensor and brake, etc.
[0053] The signal processing unit is the core processing unit that processes all sensory data, and specifically includes the central processing unit (CPU) and the graphics processing unit (GPU);
[0054] The central processing unit is responsible for processing program instructions, managing software operations and other computing tasks, and the graphics processing unit (GPU) is used for image analysis and machine vision, such as NVIDIA's Jetson series.
[0055] The central processing unit (CPU) can use conventional NVIDIA Jetson AGX Xavier, Intel Core i7-1185G7, AMD Ryzen 9 5900HX, Qualcomm Snapdragon 888, Apple M1, etc.
[0056] The graphics processing unit GPU can use conventional NVIDIA GeForce RTX 3080, AMD Radeon RX 6800XT, Intel Iris Xe Graphics G7, Qualcomm Adreno 660, Apple M1 GPU, etc.
[0057] The graphics processing unit and the visual sensor constitute a visual video tracking system. The visual sensor captures the visual information of the surrounding environment and transmits it to the graphics processing unit for analysis, recognition and tracking.
[0058] The graphics processing unit has a built-in image processing algorithm. The visual sensor captures continuous image frames and then uses the image processing algorithm to identify and extract the required image or object. The specific steps include:
[0059] Step S1: The visual sensor captures real-time images or videos in the environment and transmits them to the graphics processing unit;
[0060] Step S2: The image processing algorithm pre-processes the captured image;
[0061] Preprocessing includes resizing, normalization, denoising, etc. to enhance the image to improve the effect and efficiency of subsequent processing steps;
[0062] Step S3: Feature extraction: extracting specific information from the preprocessed image can help identify the objects in the image;
[0063] The feature extraction methods include edge detection based on edge detection algorithm, corner detection based on corner detection algorithm, texture analysis based on texture analysis algorithm, etc.
[0064] Step S4: object recognition;
[0065] Identify and classify objects in images using a given machine learning model;
[0066] Machine learning models have been trained on large amounts of labeled data to recognize different types of objects;
[0067] Step S5: object positioning and tracking;
[0068] After object recognition, the specific location of the object in the image is determined based on bounding boxes, which frame each recognized object in the image to locate it;
[0069] Step S6: output the result;
[0070] The graphics processing unit outputs the results of recognition and positioning to the perception center, which analyzes the results to determine whether to generate a corresponding decision. If so, it adjusts the parameters and executes step S2 to process the newly captured image. If not, it adjusts the model and calls a deep learning model (such as a convolutional neural network) to execute step S4.
[0071] Step S4 object recognition is achieved through machine learning and computer vision technology. Common methods include:
[0072] Object classification: Classify objects in an image into predefined categories.
[0073] Object Detection: Detecting object locations and bounding boxes in an image.
[0074] Object tracking: Tracking the motion trajectory of a specific object in consecutive image frames.
[0075] ——The actions that can be triggered according to different images depend on pre-defined programs or algorithms. For example, if the image recognition system detects a face, the robot head may turn and face the face; if a specific object is detected, the robot may perform tasks or actions related to the object, such as grasping, moving, etc.
[0076] This embodiment proposes a mathematical model of an image recognition algorithm based on neural network multi-model feature fusion, integrating image features from multiple sensors to improve the accuracy of image recognition. Specifically:
[0077] Assuming the number of sensors is n, each sensor can capture images and extract image features. Let the image features captured by the i-th sensor be expressed as Where i∈{1, 2, ..., n}. These image features are integrated into a feature vector, represented as x∈R m ,in
[0078] A deep neural network is established for image recognition, and its structure is as follows:
[0079] h=σ(W1x+b1)
[0080] y=softmax(W2h+b2)
[0081] Among them, is the category probability distribution predicted by the model, h represents the hidden layer bias vector, W1∈R h×m and W2∈R k×h Represents the weight matrices from the input layer to the hidden layer and from the hidden layer to the output layer, R h×m Represents the weight matrix from the input layer to the hidden layer, R k×h Represents the weight matrix from the hidden layer to the output layer, b1∈R k and b2∈R k Represent the bias vectors of the hidden layer and the output layer respectively, k represents the bias vector of the output layer, R h represents the weight matrix of the input layer, R k represents the weight matrix of the hidden layer; σ() represents the activation function, and softmax() represents the softmax function.
[0082] To train the neural network, the cross entropy loss function is used:
[0083]
[0084] Among them, y is the category probability distribution predicted by the model, is the true category label, and j represents the category index.
[0085] Minimize the loss function by gradient descent, that is, update the weights and biases:
[0086]
[0087] Among them, α is the learning rate, They are the gradients of the loss function with respect to the weight matrices W1, W2 and the bias vectors b1, b2 respectively.
[0088] Use a deep convolutional neural network (CNN) for image recognition. CNN contains multiple convolutional layers, pooling layers, and fully connected layers to extract image features and perform classification. Let the input of CNN be x and the output be y∈R k , represents the category prediction result of the image.
[0089] The output y of CNN can be expressed as:
[0090] h1=ReLU(W1x+b1)
[0091] h2=MaxPooling(h1)
[0092] h3=ReLU(W2h2+b2)
[0093] h4=MaxPooling(h3)
[0094] h5=ReLU(W3h4+b3)
[0095] h6=Flatten(h5)
[0096] y=Softmax(W4h6+b4)
[0097] Among them, b3 and b4 are bias vectors, W3 and W4 are weight matrices, ReLU() represents the rectified linear unit activation function, MaxPooling() represents the maximum pooling operation, Flatten() represents flattening multidimensional data into a one-dimensional vector, and Softmax() represents the softmax function.
[0098] The cross entropy loss function is used to measure the difference between the model prediction and the true label:
[0099]
[0100] Among them, y is the category probability distribution predicted by the model, is the true category label, k represents the number of categories, and j is the starting parameter value of the model.
[0101] Minimize the loss function by gradient descent, that is, update the weights and biases:
[0102]
[0103] Among them, α is the learning rate, and They represent the gradients of the loss function with respect to the weight matrix W1 and the bias vector b1 respectively.
[0104] Calculation example:
[0105] Assume that there are three sensors, capturing the color features, texture features, and shape features of the image respectively. The output feature vector of each sensor is x1∈R 10 x1∈R10、x2∈R 15 x2∈R15 and x3∈R 20 x3∈R20.
[0106] The goal is to use a deep convolutional neural network to identify the category of the image, assuming there are 5 categories.
[0107] First, initialize the parameters of the neural network. Assume that the hidden layer contains 20 neurons and the learning rate is 0.01.
[0108] [W1∈R 20×(10+15+20) , W2∈R 5×20 ][b1∈R 20 , b2∈R 5 ]
[0109] Next, prepare the training data. Assume there are 1000 samples, and the label of each sample is represented by one-hot encoding. Stochastic gradient descent will be used to update the parameters, using one sample at a time.
[0110] Then, you can train and predict the model. Suppose you have completed 100 rounds of iterative training and want to make a classification prediction for a new image.
[0111] Finally, the prediction results will be output and the accuracy of the model will be calculated.
[0112] Operation process:
[0113] Assume that the hidden layer contains 20 neurons and the learning rate is 0.01. The initialization parameters are as follows:
[0114] [W1∈R 20×45 , W2∈R 5×20 \]\[b1∈R 20 , b2∈R 5 \]
[0115] Prepare training data: Assume there are 1000 samples and the label of each sample is represented by one-hot encoding.
[0116] Use stochastic gradient descent to update parameters, using one sample at a time for training. Complete 100 rounds of iterative training.
[0117] Based on the above premise, suppose there is a new image that needs classification prediction. The output result is category 3 with an accuracy of 80%.
[0118] The above-mentioned sound sensor can adopt conventional Knowles SPH0641LU4H-1 MEMS microphone, PULAudio POM-3245L-CR microphone, Infineon IM69D130 MEMS microphone, STMicroelectronics MP23ABO2B MEMS microphone, TDK InvenSense ICS-41350 MEMS microphone, etc.
[0119] The visual sensor may adopt conventional Sony IMX477 CMOS image sensor, OmniVision OV5670 CMOS image sensor, ON Semiconductor AR0144 CMOS image sensor, Samsung S5K4H7YX CMOS image sensor, Canon 5D Mark IV CMOS image sensor, etc.
[0120] The sensory system also includes a sensory central anti-electromagnetic interference system and a sensory central storage chip family. The sensory central anti-electromagnetic interference system implements electromagnetic interference (EMI) protection. The sensory central anti-electromagnetic interference system uses EMI shielding materials and covers sensitive components with conductive or magnetic materials to reduce the impact of external electromagnetic waves; the sensory central storage chip family is used to store memory and storage devices for data collected from various sensor networks, including random access memory (RAM): fast-access memory for temporary storage of data in processing and solid-state drive (SSD): used for long-term data storage with fast read and write speeds and high durability.
[0121] The sensory central storage chip family can use conventional Samsung PM9A3 E1.S SSD, Western Digital WD Black SN850 NVMe SSD, SK Hynix Gold P31 NVMe SSD, Crucial P5 Plus NVMe SSD, Intel Optane SSD 905P U.2 SSD, etc.
[0122] An inspection hatch 11 is provided on the rear side of the head structure 1, and the inspection hatch 11 is hinged to the head structure 1. Such hinged structure allows the inspection hatch 11 to be opened and stay at a certain angle, so as to facilitate access to the internal structure of the head structure 1. The inspection hatch 11 facilitates replacement and maintenance of the robot's sensory system. The design of the inspection hatch 11 enables technicians to easily replace or upgrade these components, as well as perform routine maintenance and fault diagnosis.
[0123] A memory chip card slot and a hatch 12 are also provided on the rear side of the skull structure 1. The hatch 12 can be detachably installed in the memory chip card slot, and the sensory central memory chip family is installed in the memory chip card slot.
[0124] The skull structure 1 is provided with an oral space, and the oral system 2 is installed in the oral space. The oral system 2 includes a jaw structure 20, a tooth structure, a lip structure 21 and a tongue structure 23. The oral system 2 is connected to the sensory system to simulate the function and perception process of the human oral cavity.
[0125] The jaw structure 20 includes an upper jaw shell 201, a jaw movable component and a lower jaw shell 202. The left and right ends of the upper jaw shell 201 and the lower jaw shell 202 are rotatably connected by upper and lower jaw fixed shafts 203. The upper jaw shell 201 and the lower jaw shell 202 can rotate in the up and down directions relative to the upper and lower jaw fixed shafts 203. The upper and lower jaw fixed shafts 203 are installed on the skull structure 1, thereby realizing the upper and lower jaw shells 201 and the lower jaw shell 202 relative to the skull structure 1. Movement.
[0126] The upper jaw shell 201 and the lower jaw shell 202 are made of ABS material.
[0127] There are two groups of jaw movable components, which are symmetrically arranged on the left and right sides. The jaw movable components connect the upper jaw shell 201 and the lower jaw shell 202. The jaw movable components receive the control signal sent by the execution unit to control the upper jaw shell 201 and the lower jaw shell 202 to perform corresponding movements.
[0128] The jaw movable assembly includes a brushless servo 204, a jaw rear connecting rod structure 205 and a jaw front connecting rod structure 206, the jaw rear connecting rod structure 205 and the jaw front connecting rod structure 206 are respectively connected to the brushless servo 204, the jaw rear connecting rod structure 205 and the jaw front connecting rod structure 206 respectively include an upper movable joint 2051, a middle connecting rod 2052 and a lower movable joint 2053, the upper movable joint 2051 and the lower movable joint 2053 are connected to both ends of the middle connecting rod 2052 The upper movable joint 2051 is connected to the fixing pin of the upper jaw shell 201, the lower movable joint 2053 is connected to the fixing pin of the lower jaw shell 202, the middle connecting rod 2052 is set as a telescopic rod, and the brushless servo 204 is connected to the upper movable joint 2051 to realize the motion control of the upper jaw shell 201 and the lower jaw shell 202, and realize the forward and backward movement speed control of the upper jaw shell 201 and the lower jaw shell 202 relative to the skull structure 1, and can also realize local micro-motion.
[0129] The brushless servo 204 is of 15kg grade, providing sufficient power for the movement of the jaw.
[0130] In this embodiment, the upper jaw shell 201 and the lower jaw shell 202 are controlled by the jaw movable component to perform corresponding movements. The jaw movable component is connected to the execution unit. The execution unit receives the jaw movement trigger instruction issued by the perception center and processes it, and then sends a jaw control signal to the jaw movable component, thereby controlling the movement of the upper jaw shell 201 and the lower jaw shell 202.
[0131] The jaw movement triggering command includes the command for triggering two kinds of activities of the jaw, namely the jaw movement triggering command and the jaw movement triggering command;
[0132] Jaw up and down movement trigger commands include
[0133] Chewing instruction: When the perception center detects food particles or other objects in the oral cavity, it triggers the upper jaw shell 201 to move up and down to simulate the chewing action and send food into the oral cavity;
[0134] Voice command: When the humanoid robot receives a specific verbal command, it triggers the up and down movement of the jaw, such as "mouth closed", "mouth open" and other commands;
[0135] Voice interaction instructions: When the humanoid robot has a voice conversation with the user, it triggers the up and down movement of the upper jaw, such as the start and end of utterance.
[0136] Visual recognition commands: When a specific gesture or object movement is detected by the visual sensor, the upper jaw is triggered to move up and down, such as detecting the movement of the hand or the object approaching the mouth;
[0137] Emotional expression instructions: When the humanoid robot needs to simulate expressions such as smiling or crying, the upper jaw moves up and down according to the emotion recognition algorithm;
[0138] The left and right jaw micro-movement trigger commands include
[0139] Visual tracking command: When the visual sensor detects a specific object or person moving left or right, it triggers a small left or right movement of the upper jaw to simulate a person's attention or observation action;
[0140] Voice interaction commands: When the humanoid robot has a voice conversation with the user, it triggers a slight left-right movement of the upper jaw according to the directionality of the voice command to express the intention to listen or respond;
[0141] Emotional simulation instructions: Analyze emotional signals in the environment based on the emotion recognition algorithm. For example, when a user's smile or nervousness is detected, the upper jaw will be triggered to move slightly to establish a closer emotional connection with the user.
[0142] The tooth structure includes an upper gum 241, upper teeth 242, lower teeth 243 and a lower gum 244. The upper gum 241 and the lower gum 244 are fixed to the skull structure 1. The upper teeth 242 are fixed to the upper gum 241, and the lower teeth 243 are fixed to the lower gum 244. The upper gum 241 is fixed to the upper jaw shell 201, and the lower jaw 244 is fixed to the lower jaw shell 202. The upper teeth 242 and the lower teeth 243 move synchronously with the upper jaw shell 201 and the lower jaw shell 202. The 15kg-level brushless servo 204 transmits torque to provide sufficient bite force for the tooth structure.
[0143] In this embodiment, the upper gum 241, upper teeth 242, lower teeth 243 and lower gum 244 are all made of high molecular polymer, and the number of teeth of the upper teeth 242 and lower teeth 243 can be set as needed.
[0144] The lip structure 21 includes an upper lip 211, a lip movable component and a lower lip 212. The upper lip 211 is movably connected to the upper jaw shell 201, and the lower lip 212 is movably connected to the lower jaw shell 202. The lip movable component is connected to the upper lip 211 and the lower lip 212, and the up and down movement and left and right movement of the upper lip 211 and the lower lip 212 are controlled by the lip movable component.
[0145] There are eight groups of lip movable components, four groups of lip movable components are used for connecting the upper lip 211 and the upper jaw shell 201, the four groups of lip movable components are evenly distributed along the length direction of the upper lip 211, specifically distributed at both ends of the upper lip 211 and the middle of the upper lip 211, and the other four groups of lip movable components are used for connecting the lower lip 212 and the lower jaw shell 202. The distribution method is the same as above and will not be repeated here.
[0146] The lip movable assembly includes a servo servo 213, a movable connecting rod 214 and a movable telescopic rod 215. The servo servo 213 is installed on the upper jaw shell 201, the movable connecting rod 214 is connected to the servo servo 213, one end of the movable telescopic rod 215 is connected to the movable connecting rod 214, and the other end of the movable telescopic rod 215 is rotatably connected to the upper lip 211. The lower lip 212 is movably connected to the lower jaw shell 202 through a movable assembly, which is the same as above and will not be repeated here.
[0147] The servo actuator 213 provides power for the movement of the upper lip 211 and the lower lip 212. The servo actuator 213 and the movable connecting rod 214, as well as the movable connecting rod 214 and the movable telescopic rod 215 are connected by universal balls to achieve 360° rotation. When the two groups of lip movable components located in the middle of the upper lip 211 and / or the lower lip 212 are started, the upper lip 211 and / or the lower lip 212 can move up and down. When the lip movable components located at both ends of the upper lip 211 and / or the lower lip 212 are started, the upper lip 211 and / or the lower lip 212 can move left and right.
[0148] In this embodiment, the upper lip 211 and the lower lip 212 are made of high temperature resistant silicone polymer.
[0149] In this embodiment, the upper lip 211 and the lower lip 212 are controlled by the lip movable component to perform corresponding movements. The lip movable component is connected to the execution unit. After receiving the lip movement trigger instruction, the execution unit sends a lip control signal to the lip movable component, thereby controlling the movement of the upper lip 211 and the lower lip 212.
[0150] The lip movement triggering instructions include instructions for triggering two kinds of activities of the lip movement, namely, the lip movement triggering instructions for up and down movement and the lip movement triggering instructions for left and right movement;
[0151] Lip up and down movement trigger commands include
[0152] Voice command: When the humanoid robot receives a specific verbal command, it triggers the up and down movement of the lips, such as "mouth closed", "mouth open" and other commands;
[0153] Visual recognition commands: When a specific gesture or object movement is detected by the visual sensor, the up and down movement of the lips is triggered, such as detecting the movement of the hand or the object approaching the mouth;
[0154] Emotional expression instructions: When the humanoid robot needs to simulate a smile, expression change, or vocalize, the up and down movement of the lips is triggered according to the emotion recognition algorithm;
[0155] Lip left and right micro-movement trigger commands include
[0156] Visual tracking command: When the visual sensor detects a specific object or person moving left or right, it triggers a small left or right movement of the lips to simulate a person's attention or observation action;
[0157] Voice interaction commands: When the humanoid robot has a voice conversation with the user, it triggers small left and right movements of the lips according to the directionality of the voice command to express the intention to listen or respond;
[0158] Emotional simulation instructions: Analyze emotional signals in the environment based on the emotion recognition algorithm. For example, when a user's smile or nervousness is detected, the lip movements are triggered to establish a closer emotional connection with the user.
[0159] A mounting frame 230 is fixedly provided in the oral space, and the tongue structure 23 is installed on the mounting frame 230. The tongue structure 23 includes a controllable telescopic servo, a tongue tire 231 and a tongue segment structure 232. The tongue segment structure 232 is composed of a plurality of tongue segments 233. The plurality of tongue segments 233 are installed in sequence and connected through connecting pipelines to form a snake-like chain structure, which can realize telescopic and curling movements. One end of the tongue segment structure 232 is connected to the mounting frame 230 and to the controllable telescopic servo. The controllable telescopic servo ensures that the tongue segment structure 232 moves according to instructions. The other end of the tongue segment structure 232 is a free end. The tongue tire 231 is nested on the outside of the tongue segment structure 232. The tongue tire 231 adopts APS+PC material, and a layer of sheath can be nested on the outside of the tongue tire 231.
[0160] The tongue 231 sheath is made of special polymer, specifically engineering plastic polyphenylene sulfide (PPS). PPS has very good thermal stability (can be used at 200°C for a long time and can withstand up to 260°C for a short time), and is also very resistant to chemical corrosion and oxidation. PPS has high mechanical strength and certain elasticity.
[0161] The tongue structure 23 also includes a taste perception system, which simulates the human taste perception ability. The taste perception system includes a number of high temperature resistant taste sensors, which are evenly distributed on the surface of the tongue 231 and are used to perceive five basic tastes: sour, sweet, bitter, salty, and fresh. Each taste has two sensors distributed on the left and right to ensure comprehensive coverage and accurate perception of taste stimuli.
[0162] The humanoid robot of this embodiment is equipped with a taste perception system to implement testing and evaluation of food without a human tester. The taste perception system can simulate the function of the human taste perception system, thereby automatically evaluating the taste of food.
[0163] The high temperature resistant taste sensor integrates a taste sensor module, which can detect the taste in the mouth and convert it into an electrical signal. Taste signal lines are distributed inside the tongue structure 232. The electrical signal is transmitted to the signal processing unit through the taste signal line for processing. The signal processing unit converts the original signal into a digital signal, and filters and processes it to enhance the accuracy of the signal. The processed signal is sent to the perception center, which analyzes, decodes and distinguishes the signal. It can identify the characteristic patterns of five basic tastes and compare them with pre-stored patterns to determine the type and degree of taste.
[0164] The process of the perception center analyzing the signal is as follows:
[0165] The perception center receives digital signals from the signal processing unit. These signals are first analyzed to determine their strength and pattern. This step is the first step in identifying the taste. By amplifying and filtering these signals, the perception center is able to distinguish which signals are relevant and which are background noise or irrelevant information.
[0166] The process of decoding the signal by the perception center is:
[0167] The decoding step involves converting the electrical signals obtained in the analysis phase into specific taste information. This process requires the use of known neural coding patterns, which are pre-set through extensive sensory testing and data analysis. During the decoding process, the nervous system matches the received signals with these pre-set patterns, such as sweet, sour, bitter, salty and umami.
[0168] The process of the perception center judging the signal is as follows:
[0169] In the discrimination phase, the perception center uses the previous decoding results to determine the specific flavor type and intensity. This step involves comparing the decoded flavor pattern with the standard flavor patterns stored in the database. Through this comparison, the perception center can accurately identify the flavor being experienced and evaluate its intensity.
[0170] The analysis, decoding and discrimination of signals by the perception center is a highly integrated and automated process that relies on advanced neural network and machine learning techniques to improve the accuracy and efficiency of recognition. In artificial systems, algorithms and neural network models that simulate these biological processes are used to perform similar tasks.
[0171] In this embodiment, the taste level is divided into 0-10 levels to indicate the robot's acceptance of different tastes. The threshold value may also be adjusted according to specific circumstances.
[0172] The perception center determines the type and degree of taste and sends corresponding control signals to the execution unit to perform corresponding actions.
[0173] The robot expresses its feelings about different tastes through mouth movements, sounds or expressions according to the detected tastes and taste levels. That is, the robot's execution unit responds and makes decisions according to the preset mode and task requirements for different tastes, so that the tongue structure performs specific actions, such as opening and closing, chewing, swallowing, etc., so that the humanoid robot can simulate the human taste perception process more realistically. Specifically:
[0174] Sour: The tongue is slightly curled or makes a sour expression, and there is a certain sharpness in the sound.
[0175] Sweet: tongue licking lips or smiling, with a pleasant note in the voice.
[0176] Bitter: The tongue may contract slightly or make an unhappy expression, and there may be a tone of depression or discomfort in the voice.
[0177] Salty: Sticking out the tongue or licking the lips, making a gesture of wanting to drink water, and the sound carries a sense of thirst.
[0178] Fresh: Slightly open your mouth and lick your lips, with a fresh or energetic tone in your voice.
[0179] The emotional tones or feelings in the above sounds are obtained by establishing a sour / sweet / bitter / salty / fresh model using conventional methods. The model establishes an association between taste and sound, and the corresponding sound is called through the recognized taste.
[0180] In this embodiment, the tongue structure 23 is controlled by a controllable telescopic servo to perform corresponding movements. The controllable telescopic servo is connected to an execution unit. After receiving and processing the tongue activity trigger instruction, the execution unit sends a tongue control signal to the controllable telescopic servo, thereby controlling the movement of the tongue structure 23.
[0181] The tongue movement triggering instructions include instructions for triggering two activities of the tongue extension and curling movements, namely, the tongue extension and curling movement triggering instructions and the tongue curling movement triggering instructions;
[0182] The tongue extension and retraction movement triggering instructions include
[0183] Voice command: When the humanoid robot receives a specific verbal command, it triggers the tongue to extend and retract, such as "tongue out", "tongue retract", etc.
[0184] Visual recognition commands: trigger tongue extension and retraction movements when specific gestures or object movements are detected by the visual sensor, such as detecting hand movements or objects approaching the mouth;
[0185] Emotional expression instructions: When the humanoid robot needs to simulate swallowing, spitting out objects, or changing facial expressions, the tongue extension and retraction movement is triggered according to the emotion recognition algorithm;
[0186] The tongue curling movement trigger commands include
[0187] Voice interaction commands: When the humanoid robot has a voice conversation with the user, it will perform tongue curling movements according to the content and emotion of the voice command, such as simulating spitting, licking, etc.
[0188] Emotional simulation instructions: Analyze emotional signals in the environment based on the emotion recognition algorithm. For example, when a user's smile, surprise or anger is detected, the tongue will be triggered to move slightly or curl to establish a closer emotional connection with the user.
[0189] The oral space is also equipped with sensory organs, which are devices that simulate the functions of the human oral cavity. The sensory organs perceive stimuli inside the oral cavity, such as taste, temperature, texture, etc., by capturing various stimulation signals inside the oral cavity. The sensory organs integrate a variety of sensors to simulate and evaluate the sensory experience of humans during the food intake process, including the above-mentioned taste sensors, tactile sensors, temperature sensors, pressure sensors, etc. Several types of sensors work together to fully simulate the taste, temperature, texture and other properties of food.
[0190] The installation location of the sensory organs will simulate the structural layout of the human oral cavity. Taste sensors and temperature sensors are distributed on the surface of the tongue, touch sensors and pressure sensors are distributed on the tooth structure, and temperature sensors are distributed in different parts of the simulated oral cavity to fully capture the experience of food intake.
[0191] The taste of food can be captured by taste sensors, and then processed by signal processing units. The perception center analyzes the signal to obtain the taste. The temperature of food can be captured by temperature sensors. The texture of food refers to the physical composition and feel of food, such as hardness, viscosity, humidity, graininess, etc. The texture is captured by tactile sensors. For example, the tooth structure can evaluate the hardness of food through pressure sensors, and tactile sensors can evaluate the viscosity of food or the roughness of the surface.
[0192] The skull structure is provided with two groups of visual mounting holes which are symmetrical on the left and right, and the visual system 3 is provided with two groups, and the two groups of visual systems 3 are respectively arranged at the positions of the two groups of visual mounting holes, and the two groups of visual systems 3 have the same structure. In this embodiment, the structure of one group of visual systems 3 is described, and the visual system 3 includes an eye shell 30 and an eyebrow structure 31, an eyelid structure 32 and an eyeball 33 which are installed in the eye shell 3, the eyebrow structure 31 is arranged above the visual mounting hole, the eyelid structure 32 and the eyeball 33 are installed in the visual mounting hole, the eyebrow structure 31 includes an eyebrow movable bracket 311 and two groups of eyebrow movable components 312, and the two groups of eyebrow The movable components 312 are respectively connected to the eyebrow movable brackets 311, and the up and down movement of the eyebrow movable brackets 311 is realized through the eyebrow movable components 312. The eyelid structure 32 includes an upper eyelid bracket 321, an upper eyelid 322, a lower eyelid 323, a lower eyelid bracket 324 and an eyelid movable component. The eyelid movable component controls the opening and closing movement of the upper eyelid bracket 321 and the lower eyelid bracket 324. The eyeball 33 includes a main body assembly 331 and an eye movable component 332. The main body assembly 33 is movably installed in the middle of the eyelid structure 32. The eye movable component 332 controls the up and down movement and rotation of the main body assembly 331.
[0193] The eye housing 30 serves as an outer structure that provides protection and support for the internal components of the eye.
[0194] The eyebrow movable support 31, the upper eyelid support 321, and the lower eyelid support 324 are made of ABS material; the upper eyelid 322 and the lower eyelid 323 are made of high-temperature resistant silicone polymer, on which eyelashes can be installed.
[0195] The eyebrow movable component 312 includes a servo actuator 3121 and an eyebrow connecting rod 3122. The upper end of the eyebrow connecting rod 3122 is connected to an upper movable joint 3123, and the lower end of the eyebrow connecting rod 3122 is connected to a lower movable joint 3124. The servo actuator 3121 is installed on the skull structure. The upper movable joint 3123 is connected to the servo actuator 3121 through an eyebrow movable pin 3125, and can achieve a rotation similar to a universal ball. The eyebrow movable bracket 311 is provided with an eyebrow fixing pin, and the lower movable joint 3124 is connected to the eyebrow fixing pin.
[0196] The servo motor 3121 starts to control the rotation of the eyebrow movable pin 3125. When the eyebrow movable pin 3125 rotates from the bottom to the top, it drives the eyebrow connecting rod 3122 to move upward, and then drives the eyebrow movable bracket 311 to move upward. When the eyebrow movable pin 3125 rotates from the top to the bottom, it drives the eyebrow connecting rod 3122 to move downward, and then drives the eyebrow movable bracket 311 to move downward, thereby controlling the up and down movement of the eyebrow movable bracket 311.
[0197] The two sets of eyebrow movable supports 311 of the two sets of visual systems 3 can perform synchronous or asynchronous movement through the drive of the eyebrow movable component 312, thereby realizing the expression of various emotions by the robot.
[0198] In this embodiment, the eyebrow movable support 311 is controlled by two groups of eyebrow movable components 312 to perform corresponding movements. The eyebrow movable component 312 is connected to the execution unit. After receiving the eyebrow movable support activity trigger instruction, the execution unit sends an eyebrow movable support control signal to the eyebrow movable component 312, thereby controlling the movement of the eyebrow movable support structure 23.
[0199] The eyebrow movable support activity triggering instruction is the eyebrow movable support up and down movement triggering instruction and the asynchronous movement instruction of two groups of eyebrow movable support activities at the left and right positions, including:
[0200] Visual recognition instructions: When the user's eye expression or gesture is detected by a camera or other visual sensor, the up and down movement of the right eyebrow is triggered according to the user's eye direction or gesture. For example, when the user raises his eyebrows upward or frowns downward, the corresponding movement of the right eyebrow is triggered;
[0201] Voice interaction commands: When the humanoid robot has a voice conversation with the user, the up and down movement of the right eyebrow is triggered according to the content and directionality of the voice command to express the robot's attention or response action.
[0202] The asynchronous motion triggering instructions for the two sets of eyebrow movable bracket activities include
[0203] Visual tracking instructions: Track the user's head movement or facial expression through a camera or other visual sensors. When the user's head or eyes turn left or right, the right eyebrow will be triggered to move slightly left or right to simulate the person's eye gaze or expression changes.
[0204] Voice interaction commands: When the humanoid robot has a voice conversation with the user, the right eyebrow will be triggered to move slightly left and right according to the content and direction of the voice command to express the robot's attention or response action.
[0205] The upper eyelid 322 is fixed on the upper eyelid support 321, and the lower eyelid 323 is fixed on the lower eyelid support 324. The upper eyelid support 321 and the lower eyelid support 324 are respectively provided with tooth ends at the left and right ends. The two tooth ends are coaxial and meshingly arranged inside and outside. When one of the tooth ends is driven to rotate, the upper eyelid support 321 and the lower eyelid support 324 realize opposite rotation due to their meshing movement, which can make the upper eyelid support 321 and the lower eyelid support 324 perform opening and closing movements to realize blinking operations, and the opening and closing frequency can be set according to needs.
[0206] The eyelid movable component provides power for the eyelid movable component. Specifically, the eyelid movable component includes an eyelid opening and closing rotating shaft 325 and a servo servo. The servo servo is installed on the skull structure. The output shaft of the servo servo is fixedly connected to the eyelid opening and closing rotating shaft 325. The eyelid opening and closing rotating shaft 325 is fixedly connected to the tooth end located on the inner side. When the servo servo is started, power is transmitted to the inner tooth end through the eyelid opening and closing rotating shaft 325, thereby realizing the opening and closing movement of the upper eyelid bracket 321 and the lower eyelid bracket 324.
[0207] The two sets of eyelid structures 32 of the two sets of visual systems 3 can perform synchronous opening and closing movements or asynchronous opening and closing movements through the driving of the eyelid movable components.
[0208] In this embodiment, the upper eyelid support 321 and the lower eyelid support 324 are controlled by the eyelid movable component to perform corresponding movements. The servo actuator is connected to the execution unit. After receiving the eyelid support trigger instruction, the execution unit sends the eyelid support control signal to the servo actuator, thereby controlling the movement of the eyelid support.
[0209] The eyelid support triggering instruction includes an instruction for triggering the opening and closing movement of the upper eyelid support 321 and the lower eyelid support 324;
[0210] The upper eyelid support 321 and the lower eyelid support 324 opening and closing movement triggering instructions include
[0211] Visual recognition command: When the user's eye expression or gesture is detected by a camera or other visual sensor, the opening and closing movement of the right upper eyelid bracket is triggered according to the user's eye state. For example, when the user closes or opens his eyes, the corresponding action of the right upper eyelid bracket is triggered;
[0212] Emotion simulation instructions: Analyze emotional signals in the environment based on the emotion recognition algorithm. For example, when the user's tiredness, surprise or joy is detected, the opening and closing movement of the right upper eyelid bracket is triggered to simulate the corresponding changes in eye expression.
[0213] The main body assembly 331 includes an eyeball-shaped shell 3315, an eyeball-shaped shell back cover 3316 and a back cover wire collection slot locking nut 3317. A visual sensor 3311, a variable-focus movable motor seat 3312, an integrated circuit board 3313 and a rubber ring 3314 for fixing a connecting rod seat are fixed in sequence inside the eyeball-shaped shell 3315. The eyeball-shaped shell back cover 3316 is fixed to the end of the eyeball-shaped shell 3315 and is locked by the back cover wire collection slot locking nut 3317.
[0214] The visual sensor 3311 can use a miniature high-definition camera with functions such as automatic exposure / auto focus / auto white balance / auto fill light to capture visual information of the surrounding environment.
[0215] The visual sensor 3311 is part of the visual video tracking system and transmits the captured high-definition video and images to the graphics processing unit for processing.
[0216] The eye movement component 332 is used to control the movement of the eyeball so that it can rotate in the horizontal and vertical directions, thereby changing the direction of the line of sight. The eye movement component 332 includes an eyeball support rod 3320, an eyeball connecting rod 3322 and a servo actuator 3325. The eyeball support rod 3320 is located in the middle and connects the eyeball 33 and the eyeball connecting block 3324. The eyeball connecting block 3324 is installed on the skull structure. The servo actuator 3325 is installed on the eyeball connecting block 3324. Both ends of the eyeball connecting rod 3322 are provided with movable joints. The movable joint at one end is connected to the eyeball contour shell 3315 through a fixing pin, and the other end is connected to the servo actuator 3325 through a movable pin. A group of servo actuators 3325 controls the movement of a group of eyeball connecting rods 3322.
[0217] When a number of servo actuators 3325 control the main body assembly 331 to move forward and the remaining number of servo actuators 3325 control the main body assembly 331 to move backward, the main body assembly 331 can be turned up or down. This embodiment controls different servo actuators 3325 to drive different eyeball connecting rods 3322 to achieve horizontal and vertical rotation of the eyeball, thereby changing the direction of the line of sight.
[0218] In this embodiment, when the robot detects that a person or object has entered its surroundings, the visual video tracking system will begin to capture images and transmit them to the perception center for analysis. If the perception center determines that the line of sight needs to be adjusted to track the target, it sends a corresponding control instruction to the eye movement component 332 to trigger the eye movement;
[0219] The triggering of eye movement mainly includes target detection, target importance assessment, target movement analysis, and the priority and requirements of the current task, which specifically includes the following steps:
[0220] Step A1. Target detection and confirmation:
[0221] 1.1. Detection: The visual video tracking system detects people or objects in the image;
[0222] 1.2. Confirmation: Confirm whether the detected object meets the tracking conditions and characteristics, such as a specific shape, color, or known markers.
[0223] Step A2. Goal importance and priority assessment:
[0224] 2.1. Importance: Assessing the importance of a target depends on the specific requirements of the task (e.g., security monitoring, interactive tasks, or important objects in a specific scenario).
[0225] 2.2 Priority: In a multi-target environment, determine which targets have a higher tracking priority based on the target’s dynamic behavior, mission relevance, or preset rules.
[0226] Step A3. Target dynamic analysis:
[0227] 3.1 Motion tracking: Analyze the target's motion trajectory to determine whether it is moving, its speed and direction.
[0228] 3.2 Predict future position: Use the motion estimation model to predict the target’s future position to determine how to adjust the line of sight most effectively.
[0229] Step A4. Task and environment requirements:
[0230] 4.1 Task requirements: Decide whether to adjust the line of sight based on the robot’s current task requirements (e.g., whether it needs to continuously monitor an area or object) or environmental changes (e.g., other moving objects, light changes, etc.).
[0231] Step A5. Sight adjustment strategy:
[0232] 5.1 Instant adjustment: If the position or speed of the target does not match the predetermined tracking strategy, immediately adjust the sight direction to keep the target in the center of the field of view or in an appropriate position.
[0233] 5.2 Predictive adjustment: Based on the prediction model of target movement, adjust the line of sight direction in advance to reduce reaction delay and improve the continuity and accuracy of tracking.
[0234] The trigger instructions may include
[0235] Visual tracking instructions: The user's head movement or facial expression is captured in real time through a camera or other visual sensors, and the left and right eyeballs are triggered to move synchronously according to the user's head movement. When the user's head is raised up, lowered, turned left or right, the left and right eyeballs are triggered to move accordingly to simulate the changes in a person's line of sight.
[0236] Gesture recognition commands: The humanoid robot is equipped with gesture recognition technology that triggers eye movements based on the user's hand movements. For example, when the user's finger points in a specific direction, the eye moves in the corresponding direction.
[0237] Context-aware instructions: Through environmental perception technologies such as sound sensors or depth cameras, the sound or object position in the surrounding environment is sensed, thereby triggering the asynchronous movement of the left and right eyeballs. For example, when a humanoid robot detects that the sound comes from the left, the left eyeball may turn to the left, while the right eyeball remains still to simulate the direction of human attention.
[0238] Emotional simulation instructions: Analyze the user's emotional changes based on the emotion recognition algorithm and trigger asynchronous eye movement based on the emotional changes. For example, when the user is detected to be anxious or surprised, the left and right eyeballs may show asynchronous up and down and left and right movements to express the robot's response to the user's emotions.
[0239] When the robot needs to change the direction of its sight, the perception center will analyze the current environment and determine the target position that needs to be adjusted. Specifically, based on the target position and the current position of the eyeball, the execution unit will send a corresponding control signal to the eye movement component 332 to adjust the direction and position of the eyeball to align with the target.
[0240] The direction in which the robot needs to change its line of sight is related to target tracking, task requirements, and environmental changes.
[0241] Target tracking: Deviation of the target position detected by the vision sensor from the expected or planned path;
[0242] Task requirements: For example, the user needs to change their line of sight to adapt to changes in the environment or focus when navigating, avoiding obstacles, or interacting with people.
[0243] Environmental changes: Changes in the environment may require the robot to adjust its line of sight to gain more information or better understand its surroundings.
[0244] The specific steps include:
[0245] Step B1: Determination of target location;
[0246] The specific location or coordinates of the target are usually determined by the following steps:
[0247] 1.1. Image Capture: First, the visual sensor captures the image of the current field of view;
[0248] 1.2. Image processing: Identify objects in the image through the signal processing unit;
[0249] 1.3. Object positioning: Locate the position of an object in three-dimensional space based on a deep learning model;
[0250] 1.4. Coordinate transformation: According to the coordinates of the robot's own position and orientation, the coordinates of the target object are transformed into the position in the global coordinate system;
[0251] Step B2: Adjust the eye position and sight direction;
[0252] The adjustment process includes:
[0253] 2.1. Error calculation: Calculate the deviation between the target position and the current line of sight center;
[0254] 2.2. Motion planning: Calculate the angle and direction of eye movement based on the deviation;
[0255] 2.3. Perform adjustment: By controlling the eye movement component 332, adjust the line of sight so that the target is located in the center of the visual field or other predetermined position.
[0256] The nasal system 4 includes a nose bridge 41 and an olfactory sensor 42. The nose bridge 41 is provided with a left nostril and a right nostril. The olfactory sensor 42 is installed in the left nostril and the right nostril respectively. The olfactory sensor 42 is an important component for receiving and analyzing odor information in the surrounding environment. It simulates the olfactory perception mechanism in the human nasal cavity, can perceive different odor molecules, and convert them into information that can be understood and responded to by the system.
[0257] The olfactory sensor 42 can achieve the following functions:
[0258] Environmental perception: The olfactory sensor 42 can sense various odors in the surrounding environment, such as the fragrance of flowers, the aroma of food, the smell of smoke, etc., which enables the humanoid robot to respond promptly to changes in the odor of the surrounding environment.
[0259] Safety protection: The olfactory sensor 42 can also detect dangerous odors, such as gas leaks, the smell of gas, or the presence of toxic gases. Once these dangerous odors are detected, the system can trigger corresponding safety protection measures, such as stopping the robot operation, alarming, or moving to a safe area.
[0260] Health monitoring: By detecting odors in the surrounding environment through the olfactory sensor 42, the humanoid robot can monitor some health-related information, such as air quality, environmental pollution level, etc. This helps to protect the health of the robot and the human user.
[0261] Mission execution: In some specific tasks, odor information may be one of the important information necessary to perform the mission. For example, in search and rescue missions, robots may need to use odor to find the location of trapped people; in the agricultural field, robots may need to use odor to detect the maturity of crops or the presence of pests and diseases.
[0262] Emotional expression: In some cases, scent can also be used as a form of emotional expression. For example, a robot might be designed to “smell” flowers and express joy, or “smell” smoke and express worry or nervousness.
[0263] The role of the olfactory sensor 42 is not only to provide environmental odor information, but also to convert this information into meaningful data for the robot, thereby helping the robot to better understand and adapt to its surrounding environment. Through the olfactory sensor, the bionic humanoid robot can achieve more intelligent and humanized interaction and task execution, enhancing its applicability and practicality in various application scenarios.
[0264] The ear system 5 includes ears 51 and auditory sensors 52. The ears 51 are provided with two groups, symmetrically located on both sides of the skull structure. The auditory sensors 52 are fixed to the ears 51 to receive sounds in the surrounding environment. The signal processing unit receives and processes the sounds to obtain sound signals that can be recognized by the system, and the sound signals are transmitted to the sensory center.
[0265] The steps for the signal processing unit to receive and process the sound are as follows:
[0266] Step 1: Sound capture: The auditory sensor 52 captures sound signals in the surrounding environment and converts the sound waves into electrical signals;
[0267] Step 2: Preprocessing: De-noise and filter the electrical signal to eliminate environmental noise and enhance the clarity and recognizability of the sound;
[0268] Step 3: Extract sound features: Extract sound features from the preprocessed signal;
[0269] Sound features include frequency, amplitude, and duration of sound, which help to analyze and identify the sound;
[0270] Step 4: Sound recognition: Use machine learning algorithms or pattern recognition technology to analyze and compare sound features to identify the type and meaning of the sound. The training model can recognize different sound categories, such as voice commands, environmental noise, alarm sounds, etc., and classify and process them.
[0271] Step 5: Action triggering: The sensory center sends instructions to the execution unit based on the recognized sound type and meaning, and the execution unit triggers the corresponding action or response.
[0272] For example, if it is recognized as a user's voice command, the execution unit may perform a corresponding task or action; if it is recognized as an emergency alarm sound, the execution unit may take corresponding safety measures or issue a warning signal.
[0273] Different sounds can trigger different actions or responses, and the specific triggering conditions depend on the system design and the preset task requirements. For example, a user speaking a specific voice command can trigger the robot to perform a corresponding operation, such as moving to a specific location, performing a specific task, etc.
[0274] When an alarm or unusual sound is captured, the system may trigger the robot to take emergency measures, such as stopping the current task and moving to a safe location.
[0275] After receiving the sound, the ear sensor 52 of the humanoid robot converts the sound into understandable information through the above processing steps, and triggers corresponding actions or responses according to the recognition results to achieve interaction with the environment and task execution.
[0276] The voice instructions and voice interaction instructions in this embodiment are captured by the auditory sensor, processed by the signal processing unit, and issued by the perception center to the execution unit; the visual tracking / recognition instructions and emotion simulation instructions are captured by the visual sensor, processed by the signal processing unit, and issued by the perception center to the execution unit; and the emotion expression instructions are issued by the perception center to the execution unit.
[0277] Obviously, the described embodiments are only some embodiments of the utility model, not all embodiments. Based on the embodiments of the utility model, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the utility model.
Claims
1. The head structure of a bionic service humanoid robot is characterized by: It includes a skull structure, a sensory system, an oral system connected to the sensory system, a visual system, a nasal system and an ear system. The oral system, the visual system, the nasal system and the ear system are installed on the skull structure. The oral system includes the jaw structure, the tooth structure, the lip structure and the tongue structure. The visual system includes the eyebrow structure, the eyelid structure and the eyeball. The nasal system includes the nose bridge and the olfactory sensor installed on the nose bridge. The ear system includes the ear and the auditory sensor installed on the ear.
2. The head structure of the bionic service humanoid robot according to claim 1, characterized in that: The jaw structure comprises an upper jaw shell, a jaw movable component and a lower jaw shell. The left and right ends of the upper jaw shell and the lower jaw shell are rotatably connected through upper and lower jaw fixed shafts, and the upper and lower jaw fixed shafts are installed on the skull structure.
3. The head structure of the bionic service humanoid robot according to claim 2, characterized in that: The jaw movable assembly includes a brushless servo, a jaw rear connecting rod structure and a jaw front connecting rod structure, the jaw rear connecting rod structure and the jaw front connecting rod structure are respectively connected to the brushless servo, the jaw rear connecting rod structure and the jaw front connecting rod structure respectively include an upper movable joint, a middle connecting rod and a lower movable joint, the upper movable joint and the lower movable joint are connected to the two ends of the middle connecting rod, the upper movable joint is connected to the fixing pin of the upper jaw shell, the lower movable joint is connected to the fixing pin of the lower jaw shell, and the brushless servo is connected to the upper movable joint.
4. The head structure of the bionic service humanoid robot according to claim 1, characterized in that: The tooth structure comprises an upper gum, upper teeth, lower teeth and a lower gum. The upper gum and the lower gum are fixed on the skull structure. The upper teeth are fixed on the upper gum, and the lower teeth are fixed on the lower gum. The upper gum is fixed on the upper jaw shell, and the lower gum is fixed on the lower jaw shell.
5. The head structure of the bionic service humanoid robot according to claim 1, characterized in that: The lip structure includes an upper lip, a lip movable component and a lower lip. The upper lip is movably connected to the upper jaw shell, and the lower lip is movably connected to the lower jaw shell. The lip movable component is connected to the upper lip and the lower lip, and the up and down movement and left and right movement of the upper lip and the lower lip are controlled by the lip movable component.
6. The head structure of the bionic service humanoid robot according to claim 5, characterized in that: The lip movable assembly includes a servo steering engine, a movable connecting rod and a movable telescopic rod. The servo steering engine is installed on the upper jaw shell, the movable connecting rod is connected to the servo steering engine, one end of the movable telescopic rod is connected to the movable connecting rod, and the other end of the movable telescopic rod is rotatably connected to the upper lip.
7. The head structure of the bionic service humanoid robot according to claim 1, characterized in that: The oral system is provided with an oral space, a mounting frame is fixedly provided in the oral space, a tongue structure is installed on the mounting frame, the tongue structure includes a controllable telescopic servo, a tongue tire and a tongue segment structure, the tongue segment structure is composed of a plurality of tongue segments, the plurality of tongue segments are installed in sequence and connected through connecting pipelines, one end of the tongue segment structure is connected to the mounting frame and to the controllable telescopic servo, the other end of the tongue segment structure is a free end, and the tongue tire is nested outside the tongue segment structure.
8. The head structure of the bionic service humanoid robot according to claim 7, characterized in that: The tongue structure also includes a taste perception system, which includes a plurality of high-temperature resistant taste sensors distributed on the surface of the tongue.
9. The head structure of the bionic service humanoid robot according to claim 1, characterized in that: The sensory system includes a perception center, a sensor network, a signal processing unit and an execution unit. The sensor network is connected to the signal processing unit, the signal processing unit is connected to the perception center, and the perception center is connected to the execution unit. The sensory system also includes a sensory central electromagnetic interference protection system and a sensory central storage chip family. The sensory central electromagnetic interference protection system realizes electromagnetic interference protection. The sensory central electromagnetic interference protection system uses conductive or magnetic materials to cover sensitive components. The sensory central storage chip family is a memory and storage device for storing data collected from the sensor network, including random access memory, fast access memory and solid state hard disk. The fast access memory is used to temporarily store data in processing, and the solid state hard disk is used for long-term data storage.
10. The head structure of the bionic service humanoid robot according to claim 1, characterized in that: The head structure is provided with an inspection hatch, a memory chip card slot and a hatch. The hatch is detachably installed in the memory chip card slot. The sensory central memory chip family is installed in the memory chip card slot. The inspection hatch is hinged to the head structure.
Citation Information
Cited By
Bionic service type humanoid robot head structure
CN118952165A
Bionic service type humanoid robot head structure
CN118952165B
Bionic robot tongue structure, head system, robot and control method
CN121340372A