Multi-modal interaction control method and system of multifunctional teaching assisting robot and robot

By integrating multimodal interaction modules and unified data processing, the multifunctional teaching assistant robot system solves the problem of single interaction modality in teaching assistant robots, realizes stable interaction and efficient device linkage in complex environments, and improves the recognition accuracy of teaching scenarios and the reusability of the system.

CN121572277APending Publication Date: 2026-02-27SHENZHEN YUXIN DIGITAL TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202610001843.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-04
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing teaching assistant robots have a single interaction modality, resulting in insufficient recognition accuracy and interaction stability in noisy, backlit, or multi-person scenarios. Incomplete peripheral interconnection interfaces or unclear protocol stack layers lead to high device access costs, poor compatibility, and difficulty in reusing linkage logic.

Method used

The system employs a multi-functional teaching assistant robot system, integrating modules for image acquisition, audio input, audio output, human-computer touch control, and display interaction. It uses a tablet AI processor to uniformly schedule and process multi-source data, and combines multi-modal collaboration and redundancy mechanisms to achieve multi-channel voice signal processing, face recognition, environmental monitoring, and device linkage. It also supports WIFI, Bluetooth, and Ethernet communication.

Benefits of technology

It improves the continuity and stability of teaching interaction, enhances the accuracy and consistency of identity recognition, reduces the engineering complexity of system expansion and deployment, and strengthens the versatility and reusability of teaching assistant robots in multiple teaching scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121572277A_ABST
    Figure CN121572277A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent education and Internet of Things fusion, in particular to a multi-modal interaction control method and system of a multifunctional teaching-assistant robot and the robot. According to the system, a tablet AI processor runs an Android system as a core, a display module, a man-machine interaction module, an audio input module, an audio output module, an image acquisition module, a temperature and humidity sensor module, a WIFI or Bluetooth module, an Ethernet module and an Internet of Things module are connected, and the Internet of Things module supports RS485 wired access and Zigbee wireless access at the same time; the tablet AI processor executes voice interaction, video call, face recognition, environment monitoring, network interaction and equipment linkage, and provides a face recognition alignment and depth feature matching algorithm and an annular microphone array beam forming and sound source direction estimation algorithm, so that multi-modal interaction and multi-equipment linkage in a teaching scene are realized. The convenience and the safety are improved; and the equipment access cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent education, artificial intelligence and Internet of Things, and in particular to a multi-modal interaction control method and system of a multifunctional teaching robot and the robot. BACKGROUND

[0002] In a teaching scene, a teaching robot not only needs to undertake human-computer interaction tasks such as teaching interaction, information display, question answering and guidance, but also is often required to interconnect with intelligent education devices, smart home devices or peripheral control devices in a classroom to realize device linkage, environment regulation and safety management. The existing solutions generally have the problem of single interaction mode, usually only relying on single voice or single camera to complete the interaction, resulting in insufficient recognition accuracy and interaction stability in a noisy environment, a backlight environment or a multi-person scene, and incomplete peripheral interconnection interfaces or unclear protocol stack layering, resulting in high device access cost, poor compatibility and difficult reuse of linkage logic.

[0003] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0004] The present application aims to provide a multi-modal interaction control method and system of a multifunctional teaching robot and the robot to overcome the deficiencies of the prior art.

[0005] The present application achieves the above-mentioned purpose through the following technical solutions: a multi-modal interaction control method based on a multifunctional teaching robot, comprising: System starting step: starting an Android system and initializing a display module, a human-computer interaction module, an audio input module, an audio output module, an image acquisition module, a temperature and humidity sensor module, an Internet of Things module, a WIFI or Bluetooth module and an Ethernet module by a tablet AI processor; Interaction presentation step: outputting a human-computer interaction interface by the display module and receiving touch interaction input information or key interaction input information from the human-computer interaction module; Voice acquisition step: acquiring multi-channel voice signals by the audio input module and transmitting them to the tablet AI processor; Voice processing step: after the voice acquisition step, performing short-time Fourier transform on the multi-channel voice signals and performing beamforming based on a steering vector to obtain enhanced voice signals; Voice interaction step: after the voice processing step, performing voice wake-up or voice recognition based on the enhanced voice signals to generate interaction response information; Audio broadcast step: after the voice interaction step, outputting voice or multimedia audio corresponding to the interaction response information by the audio output module; Video processing step: collecting video images through the image acquisition module and performing video call processing or face recognition processing to generate identity results; Environment monitoring step: collecting environmental temperature and humidity through the temperature and humidity sensor module and performing threshold judgment to generate environmental state information; Network interaction step: data interaction with the server through WIFI or Bluetooth module or Ethernet module to upload the identity results or the environmental state information and receive control instructions; Device linkage step: after the network interaction step, sending device control instructions to external intelligent education devices or smart home devices through the Internet of Things module and receiving device state feedback.

[0006] Further, in the speech processing step, phase transformation weighting is performed on the multi-channel speech signal, and direction response power calculation is performed based on the candidate direction angle to obtain the sound source direction estimation result; After obtaining the sound source direction estimation result, it also includes constructing a steering vector based on the sound source direction estimation result and performing delay-sum beamforming to obtain an enhanced speech signal; Between the speech processing step and the speech interaction step, there are also echo cancellation steps, speech activity detection steps, and frequency domain noise suppression steps.

[0007] Further, in the video processing step, the face recognition processing includes the following sub-steps: Face detection and key point positioning step: performing face detection on the collected video images and outputting a set of key point coordinates; Face alignment step: performing rotation, scaling and translation on the face image based on the set of key point coordinates to obtain an aligned face image; Feature extraction step: input the aligned face image into a deep feature extraction network to obtain a feature vector and perform two-norm normalization; Feature matching step: performing similarity calculation on the normalized feature vector and the normalized feature vector in the feature library, and outputting an identity result or a rejection result based on a threshold.

[0008] Further, in the environment monitoring step, when the environmental temperature is greater than or equal to 60 degrees Celsius or the environmental humidity is greater than or equal to 90% relative humidity, an alarm event is generated and the audio broadcast step and the interaction presentation step output alarm information are triggered; In the device linkage step, the device control instructions are sent to wired devices through RS485 communication or to wireless devices through Zigbee communication.

[0009] A multi-modal interaction control system based on a multifunctional teaching robot, comprising: System startup unit: Used by the tablet AI processor to start the Android system and initialize the display module, human-computer interaction module, audio input module, audio output module, image acquisition module, temperature and humidity sensor module, IoT module, WIFI or Bluetooth module and Ethernet module; Interactive presentation unit: used to output the human-computer interaction interface through the display module and receive touch interaction input information or button interaction input information from the human-computer interaction module; Voice acquisition unit: used to acquire multi-channel voice signals through the audio input module and transmit them to the tablet AI processor; Speech processing unit: used to perform short-time Fourier transform on multi-channel speech signals and perform beamforming based on steering vectors to obtain enhanced speech signals; Voice interaction unit: used to perform voice wake-up or voice recognition based on the enhanced voice signal to generate interactive response information; Audio broadcasting unit: used to output voice or multimedia audio corresponding to the interactive response information through the audio output module; Video processing unit: used to acquire video images through the image acquisition module and perform video call processing or face recognition processing to generate identity results; Environmental monitoring unit: used to collect ambient temperature and humidity through temperature and humidity sensor modules and perform threshold judgment to generate environmental status information; Network interaction unit: used to interact with the server via WIFI, Bluetooth or Ethernet module to upload the identity result or the environmental status information and receive control commands; Device linkage unit: Used to send device control commands to external smart education devices or smart home devices and receive device status feedback via IoT module.

[0010] A multi-functional teaching assistant robot, comprising: A tablet AI processor, used to run the Android system and carry peripheral function modules; The display module is connected to the tablet AI processor and is used to display the human-computer interaction interface; The human-computer interaction module is connected to the tablet AI processor and is used to output touch interaction input information or button interaction input information; An audio output module, connected to the tablet AI processor, is used to output multimedia audio or human-computer interaction voice. An image acquisition module, connected to the tablet AI processor, is used to acquire video images to support video calls or facial recognition; An audio input module, connected to the tablet AI processor, is used to collect voice signals to support voice wake-up or voice interaction; The Internet of Things (IoT) module is connected to the tablet AI processor and is used to communicate with external smart education devices or smart home devices to achieve device linkage. A WIFI or Bluetooth module is connected to the tablet AI processor to connect to the wireless network and interact with the server for data exchange. An Ethernet module, connected to the tablet AI processor, is used to connect to a wired network and interact with the server for data exchange; A temperature and humidity sensor module is connected to the tablet AI processor to detect ambient temperature and humidity.

[0011] Furthermore, the audio input module includes a six-element circular microphone array, the array radius of which is in the range of 30 mm to 45 mm; The tablet AI processor samples the multi-channel speech signal of the six-element ring microphone array at a sampling rate of 48000 Hz and performs a short-time Fourier transform. The frame length of the short-time Fourier transform is 1024 sampling points and the frame shift is 256 sampling points. The tablet AI processor performs SRP-PHAT sound source direction estimation on the multi-channel complex spectrum observation obtained by the short-time Fourier transform to obtain candidate sound source direction angle information, and constructs a steering vector based on the candidate sound source direction angle information and performs delayed summation beamforming to output an enhanced speech signal. The tablet AI processor sequentially performs echo cancellation processing, voice activity detection processing, and frequency domain noise suppression processing on the enhanced voice signal, and uses the processed voice signal for the voice wake-up interaction.

[0012] Furthermore, the image acquisition module includes a camera with a resolution of at least 8 megapixels and supports autofocus and dynamic ISP adjustment functions; The display module includes a MIPI interface display screen and supports outputting 1080P video signals to an external display via an HDMI interface; The audio output module includes a digital amplifier and two 4-ohm 5-watt speakers. The IoT module includes an RJ12 interface for RS485 communication and a Zigbee gateway module for Zigbee communication. The Zigbee gateway module supports 802.15.4 MAC and PHY and operates on channels 11 to 26 in the 2.400GHz to 2.483GHz frequency band, with an air interface rate of 250Kbps and supports AES128 or AES256 hardware encryption.

[0013] Furthermore, the temperature and humidity sensor module includes: The SHT40 temperature and humidity sensor is connected to the tablet AI processor via the IIC interface and is used to periodically report ambient temperature and humidity sampling values. The threshold judgment module is electrically connected to the SHT40 temperature and humidity sensor and has a built-in threshold judgment program. The threshold judgment program is used to generate an alarm event when the ambient temperature sampling value is greater than or equal to 60 degrees Celsius or the ambient humidity sampling value is greater than or equal to 90% relative humidity. The alarm interface is output through the display module and the alarm voice is output through the audio output module.

[0014] The SHT40 temperature and humidity sensor is replaced with an SHT70 temperature and humidity sensor, which is connected to the tablet AI processor via an IIC interface.

[0015] Furthermore, the tablet AI processor is configured to perform face recognition processing according to the following processing chain: Perform face detection on the images acquired by the image acquisition module and output a set of key point coordinates to generate alignment transformation parameters; Based on the alignment transformation parameters, the face image is rotated, scaled, and translated to obtain an aligned face image; The aligned face image is input into a deep feature extraction network to obtain a feature vector, and the feature vector is normalized by L2 to obtain a normalized feature vector. The cosine similarity score is calculated based on the normalized feature vector and the normalized feature vector in the feature library, and the identity result or rejection result is output based on the threshold.

[0016] The beneficial effects of this invention are: Compared to existing teaching assistant robot solutions that rely solely on a single voice or visual interaction method, this project integrates multiple interaction modules within a single system, including image acquisition, audio input, audio output, human-computer touch control, and display interaction. A tablet AI processor then uniformly schedules and processes the multi-source data, enabling teaching interaction to no longer depend on a single sensory channel. Even when a sensory channel is affected by environmental noise, lighting changes, or human occlusion, the system can still complete command input and information feedback through other interaction channels. This creates a multimodal collaboration and redundancy mechanism in actual teaching scenarios, improving the continuity and stability of the human-computer interaction process.

[0017] In terms of identity recognition and teaching management, this case constructs a complete face recognition processing chain within a tablet AI processor. It sequentially performs face detection, key point localization, geometric alignment, deep feature extraction, and similarity calculation on the acquired face images, and outputs an identity result or rejection result based on a threshold, ensuring a clear data flow and judgment basis for the identity recognition process. This processing method avoids misidentification problems caused by relying solely on raw images or simple feature comparisons. Especially in multi-person scenarios or dynamic teaching environments, it can improve the accuracy and consistency of identity recognition, providing a reliable foundation for personalized teaching and device integration.

[0018] Furthermore, this project simultaneously incorporates wired communication and wireless IoT modules at the system level, enabling external smart education or smart home devices to access and be controlled through a unified processing core. Combined with a clear software architecture and device interface logic, various peripherals can report status and issue control commands without requiring additional control nodes or complex protocol conversions during the connection process. This reduces the engineering complexity of system expansion and deployment, and enhances the versatility and reusability of the teaching assistant robot across multiple teaching scenarios. Attached Figure Description

[0019] Figure 1 This is a schematic flowchart of one method of the present invention.

[0020] Figure 2 This is a schematic block diagram of a system module of the present invention.

[0021] Figure 3 This is a schematic block diagram of system modules according to an embodiment of the present invention.

[0022] Figure 4 This is a schematic diagram of the software architecture of the present invention.

[0023] Figure 5 This is a schematic diagram of the hardware framework of the present invention.

[0024] Figure 6 This is a schematic diagram of the Zigbee device docking logic of the present invention. Detailed Implementation

[0025] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0026] In related technologies, especially in teaching scenarios, teaching assistant robots not only undertake human-computer interaction tasks such as teaching interaction, information display, and Q&A guidance, but are also often required to interconnect with intelligent educational devices, smart home devices, or peripheral control devices in the classroom to achieve device linkage, environmental adjustment, and safety management. Existing solutions generally suffer from the problem of single interaction modality, usually relying on a single voice or a single camera to complete the interaction. This results in insufficient recognition accuracy and interaction stability in noisy environments, backlit environments, or multi-person scenarios. At the same time, incomplete peripheral interconnection interfaces or unclear protocol stack layering lead to high device access costs, poor compatibility, and difficulty in reusing linkage logic.

[0027] Therefore, the inventors discovered through research that the industry urgently needs a system and method that modularly integrates display touch control, voice pickup and playback, image acquisition and recognition, temperature and humidity sensing, Ethernet and wireless network communication, RS485 wired IoT, and Zigbee wireless IoT on the same platform, and provides a clear software architecture and device docking logic, in order to improve the convenience and security of teaching interaction and reduce the engineering implementation threshold of multi-device linkage.

[0028] Therefore, such as Figures 1-6 As shown, this invention provides a multimodal interactive control method based on a multifunctional teaching assistant robot, a multimodal interactive control system based on a multifunctional teaching assistant robot, and a multifunctional teaching assistant robot.

[0029] like Figure 2As shown, the multifunctional teaching assistant robot system in this embodiment uses a tablet AI processor as its core computing and control unit. The tablet AI processor runs an Android system and outputs a human-computer interaction interface to the display module after system startup. Simultaneously, it receives touch or button input information from the human-computer interaction module. Within the same Android system, the tablet AI processor performs voice wake-up and voice interaction processing on the multi-channel voice signals collected by the audio input module, and outputs human-computer interaction voice or multimedia audio to the audio output module. The tablet AI processor also receives video images collected by the image acquisition module and performs video call processing or face recognition processing. Simultaneously, the temperature and humidity sensor module reports the ambient temperature and humidity to the tablet AI processor. The tablet AI processor performs threshold judgments and triggers the display module and audio output module to synchronously output alarm information when alarm conditions are met. In order to realize data interaction with the server, the system is equipped with a WIFI or Bluetooth module and an Ethernet module. The tablet AI processor can interact with the server through wireless network or wired network to receive control commands or upload identity results and environmental status information. In order to realize device linkage, the system also includes an IoT module. The IoT module communicates with external smart education devices or smart home devices and transmits the device status back to the tablet AI processor, so that the teaching assistant robot can simultaneously complete a comprehensive functional closed loop of voice interaction, visual recognition, environmental monitoring, network interaction and multi-device linkage in the teaching scenario.

[0030] like Figure 3As shown, in an engineering selection implementation, the flat-panel AI processor can select the Allwinner A733 chip as the system-level main control and run the Android system. The WIFI or Bluetooth module can select the AW869C wireless communication module to support 2.4G and 5G dual-band wireless networks and support Bluetooth 5.4, so as to maintain stable data interaction with the server through the wireless network in the classroom scenario; the Ethernet module is connected to the flat-panel AI processor to provide a wired network link, so that the wired network can be preferentially selected for data interaction in scenarios that are more sensitive to latency and packet loss; the display module can drive the external LCD module by outputting the MIPI signal through the TCON of the flat-panel AI processor to display the human-computer interaction interface. At the same time, the flat-panel AI processor can output a 1080P video signal to the external display through HDMI to support a larger-screen classroom display; the human-computer interaction module can communicate with the external touch controller GT9271 through IIC to form CTP multi-touch input, and can further superimpose key interaction input information to meet the classroom fast control requirements; the audio output module can use the digital power amplifier AD82584F to amplify the digital audio signal output by the flat-panel AI processor and drive two 4-ohm 5-watt speakers, so as to avoid sound distortion and ensure the clarity of classroom broadcasts by limiting and dynamic control when working at the maximum volume; the image acquisition module can use the GC08A8 camera and connect it to the flat-panel AI processor through MIPI CSI. The camera pixel is not less than 8M and supports automatic focus and dynamic adjustment of the ISP function to improve the image quality of video calls and face recognition in the classroom lighting change scenario; the temperature and humidity sensor module can use the SHT40 temperature and humidity sensor and connect it to the flat-panel AI processor through IIC to realize the periodic sampling and reporting of environmental temperature and environmental humidity. At the same time, in another replaceable implementation, the temperature and humidity sensor module can use the SHT70 temperature and humidity sensor and keep the IIC interface and software sampling interface consistent, so as to complete the sensor replacement without changing the main control software; the Internet of Things module can be connected to the flat-panel AI processor through SP3485EN to realize RS485 communication, so as to control external wired intelligent education devices. The Internet of Things module can also be connected to the flat-panel AI processor through the Tuya TYZS13 Zigbee gateway to realize Zigbee communication, so as to control external wireless smart home devices and form an execution path for classroom one-key linkage.

[0031] To ensure the system can be implemented according to the methodology, within a typical software runtime cycle, after system startup, the tablet AI processor completes the loading of peripheral drivers and service registration within the Android system. Subsequently, the display module outputs the human-computer interaction interface and enters the event loop to receive touch or button input information. When the user triggers voice interaction, the audio input module begins outputting multi-channel audio signals. The tablet AI processor buffers, segments, and performs time-frequency transformation on the multi-channel audio signals. After obtaining the enhanced audio signal, it enters the voice wake-up or voice recognition process to generate interactive response information. This response information further drives the audio output module to output human-computer interaction voice or drives the display module to refresh the interface content. When the user triggers a video call or face recognition, the image acquisition module continuously acquires video images, and the tablet AI processor will... Video images are used for encoding and transmission to form video call processing, or input into the facial recognition processing link to generate identity results and be used for personalized permissions and linkage strategies; the temperature and humidity sensor module outputs ambient temperature and humidity periodically, and the tablet AI processor performs threshold judgment and drives the display module and audio output module to output alarm information synchronously when alarm conditions are met; in terms of network interaction, the tablet AI processor interacts with the server through WIFI, Bluetooth or Ethernet modules to upload identity results or environmental status information and receive control commands issued by the server; in terms of device linkage, the tablet AI processor generates device control commands based on control commands or local policies, sends device control commands to external smart education devices or smart home devices through the IoT module and receives device status feedback to achieve a closed-loop linkage.

[0032] To ensure the reproducibility and direct engineering implementation of the voice processing link, one implementation employs a six-element circular microphone array for the audio input module. The array radius of the six-element circular microphone array is set within the range of 30 mm to 45 mm, and in this embodiment, the array radius is set to 35 mm to fit the space of a typical tablet form factor enclosure. The tablet AI processor synchronously samples the multi-channel voice signals of the six-element circular microphone array at a sampling rate of 48000 Hz, and performs a short-time Fourier transform on each channel. The short-time Fourier transform frame length is set to 1024 sampling points and the frame shift is set to 256 sampling points, thereby ensuring real-time performance while obtaining sufficient frequency resolution for sound source direction estimation and beamforming. After obtaining the time-frequency domain complex spectrum of each channel, the tablet AI processor first performs phase transformation weighting to suppress the influence of reverberation and amplitude instability on direction estimation. Then, it performs pointing response power calculation based on candidate direction angles to obtain the sound source direction estimation result. After obtaining the sound source direction estimation result, it constructs the steering vector and performs delayed summation beamforming to obtain enhanced speech signal. Furthermore, in order to improve the success rate of voice wake-up in the classroom echo environment, the tablet AI processor can insert echo cancellation processing, voice activity detection processing, and frequency domain noise suppression processing between speech processing and voice interaction to reduce crosstalk between speaker playback and microphone pickup and reduce interference of non-speech frames on the recognition algorithm.

[0033] The geometric modeling and delay relationship of a six-element circular microphone array can be determined in a computable implementation as follows, so that ordinary technicians can reproduce the guide vector and delay compensation according to the formula.

[0034] Satisfying formula (1): Among them, the angle of the array element Indicates the first The polar angles of the microphones on the circumference, and the array element numbers. Channel numbers used to identify a six-element circular microphone array.

[0035] Satisfying formula (2): Wherein, the array radius (R) represents the geometric radius of the six-element circular microphone array, and the element position vector... Indicates the first The position of each microphone in a plane coordinate system.

[0036] Satisfying formula (3): ; where, incident direction angle Indicates the incident direction of the sound source relative to the array, and the speed of sound. This represents the speed of sound in air, the th Channel arrival delay Indicates direction angle The plane wave reaches the first The time offset of each microphone relative to the center of the array.

[0037] In engineering implementation, the tablet AI processor is based on arrival latency. A steering vector is constructed and used for delay summation beamforming. The frequency domain form of the steering vector can be calculated from the arrival delay and angular frequency, thereby transforming delay compensation into phase compensation and aligning and superimposing the complex spectra of each channel.

[0038] In terms of the face recognition processing chain, to ensure that the system side has a feasible implementation point, the tablet AI processor is configured in software to perform face recognition processing according to a fixed processing chain. The processing chain includes face detection, key point output, alignment transformation parameter generation, aligned image generation, deep feature extraction, L2 normalization, cosine similarity calculation, and threshold judgment to output identity result or rejection result. Among them, key point output is used to generate alignment transformation parameters, which are used to unify faces with different poses and scales to a standard template coordinate system to improve the consistency of feature extraction. The deep feature extraction network is used to map the aligned face image into feature vectors. L2 normalization is used to map the feature vectors to a unit hypersphere to make the cosine similarity have a stable geometric meaning. Threshold judgment is used to output rejection result when the similarity is insufficient, thereby avoiding false recognition.

[0039] To enable the face alignment steps described above to be explicitly calculated and directly reproduced, in one implementation, the tablet AI processor performs face detection on the acquired image and outputs a set of key point coordinates. The set of key point coordinates corresponds one-to-one with the set of key points in the standard template. The similarity transformation parameters are solved by least squares fitting to obtain the alignment transformation parameters, and the face image is rotated, scaled, and translated.

[0040] Satisfying formula (4): Among them, the current key points Indicates the number of images captured. Two-dimensional coordinate vectors of key points, template key points Indicates the first in the standard template Two-dimensional coordinate vectors of key points, scaling factor Represents the overall scaling factor, rotation matrix Represents a two-dimensional rotation transformation, translation vector This represents a two-dimensional translation.

[0041] Satisfying formula (5): ; where, rotation angle The rotation matrix represents the rotation angle required for alignment. Used to rotate the current face pose to a pose consistent with the standard template.

[0042] In engineering implementation, the tablet AI processor performs rotation, scaling, and translation on the face image according to the alignment transformation parameters and crops it to obtain an aligned face image. The aligned face image is input into a deep feature extraction network to obtain a feature vector and then normalized by the L2 norm to obtain a normalized feature vector. Subsequently, a cosine similarity score is calculated based on the normalized feature vector and the normalized feature vector in the feature library, and the identity result or rejection result is output based on the threshold. Thus, the face recognition processing link on the system side has a clear input and output and a repeatable calculation process.

[0043] Satisfying formula (6): Among them, the feature vector The L2 norm represents the original feature vector output by the deep feature extraction network. Represents the Euclidean length of the eigenvector, and the normalized eigenvector. This represents the unit vector after performing L2 norm normalization.

[0044] Satisfying formula (7): Among them, the cosine similarity score Represents the normalized feature vector of the current face. With the feature library Individual normalized feature vector Similarity between them, dot product When both vectors are unit vectors, it is equivalent to cosine similarity.

[0045] During the threshold determination stage, the tablet AI processor compares the cosine similarity score with a preset threshold. When the cosine similarity score meets the threshold condition, it outputs the identity result; when the cosine similarity score does not meet the threshold condition, it outputs the rejection result. This avoids forcibly matching strangers as individuals in the database in high-traffic classroom scenarios.

[0046] Regarding the implementation of the IoT module, to ensure clear communication support for device linkage steps, the IoT module simultaneously possesses both RS485 and Zigbee communication paths. In one implementation, the IoT module includes an RJ12 interface for RS485 communication. The RJ12 interface connects to the tablet AI processor via SP3485EN to achieve half-duplex differential communication with wired devices. The IoT module also includes a Zigbee gateway module, which supports 802.15.4 MAC and PHY and operates on channels 11 to 26 in the 2.400GHz to 2.483GHz frequency band, with an air interface rate of [missing information]. With a bandwidth of 250Kbps and support for AES128 or AES256 hardware encryption, wireless devices can be configured for network, encrypted communication, and status feedback. In the device linkage process, the tablet AI processor generates device control commands based on server control instructions or local policies. For wired devices, the device control commands are encapsulated into RS485 communication frames and sent through the RJ12 interface. For wireless devices, the device control commands are converted into Zigbee application layer commands and sent through the Zigbee gateway module. The device status feedback returned by the external device is sent back to the IoT module through the corresponding path and reported to the tablet AI processor to form a linkage closed loop.

[0047] like Figure 6 As shown, in one implementation of the Zigbee device docking logic, the tablet AI processor runs a Zigbee device management service within the Android system. The Zigbee device management service communicates with the Zigbee gateway module through a serial interface and maintains device identifiers, device capabilities, and device status. During the initial docking, the system performs device network access and key negotiation and writes the device identifier into the local device table. Subsequently, in a classroom linkage scenario, device control commands are generated based on interactive response information or server control commands and sent to the Zigbee gateway module. The Zigbee gateway module selects the target device according to the device identifier and sends control commands. The target device returns the execution result and device status. The Zigbee gateway module returns the device status to the tablet AI processor, which then uploads it to the server via network interaction steps, thus realizing a closed-loop link from human-computer interaction to device control and then to status return. The above logic and the RS485 channel can coexist in parallel, enabling both wired smart education devices and wireless smart home devices to be linked in the same teaching scenario.

[0048] Regarding the alarm implementation of the temperature and humidity sensor module, to ensure that the environmental monitoring steps have directly executable judgment conditions, in one implementation, the temperature and humidity sensor module uses an SHT40 temperature and humidity sensor and connects to a tablet AI processor via an IIC interface. The tablet AI processor periodically collects ambient temperature and humidity data and performs threshold judgment. When the ambient temperature is greater than or equal to 60 degrees Celsius or the ambient humidity is greater than or equal to 90% relative humidity, an alarm event is generated. The alarm event simultaneously triggers the display module to overlay an alarm prompt on the current human-machine interface and triggers the audio output module to play an alarm voice message, thereby promptly reminding management personnel to ventilate and cool down when the classroom air conditioner fails or when the classroom is stuffy due to overcrowding. In another alternative implementation, the temperature and humidity sensor module uses an SHT70 temperature and humidity sensor and maintains the same IIC interface and software sampling period as the SHT40 temperature and humidity sensor, thereby completing the device replacement without changing the threshold judgment logic and alarm output logic.

[0049] like Figure 4 As shown, in one software architecture implementation, the upper layer of the Android system sets up a human-computer interaction application layer for interface presentation and event handling, a voice processing service for multi-channel voice signal acquisition, time-frequency transformation, sound source direction estimation, beamforming, and voice wake-up interaction, a vision processing service for video image acquisition, video call processing, and face recognition processing chain execution, an environmental monitoring service for temperature and humidity acquisition and threshold alarms, an IoT control service for RS485 and Zigbee device interfacing, control command issuance, and status feedback, and a network communication service for data interaction with the server via WIFI, Bluetooth, or Ethernet modules, and for uploading identity results and environmental status information, as well as receiving and distributing control commands. The above services transmit interactive response information and device control commands through the inter-process communication mechanism of the Android system, so that the system startup, interactive presentation, voice acquisition, voice processing, voice interaction, audio broadcasting, video processing, environmental monitoring, network interaction, and device linkage in the method flow can run in a closed loop in sequence under the same software framework and can be directly deployed and implemented in engineering.

[0050] As can be seen from the above embodiments, the tablet AI processor, display module, human-computer interaction module, audio input module, audio output module, image acquisition module, temperature and humidity sensor module, IoT module, WIFI or Bluetooth module, and Ethernet module in this embodiment have clear connections in hardware and clear service divisions in software. The voice link achieves sound source direction estimation and beamforming through a six-element ring microphone array and short-time Fourier transform parameterized configuration, thereby obtaining enhanced voice signals and driving voice wake-up interaction. The visual link achieves key point detection and alignment transformation parameter generation, as well as depth feature extraction and L2 norm. The system establishes a stable face recognition processing link through normalization and outputs identity verification or rejection results based on cosine similarity scores and thresholds. The IoT link enables control and status feedback of external wired and wireless devices through dual RS485 and Zigbee channels. The environmental link provides classroom environment anomaly alerts through temperature and humidity sampling and threshold alarms. The network link maintains data interaction with the server and receives control commands through dual wireless and wired channels. This allows the multi-functional teaching assistant robot to simultaneously complete multimodal interaction and multi-device linkage in the classroom scenario using the same system, and each link has a directly implementable and repeatable engineering implementation basis.

[0051] In detail, in a more specific implementation: like Figure 2 As shown, the multifunctional teaching assistant robot system of this invention centers on a tablet AI processor. The tablet AI processor runs the Android system and executes application logic and algorithm logic. The tablet AI processor is connected to a display module, a human-computer interaction voice control module, an audio output module, an image acquisition module, an audio input module, an IoT module, a WIFI or Bluetooth module, an Ethernet module, and a temperature and humidity sensor module. The display module displays the human-computer interaction interface and presents course content, interactive prompts, and device status information to the user. The human-computer interaction voice control module generates interactive control commands and submits interactive events to the tablet AI processor. The audio output module plays multimedia audio and interactive voice. The image acquisition module acquires video images to support video calls and facial recognition. The audio input module acquires voice signals to support recording and intelligent voice pickup. The IoT module interconnects with external intelligent educational devices or smart home devices. The WIFI or Bluetooth module accesses a wireless network and interacts normally with the server. The Ethernet module accesses a wired network and interacts normally with the server. The temperature and humidity sensor module detects ambient temperature and humidity and reports the sampled data to the tablet AI processor to support environmental monitoring and early warning. Through the combination of the above modules, the system simultaneously possesses multimodal perception, interactive presentation, network communication, and IoT control capabilities on the same platform, thus forming an AI-enabled and interconnected teaching assistant robot system for teaching scenarios.

[0052] like Figure 3 As shown, in one specific embodiment, the tablet AI processor can be implemented using the Allwinner SOCA733 as the core chip. The display side can simultaneously support MIPI displays and external HDMI displays. The tablet AI processor outputs display signals to the MIPI display through its internal display control output channel to support the human-computer interaction interface. Simultaneously, the tablet AI processor outputs 1080P video signals to an external display through the HDMI output interface to adapt to classroom large-screen display scenarios. The human-computer interaction side can use the GT9271 touch controller. The GT9271 communicates with the tablet AI processor through the IIC interface, and combined with the touchscreen panel, it enables multi-touch input. A button interaction module serves as supplementary input for triggering frequently used functions or emergency operations. The audio side can use the AD82584F digital amplifier as the audio output device. The tablet AI processor outputs digital audio streams to the AD82584F digital amplifier through the IIS interface, and the AD82584F digital amplifier then drives speakers to play teaching audio, prompts, and interactive voice. The audio input side can use a ring microphone array, connected through the audio input channel of the tablet AI processor, to achieve far-field sound pickup, wake-up, and interaction. On the image side, a GC08A8 camera can be used, connecting to the tablet AI processor via the MIPICSI interface to capture video streams. The tablet AI processor executes facial recognition or QR code scanning algorithms on the video streams to support identity recognition and interactive teaching. On the network side, an RTL8211F-CG Ethernet device can be used, connecting to the tablet AI processor via the RGMII interface for wired network access. On the wireless side, an AW869C network module can be used, connecting to the tablet AI processor via the SDIO or UART interface for Wi-Fi and Bluetooth communication. On the IoT side, an SP3485EN485 transceiver can be used, connecting to the tablet AI processor via the UART interface and outputting RS485 bus signals to control external wired devices or communicate with external smart education devices. Additionally, a TYZS13 Zigbee gateway module can be used, connecting to the tablet AI processor via the UART interface for Zigbee device access and control. The environmental sensing side can use the SHT40 temperature and humidity sensor or an equivalent temperature and humidity sensor device. The temperature and humidity sensor is connected to the tablet AI processor through the IIC interface and periodically reports the environmental temperature and humidity sampling values ​​to support high temperature warning, environmental comfort prompts or linkage control strategies.

[0053] like Figure 3As shown, in one engineering implementation, the temperature and humidity sensor module uses an SHT40 temperature and humidity sensor. The SHT40 temperature and humidity sensor is electrically connected to the tablet AI processor via an IIC interface. The tablet AI processor reads the ambient temperature and ambient humidity sampling values ​​at a 1-second sampling period and writes them to the local status register. Simultaneously, an alarm event is triggered when the ambient temperature sampling value is greater than or equal to 60 degrees Celsius or the ambient humidity sampling value is greater than or equal to 90% relative humidity. The alarm event drives the display module to output an alarm interface and drives the audio output module to output an alarm voice to remind the user to perform ventilation or cooling measures. In another alternative implementation, the temperature and humidity sensor module uses an SHT70 temperature and humidity sensor and maintains the same IIC interface and sampling period, thereby achieving device replacement without changing the software interface.

[0054] like Figure 4 and Figure 5 As shown, the software architecture of this invention is based on a hardware abstraction layer. Within this layer, peripheral interfaces such as GPIO, PWM, UART, IIC, SPI, and FLASH are uniformly encapsulated, enabling upper-layer protocols and applications to consistently access underlying resources. Above this layer, a core Zigbee protocol stack module is built. This core module, along with the cluster management module, attribute management module, and RF transceiver module, works collaboratively to enable Zigbee devices to join the network, discover, bind, send and receive commands, and synchronize their states. An event-driven timing module is placed between the protocol stack and the application for scheduled task scheduling and event triggering. The Zigbee protocol functional layer includes a basic device discovery and network entry module, a link management module, a battery management module, and a Zigbee cluster command transmission and reception module, enabling the system to perform unified control and status feedback of lighting fixtures, switches, and scene devices in classroom or home device scenarios.

[0055] like Figure 5As shown, in one engineering implementation, the audio input module uses a six-element circular microphone array with an array radius of 35 mm. The tablet AI processor synchronously samples the six-channel audio signals at a sampling rate of 48000 Hz and performs short-time Fourier transform processing. The short-time Fourier transform frame length is 1024 sampling points, the frame shift is 256 sampling points, and a Hanning window is used. The tablet AI processor first performs PHAT normalized cross-spectrum calculation on the multi-channel complex spectrum observations and performs SRP-PHAT sound source direction estimation based on the candidate direction grid to obtain the candidate sound source direction angle information. Then, based on... Candidate sound source direction angle information is used to construct a steering vector and perform delayed summation beamforming to output an enhanced speech signal. The enhanced speech signal is then further processed sequentially with echo cancellation, speech activity detection, and frequency domain noise suppression before being input into the voice wake-up interaction module. Echo cancellation is used to suppress playback echoes generated by the audio output module, speech activity detection is used to eliminate non-speech frames and reduce invalid calculations, frequency domain noise suppression is used to reduce the impact of classroom ambient noise on wake-up and recognition, and the sound source direction estimation results are used to update the steering vector direction once per second to adapt to the user's movement in the classroom.

[0056] like Figure 6 As shown, the Zigbee device docking logic of this invention is embodied in a multi-network fusion path. Multiple Zigbee terminal devices establish a wireless connection with a Zigbee gateway through a Zigbee network. The Zigbee gateway then connects to a router via Ethernet or Wi-Fi and establishes a connection with a cloud platform. After the mobile application interacts with the cloud platform, it can send control commands to the Zigbee gateway, which forwards the commands to the target Zigbee terminal device and obtains the execution status feedback. The system can also support BLEmesh networks in parallel. BLEmesh terminal devices connect to a router through a BLEmesh gateway and interact with the cloud platform, thereby achieving unified management of Zigbee terminal devices and BLEmesh terminal devices under the same cloud platform and the same mobile application. Through this logic, the teaching assistant robot can serve as both a local interaction and computing center and an IoT control entry point, realizing multi-device linkage and unified control within the teaching space.

[0057] To enable face recognition to be implemented directly, this invention provides a mathematical processing pipeline for face recognition and gives the definition of key formulas, which are expressed in the following fixed format.

[0058] Satisfying formula (1): Among them, the set of key points This represents the set of coordinates of all keypoints detected in the current face image, and the number of keypoints. This indicates the number of key points and their coordinates. Indicates the first The position of each key point in the current image coordinate system, with the horizontal coordinate... Represents the column direction coordinates and vertical coordinates of the image. Represents the row direction coordinates of the image.

[0059] Satisfying formula (2): Among them, the template key point set This represents the set of keypoint coordinates on a standard template face. Indicates the first in the template The location of each key point, and the coordinates of the current key point. Coordinates of key points in the template One-to-one correspondence to satisfy the alignment of semantic points with the same name.

[0060] Satisfying formula (3): Among them, the scaling factor Indicates the alignment scaling ratio, rotation matrix Represents alignment and rotation transformation, translation vector Equation (3) represents the alignment translation amount and is used to adjust the coordinates of the current key points. Mapping to template key point coordinates via similarity transformation The similarity transformation parameters can be solved by least-squares Protodyakonov analysis to achieve face straightening and scale uniformity.

[0061] Satisfying formula (4): Among them, rotation angle The rotation matrix represents the rotation angle required to rotate the current face pose to the template pose. satisfy To maintain the shape and only change the direction.

[0062] Satisfying Formula (5): Formula (5): ; where the input image Represents the aligned face image, deep feature extraction network Represents the backbone network and feature vectors of a convolutional neural network. Indicates network output Dimensional embedding features.

[0063] Satisfying formula (6): ; where, normalized eigenvectors Represents the eigenvector A unit vector after L2 normalization, L2 norm This represents the Euclidean length of the feature vector. Normalization is used to make subsequent similarity measurements more stable.

[0064] Satisfying formula (7): Among them, Euclidean distance Represents the normalized feature vectors of two faces with normalized eigenvectors exist The distance in 3D space indicates a higher degree of similarity; the smaller the distance, the higher the similarity.

[0065] Satisfying formula (8): Among them, cosine similarity Let cosine be the angle between the normalized feature vectors of two faces. The larger the value, the higher the similarity. Under the condition of L2 normalization, the cosine similarity is equal to the inner product.

[0066] Satisfying formula (9): ; where the input feature vector This represents the normalized feature vector of the currently aligned face, and the feature vector within the database. Indicates the first in the feature library Normalized feature vectors of each identity, similarity scores Indicates the input and the first in the library The cosine similarity of each identity, the recognition result This represents the identity index with the highest similarity, and can be combined with a threshold to achieve external identity rejection.

[0067] In one implementation, the deep embedding model can be trained using angular margin loss to improve inter-class separability. The core form of angular margin loss can be written as follows.

[0068] Satisfying formula (10): Among them, the scaling factor The scaling factor for the logarithmic values ​​of the classification, angular interval. This represents the angle penalty introduced by the true category, and the category angle. This represents the angle between the input features and the true class weight vector, not the true class angle. The loss function represents the angle between the input features and the non-true class weight vector. Enhance feature discriminative power by forcing the true class angle to be smaller.

[0069] To enable far-field sound pickup and sound source localization, this invention provides key formula definitions for the geometry and signal processing of a ring microphone array.

[0070] Satisfying formula (11): Among them, the angle of the array element Indicates the number of microphones in a circular microphone array The polar angle of each microphone relative to the center of the array, element number This indicates the microphone index; the total number of array elements is 6.

[0071] Satisfying formula (12): Among them, the array radius The position vector represents the distance from the microphone to the center of the array. Indicates the first The position of each microphone in a plane coordinate system.

[0072] Satisfying formula (13): Among them, the angle of incidence Indicates the direction angle of the sound source and the speed of sound. Represents the speed constant of sound in air, the first Microphone delay Indicates the direction from which it originates. The plane wave reaches the first relative to the center of the array. Arrival delay of each microphone.

[0073] Satisfying Formula (14): Formula (14): Among them, relative delay Indicates the first The microphone relative to the first The arrival delay difference of each microphone is used for beamforming delay compensation and sound source direction estimation.

[0074] Satisfying formula (15): Among them, angular frequency Angular frequency representing the analysis frequency, and the pilot component. Indicates the direction from which it originates. Narrowband plane waves in the first Complex gain on each microphone, imaginary unit satisfy .

[0075] Satisfying formula (16): Among them, the guiding vector Indicates the array at frequency With direction The spatial fingerprint is used for beamforming and sound source direction estimation.

[0076] Satisfying formula (17): Among them, the short-time frequency domain observation vector Indicates the first The frequency band and the first Multi-channel complex spectrum observation at frame time, target complex spectrum This represents the complex spectral amplitude and phase of the target speech at that time-frequency point, and the target direction. Represents the target sound source direction angle, noise vector This represents the multichannel complex spectrum of noise and interference at that time and frequency point.

[0077] Satisfying formula (18): Among them, the covariance matrix Indicates the first Spatial covariance matrix of each frequency band, Hermitian transpose Represents the conjugate transpose, expectation operator This represents the statistical expectation, which can be estimated using a time sliding window average to meet engineering requirements.

[0078] Satisfying formula (19): Among them, the delayed summation weight vector Indicates the direction of aiming at the target. Delayed summation beamforming weights, complex conjugate This indicates that the beam output complex spectrum is obtained by taking the conjugate of the guide vector element by element. This represents the single-channel output complex spectrum after weighted summation of multi-channel complex spectrum observations. The output time-domain signal can be obtained by inverse short-time Fourier transform.

[0079] like Figures 4-6 As shown, the IoT module of this invention covers both RS485 and Zigbee at the system level. In RS485 mode, a differential bus signal can be output via an RJ12 interface to connect to external wired devices. In Zigbee mode, access is achieved through a Zigbee gateway, which sends commands to Zigbee terminal devices and receives status feedback. The Zigbee gateway can be implemented using a built-in low-power 32-bit ARM Cortex-M4 processor, covering channels 11 to 26 of the 2.4GHz band and supporting 802.15.4 MAC and PHY. It has hardware encryption capabilities and supports AES128 or AES256, meeting the security requirements of classroom equipment control.

[0080] This invention further specifies at the engineering parameter level that the WIFI or Bluetooth module can include WIFI 2.4GHz and WIFI 5GHz and support Bluetooth 5.4; the image acquisition module's camera pixel is not less than 8M and supports autofocus and dynamic adjustment ISP function; the audio output module can include two 4-ohm 5-watt speakers and does not produce distortion when operating at maximum volume; the audio input module can adopt a ring microphone array and meet the requirements of sensitivity not less than -32dBm and wake-up and interaction distance not less than 5 meters, while also having 360-degree omnidirectional sound pickup and sound source localization capabilities; the temperature and humidity sensor module can have a temperature accuracy of ±0.2 degrees Celsius and a humidity accuracy of ±1.8% relative humidity, so as to realize refined environmental monitoring and linkage strategies in the teaching environment.

[0081] In summary, in real-world teaching scenarios, teaching assistant robots face environments characterized by dense crowds, complex noise levels, variable lighting conditions, and multiple devices operating in parallel. Relying solely on a single perception method or simple interaction logic can easily lead to inaccurate recognition, interrupted interaction, or system instability. Therefore, the core technical problem this invention aims to solve is how to enable teaching assistant robots to stably and accurately recognize interactive objects in real-world teaching environments, maintain continuous and natural human-computer interaction capabilities under various interference conditions, and simultaneously possess good system scalability and device linkage capabilities.

[0082] To address the aforementioned technical challenges, this invention does not employ a single sensing method for enhancement. Instead, it constructs a unified architecture for multimodal sensing and processing within the same system, enabling parallel acquisition and centralized processing of image, voice, touch input, and environmental sensing information. The system uses a tablet AI processor as the unified processing core, with each functional module establishing a clear data interaction relationship with it. The tablet AI processor schedules, analyzes, and makes decisions regarding data from different modules, thus avoiding logical confusion caused by direct coupling between modules. This structure allows the system to maintain normal human-computer interaction even when any sensing channel is affected by environmental factors, through other interaction channels, thereby improving the overall stability of the system.

[0083] In terms of identity recognition, this invention constructs a complete and continuous face recognition processing chain within a tablet AI processor, enabling the face recognition process to move beyond simple image comparison and instead be based on a defined mathematical model for judgment. Specifically, the face image acquired by the image acquisition module is first processed by the tablet AI processor to perform face detection and extract a set of facial key point coordinates. Based on this set of key point coordinates, alignment transformation parameters are calculated, and the original face image is rotated, scaled, and translated to obtain an aligned face image with uniform pose. Subsequently, the aligned face image is input into a deep feature extraction network to obtain corresponding feature vectors, and these feature vectors are normalized using the L2 norm to eliminate the influence of scale differences. Finally, the cosine similarity score between the normalized feature vector and the normalized feature vectors in the feature library is calculated and combined with a preset threshold for judgment, outputting an identity recognition result or a rejection result. Through the above processing flow, the identity recognition process has a clear data flow and judgment criteria, thereby improving the accuracy and consistency of recognition in multi-person scenarios and dynamic environments.

[0084] In terms of voice interaction, the present invention sets the audio input module as an annular microphone array structure, enabling the voice acquisition process to not only obtain the sound signal itself but also the spatial information of the sound source. Based on the time difference of the same sound source signal received by each microphone in the array, the tablet AI processor can estimate the direction of the sound source and enhance the voice in the target direction through corresponding signal processing methods, while suppressing the noise and reverberation interference in non-target directions. Thereby, voice wake-up and voice interaction have higher stability in the teaching environment, avoiding problems such as false triggering or missed recognition caused by environmental noise or multi-source interference.

[0085] In addition, the present invention sets a wired communication module and an Internet of Things module at the system level, enabling external intelligent education devices or smart home devices to access the system through a unified processing core. By uniformly managing the device access process, status reporting process, and control instruction issuing process in the tablet AI processor, different types of peripherals can complete linkage control without additional control nodes or complex protocol conversions during the access process, thereby enhancing the scalability of the system and the maintainability of engineering implementation.

[0086] In summary, this case provides a multifunctional teaching assistant robot system and method integrating multi-modal human-computer interaction, biometric recognition, and Internet of Things device linkage, belonging to the technical category of manufacturing digital home intelligent terminal devices, intelligent perception and control devices, and other intelligent consumer devices. In actual teaching and home intelligent application scenarios, the teaching assistant robot not only needs to have stable human-computer interaction capabilities but also needs to be able to accurately identify the user's identity and form effective linkage with external intelligent education devices or smart home devices. Therefore, the present invention constructs a teaching assistant robot system architecture with a tablet AI processor as the core around the collaborative acquisition and centralized processing of multi-modal perception information. By integrating various interaction methods such as image acquisition, voice input, human-computer touch, and environmental perception on a unified processing platform, the system has good interaction stability and engineering scalability.

[0087] The system described in the present invention runs on the Android operating system at the software level and constructs a function scheduling and data processing mechanism for artificial intelligence applications on it, enabling functions such as image recognition, voice interaction, and device linkage to run in a modular manner. The overall technical solution is consistent with the technical forms of artificial intelligence optimized operating systems, artificial intelligence middleware, and function libraries. On this basis, by introducing a face recognition processing link and a voice signal processing link into the system, the system has comprehensive processing capabilities for computer vision and computer audition. The processes of face feature extraction, similarity calculation, and identity determination involved belong to the scope of application software development such as typical computer vision and audition software, biometric recognition software, etc.

[0088] In addition, by setting up an Internet of Things module, the present invention enables communication and linkage control between the teaching assistant robot and external intelligent education devices or smart home devices, allowing the system to not only operate as an independent interactive terminal but also be deployed as a centralized control node among multiple devices. The overall technical solution conforms to the system integration characteristics of information system integration services such as artificial intelligence systems and smart home systems in the production field. Through the collaborative design of software and hardware, the teaching assistant robot has both the attributes of a smart terminal and system integration in educational and home scenarios, thereby enhancing the generality and scalability of the overall application.

Claims

1. A multimodal interactive control method based on a multifunctional teaching assistant robot, characterized in that, include: System startup steps: The tablet AI processor starts the Android system and initializes the display module, human-computer interaction module, audio input module, and audio output module; Interactive presentation steps: Output the human-computer interaction interface through the display module and receive touch interaction input information or button interaction input information from the human-computer interaction module; Voice acquisition steps: Acquire multi-channel voice signals through the audio input module and transmit them to the tablet AI processor; Speech processing steps: Perform short-time Fourier transform on the multi-channel speech signal and perform beamforming based on the steering vector to obtain an enhanced speech signal; Voice interaction steps: Perform voice wake-up or voice recognition based on the enhanced voice signal to generate interactive response information; Audio broadcasting steps: Output the voice or multimedia audio corresponding to the interactive response information through the audio output module.

2. The multimodal interactive control method based on a multifunctional teaching assistant robot according to claim 1, characterized in that, The system startup steps are as follows: The tablet AI processor starts the Android system and initializes the display module, human-computer interaction module, audio input module, audio output module, image acquisition module, temperature and humidity sensor module, IoT module, WIFI or Bluetooth module and Ethernet module; The audio broadcasting steps are followed by: Video processing steps: The image acquisition module acquires video images and performs video call processing or face recognition processing to generate identity results; Environmental monitoring steps: Collect ambient temperature and humidity data through a temperature and humidity sensor module and perform threshold judgment to generate environmental status information; Network interaction steps: Interact with the server via WIFI, Bluetooth or Ethernet module to upload the identity result or the environmental status information and receive control commands; Device linkage steps: Send device control commands to external smart education devices or smart home devices through the IoT module and receive device status feedback.

3. The multimodal interactive control method based on a multifunctional teaching assistant robot according to claim 2, characterized in that, In the speech processing step, phase transformation weighting is performed on the multi-channel speech signal and directional response power calculation is performed based on the candidate direction angle to obtain the sound source direction estimation result; After obtaining the sound source direction estimation result, the method further includes constructing a steering vector based on the sound source direction estimation result and performing delayed summation beamforming to obtain an enhanced speech signal; Between the speech processing step and the speech interaction step, there are also echo cancellation step, speech activity detection step, and frequency domain noise suppression step.

4. The multimodal interactive control method based on a multifunctional teaching assistant robot according to claim 3, characterized in that, Performing the face recognition process in the video processing steps includes the following sub-steps; Face detection and key point localization steps: Perform face detection on the acquired video images and output a set of key point coordinates; Face alignment step: Based on the set of key point coordinates, rotate, scale, and translate the face image to obtain an aligned face image; Feature extraction steps: Input the aligned face image into a deep feature extraction network to obtain feature vectors and perform L2 normalization; Feature matching steps: Perform similarity calculation between the normalized feature vector and the normalized feature vector in the feature library, and output the identity result or rejection result based on the threshold.

5. The multimodal interactive control method based on a multifunctional teaching assistant robot according to claim 4, characterized in that, In the environmental monitoring step, when the ambient temperature is greater than or equal to 60 degrees Celsius or the ambient humidity is greater than or equal to 90% relative humidity, an alarm event is generated and the audio broadcasting step and the interactive presentation step are triggered to output alarm information. In the device linkage step, the device control command is sent to the wired device via RS485 communication or to the wireless device via Zigbee communication.

6. A multimodal interactive control system based on a multifunctional teaching assistant robot, comprising: System startup unit: Used by the tablet AI processor to start the Android system and initialize the display module, human-computer interaction module, audio input module, audio output module, image acquisition module, temperature and humidity sensor module, IoT module, WIFI or Bluetooth module and Ethernet module; Interactive presentation unit: used to output the human-computer interaction interface through the display module and receive touch interaction input information or button interaction input information from the human-computer interaction module; Voice acquisition unit: used to acquire multi-channel voice signals through the audio input module and transmit them to the tablet AI processor; Speech processing unit: used to perform short-time Fourier transform on multi-channel speech signals and perform beamforming based on steering vectors to obtain enhanced speech signals; Voice interaction unit: used to perform voice wake-up or voice recognition based on the enhanced voice signal to generate interactive response information; Audio broadcasting unit: used to output voice or multimedia audio corresponding to the interactive response information through the audio output module; Video processing unit: used to acquire video images through the image acquisition module and perform video call processing or face recognition processing to generate identity results; Environmental monitoring unit: used to collect ambient temperature and humidity through temperature and humidity sensor modules and perform threshold judgment to generate environmental status information; Network interaction unit: used to interact with the server via WIFI, Bluetooth or Ethernet module to upload the identity result or the environmental status information and receive control commands; Device linkage unit: Used to send device control commands to external smart education devices or smart home devices and receive device status feedback via IoT module.

7. A multifunctional teaching assistant robot, characterized in that, include: A tablet AI processor, used to run the Android system and carry peripheral function modules; The display module is connected to the tablet AI processor and is used to display the human-computer interaction interface; The human-computer interaction module is connected to the tablet AI processor and is used to output touch interaction input information or button interaction input information; An audio output module, connected to the tablet AI processor, is used to output multimedia audio or human-computer interaction voice. An image acquisition module, connected to the tablet AI processor, is used to acquire video images to support video calls or facial recognition; An audio input module, connected to the tablet AI processor, is used to collect voice signals to support voice wake-up or voice interaction.

8. The multifunctional teaching assistant robot according to claim 7, characterized in that, Also includes: The Internet of Things (IoT) module is connected to the tablet AI processor and is used to communicate with external smart education devices or smart home devices to achieve device linkage. A WIFI or Bluetooth module is connected to the tablet AI processor to connect to the wireless network and interact with the server for data exchange. An Ethernet module, connected to the tablet AI processor, is used to connect to a wired network and interact with the server for data exchange; A temperature and humidity sensor module is connected to the tablet AI processor to detect ambient temperature and humidity.

9. The multifunctional teaching assistant robot according to claim 8, characterized in that, The audio input module includes a six-element circular microphone array, the array radius of which is in the range of 30 mm to 45 mm; The tablet AI processor samples the multi-channel speech signal of the six-element ring microphone array at a sampling rate of 48000 Hz and performs a short-time Fourier transform. The frame length of the short-time Fourier transform is 1024 sampling points and the frame shift is 256 sampling points. The tablet AI processor performs SRP-PHAT sound source direction estimation on the multi-channel complex spectrum observation obtained by the short-time Fourier transform to obtain candidate sound source direction angle information, and constructs a steering vector based on the candidate sound source direction angle information and performs delayed summation beamforming to output an enhanced speech signal. The tablet AI processor sequentially performs echo cancellation processing, voice activity detection processing, and frequency domain noise suppression processing on the enhanced voice signal, and uses the processed voice signal for the voice wake-up interaction. The IoT module includes an RJ12 interface for RS485 communication and a Zigbee gateway module for Zigbee communication. The Zigbee gateway module supports 802.15.4 MAC and PHY and operates on channels 11 to 26 in the 2.400GHz to 2.483GHz frequency band, with an air interface rate of 250Kbps and supports AES128 or AES256 hardware encryption.

10. The multifunctional teaching assistant robot according to claim 9, characterized in that, The temperature and humidity sensor module includes: The SHT40 temperature and humidity sensor is connected to the tablet AI processor via the IIC interface and is used to periodically report ambient temperature and humidity sampling values. The threshold judgment module is electrically connected to the SHT40 temperature and humidity sensor and has a built-in threshold judgment program. The threshold judgment program is used to generate an alarm event when the ambient temperature sampling value is greater than or equal to 60 degrees Celsius or the ambient humidity sampling value is greater than or equal to 90% relative humidity. The alarm interface is output through the display module and the alarm voice is output through the audio output module. The tablet AI processor is configured to perform face recognition processing as follows: Perform face detection on the images acquired by the image acquisition module and output a set of key point coordinates to generate alignment transformation parameters; Based on the alignment transformation parameters, the face image is rotated, scaled, and translated to obtain an aligned face image; The aligned face image is input into a deep feature extraction network to obtain a feature vector, and the feature vector is normalized by L2 to obtain a normalized feature vector. The cosine similarity score is calculated based on the normalized feature vector and the normalized feature vector in the feature library, and the identity result or rejection result is output based on the threshold.

Citation Information

Cited By

  • AI teaching assistant robot supporting synchronization and interaction of multiple paths of media streams

    CN122024544A

  • Mechanical arm security control system for fixed wrench production

    CN122100190A