AI teaching assistant robot supporting synchronization and interaction of multiple paths of media streams

The AI ​​teaching assistant robot, optimized through multimodal timestamp alignment, ant colony algorithm, and Q-learning algorithm, solves the problems of timestamp disorder and resource competition in the acquisition of multiple audio and video streams. It achieves accurate alignment of speaker identity with voice content and stability of interaction, and improves the response speed and smoothness of multi-person turn-taking interaction.

CN122024544APending Publication Date: 2026-05-12HEBEI HUAFA EDUCATION TECH CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610427967.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-02
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In a classroom setting, disordered timestamps from multiple audio and video streams make it difficult for the robot to accurately align the speaker's identity with the voice content, resulting in abnormal interactive feedback, intermittent voice pickup, and loss of eye contact. When multiple media streams interact concurrently, resource competition is fierce, leading to delayed command response or frame drops, which affects the interactive experience of multiple people taking turns naturally.

Method used

A multimodal timestamp alignment algorithm is used to fuse optical flow and phase cross-correlation audio delay estimation to dynamically calibrate hardware clock drift; an ant colony algorithm is introduced to optimize microphone array beamforming and visual Kalman filter for joint audio-visual tracking; a Q-learning algorithm is used to dynamically schedule resource priorities and combined with a visual attention mechanism to perceive interaction hotspots to ensure accurate resource allocation.

Benefits of technology

It achieves precise alignment of multiple audio and video streams at the acquisition source, ensuring accurate binding of speaker identity and voice content, continuity of voice pickup and eye-tracking, reducing resource competition, and improving the response speed and smoothness of multi-person interactive sessions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024544A_ABST
    Figure CN122024544A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent interaction robots, and particularly discloses an AI teaching-aiding robot supporting synchronization and interaction of multiple media streams, which comprises a robot body, an AI teaching-aiding interaction system is arranged in the robot body, and the AI teaching-aiding interaction system comprises a streaming media synchronization module, a streaming media interaction module, a streaming media interaction module, a streaming media interaction module and a streaming media interaction module, carrying out deep fusion on the video frame motion characteristics based on the optical flow method and the audio time delay estimation based on phase cross-correlation; according to the invention, through a multi-modal timestamp alignment algorithm, combined modeling is carried out on lip motion features based on an optical flow method and audio time delay estimation based on phase cross-correlation in a deep fusion network, hardware clock drift between a camera and a microphone array is learned and compensated, the problem of disorder timestamps of multiple paths of audio and video streams is solved, and the robustness of the system is improved. Each frame of image is accurately aligned with the corresponding voice segment on the time axis, and it is ensured that the robot can accurately capture the complete expression intention of the spokesman in the complex classroom environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent interactive robot technology, and in particular to an AI teaching assistant robot that supports the synchronization and interaction of multiple media streams. Background Technology

[0002] With the advancement of technology, traditional teaching models are facing higher demands for efficiency, interactivity, and personalized educational services. AI teaching assistant robots utilize artificial intelligence technology to simulate teachers' teaching behaviors through methods such as deep learning and natural language processing, providing students with intelligent tutoring.

[0003] For example, the multimodal interactive control method, system and robot of the multifunctional teaching assistant robot in Chinese Patent Publication No. CN121572277A can still complete the instruction input and information feedback through other interactive channels when a certain sensing channel is affected by environmental noise, changes in lighting or people blocking it. This forms a multimodal collaboration and redundancy mechanism in actual teaching scenarios, and improves the continuity and stability of the human-computer interaction process.

[0004] In existing technologies, when the timestamps of multiple audio and video streams are disordered in a classroom setting, it becomes difficult for the robot to accurately align the speaker's identity with the voice content, resulting in abnormal interactive feedback. Furthermore, it is difficult to track the fixed beam direction of a moving speaker in real time, leading to intermittent voice pickup and loss of eye contact. In addition, during the joint audio-visual tracking process, there is a problem of intense resource competition when multiple media streams interact concurrently, which can further lead to instruction response delays or frame drops, affecting the interactive experience of natural rotation among multiple users. To address these issues, we propose an AI teaching assistant robot that supports the synchronization and interaction of multiple media streams. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, the present invention provides an AI teaching assistant robot that supports the synchronization and interaction of multiple media streams, which can effectively solve the problems involved in the prior art.

[0006] The objective of this invention can be achieved through the following technical solution: This invention provides an AI teaching assistant robot that supports simultaneous and interactive multi-channel media streams, including a robot body. The robot body has a built-in AI teaching assistant interaction system, which includes the following modules: The streaming media synchronization module is used to employ a multimodal timestamp alignment algorithm to deeply fuse video frame motion features based on optical flow with audio delay estimation based on phase cross-correlation. This enables simultaneous locking of the speaker's identity and voice content at the source of acquisition, achieving microsecond-level synchronization at the audio and video acquisition source and eliminating audio-visual misalignment caused by timestamp disorder. The audio-visual joint tracking module is used to dynamically optimize the beamforming direction of the microphone array by introducing an ant colony algorithm on the basis of accurate audio-visual synchronization, and to predict the target motion trajectory by integrating a visual Kalman filter. This drives the robot's audio-visual perception function to achieve joint active tracking, ensuring the continuity of voice pickup and the following of the gaze. The priority learning and scheduling module is used to address resource contention when multiple media streams are running concurrently. It employs the Q-learning algorithm in reinforcement learning to dynamically schedule the processing priorities of various data streams. Through trial and error learning, it obtains the optimal scheduling strategy, thereby minimizing resource contention through autonomous learning. The interactive hotspot perception and allocation module is used to integrate the visual attention mechanism to analyze the visual focus in the scene, perceive the current interactive hotspot, and use it as the state input of Q-learning to guide the scheduling algorithm to allocate resources to the real interactive subject, ensuring low latency response and frame-free interactive experience in multi-person natural rotation scenarios. The control module receives audio-visual data, physical state data, and resource data from each module, constructs a complete context, makes interactive decisions, and translates them into specific control commands to obtain stable resource support. This ensures that the robot's feedback is timely, natural, and coherent in complex interactions involving multiple people taking turns.

[0007] Preferably, the streaming media synchronization module includes a multimodal timestamp alignment unit and an audio-visual identity synchronization locking unit; The multimodal timestamp alignment unit is used to deeply fuse video frame motion features based on optical flow and audio delay estimation algorithm based on phase cross-correlation at the source of audio and video acquisition. By analyzing the time difference between lip micro-movement and sound wave arrival, it dynamically calibrates timestamps from different hardware (camera, microphone), eliminates timestamp disorder from multiple streams, and aligns each frame of image with the corresponding speech segment in time, thereby ensuring strict alignment between each frame of image and the corresponding speech segment. The audio-visual identity synchronization locking unit binds synchronized audio features (voiceprint) and video features (face / lip shape) based on the accurate synchronization data provided by the multimodal timestamp alignment unit. Through joint feature locking, it associates the speaker's identity with the voice content, thereby achieving accurate binding between the speaker and the voice content and eliminating identity recognition errors in multi-person scenarios.

[0008] Preferably, the multimodal timestamp alignment unit performs the following steps: The video frame sequence output by the robot's camera and the audio stream picked up by the microphone array are acquired simultaneously. The optical flow motion feature vector of the mouth region in each frame image and the phase information of each audio segment in the frequency domain are extracted respectively. This lays the foundation for subsequent accurate analysis of the temporal correspondence between lip movement and speech, and ensures the synchronicity and completeness of feature extraction. The extracted optical flow motion features and the audio delay estimate calculated based on the phase cross-correlation algorithm are input into the deep fusion network. By analyzing the difference between the start time of lip movement and the arrival time of sound waves, a dynamic time offset model is constructed, enabling the network to autonomously learn the drift pattern of the hardware clock, generate accurate time offset correction, and achieve dynamic matching of audio and video timing. By using a dynamic time offset model to calibrate the hardware clock drift of the camera and microphone in real time, the timestamps of video frames and audio packets are dynamically compensated, eliminating timestamp disorder at the source of multiple streams, eliminating audio and video asynchrony from the root, and ensuring that subsequent identity locking and tracking functions operate stably based on aligned data streams.

[0009] Preferably, the audio-visual identity synchronization locking unit performs the following steps: Receive the synchronized audio and video streams output by the multimodal timestamp alignment unit, extract the voiceprint feature vectors from the synchronized audio segments and the face / mouth structure features from the corresponding video frames, and ensure that the core biometric features that can be used for identity verification are accurately extracted from the source synchronized audio and video data. Construct a cross-modal feature joint embedding space, measure the similarity and match the voiceprint features with the mouth shape dynamic features to form an audio-visual identity binding mapping relationship, realize the deep association between voice and facial dynamics, and establish a stable and unique identity credential. The system locks the current speaker's identity based on the binding mapping relationship and associates the collected voice content with the corresponding identity identifier in real time, so that the speaker and the voice content are accurately matched, completely solving the problems of identity confusion and voice content mismatch in multi-person interaction scenarios.

[0010] Preferably, the audio-visual joint tracking module includes a beamforming dynamic optimization unit and a vision-motion prediction driving unit; The beamforming dynamic optimization unit is used to dynamically optimize the beamforming direction of the microphone array in response to the random movement of the speaker by introducing an ant colony algorithm. In a complex sound field environment, it can quickly find the optimal path of the sound source and dynamically adjust the sound pickup direction to ensure that clear and high signal-to-noise ratio speech can be continuously captured even when multiple people are moving around. This enables rapid locking and continuous tracking of the sound source in a dynamic sound field, ensuring the continuity of speech pickup. The vision-motion prediction drive unit is used to predict the target's motion trajectory by integrating a visual Kalman filter while tracking audio, and to drive the robot's chassis to rotate smoothly in advance to achieve joint active tracking. This enables the robot to follow smoothly and predictively, and completely solves the problems of intermittent voice pickup and line-of-sight loss.

[0011] Preferably, the beamforming dynamic optimization unit performs the following steps: Initialize the beamforming parameters of the microphone array, regard the beam pointing in each direction as the search path in the ant colony algorithm, use the signal-to-noise ratio of the picked signal as the pheromone concentration evaluation index, establish an initial search space for rapid sound source localization, and ensure that the beamforming algorithm has global optimization capability. Simulating the movement of ant colonies in a sound field environment, through pheromone updates and path selection probability calculations, iterative search enables the beamforming direction to converge rapidly to the optimal direction of the current sound source, simulating the intelligent behavior of biological groups, quickly locking the sound source location in complex sound fields, and improving the convergence speed of beam pointing. The array weighting coefficients are dynamically adjusted based on the optimal beam direction output by the ant colony algorithm. The pickup beam direction is updated in real time in environments with multiple people moving around and noise interference, continuously capturing high signal-to-noise ratio speech. The pickup beam direction is adjusted in real time to ensure that high-quality speech signals can still be stably captured when the sound source moves.

[0012] Preferably, the vision-motion prediction driving unit performs the following steps: The sound source region locked by the receiving beamforming dynamic optimization unit is simultaneously acquired by the target speaker image sequence collected by the visual sensor, and the position information of the target in the image coordinate system is extracted to ensure that the sound source localization and the spatial position of the visual target are accurately corresponded. The target position sequence is input into the visual Kalman filter, and the position and velocity of the next moment are estimated and predicted by combining the preset target motion model. Smooth motion trajectory prediction data is generated, effectively filtering out detection jitter interference and outputting a continuous and stable motion trajectory. Based on motion trajectory prediction data, the required rotation angle and angular velocity of the robot chassis are calculated in advance, and smooth following instructions are generated to drive the robot to perform joint active tracking, realizing the coordinated control of sound and image perception and motion, enabling the robot to smoothly follow the moving speaker, and eliminating line-of-sight loss and intermittent sound pickup.

[0013] Preferably, the priority learning and scheduling module performs the following steps: The various media stream types that coexist in the system are defined as a set of scheduling objects. An initial priority queue is set for each type of stream, and the state space and action space of the Q-learning algorithm are initialized to establish a basic data model and decision framework for subsequent dynamic priority adjustment. Real-time monitoring of data volume, processing latency, and system resource utilization of each media stream; using the current resource status as the input state of the Q-learning algorithm; selecting priority adjustment actions based on the ε-greedy strategy; achieving real-time perception and adaptive priority adjustment of system load changes. After performing the priority adjustment action, the reward value is calculated based on the changes in processing latency and frame loss. The scheduling strategy is optimized by iteratively updating the Q table, gradually converging to the optimal priority allocation scheme that minimizes resource contention, ensuring that critical media streams are given priority processing when resources are scarce, and guaranteeing smooth interaction.

[0014] Preferably, the interactive hotspot sensing and allocation module performs the following steps: The system collects classroom scene images from cameras, extracts salient areas in the scene through a visual attention mechanism, analyzes the distribution of visual focus and the direction of human eye gaze in each area, and realizes real-time visual attention perception in classroom scenes with multiple students, accurately locating the areas where students' attention is concentrated. By integrating the visual focus distribution characteristics with the speaking status in each area, the interactive hot spots and corresponding hot figures in the current scene are identified, the popularity weight of each interactive subject is quantified, and the passive listeners and active interactors in the classroom are effectively distinguished, thus avoiding non-focus personnel from occupying system resources. The popularity weights of each interactive entity are transmitted as state inputs to the priority learning and scheduling module, guiding the Q-learning algorithm to allocate computing and transmission resources to the most popular real interactive entities, ensuring that the audio and video data of popular figures are processed first, and improving the real-time interaction response in multi-person rotation scenarios.

[0015] Preferably, the control module performs the following steps: By collecting the output audio-visual binding data, physical status data, and resource allocation data, a classroom interaction context containing multi-dimensional information is constructed to ensure that the interaction decision engine obtains millisecond-level consistent comprehensive status perception, laying the foundation for accurate decision-making. The constructed interactive context is input into the interactive decision engine, which generates interactive feedback content to be executed and corresponding robot action sequences based on the preset teaching assistant interaction rules, thereby realizing rule-driven intelligent and scenario-based interactive response generation in the classroom setting. The system requests stable resource guarantees for the execution of action sequences from the priority learning and scheduling module, converts interactive feedback content into specific voice broadcast commands and chassis movement commands for execution, and achieves coherent and natural multi-person rotation interactive feedback. This ensures that interactive commands are executed first in resource contention and achieves coherent and smooth physical feedback.

[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. This AI teaching assistant robot supports the synchronization and interaction of multiple media streams. Through a multimodal timestamp alignment algorithm, it jointly models lip movement features based on optical flow and audio delay estimation based on phase cross-correlation in a deep fusion network. It dynamically learns and compensates for hardware clock drift between the camera and microphone array, solving the problem of timestamp disorder of multiple audio and video streams from the source of acquisition. This ensures that each frame of image and the corresponding speech segment are accurately aligned on the timeline, ensuring that the robot can accurately capture the speaker's complete expressive intent in complex classroom environments.

[0017] 2. This AI teaching assistant robot, which supports simultaneous and interactive multi-channel media streams, introduces an ant colony algorithm to dynamically optimize the beamforming direction of the microphone array. By simulating the foraging behavior of ants, it quickly converges to the optimal direction of the sound source in a complex sound field environment. At the same time, it integrates a visual Kalman filter to predict the speaker's movement trajectory and drive the robot chassis to rotate smoothly in advance. This effectively solves the problem of fixed beams being difficult to follow in scenarios with moving speakers, ensuring the continuity of voice pickup and the following of the camera's line of sight. It keeps the speaker in the main lobe of the sound pickup beam and the center of the field of vision, ensuring the stability of the interaction process.

[0018] 3. This AI teaching assistant robot, which supports simultaneous and interactive multi-channel media streams, employs the Q-learning algorithm to dynamically prioritize and schedule concurrent media streams of various types. It uses system resource status as the state space and priority adjustment as the action space. By handling latency and frame loss, a reward function is constructed. Through iterative learning, it converges to the optimal allocation strategy that minimizes resource contention. This allows for intelligent balancing of computing, memory, and bandwidth resources during concurrent interaction of multiple audio / video streams, control streams, and command streams, effectively avoiding command response latency and data frame loss, and providing stable resource guarantees for real-time interaction. Attached Figure Description

[0019] Figure 1 This is a schematic diagram illustrating the workflow of an AI teaching assistant robot that supports multi-channel media stream synchronization and interaction according to the present invention. Figure 2 This is a schematic diagram of the module data flow of an AI teaching assistant robot that supports multi-channel media stream synchronization and interaction according to the present invention; Figure 3 This is a schematic diagram of the external structure of an AI teaching assistant robot that supports multi-channel media stream synchronization and interaction according to the present invention.

[0020] In the picture: 1. Robot body. Detailed Implementation

[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0022] Example 1, please refer to Figure 1 , Figure 2 This invention provides a technical solution: an AI teaching assistant robot that supports simultaneous and interactive multi-channel media streams, comprising a robot body 1, wherein the robot body 1 has a built-in AI teaching assistant interaction system, the AI ​​teaching assistant interaction system comprising the following modules: The streaming media synchronization module is used to employ a multimodal timestamp alignment algorithm to deeply fuse video frame motion features based on optical flow with audio delay estimation based on phase cross-correlation. This enables simultaneous locking of speaker identity and voice content at the source of audio and video acquisition, achieving microsecond-level synchronization at the source of audio and video acquisition and eliminating audio-visual misalignment caused by timestamp disorder. The streaming media synchronization module includes a multimodal timestamp alignment unit and an audio-visual identity synchronization locking unit. The multimodal timestamp alignment unit is used at the source of audio and video acquisition to deeply fuse video frame motion features based on optical flow and audio delay estimation algorithms based on phase cross-correlation. By analyzing the time difference between lip micro-movement and sound wave arrival, it dynamically calibrates timestamps from different hardware (camera, microphone), eliminates timetamp disorder from multiple streams, and aligns each frame of image with its corresponding audio segment in time. This ensures strict alignment between each frame of image and its corresponding audio segment, providing a precise time reference for subsequent identity binding. Simultaneously, it acquires video frame sequences output by the robot's camera and audio streams picked up by the microphone array, extracting optical flow motion feature vectors of the mouth region in each frame of image, as well as phase information of each audio segment in the frequency domain, providing a basis for subsequent accurate analysis of lip movements. The time correspondence between motion and speech lays the foundation, ensuring the synchronization and integrity of feature extraction. The extracted optical flow motion features and the audio delay estimate calculated based on the phase cross-correlation algorithm are input into the deep fusion network. By analyzing the time difference between the start of lip movement and the arrival time of sound waves, a dynamic time offset model is constructed, enabling the network to autonomously learn the drift pattern of the hardware clock and generate accurate time offset correction. This achieves dynamic matching of audio and video timing. The dynamic time offset model is used to calibrate the hardware clock drift of the camera and microphone in real time, dynamically compensate for the timestamps of video frames and audio packets, eliminate timestamp disorder at the acquisition source of multiple streams, eliminate audio and video asynchrony from the root, and ensure that subsequent identity locking and tracking functions operate stably based on aligned data streams. It should be noted that the robot's binocular vision camera captures video sequences at a frame rate of 30 frames per second, with a resolution of 1920×1080 pixels. Simultaneously, the microphone array picks up the audio stream at a sampling rate of 16kHz and a quantization precision of 16bit. First, face detection and alignment are performed on the captured video frame sequence to locate the region of interest (ROI) around the mouth. Then, the Farneback dense optical flow algorithm is used to calculate the motion vector field of the mouth region between adjacent frames, extracting a 64-dimensional optical flow feature vector representing the lip movement trajectory. At the same time, the audio stream is processed by frame segmentation, with the frame length set... With a frame shift of 16 milliseconds and a 32-millisecond frame rate, a Fast Fourier Transform (FFT) is performed on each audio frame to obtain the frequency domain phase spectrum. The arrival time difference of the audio signals between each microphone channel is then calculated using a phase cross-correlation algorithm to obtain an initial audio delay estimate. A deep fusion network receives the optical flow feature vector and the audio delay estimate as input. This network employs a bidirectional Long Short-Term Memory (LSTM) structure, containing two hidden layers with 128 neurons each. By analyzing the sequence of lip movement initiation and the time difference between the sound wave arrival at the microphone, the network learns the dynamic clock offset between the camera and the audio acquisition hardware. Specifically, the network output layer adopts a fully connected structure, regresses to generate the time offset correction amount at the current moment, and sets the time offset correction update frequency to 10 times per second to adapt to the slow change characteristics of hardware clock drift. During the training phase, mean squared error is used as the loss function, and supervised learning is performed using a dataset containing 5000 sets of lip movement-speech alignment samples to ensure that the model's estimation error of the time delay between the lip movement start point and the sound wave arrival time is less than 5 milliseconds. Based on the correction amount output by the dynamic time offset model, the timestamps of video frames and audio packets are dynamically compensated. The video frame timestamp compensation accuracy is controlled within 1 millisecond, and the audio packet timestamp compensation is based on audio frames. The compensation amount corresponding to each audio frame is calculated independently. The compensated audio and video streams are temporarily stored in a circular buffer with a capacity designed to accommodate 2 seconds of synchronization data to cope with network jitter or processing delay. The timestamp calibration process is supplemented by a hardware clock synchronization protocol. The initial value of the clock deviation between the camera and microphone array is obtained by periodically exchanging synchronization messages. The dynamic time offset model is fine-tuned and corrected on this basis to achieve a synchronization error of less than one video frame interval at the acquisition source for multiple streams. The audio-visual identity synchronization and locking unit, based on the accurate synchronization data provided by the multimodal timestamp alignment unit, binds the synchronized audio features (voiceprint) with video features (face / lip shape). Through joint feature locking, it associates the speaker's identity with the voice content, thereby achieving accurate binding between the speaker and the voice content, eliminating identity recognition errors in multi-person scenarios. It receives the synchronized audio and video streams output by the multimodal timestamp alignment unit, extracts the voiceprint feature vector from the synchronized audio segment and the face / lip shape structural features from the corresponding video frame, ensuring that the core biometric features that can be used for identity authentication are accurately extracted from the source synchronized audio and video data. It constructs a cross-modal feature joint embedding space, performs similarity measurement and association matching between voiceprint features and dynamic lip shape features, and forms an audio-visual identity binding mapping relationship. This achieves a deep association between voice and facial dynamics, establishes a stable and unique identity credential, locks the current speaker's identity based on the binding mapping relationship, and associates the collected voice content with the corresponding identity identifier in real time, so that the speaker and the voice content correspond accurately, completely solving the problems of identity confusion and voice content mismatch in multi-person interaction scenarios. It should be noted that the system receives audio and video streams output from the multimodal timestamp alignment unit. The video stream is a 30fps, 1920×1080 resolution RGB image sequence, and the audio stream is PCM data with a 16kHz sampling rate and 16-bit quantization precision. Face detection is performed on each frame of the video stream, and after locating the key mouth region, a 3D convolutional neural network is used to extract a 128-dimensional dynamic feature vector representing lip movement. Simultaneously, the audio stream is processed into frames, with a frame length of 32ms and a frame shift of 16ms. Each frame is processed rapidly. After Fourier transform, 40-dimensional Mel-frequency cepstral coefficients are extracted and mapped to a 512-dimensional voiceprint feature vector using a Gaussian mixture model-general background model algorithm, ensuring speaker-discriminatory and environment-robust voiceprint features. The extracted 128-dimensional lip-shape dynamic features and 512-dimensional voiceprint features are input into a two-stream neural network architecture. Two fully connected layers map the two types of features to a 256-dimensional common embedding space. In this space, cosine similarity is used to calculate the matching score between lip-shape features and voiceprint features, with a similarity threshold of 0.75. When the score exceeds the threshold, it is determined to be the same speaker. For audio and video segments that meet the matching conditions for 30 consecutive frames (corresponding to 1 second in duration), an audio-visual identity binding mapping relationship is established, and the speaker's unique ID and corresponding voiceprint template and mouth shape dynamic template are recorded to form a temporary speaker registration library in the classroom scenario. Based on the established audio-visual identity binding mapping relationship, the input synchronized audio and video stream is locked in real time. When a new voice activity is detected, the voiceprint features of the current 1-second audio segment are extracted and compared with the similarity of all speaker templates in the registration library. The ID with the highest similarity and greater than 0.75 is selected as the current speaker. At the same time, the mouth area image and audio waveform data of the speaker in the video frame within the corresponding time window are packaged to generate a structured message containing identity, timestamp, audio and video data, which is transmitted to the upper-layer interaction module. This process runs continuously and supports dynamic updates of the registration library. When a new speaker speaks for 10 consecutive seconds and cannot match the existing template, a new ID is automatically created for him and added to the registration library to achieve accurate identity tracking in the scenario of multiple people taking turns. The audio-visual joint tracking module is used to dynamically optimize the beamforming direction of the microphone array by introducing an ant colony algorithm on the basis of precise audio-visual synchronization, and to predict the target motion trajectory by integrating a visual Kalman filter. This drives the robot's audio-visual perception function to achieve joint active tracking, ensuring the continuity of voice pickup and the following of the gaze, building an active tracking capability of audiovisual collaboration, and ensuring that the speaker is always in the optimal perception area. The audio-visual joint tracking module includes a beamforming dynamic optimization unit and a vision-motion prediction driving unit. The beamforming dynamic optimization unit is used to dynamically optimize the beamforming direction of the microphone array in response to the random movement of the speaker. This involves incorporating an ant colony algorithm to quickly find the optimal path to the sound source in complex sound environments, dynamically adjusting the pickup direction to ensure continuous capture of clear, high signal-to-noise ratio speech even when multiple people are moving around. This enables rapid locking and continuous tracking of the sound source in dynamic sound fields, ensuring the continuity of speech pickup. The unit initializes the beamforming parameters of the microphone array, treating the beam direction in each direction as the search path in the ant colony algorithm, and using the signal-to-noise ratio of the picked-up signal as the pheromone concentration evaluation index to establish a rapid sound source localization mechanism. An initial search space is established to ensure that the beamforming algorithm has global optimization capabilities. The movement process of ant colonies in a sound field environment is simulated. Through pheromone updates and path selection probability calculations, iterative search enables the beamforming direction to converge quickly to the optimal direction of the current sound source. Simulating the intelligent behavior of biological groups, the beamforming algorithm can quickly lock the location of the sound source in a complex sound field, thereby improving the convergence speed of the beam direction. The weighting coefficients of the array are dynamically adjusted according to the optimal beam direction output by the ant colony algorithm. The pickup beam direction is updated in real time under the conditions of multiple people walking and noise interference, continuously capturing high signal-to-noise ratio speech. The pickup beam direction is adjusted in real time to ensure that high-quality speech signals can still be stably captured when the sound source moves. It should be noted that the microphone array adopts a uniform circular array structure with 8 array elements and an array radius of 4.5 cm to meet the frequency response requirements for voice signal acquisition in a classroom setting. During the initialization phase, the horizontal 360-degree range is divided into 72 candidate beam pointing directions at 5-degree intervals. Each direction corresponds to a set of beamforming weighting coefficients. The initial weighting coefficients are calculated based on a delay accumulation algorithm. Each candidate direction is considered a search path in the ant colony algorithm. The signal-to-noise ratio (SNR) of the signal picked up in that direction is used as the evaluation index for pheromone concentration. The SNR is calculated as the ratio of speech segment energy to noise segment energy. The noise segment is estimated using the non-speech interval identified by the speech activity detector. The initial pheromone concentration is set to 0.1, and the pheromone evaporation coefficient is set to 0.5 to balance the relationship between historical information and the current search, ensuring the algorithm can quickly respond to changes in the sound source location. When simulating the movement of the ant colony in the sound field environment, 20 artificial ants are used. Each ant represents an initial beam pointing attempt. The ants calculate based on the pheromone concentration on each path and heuristic information. The path selection probability is determined by the guiding power of the signal picked up in the current direction, which is obtained by calculating the spatial spectrum of that direction. In each iteration, after the ants complete a round of path selection, the pheromone concentration is updated according to the actual signal-to-noise ratio (SNR) measurement of the selected direction. For every 1 dB increase in SNR, the pheromone increment increases by 0.05. The number of iterations is set to 15. After multiple iterations, the path with the highest pheromone concentration is determined as the optimal beam pointing direction of the current sound source. Based on the optimal beam pointing direction output by the ant colony algorithm, the corresponding microphone array weighting coefficients are calculated in real time. Specifically, a minimum variance distortionless response beamformer is used to minimize the total output power under the constraint of ensuring distortionless output of the signal in the desired direction, thereby forming nulls in the interference direction. The weighting coefficient update frequency is set to once every 50 milliseconds to match the short-term stationary characteristics of the speech signal. In environments with multiple people walking and noise interference, the changes in SNR in each direction are continuously monitored. When the SNR of the optimal pointing direction is lower than the preset threshold of 6 dB, the ant colony algorithm is triggered to search again to ensure that the beam always tracks the movement of the sound source. The vision-motion prediction drive unit is used to predict the target's motion trajectory by integrating a visual Kalman filter while tracking audio. This allows the robot's chassis to rotate smoothly in advance, achieving joint active tracking. This enables the robot to smoothly predictively follow the target, completely solving the problems of intermittent voice pickup and line-of-sight loss. The receiving beamforming dynamic optimization unit locks the sound source region and simultaneously acquires the target speaker image sequence collected by the visual sensor. It extracts the target's position information in the image coordinate system to ensure that the sound source localization and the spatial position of the visual target are accurately matched. The target position sequence is input into the visual Kalman filter, which, combined with a preset target motion model, estimates and predicts the position and velocity at the next moment, generating smooth motion trajectory prediction data. This effectively filters out detection jitter interference and outputs a continuous and stable motion trajectory. Based on the motion trajectory prediction data, it calculates the required rotation angle and angular velocity of the robot's chassis in advance, generating smooth following instructions to drive the robot to perform joint active tracking. This achieves coordinated control of sound and image perception and motion, enabling the robot to smoothly follow the moving speaker and eliminate line-of-sight loss and intermittent sound pickup. It should be noted that the robot's control system receives the azimuth data of the sound source region locked by the beamforming dynamic optimization unit in real time, updating it every 50 milliseconds, specifically pointing to the horizontal angle of the current speaker. Simultaneously, the control system acquires image sequences from a binocular vision camera operating at 30 frames per second and a resolution of 1920×1080 pixels. A target detection algorithm is used to identify and locate the speaker in each frame, extracting their pixel coordinates in the image coordinate system. This is then combined with depth information and converted into distance in a polar coordinate system centered on the robot. With visual azimuth Horizontal angle With visual azimuth At the decision-making level, a fusion verification is performed. When the deviation between the two is less than a preset 10-degree threshold, the target position is confirmed as valid, and its coordinates are input as observations to the subsequent tracking filter. The confirmed target position observation sequence is then input into a visual Kalman filter for state estimation and prediction. This filter employs a constant velocity motion model, and the state vector is defined as follows: ,in Let these be the robot's relative coordinates in a two-dimensional plane. This is the transpose of the vector. For the corresponding linear velocity component, the filter parameters are set to include the process noise covariance matrix and the measurement noise covariance matrix. The update frequency is synchronized with the visual frame rate, i.e., 30 times per second. Through iterative calculations in the prediction and update stages, the filter removes noise caused by detection jitter or sound source localization errors, generates a smooth target motion trajectory, and outputs the predicted position and velocity of the target at the next moment (i.e., 33.3 milliseconds later). Based on the predicted position and velocity of the target output by the Kalman filter, the control system calculates the motion control commands for the robot chassis. Specifically, it first calculates the current azimuth angle of the target relative to the robot. With predicted azimuth This allows us to obtain the desired steering angle of the chassis. To achieve smooth following, the robot employs a proportional-integral-derivative control algorithm, with the proportional coefficient... Set to 1.2, integral coefficient Set to 0.1, differential coefficient Set to 0.05 to convert the steering angle error into chassis steering angular velocity. The unit is radians per second, and it is also based on the target distance. With speed of movement Calculate the robot's linear velocity The unit is meters per second, and the angular velocity and linear velocity commands are sent to the chassis drive motor through the controller local area network bus every 50 milliseconds to realize the coordinated control of sound and image perception and motion, ensuring that the speaker is always in the main lobe of the robot's sound pickup beam and the center of the field of vision. The priority learning scheduling module is used to address resource contention when multiple media streams are running concurrently. It employs the Q-learning algorithm in reinforcement learning to dynamically schedule the processing priorities of various data streams. Through trial and error learning, it obtains the optimal scheduling strategy, thereby minimizing resource contention through autonomous learning and fundamentally reducing instruction response latency and frame drop rate. The interactive hotspot perception and allocation module is used to integrate visual attention mechanism to analyze the visual focus in the scene, perceive the current interactive hotspot, and use it as the state input of Q-learning to guide the scheduling algorithm to allocate resources to the real interactive subject, ensuring low latency response and frame-free interactive experience in multi-person natural rotation scenarios. In this way, resources are accurately allocated to the interactive focus to ensure the smoothness and real-time performance of core interactions. The control module receives audio-visual data, physical state data, and resource data from each module, constructs a complete context, makes interactive decisions, and translates them into specific control commands to obtain stable resource support. This ensures that the robot's feedback is timely, natural, and coherent in complex interactions involving multiple users, enabling the robot to make coherent decisions and respond smoothly in complex interactions, and significantly improving the naturalness of multi-user interactions.

[0023] Example 2, as Figure 1 , Figure 2 As shown, based on Embodiment 1, the present invention provides a technical solution: the priority learning scheduling module performs the following steps: defining multiple media stream types existing concurrently in the system as a set of scheduling objects, setting an initial priority queue corresponding to each type of stream, and initializing the state space and action space of the Q-learning algorithm to establish a basic data model and decision framework for subsequent dynamic priority adjustment, monitoring the data volume, processing latency and system resource occupancy of each media stream in real time, taking the current resource status as the input state of the Q-learning algorithm, selecting priority adjustment actions according to the ε-greedy strategy, realizing real-time perception and adaptive priority adjustment of system load changes, calculating the reward value according to the processing latency change and frame drop situation after executing the priority adjustment action, optimizing the scheduling strategy by iteratively updating the Q table, and gradually converging to the optimal priority allocation scheme that minimizes resource contention, ensuring that key media streams are given priority processing when resources are scarce, and ensuring smooth interaction; It should be noted that the media stream types include synchronized audio and video streams, control streams for audio-visual joint tracking, interactive decision command streams, and system log streams, totaling four categories. For each type of media stream, a corresponding initial priority queue is set, with the initial priority of audio and video streams set to the highest level (3), control streams to the second highest level (2), command streams to the second highest level (2), and log streams to the lowest level (1). Simultaneously, the state space and action space of the Q-learning algorithm are initialized. The state space consists of the real-time data volume, processing latency, and system resource utilization of each media stream, with a 12-dimensional dimension. The action space is defined as priority adjustment operations, including raising the priority by one level, lowering the priority by one level, or maintaining the current priority. The algorithm remains unchanged, offering a total of 12 selectable actions. The discretization granularity of the state space and action space has been optimized to ensure real-time operation on the embedded platform. The initial Q-table is set to zero, awaiting subsequent iteration updates. In the real-time monitoring and decision execution phase, the running status data of each media stream is collected every 100 milliseconds. Specifically, the data volume of the audio and video streams is obtained through interfaces provided by the operating system. Processing latency is calculated using timestamps, recording the time elapsed from data enqueueing to processing completion in milliseconds. System resource utilization includes CPU utilization percentage and memory utilization percentage. These multi-dimensional states are concatenated to form the current state vector. As input to the Q-learning algorithm, priority adjustment actions are selected according to the ε-greedy strategy. The exploration rate ε is initially set to 0.3, gradually decreasing to 0.05 with the number of iterations to balance exploration and utilization. After selecting an action, the thread priority or queue weight of the corresponding media stream is dynamically adjusted through system calls to achieve resource reallocation. In the reward calculation and strategy optimization phase, after each priority adjustment action, the system waits for the next monitoring cycle to arrive and collects the adjusted system state. The reward value calculation comprehensively considers the change in processing latency and frame drop: for every 10 milliseconds of processing latency reduction, the reward increases by 0.1; for every 10 milliseconds of latency increase... A penalty of -0.1 is added every 10 milliseconds; frame loss is measured by the number of data packets lost per second, with a reward of +0.5 for no frame loss and a penalty of -0.2 for each lost data packet. The combined reward value is used to update the Q value of the corresponding state-action pair in the Q table. The update formula adopts the standard Q-learning algorithm, with a learning rate set to 0.1 and a discount factor γ set to 0.9. When the reward value fluctuation is less than 0.01 after 500 iterations or 50 consecutive iterations, the algorithm is considered to have converged. The priority allocation strategy obtained at this time is the optimal solution that minimizes resource contention, and the algorithm enters the fixed strategy execution phase, triggering a new round of learning only when the resource state changes significantly. The interactive hotspot perception and allocation module performs the following steps: it acquires classroom scene images from cameras, extracts salient areas in the scene through a visual attention mechanism, analyzes the distribution of visual focus and the direction of human eye gaze in each area, realizes real-time visual attention perception in a multi-person classroom setting, accurately locates the students' attention concentration areas, integrates the visual focus distribution characteristics with the speaking status in each area, identifies interactive hotspot areas and corresponding hotspots in the current scene, quantifies the heat weight of each interactive subject, effectively distinguishes between passive listeners and active interactors in the classroom, avoids non-focus personnel occupying system resources, and transmits the heat weight of each interactive subject as a state input to the priority learning and scheduling module, guiding the Q-learning algorithm to allocate computing and transmission resources to the real interactive subject with the highest heat, ensuring that the audio and video data of the hotspots are processed first, and improving the real-time interactive response in multi-person rotating scenarios. It should be noted that the classroom scene images acquired by the camera are used to extract salient regions in the scene through a visual attention mechanism. Specifically, the binocular vision camera acquires a sequence of 1920×1080 pixel RGB images at a rate of 30 frames per second. Each frame is input into a deep learning-based saliency detection network. This network adopts an encoder-decoder structure. The encoder part uses a ResNet-50 network pre-trained on the ImageNet dataset, and the decoder part gradually restores the spatial resolution through deconvolution layers. Finally, it outputs a saliency probability map with the same resolution as the input image. The value of each pixel in the saliency probability map represents the probability that the location will attract visual attention. Based on this, connected component analysis is performed on the saliency probability map, and connected regions with an area greater than 100 pixels are extracted as candidate saliency regions. For each saliency region, an eye keypoint detection algorithm is used to locate the coordinates of the center of the person's eye and the center of the pupil within the region. The direction of the line connecting the two is calculated to obtain the direction vector of the person's gaze. At the same time, the total number of people in the region and the number of people gazing in the same direction are counted to form a visual focus distribution feature vector. The visual focus distribution feature is fused with the speaking status in each region to identify the interactive hotspots and corresponding hot figures in the current scene. Specifically, the azimuth angle of the current speaker is obtained in real time from the beamforming dynamic optimization unit. According to the data, which is updated every 50 milliseconds, the speech state is spatially aligned with the visual focus by combining the salient region location information output by the visual attention mechanism. For each salient region, the confidence level of the speech state within that region is calculated. The confidence level is determined by whether the region contains the current sound source localization result and the distance between the sound source localization result and the center point of the region. A distance less than 0.5 meters is considered a strong correlation, with a confidence level of 0.9; a distance between 0.5 and 1.5 meters is considered a weak correlation, with a confidence level of 0.5; and a distance greater than 1.5 meters is considered unrelated, with a confidence level of 0.1. The visual focus distribution characteristics are then compared with the speech state confidence level. The interaction heat value of each region is calculated by weighted fusion, with the visual focus weight set to 0.4 and the speaking status weight set to 0.6. The region with the highest interaction heat value is selected as the current interaction hotspot region. The person with the highest heat value in this region is the hotspot person. The heat weight is determined by the interaction heat value of the region where the person is located, their gaze participation in the region, as well as the preset base participation coefficient and gaze participation gain coefficient. The gaze participation is quantified by whether the person is looking at the robot or the speaker. A participation value of 1 indicates that the person is looking, a value of 0.5 indicates that the person is looking occasionally, and a value of 0 indicates that the person is not looking. The final heat weight of the hotspot person ranges from 0 to 1. The expression for the popularity weight of trending figures is as follows: ; In the formula: Indicates hotspot areas Insiders The popularity weight is used to quantify the priority of attention given to this person in the current interaction scenario; Indicates hotspot areas The interaction heat value is calculated by weighted fusion of the visual focus features of the area and the confidence level of the speaking state, reflecting the overall heat level of the area as the current interaction focus; This represents the baseline engagement coefficient, with a value of 0.4, which indicates the proportion of basic engagement that individuals retain within the area even when they are not actively looking at the target. This represents the gaze engagement gain coefficient, with a value of 0.6. It is used to amplify the contribution of gaze behavior to the popularity weight, reflecting the role of active gaze in improving interaction engagement. Personnel The gaze engagement quantifies the degree of visual attention a person pays to the current interaction, and the specific value rules are as follows; Personnel Watching the robot or the current speaker. Personnel Occasionally glancing at the robot or the speaker, Personnel The robot or speaker was not being observed. The popularity weights of each interactive entity are transmitted as state input to the priority learning and scheduling module, guiding the Q-learning algorithm to allocate computational and transmission resources preferentially to the real interactive entity with the highest popularity. Specifically, the priority learning and scheduling module receives a popularity weight vector from the interaction hotspot perception and allocation module every 100 milliseconds. This vector contains the popularity weight values ​​of all registered speakers in the current scene, with dimensions consistent with the speaker registration database, set to a maximum of 30 dimensions. This popularity weight vector is concatenated with the original system state vector to form an expanded state space. The expanded state vector has a total of 42 dimensions, including the original 12-dimensional system resource state and the newly added 30-dimensional popularity weight state. The Q-learning algorithm dynamically... During the selection process, the popularity weight is used as a priori guide for the reward signal. For interactive subjects with a popularity weight higher than 0.8, their corresponding media streams receive a positive bias when adjusting priorities. Specifically, a popularity bonus coefficient of 0.2 is added when calculating the Q value, making the algorithm more inclined to choose actions that increase the priority of the media streams related to that subject. Through the above mechanism, limited CPU processing resources, memory bandwidth, and network transmission resources can be prioritized for allocation to the real interactive focus in the current classroom scenario, ensuring that the audio and video data of the hot subjects receive the lowest latency processing and the highest transmission quality. Background data in non-hot areas will have their processing priority appropriately reduced when resources are scarce, thus achieving efficient and accurate resource allocation in complex interactive environments where multiple people take turns naturally. The control module performs the following steps: it collects the output audio-visual binding data, physical state data, and resource allocation data to construct a classroom interaction context containing multi-dimensional information, ensuring that the interaction decision engine obtains millisecond-level consistent comprehensive state perception, laying the foundation for accurate decision-making; it inputs the constructed interaction context into the interaction decision engine, generates interactive feedback content to be executed and corresponding robot action sequences according to the preset teaching assistant interaction rules, realizes rule-driven intelligent and scenario-based interactive response generation in the classroom scenario, requests stable resource guarantees required for executing action sequences from the priority learning and scheduling module, converts the interactive feedback content into specific voice broadcast instructions and chassis movement instructions for execution, realizes coherent and natural multi-person rotation interactive feedback, ensures that interactive instructions are executed first in resource competition, and realizes coherent and smooth physical feedback; It should be noted that the system aggregates multi-source data from various functional modules in real time to construct a multi-dimensional interactive context encompassing the complete classroom state. Specifically, the audio-visual identity synchronization locking unit outputs the bound speaker's identity identifier and its corresponding audio-visual data package at a frequency of 30 times per second, which includes a 32-dimensional voiceprint feature template and a 128-dimensional lip movement feature vector; the audio-visual joint tracking module updates the robot's current chassis pose data every 50 milliseconds, including the speaker's polar coordinate position centered on the robot and the chassis's own angular and linear velocities; and the priority learning and scheduling module outputs the current resource allocation status of each media stream every 100 milliseconds, including audio... Video stream processing latency, CPU utilization, and memory usage are aligned using a unified system timestamp, encapsulated into a structured context message, and transmitted to the interactive decision engine via a shared memory mechanism to ensure millisecond-level time consistency of the state information upon which decisions are based. After receiving the complete classroom interaction context, the interactive decision engine performs inference and decision-making based on a pre-set teaching assistant interaction rule base. This rule base uses a production rule representation and covers typical scenarios such as classroom question and answer, speaker rotation management, and knowledge point explanation triggers. For example, if the current speaker's identity remains unchanged for 3 consecutive seconds and their azimuth angle α overlaps with the visual focus area beyond a certain threshold... When the robot reaches 80% completion, it enters a focused explanation state. The interactive decision engine generates corresponding interactive feedback content, including the speech text to be synthesized (no more than 20 characters) and a sequence of fine-grained actions the robot should execute. The action sequence is planned with a control cycle of 50 milliseconds and includes chassis movement commands. The generated interactive feedback content and action sequence are packaged into a task description file and sent via inter-process communication for resource request and command execution. Before executing the interactive feedback, the control module submits a resource guarantee request to the priority learning and scheduling module, specifying the computational and transmission resource requirements for the action sequence to be executed. The priority learning and scheduling module determines the resource requirements based on the current 42-dimensional state space (including 12 dimensions). The system uses Q-learning inference (based on the status of system resources and the popularity weights of 30 trending figures) to allocate stable processing resources for the upcoming audio and video broadcast streams and control command streams. After resource assurance is confirmed, the control module converts the interactive feedback content into specific execution commands: the voice broadcast commands generate PCM audio data streams with a 16kHz sampling rate and 16-bit quantization precision, and output them to the speakers through the audio driver interface; the chassis motion commands calculate the angular velocity and linear velocity according to the PID control law, encapsulate them into CAN bus messages, and send them to the motor drive unit, realizing a smooth, coherent, and low-latency physical response of the robot in multi-person rotating interaction scenarios.

[0024] The following is a detailed explanation of the workflow of this AI teaching assistant robot that supports simultaneous and interactive multi-media streams.

[0025] First, the robot's binocular vision camera and microphone array respectively collect video frame sequences and audio streams in the classroom scene in real time. The streaming media synchronization module starts multimodal timestamp alignment at the source of acquisition. By deeply fusing lip movement features based on optical flow and audio delay estimation based on phase cross-correlation, a dynamic time offset model is constructed. The clock drift caused by hardware differences between the camera and microphone is calibrated in real time, and the timestamp disorder of multiple streams is eliminated. After completing the microsecond-level alignment of audio and video frames, the audio-visual identity synchronization locking unit extracts the synchronized voiceprint features and lip shape dynamic features, maps them to the public embedding space for similarity matching, establishes a stable audio-visual identity binding relationship, forms a temporary classroom speaker registration library, and realizes the accurate association between the speaker and the voice content. Based on the precise synchronization of identity and content, the audio-visual joint tracking module initiates a dynamic tracking process. The beamforming dynamic optimization unit treats each candidate beam direction of the microphone array as a search path, introduces an ant colony algorithm, and uses the signal-to-noise ratio of the picked-up signal as the pheromone concentration evaluation index. By simulating the foraging behavior of ants, it iterative optimization is performed to quickly converge to the optimal beam direction of the current moving sound source, dynamically adjusting the sound pickup direction to ensure continuous capture of high signal-to-noise ratio speech under the interference of multiple people walking and noise. At the same time, the vision-motion prediction drive unit receives the locked sound source area information and simultaneously acquires the target speaker image sequence collected by the vision sensor. The target position observation value is input into the visual Kalman filter. This filter adopts a constant velocity motion model to perform state estimation and smoothing filtering on the target's motion trajectory and outputs the predicted position and motion speed for the next moment. The control system calculates the motion control command of the chassis in advance based on this, driving the robot to achieve smooth joint active tracking and ensuring that the speaker is always in the main lobe of the sound pickup beam and the center of the field of vision. To address the resource contention issue during concurrent interaction of multiple media streams, the interaction hotspot perception and allocation module and the priority learning and scheduling module work together. The interaction hotspot perception and allocation module analyzes scene images through a visual attention mechanism, extracts salient regions, and integrates the current speaking state to quantify the popularity weight of each interactive subject. The priority learning and scheduling module receives this popularity weight vector, concatenates it with the system's real-time resource state to form an extended state space, and uses the Q-learning algorithm for dynamic scheduling decisions. The algorithm selects priority adjustment actions based on the ε-greedy strategy, calculates reward values ​​based on post-execution processing delays and frame drops, and iteratively updates the Q-table, ultimately converging to the optimal priority allocation scheme, tilting computing and transmission resources towards the real interactive subject with the highest popularity. Finally, the control module gathers audio-visual data, physical state, and resource allocation information from various modules to construct a complete classroom interaction context. The input interaction decision engine generates interactive feedback content and action sequences based on the rule base, and after obtaining resource guarantees, converts the instructions into specific voice broadcasts and chassis movements, and issues them for execution, achieving coherent and natural human-computer interaction.

[0026] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An AI teaching assistant robot supporting multi-channel media stream synchronization and interaction, comprising a robot body (1), characterized in that, The robot body (1) has an AI teaching assistant interaction system built in, which includes the following modules: The streaming media synchronization module is used to employ a multimodal timestamp alignment algorithm to deeply fuse video frame motion features based on optical flow with audio delay estimation based on phase cross-correlation, thereby simultaneously locking the speaker's identity and voice content at the source of acquisition. The audio-visual joint tracking module is used to introduce ant colony algorithm to dynamically optimize the beamforming direction of microphone array, and integrate visual Kalman filter to predict target motion trajectory, driving the robot's audio-visual perception function to jointly and actively track. The priority learning and scheduling module is used to dynamically schedule the processing priorities of various data streams using the Q-learning algorithm in reinforcement learning, and obtain the optimal scheduling strategy through trial and error learning. The interactive hotspot perception and allocation module is used to integrate the visual attention mechanism to analyze the visual focus in the scene, perceive the current interactive hotspot, and use it as the state input of Q-learning to guide the scheduling algorithm to allocate resources to the real interactive subject. The control module receives audio-visual data, physical status data, and resource data from each module, constructs a complete context, makes interactive decisions, and translates them into specific control commands to obtain stable resource support.

2. The AI ​​teaching assistant robot supporting multi-channel media stream synchronization and interaction according to claim 1, characterized in that: The streaming media synchronization module includes a multimodal timestamp alignment unit and an audio-visual identity synchronization locking unit; The multimodal timestamp alignment unit is used to deeply fuse video frame motion features based on optical flow and audio delay estimation algorithm based on phase cross-correlation at the source of audio and video acquisition. By analyzing the time difference between lip micro-movement and sound wave arrival, it dynamically calibrates timestamps from different hardware, eliminates timestamp disorder from multiple streams, and aligns each frame of image with the corresponding speech segment in time. The audio-visual identity synchronization locking unit binds synchronized audio features and video features based on the accurate synchronization data provided by the multimodal timestamp alignment unit. Through joint feature locking, it associates the speaker's identity with the voice content.

3. The AI ​​teaching assistant robot supporting multi-channel media stream synchronization and interaction according to claim 2, characterized in that: The multimodal timestamp alignment unit performs the following steps: The video frame sequence output by the robot's camera and the audio stream picked up by the microphone array are acquired simultaneously. The optical flow motion feature vector of the mouth region in each frame image and the phase information of each audio segment in the frequency domain are extracted respectively. The extracted optical flow motion features and the audio time delay estimate calculated based on the phase cross-correlation algorithm are input into the deep fusion network, and a dynamic time offset model is constructed by analyzing the lip movement start time and the sound wave arrival time difference. The hardware clock drift of the camera and microphone is calibrated in real time using a dynamic time offset model, and the timestamps of video frames and audio packets are dynamically compensated to eliminate the timestamp disorder of multiple streams at the acquisition source.

4. The AI ​​teaching assistant robot supporting multi-channel media stream synchronization and interaction according to claim 2, characterized in that: The audio-visual identity synchronization locking unit performs the following steps: Receive synchronized audio and video streams output by the multimodal timestamp alignment unit, and extract the voiceprint feature vectors from the synchronized audio segments and the face / mouth structure features from the corresponding video frames; A cross-modal feature joint embedding space is constructed, and similarity measurement and association matching are performed between voiceprint features and dynamic lip shape features to form an audio-visual identity binding mapping relationship; The speaker's identity is locked based on the binding mapping relationship, and the collected voice content is associated with the corresponding identity identifier in real time, so that the speaker and the voice content are accurately matched.

5. The AI ​​teaching assistant robot supporting multi-channel media stream synchronization and interaction according to claim 2, characterized in that: The audio-visual joint tracking module includes a beamforming dynamic optimization unit and a vision-motion prediction driving unit. The beamforming dynamic optimization unit is used to dynamically optimize the beamforming direction of the microphone array by introducing an ant colony algorithm in response to the speaker's random movement, so as to quickly find the optimal path of the sound source in the sound field environment and dynamically adjust the sound pickup direction. The vision-motion prediction drive unit is used to fuse a visual Kalman filter to predict the motion trajectory of the target and drive the robot's chassis to rotate smoothly in advance.

6. The AI ​​teaching assistant robot supporting multi-channel media stream synchronization and interaction according to claim 5, characterized in that: The beamforming dynamic optimization unit performs the following steps: Initialize the beamforming parameters of the microphone array, regard the beam pointing in each direction as the search path in the ant colony algorithm, and use the signal-to-noise ratio of the picked signal as the pheromone concentration evaluation index. The movement process of ant colony in sound field environment is simulated. Through pheromone update and path selection probability calculation, iterative search enables the beamforming direction to converge quickly to the optimal direction of the current sound source. The array weighting coefficients are dynamically adjusted based on the optimal beam direction output by the ant colony algorithm. The pickup beam direction is updated in real time in environments with multiple people moving around and noise interference, continuously capturing high signal-to-noise ratio speech.

7. An AI teaching assistant robot supporting multi-channel media stream synchronization and interaction according to claim 5, characterized in that: The vision-motion prediction driving unit performs the following steps: The sound source region locked by the receiving beamforming dynamic optimization unit is simultaneously acquired by the visual sensor to obtain the target speaker image sequence and extract the target's position information in the image coordinate system. The target position sequence is input into a visual Kalman filter, and the position and velocity at the next moment are estimated and predicted by combining the preset target motion model to generate smooth motion trajectory prediction data. Based on motion trajectory prediction data, the required rotation angle and angular velocity of the robot chassis are calculated in advance, and smooth following instructions are generated to drive the robot to perform joint active tracking.

8. An AI teaching assistant robot supporting multi-channel media stream synchronization and interaction according to claim 5, characterized in that: The priority learning and scheduling module performs the following steps: Define the various media stream types that coexist concurrently in the system as a set of scheduling objects, set the initial priority queues corresponding to each type of stream, and initialize the state space and action space of the Q-learning algorithm; Real-time monitoring of data volume, processing latency, and system resource utilization of each media stream; using the current resource status as input to the Q-learning algorithm; and prioritizing actions based on the ε-greedy strategy. After performing the priority adjustment action, the reward value is calculated based on the changes in processing latency and frame loss. The scheduling strategy is optimized by iteratively updating the Q table, and gradually converges to the optimal priority allocation scheme that minimizes resource contention.

9. An AI teaching assistant robot supporting multi-channel media stream synchronization and interaction according to claim 8, characterized in that: The interactive hotspot perception and allocation module performs the following steps: Classroom scene images acquired by cameras are collected, and salient regions in the scene are extracted through visual attention mechanisms. The distribution of visual focus and the direction of human eye gaze in each region are analyzed. By fusing visual focus distribution features with speaking status in each area, interactive hotspots and corresponding hot figures in the current scene are identified, and the popularity weight of each interactive subject is quantified. The popularity weights of each interactive entity are transmitted as state inputs to the priority learning and scheduling module, guiding the Q-learning algorithm to allocate computation and transmission resources preferentially to the real interactive entity with the highest popularity.

10. An AI teaching assistant robot supporting multi-channel media stream synchronization and interaction according to claim 9, characterized in that: The control module performs the following steps: By collecting the output audio-visual binding data, physical status data, and resource allocation data, a classroom interaction context containing multi-dimensional information is constructed. The constructed interaction context is input into the interaction decision engine, which generates the interaction feedback content to be executed and the corresponding robot action sequence according to the preset teaching assistant interaction rules. The system requests stable resource guarantees from the priority learning and scheduling module to execute the action sequence, and converts the interactive feedback content into specific voice broadcast commands and chassis movement commands for execution, enabling multi-person rotation interactive feedback.