Control method of teaching device, control device, teaching system, and storage medium

By collecting image and audio data in the teaching space and extracting visual and auditory information, the problem of poor convenience caused by the single control method of teaching equipment is solved, realizing multi-dimensional equipment control and improving the convenience and accuracy of operation.

CN114779922BActive Publication Date: 2026-05-19IFLYTEK LINGZHI (JIANGSU) TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK LINGZHI (JIANGSU) TECH CO LTD
Filing Date
2022-03-11
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

The existing control methods for teaching equipment mainly rely on contact operation, which restricts the user's activity space, makes it difficult to meet the needs for convenience, and easily leads to equipment unavailability.

Method used

By collecting image and audio data in the teaching space, visual and auditory information is extracted, and teaching equipment is controlled based on this information to achieve multi-dimensional intent recognition and equipment operation.

Benefits of technology

It improves the convenience and accuracy of controlling teaching equipment, makes operation simple for users, and allows for quick switching between different devices, reducing the number of times equipment is unavailable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114779922B_ABST
    Figure CN114779922B_ABST
Patent Text Reader

Abstract

The application discloses a teaching equipment control method, a control device, a teaching system and a storage medium. The teaching equipment control method comprises the following steps: image and audio of a target in a teaching space are collected to obtain image data and audio data of the target, wherein the teaching space comprises the teaching equipment; visual information of the target is extracted by using the image data of the target, and auditory information of the target is extracted by using the audio data of the target; and the teaching equipment is controlled based on the visual information and the auditory information of the target. In the foregoing manner, the convenience of the teaching equipment control can be improved, and the accuracy is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent teaching technology, and in particular to a control method, control device, teaching system, and storage medium for teaching equipment. Background Technology

[0002] With the rise of smart classrooms, education is developing towards intelligence. Leveraging artificial intelligence, smart classrooms are equipped with an increasing number of teaching devices, and the intelligence of these devices brings convenience to teaching.

[0003] However, currently, the main control method for teaching equipment is contact-based, such as operating individual teaching equipment through contact control panels, computer monitors, wireless remote controls, etc. This control method requires users to make contact with the equipment individually, which restricts the user's hand space and makes it difficult to meet the increasingly diverse needs for convenient control of teaching equipment. Summary of the Invention

[0004] The main technical problem addressed by this application is to provide a control method, control device, teaching system, and storage medium for teaching equipment, which can improve the convenience of controlling teaching equipment while maintaining high accuracy.

[0005] To address the aforementioned technical problems, the first aspect of this application provides a method for controlling teaching equipment, comprising: acquiring images and audio data of a target in a teaching space to obtain image data and audio data of the target, wherein the teaching space includes teaching equipment; extracting visual information of the target using the image data of the target, and extracting auditory information of the target using the audio data of the target; and controlling the teaching equipment based on the visual and auditory information of the target.

[0006] To address the aforementioned technical problems, a second aspect of this application provides a control device, comprising: an acquisition module for acquiring images and audio data of a target in a teaching space, wherein the teaching space includes teaching equipment; an extraction module for extracting visual information of the target using the image data and extracting auditory information of the target using the audio data; and a control module for controlling the teaching equipment based on the visual and auditory information of the target.

[0007] To address the aforementioned technical problems, a third aspect of this application provides a control device comprising a memory and a processor coupled to each other, wherein the memory stores program data and the processor executes the program data to implement the aforementioned method.

[0008] To address the aforementioned technical problems, a fourth aspect of this application provides a teaching system, which includes the aforementioned control device and teaching equipment, wherein the control device is communicatively connected to the teaching equipment and is used to control the teaching equipment.

[0009] To address the aforementioned technical problems, a fifth aspect of this application provides a computer-readable storage medium storing program data, which, when executed by a processor, is used to implement the aforementioned method.

[0010] The beneficial effects of this application are as follows: Unlike the prior art, this application acquires image and audio data of targets in a teaching space, including teaching equipment. Then, it extracts visual information of the targets using the image data and auditory information using the audio data. Finally, it controls the teaching equipment based on the visual and auditory information of the targets. By integrating visual and auditory information, it identifies the intentions of the targets from multiple dimensions, accurately identifying the target intentions and thus enabling rapid and accurate control of the teaching equipment. In addition, unlike contact control methods, it allows for timely switching of control over different teaching equipment through visual and auditory intention recognition, simplifying user operation and improving the convenience of controlling the teaching equipment. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them:

[0012] Figure 1 This is a flowchart illustrating an embodiment of the control method for the teaching equipment of this application;

[0013] Figure 2 This is a schematic diagram of the teaching space in this application;

[0014] Figure 3 This is a flowchart illustrating another embodiment of the control method for the teaching equipment of this application;

[0015] Figure 4 This is a schematic diagram of the target's 3D point cloud;

[0016] Figure 5 yes Figure 3 A flowchart illustrating the process of extracting attitude information in step S24.

[0017] Figure 6 yes Figure 3 A flowchart illustrating the process of extracting visual information in step S24.

[0018] Figure 7 This is a schematic diagram of the image after face detection.

[0019] Figure 8 This is a schematic diagram of the key facial features of the target.

[0020] Figure 9 This is a schematic diagram of the target's eye movement vector;

[0021] Figure 10 This is a schematic diagram of a head pose estimation scenario;

[0022] Figure 11 This is a schematic diagram of the coordinate system for calculating the Euler angles of the head in head pose estimation;

[0023] Figure 12 yes Figure 3 A flowchart illustrating the process of extracting gesture information in step S24.

[0024] Figure 13 yes Figure 3 A flowchart illustrating the process of extracting auditory information in step S24.

[0025] Figure 14 yes Figure 3 A flowchart illustrating another embodiment of step S25;

[0026] Figure 15 It is the angle of line of sight deflection of the target looking at the edges of the control equipment;

[0027] Figure 16 This is a floor plan of the teaching space;

[0028] Figure 17 This is a flowchart illustrating yet another embodiment of the control method for the teaching equipment of this application;

[0029] Figure 18 This is a schematic block diagram of the structure of an embodiment of the control device of this application;

[0030] Figure 19 This is a schematic block diagram of another embodiment of the control device of this application;

[0031] Figure 20 This is a schematic block diagram of the structure of an embodiment of the teaching system of this application;

[0032] Figure 21 This is a schematic block diagram of an embodiment of a computer-readable storage medium of this application. Detailed Implementation

[0033] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0034] The terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0036] Traditionally, the use of contact-based control methods in teaching spaces, operating individual control devices, leads to complex operations and poor control convenience. Furthermore, situations frequently arise where remote controls are lost or batteries run out, preventing timely control of the teaching equipment. Therefore, this application provides a method that integrates visual and auditory information to accurately understand user intent and activate and control hardware in the teaching space via control devices. This method improves the convenience of controlling teaching equipment, while also offering higher accuracy and a better user experience.

[0037] Please see Figures 1 to 2 , Figure 1 This is a flowchart illustrating an embodiment of the control method for the teaching equipment of this application. Figure 2 This is a schematic diagram of the teaching space described in this application. The implementing entity of this application is the control equipment.

[0038] The method may include the following steps:

[0039] Step S11: Collect images and audio data of the target in the teaching space to obtain image data and audio data of the target. The teaching space includes teaching equipment.

[0040] The teaching space refers to the space used for teaching, such as classrooms and laboratories. Teaching equipment may include, but is not limited to: video equipment, display equipment, audio equipment, lighting equipment, and light-blocking devices. Video equipment may include recording and broadcasting hosts, 4K cameras, PTZ cameras, and whiteboard cameras. Display equipment may include nano blackboards, smart screens, interconnected blackboards, and projectors. Audio equipment may include audio hosts, noise-canceling microphones, hanging microphones, wireless microphones, all-in-one audio systems, and wireless microphones. Lighting equipment may include smart lights and smart desk lamps. Light-blocking devices may include smart curtains. All of the above teaching equipment can communicate with and be controlled by a control device (wired or wireless).

[0041] like Figure 2 The diagram shown is a schematic of a teaching space, which includes the following teaching equipment: smart access control 10, electronic class sign 11, camera 12, electronic whiteboard 13, screen 14, air conditioner 15, all-in-one machine 16, projector 17, smart curtains 18, and smart lighting 19. In addition, it may include wireless apps (not shown), microphones, etc., which are not limited here.

[0042] Image and audio data are acquired synchronously. The control device, located in the teaching space, can acquire images of targets within the space using image sensors and simultaneously acquire audio data from targets using microphones. Image data can include multiple frames, and the target can be a person. Specifically, the control device can capture and receive video streams in real time using image sensors, then extract image frames based on the video stream, while simultaneously acquiring audio data from the target using the microphone. The teaching space can include one or more image sensors, where the type and location of each image sensor can be configured according to actual needs. In one example, the teaching space includes six image sensors positioned at different locations within the teaching space to capture images of the entire teaching space.

[0043] In some implementations, the control device can enhance the acquired image data to improve image quality and recognizability, facilitating further image analysis and processing. For example, when the image sensor's exposure rate cannot be automatically adjusted, and the lighting in the teaching space changes, the image may be underexposed, necessitating image enhancement algorithms such as histogram equalization, Laplacian operator, Log function, and gamma transform. Furthermore, when images are blurred due to interference from smoke, dust, or other factors, dehazing algorithms can be used to enhance the image data and obtain a clearer image.

[0044] Step S12: Extract visual information of the target using the target's image data, and extract auditory information of the target using the target's audio data.

[0045] Specifically, image processing is performed on the target's image data to obtain the target's visual information. This visual information may include, but is not limited to, at least one of: posture information, gaze information, gesture information, and lip information. Posture information is used to record the target's posture category, gaze information is used to record the target's gaze direction, gesture information is used to record the target's gesture category, and lip information is used to record the target's lip movement state. Different image processing algorithms or methods can be used to obtain different visual information; specific implementation methods are described in the following examples.

[0046] Specifically, extracting auditory information includes: extracting the acoustic features of the target using its audio data, and then using the extracted acoustic features to perform speech recognition, thereby obtaining the auditory information of the target. Auditory information can be speech information, such as an audio clip.

[0047] Step S13: Control the teaching equipment based on the visual and auditory information of the target.

[0048] In one embodiment, the teaching equipment can be controlled when both the visual and auditory information of the target meet preset requirements, thereby increasing the accuracy of control over the teaching equipment. Alternatively, the teaching equipment can be controlled when either the visual or auditory information of the target meets preset requirements, thereby increasing the accuracy of control over the teaching equipment.

[0049] The above solution acquires image and audio data of targets in the teaching space, including teaching equipment. It then extracts visual information from the target's image data and auditory information from its audio data. Finally, based on this visual and auditory information, the teaching equipment is controlled. By integrating visual and auditory information, the solution identifies the target's intent from multiple dimensions, enabling accurate identification and rapid, accurate control of the teaching equipment. Furthermore, unlike contact control methods, visual and auditory intent recognition allows for timely switching between different teaching devices, simplifying user operation and improving the convenience of controlling the teaching equipment.

[0050] Please see Figures 3 to 4 , Figure 3 This is a flowchart illustrating another embodiment of the control method for the teaching equipment described in this application. Figure 4 This is a schematic diagram of the target's three-dimensional point cloud.

[0051] Step S21: Collect images and audio data of the target in the teaching space to obtain image data and audio data of the target. The teaching space includes teaching equipment.

[0052] Step S22: Identify and track the target in the image data to obtain the initial trajectory of the target.

[0053] In one embodiment, target detection can be performed using the target's image data to obtain a target bounding box for at least one target. Then, 3D point cloud reconstruction can be performed using the target's image data to obtain a point cloud of at least one subject. Target recognition can then be performed on the point clouds of the at least one subject to determine the point cloud corresponding to at least one target. Finally, the target's point cloud and target bounding box can be used to track the target and obtain its initial trajectory. Figure 4 The diagram shows a 3D point cloud when the target is a human body.

[0054] In this process, object detection algorithms can be used to process image data to obtain target boxes for at least one target, wherein each target box contains only one target.

[0055] Among these methods, the 3D point cloud of the target can be constructed based on the SFM algorithm, achieving higher accuracy. Traditional SLAM algorithms assume the image sensor is located at the target's position, but in this application, the image sensor and target are located at different positions, thus failing to meet the requirements of this application. Specifically, images of the target from multiple angles within the teaching space can be captured using an image sensor, and the 3D point cloud of the target can be reconstructed through image feature point matching. When the image sensor is a depth camera, the spatial positional differences of the point cloud can be determined more accurately, resulting in more precise relative positions of the point cloud. The subject can include targets (e.g., people) and non-targets (e.g., objects, including tables, chairs, etc.).

[0056] The calculation of the point cloud world coordinate system first requires establishing a coordinate system based on the teaching space. For example, the origin can be the center point of the blackboard in the classroom, the positive X-axis can be the direction horizontally to the right of the origin, the positive Y-axis can be the direction perpendicular to the origin, and the positive Z-axis can be the direction perpendicular to the XY plane pointing towards the target. After the 3D point cloud is reconstructed, each point on the target body can be represented by coordinates. Since the multiple image sensors are located in different spatial positions in the teaching space, they need to be transformed to the world coordinate system. Therefore, it is necessary to collect relevant parameters (such as intrinsic parameters) of the image sensors, and then transform them based on the intrinsic parameters and coordinates of the image sensors to obtain the world coordinates of the point cloud.

[0057] After obtaining the point cloud of at least one subject, a trained 3D object recognition model can be used to identify the target within the point cloud of the at least one subject. The 3D object recognition model is a deep learning model; therefore, it needs to be trained before practical application. Specifically, data of the desired category (human body in this example) can be collected, preprocessed, and labeled to train a deep learning model for classification. In addition, the categories of the collected data can also include other objects, such as tables and chairs, so that the trained 3D recognition model can also recognize other objects in the scene, effectively providing scalability.

[0058] In summary, after obtaining the initial point cloud and bounding box of the target using image data, target tracking can be achieved based on the positional changes of the point cloud and the positional changes of the bounding box in each frame of the image, thus eliminating the need for separate remodeling of the teaching space. Furthermore, tracking the target simultaneously based on the point cloud and bounding box is suitable for long-distance target tracking throughout the entire teaching space, enabling rapid and accurate target localization in subsequent frame image data.

[0059] Step S23: Based on the initial trajectory of the target, associate the target in the image data.

[0060] After identifying and tracking the target, the initial trajectory of the target can be obtained. Based on the initial trajectory, the target in each frame of the image data can be associated. Furthermore, based on the initial trajectory, the position of the target in subsequent frames can be determined, and the target in subsequent frames can be associated with the initial trajectory to update the target's trajectory.

[0061] Step S24: Extract visual information of the target using the target's image data, and extract auditory information of the target using the target's audio data.

[0062] Please see Figure 5 , Figure 5 yes Figure 3 A flowchart illustrating the process of extracting attitude information in step S24.

[0063] like Figure 5 As shown, when the visual information includes pose information, extracting the visual information of the target using the target's image data includes steps S2411 to S2415:

[0064] Step S2411: Establish a corresponding tracking sequence for the target in the image data.

[0065] Step S2412: Update and record the target bounding box in each frame of the tracking sequence.

[0066] Specifically, target tracking algorithms can be used to establish corresponding tracking sequences for targets in image data.

[0067] In some implementations, when the image data contains multiple targets, a multi-target tracking algorithm can be used to establish a tracking sequence for each target in the image data, and then the target bounding box of the target can be updated and recorded in each frame of the image corresponding to the tracking sequence.

[0068] Step S2413: Crop the region corresponding to the target box in each frame image to obtain at least one target box region image of the target.

[0069] Step S2414: Perform pose estimation using each target bounding box region image to obtain at least one key point image of the target.

[0070] Specifically, pose estimation algorithms can be used to estimate the pose of the target bounding box image, output the keypoint image of the target, and save it. Commonly used pose estimation algorithms generally adopt two approaches: top-down and bottom-up. Top-down algorithms include CPM and Hourglass, while bottom-up algorithms include OpenPose and HigherHRNet.

[0071] Step S2415: Use a preset number of key point images to perform behavior recognition and obtain the target's pose information.

[0072] Specifically, the target's pose information can be obtained by using a preset number of keypoint images from the previous frame at the current time point saved in each tracking sequence. The preset number can be set according to the actual situation, such as 10 frames or 20 frames.

[0073] In some implementations, a preset number of keypoint images can be input into a skeleton-based behavior recognition algorithm (such as GCN, Potion, etc.) for behavior recognition, so as to output the category of the target's behavior as posture information. Posture information may include, but is not limited to: sitting, standing, walking, squatting, etc.

[0074] Please see Figures 6 to 9 , Figure 6 yes Figure 3 A flowchart illustrating the process of extracting line-of-sight information in step S24. Figure 7 This is a schematic diagram after face detection is performed on the image. Figure 8 This is a schematic diagram of the key facial features of the target. Figure 9 This is a schematic diagram of the target's eye movement vector.

[0075] like Figure 6As shown, when visual information includes gaze information, extracting the visual information of the target using the target's image data includes sub-steps S2421 to S2424:

[0076] Step S2421: Perform face detection on each frame of the target's image data to obtain the target's face image and facial key points.

[0077] like Figure 7 The image shown is a schematic diagram after face detection is performed on a frame of image. Face detection can detect valid faces in the image, while invalid faces that are occluded will not be detected.

[0078] Specifically, face detection algorithms (such as RetinaFace_R50.whole) can be used to detect faces in each frame of the image, and then the target's face bounding box and facial landmarks (such as...) can be obtained. Figure 8 As shown in the diagram, the face bounding box includes four feature points (the four vertices of the face bounding box) and a face confidence score (i.e., 4 positions + 1 score). Facial keypoints can include five localizations (such as face shape, facial features, inner corner of the eye, outer corner of the eye, pupil, etc.), including the iris center and inner corner of the eye localization. After recognizing the face bounding box, the face image can be obtained by cropping the image within the face bounding box region.

[0079] Step S2422: Perform face alignment using the target's face image and facial landmarks to obtain an aligned face image.

[0080] In some implementations, the target face image and facial landmarks can be used as input to a face alignment algorithm (e.g., ArcFace_R50.backbone) to output a face image aligned to standard facial landmarks. Similarity transformations can be employed in the face alignment algorithm. Optionally, the face alignment algorithm may include, but is not limited to: ASM (Active Shape Model), AAM (Active Appearance Model), CLM (Constrained Local Model), and SDM (Supervised Descent Method).

[0081] Step S2423: Embed features into the aligned face images to obtain the target face feature vector.

[0082] The purpose of feature embedding is to transform (reduce dimensionality) data into a fixed-size feature representation (vector) to facilitate processing and computation (such as distance calculation). Aligned face images, after feature embedding, can be transformed into fixed-dimensional feature vectors.

[0083] Step S2424: Perform feature matching using the target's facial feature vector and use the obtained target's eye movement vector as gaze information.

[0084] In this algorithm, the target's facial feature vector can be used as input to the feature matching algorithm to output the target's eye movement vector. For example... Figure 9 As shown, the direction of the eye movement vector is from the inner corner of the target eye (x0, y0) to the center of the iris (x1, y1).

[0085] Furthermore, after obtaining the key points of facial features (including lips), a relevant lip recognition model can be used for identification to obtain the lip information of the target. This lip recognition model employs neural networks (e.g., 3D convolutional and residual network structures). During the training phase, an RGB image region centered on the lip shape is cropped, and a keypoint mask is generated based on the lip key points to extract lip movement features. Various types of noise data are also collected to add noise to the audio data during model training.

[0086] Please see Figures 10 to 11 , Figure 10 This is a schematic diagram of a head pose estimation scenario. Figure 11 This is a schematic diagram of the coordinate system used for calculating the Euler angles of the head in head pose estimation.

[0087] In some implementations, the intrinsic relationship between head pose estimation and visual gaze estimation can be leveraged. Head pose can roughly provide the direction of the target's gaze, and can be used as the gaze direction when the eyes are not visible (e.g., in low-resolution images, when the eyes are obstructed by objects like sunglasses, or when a face is not detected (e.g., when wearing a mask)). Implementation has shown that the head direction contributes an average of 68.9% to the overall gaze direction; therefore, head pose estimation can be combined with this method to determine the target's gaze information. Specific steps may include: estimating the head pose of the target's image data to obtain the target's head deflection angle; then combining the target's eye movement vector and the head deflection angle to obtain the target's gaze information. The target's image data can be used as input to a head pose Euler angle prediction model to output the head deflection angle, i.e., the Euler angle. The head pose Euler angle prediction model is structured as a multi-loss convolutional neural network.

[0088] In some implementations, when human eyes cannot be detected, it is difficult to accurately obtain the target's eye movement vector. In this case, the target's head deflection angle can be directly used as the target's gaze information to compensate for the inability to detect facial features.

[0089] Head pose estimation: Obtaining the head's tilt angle from an image containing the face. For example... Figure 10 The image shown is a real-world scene diagram after head pose estimation from a single frame of image data. It can be seen that head pose estimation can essentially identify the head tilt angle of each person, thus compensating for situations where facial features are not detected and using this as the target's gaze direction. Figure 11 As shown, in 3D space, the rotation of an object can be represented by three Euler angles: Pitch (rotation around the y-axis), Yaw (rotation around the z-axis), and Roll (rotation around the x-axis), formally known as pitch, yaw, and roll angles, respectively. In simpler terms, these correspond to head tilting (bowing up and down), head shaking (or tilting left and right), and head turning. Below is a method for training a head pose Euler angle prediction model, which combines classification and regression losses to predict Euler angles:

[0090] First, Euler angles are classified according to their angle ranges. For example, if the angle range is 3 degrees, then the Yaw range is -180 to +180 degrees, which can be divided into 360 / 3 = 120 categories. The Pitch and Roll ranges are both -99 to +99 degrees, which can be divided into 66 categories. Thus, a classification task can be performed. Specifically, for each category of Euler angles, the classification loss can be calculated by comparing the predicted classification result with the actual classification result, thereby obtaining the classification loss for Pitch, Yaw, and Roll. The cross-entropy loss function can be used to calculate the classification loss for each Euler angle.

[0091] Then, the classification results can be restored to the actual angles (e.g., category * 3 - 90), and the regression loss can be calculated by comparing them with the actual angles. This yields the regression losses for Pitch, Yaw, and Roll. The MSE (mean squared error) function can be used to calculate the regression loss for each Euler angle.

[0092] Finally, the regression loss and classification loss are combined to obtain the total loss. Training of the head pose Euler angle prediction model is stopped when the total loss for each Euler angle is less than a preset loss threshold or the number of training iterations exceeds a preset number of training iterations. The preset loss threshold and preset number of training iterations can be set according to actual conditions, for example, 0.1 and 1000 respectively. In summary, using classification and regression paradigms for constraints can improve the accuracy of head pose estimation.

[0093] Please see Figure 12 , Figure 12 yes Figure 3 A flowchart illustrating the process of extracting gesture information in step S24.

[0094] like Figure 12As shown, when the visual information includes gesture information, extracting the visual information of the target using the target's image data includes sub-steps S2431 to S2433:

[0095] Step S2431: Identify each frame of the image data to obtain the hand region of the target.

[0096] Specifically, hand and human detection algorithms can be pre-built and then trained to obtain hand and human detection models. Each frame of the image data is used as input to the hand and human detection model, and the output is the hand region of the target in each frame. The hand and human detection model can be a deep learning model.

[0097] Step S2432: Process the image of the target's hand region to obtain the target's gesture action feature vector.

[0098] Step S2433: Recognize the target's gesture action feature vector to obtain the target's gesture information.

[0099] Specifically, a gesture recognition algorithm can be pre-built and trained to obtain a gesture recognition model. The target's hand region image is used as input to the gesture recognition model to output the target's gesture feature vector. Finally, a classifier is used to recognize the gesture feature vector to obtain the target's gesture category, i.e., gesture information. Thus, based on the trained gesture recognition model, both the target's gesture and action can be recognized simultaneously.

[0100] Before training the gesture recognition model, the gesture and action categories can be determined first. Then, data for the required categories can be collected, preprocessed, and labeled to train a deep learning model for classification.

[0101] After extracting all visual information (pose, gaze, gesture, etc.) of the target over a period of time, it is necessary to integrate them. This requires verifying the target corresponding to each tracking sequence in the behavior recognition statistics. First, the target bounding boxes recorded in the tracking sequences on the same frame are matched with the face bounding boxes recorded in face recognition using the Hungarian algorithm. This merges the frontal faces recorded in the face sequences into each tracking sequence. Verification using these frontal faces yields the identity of each tracking sequence, and tracking sequences belonging to the same identity are merged according to their chronological order. This allows the visual information, such as pose and gaze, of each target to be combined.

[0102] Please see Figure 13 , Figure 13 yes Figure 3 A flowchart illustrating the process of extracting auditory information in step S24.

[0103] In this embodiment, extracting auditory information of the target using the target's audio data may include steps S2441 to S2443:

[0104] Step S2441: Extract the acoustic features of the target using the target's audio data.

[0105] Acoustic features may include, but are not limited to, features related to energy, fundamental frequency, sound quality, and spectrum. In this embodiment, the target's speech spectrum features, such as Mel-frequency cepstral coefficients (MFCCs), can be extracted using the target's audio data.

[0106] Step S2442: Use a language model to process the audio data to obtain the probability of sentences in the audio data.

[0107] Specifically, language models (LMs) play a crucial role in natural language processing. Their task is to predict the probability of a sentence appearing in a language. Therefore, by processing audio data using language models, the probabilities of sentences within the audio data can be obtained. Language models employ recurrent neural networks and attention mechanisms.

[0108] Step S2443: Perform speech recognition based on the acoustic features of the target and the probability of sentences in the audio data to obtain the auditory information of the target.

[0109] For speech recognition, the acoustic features must first be analyzed to obtain ordered feature vectors. These feature vectors can then be used as input to the speech recognition model, which reads the features sequentially and outputs the corresponding text. The speech recognition model is a deep learning model.

[0110] Step S25: Control the teaching equipment based on the visual and auditory information of the target.

[0111] Please see Figures 14 to 16 , Figure 14 yes Figure 3 A flowchart illustrating another embodiment of step S25 is shown. Figure 15 It refers to the angle of line of sight of the target looking towards the edges of the control equipment. Figure 16 This is a floor plan of the teaching space.

[0112] When the visual information includes gesture information and gaze information, step S25 may include sub-steps S251 to S252:

[0113] Step S251: Determine whether the target's gestures meet the first requirement based on gesture information, determine whether the target's gaze meets the second requirement based on gaze information, and determine whether the target's speech meets the third requirement based on auditory information.

[0114] Specifically, determining whether a target's gesture meets the first requirement based on gesture information can be done by judging whether the target's gesture category is the same as a preset gesture. If they are the same, the target's gesture is determined to meet the first requirement; otherwise, the target's gesture is determined not to meet the first requirement. A preset gesture, for example, is an extended index finger with the other four fingers bent. If the target's gesture is this preset gesture, it indicates that the target's fingers are intentionally pointing at an object, meaning they intend to control that object.

[0115] The process of determining whether the target's line of sight meets the second requirement based on line-of-sight information includes: determining the target's line-of-sight threshold range, and then, based on the line-of-sight information, judging whether the target's line of sight is within the line-of-sight threshold range. If it is, the target's line of sight is determined to meet the second requirement; otherwise, the target's line of sight is determined not to meet the second requirement. To prevent the target's finger from accidentally pointing at the control device and performing operation, it is necessary to confirm the line-of-sight threshold range at this time.

[0116] One method for determining whether the target's speech meets the third requirement based on auditory information is to judge whether the target's speech contains keywords or key phrases. If it does, the target's speech meets the third requirement; otherwise, it does not. Alternatively, semantic recognition can be performed on the auditory information, and semantic analysis can be used to determine whether the target intends to control the teaching equipment. If so, the target's speech meets the third requirement; otherwise, it does not.

[0117] In some implementations, determining the target's line-of-sight threshold range includes: determining a first line-of-sight threshold range in the horizontal direction based on the horizontal distance between the target and the teaching equipment and the length of the teaching equipment; and determining a second line-of-sight threshold range in the vertical direction based on the horizontal distance between the target and the teaching equipment, the target's line-of-sight height, and the width of the teaching equipment.

[0118] In a specific example, such as Figure 15 As shown, firstly, a coordinate system is established based on the teaching space. For example, the center point of the classroom blackboard (or electronic whiteboard) is taken as the origin, the direction horizontally to the right of the origin is the positive X-axis, the direction perpendicular to the origin is the positive Y-axis, and the direction perpendicular to the XY plane pointing towards the target is the positive Z-axis. Let the length of the control device be h, the width be d, and the coordinates of the target's head center point be F(x, y, z). Then, the formula for calculating the line-of-sight threshold range is as follows:

[0119]

[0120]

[0121] Here, α1, α2, β1, and β2 are the angular thresholds for abnormal gaze deflection, (α1, α2) is the gaze threshold range of the target in the θYaw direction, and (β1, β2) is the gaze threshold range of the target in the θPitch direction. When the head rotation range exceeds the threshold, the target's gaze is considered to be outside the control device, and eye movement is deemed invalid.

[0122] like Figure 16 As shown, point A is on the central axis of the control device, and points B, C, and D are located at the leftmost, middle, and rightmost edges of the first row of the teaching space, respectively. When the target looks towards the left and right edges of the control device from points B and D in the first row of the teaching space, this is the maximum head rotation range in the θYaw direction (i.e., the line-of-sight threshold range in the θYaw direction), denoted by formula (1). When the target looks towards the top and bottom edges of the control device from point C, this is the maximum head rotation range in the θPitch direction (i.e., the line-of-sight threshold range in the θPitch direction), denoted by formula (2). The line-of-sight threshold range can be obtained by calculating using inverse trigonometric functions.

[0123] Step S252: When the third requirement is met, and the first and / or second requirements are met, control the teaching equipment.

[0124] Specifically, the teaching equipment is controlled when the third and first requirements are met, or when the third and second requirements are met, or only when all three requirements are met simultaneously.

[0125] In one embodiment, the teaching equipment can be controlled based on control commands determined by auditory information, such as the target saying "turn off the projector", or based on control commands determined by visual information, such as "raise your hand - close / open the curtains", etc. There are no limitations here.

[0126] Please see Figure 17 , Figure 17 This is a flowchart illustrating another embodiment of the control method for the teaching equipment of this application.

[0127] In some embodiments, after obtaining the target audio data, the following steps may also be included:

[0128] Step S26: Convert the target's audio data into text information.

[0129] Specifically, the input audio data is used to extract acoustic features on the one hand, and can be converted into text for semantic prosodic analysis on the other.

[0130] Before step S27, the obtained text information can be preprocessed, such as semantic understanding, to eliminate ambiguity to the greatest extent. Then, the text information is segmented and prosodic generated according to semantic prosodic analysis.

[0131] Step S27: Perform semantic prosodic analysis on the text information to obtain the prosodic information of the audio data.

[0132] Among them, prosodic information is used to record the prosody in the target's audio data.

[0133] Specifically, prosodic processing models can be used to process textual information to obtain prosodic information from audio data. For semantic and prosodic analysis, it is necessary to fully identify and locate the content and prosody of polyphonic characters, pauses, intonation, stress, tone, and punctuation in sentences, in conjunction with the context or even long texts. At the same time, it is necessary to exaggerate the prosodic aspects such as tone and stress. Therefore, the prosodic processing model used in practice also needs to be adjusted accordingly to achieve accurate positioning of semantics and prosody.

[0134] Step S28: Based on text information and prosodic information, synthesized speech is obtained.

[0135] Specifically, deep learning models can be used to generate waveforms based on textual and prosodic information to obtain synthesized speech. For details on the specific steps involved in speech synthesis, please refer to other related technologies.

[0136] Step S29: Play the synthesized speech.

[0137] Specifically, synthesized speech can be played through the speaker or Bluetooth device in the teaching equipment. By playing the synthesized speech, the target can confirm whether the speech recognition is accurate.

[0138] In one application scenario, control devices are used to wake up and operate the hardware in the teaching space. These control devices refer to the equipment that centrally manages and controls various devices in the smart classroom, such as sound, light, and electricity. The teaching system includes intelligent central control equipment, audio / video matrix, switches, wireless projection, and streaming media processing units. Several methods for controlling teaching equipment are illustrated below:

[0139] (1) Using a smart microphone and a teacher's personal computer (PC), the smart classroom teaching platform management system can be opened by voice control to record and broadcast the teaching process;

[0140] (2) Use voice combined with eye contact, lip movements, and gestures to further clarify the operation object, such as controlling the expansion and contraction of the blackboard, opening or closing the electronic whiteboard, opening or closing the projector or the small screen for extracting blackboard writing, adjusting the focus of the camera, etc.

[0141] (3) Use voice or touch operation interface to switch between smart classroom or exam inspection mode. The smart classroom analyzes students’ classroom behavior in real time [sleeping, playing on mobile phones, etc.] and reminds and records abnormal behavior. Among them, the exam inspection mode can perform real-time automatic analysis of suspected cheating without any blind spots.

[0142] Furthermore, this application utilizes multimodal interaction in teaching spaces to effectively eliminate ambiguity. For human-computer interaction scenarios requiring a precise, complete statement such as "I want to open the recording software to record" or "I want to conduct contactless attendance," multimodal interaction simply requires pointing a finger at the interactive object's cloud desktop and then overlaying speech with eye tracking to accurately achieve the desired interaction. For pronouns like "that" and "this," which are often used verbally and can easily lead to semantic ambiguity, gestures avoid this problem. For example, suppose a teacher standing in the middle of the classroom wants to switch to AI exam mode using a camera at the front of the classroom. They need to point to the control device, turn their eyes to the device, and then output "Start exam mode" to accurately and quickly switch to exam mode.

[0143] In summary, this application captures and receives video streams in real time through a camera, performs target detection and tracking, confirms targets through facial recognition and models them, and integrates multimodal information such as posture, gestures, eye contact, lip movements and voice. Through the interrelation of multi-source information, it can further understand the user's true intentions and realize the wake-up and precise and rapid operation control of intelligent teaching equipment.

[0144] Please see Figure 18 , Figure 18 This is a schematic block diagram of the structure of an embodiment of the control device of this application.

[0145] The control device 100 includes an acquisition module 110, an extraction module 120, and a control module 130. The acquisition module 110 acquires images and audio data of targets in the teaching space, which includes teaching equipment. The extraction module 120 extracts visual information from the target's image data and auditory information from its audio data. The control module 130 controls the teaching equipment based on the target's visual and auditory information.

[0146] In some implementations, the visual information includes at least one of posture information, gaze information, gesture information, and lip information.

[0147] In some implementations, the image data includes multiple frames of images. Before extracting the visual information of the target using the image data, the extraction module 120 is also used to identify and track the target in the image data to obtain the initial trajectory of the target; based on the initial trajectory of the target, the target in the image data is associated.

[0148] In some implementations, identifying and tracking targets in image data to obtain the initial trajectory of the targets includes: performing target detection using the target image data to obtain the target bounding box of at least one target; performing three-dimensional point cloud reconstruction using the target image data to obtain the point cloud of at least one subject; performing target identification on the point cloud of at least one subject to determine the point cloud corresponding to at least one target; and tracking the target using the target point cloud and the target bounding box to obtain the initial trajectory of the target.

[0149] In some implementations, when the visual information includes pose information, the extraction module 120 is further configured to establish a corresponding tracking sequence for the target in the image data; update and record the target bounding box of the target in each frame of the tracking sequence; crop the region corresponding to the target bounding box of each frame of the image to obtain at least one target bounding box region image of the target; perform pose estimation using each target bounding box region image to obtain at least one key point image of the target; and perform behavior recognition using a preset number of frames of key point images to obtain the pose information of the target.

[0150] In some implementations, when the visual information includes gaze information, the extraction module 120 is further configured to perform face detection on each frame of the target's image data to obtain the target's face image and facial key points; perform face alignment using the target's face image and facial key points to obtain an aligned face image; embed features into the aligned face image to obtain the target's face feature vector; perform feature matching using the target's face feature vector, and use the obtained target's eye movement vector as gaze information.

[0151] In some implementations, feature matching is performed using the target's facial feature vector to obtain the target's eye movement vector as gaze information, including: estimating the head pose of the image in the target's image data to obtain the target's head deflection angle; and combining the target's eye movement vector and head deflection angle to obtain the target's gaze information.

[0152] In some implementations, when the visual information includes gesture information, the extraction module 120 is further configured to identify each frame of the image data to obtain the target's hand region; process the target's hand region image to obtain the target's gesture action feature vector; and perform gesture recognition based on the target's gesture action feature vector to obtain the target's gesture information.

[0153] In some implementations, the auditory information of the target is extracted using the target's audio data, including: extracting the target's acoustic features using the target's audio data; processing the audio data using a language model to obtain the probability of sentences in the audio data; and performing speech recognition based on the target's acoustic features and the probability of sentences in the audio data to obtain the target's auditory information.

[0154] In some implementations, when the visual information includes gesture information and gaze information, the control module 130 is further configured to determine whether the target's gesture meets the first requirement based on the gesture information, whether the target's gaze meets the second requirement based on the gaze information, and whether the target's voice meets the third requirement based on the auditory information; when the third requirement is met, and the first requirement and / or the second requirement are met, the teaching equipment is controlled.

[0155] In some implementations, determining whether the target's line of sight meets the second requirement based on line of sight information includes: determining the target's line of sight threshold range; determining whether the target's line of sight is within the line of sight threshold range based on the line of sight information; if so, determining that the target's line of sight meets the second requirement.

[0156] In some implementations, determining the target's line-of-sight threshold range includes: determining a first line-of-sight threshold range in the horizontal direction based on the horizontal distance between the target and the teaching equipment and the length of the teaching equipment; and determining a second line-of-sight threshold range in the vertical direction based on the horizontal distance between the target and the teaching equipment, the target's line-of-sight height, and the width of the teaching equipment.

[0157] For a detailed explanation of the above steps, please refer to the corresponding section in the previous method embodiments; it will not be repeated here.

[0158] Please see Figure 19 , Figure 19 This is a schematic block diagram of another embodiment of the control device of this application.

[0159] The control device 200 may include a memory 210 and a processor 220 coupled to each other. The memory 210 is used to store program data, and the processor 220 is used to execute the program data to implement the steps in any of the above method embodiments. The control device 200 may include, but is not limited to, personal computers (e.g., desktop computers, laptop computers, tablet computers, PDAs, etc.), mobile phones, servers, wearable devices, and augmented reality (AR) and virtual reality (VR) devices, televisions, etc., without limitation.

[0160] Specifically, processor 220 controls itself and memory 210 to implement the steps in any of the above method embodiments. Processor 220 may also be referred to as a CPU (Central Processing Unit). Processor 220 may be an integrated circuit chip with signal processing capabilities. Processor 220 may also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 220 may be implemented by multiple integrated circuit chips.

[0161] Please see Figure 20 , Figure 20 This is a schematic block diagram of the structure of an embodiment of the teaching system of this application.

[0162] The teaching system 300 may include a control device 310 and a teaching device 320 in any of the above embodiments. The control device 310 is communicatively connected to the teaching device 320 and is used to control the teaching device 320.

[0163] The teaching equipment 320 includes at least one of the following: camera equipment, display equipment, audio equipment, lighting equipment, and light-blocking equipment. Camera equipment may include a recording host, a 4K camera, a PTZ camera, and a whiteboard camera. Display equipment may include a nano blackboard, a smart screen, an interconnected blackboard, and a projector. Audio equipment may include an audio host, a noise-canceling microphone, a hanging microphone, a wireless microphone, an all-in-one audio system, and a wireless microphone. Lighting equipment may include smart lights and smart desk lamps. Light-blocking equipment may include smart curtains. The control device 310 can connect to the teaching equipment 320 via a wireless app.

[0164] Please see Figure 21 , Figure 21 This is a schematic block diagram of an embodiment of a computer-readable storage medium of this application.

[0165] The computer-readable storage medium 400 stores program data 410, which, when executed by a processor, is used to implement the steps in any of the above method embodiments.

[0166] The computer-readable storage medium 400 can be a medium that can store computer programs, such as a USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk. It can also be a server that stores the computer program, which can send the stored computer program to other devices for execution or run the stored computer program itself.

[0167] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0168] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0169] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0170] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0171] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for controlling teaching equipment, characterized in that, include: Images and audio data of targets in a teaching space are acquired to obtain image data and audio data of the targets. The teaching space includes teaching equipment, and the image data includes multiple frames of images. The system performs target detection using the image data of the target to obtain a target bounding box for at least one target; it then performs 3D point cloud reconstruction using the image data of the target to obtain a point cloud of at least one subject; it performs target recognition on the point clouds of at least one subject to determine a point cloud corresponding to at least one target; and it tracks the target using the point cloud and the target bounding box to obtain an initial trajectory of the target; based on the initial trajectory of the target, it associates the target in the image data. Visual information of the target is extracted using image data of the target, and auditory information of the target is extracted using audio data of the target; The teaching equipment is controlled based on the visual and auditory information of the target. Wherein, when the visual information includes pose information, the step of extracting the visual information of the target using the image data of the target includes: establishing a corresponding tracking sequence for the target in the image data; updating and recording the target bounding box of the target in each frame of the tracking sequence; cropping the region corresponding to the target bounding box in each frame of the image to obtain at least one target bounding box region image of the target; performing pose estimation using each target bounding box region image to obtain at least one key point image of the target; and performing behavior recognition using a preset number of frames of key point images to obtain the pose information of the target.

2. The method according to claim 1, characterized in that, The visual information includes at least one of gaze information, gesture information, and lip information.

3. The method according to claim 2, characterized in that, When the visual information includes gaze information The step of extracting visual information of the target using image data of the target includes: Face detection is performed on each frame of the image data of the target to obtain the face image and facial key points of the target; Face alignment is performed using the target's face image and facial key points to obtain an aligned face image; The aligned face image is then used for feature embedding to obtain the target's face feature vector; Feature matching is performed using the facial feature vector of the target, and the resulting eye movement vector of the target is used as gaze information.

4. The method according to claim 3, characterized in that, The step of performing feature matching using the facial feature vector of the target and using the resulting eye movement vector of the target as gaze information includes: The head pose of the target is estimated by performing head pose estimation on the image data of the target to obtain the head deflection angle of the target; By combining the target's eye movement vector and head tilt angle, the target's gaze information is obtained.

5. The method according to claim 2, characterized in that, When the visual information includes gesture information The step of extracting visual information of the target using image data of the target includes: Each frame of the image data is identified to obtain the hand region of the target; The hand region image of the target is processed to obtain the hand gesture feature vector of the target; Gesture recognition is performed based on the gesture action feature vector of the target to obtain the gesture information of the target.

6. The method according to claim 1, characterized in that, The step of extracting auditory information of the target using the target's audio data includes: The acoustic features of the target are extracted using the target's audio data; The probability of sentences in the audio data is obtained by processing the audio data using a language model; Speech recognition is performed based on the acoustic features of the target and the probability of sentences in the audio data to obtain the auditory information of the target.

7. The method according to claim 2, characterized in that, When the visual information includes gesture information and gaze information... The control of the teaching equipment based on the visual and auditory information of the target includes: Based on the gesture information, determine whether the target's gesture meets the first requirement; based on the gaze information, determine whether the target's gaze meets the second requirement; and based on the auditory information, determine whether the target's voice meets the third requirement. When the third requirement is met, and the first requirement and / or the second requirement are met, the teaching equipment is controlled.

8. The method according to claim 7, characterized in that, The step of determining whether the target's line of sight meets the second requirement based on the line of sight information includes: Determine the line-of-sight threshold range of the target; Based on the line-of-sight information, determine whether the target's line of sight is within the line-of-sight threshold range; If so, then the line of sight of the target is determined to meet the second requirement.

9. The method according to claim 8, characterized in that, Determining the line-of-sight threshold range for the target includes: Based on the horizontal distance between the target and the teaching equipment and the length of the teaching equipment, a first line-of-sight threshold range for the target in the horizontal direction is determined; Based on the horizontal distance between the target and the teaching equipment, the line-of-sight height of the target, and the width of the teaching equipment, a second line-of-sight threshold range for the target in the vertical direction is determined.

10. A control device, characterized in that, include: The acquisition module is used to acquire images and audio data of targets in the teaching space, and to obtain image data and audio data of the targets. The teaching space includes teaching equipment, and the image data includes multiple frames of images. An extraction module is configured to: perform target detection using the image data of the target to obtain a target bounding box for at least one target; perform 3D point cloud reconstruction using the image data of the target to obtain a point cloud of at least one subject; perform target recognition on the point clouds of at least one subject to determine a point cloud corresponding to at least one target; track the target using the point cloud and the target bounding box to obtain an initial trajectory of the target; associate the target in the image data based on the initial trajectory of the target; extract visual information of the target using the image data of the target; and extract auditory information of the target using the audio data of the target. A control module is used to control the teaching equipment based on the visual and auditory information of the target; wherein, when the visual information includes posture information, the step of extracting the visual information of the target using the image data of the target includes: establishing a corresponding tracking sequence for the target in the image data; updating and recording the target bounding box of the target in each frame image corresponding to the tracking sequence; cropping the region corresponding to the target bounding box in each frame image to obtain at least one target bounding box region image of the target; performing posture estimation using each target bounding box region image to obtain at least one key point image of the target; and performing behavior recognition using a preset number of frames of key point images to obtain the posture information of the target.

11. A control device, characterized in that, The control device includes a memory and a processor coupled to each other, the memory being used to store program data, and the processor being used to execute the program data to implement the method as described in any one of claims 1-9.

12. A teaching system, characterized in that, It includes the control device as described in claim 11, and the teaching device, wherein the control device is communicatively connected to the teaching device and is used to control the teaching device.

13. The teaching system according to claim 12, characterized in that, The teaching equipment includes at least one of the following: camera equipment, display equipment, audio equipment, lighting equipment, and light-shielding equipment.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program data, which, when executed by a processor, is used to implement the method as described in any one of claims 1-9.