Classroom automatic director method and electronic equipment
By using microphone arrays and video analysis technology, combined with sound source localization, lip movement and motion recognition, high-precision recognition of speakers and interactive objects in the classroom and natural and accurate screen switching are achieved. This solves the problems of intelligence and accuracy in existing classroom broadcasting systems and improves the intelligence level and viewing experience of teaching recording and live broadcasting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU KINDLINK INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2025-11-28
- Publication Date
- 2026-04-14
Smart Images

Figure CN121865087A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of educational informatization and artificial intelligence technology, and in particular to an automatic classroom broadcasting method and electronic equipment. Background Technology
[0002] In smart classrooms or recording studios, there are usually multiple cameras. The cameras can automatically or manually switch to the appropriate camera feed based on the classroom scenario (teacher lecturing, student questions, presentation slides, etc.) to output the footage to the recording system or live streaming platform. Typically, a dedicated person operates the switcher or directing software to manually select the camera based on the classroom situation; alternatively, simple voice detection may trigger the camera switching.
[0003] However, existing classroom broadcasting systems lack the ability to comprehensively understand audio and video information, making it difficult to accurately perceive and intelligently respond to real classroom interaction scenarios. Summary of the Invention
[0004] One objective of this application is to provide an automated classroom broadcasting method, device, and electronic device to address how to improve the accuracy and intelligence of classroom broadcasting, thereby enabling precise identification of speakers and interactive behaviors, and achieving natural and accurate screen switching.
[0005] To address the aforementioned technical problems, one technical solution adopted in this application is: providing an automatic classroom broadcasting method, comprising: acquiring audio and video data of the current classroom scene; determining the current sound source direction by locating the sound source direction using a microphone array based on the audio data; extracting facial key points from the video data and performing lip movement recognition based on the facial key points to obtain lip movement recognition results; extracting human body key points from the video data and performing action recognition based on the human body key points to obtain target action recognition results; controlling the current shooting screen of the camera device based on the sound source direction, lip movement recognition results, and target action recognition results, wherein the current shooting screen of the camera device is the broadcasting screen; and outputting and displaying the broadcasting screen.
[0006] This automated classroom broadcasting method uses a microphone array to locate the sound source direction, accurately determining the direction of the sound. Combined with lip movement recognition of facial key points and motion recognition of human body key points, it effectively distinguishes the speaker from the interacting audience. This method integrates multimodal audio and video information, achieving high-precision recognition of speakers and interacting audiences in classroom scenarios and natural, accurate automatic screen switching. In complex classroom environments, this method can stably track and intelligently focus on target objects, ensuring the continuity and naturalness of screen transitions, thereby significantly improving the intelligence and viewing experience of teaching recording, remote teaching, and other related applications.
[0007] Optionally, based on audio data, sound source direction localization is performed using a microphone array to determine the current sound source direction. This includes: acquiring multi-channel audio signals from the microphone array based on the audio data; performing windowing processing on each channel of the multi-channel audio signal to obtain a windowed multi-channel audio signal; applying a preset window function to each channel audio signal and performing weighted processing; performing Fourier transform on each windowed multi-channel audio signal to obtain the time-frequency representation of each channel audio signal; obtaining the first propagation delay between any two microphones in the microphone array based on the time-frequency representation of each channel audio signal; determining candidate sound source directions based on the first propagation delay and the spatial geometric position of the microphone array, and obtaining the spatial probability distribution corresponding to the candidate sound source directions; selecting the candidate sound source direction with the largest spatial probability distribution as the current sound source direction. This method can accurately identify the direction of a sound source in space, providing a reliable basis for real-time sound source localization.
[0008] Optionally, based on the time-frequency representation of each audio signal channel, the first propagation delay between any two microphones in the microphone array is obtained, including: selecting audio signals belonging to the same time window based on the time-frequency representation of each audio signal channel; for audio signals belonging to the same time window, calculating the correlation between the first audio signal and the second audio signal at different time offsets to obtain a similarity metric corresponding to the time offset; based on the similarity metrics obtained for all time offsets, selecting the time offset with the largest similarity metric, and using the largest time offset as the first propagation delay between the two microphones; the two microphones refer to the microphones corresponding to the first audio signal and the second audio signal. This method can accurately measure the sound wave propagation time difference between any two microphones in the microphone array, providing a reliable time basis for sound source localization.
[0009] Optionally, based on the first propagation delay and the spatial geometric position of the microphone array, candidate directions of the sound source are determined, and the spatial probability distribution corresponding to the candidate directions is obtained. This includes: determining the spatial range of the sound source based on the spatial geometric position of the microphone array, and generating multiple candidate directions within the spatial range; obtaining the propagation time of the sound wave reaching each microphone under each candidate direction based on the candidate directions and the spatial geometric position of the microphone array, and obtaining the second propagation delay for each pair of microphones based on the propagation time; comparing the first propagation delay with the second propagation delay to calculate the matching degree between each candidate direction and the actual observation; and normalizing all candidate directions based on the matching degree to obtain the spatial probability distribution corresponding to the candidate directions. This method can match the actual observed microphone delay with the theoretical delay of the candidate directions, accurately assess the possible location of the sound source in space, and thus provide reliable probabilistic information for sound source localization.
[0010] Optionally, facial key points are extracted from video data, and lip movement recognition is performed based on these key points to obtain lip movement recognition results. This includes: for each frame of the video data, a pre-defined face detection model is used to detect and determine the spatial location region of the face in the image; the detected spatial location region of the face is input into a facial key point detection network, and the coordinate information of the corresponding facial key points is output; based on the facial key point coordinate information, lip geometric feature parameters are calculated; these parameters include the distance between the upper and lower lips, the distance between the inner and outer corners of the mouth, and the area of the mouth contour; based on a sliding window of time series, the lip geometric feature parameters of consecutive frames are statistically analyzed to obtain the lip movement amplitude variation results; based on the lip movement amplitude variation results, the presence of a lip movement event in the current video frame is detected. This method can accurately identify the lip movement state of a speaker in a video, achieving real-time detection and quantitative analysis of speaking activities.
[0011] Optionally, human key points are extracted from video data, and action recognition is performed based on these key points to obtain target action recognition results. This includes: detecting each frame of the video data using a pre-defined human detection model to determine the human spatial region of the target human in the image; estimating the pose of the human spatial region and outputting the coordinate information of the human key points corresponding to the target human; constructing a temporal sequence of human key points based on the coordinate information of the human key points; performing global attention modeling on the temporal sequence of human key points based on the temporal sequence of human key points to obtain a high-dimensional temporal feature representation of the dynamic features of human actions; inputting the high-dimensional temporal feature representation of the dynamic features of human actions into a classification layer to obtain the action recognition results of the target human; the action recognition results include standing up, sitting down, and remaining still. This method provides a reliable data foundation for action recognition in classroom management or multimodal interaction, improving the accuracy and real-time performance of action analysis.
[0012] Optionally, based on the direction of the sound source, lip movement recognition results, and target action recognition results, the current shooting screen of the camera device is controlled, including: obtaining candidate seat areas corresponding to the direction of the sound source based on the direction of the sound source and the seating layout of the classroom; when the current speaker is detected within the candidate seat area according to the lip movement recognition results, the current shooting screen of the camera device is switched to a close-up shot of the current speaker; when an interactive object is detected according to the target action recognition results, the current shooting screen of the camera device is switched to a close-up shot of the interactive object; when both the current speaker and the interactive object are detected according to the target action recognition results, the current shooting screen of the camera device is first switched to a close-up shot of the interactive object, and then switched to a close-up shot of the current speaker; if the current speaker is not detected, the current shooting screen of the camera device is switched to the overall view of the classroom or the teacher's view. This method, by integrating the direction of the sound source, lip movement state, and action recognition, achieves dynamic tracking and automatic screen switching of key figures in the classroom, improves the naturalness of the broadcast presentation, and enhances the intelligence and real-time performance of the broadcast system.
[0013] Optionally, the method further includes: acquiring a backup screen, which is a preparatory output screen following the close-up screen of the current speaker or the close-up screen of the interactive object, and displaying it as the screen to be switched during subsequent broadcast switching. This solution can prepare the next output screen in advance, making the broadcast switching process more continuous and smooth and reducing switching delays.
[0014] To address the aforementioned technical problems, one technical solution adopted in this application is as follows: An automatic classroom broadcast control device is provided, comprising: an audio and video data acquisition module for acquiring audio and video data of the current classroom scene; a sound source direction determination module for locating the sound source direction based on the audio data using a microphone array, and determining the current sound source direction; a lip movement recognition module for extracting facial key points from the video data and performing lip movement recognition based on the facial key points to obtain a lip movement recognition result; a target action recognition module for extracting human body key points from the video data and performing action recognition based on the human body key points to obtain a target action recognition result; a broadcast control screen determination module for controlling the current shooting screen of a camera device based on the sound source direction, the lip movement recognition result, and the target action recognition result, wherein the current shooting screen of the camera device is the broadcast control screen; and a broadcast control screen output module for outputting and displaying the broadcast control screen.
[0015] To solve the above-mentioned technical problems, one technical solution adopted in this application is: to provide an electronic device, including: a memory and a processor, the memory being connected to the processor, the processor being used to execute one or more computer programs stored in the memory, and when the processor executes one or more computer programs, causing the electronic device to implement a classroom automatic broadcasting method applied to the electronic device.
[0016] To solve the above-mentioned technical problems, one technical solution adopted in this application is to provide a non-volatile computer-readable storage medium that stores computer-executable instructions. When the computer-executable instructions are executed by an electronic device, the electronic device executes the classroom automatic broadcasting method described above. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application; Figure 2 This is a flowchart of an automatic classroom broadcasting method provided in an embodiment of this application; Figure 3 This is a flowchart of a method for determining the current sound source direction based on audio data and using a microphone array, according to an embodiment of this application. Figure 4 This is a schematic diagram of the structure of an automatic classroom broadcasting device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0020] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first" and "second" used in this application do not limit the data or execution order, but only distinguish between identical or similar items with essentially the same function and effect.
[0021] In various scenarios such as meetings, speeches, and lectures, it is usually necessary to have a director or a directing system to coordinate and switch between multiple camera feeds. The director's main responsibility is to comprehensively judge the current key speaker, interactive objects, or presentation content based on the situation on site, and promptly select the most appropriate camera feed for output, thereby ensuring the continuity and visual appeal of the overall footage.
[0022] Taking smart classrooms as an example, multiple cameras are typically configured in smart classroom or recording studio scenarios to provide comprehensive coverage of the teacher, students, and courseware display area, meeting the needs of multi-angle and multi-scenario teaching recording and live streaming. Current systems generally rely on manual or semi-automatic methods for switching footage. For example, a specialist selects appropriate shots in real-time based on the class progress using a control console or software; or a simple voice detection algorithm triggers the corresponding camera switch when a speaking source is detected. However, these methods have significant limitations: manual operation relies heavily on the experience and reaction speed of the director, making it difficult to guarantee the timeliness and naturalness of footage switching; while automatic switching based on single voice detection is easily affected by environmental noise, multiple people speaking simultaneously, etc., leading to misjudgments or frequent switching, impacting the viewing experience.
[0023] In view of this, the embodiments of this application provide a classroom automatic broadcasting method, device and electronic device that, by integrating audio and video multimodal information, achieves high-precision recognition of speakers and interactive objects in the classroom and natural and accurate screen switching, and can automatically complete camera control and target screen output without human intervention, thereby improving the intelligence level of recording and live broadcasting systems.
[0024] The automatic classroom broadcasting method proposed in this application is described below through specific embodiments.
[0025] Please see Figure 1 , Figure 1This is a schematic diagram of an application scenario provided by an embodiment of this application. Taking a smart classroom as an example, the application scenario specifically includes: a camera device 10, a microphone array 20, a processing module 30, and a display module 40. The processing module 30 is connected to the camera device 10, the microphone array 20, and the display module 40 respectively. These four modules work together to form a complete intelligent broadcasting system, capable of multimodal perception, intelligent analysis, and natural image switching in the classroom scene.
[0026] The camera device 10 can consist of multiple high-definition cameras distributed in different locations within the classroom, such as the teacher's area, student area, and blackboard or presentation area. The camera device 10 is used to collect classroom video data from multiple angles. The cameras can be configured with pan-tilt-zoom (PTZ) functionality to achieve automatic rotation, zoom, and tracking, automatically focusing on the speaker or interacting object based on recognition results. The microphone array 20 can be installed at the front of the classroom, on the ceiling, or on the podium area to collect multi-channel audio signals. Through relevant array signal processing algorithms, it can achieve sound source direction localization, speech enhancement, and noise suppression, thereby accurately determining the current sound source location. The processing module 30 is the core computing unit and can be composed of a processor or GPU with AI computing capabilities. This processing module 30 is used to fuse and analyze the collected audio and video data, including sound source direction recognition, facial key point detection, lip movement recognition, human body key point extraction, and action recognition. Based on the multimodal recognition results, the processing module 30 can comprehensively determine the current speaker or interacting object, automatically generating camera control and scene switching commands to achieve natural and accurate camera transitions. The display module 40 is used to present the final broadcast screen and can be connected to a recording system, classroom display screen, or remote teaching platform, etc.
[0027] In some embodiments, the camera device 10, microphone array 20, and processing module 30 can be integrated into a single physical device to form an integrated intelligent broadcast control terminal. This intelligent broadcast control terminal can adopt a modular design, encapsulating the camera module, microphone array board, and processing motherboard within the same housing. The three components achieve real-time transmission and synchronous processing of audio and video data via a high-speed bus or inter-board communication interface. The camera module is used to acquire video signals, and the microphone array is used for multi-channel audio acquisition and sound source localization. All acquired data is directly input to the processing module. The processing module incorporates a high-performance AI chip or GPU unit to execute algorithms such as face detection, lip movement recognition, action recognition, and multimodal fusion, and automatically generates screen switching instructions based on the recognition results. This integration method achieves tight coupling between audio and video acquisition and intelligent analysis, significantly reducing system latency and wiring complexity.
[0028] Of course, in some complex or large-space application scenarios, the camera device 10, microphone array 20, and processing module 30 can also be deployed separately and interconnected and synchronized through wired or wireless networks. This distributed architecture facilitates flexible placement of cameras and microphones, covering a wider acquisition area to adapt to smart classrooms or recording systems of different sizes and functional requirements.
[0029] Figure 1 The application scenarios shown are merely examples to illustrate the principles and implementation of the methods in this application, and do not constitute a limitation on the scope of application. In fact, the automatic broadcasting method described in this application is not only applicable to smart classroom scenarios, but can also be flexibly applied to various occasions such as meetings, speeches, distance education, seminars and training, online live streaming, and enterprise video conferencing, depending on actual needs. In different application environments, the number, installation location, and interaction methods of camera equipment, microphone arrays, and processing modules can be adjusted or expanded according to spatial layout and usage requirements. An integrated structure can be used for unified deployment, or a distributed architecture can be used to cover a larger acquisition area, thereby achieving natural and accurate target recognition and automatic image switching in various environments.
[0030] Please see Figure 2 , Figure 2 This is a flowchart of an automatic classroom broadcasting method provided in an embodiment of this application. Specifically, this method can be executed by the aforementioned processing module 30. The method includes, but is not limited to, the following steps: S11. Obtain the audio and video data of the current classroom scene.
[0031] The video data of the current classroom scene can be acquired by the aforementioned camera device 10. Specifically, the video data can be acquired in real time by a close-up camera (e.g., a 4K@60fps industrial camera or a smart classroom front-end camera module), which has a high-resolution sensor and high frame rate acquisition capability (e.g., 60 frames / second). The processing module 30 executing this method can call the camera acquisition interface through the relevant driver interface or video acquisition module to receive the video stream data in real time. Each frame of video data can be numbered by timestamp. Each frame of video data contains visual information about people and the environment in the classroom scene.
[0032] The audio data for the current classroom scenario can be acquired by the aforementioned microphone array 20. This microphone array 20 supports multi-channel synchronous acquisition, with a sampling rate of up to 48kHz, and supports either a USB interface or a dedicated array acquisition module interface. Audio and video data frames can be buffered via shared memory or a buffer queue for use by the processing module 30.
[0033] In some embodiments, acquiring audio and video data of the current classroom scene refers to the process of real-time audio and video data acquisition of the classroom environment through a microphone array and a camera. Specifically, multi-channel audio signals of the classroom environment are acquired in real time through a microphone array at a sampling rate of 48kHz, and video streams of the classroom scene are acquired in real time through a close-up camera at 4K resolution and a frame rate of 60 frames per second; both audio and video data are timestamped to achieve synchronous acquisition and real-time acquisition of audio and video.
[0034] S12. Based on audio data, the direction of the sound source is located using a microphone array to determine the current direction of the sound source.
[0035] In this embodiment, microphones are deployed at different locations in the current application scenario, and these microphones together form a microphone array. Since there are slight differences in the time it takes for sound waves to arrive at each microphone (i.e., time delay), by calculating these arrival time differences and combining them with the spatial geometric position of the microphone array, the direction of sound wave propagation can be analyzed and estimated, and the approximate direction or azimuth of the current sound source can be determined.
[0036] Specifically, such as Figure 3 As shown, based on audio data, sound source direction is located using a microphone array to determine the current sound source direction, including: S121. Based on audio data, acquire multi-channel audio signals from the microphone array.
[0037] This audio data consists of sound data simultaneously collected by the microphones in the microphone array from the same application scenario. Each microphone independently samples the sound signal, forming an audio channel. The sound signals collected by all the microphones together constitute a multi-channel audio signal.
[0038] S122. Windowing is applied to each channel of the multi-channel audio signal to obtain the windowed multi-channel audio signal. This windowing process applies a preset window function to each channel of the audio signal and performs weighted processing.
[0039] The windowing process is part of the audio signal preprocessing stage, primarily aimed at providing a smooth and stable signal input for subsequent frequency domain analysis or sound source localization. Specifically, each audio signal in a multi-channel audio signal is divided into several fixed-length time frames (e.g., 20ms or 40ms), and a preset window function (e.g., Hanning window, Hamming window, or rectangular window) is applied to each frame. This window function is a weighting function that applies a smooth transition at both ends of the signal frame, preventing abrupt changes when the signal is truncated in the time domain, thus reducing "spectral leakage" during frequency domain transformation. This weighting process preserves the main energy of the signal while suppressing boundary effects, ensuring smooth transitions between different frames. Windowing makes the signal more natural and accurate during framing and transform analysis. Hanning, Hamming, and rectangular windows are all different types of window functions; they all perform weighted smoothing on each frame of data after signal framing to reduce spectral leakage. The appropriate window function can be selected based on application requirements.
[0040] S123. Perform Fourier transform on the windowed multi-channel audio signals to obtain the time-frequency representation of each channel audio signal.
[0041] One approach is to use the Short-Time Fourier Transform (STFT), which transforms the windowed multi-channel audio signal from the time domain to the frequency domain to analyze the characteristics of sound changes in time and frequency.
[0042] In the steps described above, each microphone's audio signal, after framing and windowing, is divided into many short time segments (e.g., 20ms per frame). In this step, a Fourier transform is performed on the windowed signal of each frame, converting the frame's time-domain waveform into amplitude and phase information in the frequency domain. After all frames are transformed sequentially, the time-frequency representation of that channel is obtained—a two-dimensional feature where time and frequency change together. Therefore, time-frequency representation is a form of representation that combines the time and frequency information of a signal. For multi-channel audio signals, each channel undergoes this processing separately to obtain multi-channel time-spectrum data. The obtained time-frequency representation reflects the frequency components of the signal within different time segments and provides characteristic information for subsequent time delay estimation, sound source localization, etc.
[0043] S124. Based on the time-frequency representation of the audio signal of each channel, obtain the first propagation delay between any two microphones in the microphone array.
[0044] Because sound waves from the same source arrive at different microphones at different times during propagation, the audio signals captured by different microphones exhibit slight time delays. In this step, within the same time window, the correlation between any two audio signals is calculated at different time offsets. By comparing the similarity metrics corresponding to different time offsets, the optimal matching position for time alignment between the two signals can be obtained. The time offset with the largest similarity metric corresponds to the propagation delay between the signals, which is the time difference between the sound wave propagating from one microphone to another; this time difference is the first propagation delay. In other words, the first propagation delay is the actual time it takes for the sound wave to propagate between the two microphones.
[0045] Specifically, based on the time-frequency representation of each channel's audio signal, audio signals belonging to the same time window are selected; for audio signals belonging to the same time window, the correlation between the first audio signal and the second audio signal is calculated at different time offsets to obtain the similarity metric corresponding to the time offset; based on the similarity metrics corresponding to all obtained time offsets, the time offset with the largest similarity metric is selected, and the largest time offset is used as the first propagation delay between the two microphones; the two microphones refer to the microphones corresponding to the first audio signal and the second audio signal.
[0046] The time offset is an artificially introduced time alignment difference. It involves fixing one signal (e.g., the first audio signal) while shifting (delaying or advancing) another signal (the second audio signal) along the time axis by several sampling points (i.e., a discrete representation of the time offset), and then calculating the similarity between the two signals at that offset. In this step, the correlation is calculated for a series of candidate time offsets, and the offset with the highest similarity metric is found. This highest offset represents the propagation delay between the microphones corresponding to the first and second audio signals. This correlation calculation measures the similarity between the two signals at different time offsets. It can quantify the similarity between the signals acquired by the two microphones at different time offsets using a cross-correlation function or its improved methods (such as normalized cross-correlation). The first audio signal is the audio signal acquired by one microphone in the microphone array, and the second audio signal is the audio signal acquired by another microphone in the same microphone array.
[0047] A higher correlation metric indicates a greater similarity between the two signals at that time offset, meaning that the offset most likely corresponds to the actual time difference of sound wave propagation. Therefore, the time offset with the highest similarity metric is chosen because it corresponds to the optimal alignment of the two signals in time, representing the actual time difference of sound wave propagation from one microphone to another. This provides a reliable foundation for subsequent sound source localization and is a crucial prerequisite for ensuring the accuracy and robustness of the localization results.
[0048] It should be noted that in a microphone array, any two microphones can be selected as the calculation objects, and the first propagation delay between the two microphones can be obtained using the method described above. This process can be repeated for all possible microphone pairs in the microphone array to obtain the first propagation delay of all microphone pairs.
[0049] S125. Based on the first propagation delay and the spatial geometric position of the microphone array, determine the candidate direction of the sound source and obtain the spatial probability distribution corresponding to the candidate direction of the sound source.
[0050] Determining candidate directions of a sound source and obtaining the corresponding spatial probability distributions is equivalent to determining the possible locations of the sound source and their probability distributions within space. Candidate directions are the most likely directions or locations where the sound source will occur. By quantifying the confidence level of each direction, the probability distribution of the sound source in space can be obtained.
[0051] Specifically, based on the spatial geometric position of the microphone array, the spatial range of the sound source is determined, and multiple candidate directions are generated within the spatial range. According to the candidate directions and the spatial geometric position of the microphone array, the propagation time of the sound wave to each microphone under each candidate direction is obtained, and the second propagation delay of each pair of microphones is obtained based on the propagation time. The first propagation delay and the second propagation delay are compared to calculate the matching degree between each candidate direction and the actual observation. Based on the matching degree, all candidate directions are normalized to obtain the spatial probability distribution corresponding to the candidate directions.
[0052] The spatial range of the sound source can be determined by defining its possible azimuth, elevation, or planar coordinate range based on the microphone array's installation location, direction, and visible / audible range, as well as the actual application scenario (e.g., a classroom). Candidate directions are a set of possible sound source directions (or angles) pre-generated within the sound source's spatial range for localization. The spatial range can be discretized into several uniform or non-uniform angle points (specifically, candidate azimuth and elevation angles) to determine candidate directions. For example, for a horizontal angle from 0° to 180°, one candidate direction can be generated for every 1°, resulting in 181 candidate directions. The second propagation delay is the theoretical propagation time difference of the sound wave from the assumed sound source to each microphone in the microphone array for a given candidate direction. The first propagation delay is the actual measured arrival time difference of the sound wave between the microphone pairs. The first propagation delay of each pair of microphones is compared with the second propagation delay, for example, by calculating the sum of squared errors or a similarity metric; a higher degree of matching indicates that the candidate direction is more likely to be the actual sound source direction. Finally, the matching degree of all candidate directions is normalized (e.g., using softmax) to obtain the sound source probability corresponding to each candidate direction. Ultimately, a probability distribution map of the sound source in space can be obtained, with the direction with the highest probability being the most likely sound source direction.
[0053] This embodiment can match the actual observed microphone delay with the theoretical delay of the candidate direction, accurately assess the possible location of the sound source in space, and thus provide reliable probabilistic information for sound source localization.
[0054] S126. Select the candidate direction of the sound source with the largest spatial probability distribution and use it as the current sound source direction.
[0055] Besides the aforementioned method based on microphone array delay matching for candidate directions, other techniques exist for determining sound source direction. For example, the phase difference between microphone pairs can be analyzed at different frequencies in the short-time Fourier transform domain to accumulate and estimate the sound source direction. Neural networks can also be used to directly predict the sound source azimuth or probability distribution from multi-channel audio or spectrogram features. Furthermore, these methods can be combined to balance real-time performance, accuracy, and robustness.
[0056] S13. Extract facial key points from video data, and perform lip movement recognition based on facial key points to obtain lip movement recognition results.
[0057] This lip movement recognition result is an analysis of the speaker's lip movement state, used to determine whether the speaker is currently speaking (open / closed mouth state). Specifically, for each frame of the video data, a preset face detection model is used to detect and determine the spatial location region of the face in the image; the detected spatial location region of the face is input into a facial landmark detection network, which outputs the coordinate information of the facial landmarks corresponding to the face; based on the coordinate information of the facial landmarks, lip geometric feature parameters are calculated, including the distance between the upper and lower lips, the distance between the inner and outer corners of the mouth, and the area of the mouth contour; based on a sliding window of time series, the lip geometric feature parameters of consecutive frames are statistically analyzed to obtain the lip movement amplitude change results; based on the lip movement amplitude change results, the presence of a lip movement event in the current video frame is detected.
[0058] The preset face detection model can be a deep learning model (such as YOLOv5-face, YOLOv8-face, etc.). The spatial location region of the face refers to the specific location range of the face in the image within a video frame, which can be represented by two-dimensional coordinates. The spatial location region of the face provides the target area for keypoint detection, avoiding searching the entire image. The face keypoint detection network can use the HRNet (High-Resolution Network) model, which can provide high-precision face keypoint information; other models can also be used, such as FAN (Face Alignment Network, face keypoint detection based on convolutional neural networks). The output face keypoint coordinate information refers to a set of two-dimensional or three-dimensional coordinate points output by the face keypoint detection network, representing the key positions of the face in the image or space, with the total number of keypoints such as 68, 98, etc.
[0059] Next, facial landmark information is transformed into quantifiable lip movement features, including the distance between the upper and lower lips, the distance between the inner and outer corners of the mouth, and the area of the mouth contour. Point sets representing the positions of the upper and lower lips and the corners of the mouth can be extracted from the lip landmark coordinates obtained by the facial landmark detection network. Then, the distance between the upper and lower lips, i.e., the average vertical distance between the upper and lower lip landmarks, is calculated to measure the degree of mouth opening and closing; the distance between the inner and outer corners of the mouth, i.e., the horizontal or spatial distance between the left and right corners of the mouth, is also calculated to describe the lateral opening or smiling amplitude of the mouth; simultaneously, polygons are generated using lip contour landmarks, and their contour areas are calculated to reflect the overall size of the lips opening or closing. Through these geometric feature parameters, lip shape and movement can be quantified into numerical indicators.
[0060] Next, the lip geometric feature parameters of consecutive frames are converted into lip movement information that varies over time to capture the speaker's mouth movement patterns. Specifically, a sliding time window (such as several frames or several milliseconds) can be set in the video sequence, and the lip geometric feature parameters (distance between the upper and lower lips, distance between the corners of the mouth, area of the mouth contour, etc.) of consecutive frames within each window can be statistically analyzed, for example, calculating the mean, variance, extreme values, or amplitude of change. By processing the entire video sequence through continuous sliding windows, the curve of lip movement over time can be obtained, that is, the result of the change in lip movement amplitude.
[0061] Finally, based on the change in lip movement amplitude, the presence of a lip movement event in the current video frame is detected. A lip movement event refers to a time segment or video frame state in which the speaker's lips move significantly, corresponding to actions such as opening and closing the mouth or changes in mouth shape. For example, if the lip movement amplitude exceeds a preset threshold (such as a change in the distance between the upper and lower lips or the area of the mouth contour exceeding a certain value), then a lip movement event is considered to exist in that frame; otherwise, no lip movement event is considered to exist. This process can output a binary result indicating whether a lip movement event exists in each frame.
[0062] This embodiment achieves quantitative monitoring of a speaker's mouth movements by extracting facial key points from continuous video frames, calculating lip geometric features, and performing time series analysis. This method accurately captures lip opening and closing and movement states, providing a reliable and quantifiable data foundation for speaker recognition and improving the accuracy and real-time performance of speaker behavior analysis.
[0063] S14. Extract human body key points from video data, and perform action recognition based on human body key points to obtain target action recognition results.
[0064] Specifically, for each frame of the video data, a pre-defined human detection model is used to detect and determine the human spatial region of the target human body in the image; pose estimation is performed on the human spatial region, outputting the coordinate information of the corresponding human key points; based on the human key point coordinate information, a temporal sequence of human key point fragments is constructed in chronological order; based on the temporal sequence of human key point fragments, global attention modeling is performed on the temporal dimension of the human key point sequence fragments to obtain a high-dimensional temporal feature representation of the dynamic features of human actions; the high-dimensional temporal feature representation of the dynamic features of human actions is input into the classification layer to obtain the action recognition result of the target human body. This action recognition result includes standing up, sitting down, and remaining still.
[0065] The above scheme can accurately identify the motion state of a target human body in dynamic scenes. First, a pre-defined human detection model is used for each frame of the video data to determine the spatial region of the target human body in the image, providing a precise target region for subsequent pose analysis. Specifically, a panoramic camera can be used to detect the human body in the video frames, obtaining a sequence of bounding boxes (bboxes), which serves as input data for motion analysis. Then, pose estimation is performed within the detected human body region, outputting the coordinates of key points. These key points cover the positions of the main joints of the human body, providing basic data for motion feature extraction. Next, the key point coordinates are constructed in chronological order to form a temporal sequence of human key point fragments, ensuring the continuity and integrity of the motion dynamic information in the temporal dimension. A temporal Transformer Encoder can be used to perform global attention modeling on the key point sequence, capturing the dependencies between different time steps and joints, thereby generating a high-dimensional temporal feature representation that comprehensively reflects the dynamic changes of human motion. Finally, these high-dimensional temporal features are input into the classification layer to achieve the classification and recognition of target human actions, including actions such as "standing up," "sitting down," and "remaining still." At the same time, spatial positioning can be performed by combining key point location information (such as the seat number corresponding to the target standing up).
[0066] This method can accurately identify the action states of a target human body in a video in real time, including "standing up," "sitting down," and "stillness," and can determine the specific location of the target human body by combining spatial location information. This provides a reliable data foundation for action recognition in classroom management or multimodal interaction, improving the accuracy and real-time performance of action analysis. Besides the temporal Transformer method based on human keypoints, there are other human action recognition schemes, such as CNN methods based on image / video frames. These methods input video frames into a 2D / 3D convolutional network for action classification, extracting features in the spatial and temporal dimensions to directly generate action categories.
[0067] S15. Based on the direction of the sound source, the lip movement recognition result, and the target action recognition result, control the current shooting screen of the camera device, wherein the current shooting screen of the camera device is the director's screen.
[0068] By integrating multimodal information such as sound source localization, lip movement detection, and action recognition, the system comprehensively judges speaking and interactive behaviors in classroom scenarios, automatically controls camera equipment to select the optimal shooting frame, and thus presents real-time broadcast footage.
[0069] The people highlighted in the image, such as a teacher speaking or a student asking a question, include the current speaker and / or the person interacting with the speaker. The current speaker is the person currently speaking, which can be determined using the sound source direction localization and lip movement recognition results described above. For example, if the microphone array detects a sound source in a certain direction, and the lips of a person in that direction are moving, that person can be identified as the speaker. The person interacting with the speaker refers to those who interact with the speaker, such as the questioner, someone raising their hand to respond, or someone participating in the discussion. This can be determined through action recognition (such as standing up, raising a hand, etc.) combined with positional judgment.
[0070] Specifically, based on the direction of the sound source and the corresponding seating arrangement in the classroom, candidate seating areas corresponding to the direction of the sound source are obtained; within the candidate seating areas: When the current speaker is detected based on the lip movement recognition results, the current shooting screen of the camera device is switched to a close-up shot of the current speaker; When an interactive object is detected based on the target action recognition result, the current shooting screen of the camera device is switched to the close-up screen corresponding to the interactive object; When the current speaker is detected based on the lip movement recognition result and the interactive object is detected based on the target action recognition result, the current shooting screen of the camera device is first switched to the close-up screen corresponding to the interactive object, and then switched to the close-up screen corresponding to the current speaker. If the current speaker is not detected, the camera will switch the current view to the overall view of the classroom or the teacher's view.
[0071] The process involves using a microphone array to determine the direction of the sound source, combined with the classroom seating arrangement, to map the positions of potential speakers (students or teachers) into a candidate seating area, thus narrowing down the scope of subsequent analysis. Within this candidate area, lip movement recognition is used to determine if an individual is speaking, identifying them as the current speaker and selecting their corresponding camera feed for output, thus focusing the image on the current speaker. Simultaneously, human motion recognition results can be used within the candidate area to detect actions such as standing up and raising a hand, identifying these individuals as interactive subjects and displaying their corresponding camera feeds. If both the current speaker and an interactive subject are detected, the interactive subject is prioritized for display, and the feed switches to a close-up of the current speaker after the interaction has been shown, ensuring the classroom interaction is fully presented and maintaining narrative continuity. If the current speaker is not detected within the candidate area, the system automatically reverts to the overall classroom view or the teacher's view. A close-up view refers to a shot of the target person at close range, where the subject occupies a large proportion of the frame and clearly shows facial expressions, lip movements, and detailed actions.
[0072] In some embodiments, during the process of switching the current camera view to a close-up view of the current speaker, or switching the current camera view to a close-up view of the interactive object, or switching the current camera view first to a close-up view of the interactive object and then to a close-up view of the current speaker, or switching the current camera view to a global view or teacher view of the classroom, the pan-tilt-zoom (PTZ) command can be used to finely adjust the pan-tilt angle, horizontal rotation direction, and lens focal length of the close-up camera. This allows the camera to quickly, smoothly, and accurately aim at the target object (the current speaker and / or interactive object), achieving intelligent broadcast control. Simultaneously, the system can output the PTZ-controlled footage as a real-time PGM (Program) main screen, and can overlay prompts, transition effects, or record synchronously as needed to improve the smoothness and professionalism of the live / recorded classroom footage.
[0073] When the processing module acts as the execution subject of the above method, it generates a screen switching instruction by analyzing the sound source direction, lip movement recognition result and target action recognition result, and sends the instruction to the camera device or its control interface to control the camera device to adjust the PTZ angle, direction and focal length to achieve close-up or panoramic shooting of the target object.
[0074] This embodiment achieves dynamic tracking and automatic screen switching of key figures in the classroom by integrating sound source direction, lip movement state, and action recognition, which improves the naturalness of the broadcast presentation and enhances the intelligence and real-time performance of the broadcast system.
[0075] S16. Output and display the above broadcast screen.
[0076] After generating the PTZ-adjusted broadcast screen, this broadcast screen is encoded and output as the main PGM (Program) screen, and simultaneously transmitted to the display device (such as the display module 40 mentioned above). The display device can include a broadcast console monitor, a large screen display, a teaching recording screen, or a preview / live window of a remote live streaming platform. During the output process, video signal formatting, frame rate and resolution adaptation, and image transition effect rendering can be automatically completed, ensuring that the final broadcast screen is presented stably, continuously, and clearly on the corresponding display device. This ensures natural classroom screen transitions, professional visual effects, and facilitates real-time viewing for teachers, students, or remote viewers.
[0077] In some embodiments, the method further includes: acquiring a backup screen. This backup screen is a pre-selected output screen following a close-up shot of the current speaker or a close-up shot of an interactive object, and is displayed as the screen to be switched during subsequent broadcast switching. Specifically, the backup screen can be the global screen corresponding to the classroom or the teacher's screen. This embodiment can prepare the next output screen in advance, making the broadcast switching process more continuous, smoother, and reducing switching delays.
[0078] For example, in a smart classroom, after a teacher asks a question, the microphone array detects that a student has spoken and locates the sound source at position 2, row 3 of the seating arrangement using a sound source algorithm. The close-up camera footage corresponding to that position is retrieved, facial landmark detection is performed on the student's face, and lip geometry is analyzed. It is confirmed that the lip movement lasted for 3 seconds, indicating that the student was speaking. Subsequently, the broadcast control module receives the signal and switches the main PGM screen to the close-up of that student, while the PV preview maintains a panoramic view of other areas. This achieves natural focusing and switching of the classroom footage, allowing remote or pre-recorded viewers to clearly see the student currently speaking.
[0079] For example, in a smart classroom, when a teacher asks a student to stand up to answer a question, the temporal Transformer model is used to perform motion recognition on the classroom video, detecting that the student in row 2, seat 4 has stood up. Simultaneously, the microphone array uses a sound source algorithm to detect that the direction of the sound source matches the student's position, confirming that the student is speaking and identifying them as the current speaker. The broadcast control module then switches the main PGM screen to a close-up camera view of the student, allowing remote or pre-recorded viewers to clearly see the student standing to answer the question. Meanwhile, the PV preview maintains a panoramic view or a view of the teacher, achieving natural focusing and dynamic switching of the classroom footage.
[0080] For example, in a smart classroom, the microphone array detects the main sound source from the left side of the classroom, and the system, combined with the seating layout, determines the candidate student area. Within this area, lip movement recognition detects that a student in row 3, seat 1 is continuously speaking; simultaneously, the target action recognition module detects obvious interactive actions such as raising a hand and nodding in the seat directly in front of the student (row 3, seat 2). Based on this, it is determined that not only is there a speaker, but also an interactive object related to the content of the speech. To naturally present the classroom interaction process, the system first switches the PGM main screen to a close-up of the interactive object (student in row 3, seat 2) to highlight their interactive behavior; then, it smoothly switches back to a close-up of the current speaker (student in row 3, seat 1) to display their complete speech. Throughout the process, the PV preview continuously displays the global screen to assist in verifying the director's logic, achieving accurate capture and reasonable presentation of classroom interaction relationships under multimodal fusion recognition.
[0081] It should be noted that in the above embodiments, there is no necessarily a certain order between the steps. Those skilled in the art can understand from the description of the embodiments of this application that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0082] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an automatic classroom broadcasting device provided in an embodiment of this application. The automatic classroom broadcasting device 400 includes: The audio and video data acquisition module 401 is used to acquire audio and video data of the current classroom scene. The sound source direction determination module 402 is used to locate the sound source direction using a microphone array based on the audio data, thus determining the current sound source direction. The lip movement recognition module 403 is used to extract facial key points from the video data and perform lip movement recognition based on these key points to obtain lip movement recognition results. The target action recognition module 404 is used to extract human body key points from the video data and perform action recognition based on these key points to obtain target action recognition results. The director's screen determination module 405 is used to control the current shooting screen of the camera device based on the sound source direction, lip movement recognition results, and target action recognition results; the current shooting screen of the camera device is the director's screen. The director's screen output module 406 is used to output and display the director's screen.
[0083] The aforementioned automatic classroom broadcasting device 400 can be a software module. This software module includes several instructions, which are stored in a memory. The processor can access the memory, call the instructions, and execute them to complete the automatic classroom broadcasting methods described in the above embodiments.
[0084] In some embodiments, the aforementioned automatic classroom broadcast control device 400 can also be constructed from hardware devices. For example, the automatic classroom broadcast control device 400 can be constructed from one or more chips, and the chips can work in coordination to complete the automatic classroom broadcast control method described in the various embodiments above. As another example, the aforementioned automatic classroom broadcast control device 400 can also be constructed from various logic devices, such as general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontrollers, ARM (Acorn RISC Machine) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0085] It should be noted that the aforementioned automatic classroom broadcasting device 400 can execute the automatic classroom broadcasting method for electronic devices provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the embodiments of the automatic classroom broadcasting device 400 can be found in the aforementioned automatic classroom broadcasting method for electronic devices provided in the embodiments of this application.
[0086] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 50 includes one or more processors 51 and a memory 52. The memory 52 is connected to one or more processors 51, for example, via a bus.
[0087] Processor 51 is configured to support the electronic device 50 in performing the corresponding functions in the methods described in the above method embodiments. Processor 51 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0088] Memory 52 is used to store program code, etc. Memory 52 may include volatile memory (VM), such as random access memory (RAM); memory may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory 52 may also include combinations of the above types of memory.
[0089] The memory 52 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the automatic classroom broadcasting method in the embodiments of this application. The processor 51 executes various functional applications and data processing of the automatic classroom broadcasting method and automatic classroom broadcasting device by running the non-volatile software programs, instructions, and modules stored in the memory 52, that is, it realizes the functions of each module or unit of the automatic classroom broadcasting method and automatic classroom broadcasting device provided in the above method embodiments.
[0090] The memory 52 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function. The data storage area may store data created based on the use of the automated classroom broadcasting device. In some embodiments, the memory 52 may include memory remotely located relative to the processor 51, and this remote memory may be connected to the automated classroom broadcasting device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0091] The one or more modules are stored in the memory 52. When executed by the one or more processors 51, they execute the automatic classroom broadcasting method in any of the above method embodiments. For example, they execute the method steps described in the above method embodiments to realize the functions of the modules described in the above device embodiments.
[0092] The electronic device 50 in this application embodiment can be the processing module in the above embodiment, specifically it can be the host of the automatic broadcasting system of a smart classroom / recording classroom, or it can be a smart terminal (such as an all-in-one machine) that integrates microphone array, camera control and AI processing capabilities.
[0093] This application provides a non-volatile computer-readable storage medium storing computer-executable instructions that are executed by one or more processors, for example... Figure 5 One of the processors 51 can enable the one or more processors to execute the automatic classroom broadcasting method in any of the above method embodiments, for example, to execute the above-described method. Figure 2 Method steps S11 to S16, Figure 3 Steps S121 to S126 in the method are implemented. Figure 4 The functions of modules 401-406 in the document.
[0094] This application provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions, which, when executed by an electronic device, enable the electronic device to perform the automatic classroom broadcasting method in any of the above method embodiments, for example, to perform the above-described... Figure 2 Method steps S11 to S16, Figure 3 Steps S121 to S126 in the method are implemented. Figure 4 The functions of modules 401-406 in the document.
[0095] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0096] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A method for automatic classroom broadcasting, characterized in that, include: Acquire audio and video data for the current classroom setting; Based on the audio data, the sound source direction is located using a microphone array to determine the current sound source direction. Facial key points are extracted from the video data, and lip movement recognition is performed based on the facial key points to obtain lip movement recognition results; Human body key points are extracted based on the video data, and action recognition is performed based on the human body key points to obtain the target action recognition result; Based on the sound source direction, the lip movement recognition result, and the target action recognition result, the current shooting screen of the camera device is controlled, and the current shooting screen of the camera device is the director's screen; Output and display the broadcast screen.
2. The method according to claim 1, characterized in that, The step of locating the sound source direction using a microphone array based on the audio data to determine the current sound source direction includes: Based on the audio data, acquire multi-channel audio signals from the microphone array; Each channel of the multi-channel audio signal is windowed to obtain a windowed multi-channel audio signal; the windowing process is used to apply a preset window function to each channel of the audio signal and perform weighted processing. The windowed multi-channel audio signals are subjected to Fourier transform to obtain the time-frequency representation of each channel audio signal; Based on the time-frequency representation of each channel's audio signal, the first propagation delay between any two microphones in the microphone array is obtained; Based on the first propagation delay and the spatial geometric position of the microphone array, the candidate direction of the sound source is determined, and the spatial probability distribution corresponding to the candidate direction of the sound source is obtained; Select the candidate direction of the sound source with the largest spatial probability distribution, and use it as the current sound source direction.
3. The method according to claim 2, characterized in that, The method of obtaining the first propagation delay between any two microphones in the microphone array based on the time-frequency representation of the audio signal of each channel includes: Based on the time-frequency representation of each channel's audio signal, select audio signals belonging to the same time window; For audio signals belonging to the same time window, at different time offsets, the correlation between the first audio signal and the second audio signal is calculated to obtain the similarity measure corresponding to the time offset; Based on the similarity metric corresponding to all the obtained time offsets, the time offset with the largest similarity metric is selected, and the largest time offset is used as the first propagation delay between the two microphones; the two microphones refer to the microphones corresponding to the first audio signal and the second audio signal.
4. The method according to claim 2, characterized in that, The step of determining the candidate direction of the sound source based on the first propagation delay and the spatial geometric position of the microphone array, and obtaining the spatial probability distribution corresponding to the candidate direction of the sound source, includes: Based on the spatial geometric position of the microphone array, the spatial range of the sound source is determined, and multiple candidate directions are generated within the spatial range. Based on the candidate direction and the spatial geometric position of the microphone array, the propagation time of the sound wave to each microphone under each candidate direction is obtained, and the second propagation delay of each pair of microphones is obtained based on the propagation time. The first propagation delay is compared with the second propagation delay to calculate the degree of matching between each candidate direction and the actual observation; Based on the matching degree, all candidate directions are normalized to obtain the spatial probability distribution corresponding to the candidate directions.
5. The method according to claim 1, characterized in that, The step of extracting facial key points from the video data and performing lip movement recognition based on the facial key points to obtain lip movement recognition results includes: For each frame of the video data, a preset face detection model is used to detect and determine the spatial location region of the face in the image; The detected spatial location region of the face is input into the facial landmark detection network, and the coordinate information of the facial landmark corresponding to the face is output. Based on the facial key point coordinate information, calculate the lip geometric feature parameters; the lip geometric feature parameters include the distance between the upper and lower lips, the distance between the inner and outer corners of the mouth, and the area of the mouth outline; Based on a time-series sliding window, the geometric feature parameters of the lips in consecutive frames are statistically analyzed to obtain the results of changes in lip movement amplitude. Based on the changes in lip movement amplitude, detect whether there is a lip movement event in the current video frame.
6. The method according to claim 1, characterized in that, The step of extracting human body key points based on the video data and performing action recognition based on the human body key points to obtain target action recognition results includes: For each frame of the video data, a preset human detection model is used to detect and determine the human spatial region of the target human body in the image; Perform pose estimation on the human body spatial region and output the coordinate information of the human body key points corresponding to the target human body; Based on the coordinate information of the human body key points, a temporal sequence of human body key points is constructed in chronological order. Based on the temporal human key point sequence fragments, global attention modeling is performed on the human key point sequence fragments in the time dimension to obtain a high-dimensional temporal feature representation of human motion dynamic features. The high-dimensional temporal feature representation of the human body's dynamic features is input into the classification layer to obtain the action recognition result of the target human body; the action recognition result includes standing up, sitting down, and remaining still.
7. The method according to claim 1, characterized in that, The step of controlling the current shooting frame of the camera device based on the sound source direction, the lip movement recognition result, and the target action recognition result includes: Based on the direction of the sound source and the corresponding seating layout in the classroom, candidate seating areas corresponding to the direction of the sound source are obtained. When within the candidate seating area: When the current speaker is detected based on the lip movement recognition result, the current shooting screen of the camera device is switched to the close-up screen corresponding to the current speaker; When an interactive object is detected based on the target action recognition result, the current shooting screen of the camera device is switched to a close-up shot of the interactive object; When the current speaker is detected based on the lip movement recognition result and the interactive object is detected based on the target action recognition result, the current shooting screen of the camera device is first switched to the close-up screen corresponding to the interactive object, and then switched to the close-up screen corresponding to the current speaker. If the current speaker is not detected, the camera's current view will be switched to the global view or the teacher's view corresponding to the classroom.
8. The method according to claim 7, characterized in that, The method further includes: Obtain a backup screen, which is a preparatory output screen following the close-up screen corresponding to the current speaker or the close-up screen corresponding to the interactive object, and is displayed as the screen to be switched during subsequent broadcast switching.
9. An electronic device, characterized in that, include: A memory and a processor, the memory being connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, the processor causing the electronic device to perform the method as described in any one of claims 1 to 8 when executing the one or more computer programs.
10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions that, when executed by an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 8.
Citation Information
Cited By
A financial service video post-event quality inspection method, device, equipment and medium
CN122290225A