Classroom summary generation method and system, wearable device, storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2026-08-11
AI Technical Summary
目前,多数智能眼镜在课堂记录过程中生成的课堂纪要采用固定的结构化模板,手动调整格式需要复杂操作,难以适应学生的使用场景
[0007] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described classroom minutes generation method.
Smart Images

Figure CN120640103B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of intelligent device control technology, and more specifically, relates to a method and system for generating classroom minutes, a wearable device, and a storage medium. Background Technology
[0002] With the development of technology, wearable devices such as smart glasses are finding increasingly wider applications. Especially in classroom settings, smart glasses can record lecture content for easy review after class. Currently, most smart glasses generate class notes using fixed structured templates, and manually adjusting the format requires complex operations, making them unsuitable for students' usage scenarios. Summary of the Invention
[0003] The purpose of this application is to provide a method and system for generating classroom minutes, a wearable device, and a storage medium to improve the reliability of classroom minutes generation.
[0004] A first aspect of this application provides a method for generating classroom minutes, applied to a wearable device, comprising: Acquire audio and video information of the target area in the classroom, and extract features from the audio and video information to obtain audio and video features; Acquire behavioral information of users wearing wearable devices, and extract features from the behavioral information to obtain behavioral features; Based on the correlation between audio / video features and behavioral features, the teaching scenario corresponding to the audio / video information is determined; The audio and video information is processed according to the classroom minutes extraction strategy corresponding to the teaching scenario to generate classroom minutes corresponding to the audio and video information.
[0005] A second aspect of this application provides a classroom minutes generation system for wearable devices, comprising: The first acquisition module is used to acquire audio and video information of the target area in the classroom, and to extract features from the audio and video information to obtain audio and video features; The second acquisition module is used to acquire the behavioral information of users wearing wearable devices and extract features from the behavioral information to obtain behavioral features. The first control module is used to determine the teaching scenario corresponding to the audio and video information based on the correlation between audio and video features and behavioral features; The second control module is used to process audio and video information according to the classroom minutes extraction strategy corresponding to the teaching scenario, and generate classroom minutes corresponding to the audio and video information.
[0006] A third aspect of this application provides a wearable device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described classroom minutes generation method.
[0007] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described classroom minutes generation method.
[0008] A fifth aspect of this application provides a computer program product, including a computer program or computer executable instructions, wherein when the computer program or computer executable instructions are executed by a processor, the steps of the above-described classroom minutes generation method are implemented.
[0009] The beneficial effects of the classroom minutes generation method and system, wearable device, and storage medium provided in this application embodiment are as follows: This application embodiment determines classroom minutes extraction strategies for different teaching scenarios by determining the correlation between the audio-visual features of audio-visual information in a specified area of the classroom and user behavior features. Audio-visual information is processed according to different classroom minutes extraction strategies to generate personalized classroom minutes that meet the actual needs of customers, thereby improving the reliability of generated classroom minutes. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a classroom minutes generation method provided in an embodiment of this application; Figure 2 A structural block diagram of a classroom minutes generation system provided in an embodiment of this application; Figure 3 This is a schematic block diagram of a wearable device provided in an embodiment of this application. Detailed Implementation
[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0014] In this embodiment, the wearable device can be smart glasses or other head-mounted smart devices. In this embodiment, the wearable device is smart glasses.
[0015] The smart glasses in this embodiment are equipped with an audio and video acquisition module, which includes an audio acquisition unit and a video acquisition unit. The audio acquisition unit can acquire ambient audio around the smart glasses, and the video acquisition unit can acquire ambient video around the smart glasses.
[0016] Specifically, the smart glasses integrate an eye-tracking module, a video tracking module, and a microphone module. The eye-tracking module tracks the user's eye movements, collects and analyzes this information to determine the user's learning state. The video tracking module includes a camera that can record video of the area the user is looking in or facing. The microphone module collects and filters ambient sound.
[0017] Furthermore, the smart glasses can connect to the user's electronic device via Bluetooth. This device can then connect to a server, which can process and analyze the received information. In online mode, the server can perform detailed analysis of the audio and video information collected by the smart glasses to generate practical classroom minutes. Additionally, whether online or offline, the smart module within the smart glasses can adaptively analyze the collected audio and video information to generate classroom minutes tailored to the user's habits. This will be explained in detail below.
[0018] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a classroom minutes generation method provided in an embodiment of this application. The method can be executed by a wearable device and may include steps S101 to S104.
[0019] S101: Obtain audio and video information of the target area in the classroom, and extract features from the audio and video information to obtain audio and video features.
[0020] In this embodiment, the audio and video information includes audio information and video information. Audio information of a target area in the classroom can be collected using the audio acquisition unit in the smart glasses, and video information of the target area in the classroom can be collected using the video acquisition unit in the smart glasses.
[0021] The target area is a designated area in the classroom, which can be the area where the teacher is located, including the blackboard. The video capture unit can capture the teacher's behavior and posture in the target area, as well as the PPT slides or handwriting on the blackboard. The audio capture unit can capture the teacher's voice in the target area.
[0022] Alternatively, the target area can be the area where the teacher walks under the podium. The video capture unit can capture the teacher's behavior, posture, and expressions within the target area. The audio capture unit can capture the teacher's voice within the target area.
[0023] Optionally, the smart glasses can automatically locate target areas in the classroom by extracting features from the collected video information, comparing these features with the features of a pre-set target area, and calculating the similarity. If the similarity reaches a certain level, the area can be identified as the target area. Alternatively, the smart glasses can locate the target area based on the user's gaze pattern; if the user gazes for a certain duration, the gazed area can be identified as the target area.
[0024] Optionally, embodiments of this application can extract low-level features such as edges, textures, and colors from video frames using computer vision algorithms, and compare them frame-by-frame with pre-stored target area templates (such as the geometric contours or color distribution of a blackboard or projection screen). When the similarity of local area features exceeds a threshold, it is determined to be a target area. Alternatively, a deep learning model can be used to perform semantic segmentation on the video to identify semantic objects such as "blackboard" and "projection screen." High-probability target areas can be identified through scene understanding (such as prior knowledge of classroom layout).
[0025] Optionally, embodiments of this application can use an eye tracker built into the glasses to capture the pupil position and gaze direction in real time. When the duration of continuous fixation on a certain area exceeds a preset threshold, it is determined as the target area after calibration with head posture data. Alternatively, a heatmap can be generated by collecting fixation data from multiple users. When the proportion of a certain area in the group fixation heatmap exceeds a threshold, the target area is determined by combining the individual fixation duration.
[0026] S102, acquire the behavioral information of the user wearing the wearable device, and extract features from the behavioral information to obtain behavioral features.
[0027] In this embodiment, eye-tracking modules can be used to collect eye-tracking information of wearable device users as user behavior information.
[0028] Specifically, the eye-tracking module can be an infrared eye tracker, capable of high-speed capturing the displacement of the pupil's reflection point. Combining this with the corneal reflection principle, it calculates the direction of the gaze and generates a real-time gaze point coordinate sequence. After noise removal using Kalman filtering, the trajectory is segmented to obtain the user's behavioral characteristics.
[0029] For example, if the duration of sustained focus on a certain area exceeds a preset time, it is considered valid focus. If the number of gaze shifts per unit time exceeds a preset number, it is considered distracted. If the percentage of focus on the blackboard area exceeds a preset percentage and the scanning frequency is less than a preset frequency, it can be considered focused listening. If a single gaze on a non-blackboard area exceeds a set duration, it can be considered a state of distraction.
[0030] S103. Based on the correlation between audio / video features and behavioral features, determine the teaching scenario corresponding to the audio / video information.
[0031] The embodiments of this application can construct an association model to calculate the association degree, and then determine the corresponding teaching scenario based on the association degree.
[0032] The embodiments of this application can determine the teaching scenario through a multidimensional feature weighted association model.
[0033] Specifically, audio and video features can include both audio and video features. Audio features can include the audio energy of the target area. Video features can include text density in the target area, teacher movement trajectory, etc. User behavioral features can include the degree of overlap between the gaze point and the blackboard, the user's head orientation, the user's blinking characteristics, etc. Audio and video features and behavioral features can be normalized to the same vector dimension, and then the correlation between them can be calculated using a weighted Euclidean distance formula.
[0034] When the correlation is greater than a certain preset threshold, and the teacher stays in the target area for a longer period of time than a certain preset threshold, the current teaching scenario can be determined as "focused listening to the blackboard explanation".
[0035] When the correlation is greater than another preset threshold, and the teacher's movement trajectory covers the area under the podium, and the student's gaze follows the teacher's movement at a speed greater than a certain preset speed, the current teaching scenario can be determined as "focused listening and roving explanation".
[0036] For example, when the correlation is greater than 0.4 and less than or equal to 0.5, and the teacher spends more than 80% of the time in a fixed area on the podium, it is judged as "focused listening to the blackboard explanation"; If the correlation is greater than 0.5 but the teacher's movement trajectory covers the area under the podium, and the student's gaze follows the teacher's movement at a speed greater than 10 pixels per second, it is judged as "focused listening and circulating explanation".
[0037] S104: Process the audio and video information according to the classroom minutes extraction strategy corresponding to the teaching scenario, and generate classroom minutes corresponding to the audio and video information.
[0038] In this embodiment, the smart glasses pre-store classroom minutes extraction strategies for different teaching scenarios, and different classroom minutes extraction strategies extract audio and video information in different ways.
[0039] In scenarios where students are focused on listening to a lecture from the blackboard, each frame of the video can be processed to generate corresponding audio. Then, the audio is transcribed into text to create class notes. In scenarios where students are focused on listening to a roving lecture, less important video frames can be discarded, retaining only the essential ones. The audio is then matched to the processed video. Finally, the audio is transcribed into text to generate class notes.
[0040] In scenarios where students are focused on listening to lectures delivered on the blackboard, the content written on the board and the audio of the lecture are highly synchronized when the teacher is explaining in a fixed area (e.g., the steps of formula derivation correspond to the audio explanation). Frame-by-frame video processing can ensure the integrity of the OCR recognition of the written text on the board (avoiding missing frames that could lead to missing formulas). Combined with full audio-to-text conversion, it can generate coherent minutes, which meets the teaching requirement that "static content must be recorded completely".
[0041] In scenarios involving focused listening and roving explanations, the image is prone to shaking when the teacher moves, and non-keyframes (such as transitional walking frames) contain less teaching information. Discarding these frames can reduce the amount of video data by more than 50% (experimental data shows that redundant frames account for about 60%). At the same time, by matching audio with keyframes (such as the image when the teacher stops to explain), we can ensure that key points are recorded (such as the audio corresponding to the steps in a close-up of an experimental operation), thereby improving storage efficiency (saving about 30% of storage space) while ensuring the integrity of information.
[0042] For example, when the scene is identified as a "blackboard explanation" scenario, full-frame processing mode is activated: First, the video stream is preprocessed frame by frame using an integrated neural network accelerator (such as an NPU) to enhance the whiteboard area through edge detection, while audio is recorded at a sampling rate of 44.1kHz to meet the clarity requirements of speech-to-text conversion.
[0043] Subsequently, using timestamp synchronization technology, an index mapping is established between the OCR text results (such as whiteboard content) of each video frame and the corresponding audio segment (such as "Next, let's look at this formula") to generate structured intermediate data ({timestamp: [text block, audio segment]}).
[0044] Finally, the intermediate data is reorganized according to the "whiteboard writing sequence + voice logic" through a lightweight natural language processing module: repeated explanations are automatically filtered out (such as only the first repetition of the same formula), text blocks are arranged in row and column order, audio-to-text results are inserted into the corresponding positions in segments, and finally a Markdown format summary is generated (titles, formula blocks, and voice segments are automatically layered).
[0045] For example, when the scenario is determined to be a "circuitous explanation" scenario, moving target detection is initiated: First, the teacher's human body outline is set as the region of interest. When the teacher's movement speed is greater than 5 pixels / second, the current frame is marked as a "transition frame" and discarded. When a stationary action is detected (speed < 1 pixel / second and lasts for 2 seconds), the frame is retained and a close-up of the local area is extracted (such as the operation screen of experimental equipment).
[0046] Then, the built-in Voice Activity Detection (VAD) module is used to segment the audio stream, and keyword recognition (such as "pay attention here" and "next step") is used to locate key segments for explanation. Environmental noise is filtered through short-time energy analysis (50ms window) to retain effective speech segments with a signal-to-noise ratio >20dB.
[0047] Finally, the Dynamic Time Warping (DTW) algorithm is used to align the retained keyframe timestamps with the audio semantic blocks (allowing a deviation of ±500ms), and the matched audio and video pairs are compressed using H.265 encoding (compression ratio of 10:1). At the same time, a lightweight summary is generated: keyframe thumbnails are embedded in text fields, and the audio-to-text result retains only sentences containing keywords.
[0048] This application's embodiments determine classroom minutes extraction strategies for different teaching scenarios by assessing the correlation between audio-visual features of designated areas in the classroom and user behavior characteristics. Audio-visual information is processed according to different extraction strategies to generate personalized classroom minutes that meet the actual needs of the client, thus improving the reliability of the generated minutes.
[0049] In one embodiment of this application, the teaching scenario includes a first teaching scenario, a second teaching scenario, or a third teaching scenario; the audio and video information includes audio information and video information.
[0050] The audio and video information is processed according to the classroom summary extraction strategy corresponding to the teaching scenario to generate classroom summaries corresponding to the audio and video information, including: Responding to the teaching scenario as the primary teaching scenario, an audio classroom summary corresponding to the audio information is generated. The audio classroom summary is then matched with the video information to serve as the classroom summary corresponding to the audio and video information.
[0051] Responding to the teaching scenario as the second teaching scenario, an audio classroom summary corresponding to the audio information is generated; multiple key frames are extracted from the video information; and the audio classroom summary is matched with multiple key frames to serve as the classroom summary corresponding to the audio and video information.
[0052] Responding to the teaching scenario as the third teaching scenario, the system extracts key audio information corresponding to the audio information and generates key audio class notes corresponding to the key audio information; it also extracts multiple key frames from the video information; and matches the key audio class notes with multiple key frames to obtain the class notes corresponding to the audio and video information.
[0053] In this embodiment, the first teaching scenario is a crucial scenario. In this scenario, the teacher explains core knowledge points, formula derivations, or experimental principles in the target area, requiring detailed recording. This includes frequently updated blackboard writing and audio containing complex technical terms. At this point, a complete audio class summary (full audio-to-text conversion) needs to be generated and matched frame-by-frame with the video information to ensure complete synchronization between the blackboard text, teacher gestures, and other visual information and the audio explanation (timestamp alignment accuracy ±50ms), meeting the requirement of recording "no omissions of important content."
[0054] The second teaching scenario is of secondary importance. In this scenario, the blackboard writing remains unchanged for an extended period (e.g., displaying a fixed chart), and the teacher primarily circulates and explains. The blackboard area in the video footage shows no substantial updates (the OCR results are identical for 10 consecutive frames). In this case, redundant video frames need to be discarded (only the initial blackboard writing and frames marked by the teacher's key points are retained), reducing the video data volume by more than 60%. Simultaneously, the explanation content is recorded using audio-to-text conversion. Matching key frames with audio can be achieved through a "voice keyword - screen marker" association (e.g., automatically marking the corresponding frame when the teacher says "Please look at this diagram").
[0055] The third teaching scenario is a less important scenario. In this scenario, the teacher is silent waiting for students to answer (audio energy <40dB and lasts for more than 10 seconds), during breaks, or engaged in non-teaching activities (such as distributing homework, with no updates to the blackboard in the video and the teacher's actions unrelated to teaching). In this case, only key audio (such as the audio segment when the teacher asks "Who will answer?") and keyframes (such as the moment the teacher points to the blackboard) can be extracted, filtering out invalid information (such as audio during silent periods). This allows the minutes to focus on the core teaching elements, reducing the file size by 80% compared to the full recording.
[0056] Specifically, in the first teaching scenario, this embodiment of the application can record audio using an 8kHz sampling rate + 16-bit quantization, segment effective speech segments using a VAD (Voice Activity Detection) module, and perform cloud-to-text conversion to identify effective text. Subsequently, professional terms can be automatically annotated (e.g., "Lorentz force" can be added in italics) to generate a text stream with timestamps (e.g., "05:12 Next, the Lorentz force formula is derived..."). Finally, each frame of video can be used for whiteboard OCR recognition to align text blocks with the audio text stream: when "formula (1)" appears in the audio, the corresponding formula area in the video at that moment is automatically extracted to generate a mixed text and image summary (video screenshots are embedded on the right side of the text paragraphs).
[0057] In the second teaching scenario, the embodiments of this application can calculate the difference between consecutive frames (based on the histogram intersection algorithm), and determine that a frame is a duplicate when the difference is less than 5%. Subsequently, the keyframes can be arranged in chronological order, and an audio segment index can be added to each frame (e.g., "keyframe 3 corresponds to the audio segment from 08:25 to 08:40") to generate a Markdown format summary.
[0058] In the third teaching scenario, this application embodiment can establish a non-teaching state detection model: when the audio energy is <40dB and lasts for more than 15 seconds, or when keywords such as "take a break" are detected, it is marked as a non-teaching period. Only audio segments containing teaching intent can be retained: key audio is located through keyword matching (such as "question" and "answer") and speech fundamental frequency mutation (fundamental frequency rises >8Hz when asking a question), extracted, converted to text, and labeled with tags such as "[Teacher asks a question]". Subsequently, an action recognition model can be used to detect effective teaching actions of the teacher, and then the key audio and keyframes are semantically associated to generate corresponding classroom minutes.
[0059] This application's embodiments employ corresponding summary extraction strategies for different teaching scenarios, improving summary generation efficiency and practicality, optimizing smart glasses storage and power consumption, and better meeting user needs. In the first teaching scenario, full recording ensures no important content is missed, and frame-by-frame matching synchronizes audio and video, facilitating the review of complex knowledge. In the second teaching scenario, redundant frames are discarded, reducing data volume, and key frames can be marked with voice keywords, saving storage space while highlighting key points. In the third teaching scenario, key audio and video are extracted, filtering out invalid information, allowing the summary to focus on the core content, and significantly reducing file size.
[0060] In one embodiment of this application, the teaching scenario corresponding to the audio and video information is determined based on the correlation between audio / video features and behavioral features, including: Calculate the similarity between the audio / video feature and each standard audio / video feature in the standard audio / video feature library, and select the standard audio / video feature with the highest similarity as the target audio / video feature; the teacher behaviors corresponding to each standard audio / video feature are different; Calculate the similarity between the behavioral feature and each standard behavioral feature in the standard behavioral feature library, and select the standard behavioral feature with the highest similarity as the target behavioral feature; each standard behavioral feature corresponds to a different user behavior; Calculate the correlation between the target audio and video features and the target behavioral features, and determine the teaching scenario corresponding to the audio and video information based on the correlation.
[0061] In this embodiment, the standard audio-visual feature library can store multiple sets of audio-visual feature data for typical teaching scenarios, i.e., multiple standard audio-visual features. Different standard audio-visual features or their combinations can correspond to specific teacher behaviors. For example, this can include blackboard explanation and roving explanation. The standard behavior feature library can store multiple sets of typical student classroom behavior feature data, i.e., multiple standard behavior features. Different standard behavior features or their combinations can correspond to specific learning behaviors. For example, this can include attentive listening and distraction.
[0062] This application embodiment calculates the similarity between the extracted audio and video features and each other, and selects the standard audio and video feature with the highest similarity as the target audio and video feature to determine the teacher scene corresponding to the target audio and video feature. The teacher scene may include: lecturing, interaction, questioning, silence, etc.
[0063] This application's embodiments calculate the similarity between behavioral features and various standard behavioral features, and select the standard behavioral feature with the highest similarity as the target behavioral feature to determine the student scenario corresponding to the target behavioral feature. Student scenarios may include listening, eye contact, answering questions, and distraction.
[0064] After obtaining the teacher scenario and the student scenario, the correlation between the target audio and video features and the target behavioral features can be calculated, that is, the correlation between the teacher scenario and the student scenario can be calculated, and the corresponding teaching scenario can be determined based on the correlation.
[0065] This application's embodiments achieve accurate determination of teaching scenarios through dual-feature library matching and correlation calculation. First, audio-visual features are matched with teacher behavior, and behavioral features are matched with student status. Then, the correlation between the two is comprehensively judged, taking into account both teacher teaching behavior and student learning status. On the one hand, similarity matching based on a standard feature library can improve the accuracy of scenario recognition by utilizing historical typical scenario data. On the other hand, through correlation analysis between teacher and student scenarios, the actual classroom interaction can be more comprehensively reflected, ensuring that the determined teaching scenario considers both teacher behavior and student status. This provides a more accurate basis for subsequent classroom minutes extraction strategies, improving the relevance and practicality of minutes generation.
[0066] In one embodiment of this application, the correlation degree between target audio / video features and target behavioral features is calculated, and the teaching scenario corresponding to the audio / video information is determined based on the correlation degree, including: Calculate the correlation between the target audio / video features and the target behavioral features; If the correlation degree is greater than or equal to the first preset correlation degree, then the teaching scenario corresponding to the audio and video information is determined to be the first teaching scenario; If the correlation is less than the first preset correlation, but greater than or equal to the second preset correlation, then the teaching scenario corresponding to the audio and video information is determined to be the second teaching scenario. If the correlation is less than the second preset correlation, then the teaching scenario corresponding to the audio and video information is determined to be the third teaching scenario; Among them, the first preset correlation degree is greater than the second preset correlation degree.
[0067] This application embodiment can convert target audio / video features and target behavioral features into normalized vectors. For example, target audio / video features include whiteboard update frequency f1 and audio term density f2. Target behavioral features may include gaze overlap rate s1 and head nodding frequency s2.
[0068] F=[f1,f2] S=[s1,s2] For example, in a blackboard explanation scenario, F=[0.8,0.7] indicates a high frequency of writing on the blackboard and a concentrated density of terminology. S=[0.9,0.6] indicates a high rate of eye contact overlap and a moderate frequency of nodding.
[0069] Then, in this embodiment of the application, a feature weight matrix can be constructed, and the degree of association can be determined by cosine similarity calculation.
[0070] Feature weight matrix W=[w1,w2,w3,w4] Merged vector V = [f1, f2, s1, s2] Correlation formula:
[0071] in, Let C be the standard feature vector, and its value range is [0,1]. Let wi represent the data in the i-th feature weight matrix.
[0072] In this embodiment of the application, the correlation distribution of each scenario can be statistically analyzed based on historical data: for example, the mean correlation of the first scenario is 0.82 and the standard deviation is 0.1. The first preset value is set to mean - 0.12 = 0.7 and the second preset value is set to mean - 0.42 = 0.4 to ensure that the threshold is adapted to different classroom scenarios. The specific threshold can be adjusted according to the actual situation.
[0073] In this embodiment of the application, a first teaching scenario is defined as a correlation degree ≥ a first preset value (e.g., 0.7). This indicates that the teacher is explaining core knowledge (e.g., formula derivation) and the students are highly focused, requiring full recording of audio and video.
[0074] If the second preset value ≤ relevance < the first preset value (e.g., 0.4-0.7), it is determined to be the second teaching scenario. This indicates that the teacher is explaining important content but the students' attention is moderate (e.g., static blackboard writing and circulating explanation), and only keyframes and audio are retained.
[0075] If the correlation score is less than the second preset value (e.g., less than 0.4), it is judged as the third teaching scenario. This indicates that the teacher is in a non-teaching state (e.g., waiting silently) or the students are distracted, and only key information is extracted and recorded.
[0076] The embodiments of this application can record all core knowledge in highly relevant scenarios to ensure no omissions, discard 60% of redundant video frames in medium-relevance scenarios but retain key frames related to keywords, and record only 20% of the core content in low-relevance scenarios, so that the size of the class minutes file can be dynamically adjusted according to the importance of the scenario (maximum compression of 80%).
[0077] In one embodiment of this application, after generating class minutes corresponding to the audio and video information, the method further includes: Get the device battery level of the wearable device; If the device's battery level is greater than or equal to the first preset battery level, the class notes will be adjusted based on the user's historical learning characteristics, and the adjusted class notes will be stored.
[0078] If the device's battery level is lower than the first preset battery level, the class minutes will not be adjusted.
[0079] In this embodiment, the wearable device may have a pre-installed user learning preference model stored in local flash memory. Learning preferences may include content preferences and format preferences. For example, content preferences may include marking frequently used knowledge points in historical minutes and highlighting areas of the blackboard that the user focuses on. Format preferences may include the user's editing habits for class minutes.
[0080] In this embodiment, if the wearable device's battery level is greater than or equal to a first preset battery level, it indicates that the current battery level is high, and the user learning preference model can be invoked to optimize the currently generated class minutes without waiting for subsequent optimization. When the wearable device's battery level is low, the user learning preference model can be omitted, thus improving battery life.
[0081] In one embodiment of this application, obtaining audio and video information of a target area in a classroom includes: In response to the target area being a blackboard area, the audio and video acquisition module in the wearable device is controlled to acquire audio and video information of the blackboard area in the first working mode; In response to the target area being a non-blackboard area, the audio and video acquisition module in the wearable device is controlled to acquire audio and video information of the non-blackboard area in the second working mode; In the first working mode, the audio and video acquisition resolution is greater than that in the second working mode.
[0082] In classroom teaching, blackboards are typically used to write core knowledge points, formula derivations, and other key content. High-resolution recording is required to ensure the accuracy and completeness of text, charts, and other information, facilitating the accurate generation of subsequent class minutes. Non-blackboard areas, such as areas where the teacher moves or where students sit, often contain supplementary information or non-teaching scenarios, and therefore have relatively lower requirements for image resolution.
[0083] The smart glasses' built-in audio and video acquisition module is adjustable. When the target area is determined to be a blackboard area through techniques such as region feature matching or eye tracking (e.g., the aforementioned method based on region feature matching combined with user gaze patterns), the acquisition module switches to the first working mode, acquiring audio and video at a higher resolution (e.g., 1920×1080 pixels) to ensure that the details of the written content are clearly discernible. When the area is determined to be a non-blackboard area, it automatically switches to the second working mode, acquiring data at a lower resolution (e.g., 1280×720 pixels or lower), reducing the amount of data while meeting basic information recording requirements.
[0084] This application embodiment can effectively reduce the amount of data collected by lowering the acquisition resolution in non-blackboard areas, thereby reducing the workload of the processor and storage module, reducing power consumption, and thus extending the battery life of smart glasses and ensuring the stable operation of the device throughout the class.
[0085] Corresponding to the classroom minutes generation method in the above embodiment, Figure 2 This is a structural block diagram of a classroom minutes generation system provided in one embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 2 The classroom minutes generation system 20 is applied to wearable devices and includes: a first acquisition module 201, a second acquisition module 202, a first control module 203, and a second control module 204. The first acquisition module 201 is used to acquire audio and video information of the target area in the classroom, and to extract features from the audio and video information to obtain audio and video features; The second acquisition module 202 is used to acquire the behavioral information of users wearing wearable devices and extract features from the behavioral information to obtain behavioral features. The first control module 203 is used to determine the teaching scenario corresponding to the audio and video information based on the correlation between audio and video features and behavioral features; The second control module 204 is used to process audio and video information according to the classroom minutes extraction strategy corresponding to the teaching scenario, and generate classroom minutes corresponding to the audio and video information.
[0086] In one embodiment of this application, the teaching scenario includes a first teaching scenario, a second teaching scenario, or a third teaching scenario; the audio and video information includes audio information and video information; The second control module 204 is specifically used to respond to the teaching scenario as the first teaching scenario, generate audio classroom minutes corresponding to the audio information, and match the audio classroom minutes with the video information to serve as the classroom minutes corresponding to the audio and video information; Responding to the teaching scenario as the second teaching scenario, an audio classroom summary corresponding to the audio information is generated; multiple key frames are extracted from the video information; the audio classroom summary is matched with multiple key frames to serve as the classroom summary corresponding to the audio and video information; Responding to the teaching scenario as the third teaching scenario, the system extracts key audio information corresponding to the audio information and generates key audio class notes corresponding to the key audio information; it also extracts multiple key frames from the video information; and matches the key audio class notes with multiple key frames to obtain the class notes corresponding to the audio and video information.
[0087] In one embodiment of this application, the first control module 203 is specifically used to calculate the similarity between the audio and video features and each standard audio and video feature in the standard audio and video feature library, and select the standard audio and video feature with the highest similarity as the target audio and video feature; the teacher behaviors corresponding to each standard audio and video feature are different; Calculate the similarity between the behavioral feature and each standard behavioral feature in the standard behavioral feature library, and select the standard behavioral feature with the highest similarity as the target behavioral feature; each standard behavioral feature corresponds to a different user behavior; Calculate the correlation between the target audio and video features and the target behavioral features, and determine the teaching scenario corresponding to the audio and video information based on the correlation.
[0088] In one embodiment of this application, the first control module 203 is specifically used to calculate the correlation between target audio and video features and target behavior features; If the correlation degree is greater than or equal to the first preset correlation degree, then the teaching scenario corresponding to the audio and video information is determined to be the first teaching scenario; If the correlation is less than the first preset correlation, but greater than or equal to the second preset correlation, then the teaching scenario corresponding to the audio and video information is determined to be the second teaching scenario. If the correlation is less than the second preset correlation, then the teaching scenario corresponding to the audio and video information is determined to be the third teaching scenario; Among them, the first preset correlation degree is greater than the second preset correlation degree.
[0089] In one embodiment of this application, the system may further include: The third control module is used to obtain the device battery level of the wearable device after generating the classroom minutes corresponding to the audio and video information; If the device's battery level is greater than or equal to the first preset battery level, the class notes will be adjusted based on the user's historical learning characteristics, and the adjusted class notes will be stored.
[0090] In one embodiment of this application, the third control module is further configured to not adjust the class minutes if the device battery level is less than a first preset battery level.
[0091] In one embodiment of this application, the first acquisition module 201 is specifically used to control the audio and video acquisition module in the wearable device to acquire audio and video information of the blackboard area in a first working mode in response to the target area being a blackboard area; In response to the target area being a non-blackboard area, the audio and video acquisition module in the wearable device is controlled to acquire audio and video information of the non-blackboard area in the second working mode; In the first working mode, the audio and video acquisition resolution is greater than that in the second working mode.
[0092] See Figure 3 , Figure 3 This is a schematic block diagram of a wearable device provided in one embodiment of this application. Figure 3 The wearable device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the aforementioned device embodiments, for example... Figure 2 The functions of the first acquisition module 201, the second acquisition module 202, the first control module 203, and the second control module 204 are shown.
[0093] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0094] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.
[0095] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory.
[0096] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation method described in the classroom minutes generation method provided in the embodiments of this application, or they can execute the implementation method of the wearable device described in the embodiments of this application, which will not be elaborated here.
[0097] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0098] The computer-readable storage medium can be an internal storage unit of the wearable device in any of the foregoing embodiments, such as a hard drive or memory of the wearable device. The computer-readable storage medium can also be an external storage device of the wearable device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the wearable device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the wearable device. The computer-readable storage medium is used to store computer programs and other programs and data required by the wearable device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0099] This application provides a computer program product, which includes computer-executable instructions or a computer program. The computer-executable instructions or computer program are stored in a computer-readable storage medium. The processor of a wearable device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the wearable device to perform the classroom minutes generation method described above in this application.
[0100] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0101] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the wearable device and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0102] In the several embodiments provided in this application, it should be understood that the disclosed wearable devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and other division methods may exist in actual implementation. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces or units, or they may be electrical, mechanical, or other forms of connection.
[0103] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0104] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for generating classroom minutes, characterized in that, Applications in wearable devices, including: Acquire audio and video information of a target area in the classroom, and extract features from the audio and video information to obtain audio and video features. The target area is the area where the teacher is located in the classroom. The behavior information of the user wearing the wearable device is obtained, and the behavior information is used to extract features to obtain behavior features; The teaching scenario corresponding to the audio and video information is determined based on the correlation between the audio and video features and the behavioral features; The audio and video information is processed according to the classroom minutes extraction strategy corresponding to the teaching scenario to generate classroom minutes corresponding to the audio and video information; The teaching scenarios include a first teaching scenario, a second teaching scenario, and a third teaching scenario; the audio and video information includes audio information and video information. The step of processing the audio and video information according to the classroom minutes extraction strategy corresponding to the teaching scenario to generate classroom minutes corresponding to the audio and video information includes: In response to the teaching scenario being the first teaching scenario, an audio class summary corresponding to the audio information is generated, and the audio class summary is matched with the video information to serve as the class summary corresponding to the audio and video information; In response to the teaching scenario being a second teaching scenario, an audio class summary corresponding to the audio information is generated; multiple key frames are extracted from the video information; and the audio class summary is matched with the multiple key frames to obtain the class summary corresponding to the audio and video information. In response to the teaching scenario being a third teaching scenario, key audio information corresponding to the audio information is extracted, and key audio class notes corresponding to the key audio information are generated; multiple key frames are extracted from the video information; and the key audio class notes are matched with the multiple key frames to serve as the class notes corresponding to the audio and video information.
2. The classroom minutes generation method as described in claim 1, characterized in that, The step of determining the teaching scenario corresponding to the audio-visual information based on the correlation between the audio-visual features and the behavioral features includes: Calculate the similarity between the audio / video feature and each standard audio / video feature in the standard audio / video feature library, and select the standard audio / video feature with the highest similarity as the target audio / video feature; the teacher behaviors corresponding to each standard audio / video feature are different; Calculate the similarity between the stated behavioral feature and each standard behavioral feature in the standard behavioral feature library, and select the standard behavioral feature with the highest similarity as the target behavioral feature; the user behaviors corresponding to each standard behavioral feature are different; Calculate the correlation degree between the target audio and video features and the target behavioral features, and determine the teaching scenario corresponding to the audio and video information based on the correlation degree.
3. The classroom minutes generation method as described in claim 2, characterized in that, The step of calculating the correlation degree between the target audio / video features and the target behavioral features, and determining the teaching scenario corresponding to the audio / video information based on the correlation degree, includes: Calculate the correlation degree between the target audio / video features and the target behavior features; If the correlation degree is greater than or equal to the first preset correlation degree, then the teaching scenario corresponding to the audio and video information is determined to be the first teaching scenario; If the correlation degree is less than the first preset correlation degree and greater than or equal to the second preset correlation degree, then the teaching scenario corresponding to the audio and video information is determined to be the second teaching scenario. If the correlation degree is less than the second preset correlation degree, then the teaching scenario corresponding to the audio and video information is determined to be the third teaching scenario; Wherein, the first preset correlation degree is greater than the second preset correlation degree.
4. The method for generating classroom minutes as described in claim 1, characterized in that, After generating the class minutes corresponding to the audio and video information, the method further includes: Obtain the device power level of the wearable device; If the device's battery level is greater than or equal to a first preset battery level, the class minutes are adjusted according to the user's historical learning characteristics, and the adjusted class minutes are stored.
5. The classroom minutes generation method as described in claim 4, characterized in that, Also includes: If the device's battery level is lower than the first preset battery level, the class minutes will not be adjusted.
6. The method for generating classroom minutes as described in claim 1, characterized in that, The acquisition of audio and video information of the target area in the classroom includes: In response to the target area being a blackboard area, the audio and video acquisition module in the wearable device is controlled to acquire audio and video information of the blackboard area in a first working mode; In response to the target area being a non-blackboard area, the audio and video acquisition module in the wearable device is controlled to acquire audio and video information of the non-blackboard area in a second working mode; The audio and video acquisition resolution in the first working mode is greater than that in the second working mode.
7. A classroom minutes generation system, characterized in that, Applications in wearable devices, including: The first acquisition module is used to acquire audio and video information of a target area in the classroom, and to extract features from the audio and video information to obtain audio and video features, wherein the target area is the area where the teacher is located in the classroom; the second acquisition module is used to acquire behavioral information of the user wearing the wearable device, and to extract features from the behavioral information to obtain behavioral features; The first control module is used to determine the teaching scenario corresponding to the audio and video information based on the correlation between the audio and video features and the behavioral features; the teaching scenario includes a first teaching scenario, a second teaching scenario, and a third teaching scenario; the audio and video information includes audio information and video information. The second control module is used to process the audio and video information according to the classroom minutes extraction strategy corresponding to the teaching scenario, and generate classroom minutes corresponding to the audio and video information; Specifically, the second control module is used to: in response to the teaching scenario being the first teaching scenario, generate an audio class summary corresponding to the audio information, and match the audio class summary with the video information to serve as the class summary corresponding to the audio and video information; In response to the teaching scenario being a second teaching scenario, an audio class summary corresponding to the audio information is generated; multiple key frames are extracted from the video information; and the audio class summary is matched with the multiple key frames to obtain the class summary corresponding to the audio and video information. In response to the teaching scenario being a third teaching scenario, key audio information corresponding to the audio information is extracted, and key audio class notes corresponding to the key audio information are generated; multiple key frames are extracted from the video information; and the key audio class notes are matched with the multiple key frames to serve as the class notes corresponding to the audio and video information.
8. A wearable device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-modal classroom summary automatic generation method and system based on improved PEGASUS model
CN118035474A
Conference multimedia summary support system and method
US5572728A