Classroom summary generation method and system, wearable device and storage medium

Smart glasses are used to obtain audio, video and user behavior characteristics, determine teaching scenarios based on correlation, and generate personalized classroom minutes. This solves the problem that the minutes generated in existing technologies do not meet students' needs, and achieves more efficient and reliable classroom minute generation.

CN120640103AActive Publication Date: 2025-09-12BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 19 Cites 0 Cited by

Patent Information

Application Number
CN202510952574.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-09-12
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

The classroom minutes generated by existing smart glasses during the classroom recording process use fixed structured templates, which are difficult to adapt to students' usage scenarios, resulting in complicated operations and insufficient personalization of the generated minutes.

Method used

By obtaining the audio and video information and user behavior information of the target area in the classroom, features are extracted, and the teaching scene is determined based on the correlation between the audio and video features and the behavior features. The corresponding classroom minutes extraction strategy is used to generate personalized classroom minutes.

Benefits of technology

The reliability of classroom minutes generation has been improved, the generated minutes meet the actual needs of users, and the ease of operation and the degree of personalization of generation have been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640103A_ABST
    Figure CN120640103A_ABST
Patent Text Reader

Abstract

The invention provides a classroom summary generation method and system, wearable equipment and a storage medium, and belongs to the field of intelligent equipment control, and the method comprises the steps: obtaining audio and video information of a target area in a classroom, and carrying out the feature extraction of the audio and video information, and obtaining audio and video features; acquiring behavior information of a user wearing the wearable device, and performing feature extraction on the behavior information to obtain behavior features; determining a teaching scene corresponding to the audio and video information according to the association degree between the audio and video features and the behavior features; and processing the audio and video information according to a classroom summary extraction strategy corresponding to the teaching scene to generate a classroom summary corresponding to the audio and video information. According to the classroom summary generation method and system, the wearable device and the storage medium provided by the invention, the classroom summary generation accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of intelligent device control technology, and more specifically, relates to a classroom minutes generation method and system, a wearable device, and a storage medium. Background Art

[0002] With the advancement of technology, the application scenarios of wearable devices such as smart glasses are becoming increasingly diverse. This is particularly true in classroom settings, where smart glasses can record teachers' lectures for easy post-class review. Currently, most smart glasses use fixed, structured templates to generate classroom minutes. Manually adjusting the format requires complex operations, making them difficult to adapt to student usage scenarios. Summary of the Invention

[0003] The purpose of this application is to provide a classroom minutes generation method and system, a wearable device, and a storage medium to improve the reliability of classroom minutes generation.

[0004] According to a first aspect of an embodiment of the present application, a method for generating classroom minutes is provided, which is applied to a wearable device and includes: Obtain audio and video information of the target area in the classroom, and perform feature extraction on the audio and video information to obtain audio and video features; Obtaining behavioral information of users wearing wearable devices, and extracting features from the behavioral information to obtain behavioral features; Determine the teaching scenario corresponding to the audio and video information based on the correlation between the audio and video features and the behavioral features; The audio and video information is processed according to the classroom minutes extraction strategy corresponding to the teaching scenario to generate classroom minutes corresponding to the audio and video information.

[0005] A second aspect of an embodiment of the present application provides a classroom minutes generation system, which is applied to a wearable device and includes: The first acquisition module is used to obtain audio and video information of the target area in the classroom and perform feature extraction on the audio and video information to obtain audio and video features; The second acquisition module is used to obtain behavior information of the user wearing the wearable device and perform feature extraction on the behavior information to obtain behavior features; A first control module is used to determine the teaching scene corresponding to the audio and video information based on the correlation between the audio and video features and the behavioral features; The second control module is used to process the audio and video information according to the classroom minutes extraction strategy corresponding to the teaching scene, and generate classroom minutes corresponding to the audio and video information.

[0006] According to a third aspect of an embodiment of the present application, a wearable device is provided, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the above-mentioned method for generating classroom minutes when executing the computer program.

[0007] In a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned classroom minutes generation method are implemented.

[0008] In a fifth aspect of an embodiment of the present application, a computer program product is provided, comprising a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the steps of the above-mentioned classroom minutes generation method are implemented.

[0009] The beneficial effects of the classroom minutes generation method and system, wearable device, and storage medium provided in the embodiments of the present application are as follows: the embodiments of the present application determine classroom minutes extraction strategies for different teaching scenarios based on the correlation between the audio and video characteristics of the audio and video information in a specified area of ​​the classroom and the user's behavioral characteristics. The audio and video information is processed according to different classroom minutes extraction strategies to generate personalized classroom minutes that meet the actual needs of the customer, thereby improving the reliability of the generated classroom minutes. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A flowchart of a method for generating classroom minutes provided in one embodiment of the present application; Figure 2 A structural block diagram of a classroom minutes generation system provided in one embodiment of the present application; Figure 3 A schematic block diagram of a wearable device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0012] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0013] In order to make the purpose, technical solutions and advantages of this application clearer, specific embodiments will be described below with reference to the accompanying drawings.

[0014] In the embodiment of the present application, the wearable device may be smart glasses or other head-mounted smart devices. In the embodiment of the present application, the wearable device is smart glasses.

[0015] The smart glasses in the embodiment of the present application are provided with an audio and video acquisition module, which includes an audio acquisition unit and a video acquisition unit. The audio acquisition unit can acquire the ambient audio around the smart glasses, and the video acquisition unit can acquire the ambient video around the smart glasses.

[0016] Specifically, the smart glasses integrate an eye tracking module, a video tracking module, and an audio module. The eye tracking module can track the user's eye movements, collect eye information, and analyze it to determine the user's learning state. The video tracking module is equipped with a camera that can track and record the direction the user is looking or the area they are facing. The audio module can collect and filter surrounding sounds.

[0017] In addition, the smart glasses can also connect to the user's electronic device via Bluetooth, which can be connected to a server, which can process and analyze the received information. When online, the server can perform a detailed analysis of the audio and video information collected by the smart glasses to generate classroom minutes that are tailored to the actual situation. In addition, whether online or offline, the intelligent module in the smart glasses can perform adaptive analysis of the collected audio and video information to generate classroom minutes that are tailored to the user's habits. Detailed explanation is given below.

[0018] Please refer to Figure 1 , Figure 1 A flow chart of a method for generating classroom minutes provided in an embodiment of the present application can be executed by a wearable device. The method may include S101 to S104.

[0019] S101, obtaining audio and video information of a target area in a classroom, and performing feature extraction on the audio and video information to obtain audio and video features.

[0020] In the embodiment of the present application, the audio and video information includes audio information and video information. The audio information of the target area in the classroom can be collected by the audio collection unit in the smart glasses, and the video information of the target area in the classroom can be collected by the video collection unit in the smart glasses.

[0021] The target area is a designated area in the classroom, such as the teacher's podium area, which may include the blackboard. The video acquisition unit can capture the teacher's behavior and posture in the target area, as well as the PPT or writing on the blackboard. The audio acquisition unit can capture the teacher's voice in the target area.

[0022] Alternatively, the target area may be the area where the teacher walks under the podium. The video acquisition unit may capture the teacher's behavior, posture, and demeanor in the target area. The audio acquisition unit may capture the teacher's voice in the target area.

[0023] Optionally, the smart glasses can automatically locate the target area in the classroom by extracting features from the collected video information, extracting regional features, comparing these regional features with the regional features of a pre-set target area, and calculating similarity. If the similarity reaches a certain level, the area can be located as the target area. Alternatively, the smart glasses can locate the target area based on the user's gaze. If the user's gaze duration reaches a certain length, the gaze area can be determined as the target area.

[0024] Optionally, embodiments of the present application can use computer vision algorithms to extract underlying features such as edges, textures, and colors from video frames, and perform frame-by-frame comparisons with pre-stored target area templates (such as the geometric outline or color distribution of a blackboard or projection screen). When the similarity of local area features exceeds a threshold, it is determined to be the target area. Alternatively, a deep learning model can be used to perform semantic segmentation on the video to identify semantic objects such as "blackboard" and "projection screen." High-probability target areas can be identified through scene understanding (such as prior knowledge of classroom layout).

[0025] Optionally, embodiments of the present application can use a built-in eye tracker in glasses to capture pupil position and gaze direction in real time. When the duration of continuous gaze on a certain area exceeds a preset threshold, combined with head posture data calibration, it is determined to be a target area. Alternatively, gaze data from multiple users is collected to generate a heat map. When the proportion of a certain area in the group gaze heat map exceeds a threshold, the target area is comprehensively determined based on the individual gaze duration.

[0026] S102: Obtain behavior information of a user wearing a wearable device, and perform feature extraction on the behavior information to obtain behavior features.

[0027] In an embodiment of the present application, the eye movement information of the user of the wearable device can be collected through the eye tracking module as the user's behavior information.

[0028] Specifically, the eye tracking module can be an infrared eye tracker that can quickly capture pupil reflection point displacement, calculate gaze direction based on corneal reflection principles, and generate a real-time gaze point coordinate sequence. After using Kalman filtering to remove noise, the trajectory is segmented to obtain the user's behavioral characteristics.

[0029] For example, if it is determined that the continuous gaze time in a certain area is greater than a certain preset time, it is determined to be effective attention. If it is determined that the number of gaze shifts per unit time is greater than a certain preset number, it is determined to be distracted. If it is determined that the gaze ratio of the blackboard area is greater than a predetermined preset ratio and the scanning frequency is less than a certain preset frequency, it can be determined to be focused listening. If it is determined that a single gaze in a non-blackboard area is greater than a set time, it can be determined to be a distracted state.

[0030] S103: Determine the teaching scenario corresponding to the audio and video information according to the correlation between the audio and video features and the behavior features.

[0031] The embodiment of the present application can calculate the correlation degree by constructing an association model, and then determine the corresponding teaching scenario according to the correlation degree.

[0032] The embodiment of the present application can determine the teaching scenario through a multi-dimensional feature weighted association model.

[0033] Specifically, audio and video features can include audio features and video features. Audio features can include the audio energy of the target area. Video features can include text density in the target area and the teacher's movement trajectory. User behavioral features can include the degree of overlap between the gaze point and the blackboard, the user's head direction, and the user's blinking characteristics. The audio and video features and behavioral features can be normalized to the same vector dimension, and the correlation between them can then be calculated using the weighted Euclidean distance formula.

[0034] When the correlation is greater than a certain preset threshold and the teacher stays in the target area for longer than a certain preset threshold, it can be determined that the current teaching scene is "concentrated listening to the blackboard explanation".

[0035] When the correlation is greater than another preset threshold, and the teacher's movement trajectory covers the area under the podium, and the student's gaze follows the teacher's movement speed faster than a certain preset speed, it can be determined that the current teaching scene is "attentive listening and tour explanation".

[0036] For example, when the correlation degree is greater than 0.4 and less than or equal to 0.5, and the teacher stays in a fixed area on the podium for more than 80% of the time, it is determined as "attentively listening to the blackboard explanation"; If the correlation degree is greater than 0.5 but the teacher's movement trajectory covers the area under the podium, and the student's gaze follows the teacher's movement speed > 10 pixels / second, it is judged as "attentive listening and tour explanation".

[0037] S104: Process the audio and video information according to the classroom minutes extraction strategy corresponding to the teaching scene to generate classroom minutes corresponding to the audio and video information.

[0038] In an embodiment of the present application, the smart glasses pre-store classroom minutes extraction strategies for different teaching scenarios, and different classroom minutes extraction strategies have different ways of extracting audio and video information.

[0039] In scenarios where students are focused on listening to a lecture on the blackboard, each frame of the video can be processed and the corresponding audio generated. This audio can then be converted into text to create the corresponding class notes. In scenarios where students are focused on listening to a lecture on a circuit, unimportant frames can be discarded, retaining only the important ones. The audio can then be matched to the processed video. Finally, the audio can be converted into text to create the corresponding class notes.

[0040] In the scenario of concentrating on listening to the blackboard explanation, when the teacher explains in a fixed area, the blackboard content and the teaching audio are highly synchronized (for example, the formula derivation steps correspond to the voice explanation). Frame-by-frame video processing can ensure the integrity of the OCR recognition of the blackboard text (avoiding missed frames and missing formulas). Combined with the full audio-to-text conversion, a coherent minutes can be generated, which meets the teaching requirement of "static content needs to be fully recorded".

[0041] In scenarios where students are focused on a lecture tour, the video often jitters when the teacher moves, and non-key frames (such as walking transitions) contain less instructional information. Discarding these frames can reduce video data by over 50% (experimental data shows that redundant frames account for approximately 60%). Furthermore, by matching audio with key frames (such as when the teacher pauses to explain), we ensure that key points are captured (e.g., close-ups of experimental procedures correspond to audio recordings of the steps), improving storage efficiency while ensuring information integrity (saving approximately 30% of storage space).

[0042] For example, when it is determined to be a "blackboard explanation" scene, the full-frame processing mode is started: First, the video stream is pre-processed frame by frame using an integrated neural network accelerator (such as the NPU), and the blackboard area is enhanced through edge detection. At the same time, audio is recorded at a 44.1kHz sampling rate to meet the clarity requirements of speech-to-text conversion.

[0043] Subsequently, timestamp synchronization technology is used to establish an index mapping between the OCR text results of each frame of video (such as the content on the blackboard) and the corresponding audio clip (such as "Next, look at this formula") to generate structured intermediate data ({timestamp: [text block, audio clip]}).

[0044] Finally, through a lightweight natural language processing module, the intermediate data is reorganized according to the "blackboard writing sequence + voice logic": repeated explanations are automatically filtered out (for example, only the first repeated explanation of the same formula is retained), text blocks are arranged in row and column order, and the audio-to-text conversion results are inserted into the corresponding positions in segments, and finally a Markdown-formatted minutes is generated (titles, formula blocks, and voice paragraphs are automatically layered).

[0045] For example, when it is determined to be a "tour explanation" scene, moving target detection is started: First, the teacher's body contour is set as the region of interest. When the teacher's movement speed is greater than 5 pixels / second, the current frame is marked as a "transition frame" and discarded. When a pause action is detected (speed less than 1 pixel / second and lasts for 2 seconds), the frame is retained and a local close-up (such as the experimental equipment operation screen) is extracted.

[0046] The built-in Voice Activity Detection (VAD) module then segments the audio stream and combines it with keyword recognition (e.g., "Pay attention here," "Next step") to locate key segments. Short-term energy analysis (50ms window) filters out ambient noise, retaining valid speech segments with a signal-to-noise ratio greater than 20dB.

[0047] Finally, the dynamic time warping (DTW) algorithm is used to align the retained key frame timestamps with the audio semantic blocks (allowing a deviation of ±500ms). The matched audio and video pairs are then compressed using H.265 encoding (compression ratio 10:1). At the same time, lightweight minutes are generated: key frame thumbnails are embedded in text segments, and only sentences containing keywords are retained in the audio-to-text conversion results.

[0048] This embodiment of the application determines classroom minutes extraction strategies for different teaching scenarios based on the correlation between the audio and video features of audio and video information in a specified area of ​​the classroom and user behavior characteristics. The audio and video information is processed according to different classroom minutes extraction strategies to generate personalized classroom minutes that meet the actual user needs, thereby improving the reliability of the generated classroom minutes.

[0049] In one embodiment of the present application, the teaching scene includes a first teaching scene, a second teaching scene or a third teaching scene; the audio and video information includes audio information and video information.

[0050] The audio and video information is processed according to the classroom minutes extraction strategy corresponding to the teaching scenario to generate classroom minutes corresponding to the audio and video information, including: In response to the teaching scene being the first teaching scene, an audio classroom minutes corresponding to the audio information is generated, and the audio classroom minutes is matched with the video information as the classroom minutes corresponding to the audio and video information.

[0051] In response to the teaching scene being the second teaching scene, an audio classroom minutes corresponding to the audio information is generated; multiple key frames in the video information are extracted; and the audio classroom minutes are matched with the multiple key frames as classroom minutes corresponding to the audio and video information.

[0052] In response to the teaching scene being the third teaching scene, key audio information corresponding to the audio information is extracted, and key audio classroom minutes corresponding to the key audio information are generated; multiple key frames in the video information are extracted; and the key audio classroom minutes are matched with the multiple key frames as classroom minutes corresponding to the audio and video information.

[0053] In this embodiment, the first teaching scenario is a key one. In this scenario, the teacher explains core knowledge points, formula derivations, or experimental principles in a target area, along with other content that requires detailed recording. This includes frequent blackboard updates and audio that includes complex technical terminology. In this scenario, it's necessary to generate complete audio classroom minutes (converting all audio to text) and match them frame-by-frame with the video information, ensuring that visual information such as the blackboard text and teacher gestures are fully synchronized with the audio explanation (with timestamp alignment accuracy of ±50ms) to ensure that no important content is missed.

[0054] The second teaching scenario is less important. In this scenario, the blackboard writing remains unchanged for extended periods (e.g., displaying a fixed chart), and the teacher primarily conducts a roving lecture. The blackboard area in the video remains essentially unchanged (10 consecutive frames show the same OCR result). In this scenario, redundant video frames are discarded (retaining only the initial blackboard writing and frames where the teacher annotates key points), reducing the video data volume by over 60%. At the same time, the lecture content is recorded through audio-to-text conversion. Keyframes and audio can be matched using the "voice keyword-image tag" function (e.g., when the teacher says "Please look at this picture," the corresponding frame is automatically marked).

[0055] The third teaching scenario is a non-critical one. It involves the teacher remaining silent while waiting for a student's response (audio power <40dB for at least 10 seconds), taking a break, or engaging in non-teaching activities (such as handing out workbooks, with no updates to the blackboard in the video, and the teacher's actions unrelated to the lesson). In this scenario, we can extract only key audio (such as the audio segment where the teacher asks, "Who will answer?") and key frames (such as the moment the teacher points to the blackboard), filtering out insignificant information (such as audio during periods of silence). This allows the minutes to focus on the core aspects of the lesson, reducing file size by 80% compared to the full record.

[0056] Specifically, in the first teaching scenario, the embodiment of the present application can use an 8kHz sampling rate + 16-bit quantization to record audio, segment valid speech segments through the VAD (Voice Activity Detection) module, and perform cloud-to-text conversion to identify valid text. Subsequently, professional terms can be automatically annotated (such as adding italic format to "Lorentz force"), and a text stream with a timestamp can be generated (such as "05:12 Next, derive the Lorentz force formula..."). Finally, each frame of video can be subjected to blackboard OCR recognition, and the text block can be aligned with the audio text stream: when "Formula (1)" appears in the audio, the corresponding formula area in the video at that moment is automatically extracted, and a mixed text and image summary is generated (a video screenshot is embedded on the right side of the text paragraph).

[0057] In the second teaching scenario, this embodiment of the application can calculate the difference between consecutive frames (based on a histogram intersection algorithm) and identify duplicate frames when the difference is less than 5%. Subsequently, the key frames can be arranged in chronological order, and an audio segment index can be added to each frame (e.g., "Key frame 3 corresponds to the audio segment from 08:25 to 08:40") to generate a Markdown-formatted summary.

[0058] In the third teaching scenario, the embodiment of the present application can establish a non-teaching state detection model: when the audio energy is less than 40dB and lasts for more than 15 seconds, or when keywords such as "take a break" are detected, it is marked as a non-teaching period. Only audio segments with teaching intentions can be retained: key audio is located through keyword matching (such as "question", "answer") and voice fundamental frequency mutation (the fundamental frequency increases by >8Hz when asking questions), and after extraction, it is converted into text and annotated with labels such as "[Teacher Questions]". Subsequently, the action recognition model can be used to detect the teacher's effective teaching actions, and then the key audio and key frames can be semantically associated to generate corresponding classroom minutes.

[0059] The embodiment of the present application adopts corresponding minutes extraction strategies in different teaching scenarios, which improves the efficiency and practicality of minutes generation, optimizes the storage and power consumption of smart glasses, and better meets user needs. The first teaching scenario fully records to ensure that important content is not missed, and frame-by-frame matching synchronizes audio and video, making it convenient to review complex knowledge. The second teaching scenario discards redundant frames to reduce the amount of data, and can also mark key frames with voice keywords, which saves storage space and highlights the key points. The third teaching scenario extracts key audio and video, filters out invalid information, focuses on the core of the minutes, and greatly reduces the file size.

[0060] In one embodiment of the present application, determining the teaching scenario corresponding to the audio and video information based on the correlation between the audio and video features and the behavioral features includes: Calculate the similarity between the audio and video features and each standard audio and video feature in the standard audio and video feature library, and select the standard audio and video feature with the largest similarity as the target audio and video feature; each standard audio and video feature corresponds to different teacher behaviors; Calculate the similarity between the behavior feature and each standard behavior feature in the standard behavior feature library, and select the standard behavior feature with the greatest similarity as the target behavior feature; each standard behavior feature corresponds to different user behaviors; Calculate the correlation between the target audio and video features and the target behavior features, and determine the teaching scene corresponding to the audio and video information based on the correlation.

[0061] In an embodiment of the present application, the standard audio and video feature library may store multiple sets of audio and video feature data of typical teaching scenarios, that is, multiple standard audio and video features. Different standard audio and video features or their combinations may correspond to specific teacher behaviors. For example, they may include blackboard explanations and roving explanations. The standard behavior feature library may store multiple sets of typical classroom behavior feature data of students, that is, multiple standard behavior features. Different standard behavior features or their combinations may correspond to specific learning behaviors. For example, they may include focused listening and distraction.

[0062] This embodiment of the application calculates the similarity between the extracted audio and video features and each audio and video feature, selects the standard audio and video feature with the greatest similarity as the target audio and video feature, and determines the teacher scene corresponding to the target audio and video feature. Teacher scenes can include: lecturing, interaction, questioning, silence, etc.

[0063] The embodiment of the present application calculates the similarity between the behavior feature and each standard behavior feature, selects the standard behavior feature with the greatest similarity as the target behavior feature, and determines the student scene corresponding to the target behavior feature. The student scene may include listening, looking, answering, being distracted, etc.

[0064] After obtaining the teacher scene and the student scene, the correlation between the target audio and video features and the target behavior features can be calculated, that is, the correlation between the teacher scene and the student scene can be calculated, and the corresponding teaching scene can be determined based on the correlation.

[0065] The embodiment of the present application realizes accurate determination of teaching scenes through dual feature library matching and correlation calculation. First, the audio and video features are matched with teacher behavior, and behavioral features with student status, and then the correlation between the two is combined for comprehensive judgment, which can take into account the dual dimensions of teacher teaching behavior and student learning status. On the one hand, based on the similarity matching of the standard feature library, the accuracy of scene recognition can be improved by using historical typical scene data; on the other hand, through the correlation analysis of teacher scenes and student scenes, the actual situation of classroom interaction can be more comprehensively reflected, so that the determined teaching scene takes into account both teacher behavior and student status, thereby providing a more accurate basis for the subsequent classroom minutes extraction strategy, and improving the pertinence and practicality of minute generation.

[0066] In one embodiment of the present application, calculating the correlation between target audio and video features and target behavior features, and determining the teaching scene corresponding to the audio and video information based on the correlation, includes: Calculate the correlation between the target audio and video features and the target behavioral features; If the correlation degree is greater than or equal to the first preset correlation degree, determining that the teaching scene corresponding to the audio and video information is the first teaching scene; If the correlation degree is less than the first preset correlation degree and greater than or equal to the second preset correlation degree, determining that the teaching scene corresponding to the audio and video information is the second teaching scene; If the correlation degree is less than the second preset correlation degree, determining that the teaching scene corresponding to the audio and video information is the third teaching scene; The first preset correlation degree is greater than the second preset correlation degree.

[0067] In this embodiment, the target audio and video features and target behavioral features can be converted into normalized vectors. For example, the target audio and video features include the blackboard update frequency f1 and the audio term density f2. The target behavioral features can include the gaze overlap rate s1 and the nodding frequency s2.

[0068] F=[f1,f2] S=[s1,s2] For example, in the blackboard explanation scenario, F = [0.8, 0.7], indicating a high frequency of blackboard writing and a high density of terminology. S = [0.9, 0.6], indicating a high rate of gaze overlap and a moderate frequency of nodding.

[0069] Then, the embodiment of the present application can construct a feature weight matrix and determine the correlation degree through cosine similarity calculation.

[0070] Feature weight matrix W=[w1,w2,w3,w4] Merge vector V=[f1,f2,s1,s2] Correlation formula:

[0071] in, is a standard eigenvector, the value range of C is [0,1], and wi represents the data in the i-th feature weight matrix.

[0072] In an embodiment of the present application, the correlation distribution of each scene can be statistically calculated based on historical data: for example, the mean correlation of the first scene is 0.82, and the standard deviation is 0.1. The first preset value is set to mean -0.12=0.7, and the second preset value is set to mean -0.42=0.4 to ensure that the threshold is adapted to different classroom scenarios. The specific adjustment can be made according to the actual situation.

[0073] In this embodiment of the application, when the correlation degree is ≥ a first preset value (e.g., 0.7), it is determined to be the first teaching scenario, indicating that the teacher is explaining core knowledge (e.g., formula derivation) and the students are highly focused, and the audio and video should be fully recorded.

[0074] If the second preset value is less than or equal to the correlation value and less than the first preset value (e.g., 0.4-0.7), the second teaching scenario is determined. This indicates that the teacher is explaining important content but the students' attention is moderate (e.g., static blackboard lecture). Only key frames and audio are retained.

[0075] If the correlation is less than the second preset value (e.g., less than 0.4), it is determined to be the third teaching scenario. This indicates that the teacher is not teaching (e.g., waiting in silence) or the student is distracted, and only key information is extracted and compressed into the record.

[0076] The embodiment of the present application can fully record in high-correlation scenarios to ensure that core knowledge is not missed, discard 60% of redundant video frames in medium-correlation scenarios but retain key frames associated with keywords, and only record 20% of core content in low-correlation scenarios, so that the size of classroom minutes files can be dynamically adjusted with the importance of the scene (maximum compression of 80%).

[0077] In one embodiment of the present application, after generating the classroom minutes corresponding to the audio and video information, the method further includes: Get the power level of the wearable device; If the power level of the device is greater than or equal to the first preset power level, the class minutes are adjusted according to the user's historical learning characteristics, and the adjusted class minutes are stored.

[0078] If the device battery level is less than the first preset battery level, the class minutes will not be adjusted.

[0079] In an embodiment of the present application, a user learning preference model can be pre-set in the wearable device and stored in the local flash memory. The learning preferences can include content preferences and format preferences. For example, content preferences can include historical minutes annotating high-frequency knowledge points and focus areas on the blackboard. Format preferences can include the user's editing habits for classroom minutes.

[0080] In an embodiment of the present application, if the battery level of the wearable device is greater than or equal to a first preset battery level, it indicates that the current battery level is high, and the user learning preference model can be invoked to optimize the currently generated class minutes without waiting for subsequent optimization. When the battery level of the wearable device is low, the user learning preference model can be omitted to improve battery life.

[0081] In one embodiment of the present application, obtaining audio and video information of a target area in a classroom includes: In response to the target area being the blackboard area, controlling the audio and video acquisition module in the wearable device to acquire audio and video information of the blackboard area in a first working mode; In response to the target area being a non-blackboard area, controlling the audio and video acquisition module in the wearable device to acquire audio and video information of the non-blackboard area in a second working mode; The audio and video acquisition resolution in the first working mode is greater than the audio and video acquisition resolution in the second working mode.

[0082] In classroom instruction, the blackboard is often used to write key content, such as core knowledge points and formula derivations. High-definition recording is required to ensure the accuracy and completeness of text, diagrams, and other information, facilitating the accurate generation of subsequent class minutes. Non-blackboard areas, such as areas where the teacher moves around and where students sit, often contain auxiliary information or non-teaching scenes, requiring lower resolution capture.

[0083] The smart glasses' built-in audio and video capture module is adjustable. When the target area is determined to be a blackboard area through regional feature matching or eye tracking (such as the aforementioned regional feature matching combined with user gaze), the capture module switches to the first operating mode, capturing audio and video at a higher resolution (e.g., 1920×1080 pixels) to ensure clear details of the blackboard content. If the target area is determined to be a non-blackboard area, it automatically switches to the second operating mode, capturing at a lower resolution (e.g., 1280×720 pixels or lower), reducing the amount of data while still meeting basic information recording requirements.

[0084] The embodiment of the present application can effectively reduce the amount of data collected by lowering the acquisition resolution of non-blackboard areas, reduce the workload of the processor and storage module, and reduce power consumption, thereby extending the battery life of the smart glasses and ensuring stable operation of the equipment throughout the class.

[0085] Corresponding to the classroom minutes generation method of the above embodiment, Figure 2 This is a structural block diagram of a classroom minutes generation system provided by an embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown. Figure 2 The classroom minutes generation system 20 is applied to a wearable device and includes: a first acquisition module 201, a second acquisition module 202, a first control module 203 and a second control module 204. The first acquisition module 201 is used to acquire audio and video information of a target area in the classroom and perform feature extraction on the audio and video information to obtain audio and video features; The second acquisition module 202 is used to obtain behavior information of the user wearing the wearable device and perform feature extraction on the behavior information to obtain behavior features; A first control module 203 is configured to determine a teaching scenario corresponding to the audio and video information based on the correlation between the audio and video features and the behavioral features; The second control module 204 is configured to process the audio and video information according to a classroom minutes extraction strategy corresponding to the teaching scenario, and generate classroom minutes corresponding to the audio and video information.

[0086] In one embodiment of the present application, the teaching scene includes a first teaching scene, a second teaching scene, or a third teaching scene; the audio and video information includes audio information and video information; The second control module 204 is specifically configured to generate an audio classroom minutes corresponding to the audio information in response to the teaching scene being the first teaching scene, and match the audio classroom minutes with the video information as the classroom minutes corresponding to the audio and video information; In response to the teaching scene being the second teaching scene, generating an audio classroom minutes corresponding to the audio information; extracting multiple key frames from the video information; matching the audio classroom minutes with the multiple key frames as the classroom minutes corresponding to the audio and video information; In response to the teaching scene being the third teaching scene, key audio information corresponding to the audio information is extracted, and key audio classroom minutes corresponding to the key audio information are generated; multiple key frames in the video information are extracted; and the key audio classroom minutes are matched with the multiple key frames as classroom minutes corresponding to the audio and video information.

[0087] In one embodiment of the present application, the first control module 203 is specifically configured to calculate the similarity between the audio and video feature and each standard audio and video feature in the standard audio and video feature library, and select the standard audio and video feature with the greatest similarity as the target audio and video feature; each standard audio and video feature corresponds to a different teacher behavior; Calculate the similarity between the behavior feature and each standard behavior feature in the standard behavior feature library, and select the standard behavior feature with the greatest similarity as the target behavior feature; each standard behavior feature corresponds to different user behaviors; Calculate the correlation between the target audio and video features and the target behavior features, and determine the teaching scene corresponding to the audio and video information based on the correlation.

[0088] In one embodiment of the present application, the first control module 203 is specifically configured to calculate the correlation between the target audio and video features and the target behavior features; If the correlation degree is greater than or equal to the first preset correlation degree, determining that the teaching scene corresponding to the audio and video information is the first teaching scene; If the correlation degree is less than the first preset correlation degree and greater than or equal to the second preset correlation degree, determining that the teaching scene corresponding to the audio and video information is the second teaching scene; If the correlation degree is less than the second preset correlation degree, determining that the teaching scene corresponding to the audio and video information is the third teaching scene; The first preset correlation degree is greater than the second preset correlation degree.

[0089] In one embodiment of the present application, the system may further include: A third control module is configured to obtain a power level of the wearable device after generating a classroom summary corresponding to the audio and video information; If the power level of the device is greater than or equal to the first preset power level, the class minutes are adjusted according to the user's historical learning characteristics, and the adjusted class minutes are stored.

[0090] In one embodiment of the present application, the third control module is further configured to not adjust the class minutes if the power level of the device is less than a first preset power level.

[0091] In one embodiment of the present application, the first acquisition module 201 is specifically configured to control the audio and video acquisition module in the wearable device to acquire audio and video information of the blackboard area in a first working mode in response to the target area being the blackboard area; In response to the target area being a non-blackboard area, controlling the audio and video acquisition module in the wearable device to acquire audio and video information of the non-blackboard area in a second working mode; The audio and video acquisition resolution in the first working mode is greater than the audio and video acquisition resolution in the second working mode.

[0092] See also Figure 3 , Figure 3 This is a schematic block diagram of a wearable device provided in one embodiment of the present application. Figure 3 The wearable device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of the modules in the above-mentioned device embodiments, such as Figure 2 The functions of the first acquisition module 201, the second acquisition module 202, the first control module 203 and the second control module 204 are shown.

[0093] It should be understood that in the embodiment of the present application, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0094] The input device 302 may include a touchpad, a fingerprint collection sensor (for collecting user fingerprint information and fingerprint direction information), a microphone, etc. The output device 303 may include a display (LCD, etc.), a speaker, etc.

[0095] The memory 304 may include a read-only memory and a random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include a nonvolatile random access memory.

[0096] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiment of the present application can execute the implementation method described in the classroom minutes generation method provided in the embodiment of the present application, and can also execute the implementation method of the wearable device described in the embodiment of the present application, which will not be repeated here.

[0097] In another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, all or part of the process of the method in the above embodiment is implemented. The computer program can also be used to instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above method embodiments are implemented. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium.

[0098] The computer-readable storage medium can be the internal storage unit of the wearable device of any of the aforementioned embodiments, such as the wearable device's hard drive or memory. The computer-readable storage medium can also be an external storage device of the wearable device, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the computer-readable storage medium can include both the wearable device's internal storage unit and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the wearable device. The computer-readable storage medium can also be used to temporarily store data that has been output or is about to be output.

[0099] An embodiment of the present application provides a computer program product, which includes computer-executable instructions or a computer program, which is stored in a computer-readable storage medium. The processor of the wearable device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the wearable device executes the classroom minutes generation method described above in the embodiment of the present application.

[0100] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0101] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the wearable device and unit described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0102] In the several embodiments provided in this application, it should be understood that the disclosed wearable devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces or units, or may be an electrical, mechanical or other form of connection.

[0103] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0104] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for generating classroom minutes, characterized in that: Applied to wearable devices, including: Acquire audio and video information of a target area in the classroom, and perform feature extraction on the audio and video information to obtain audio and video features; Obtaining behavior information of a user wearing the wearable device, and performing feature extraction on the behavior information to obtain behavior features; Determining a teaching scenario corresponding to the audio and video information based on a correlation between the audio and video features and the behavioral features; The audio and video information is processed according to the classroom minutes extraction strategy corresponding to the teaching scene to generate classroom minutes corresponding to the audio and video information.

2. The method for generating classroom minutes according to claim 1, wherein: The teaching scene includes a first teaching scene, a second teaching scene or a third teaching scene; the audio and video information includes audio information and video information; The step of processing the audio and video information according to the classroom minutes extraction strategy corresponding to the teaching scenario to generate classroom minutes corresponding to the audio and video information includes: In response to the teaching scene being the first teaching scene, generating an audio classroom minutes corresponding to the audio information, and matching the audio classroom minutes with the video information as the classroom minutes corresponding to the audio and video information; In response to the teaching scene being the second teaching scene, generating an audio classroom minutes corresponding to the audio information; extracting multiple key frames from the video information; matching the audio classroom minutes with the multiple key frames as the classroom minutes corresponding to the audio and video information; In response to the teaching scene being the third teaching scene, key audio information corresponding to the audio information is extracted, and key audio classroom minutes corresponding to the key audio information are generated; multiple key frames in the video information are extracted; and the key audio classroom minutes are matched with the multiple key frames as classroom minutes corresponding to the audio and video information.

3. The method for generating classroom minutes according to claim 2, wherein: The determining, based on the correlation between the audio and video features and the behavioral features, of a teaching scenario corresponding to the audio and video information includes: Calculating the similarity between the audio and video feature and each standard audio and video feature in the standard audio and video feature library, and selecting the standard audio and video feature with the greatest similarity as the target audio and video feature; each standard audio and video feature corresponds to a different teacher behavior; Calculating the similarity between the behavior feature and each standard behavior feature in the standard behavior feature library, and selecting the standard behavior feature with the greatest similarity as the target behavior feature; each standard behavior feature corresponds to a different user behavior; The correlation between the target audio and video features and the target behavior features is calculated, and the teaching scene corresponding to the audio and video information is determined according to the correlation.

4. The method for generating classroom minutes according to claim 3, wherein: The calculating the correlation between the target audio and video features and the target behavior features, and determining the teaching scene corresponding to the audio and video information according to the correlation, includes: Calculating the correlation between the target audio and video features and the target behavior features; If the correlation degree is greater than or equal to a first preset correlation degree, determining that the teaching scene corresponding to the audio and video information is a first teaching scene; If the correlation degree is less than the first preset correlation degree and greater than or equal to the second preset correlation degree, determining that the teaching scene corresponding to the audio and video information is the second teaching scene; If the correlation degree is less than a second preset correlation degree, determining that the teaching scene corresponding to the audio and video information is a third teaching scene; The first preset correlation degree is greater than the second preset correlation degree.

5. The method for generating classroom minutes according to claim 1, wherein: After generating the classroom minutes corresponding to the audio and video information, the method further includes: Obtaining the power level of the wearable device; If the power level of the device is greater than or equal to a first preset power level, the class minutes are adjusted according to the user's historical learning characteristics, and the adjusted class minutes are stored.

6. The method for generating classroom minutes according to claim 5, wherein: Also includes: If the power level of the device is less than the first preset power level, the class minutes will not be adjusted.

7. The method for generating classroom minutes according to claim 1, wherein: The step of obtaining audio and video information of a target area in a classroom includes: In response to the target area being a blackboard area, controlling the audio and video acquisition module in the wearable device to acquire audio and video information of the blackboard area in a first working mode; In response to the target area being a non-blackboard area, controlling the audio and video acquisition module in the wearable device to acquire audio and video information of the non-blackboard area in a second working mode; The audio and video acquisition resolution in the first working mode is greater than the audio and video acquisition resolution in the second working mode.

8. A classroom minutes generation system, characterized in that: Applied to wearable devices, including: The first acquisition module is used to acquire audio and video information of a target area in the classroom and perform feature extraction on the audio and video information to obtain audio and video features; a second acquisition module, configured to acquire behavior information of a user wearing the wearable device, and perform feature extraction on the behavior information to obtain behavior features; A first control module is configured to determine a teaching scenario corresponding to the audio and video information based on a correlation between the audio and video features and the behavioral features; The second control module is used to process the audio and video information according to the classroom minutes extraction strategy corresponding to the teaching scene, and generate classroom minutes corresponding to the audio and video information.

9. A wearable device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Classroom scene-oriented human body behavior identification method

    CN110414415A

  • Gaze based classroom notes generator

    CN110476195A

  • Intelligent online note generation system

    CN111079714A

  • Classroom teaching quality evaluation method and device, equipment and storage medium

    CN111898881A

  • Method and device for automatically generating curriculum schedule, electronic equipment and storage medium

    CN112507679A