Classroom data identification method and device, electronic equipment and storage medium

By extracting professional term data from lecture notes and assisting correction of classroom audio and video transcription content, the problem of insufficient classroom transcription accuracy in the existing technology is solved, high-accurate transcription and teaching analysis are achieved, and high-quality review materials are provided.

CN119989194APending Publication Date: 2025-05-13ANHUI ZHUOZHI EDUCATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510049228.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing phonological transcription techniques show high error rates in classrooms involving a large number of professional terms, making it difficult to meet the needs of high-accuracy transcription.

Method used

By extracting relevant professional term data from the lecture notes, using term data to assist in correction of audio and video transcripts, improving the accuracy of audio transcription, and conducting teaching analysis to output high-quality review materials.

Benefits of technology

It significantly improves the accuracy of audio transcription, meets students' after-class review needs, and provides high-quality review materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989194A_ABST
    Figure CN119989194A_ABST
Patent Text Reader

Abstract

The invention discloses a classroom data identification method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining classroom data which comprises audio and video data and lecture data; obtaining term data associated with the classroom based on the lecture data; identifying the audio and video data based on the term data to obtain text data associated with the video data; and correcting the text data based on the term data, and performing teaching analysis and identification on the corrected text data to obtain a teaching identification result, the teaching identification result comprising a teaching text marked by terms. According to the classroom data identification method, related professional term data can be extracted from lectures, audio and video transcription contents are subjected to auxiliary correction by using the term data, the accuracy of audio transcription is improved, teaching analysis is performed, and high-quality review data are output, so that the after-class review requirements of students are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to technical fields such as speech recognition and classroom transcription, and in particular to a method, device, electronic device and storage medium for identifying classroom data. Background Art

[0002] Classroom video review has gradually become an important tool for college students to review. However, existing speech transcription technology often shows a high error rate in classes involving a large number of professional terms, and it is difficult to meet the needs of high-accuracy transcription.

[0003] In related technologies, in order to improve the accuracy of transcription, one method is to develop a speech recognition system for a specific field and use data from a specific field (such as finance, law, etc.) for training, which can significantly improve the accuracy of transcription in related classes. However, this method is limited to specific fields and the effect of class transcription in other fields is still unsatisfactory. The second method is a speech recognition system that can customize hot words. It uses the provided hot word list for incentive enhancement to improve the recall rate and accuracy of hot words, but the hot words for each class need to be customized manually, which is not automated enough and is not feasible for large-scale promotion. Summary of the invention

[0004] To this end, the purpose of the implementation mode of the present application is to propose a method, device, electronic device, storage medium and computer program product for identifying classroom data. The method for identifying classroom data of the present invention can extract relevant professional terminology data from handouts, use the terminology data to assist in correcting the audio and video transcription content, improve the accuracy of audio transcription, and conduct teaching analysis, and output high-quality review materials to meet the students' after-class review needs.

[0005] An embodiment of the present application provides a method for identifying classroom data, the method comprising: acquiring classroom data, wherein the classroom data comprises audio and video data and handout data; obtaining terminology data associated with the classroom based on the handout data; identifying the audio and video data based on the terminology data to obtain text data associated with the audio and video data; correcting the text data based on the terminology data, and performing teaching analysis and identification on the corrected text data to obtain a teaching identification result, wherein the teaching identification result comprises a teaching text marked with terminology.

[0006] Exemplarily, the audio and video data includes audio data and video data, and the method further includes: identifying the audio data to determine speaker result data corresponding to the audio data; matching the speaker result data based on the video data and the campus database to obtain speaker identity data corresponding to the speaker result data.

[0007] Exemplarily, the lecture data includes a lecture format, and obtaining terminology data associated with the classroom based on the lecture data includes: parsing the lecture data based on a preset method, and extracting the parsing results based on a preset named entity model and a preset vocabulary library to obtain candidate vocabulary, wherein the preset method corresponds one-to-one to the lecture format; and screening the candidate vocabulary based on a pre-trained classification model to obtain terminology data associated with the classroom.

[0008] Exemplarily, the identifying the audio data and determining the speaker result data corresponding to the audio data includes: segmenting the audio data to obtain multiple audio segments; extracting feature data of the multiple audio segments, and clustering the multiple audio segment data based on the feature data; and determining the speaker result data corresponding to the audio data based on the clustering result.

[0009] Exemplarily, the segmenting of the audio data to obtain multiple audio segments includes: segmenting the audio data based on a preset endpoint model to obtain multiple first audio segments; segmenting the multiple first audio segments based on a preset length to obtain multiple second audio segments, wherein the first audio segment is composed of multiple second audio segments; and obtaining multiple audio segments based on the multiple second audio segments.

[0010] Exemplarily, clustering the multiple audio segment data based on the feature data includes: determining a category corresponding to each of the second audio segments; if the categories corresponding to all the second audio segments in the first audio segment are the same, retaining the endpoints of the first audio segment, otherwise, retaining the endpoints of the second audio segment; and obtaining a clustering processing result based on the retained endpoints.

[0011] Exemplarily, the clustering processing result includes multiple audio endpoints, and determining speaker result data corresponding to the audio data based on the clustering processing result includes: determining that an audio segment between two adjacent endpoints belongs to the same speaker.

[0012] Exemplarily, the terminology data includes terminology pronunciation data, the text data includes text pronunciation data, and the modifying the text data based on the terminology data includes: when the text pronunciation data corresponding to the text data is the same as the terminology pronunciation data corresponding to the terminology data, if the text data is different from the terminology data, then modifying the text data to the terminology data.

[0013] Exemplarily, the text data includes multiple text segment data, the teaching recognition result includes a teaching text, and the teaching analysis and recognition of the corrected text data to obtain the teaching recognition result includes: extracting feature data of the multiple text segment data, and determining attribute information of the text segment data based on the feature data corresponding to the text segment data, wherein the attribute information includes at least one of a question attribute and a statement attribute; when the attribute information of the text segment data is a statement attribute and the text segment includes terminology data, determining that the text segment data is a teaching text.

[0014] Exemplarily, the identifying the audio and video data based on the terminology data to obtain text data associated with the audio and video data includes: determining the weight corresponding to the terminology data based on a preset speech recognition model; identifying the audio and video data based on the weight corresponding to the terminology data to obtain text data associated with the audio and video data.

[0015] Exemplarily, the extracting the parsing results based on a preset named entity model and a preset vocabulary library to obtain candidate words includes: extracting the parsing results based on a preset named entity model and a preset vocabulary library, and determining that the text line with a word count less than a preset word count is the candidate word.

[0016] Another embodiment of the present application provides a device for identifying classroom data, the device comprising: an acquisition module, used to acquire classroom data, wherein the classroom data includes audio and video data and handout data; a first acquisition module, used to obtain terminology data associated with the classroom based on the handout data; a second acquisition module, used to identify the audio and video data based on the terminology data, and obtain text data associated with the audio and video data; an identification module, used to correct the text data based on the terminology data, and perform teaching analysis and identification on the corrected text data to obtain a teaching identification result, wherein the teaching identification result includes a teaching text marked with terminology.

[0017] Another embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method of any of the above embodiments when executing the computer program.

[0018] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method of any of the above embodiments are implemented.

[0019] Another embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed by a processor of a computer device, the computer device is enabled to perform the steps of the method of any of the above embodiments.

[0020] In the above implementation, the method for identifying classroom data includes: obtaining classroom data, wherein the classroom data includes audio and video data and handout data; obtaining terminology data associated with the classroom based on the handout data; identifying the audio and video data based on the terminology data to obtain text data associated with the video data; correcting the text data based on the terminology data, and performing teaching analysis and identification on the corrected text data to obtain a teaching identification result, wherein the teaching identification result includes a teaching text marked with terminology. The method for identifying classroom data of the present invention can extract relevant professional terminology data from handouts, use terminology data to assist in correcting audio and video transcription content, improve the accuracy of audio transcription, and perform teaching analysis to output high-quality review materials to meet students' after-class review needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A flow chart of a method for identifying classroom data provided in an embodiment of the present application;

[0022] Figure 2 A flowchart of obtaining terminology data associated with a classroom provided for an embodiment of the present application;

[0023] Figure 3 A flowchart of performing person recognition on audio data provided by an embodiment of the present application;

[0024] Figure 4 A flowchart for determining speaker result data corresponding to audio data provided by an embodiment of the present application;

[0025] Figure 5 A flowchart for segmenting audio data provided in an embodiment of the present application;

[0026] Figure 6 A flowchart of clustering multiple audio segment data provided by an embodiment of the present application;

[0027] Figure 7 A flowchart for identifying audio and video data based on terminology data provided in an embodiment of the present application;

[0028] Figure 8 A flowchart for teaching analysis and identification of the corrected text data provided in the implementation mode of the present application;

[0029] Fig. 9 A software execution flow chart provided for the implementation method of this application;

[0030] Fig.10 A schematic diagram of a device for identifying classroom data provided in an embodiment of the present application;

[0031] Fig.11 A block diagram of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION

[0032] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0033] Classroom video review has gradually become an important tool for college students to review. However, existing speech transcription technology often shows a high error rate in classes involving a large number of professional terms, and it is difficult to meet the needs of high-accuracy transcription.

[0034] In related technologies, in order to improve the accuracy of transcription, one method is to develop a speech recognition system for a specific field and use data from a specific field (such as finance, law, etc.) for training, which can significantly improve the accuracy of transcription in related classes. However, this method is limited to specific fields and the effect of class transcription in other fields is still unsatisfactory. The second method is a speech recognition system that can customize hot words. It uses the provided hot word list for incentive enhancement to improve the recall rate and accuracy of hot words, but the hot words for each class need to be customized manually, which is not automated enough and is not feasible for large-scale promotion.

[0035] Based on this, this application proposes a speech transcription enhancement method based on the assistance of classroom handouts. After a class, the relevant professional terms of this class are extracted from the handouts, and a large model is used to assist in correction. The features of video and audio are combined to distinguish different speakers, improve the accuracy of audio transcription, and finally conduct teaching analysis to output high-quality review materials with teaching marks to meet students' after-class review needs.

[0036] Figure 1 This is a flowchart of a method for identifying classroom data according to an embodiment of the present application.

[0037] As an example, Figure 1 As shown, the identification methods of classroom data include:

[0038] S101, obtaining classroom data, wherein the classroom data includes audio and video data and handout data.

[0039] S102, obtaining terminology data associated with the class based on the lecture data.

[0040] S103, identifying the audio and video data based on the terminology data to obtain text data associated with the audio and video data.

[0041] S104, correcting the text data based on the terminology data, and performing teaching analysis and recognition on the corrected text data to obtain a teaching recognition result, wherein the teaching recognition result includes a teaching text marked with terminology.

[0042] Exemplarily, the classroom data recognition method of the present application can be applied to a voice transcription system. The voice transcription system is connected to a shooting device. After the class is over, the shooting device can automatically upload the classroom data to the voice transcription system. The classroom data includes audio and video data and handout data during the class. The handout data can be uploaded to the system by the teaching staff. After a class is over, the system obtains the classroom data. First, the handout data is identified and analyzed to obtain terminology data associated with the class. The audio and video data is identified based on the terminology data to obtain text data associated with the audio and video data. The present application uses the terminology data extracted from the handout data to stimulate the transcription of audio and video data, improve the recall rate and accuracy of hot words, and obtain more accurate text data associated with the video data. There is no need to manually formulate hot words, improve transcription efficiency, and save labor costs.

[0043] Exemplarily, the present application further modifies the text data based on the terminology data, that is, the transcribed content is modified according to the terminology data to obtain more accurate transcribed content. The modified text data is subjected to teaching analysis and recognition to obtain teaching recognition results, which include teaching texts marked with terminology. For example, the modified text data is analyzed, the terms in the text data are marked, and teaching texts with knowledge point marks are obtained, providing high-quality after-class review materials for students to review after class.

[0044] The classroom data recognition method of the present application extracts relevant professional terminology data from the handouts, uses the terminology data to assist in the correction of audio and video transcription content, improves the accuracy of audio transcription, and conducts teaching analysis to output high-quality review materials to meet students' after-class review needs.

[0045] As an example, Figure 2 As shown, the lecture data includes the lecture format, and term data associated with the class is obtained based on the lecture data, including:

[0046] S201, parsing the lecture data based on a preset method, and extracting the parsing results based on a preset named entity model and a preset vocabulary library to obtain candidate vocabulary, wherein the preset method corresponds to the lecture format one by one.

[0047] S202, screening candidate words based on the pre-trained classification model to obtain terminology data associated with the classroom.

[0048] Exemplarily, the speech transcription system includes a document parsing module. The document parsing module is used to parse the lecture data to obtain terminology data associated with the classroom. The lecture data includes a lecture format, for example, the lecture format includes a PDF file, a PPT file, a Word file, and the like. The lecture data is parsed based on a preset method, and the preset method corresponds to the lecture format one by one. The lecture data can be parsed through tools such as pdfminer, pptx, docx, etc. to obtain the lecture text content. For example, pdfminer is used to parse PDF files and extract text content, python-pptx is used to parse PPT files and extract text content of each page, python-docx is used to parse Word documents and extract paragraph text, and the like.

[0049] Exemplarily, the parsed results are then extracted based on a preset named entity model and a preset vocabulary library to obtain candidate vocabulary. The preset named entity model can be a NER model (Named Entity Recognition). The preset vocabulary library can be a Wikipedia vocabulary library or other vocabulary library, or a customized vocabulary library, etc. The NER model is used to process all the parsed lecture text content, and professional terms are identified and extracted therefrom. For example, using a pre-trained NER model, the parsed lecture text is input, and a vocabulary list containing entity types is output. Then, using a preset vocabulary library such as a Wikipedia title as a dictionary, the extracted vocabulary is matched to ensure the professionalism of the terminology, the identified entity is matched with the Wikipedia title, the professional terms existing in Wikipedia are determined, and candidate vocabulary is obtained.

[0050] Exemplarily, candidate words are screened based on a pre-trained classification model to obtain terminology data associated with the class. The pre-trained classification model can be a Qianwen big model. For example, the Qianwen big model is used to screen candidate words, remove sensitive words and irrelevant words, and obtain terminology data associated with the class. For example, the big model interface can be called and input: "Please select professional terms related to the course name from the candidate words." The big model will screen and return relevant terms.

[0051] As an example, the parsing results are extracted based on a preset named entity model and a preset vocabulary library to obtain candidate words, including: extracting the parsing results based on a preset named entity model and a preset vocabulary library, and determining that the text line with a word count less than a preset word count is a candidate word.

[0052] For example, the present application takes into account the particularity of lecture data, which usually contains title texts, and most of the title texts are relatively short. When extracting the parsing results based on the preset named entity model and the preset vocabulary library, the present application uses the texts in the lecture texts whose number of characters is less than the preset number of characters as candidate words. The preset number of characters can be, for example, 6 characters, for example, the texts with less than 6 characters are directly listed as candidate words.

[0053] The document parsing module of this application is able to parse lecture files in various formats, including PDF, PPT, and DOCX, and automatically extract professional terms to ensure the high relevance and professionalism of the terms. Specifically, this module uses advanced document parsing tools and named entity recognition technology to extract terms from the lecture notes that are highly relevant to the classroom content. For example, use lines of text with less than six words as titles to ensure the validity of the terms. In addition, the Qianwen model is called to input the course name and extracted vocabulary to select terms that are highly relevant to the course and exclude irrelevant and sensitive words. By counting the frequency of occurrence of terms in the lecture notes, a term dictionary and a pinyin dictionary are generated to show the importance of each term in the course.

[0054] As an example, Figure 3 As shown, the audio and video data includes audio data and video data, and the method for identifying classroom data also includes:

[0055] S301, identifying audio data to determine speaker result data corresponding to the audio data.

[0056] S302, matching the speaker result data based on the video data and the campus database to obtain speaker identity data corresponding to the speaker result data.

[0057] Exemplarily, the speech transcription system also includes a speech processing module, which is used to receive audio and video data of the classroom, extract audio tracks, and perform speech enhancement on the audio, and the audio and video data include audio data and video data. First, the audio data is recognized to determine the speaker result data corresponding to the audio data. Based on the video data and the campus database, the speaker result data is matched to obtain the speaker identity data corresponding to the speaker result data. It can be understood that the present application uses the segmentation rules of the audio segment to improve the accuracy of distinguishing different speakers, and records the speaker information by matching the video facial features.

[0058] It should be noted that the collection and use of user personal information involved in this application are carried out on the basis of obtaining user authorization.

[0059] As an example, Figure 4 As shown, the audio data is identified to determine the speaker result data corresponding to the audio data, including:

[0060] S401, segmenting the audio data to obtain multiple audio segments.

[0061] S402: extract feature data of multiple audio segments, and perform clustering processing on the multiple audio segment data based on the feature data.

[0062] S403: Determine speaker result data corresponding to the audio data based on the clustering processing result.

[0063] Exemplarily, the ffmpeg tool can be used to extract the audio track from the video, and the sox command can be used to enhance the voice. After the enhanced audio data is obtained, the audio data is segmented to obtain multiple audio segments. The audio data can be segmented by an endpoint detection model to segment the audio data into many segments of sentence-delimited audio to obtain multiple audio segments. Feature data of the multiple audio segments are extracted, and the multiple audio segment data are clustered based on the feature data. For example, each audio segment is classified into a category, and the multiple audio segment data can be clustered based on the audio segments of the same category as the same person speaking, and the speaker result data corresponding to the audio data is determined based on the clustering processing result.

[0064] As an example, Figure 5 As shown, the audio data is segmented to obtain multiple audio segments, including:

[0065] S501, segmenting audio data based on a preset endpoint model to obtain a plurality of first audio segments.

[0066] S502: Segment the plurality of first audio segments based on a preset length to obtain a plurality of second audio segments, wherein the first audio segment is composed of the plurality of second audio segments.

[0067] S503: Obtain multiple audio segments based on the multiple second audio segments.

[0068] Exemplarily, the audio data can be segmented at silent points based on a preset endpoint model to obtain multiple first audio segments. In the segmented first audio segment, there may be more than one person speaking, resulting in less accurate results. In order to improve the accuracy, the present application performs smaller segmentation on the segmented audio, i.e., the first audio segment, according to an appropriate preset length, that is, multiple first audio segments are segmented based on the preset length to obtain multiple second audio segments. It can be understood that the first audio segment is composed of multiple second audio segments. Determining the multiple second audio segments as the multiple audio segments required by this application, and the subsequent clustering processing of the audio segments are all based on the second audio segments.

[0069] As an example, Figure 6 As shown, clustering processing is performed on multiple audio segment data based on feature data, including:

[0070] S601: Determine a category corresponding to each second audio segment.

[0071] S602: If the categories corresponding to all the second audio segments in the first audio segment are the same, retain the endpoints of the first audio segment; otherwise, retain the endpoints of the second audio segment.

[0072] S603: Obtain a clustering processing result based on the reserved endpoints.

[0073] Exemplarily, feature data of each second audio segment is extracted, and the category corresponding to each second audio segment is determined based on the feature data, where the category can be in the form of a label. If the categories corresponding to all second audio segments in the first audio segment are the same, it means that the first audio segment is for the same speaker, then the endpoints of the first audio segment are retained and the endpoints of the second audio segments in the first audio segment are merged. If the categories corresponding to all second audio segments in the first audio segment are not all the same, it means that the first audio segment includes at least two speakers, then the endpoints of the second audio segment are retained. The clustering processing result is obtained based on the retained endpoints. It can be understood that if all small audios in a segmented audio are classified into one category, it means that this segmented audio is for one speaker, and the segmentation state of the segmented audio is retained. Otherwise, the segmentation breakpoints of the small audios are retained.

[0074] As an example, the clustering processing result includes multiple audio endpoints, and determining speaker result data corresponding to the audio data based on the clustering processing result includes: determining that an audio segment between two adjacent endpoints belongs to the same speaker.

[0075] Exemplarily, according to the above clustering process, a clustering process result is obtained. The clustering process result includes multiple audio endpoints. Among them, the first audio segment endpoint and the second audio segment endpoint are included. It is determined that the audio segment between two adjacent endpoints belongs to the same speaker.

[0076] After obtaining the speaker result data corresponding to the audio data, the speaker result data is matched based on the video data and the campus database to obtain the speaker identity data corresponding to the speaker result data. For example, the facial features of the speaker in focus in the video data can be extracted, compared with the campus face database, and the speaker identity information is returned to record the specific speaker of each voice clip. If the video is not in focus, it can be regarded as the collective answer of all students.

[0077] As an example, Figure 7 As shown, the audio and video data are identified based on the term data to obtain text data associated with the audio and video data, including:

[0078] S701, determining the weight corresponding to the term data based on a preset speech recognition model.

[0079] S702, identifying the audio and video data based on the weights corresponding to the term data, and obtaining text data associated with the audio and video data.

[0080] Exemplarily, the speech processing module is also used to receive classroom videos and transcribe the audio to obtain text data associated with the audio and video data. The audio and video data can be recognized by using a preset speech recognition model, and the preset speech recognition model can adopt an ASR (Automatic Speech Recognition) model. The audio segment after the above speaker recognition processing can be input into the preset speech recognition model together with the terminology data extracted by the document parsing module. The preset speech recognition model can determine the incentive weight of the term in speech recognition based on the frequency of occurrence of the terminology data, wherein after the above extraction of the terminology data, the number of times the term appears in the handout content can be counted, and a terminology dictionary and a terminology pinyin dictionary can be output, including vocabulary and its number. The audio and video data are recognized based on the weight corresponding to the terminology data to obtain text data associated with the audio and video data.

[0081] This application inputs terminology data into the ASR hot word model, performs recognition incentives according to different weights, and improves the recall rate and accuracy of professional terms in the transcribed text.

[0082] The speech processing module of the present application improves the audio quality through speech enhancement technology, and inputs the extracted professional terms into the automatic speech recognition hot word model. The specific process includes extracting the audio track from the classroom video and performing speech enhancement to improve clarity and quality. The audio is segmented through the endpoint detection model, the system cuts the audio into paragraphs, and performs cluster analysis on small audios for each paragraph to ensure accurate identification of the speaker. In addition, the facial features of the speaker are extracted in combination with the class video, and compared with the campus database to return identity information to record the specific speaker. Subsequently, the enhanced audio is input into the ASR hot word model together with the terminology dictionary extracted by the document parsing module, and the vocabulary and its weight that need to be stimulated are determined through the terminology dictionary, thereby significantly improving the recall rate and accuracy of the professional terms in the transcribed text, ensuring the high precision and high credibility of the transcription results.

[0083] As an example, the terminology data includes terminology pronunciation data, the text data includes text pronunciation data, and the text data is modified based on the terminology data, including: when the text pronunciation data corresponding to the text data is the same as the terminology pronunciation data corresponding to the terminology data, if the text data is different from the terminology data, then the text data is modified to the terminology data.

[0084] Exemplarily, the speech transcription system also includes a transcription analysis module for further improving the accuracy of the terminology, using a terminology pinyin dictionary to replace incorrect vocabulary, and further improving the accuracy of the terminology in the transcribed text. When acquiring terminology data, the terminology dictionary and the terminology pinyin dictionary can be output for correcting the transcribed text data. The transcribed text output by the speech recognition model can be first converted into a pinyin text, and when the text pronunciation data corresponding to the text data is the same as the terminology pronunciation data corresponding to the terminology data, if the text data is not the same as the terminology data, the text data is corrected to the terminology data. It can be understood that incorrect vocabulary in which the transcribed text pinyin and the terminology pinyin are the same but the characters are different is corrected to the correct terminology form to ensure the accuracy of the transcription of the terminology.

[0085] As an example, Figure 8 As shown, the text data includes multiple text segment data, and the teaching recognition result includes the teaching text. The corrected text data is subjected to teaching analysis and recognition to obtain the teaching recognition result, including:

[0086] S801, extracting feature data of a plurality of text segment data, and determining attribute information of the text segment data based on the feature data corresponding to the text segment data, wherein the attribute information includes at least one of a question attribute and a statement attribute.

[0087] S802: When the attribute information of the text segment data is a statement attribute and the text segment includes terminology data, determine that the text segment data is a teaching text.

[0088] Exemplarily, the transcription analysis module also performs teaching analysis and identification on the corrected text data to obtain a teaching identification result. The text data includes multiple text segment data, extracts feature data of the multiple text segment data, and determines the attribute information of the text segment data based on the feature data corresponding to the text segment data, and the attribute information includes at least one of a question attribute and a statement attribute. For example, for each text segment, the feature data of the previous text segment and the next text segment are spliced ​​and input into the Fasttext model to determine the attribute information of the text segment. For example, determine whether the text segment is a question statement. When the attribute information of the text segment data is a statement attribute and the text segment includes terminology data, the text segment data is determined to be a teaching text. After finding all the question text segments through a preset model, if the remaining text segments contain words in the terminology dictionary, they are determined to be teaching texts, and the teaching texts are convenient for students to review relevant knowledge points.

[0089] The transcription analysis module of this application further improves the accuracy of terminology and conducts teaching analysis on the teacher's speech text. The transcription text is converted into pinyin text, and incorrect words are replaced using the terminology pinyin dictionary to further improve the accuracy of terminology in the transcription text. For the teacher's speech text, the question text is judged by splicing the context text features, and the knowledge points are marked to determine the teaching text, so as to improve the pertinence and efficiency of students' review.

[0090] In summary, the present invention significantly improves the accuracy, efficiency and practicality of classroom video voice transcription through the synergy of the document parsing module, the speech processing module and the transcription analysis module. The system has significant technical effects and practical application value, provides strong support for students' after-class review, further improves learning efficiency, and shows a wide range of application prospects.

[0091] Fig. 9 It is a software execution flow chart of an embodiment of the present application.

[0092] like Fig. 9 As shown in the figure, after the class is over, the speech transcription system starts running, the shooting device uploads the video data, and the teacher or teaching staff uploads the handout data. The speech transcription system performs audio segmentation and voiceprint feature clustering based on the video data. For example, the audio track is extracted from the classroom video and voice enhancement is performed to improve clarity and quality. The audio is segmented by the endpoint detection model, the system cuts the audio into paragraphs, and clusters each paragraph with small audio to ensure accurate identification of the speaker. Then, the speaker's facial features are extracted in combination with the class video to match the person information and obtain the speaker's identity information. The speech transcription system uses document parsing tools and named entity recognition (NER) technology for the uploaded handout data, and extracts short text lines to obtain candidate words, and then uses the large model to screen terms, select terms with high relevance to the course, exclude irrelevant and sensitive words, and generate terminology and pinyin dictionaries. Based on the obtained terminology data, the text transcribed by speech recognition is stimulated. For example, the vocabulary and its weight that need to be stimulated are determined through the terminology data, thereby significantly improving the recall rate and accuracy of professional terms in the transcribed text. Finally, the transcribed text is converted into pinyin form, and the pinyin of the terminology data is used to correct the incorrectly transcribed terms. The text features of each paragraph are extracted to mark the question text, and the non-question text is obtained as a teaching text with terminology markings to provide students with targeted review.

[0093] This application also proposes a device for identifying classroom data.

[0094] As an example, Fig.10As shown, the device for identifying classroom data includes: an acquisition module 1001, used to acquire classroom data, wherein the classroom data includes audio and video data and handout data; a first acquisition module 1002, used to obtain terminology data associated with the classroom based on the handout data; a second acquisition module 1003, used to identify the audio and video data based on the terminology data, and obtain text data associated with the audio and video data; an identification module 1004, used to correct the text data based on the terminology data, and perform teaching analysis and identification on the corrected text data to obtain a teaching identification result, wherein the teaching identification result includes a teaching text marked with terminology.

[0095] The application also proposes a computer-readable storage medium.

[0096] In this embodiment, a computer program is stored on a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned classroom data identification method are implemented.

[0097] Fig.11 A block diagram of an electronic device provided for an embodiment of the present application.

[0098] An embodiment of the present application provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned classroom data recognition method when executing the computer program.

[0099] like Fig.11 As shown, for ease of understanding, the embodiment of the present application shows a specific electronic device.

[0100] Electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0101] like Fig.11 As shown, the device includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the electronic device can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0102] Multiple components in the electronic device are connected to the I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a disk, an optical disk, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the electronic device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0103] The computing unit 1101 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1101 performs the various methods described above, such as the identification method of classroom data. For example, in some embodiments, the identification method of classroom data may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on an electronic device via ROM 1102 and / or a communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, the identification method of classroom data described above may be executed. Alternatively, in other embodiments, the computing unit 1101 may be configured to perform the identification method of classroom data in any other appropriate manner (e.g., by means of firmware).

[0104] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this application, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing in a suitable manner if necessary, and then stored in a computer memory.

[0105] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0106] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0107] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.

[0108] In addition, the terms "first", "second", etc. used in the embodiments of the present application are only used for descriptive purposes and should not be understood as indicating or implying relative importance, or implicitly indicating the number of technical features indicated in the present embodiment. Therefore, the features defined by the terms "first", "second", etc. in the embodiments of the present application can explicitly or implicitly indicate that at least one of the features is included in the embodiment. In the description of the present application, the word "multiple" means at least two or two or more, such as two, three, four, etc., unless otherwise clearly and specifically defined in the embodiments.

[0109] In this application, unless otherwise clearly specified or limited in the embodiments, the terms "installed", "connected", "connected" and "fixed" etc. appearing in the embodiments should be understood in a broad sense. For example, the connection can be a fixed connection, a detachable connection, or an integrated connection. It can be understood that it can also be a mechanical connection, an electrical connection, etc.; of course, it can also be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal connection of two elements, or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to the specific implementation situation.

[0110] In the present application, unless otherwise clearly specified and limited, a first feature being “above” or “below” a second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, a first feature being “above”, “above”, and “above” a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being “below”, “below”, and “below” a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.

[0111] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A method for identifying classroom data, characterized in that: The method comprises: Acquire classroom data, wherein the classroom data includes audio and video data and handout data; Obtaining terminology data associated with the class based on the lecture data; Identify the audio and video data based on the terminology data to obtain text data associated with the audio and video data; The text data is modified based on the terminology data, and the modified text data is subjected to teaching analysis and recognition to obtain a teaching recognition result, wherein the teaching recognition result includes a teaching text marked with terminology.

2. The method for identifying classroom data according to claim 1, characterized in that: The audio and video data includes audio data and video data, and the method further includes: Identify the audio data and determine speaker result data corresponding to the audio data; Based on the video data and the campus database, the speaker result data is matched to obtain speaker identity data corresponding to the speaker result data.

3. The method for identifying classroom data according to claim 1, characterized in that: The lecture data includes a lecture format, and obtaining terminology data associated with a class based on the lecture data includes: Parsing the lecture data based on a preset method, and extracting the parsing results based on a preset named entity model and a preset vocabulary library to obtain candidate vocabulary, wherein the preset method corresponds to the lecture format one by one; The candidate words are screened based on a pre-trained classification model to obtain terminology data associated with the classroom.

4. The method for identifying classroom data according to claim 2, characterized in that: The step of identifying the audio data and determining speaker result data corresponding to the audio data includes: Segmenting the audio data to obtain multiple audio segments; Extracting feature data of the plurality of audio segments, and performing clustering processing on the plurality of audio segment data based on the feature data; Speaker result data corresponding to the audio data is determined based on the clustering processing result.

5. The method for identifying classroom data according to claim 4, characterized in that: The step of segmenting the audio data to obtain a plurality of audio segments includes: Segmenting the audio data based on a preset endpoint model to obtain a plurality of first audio segments; Divide the plurality of first audio segments based on a preset length to obtain a plurality of second audio segments, wherein the first audio segment is composed of a plurality of the second audio segments; A plurality of audio segments are obtained based on the plurality of second audio segments.

6. The method for identifying classroom data according to claim 5, characterized in that: The clustering process of the plurality of audio segment data based on the feature data comprises: determining a category corresponding to each of the second audio segments; If the categories corresponding to all the second audio segments in the first audio segment are the same, the endpoints of the first audio segment are reserved; otherwise, the endpoints of the second audio segment are reserved; The clustering results are obtained based on the retained endpoints.

7. The method for identifying classroom data according to claim 4, characterized in that: The clustering processing result includes a plurality of audio endpoints, and determining speaker result data corresponding to the audio data based on the clustering processing result includes: Determine whether the audio segment between two adjacent endpoints belongs to the same speaker.

8. The method for identifying classroom data according to claim 1, characterized in that: The terminology data includes terminology pronunciation data, the text data includes text pronunciation data, and the modifying the text data based on the terminology data includes: In the case where the text pronunciation data corresponding to the text data is identical to the terminology pronunciation data corresponding to the terminology data, if the text data is different from the terminology data, the text data is corrected to the terminology data.

9. The method for identifying classroom data according to claim 1, characterized in that: The text data includes a plurality of text segment data, the teaching recognition result includes a teaching text, and the teaching analysis and recognition is performed on the corrected text data to obtain the teaching recognition result, including: Extracting feature data of a plurality of the text segment data, and determining attribute information of the text segment data based on the feature data corresponding to the text segment data, wherein the attribute information includes at least one of a question attribute and a statement attribute; In a case where the attribute information of the text segment data is a statement attribute and terminology data is included in the text segment, it is determined that the text segment data is a teaching text.

10. The method for identifying classroom data according to claim 2, characterized in that: The step of identifying the audio and video data based on the terminology data to obtain text data associated with the audio and video data includes: Determining a weight corresponding to the term data based on a preset speech recognition model; The audio and video data are identified based on the weights corresponding to the term data to obtain text data associated with the audio and video data.

11. The method for identifying classroom data according to claim 3, characterized in that: The method extracts the parsing results based on the preset named entity model and the preset vocabulary library to obtain candidate vocabulary, including: The parsing result is extracted based on a preset named entity model and a preset vocabulary library, and the text whose number of characters in the line text is less than the preset number of characters is determined as the candidate vocabulary.

12. A device for identifying classroom data, characterized in that: The device comprises: An acquisition module, used to acquire classroom data, wherein the classroom data includes audio and video data and handout data; A first obtaining module, for obtaining terminology data associated with the class based on the lecture data; A second obtaining module is used to identify the audio and video data based on the terminology data to obtain text data associated with the audio and video data; The recognition module is used to modify the text data based on the terminology data, and perform teaching analysis and recognition on the modified text data to obtain a teaching recognition result, wherein the teaching recognition result includes a teaching text marked with terminology.

13. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of the method described in any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 11 are implemented.