Pre-class education intelligent management method and device, storage medium and electronic equipment

Through the use of microphone arrays and speech recognition models, full-process recording and automatic generation of meeting minutes are achieved during pre-shift education at construction sites, solving the problems of identity verification and information recording distortion, and improving management efficiency and credibility.

CN120807238APending Publication Date: 2025-10-17CHENGDU LEISHU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510957187.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies in pre-shift education management at construction sites have problems such as rough identity verification, distorted information recording, and difficulty in process tracing. In addition, the voice transcription system has a low recognition rate for construction professional terminology, leading to identity fraud and misrecording or omission of key safety instructions.

Method used

A microphone array is used to collect voice data in real time, and combined with a speech recognition model and voiceprint clustering model, the speakers are identified and meeting minutes are generated, achieving full-process recording and automatic generation of meeting minutes.

Benefits of technology

It improves the credibility of identity verification, eliminates proxy signing and fraudulent use, improves the recognition rate of construction terminology, optimizes management efficiency, realizes full process recording and convenient review and tracing, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807238A_ABST
    Figure CN120807238A_ABST
Patent Text Reader

Abstract

The invention provides a pre-class education intelligent management method and device, a storage medium and electronic equipment, and the method comprises the steps: loading preaching materials corresponding to the pre-class education when the pre-class education starts, and enabling a preset microphone array to collect the voice data of the pre-class education process in real time; when the current voice data collected by the microphone array is obtained, determining a target person who speaks currently; inputting the current voice data into a voice recognition model to obtain current text data corresponding to the current voice data; marking a marking event corresponding to the current voice data on a time axis, wherein the marking event comprises the current text data, user information of the target person and material content matched with the current text data in the preaching material; and when the pre-class education is finished, generating a conference summary corresponding to the pre-class education based on each marking event. By applying the method provided by the invention, the whole pre-class education process can be recorded, the conference summary can be automatically generated, the examination and tracing are convenient, and the labor cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of conference intelligent management, in particular to a pre-shift education intelligent management method and device, a storage medium and an electronic device. BACKGROUND

[0002] In the field of construction site management, pre-shift education as a key link of safety disclosure and task deployment directly affects the construction execution efficiency. The current pre-shift education management of construction engineering generally adopts the mode of paper sign-in combined with manual recording, which has problems such as rough identity verification, distorted pre-shift education process information recording, and difficult process tracing. The traditional paper sign-in and pre-shift education process recording mode cannot adapt to the complex environment of the construction site, leading to frequent identity fraud, proxy sign-in, and proxy participation in pre-shift education. The existing voice transcription system has low recognition rate for construction professional terms, and the key safety instructions are often misrecorded or omitted. Although the existing technology can realize basic voice recording and image capture, the voice, image, and text data are stored in isolation, and there is a lack of intelligent management of the whole process. The supervision and examination rely on manual report sorting, which is time-consuming and difficult to verify the compliance of the process. SUMMARY

[0003] Therefore, the present application provides a pre-shift education intelligent management method and device, a storage medium and an electronic device. Through the method, the whole process of pre-shift education can be recorded, conference minutes can be automatically generated, examination and tracing are convenient, and labor costs are reduced.

[0004] The present application also provides a pre-shift education intelligent management device to ensure the implementation and application of the above method in practice.

[0005] A pre-shift education intelligent management method, the method comprising: At the beginning of the pre-shift education, loading the propaganda materials corresponding to the pre-shift education, and enabling a preset microphone array to collect voice data of the pre-shift education process in real time; When the current voice data collected by the microphone array is obtained, determining the target personnel of the current speech; Inputting the current voice data into a trained voice recognition model to obtain the current text data corresponding to the current voice data; Marking a marking event corresponding to the current voice data on a preset time axis, the marking event including the current text data, user information of the target personnel, and material content in the propaganda materials matching the current text data; At the end of the pre-shift education, generating a conference minutes corresponding to the pre-shift education based on each marking event on the time axis.

[0006] The method can further include: determining a position of the person speaking based on a sound source of the current voice data; controlling a preset camera device to capture an image of the position of the person speaking; matching the image of the person speaking with a face image in a preset database to determine the target person.

[0007] The method can further include: determining an initial sound source of the current voice data based on a time delay corresponding to each microphone in the microphone array; applying a preset MVDR adaptive beamformer to correct the initial sound source to obtain a corrected sound source; determining the position of the person speaking based on the corrected sound source.

[0008] The method can further include: inputting the current voice data into a preset voiceprint clustering model to identify at least one voiceprint in the current voice data through the voiceprint clustering model.

[0009] The method can further include: obtaining a preset sample text; collecting a voice data set related to the sample text, the voice data set including a plurality of standard voice data related to the sample text and text data corresponding to each standard voice data; using each standard voice data as training data and the text data corresponding to each standard voice data as a training target to train the voice recognition model to obtain an initial model; performing voice variation processing on each standard voice data to obtain voice variation data; optimizing the initial model using each voice variation data to obtain the trained voice recognition model.

[0010] The method can further include: generating summary content corresponding to each marked event based on text data, user information and material content included in each marked event on the timeline.

[0011] The method can further include: obtaining a preset meeting minutes template; Based on the meeting minutes template, extract the core content corresponding to each of the marked events; Fill the core content into the meeting minutes template to generate the meeting minutes corresponding to the pre-class education.

[0012] A pre-class education intelligent management device, comprising: A voice collection module, configured to load the preaching material corresponding to the pre-class education when the pre-class education starts, and enable a preset microphone array to collect voice data of the pre-class education process in real time; A determination module, configured to determine target personnel of current speech when the current voice data collected by the microphone array is obtained; A voice conversion module, configured to input the current voice data into a trained voice recognition model to obtain current text data corresponding to the current voice data; A marking module, configured to mark a marked event corresponding to the current voice data on a preset time axis, the marked event including the current text data, user information of the target personnel, and material content in the preaching material matching the current text data; A generation module, configured to generate a meeting minutes corresponding to the pre-class education based on each marked event on the time axis when the pre-class education ends.

[0013] A storage medium, comprising stored instructions, wherein the instructions, when executed, control a device in which the storage medium is located to perform the pre-class education intelligent management method described above.

[0014] An electronic device, comprising a memory, and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors to perform the pre-class education intelligent management method described above.

[0015] Compared with the prior art, the present application has the following advantages: The application provides a pre-shift education intelligent management method, which comprises the following steps: when the pre-shift education starts, loading the propaganda material corresponding to the pre-shift education, and enabling a preset microphone array to collect voice data of the pre-shift education process in real time; when the current voice data collected by the microphone array is obtained, determining the target personnel of the current speech; inputting the current voice data into a trained voice recognition model to obtain current text data corresponding to the current voice data; marking a marking event corresponding to the current voice data on a preset time axis, wherein the marking event comprises the current text data, user information of the target personnel, and material content in the propaganda material matched with the current text data; and when the pre-shift education ends, generating a meeting minutes corresponding to the pre-shift education based on each marking event on the time axis. By using the method provided by the application, the whole process of the pre-shift education can be recorded, the meeting minutes can be automatically generated, the review and tracing are convenient, and the labor cost is reduced. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only constitute a part of the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on the provided drawings.

[0017] Figure 1 A method flowchart of a pre-shift education intelligent management method provided by an embodiment of the present application; Figure 2 Another method flowchart of a pre-shift education intelligent management method provided by an embodiment of the present application; Figure 3 Still another method flowchart of a pre-shift education intelligent management method provided by an embodiment of the present application; Figure 4 A device structure diagram of a pre-shift education intelligent management device provided by an embodiment of the present application; Figure 5 An electronic device structure schematic diagram provided by an embodiment of the present application. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0019] In this application, the terms such as first and second are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between such entities or operations, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "includes a" does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0020] The present application can be used in a plurality of general or special computing device environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor devices, distributed computing environments including any of the above devices or devices, etc.

[0021] The embodiment of the present application provides a pre-class education intelligent management method, the method is applied to a processor, the method flow chart of the method is as shown in Figure 1 The specific steps are as follows: S1: when the pre-class education starts, loading the propaganda material corresponding to the pre-class education, and enabling the preset microphone array to collect the voice data of the pre-class education process in real time.

[0022] Before the pre-class education starts, the face images of all team members are collected, and all face images are saved to the database. Wherein, through establishing a dynamic living body detection model, the face collected by the infrared binocular camera is identified, and the face image of the successful identification is saved, and the infrared binocular camera supports face recognition under strong inverse light environment.

[0023] When the team member needs to participate in the pre-class education, face recognition clock-in can also be performed through the infrared binocular camera, and the face scanned by the clock-in is matched with the face image in the database at each face recognition clock-in, so as to realize identity verification of the team member participating in the pre-class education.

[0024] Before the pre-class education starts, the text of the propaganda material can be pre-input, and the propaganda material can be edited after being input. If the propaganda material is in voice format, it is converted into text format.

[0025] When the team initiates the pre-class education, the pre-class education is started, and the pre-input propaganda material corresponding to the pre-class education is automatically loaded. Wherein, the team member can initiate the pre-class education by sending a start instruction.

[0026] In the embodiment of the present application, the arrangement of each microphone in the microphone array is determined according to the spatial structure of the pre-class education process. During the pre-class education process, the microphone array collects voice data in real time. During the collection of voice data, the time stamp of the voice data collected each time is recorded.

[0027] S2: When the current voice data collected by the microphone array is obtained, the target person of the current speech is determined.

[0028] In the embodiment of the present application, after the microphone array collects voice data each time, the target person of the speech is located by an infrared binocular camera.

[0029] Specifically, after the current voice data is collected, the sound source of the current voice data is determined, and the position of the person speaking is determined through the sound source. The camera device (such as an infrared binocular camera) is controlled to turn to the position of the person, and the position of the person is photographed, so as to obtain the user image. The user image is matched with the face image in the database, and then the target person of the current speech is determined.

[0030] In the process of determining the sound source, since there is a time delay in the audio signals collected by each microphone in the microphone array, based on the time delay between each microphone in the microphone array, the initial sound source of the current voice data is determined, and then the MVDR adaptive beamformer is used to correct the initial sound source, so as to determine the position of the person according to the corrected sound source. The time delay between the microphones can be calculated by using the generalized cross correlation (GCC) method. According to the distance and time delay between the microphones, the initial sound source of the current voice data is determined through geometric relationship. After the initial sound source is determined, the sound source direction corresponding to the initial sound source is taken as input, and the MVDR (Minimum Variance Distortionless Response) formula is applied to calculate the MVDR weight, so as to perform adaptive beamforming in the sound source direction, realize signal enhancement and interference suppression, and thus realize correction and strengthening of the initial sound source, and determine the corrected sound source. According to the sound source direction of the corrected sound source, the position of the person speaking is determined.

[0031] Before photographing the position of the person, the camera device can also be controlled to determine whether there is a face in the position of the person. That is, after the camera device turns to the position of the person, it is detected whether there is a face contour in the lens range; if there is a face contour, the position of the person is photographed, and if there is no face contour, the position of the person is corrected until there is a face contour in the lens, and the corrected position of the person is photographed. The way to correct the face contour can be to rotate the lens of the camera device according to a preset rotation frequency and direction until a face appears in the lens.

[0032] Optionally, during the pre-shift education process, there may be multiple people speaking at the same time. When multiple people speak at the same time, there may be a large gap between the positions of different people, which may cause positioning failure when identifying the sound source to locate the position of the person. Therefore, before determining the position of the person, it can be further identified whether multiple voiceprints exist in the current voice data, that is, it is further determined whether multiple people speak at the same time.

[0033] The process of identifying whether multiple voiceprints exist in the current voice data is: inputting the current voice data into a voiceprint clustering model to identify at least one voiceprint in the current voice data through the voiceprint clustering model. The voiceprint clustering model is a neural network model obtained by training multiple voiceprint data, and can identify different voiceprints according to sound frequency, timbre and other related data in the sound data. When the current voice data has only one voiceprint, the speaker is positioned through the sound source of the current voice data; when the current voice data has multiple voiceprints, the current voice data is split into voice data corresponding to each voiceprint, and the positions of the speakers corresponding to each voiceprint are determined according to the sound source of the split voice data, and then the user images of each position are sequentially captured by the camera device to further identify the person.

[0034] The voiceprint clustering model can be a GMM-UBM (Gaussian Mixture-Universal Background) model. The GMM-UBM model is a speaker recognition model that combines Gaussian Mixture Model (GMM) and Universal Background Model (UBM). This model combines the advantages of GMM and UBM. In the identification process, UBM is first used to model the input voice signal, and then UBM is adjusted to the GMM of a specific speaker through adaptive technology. This adaptive process can better capture the individual characteristics of the speaker and improve the recognition accuracy.

[0035] S3: input the current voice data into the trained voice recognition model to obtain the current text data corresponding to the current voice data.

[0036] In the present application, the voice recognition model can accurately identify the industry-specific vocabulary related to pre-shift education in the voice data, and can also quickly identify some ambiguous words in the voice (for example: "20 centimeters" may be quickly connected to read similar "two s centimeters" sound, etc.).

[0037] In the embodiment of the present application, the voice recognition model is trained by a large amount of data to improve the accuracy of the voice recognition model in converting voice data into text data.

[0038] Reference Figure 2 The training process of the voice recognition model is as follows: S31: Obtain a pre-set sample text.

[0039] The standard text includes the full text of national standard specifications, local regulations, and industry-specific keywords.

[0040] S32: Collect a voice data set related to the standard text.

[0041] The voice data set includes a plurality of standard voice data related to the standard text and text data corresponding to each standard voice data.

[0042] The standard voice data in the voice data set can be automatically generated by an online AI specification voice data, or standard voice data of a standard accent recorded and uploaded by a technician.

[0043] S33: Each standard voice data is used as training data, the text data corresponding to each standard voice data is used as training target, the speech recognition model is trained, and an initial model is obtained.

[0044] Specifically, the standard voice data is input into the speech recognition model as training data, the recognition result output by the speech recognition model is obtained, the recognition result is matched and calculated with the text data, the loss function corresponding to the recognition result is obtained, and the model parameters of the speech recognition model are adjusted according to the loss function. The training data is input into the model again for training until the loss function corresponding to the obtained recognition result converges, and the initial model is obtained.

[0045] S34: Each standard voice data is processed to obtain variable voice data.

[0046] It should be noted that the variable voice processing refers to converting the standard accent voice data into voice data with at least one voice characteristic, each voice characteristic can include fast connected reading characteristic, connected reading variable voice characteristic, weak reading characteristic, and accent characteristic.

[0047] S35: Each variable voice data is applied to the initial model to optimize the initial model, and a trained speech recognition model is obtained.

[0048] It can be understood that after the standard voice data is processed, the speaking habits of different persons can be simulated, each variable voice data is input into the initial model as new training data, the initial model re-executes the training process corresponding to S33 above, and the model can recognize voice data with accent, fast speaking speed, or unclear speaking when performing voice recognition and text conversion, further improving the accuracy of the model.

[0049] In the embodiment of the present application, the speech recognition model obtained by the implementation process of S31-S35 is applied to process the speech data, which can improve the accuracy of speech recognition and processing.

[0050] S4: Mark the mark event corresponding to the current speech data on the preset time axis.

[0051] The mark event includes the current text data, the user information of the target person, and the material content in the propaganda material that matches the current text data.

[0052] During the pre-shift education process, other camera equipment can be used to shoot the entire pre-shift education process, and a time axis related to the pre-shift education is generated, which is extended as the pre-shift education time length.

[0053] During the pre-shift education process, if a person makes a speech, the relevant event of the current speaker can be marked on the time axis, which is a mark event.

[0054] Optionally, if the time axis related to the pre-shift education is not generated, a timestamp corresponding to each mark event is generated to generate a speaker time axis of the pre-shift education corresponding to each mark event.

[0055] Further, after generating the mark event, a summary content corresponding to each mark event can also be generated, that is, based on the text data, user information and material content of each mark event on the time axis, a summary content corresponding to each mark event is generated. The summary content is a summary of the text data, user information and material content of the propaganda material related to the text data in the mark event. After generating the summary content, the mark event related to the summary content on the time axis can be found through the summary content, so as to facilitate subsequent backtracking of the mark event.

[0056] S5: At the end of the pre-shift education, a meeting minutes corresponding to the pre-shift education is generated based on each mark event on the time axis.

[0057] Reference Figure 3 The process of generating the meeting minutes includes the following steps: S51: Obtain a pre-set meeting minutes template.

[0058] S52: Based on the meeting minutes template, extract the core content corresponding to each mark event.

[0059] The core content can be the summary content described above, or the content required to be extracted according to the template prompt.

[0060] Optionally, in addition to extracting core content through marked events, relevant core content can also be extracted through the clock-in records related to the pre-class education stored in the database, such as: clock-in personnel, number of participants, and meeting lateness.

[0061] S53: Fill the core content into the meeting minutes template to generate meeting minutes corresponding to the pre-class education.

[0062] In an embodiment of the present invention, AI can also be used to extract the core content of the pre-class education process and generate meeting minutes based on this core content. Similarly, if a meeting minutes template has not been set up in advance, the AI ​​can be fed with presentation materials and a timeline with various marked events, and the AI ​​will generate meeting minutes based on the input content.

[0063] In the present invention, after the meeting minutes are generated, the administrator can view all the data recorded in the meeting minutes online and can also check preset check items for evaluation.

[0064] An embodiment of the present invention provides an intelligent management method for pre-shift education. Before the start of pre-shift education, the lecture materials are entered, and the system automatically converts them into audio files and associates them with the designated work group. Workers sign in through facial recognition, and the system verifies their identities in real time and records the list of attendees. At the beginning of pre-shift education, the lecture audio associated with the work group is played, and the playback status is recorded simultaneously. The entire meeting is recorded, the meeting recording is transcribed into text, the speakers are distinguished, and the meeting scene is automatically captured. At the end of pre-shift education, it is automatically archived, and the system integrates the sign-in records, recording files, transcribed texts, and captured images. Based on the content of the meeting, meeting minutes are generated through AI, which automatically extracts information and generates structured meeting minutes.

[0065] The method provided by this invention improves the credibility of identity verification during pre-shift training sign-in, eliminating proxy signing and fraudulent use. Voice processing improves accuracy and recognition of construction terminology. Management efficiency is significantly improved, with full pre-shift training process recorded and meeting minutes automatically generated, simplifying review and tracing, and reducing labor costs.

[0066] and Figure 1 Corresponding to the above method, the embodiment of the present invention also provides a pre-class education intelligent management device for Figure 1 In the specific implementation of the method, the pre-class education intelligent management device provided by the embodiment of the present invention is applied to a processor, and its structural diagram is shown as follows Figure 4 As shown, specifically including: The voice collection module 601 is used to load the corresponding teaching materials of the pre-class education at the beginning of the pre-class education and enable the preset microphone array to collect the voice data of the pre-class education process in real time; The determining module 602 is configured to determine a target person of a current speech when current voice data collected by the microphone array is obtained. The voice conversion module 603 is configured to input the current voice data into a trained voice recognition model to obtain current text data corresponding to the current voice data. The marking module 604 is configured to mark a marking event corresponding to the current voice data on a preset time axis, wherein the marking event includes the current text data, user information of the target person, and material content in the presentation material that matches the current text data. The generating module 605 is configured to generate a meeting minutes corresponding to the pre-class education based on each marking event on the time axis when the pre-class education ends.

[0067] In the device provided by the embodiment of the application, the determining module 602 determines the target person of the current speech, and is specifically configured to: determine a position of a person of the current speech based on a sound source of the current voice data; control a preset camera device to capture the position of the person to obtain a user image; match the user image with a face image in a preset database to determine the target person.

[0068] In the device provided by the embodiment of the application, the determining module 602 determines the position of the person of the current speech based on the sound source of the current voice data, and is specifically configured to: determine an initial sound source of the current voice data based on a time delay corresponding to each microphone in the microphone array; apply a preset MVDR adaptive beamformer to correct the initial sound source to obtain a corrected sound source; and determine the position of the person of the current speech based on the corrected sound source.

[0069] In the device provided by the embodiment of the application, the device further includes: A voiceprint recognition module is configured to input the current voice data into a preset voiceprint clustering model to identify at least one voiceprint in the current voice data through the voiceprint clustering model.

[0070] In the device provided by the embodiment of the application, the device further includes: The training module is configured to obtain a preset text; collect a voice data set related to the text, the voice data set including a plurality of standard voice data related to the text and text data corresponding to each standard voice data; take each standard voice data as training data and the text data corresponding to each standard voice data as a training target, train the speech recognition model to obtain an initial model; perform voice variation processing on each standard voice data to obtain voice variation data; and apply each voice variation data to the initial model to optimize the initial model and obtain the trained speech recognition model.

[0071] The generation module 605 is further configured to, before the end of the pre-class education, generate, based on the text data, the user information and the material content contained in each marked event on the timeline, summary content corresponding to each marked event. The generation module 605 is further configured to, before the end of the pre-class education, generate, based on the text data, the user information and the material content contained in each marked event on the timeline, summary content corresponding to each marked event.

[0072] The generation module 605 is further configured to, before the end of the pre-class education, generate, based on the text data, the user information and the material content contained in each marked event on the timeline, summary content corresponding to each marked event. The generation module 605 is further configured to, before the end of the pre-class education, generate, based on the text data, the user information and the material content contained in each marked event on the timeline, summary content corresponding to each marked event.

[0073] The specific working processes of the modules of the pre-class education intelligent management device disclosed in the embodiments of the present application can be referred to the corresponding content of the pre-class education intelligent management method disclosed in the embodiments of the present application, which will not be described here.

[0074] The embodiments of the present application also provide a storage medium including stored instructions, wherein the instructions, when executed, control a device where the storage medium is located to perform the pre-class education intelligent management method.

[0075] The embodiments of the present application also provide an electronic device, a structure diagram of which is shown in FIG. 7. Figure 5 The electronic device includes a memory 701 and one or more instructions 702, wherein the one or more instructions 702 are stored in the memory 701 and configured to be executed by one or more processors 703 to perform the following operations: When the pre-class education starts, load the preaching material corresponding to the pre-class education and enable a preset microphone array to collect voice data of the pre-class education in real time; When the current voice data collected by the microphone array is obtained, determine the target personnel of the current speech; input the current voice data into a trained voice recognition model to obtain current text data corresponding to the current voice data; mark a mark event corresponding to the current voice data on a preset time axis, the mark event including the current text data, user information of the target person, and material content in the propaganda material that matches the current text data; at the end of the pre-class education, generate a meeting minutes corresponding to the pre-class education based on each mark event on the time axis.

[0076] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the system or system embodiments, since it is basically similar to the method embodiments, it is described more simply, and the related parts can be referred to the part of the method embodiments. The above-described system and system embodiments are only illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0077] The skilled person can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both.

[0078] In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical scheme. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0079] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for intelligent management of pre-class education, characterized in that: include: At the beginning of pre-class education, the corresponding teaching materials of the pre-class education are loaded, and the preset microphone array is enabled to collect the voice data of the pre-class education process in real time; When obtaining the current voice data collected by the microphone array, determining the target person who is currently speaking; Inputting the current voice data into a trained voice recognition model to obtain current text data corresponding to the current voice data; Marking a marking event corresponding to the current voice data on a preset time axis, wherein the marking event includes the current text data, the user information of the target person, and the material content in the presentation material that matches the current text data; At the end of the pre-class education, a meeting minutes corresponding to the pre-class education is generated based on each marked event on the timeline.

2. The method according to claim 1, characterized in that Determining the target person of the current speech includes: Determining the location of the person currently speaking based on the sound source of the current voice data; Controlling a preset camera device to capture the position of the person and obtain a user image; The user image is matched with the face images in a preset database to determine the target person.

3. The method according to claim 2, characterized in that The determining the position of the person currently speaking based on the sound source of the current voice data includes: Determining an initial sound source of the current voice data based on a time delay corresponding to each microphone in the microphone array; Applying a preset MVDR adaptive beamformer to modify the initial sound source to obtain a modified sound source; Based on the corrected sound source, the position of the person currently speaking is determined.

4. The method according to claim 3, characterized in that Also includes: The current voice data is input into a preset voiceprint clustering model to identify at least one voiceprint in the current voice data through the voiceprint clustering model.

5. The method according to claim 1, wherein The training process of the speech recognition model includes: Get the pre-set sample text; Collecting a speech data set related to the model text, wherein the speech data set includes a plurality of standard speech data related to the model text and text data corresponding to each standard speech data; Using each of the standard speech data as training data and the text data corresponding to each of the standard speech data as a training target, the speech recognition model is trained to obtain an initial model; Performing voice modification on each of the standard voice data to obtain voice modified voice data; The initial model is optimized using each of the voice-changed speech data to obtain the trained speech recognition model.

6. The method according to claim 1, characterized in that Before the end of the pre-class education, it also includes: Based on the text data, user information and material content included in each marked event on the timeline, summary content corresponding to each marked event is generated.

7. The method according to claim 1 or 6, characterized in that Generating the meeting minutes corresponding to the pre-class education based on each marked event on the timeline includes: Get pre-set meeting minutes templates; Based on the meeting minutes template, extract the core content corresponding to each marked event; Fill the core content into the meeting minutes template to generate the meeting minutes corresponding to the pre-class education.

8. An intelligent management device for pre-class education, characterized in that: include: The voice collection module is used to load the corresponding pre-class education materials at the beginning of the pre-class education and enable the preset microphone array to collect the voice data of the pre-class education process in real time; A determination module, configured to determine a target person who is currently speaking when obtaining current voice data collected by the microphone array; A speech conversion module, configured to input the current speech data into a trained speech recognition model to obtain current text data corresponding to the current speech data; A marking module, configured to mark a marking event corresponding to the current voice data on a preset timeline, wherein the marking event includes the current text data, the user information of the target person, and the content of the presentation material that matches the current text data; A generating module is used to generate meeting minutes corresponding to the pre-class education based on each marked event on the timeline when the pre-class education is completed.

9. A storage medium, characterized in that: The storage medium includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the intelligent management method for pre-class education according to any one of claims 1 to 7.

10. An electronic device, characterized in that: The system comprises a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors to execute the one or more instructions to execute the pre-class education intelligent management method according to any one of claims 1 to 7.