A focus-driven education conference grouping interaction method and system
By acquiring audio information from remote trainers and generating subtitles, and combining this with the focus of local users, the group interaction members are dynamically adjusted. This solves the problem of low interaction efficiency caused by random grouping in existing technologies, and achieves more efficient group interaction in educational meetings.
Patent Information
- Application Number
- CN202510212257.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Existing conferencing systems in educational settings use random grouping, ignoring the differences in users' comprehension levels of the course, resulting in low interaction efficiency and affecting teaching effectiveness and user learning interest.
By acquiring audio information from remote trainers, generating subtitles using a speech recognition model, and dynamically adjusting group interaction members based on the attention levels of local users, attention-driven group interaction can be achieved.
It improves the interactive efficiency and teaching effectiveness of educational meetings, ensures accurate judgment of user focus, and enhances learning interest and training results.
Smart Images

Figure CN120087931B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interactive technology in educational conferences, and more specifically, to a focus-driven group interaction method and system for educational conferences. Background Technology
[0002] In today's education sector, with the acceleration of globalization and the continuous updating of educational philosophies, the demand for remote communication and collaboration is showing an increasing trend. Remote teaching, online seminars, and other educational activities are becoming more frequent, leading to the increasingly widespread application of conferencing systems in educational settings. However, existing conferencing systems are gradually revealing many shortcomings when applied to educational scenarios.
[0003] In the crucial step of enabling group interaction, existing conferencing systems often employ simple random grouping. This approach completely ignores the significant differences in course comprehension among individual users. Different users have varying levels of focus, meaning some may have a deep understanding of the content while others may only have a superficial grasp. Random grouping can easily result in groups with members of varying knowledge levels and learning progress, hindering effective communication and leading to low interaction efficiency. Such group interaction not only fails to achieve the desired teaching effect but may also negatively impact user interest and motivation, significantly reducing the quality and efficiency of educational activities. In conclusion, existing conferencing systems have serious problems with group interaction methods in educational scenarios, necessitating a focus-driven approach to educational conferencing group interaction to better meet the needs of remote communication and collaboration in the education field and improve the quality and effectiveness of teaching and learning. Summary of the Invention
[0004] The purpose of this invention is to provide a focus-driven group interaction method and system for educational meetings to improve the above-mentioned problems.
[0005] To achieve the above objectives, the embodiments of this application provide the following technical solutions:
[0006] On the one hand, embodiments of this application provide a focus-driven group interaction method for educational meetings, the method comprising:
[0007] Obtain first information, which includes audio information corresponding to the remote trainer;
[0008] The first information is sent to the speech recognition model to obtain the speech recognition result;
[0009] Based on the speech recognition result, the corresponding subtitle information is displayed in a preset area, which is a preset area on the local terminal device;
[0010] The second information is determined based on the subtitle information, and the second information includes the focus level of each local user.
[0011] The interaction member information is determined based on the focus of each local user, and the interaction member information includes at least two local users.
[0012] The local users who interact between the various local terminals are determined based on the interaction member information.
[0013] Secondly, embodiments of this application provide a focus-driven group interaction system for educational meetings, the system comprising:
[0014] The acquisition module is used to acquire first information, which includes audio information corresponding to the remote trainer.
[0015] The first processing module is used to send the first information to the speech recognition model to obtain the speech recognition result;
[0016] The second processing module is used to display corresponding subtitle information in a preset area according to the speech recognition result, wherein the preset area is a preset area on the local terminal device;
[0017] The third processing module is used to determine second information based on the subtitle information, the second information including the focus of each local user;
[0018] The fourth processing module is used to determine the interaction member information based on the focus of each local user, wherein the interaction member information includes at least two local users;
[0019] The interaction module is used to determine the local terminal users who interact between the various local terminals based on the interaction member information.
[0020] Thirdly, embodiments of this application provide a focus-driven group interaction device for educational meetings, the device including a memory and a processor. The memory stores a computer program; the processor executes the computer program to implement the steps of the focus-driven group interaction method for educational meetings described above.
[0021] Fourthly, embodiments of this application provide a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described attention-driven group interaction method for educational meetings.
[0022] The beneficial effects of this invention are as follows:
[0023] This invention performs speech recognition on the audio information of remote trainers to obtain speech recognition results. The speech recognition results are then used to generate corresponding subtitle information in a preset area. The local user's attention span is determined by the time the local user spends watching the subtitles, thus obtaining a second piece of information. When group interaction is required in the classroom, the interaction members are determined based on the second piece of information. This avoids the problem of low interaction efficiency and poor interaction effect between local users caused by the random grouping method used in the prior art.
[0024] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of the attention-driven group interaction method for educational meetings as described in an embodiment of the present invention.
[0027] Figure 2 This is a schematic diagram of the structure of the attention-driven group interaction system for educational meetings as described in an embodiment of the present invention.
[0028] Figure 3 This is a schematic diagram of the structure of the attention-driven group interaction device for educational meetings as described in an embodiment of the present invention.
[0029] The diagram is labeled as follows: 901, Acquisition Module; 902, First Processing Module; 903, Second Processing Module; 904, Third Processing Module; 905, Fourth Processing Module; 906, Interaction Module; 9011, First Acquisition Unit; 9012, First Processing Unit; 9013, Second Processing Unit; 9014, Third Processing Unit; 9021, Preprocessing Unit; 9022, Fifth Processing Unit; 9023, Sixth Processing Unit; 9041, Third Acquisition Unit; 9042, Seventh Processing Unit; 9043, Eighth Processing Unit. 800, Ninth Processing Unit; 90121, Second Acquisition Unit; 90122, Fourth Processing Unit; 90123, Construction Unit; 90124, Training Unit; 90421, Tenth Processing Unit; 90422, Eleventh Processing Unit; 90423, Twelfth Processing Unit; 90424, Thirteenth Processing Unit; 800, Attention-Driven Educational Conference Group Interactive Device; 801, Processor; 802, Memory; 803, Multimedia Component; 804, I / O Interface; 805, Communication Component. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0031] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. Example 1:
[0032] This embodiment provides a focus-driven group interaction method for educational meetings. It can be understood that a scenario can be set up in this embodiment, such as a scenario in a distance education meeting where a trainer on the remote end trains multiple local users.
[0033] See Figure 1The figure shows that the method includes steps S1, S2, S3, S4, S5, and S6, which specifically include:
[0034] Step S1: Obtain first information, which includes audio information corresponding to the remote trainer;
[0035] In this step, the remote trainer is the course instructor. The instructor collects audio information by setting up a microphone array in their training room and then transmits the collected audio information to the local users to be trained via a computer network.
[0036] Step S1 further includes steps S11, S12, S13, and S14, which specifically include:
[0037] Step S11: Obtain the trainee's current location information;
[0038] Step S12: Send the trainee's current location information to the trained motion trajectory prediction model to obtain the first prediction result;
[0039] Step S12 further includes steps S121, S122, S123, and S124, which specifically include:
[0040] Step S121: Obtain historical training video information;
[0041] Step S122: Determine the position of the trainee in each frame of the historical training video information to obtain historical trajectory information;
[0042] In this step, the center of the blackboard is selected as the reference point using a manual marking method, and the position coordinates of the trainee are determined in each frame to obtain historical trajectory information.
[0043] Step S123: Construct a training set based on the historical trajectory information;
[0044] Step S124: Determine the corresponding motion feature information based on the data in the training set, and train the motion trajectory prediction model to obtain the trained motion trajectory prediction model.
[0045] In this step, the motion features include velocity features, acceleration features, and turning angle. The velocity feature can be obtained by dividing the difference between the coordinates of adjacent positions of the trainee by the time interval; the acceleration can be obtained by calculating the change in velocity; and the turning angle can be obtained by the change in direction of adjacent positions. It should be noted that the present invention does not limit the network structure used in the motion trajectory prediction model, including but not limited to long short-term memory networks and recurrent neural networks.
[0046] In this embodiment, by acquiring the trainee's historical training video information, the trainee's behavioral habits are understood, and the corresponding motion feature information is determined based on the historical trajectory information. The motion trajectory prediction model is then trained to obtain the trained motion trajectory prediction model, thereby enabling real-time prediction of the trainee's trajectory.
[0047] Step S13: Determine the motion trajectory of the microphone array based on the first prediction result to obtain trajectory information;
[0048] Step S14: Control the microphone array to move according to the trajectory information and collect the first information in real time.
[0049] In remote education training scenarios, it's common for trainers at the remote end to move around the blackboard while writing, drawing, or highlighting key points. This movement poses a significant challenge to audio capture, making it difficult to determine the optimal placement of the microphone array.
[0050] The ideal installation location for a microphone array typically needs to ensure clear sound pickup of the trainee while avoiding interference from ambient noise. However, when trainees frequently move around the blackboard, it's difficult to maintain the optimal distance and angle regardless of the microphone array's location. For example, if the microphone array is installed in a fixed wall position, the sound travels a longer distance as the trainee moves to the other side of the blackboard and may be blocked by the blackboard or other objects, leading to weakened and distorted sound signals.
[0051] Since microphone arrays cannot capture high-quality audio information, this directly impacts subsequent speech recognition and caption generation. Speech recognition models are prone to errors when processing unclear or unstable audio signals, potentially leading to inaccurate captions displayed in the designated area. These errors may include spelling mistakes, semantic misunderstandings, and incomplete sentences.
[0052] For local users, they primarily rely on displayed subtitles to better understand the trainer's explanations. However, incorrect subtitles can severely disrupt their learning process, reducing their comprehension and mastery of the training content, thus diminishing the training effectiveness for local users. Local users may misunderstand due to incorrect subtitles, miss key knowledge points, or need to spend more time and effort guessing the trainer's true intentions, which undoubtedly affects their learning motivation and efficiency. In conclusion, in remote education and training meetings, the challenges of microphone array placement caused by the trainer's movement around the blackboard, and the resulting decline in audio acquisition quality and incorrect subtitle information, have a significant negative impact on the training effectiveness for local users. Therefore, solving this problem and improving audio acquisition quality and subtitle accuracy has become key to enhancing the effectiveness of distance education and training. This invention abandons the traditional approach (setting the microphone array in a fixed position) and instead predicts the trainee's movement trajectory, allowing the microphone array to move along with the trainee's path. This ensures the microphone array maintains the same distance from the trainee, neither too far nor too close. This not only eliminates concerns about sound acquisition during movement, allowing instructors to focus more on the lesson, but also ensures high-quality audio acquisition for subsequent speech recognition model processing. This effectively solves the problem of significantly negatively impacting the training effect on local users due to decreased audio acquisition quality.
[0053] Step S2: Send the first information to the speech recognition model to obtain the speech recognition result;
[0054] Step S2 further includes steps S21, S22, and S23, which specifically include:
[0055] Step S21: Preprocess the first information to obtain preprocessed first information;
[0056] In this step, the preprocessing of audio information includes pre-emphasis, frame-by-frame windowing, and noise reduction.
[0057] Step S22: Use a genetic algorithm to select speech features from the preprocessed first information to obtain the filtered speech features;
[0058] In this step, the specific process of feature selection using a genetic algorithm is as follows: set the initial parameters of the genetic algorithm, input all speech features to be selected into the genetic algorithm, set the objective function, encode each speech feature into a corresponding chromosome, initialize the population and set the initial number to 10, calculate the applicability value of 10 feature combinations; filter combinations through selection, crossover and mutation; the stopping condition is to determine whether the maximum number of iterations has been reached. If the maximum number of iterations has been reached, output the selected speech features; otherwise, recalculate the applicability value of the feature combinations and filter again through selection, crossover and mutation.
[0059] It should be noted that, in this embodiment, the selected speech features are short-time average energy and spectral mean.
[0060] Step S23: Recognize the first information based on the filtered speech features to obtain the speech recognition result.
[0061] Step S3: Display the corresponding subtitle information in a preset area according to the speech recognition result. The preset area is a preset area on the local terminal device.
[0062] In this step, the local user's terminal device includes, but is not limited to, mobile phones and tablets. The subtitle information can be displayed in a preset area on the mobile phone or tablet. By generating corresponding subtitle information, users can be prevented from missing the trainer's voice, and the user's impression of the course can be deepened from both audio and subtitle modalities, thereby improving the training effect.
[0063] Step S4: Determine the second information based on the subtitle information, the second information including the focus level of each local user;
[0064] Step S4 further includes steps S41, S42, S43, and S44, which specifically include:
[0065] Step S41: Obtain the user's eye movement trajectory information;
[0066] In this step, an eye-tracking device is used to record the eye movement trajectory of the local user when viewing the terminal device.
[0067] Step S42: Determine the time when the user's eye movement trajectory is within the preset area based on the user's eye movement trajectory information to obtain third information;
[0068] Step S42 further includes steps S421, S422, S423, and S424, which specifically include:
[0069] Step S421: Perform keyword analysis on the subtitle information to obtain keyword information;
[0070] In this step, the keyword analysis technique is used to extract keywords, which is a well-known technical solution in the field of science, so it will not be described in detail here.
[0071] Step S422: Determine the area where the keyword appears within the preset area based on the keyword information to obtain the key area;
[0072] Step S423: Divide the preset area into ordinary areas and key areas according to the key areas;
[0073] Step S424: Determine the time when the user's eye movement trajectory is located in the normal area and the time when it is located in the key area based on the user's eye movement trajectory information.
[0074] In this embodiment, a common problem in practice is the issue of local users losing focus. Typically, even when local users are distracted, their gaze may still linger on a preset area. This interferes with accurately judging the user's level of concentration. To more accurately calculate the local user's focus, this invention divides the preset area into ordinary areas and key areas using keywords. In training courses, different words have varying degrees of relevance to the topic. Words with high relevance to the topic are identified as key areas, and these key areas are given higher weight. This is because when users are truly focused on the training content, they are more likely to pay attention to key areas, as the content presented in these areas is crucial for understanding the course topic. However, when users are distracted, it is difficult for them to constantly focus on the constantly changing key areas. Compared to ordinary areas, the content in key areas is more targeted and important; it is unrealistic to expect users to pay attention to key areas as frequently when distracted as when focused.
[0075] This approach effectively avoids misjudging users as highly focused when they are distracted. Traditional focus assessment methods may rely solely on simple factors such as whether a user is looking at the screen, neglecting the user's true level of engagement with the content. The method of this invention is more in-depth and detailed, evaluating user focus from the perspective of content importance. This method more accurately reflects the true focus state of local users during training, making the calculated local user focus more accurate and thus improving the accuracy of subsequent grouping.
[0076] Step S43: Calculate the fixation time ratio information based on the third information and the total training time;
[0077] Step S44: Determine the second information based on the gaze time ratio information.
[0078] In this embodiment, the attention span of each local user can be effectively quantified based on the gaze time ratio information, thereby increasing its interpretability.
[0079] Step S5: Determine the interaction member information based on the focus of each local user, wherein the interaction member information includes at least two local users;
[0080] Determining interaction member information based on each local user's focus level specifically involves: obtaining threshold information; classifying local users based on the threshold information and their focus levels to obtain classification results, which include a first category of users and a second category of users. The first category of users corresponds to users with focus levels greater than the threshold information, and the second category of users corresponds to users with focus levels less than the threshold information; constructing an affinity graph based on the first and second categories of users; and determining the interaction member information based on the affinity graph. In this step, the focus level of each local user can determine their level of acceptance of the class. Users are divided into a first category and a second category based on their focus level. The first category represents users with good class acceptance, and the second category represents users with poor class acceptance. Taking any user in the first category as the central node, the affinity between each user in the second category and that user is calculated. Users in the second category with affinity levels higher than the affinity threshold are grouped together for interaction, which can effectively improve interaction efficiency. Furthermore, by grouping users with good class acceptance with those with poor class acceptance together, users with good class acceptance can receive secondary training, thereby improving the overall training effect of the course.
[0081] It should be noted that the method for constructing the intimacy graph is as follows: construct a first node and a second node based on users included in the first category of users and users included in the second category of users; obtain fourth information, which includes the number of chat records between each user in the first category and each user in the second category; calculate the intimacy between each user in the first category and each user in the second category based on the total number of chat records for each user in the first category and the fourth information; and establish the connection relationship between the first node and the second node based on the intimacy to obtain the intimacy graph.
[0082] Step S6: Determine the local terminal users who interact between the various local terminals based on the interaction member information. Example 2:
[0083] like Figure 2As shown, this embodiment provides a focus-driven interactive group system for educational meetings. The system includes an acquisition module 901, a first processing module 902, a second processing module 903, a third processing module 904, a fourth processing module 905, and an interaction module 906, specifically including:
[0084] The acquisition module 901 is used to acquire first information, which includes audio information corresponding to the remote trainer.
[0085] The first processing module 902 is used to send the first information to the speech recognition model to obtain the speech recognition result;
[0086] The second processing module 903 is used to display corresponding subtitle information in a preset area according to the speech recognition result, wherein the preset area is a preset area on the local terminal device;
[0087] The third processing module 904 is used to determine second information based on the subtitle information, the second information including the focus of each local user;
[0088] The fourth processing module 905 is used to determine interaction member information based on the focus of each local user, wherein the interaction member information includes at least two local users;
[0089] The interaction module 906 is used to determine the local terminal users who interact between the various local terminals based on the interaction member information.
[0090] In one specific embodiment of this disclosure, the acquisition module 901 includes a first acquisition unit 9011, a first processing unit 9012, a second processing unit 9013, and a third processing unit 9014, specifically comprising:
[0091] The first acquisition unit 9011 is used to acquire the trainer's current location information;
[0092] The first processing unit 9012 is used to send the trainee's current location information to the trained motion trajectory prediction model to obtain a first prediction result;
[0093] The second processing unit 9013 is used to determine the motion trajectory of the microphone array based on the first prediction result and obtain trajectory information.
[0094] The third processing unit 9014 is used to control the movement of the microphone array according to the trajectory information and to collect the first information in real time.
[0095] In one specific embodiment of this disclosure, the first processing unit 9012 further includes a second acquisition unit 90121, a fourth processing unit 90122, a construction unit 90123, and a training unit 90124, specifically including:
[0096] The second acquisition unit 90121 is used to acquire historical training video information;
[0097] The fourth processing unit 90122 is used to determine the position of the trainee in each frame of the historical training video information to obtain historical trajectory information;
[0098] Construction unit 90123 is used to construct a training set based on the historical trajectory information;
[0099] The training unit 90124 is used to determine the corresponding motion feature information based on the data in the training set, and to train the motion trajectory prediction model to obtain the trained motion trajectory prediction model.
[0100] In one specific embodiment of this disclosure, the first processing module 902 further includes a preprocessing unit 9021, a fifth processing unit 9022, and a sixth processing unit 9023, specifically including:
[0101] Preprocessing unit 9021 is used to preprocess the first information to obtain preprocessed first information;
[0102] The fifth processing unit 9022 is used to select speech features from the preprocessed first information using a genetic algorithm to obtain the filtered speech features;
[0103] The sixth processing unit 9023 is used to identify the first information based on the filtered speech features to obtain a speech recognition result.
[0104] In one specific embodiment of this disclosure, the third processing module 904 further includes a third acquisition unit 9041, a seventh processing unit 9042, an eighth processing unit 9043, and a ninth processing unit 9044, specifically comprising:
[0105] The third acquisition unit 9041 is used to acquire the user's eye movement trajectory information;
[0106] The seventh processing unit 9042 is used to determine the time when the user's eye movement trajectory is located within the preset area based on the user's eye movement trajectory information, and to obtain third information;
[0107] The eighth processing unit 9043 is used to calculate the gaze time ratio information based on the third information and the total training time;
[0108] The ninth processing unit 9044 is used to determine the second information based on the gaze time ratio information.
[0109] In one specific embodiment of this disclosure, the seventh processing unit 9042 further includes a tenth processing unit 90421, an eleventh processing unit 90422, a twelfth processing unit 90423, and a thirteenth processing unit 90424, specifically comprising:
[0110] The tenth processing unit 90421 is used to perform keyword analysis on the subtitle information to obtain keyword information;
[0111] The eleventh processing unit 90422 is used to determine the area where the keyword appears in the preset area based on the keyword information, and obtain the key area;
[0112] The twelfth processing unit 90423 is used to divide the preset area into a normal area and a key area according to the key area;
[0113] The thirteenth processing unit 90424 is used to determine the time when the user's eye movement trajectory is located in the normal area and the time when it is located in the key area, respectively, based on the user's eye movement trajectory information.
[0114] It should be noted that the specific methods by which each module performs operations in the system described in the above embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here. Example 3:
[0115] Corresponding to the above method embodiments, this embodiment also provides a focus-driven educational conference group interaction device. The focus-driven educational conference group interaction device described below and the focus-driven educational conference group interaction method described above can be referred to each other.
[0116] Figure 3 This is a block diagram illustrating a focus-driven interactive group work device 800 for educational meetings, according to an exemplary embodiment. Figure 3 As shown, the attention-driven interactive education conference group 800 may include a processor 801 and a memory 802. The attention-driven interactive education conference group 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0117] The processor 801 controls the overall operation of the attention-driven interactive educational conference group 800 to complete all or part of the steps in the attention-driven interactive educational conference group method described above. The memory 802 stores various types of data to support the operation of the attention-driven interactive educational conference group 800. This data may include, for example, instructions for any application or method operating on the attention-driven interactive educational conference group 800, and application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 805 is used for wired or wireless communication between the attention-driven educational conference group interaction device 800 and other devices. Wireless communication includes Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, and an NFC module.
[0118] In an exemplary embodiment, the attention-driven educational conference group interaction device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the attention-driven educational conference group interaction method described above.
[0119] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the attention-driven educational conference group interaction method described above. For example, the computer-readable storage medium may be the memory 802 including the program instructions described above, which may be executed by the processor 801 of the attention-driven educational conference group interaction device 800 to complete the attention-driven educational conference group interaction method described above. Example 4:
[0120] Corresponding to the above method embodiments, this embodiment also provides a readable storage medium. The readable storage medium described below can be referred to in conjunction with the attention-driven group interaction method for educational meetings described above.
[0121] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the attention-driven group interaction method for educational meetings described in the above method embodiments.
[0122] Specifically, the readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other readable storage medium capable of storing program code.
[0123] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0124] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A focus-driven group interaction method for educational meetings, characterized in that, include: Obtain first information, which includes audio information corresponding to the remote trainer; The first information is sent to the speech recognition model to obtain the speech recognition result; Based on the speech recognition result, the corresponding subtitle information is displayed in a preset area, which is a preset area on the local terminal device; The second information is determined based on the subtitle information, and the second information includes the focus level of each local user. The interaction member information is determined based on the focus of each local user, and the interaction member information includes at least two local users. The local users who interact between the various local terminals are determined based on the interaction member information. Obtaining the first information includes: Obtain the trainer's current location information; The trainee's current location information is sent to the trained motion trajectory prediction model to obtain a first prediction result; The motion trajectory of the microphone array is determined based on the first prediction result, and trajectory information is obtained; The microphone array is moved according to the trajectory information, and the first information is collected in real time. The second information determined based on the subtitle information includes: Obtain the user's eye movement trajectory information; Based on the user's eye movement trajectory information, the time when the user's eye movement trajectory is located within the preset area is determined to obtain third information. The preset area includes a normal area and a key area. Calculate the gaze time ratio information based on the third information and the total training time; The second information is determined based on the gaze duration ratio information.
2. The attention-driven group interaction method for educational meetings according to claim 1, characterized in that, Sending the trainee's current location information to the trained motion trajectory prediction model includes: Access historical training video information; The position of the trainee is determined in each frame of the historical training video information to obtain historical trajectory information; A training set is constructed based on the historical trajectory information; Based on the data in the training set, determine the corresponding motion feature information, and train the motion trajectory prediction model to obtain the trained motion trajectory prediction model.
3. The attention-driven group interaction method for educational meetings according to claim 1, characterized in that, The first information is sent to the speech recognition model to obtain the speech recognition result, including: The first information is preprocessed to obtain the preprocessed first information; The genetic algorithm is used to select speech features from the preprocessed first information to obtain the filtered speech features; The first information is identified based on the filtered speech features to obtain a speech recognition result.
4. The attention-driven group interaction method for educational meetings according to claim 1, characterized in that, Determining the time when the user's eye movement trajectory is within the preset area based on the user's eye movement trajectory information includes: Keyword analysis was performed on the subtitle information to obtain keyword information; Based on the keyword information, determine the area where the keyword appears within a preset area to obtain the key area; The preset area is divided into a normal area and a key area based on the key area; Based on the user's eye movement trajectory information, determine the time when the user's eye movement trajectory is located in the normal area and the time when it is located in the key area.
5. A focus-driven interactive group system for educational meetings, characterized in that, include: The acquisition module is used to acquire first information, which includes audio information corresponding to the remote trainer. The first processing module is used to send the first information to the speech recognition model to obtain the speech recognition result; The second processing module is used to display corresponding subtitle information in a preset area according to the speech recognition result, wherein the preset area is a preset area on the local terminal device; The third processing module is used to determine second information based on the subtitle information, the second information including the focus of each local user; The fourth processing module is used to determine the interaction member information based on the focus of each local user, wherein the interaction member information includes at least two local users; An interaction module is used to determine the local terminal users who interact between the various local terminals based on the interaction member information. The acquisition module includes: The first acquisition unit is used to acquire the trainer's current location information; The first processing unit is used to send the trainee's current location information to the trained motion trajectory prediction model to obtain a first prediction result; The second processing unit is used to determine the motion trajectory of the microphone array based on the first prediction result and obtain trajectory information. The third processing unit is used to control the movement of the microphone array according to the trajectory information and to collect the first information in real time. The third processing module includes: The third acquisition unit is used to acquire the user's eye movement trajectory information; The seventh processing unit is used to determine the time when the user's eye movement trajectory is located within the preset area based on the user's eye movement trajectory information, and to obtain third information, wherein the preset area includes a normal area and a key area; The eighth processing unit is used to calculate the gaze time ratio information based on the third information and the total training time; The ninth processing unit is used to determine the second information based on the gaze time ratio information.
6. The attention-driven interactive group meeting system for educational meetings according to claim 5, characterized in that, The first processing unit includes: The second acquisition unit is used to acquire historical training video information; The fourth processing unit is used to determine the position of the trainee in each frame of the historical training video information to obtain historical trajectory information; The construction unit is used to construct a training set based on the historical trajectory information; The training unit is used to determine the corresponding motion feature information based on the data in the training set, and to train the motion trajectory prediction model to obtain the trained motion trajectory prediction model.
7. The attention-driven group interaction system for educational meetings according to claim 5, characterized in that, The first processing module includes: A preprocessing unit is used to preprocess the first information to obtain preprocessed first information; The fifth processing unit is used to select speech features from the preprocessed first information using a genetic algorithm to obtain the filtered speech features. The sixth processing unit is used to identify the first information based on the filtered speech features to obtain a speech recognition result.
Citation Information
Patent Citations
Voice information receiving method and system based on robot and terminal equipment
CN109961781A
Classroom behavior feedback teaching system based on AIGC
CN117934228A
Webinar watch-party
US20230247067A1