An adaptive and active information interaction education method

By building a database of learning environment and attention status, and using machine learning models to adaptively switch audio and video presentation methods, the problem of poor information acquisition during the learning process is solved, and learning efficiency and concentration are improved.

CN116012202BActive Publication Date: 2025-09-09ZHEJIANG LAB +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211489023.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2025-09-09
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

Existing technologies fail to comprehensively consider the presentation form of learning content, interference factors in the learning environment, and the learner's attention state, resulting in poor information acquisition and low learning efficiency during the learning process.

Method used

By building a database of learning environment, content and learner status, using machine learning models to identify learning environment and attention status, adaptively switching audio and video presentation methods, generating subtitles and text descriptions, and optimizing information transmission.

Benefits of technology

It realizes adaptive information transmission in different environments and attention states, improves learning efficiency and concentration, and enhances learning effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116012202B_ABST
    Figure CN116012202B_ABST
Patent Text Reader

Abstract

The present invention relates to an adaptive active information interaction education method, which adaptively changes the playback mode according to the learner's environment and attention situation, such as converting a video into audio through content recognition and automatic voice generation, or automatically generating subtitles and pictures from the audio to generate a video. Specifically, the device can identify environmental interference characteristics based on sound signals, vibration signals, acceleration sensors, and light sensors, and judge the level of environmental noise, sound scenes, ambient light conditions, and the shaking state of the device; at the same time, the device can identify the learner's attention state and automatically push appropriate content formats to maximize the degree of attention concentration. In addition, the device can automatically generate text describing the video and image content by recognizing the video and image content, and further synthesize audio playback; the device can also recognize the audio content and automatically generate text and pictures for visual presentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent education, and in particular relates to an adaptive active information interaction education method. Background Art

[0002] Adaptive teaching methods can effectively expand the scenarios in which learners acquire knowledge, improve learners' attention levels, and thus improve learning efficiency.

[0003] With the miniaturization of smart tools and the increase in network bandwidth, the scenarios in which ordinary learners can use smart tools have been greatly expanded. Learning and educational devices that adapt to various environments can more effectively assist learners in acquiring, processing, and internalizing information, thereby assisting and improving cognitive efficiency.

[0004] The transmission of educational information should be as close as possible to the learner's environment and attention. From the perspective of enhancing information reception, the learning content should be effectively presented through multimodal, multi-channel audio, video, images, and text.

[0005] Learners may learn in distracting environments, such as on the subway or while exercising, which can hinder information reception. This can also lead to inefficient learning due to objective factors or temporary attention deficits. Proactively and adaptively adjusting information interaction methods can improve cognitive efficiency. Several solutions are currently available, such as navigation and search for video content, push notifications for related knowledge points, audio visualization, and voice commentary for video and VR content. In other words, optimizing content presentation formats from a cognitive efficiency perspective can improve learning efficiency.

[0006] However, current learning resources and course content lack comprehensive research and holistic technical solutions for information enhancement, resulting in poor information acquisition and susceptibility to interference during the learning process. For example, in terms of learning content format, existing technologies can adaptively optimize the learning order and visual presentation of knowledge points through deep learning recommendation algorithms based on behavioral data to enhance learners' cognitive efficiency. However, they fail to comprehensively consider factors that affect cognitive efficiency, such as the multi-sensory presentation format of the learning content, interference factors in the learning environment, and the learner's attention state. In terms of learner attention state, existing technologies have proposed machine learning-based face pose estimation methods and attention assessment methods that integrate gaze estimation to improve the accuracy of learner attention assessment. They then enhance learners' attention through additional multi-sensory interactive experiences preset in the learning content. However, they fail to fully consider interference factors in the learning environment and fail to achieve active and automatic enhancement of the original information in the learning content. In summary, there is no technical solution in the existing technology that comprehensively considers the content presentation format, learning environment, and the concentration level of a specific person, nor is there any research on technologies that adaptively and actively enhance the original information of the learning content [1,2,3].

[0007] Therefore, the adaptive active information interaction education method involved in the present invention has important theoretical and practical significance. Summary of the Invention

[0008] To solve the above problems, the present invention discloses an adaptive active information interaction education method, which adaptively switches audio to video presentation according to the environmental noise level and sound scene characteristics; and adaptively switches video to audio presentation according to the environmental light conditions and the shaking state of the device.

[0009] To achieve the above object, the technical solution of the present invention is as follows:

[0010] An adaptive active information interaction education method comprises the following steps:

[0011] Step 1: Build a learning environment database, divide each learning environment into scene categories, and construct a training data set; input the training data set into a learning environment recognition model based on machine learning, use the category information as supervision information of the learning environment recognition model, and train the learning environment recognition model;

[0012] Build a learning content database, establish an association between the image and sound signal of each learning content, and build a learning content recognition model; input the learning content into the learning content recognition model, use the image as the supervision information of the learning content recognition model, and train the learning content recognition model;

[0013] Construct a video database and provide a text description of the video content; input the video into a machine learning-based video recognition model, using the text description as supervision information for the video recognition model to train the video recognition model;

[0014] Construct an image database and provide text descriptions of the image content; input the images into a machine learning-based image recognition model, using the text descriptions as supervision information for the image recognition model to train the image recognition model;

[0015] Build a learner status database, label the learner's attention according to the learner's status, build a training dataset, use attention as supervision information, and train a machine learning-based attention recognition model;

[0016] Step 2: Detect the sound, light, and vibration of the learner's learning environment, and input the detection results into the trained learning environment recognition model to obtain the learner's scene. Based on the learner's scene, the sound, light, and vibration of the learning environment, adaptively switch the auxiliary presentation mode from audio to video or from visual information to audio.

[0017] At the same time, it identifies the learner's attention state in real time, combines the learner's scene, the sound and light of the learning environment, and the shaking of the learner's device, gives prompt information, and automatically switches to the appropriate scene to maximize the learner's concentration.

[0018] Furthermore, the adaptive switching of the auxiliary presentation mode from audio to video is achieved by:

[0019] Load the speech recognition engine, recognize the audio content learned by the learner, automatically generate subtitles through speech recognition, and display the subtitles on the screen of the device held by the learner; input the audio content learned by the learner into the trained learning content recognition model, obtain the picture corresponding to the audio content learned by the learner, and display it on the screen of the device held by the learner to provide auxiliary visual presentation of the audio content.

[0020] Furthermore, adaptively switching the visual information to the audio auxiliary presentation mode is achieved by the following methods:

[0021] Load the trained video recognition model or image recognition model, and automatically generate a text description of the video or image content by recognizing the video or image content learned by the learner; load the text recognition engine to recognize the text in the visual information learned by the learner; integrate the text description and the recognized text in the visual information learned by the learner on the timeline, and then use the speech synthesis engine to synthesize the audio content of the text for playback.

[0022] Furthermore, the learner's current state is identified and input into the trained attention recognition model to obtain the learner's attention concentration level; when the learner's attention concentration level is lower than the set threshold, a corresponding prompt message is given.

[0023] Furthermore, the learner's state includes facial rotation and movement trajectory, eye movement state, gaze direction estimation, facial expression, head posture, upper body posture, and screen and peripheral operation behavior.

[0024] Furthermore, the learner's attention state is identified in real time, and prompt information is given based on the learner's scene, the sound and light of the learning environment, and the shaking of the learner's device. This is achieved through the following steps:

[0025] (1) Construct a probabilistic dependency relationship between the presentation format of learning content, the learner's situation, and the learner's attention concentration;

[0026] (2) Establish the joint conditional probability density function of attention on both content form and learning environment;

[0027] (3) Push appropriate content formats to maximize attention concentration. Under the Bayesian framework, calculate the posterior probability value P(x|S1) of the environmental features for scene 1, and calculate the posterior probability value P(x|S2) of the environmental features for scene 2, where x represents the feature value, S1 represents scene 1, and S2 represents scene 2. Set weights W1 and W2, and W1+W2=1. Adjust W1 so that the scene selection result is:

[0028]

[0029] Iteratively change the weight W1 = W1 + ΔW to minimize the attention level. In the i-th iteration, detect the change in attention level ΔL and obtain the gradient g = ΔL / ΔW. Let the next iteration be W1 = W1-α*g

[0030] Here, α is a small positive number, set to 0.001 to 0.1.

[0031] Furthermore, the ambient sound signal is sensed by a microphone.

[0032] Furthermore, the vibration signal of the device is sensed by a vibration sensor and an acceleration sensor.

[0033] Furthermore, the ambient light is sensed by a light sensor.

[0034] Furthermore, the weight W1 is adjusted between 0.3 and 0.7.

[0035] The beneficial effects of the present invention are:

[0036] The adaptive active information interaction education method of the present invention can adaptively enhance the information transmission effect according to environmental interference and learner's attention state. The method is clever and novel and has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a flow chart of one of the adaptive active information interactive education methods of the present invention.

[0038] Figure 2 It is a flow chart of a method for identifying environmental interference features of the present invention.

[0039] Figure 3 It is a flow chart of the auxiliary visual presentation of audio content in the learning environment of the present invention.

[0040] Figure 4 It is a flow chart of the auxiliary audio presentation of video content in the learning environment of the present invention.

[0041] Figure 5 This is a flow chart of the adaptive selection of content presentation format in the present invention. DETAILED DESCRIPTION

[0042] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0043] like Figure 1 As shown, the present invention is an adaptive active information interactive education method, comprising the following steps:

[0044] Step 1: Build a learning environment database, divide each learning environment into scene categories, and construct a training dataset; input the training dataset into a learning environment recognition model based on machine learning, and use the category information as supervision information for the learning environment recognition model to train the learning environment recognition model;

[0045] Build a learning content database, establish an association between the image and sound signal of each learning content, and build a learning content recognition model; input the learning content into the learning content recognition model, use the image as the supervision information of the learning content recognition model, and train the learning content recognition model;

[0046] Construct a video database and provide a text description of the video content; input the video into a machine learning-based video recognition model, using the text description as supervision information for the video recognition model to train the video recognition model;

[0047] Construct an image database and provide text descriptions of the image content; input the images into a machine learning-based image recognition model, using the text descriptions as supervision information for the image recognition model to train the image recognition model;

[0048] Build a learner's status database, label the learner's attention according to the learner's status, build a training dataset, use attention as supervision information, and train a machine learning-based attention recognition model.

[0049] Among them, such as Figure 2 As shown, to detect the sound and light in the learner's learning environment, as well as the shaking of the learner's device, a microphone is used to sense the ambient sound signals, a vibration sensor and accelerometer are used to sense the device's vibration signals, and a light sensor is used to sense the ambient light. The various types of time signals collected by the sensors are sampled and quantized to form time series signals. The time series signals are then analyzed using the statistical operators (maximum, minimum, mean, variance, maximum, minimum, mean, variance of first-order differences, maximum, minimum, mean, variance of second-order differences). A short-time Fourier transform is applied to the time series signals to construct features in the frequency domain, including the Mel-frequency cepstrum parameters of the sound signal. These frequency-domain features, along with the statistical operators (maximum, minimum, mean, variance, maximum, minimum, mean, variance of first-order differences, maximum, minimum, mean, variance of second-order differences), are used as environmental signal features to establish a learning environment recognition model based on machine learning. The environmental signal features are classified into two scenario types: one is suitable for learning audio content, that is, the reception of visual information is greatly disturbed; the other is suitable for learning visual content, that is, the reception of audio information is greatly disturbed.

[0050] The scene here refers to the learner's surrounding environment and the learner's own state. For example, on a noisy subway, if the learner holds the device steady (such as the learner is sitting in a seat), then the ideal learning mode may be video transmission; on the contrary, if the learner's device is shaking violently or in a vertical position, the ideal learning mode may be audio transmission. In general, when the ambient sound field is noisy, it should be preferred to present the learning materials in video form. On the contrary, when the light is too strong or too weak, or the equipment is unstable, it should be preferred to present the learning materials in audio form.

[0051] like Figure 3 As shown, a learning content database is constructed and auxiliary visual presentation of audio content is performed.

[0052] Building a learning content database involves establishing a correspondence between audio and images. Specifically, we can extract a large amount of correspondence between audio and a series of images from massive amounts of video data. By selecting and enhancing the images, we can generate high-quality images, thereby establishing a correspondence between audio and images. After machine learning training, we can obtain a model that generates images from audio.

[0053] Thanks to the development of search engine technology, it has become possible to obtain a large amount of video and image information containing indexes and text information. This training and learning content recognition model can automatically generate highly matching image information from text information, thereby providing visual display and illustrations.

[0054] like Figure 4 As shown, a video database and an image database are constructed separately, and audio-assisted presentation of visual information is performed. Manual text descriptions of the content, including targets and scenes, are performed. Video and image data are input into a machine learning model, and the text descriptions are used as supervision information for the model to train the content understanding model. As previously described, a large number of correspondences between images and text, as well as between videos and text, are already available from existing search engines. Therefore, the video recognition model can ultimately generate an accurate text description of the video from the video information; similarly, the image recognition model can ultimately generate an accurate text description of the image from the image information.

[0055] In addition, a learner status database is constructed, categorizing attention states from 1 to 7, with 1 representing the most focused. The attention level is estimated and labeled based on facial rotation and movement trajectories, eye movement status, gaze direction estimation, facial expression, head posture, upper body posture, and screen and peripheral operation behaviors. A training dataset is constructed, using attention as supervision information to train a machine learning-based attention recognition model.

[0056] Step 2: Detect the sound, light, and vibration of the learner's learning environment, and input the detection results into the trained learning environment recognition model to obtain the learner's scene. Based on the learner's scene, the sound, light, and vibration of the learning environment, adaptively switch the auxiliary presentation mode from audio to video or from visual information to audio.

[0057] At the same time, it identifies the learner's attention state in real time, combines the learner's scene, the sound and light of the learning environment, and the shaking of the learner's device, gives prompt information, and automatically switches to the appropriate scene to maximize the learner's concentration.

[0058] The adaptive switching of the auxiliary presentation mode from audio to video is achieved by:

[0059] Load the speech recognition engine, recognize the audio content learned by the learner, automatically generate subtitles through speech recognition, and display the subtitles on the screen of the device held by the learner; input the audio content learned by the learner into the trained learning content recognition model, automatically recognize the sound event category and scene category, obtain the corresponding keyword label, retrieve the corresponding picture from the picture resource library, and display it on the screen of the device held by the learner; perform auxiliary visual presentation of the audio content.

[0060] Adaptive switching of visual information to audio-assisted presentation is achieved through the following methods:

[0061] The video or image content understanding model is loaded. By recognizing the video or image content, a text description describing the theme of the video or image content is automatically generated. A text recognition engine is loaded to perform text recognition on the text in the visual information, generating text related to the subtitles in the video. Finally, speech recognition is used to convert the audio information in the video into text. These three parts of text are summarized, integrated, and synthesized into a comprehensive text, which is then synthesized into audio content using a speech synthesis engine. Therefore, when video is not suitable for conveying information to learners, this audio content can serve as an alternative form of information, allowing learners to receive audio materials of the corresponding content.

[0062] like Figure 5 As shown, the learner's current state is identified, and the trained attention recognition model is input to obtain the learner's attention concentration level; when the learner's attention concentration level is lower than the set threshold, a corresponding prompt message is given. This is achieved through the following steps:

[0063] Construct a probabilistic dependency relationship between the presentation format of learning content, the learner's scenario, and the learner's attention concentration; establish a joint conditional probability density function for attention on both the content format and the learning environment; and based on the scenario characteristics, self-optimize the decision weights and recommend the appropriate content presentation format to maximize attention concentration. For example, in a Bayesian framework, calculate the posterior probability value P(x|S1) of the environmental characteristics for scenario 1, and calculate the posterior probability value P(x|S2) of the environmental characteristics for scenario 2, where x represents the feature value, S1 represents scenario 1, and S2 represents scenario 2. Set weights W1 and W2, and W1+W2=1. Adjust W1 so that the scenario selection result is:

[0064] Select scenario 1, when W1*P(x|S1)>W*P(x|S2)

[0065] Select Scenario 2, Other

[0066] Generally, the weight W1 is adjusted between 0.3 and 0.7, and W2 = 1-W1.

[0067] After this initial weight is determined, further iterations can be performed during subsequent operation based on the user's actual attention level detected by the system, allowing the system to automatically adapt and optimize the user experience. The parameter W is a parameter that can be iteratively optimized, and possible optimization methods include gradient descent. The weight W1 = W1 + ΔW is iteratively changed to minimize the attention level. In the i-th iteration, the change in attention level ΔL is detected, and the gradient g = ΔL / ΔW is obtained. The next iteration is set to W1 = W1 - α*g.

[0068] Here α is the learning rate, which is usually a small positive number, usually set to 0.001 to 0.1, to ensure that the overall parameters do not fluctuate violently.

[0069] In summary, the adaptive active information interaction education method proposed by the present invention can adaptively enhance the information transmission effect according to environmental interference and the learner's attention state. The method is clever and novel and has good application prospects. Compared with the existing video playback systems, although some systems already have the function of converting video into audio for listening, this function is still very simple. It only directly removes the picture information of the video and provides audio information to the reader. On the one hand, this switching requires the user to complete it manually. On the other hand, this system has no way to present the information contained in the video picture information in the audio, which may cause the user to miss a lot of useful information. The method of the present invention comprehensively considers the information of pictures, subtitles, and audio. On the other hand, it can adaptively switch various playback modes for users, so it can transmit learning information more intelligently.

[0070] The adaptive learning method of the present invention is not only suitable for learners to learn online, but also suitable for teachers or learners to prepare learning materials for a given environment in advance.

[0071] It should be noted that the above content merely illustrates the technical idea of ​​the present invention and cannot be used to limit the scope of protection of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications all fall within the scope of protection of the claims of the present invention.

Claims

1. An adaptive active information interaction education method, characterized in that: The steps include: Step 1: Build a learning environment database, divide each learning environment into scene categories, and construct a training data set; input the training data set into a learning environment recognition model based on machine learning, use the category information as supervision information of the learning environment recognition model, and train the learning environment recognition model; Build a learning content database, establish associations between images and sound signals for each learning content, and build a learning content recognition model; Input the learning content into the learning content recognition model, use the image as the supervision information of the learning content recognition model, and train the learning content recognition model; Construct a video database and provide a text description of the video content; input the video into a machine learning-based video recognition model, using the text description as supervision information for the video recognition model to train the video recognition model; Construct an image database and provide text descriptions of the image content; input the images into a machine learning-based image recognition model, using the text descriptions as supervision information for the image recognition model to train the image recognition model; Build a learner status database, label the learner's attention according to the learner's status, build a training dataset, use attention as supervision information, and train a machine learning-based attention recognition model; Step 2: Detect the sound, light, and vibration of the learner's learning environment, and input the detection results into the trained learning environment recognition model to obtain the learner's scene. Based on the learner's scene, the sound, light, and vibration of the learning environment, adaptively switch the auxiliary presentation mode from audio to video or from visual information to audio. At the same time, it identifies the learner's attention state in real time, combines the learner's scene, the sound and light of the learning environment, and the shaking of the learner's device, gives prompt information, and automatically switches to the appropriate scene to maximize the learner's concentration. Specifically, it includes: Real-time recognition of the learner's attention state, combined with the learner's scene, the sound and light of the learning environment, and the shaking of the learner's device, provides prompt information. This is achieved through the following steps: (1) Constructing a probabilistic dependency relationship between the presentation format of learning content, the learner’s situation, and the learner’s attention level; (2) Establishing the joint conditional probability density function of attention on both content form and learning environment; (3) Push appropriate content formats to maximize attention concentration. Under the Bayesian framework, calculate the posterior probability value P(x|S1) of the environmental features for scene 1 and the posterior probability value P(x|S2) of the environmental features for scene 2, where x represents the feature value, S1 represents scene 1, and S2 represents scene 2. Set weights W1 and W2, and W1+W2=1. Adjust W1 so that the scene selection result is: When W1*P(x|S1) > W2*P(x|S2), select scenario 1; Otherwise, select scenario 2; Iteratively change the weight W1=W1+ΔW to minimize the attention level. In the i-th iteration, detect the attention level change ΔL and obtain the gradient g=ΔL / ΔW. Let the next iteration be W1=W1-α*g, where α is a small positive number set to 0.001 to 0.

1.

2. The adaptive active information interactive education method according to claim 1, characterized in that: The adaptive switching of the auxiliary presentation mode from audio to video is achieved by: Load the speech recognition engine, recognize the audio content learned by the learner, automatically generate subtitles through speech recognition, and display the subtitles on the screen of the device held by the learner; input the audio content learned by the learner into the trained learning content recognition model, obtain the picture corresponding to the audio content learned by the learner, and display it on the screen of the device held by the learner to provide auxiliary visual presentation of the audio content.

3. The adaptive active information interactive education method according to claim 1, characterized in that: Adaptive switching of visual information to audio-assisted presentation is achieved through the following methods: Load the trained video recognition model or image recognition model, and automatically generate a text description of the video or image content by recognizing the video or image content of the learner; A text recognition engine is loaded to recognize the text in the visual information learned by the learner; after integrating the text description and the recognized text in the visual information learned by the learner on the timeline, the audio content of the text is synthesized through the speech synthesis engine for playback.

4. The adaptive active information interactive education method according to claim 1, characterized in that: Identify the learner's current state, input the trained attention recognition model, and obtain the learner's attention concentration level; When the learner's concentration level is lower than the set threshold, corresponding prompt information will be given.

5. The adaptive active information interactive education method according to claim 1, characterized in that: The learner's state includes facial rotation and movement trajectory, eye movement state, gaze direction estimation, facial expression, head posture, upper body posture, screen and peripheral operation behavior.

6. The adaptive active information interactive education method according to claim 1, characterized in that: The microphone senses the sound signals of the environment.

7. The adaptive active information interactive education method according to claim 1, characterized in that: The vibration signal of the device is sensed by the vibration sensor and acceleration sensor.

8. The adaptive active information interactive education method according to claim 1, characterized in that: Sense the ambient light through the light sensor.

9. The adaptive active information interactive education method according to claim 1, characterized in that: The weight W1 is adjusted between 0.3 and 0.7.

Citation Information

Patent Citations

  • Information processing method, intelligent terminal and storage medium

    CN115118812A

  • Learning device and method

    JP2002366925A