Recitation assisting method and device, electronic equipment and computer readable storage medium
The extended reality device solves the problem of low efficiency of traditional recitation by automatically identifying obstructions in recitation and generating associative images, achieving an efficient and vivid recitation experience and improving the user's memory effect.
Patent Information
- Application Number
- CN202511115597.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-11
AI Technical Summary
Traditional paper-based recitation and electronic device-assisted recitation have problems such as low recitation efficiency, inability to obtain timely feedback, and easily interrupted recitation rhythm.
The extended reality device collects the user's real-time recitation voice, automatically identifies whether the recitation is obstructed, and generates associative image displays. It uses a virtual display screen to show associative images, stimulating the user's active associative memory.
It improves recitation efficiency, reduces fatigue, enhances memory effect, improves memory efficiency through multi-sensory linkage, and avoids interruption of recitation rhythm.
Smart Images

Figure CN120632141A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of extended reality technology, and specifically to a method, device, electronic device, and computer-readable storage medium for assisting recitation. Background Art
[0002] When users memorize texts or words using traditional paper-based methods, they experience a dull experience and a lack of timely feedback. Related technologies rely on electronic devices like learning machines to assist users in memorization. However, when users forget a word, they must actively interact with the device before the correct content is prompted. This can easily disrupt the recitation rhythm and lacks contextual stimulation, resulting in low memorization efficiency. Summary of the Invention
[0003] The embodiments of the present application provide a method, device, electronic device and computer-readable storage medium for assisting recitation, which can automatically identify whether a user's recitation is obstructed, and provide the user with situational stimulation through associative images when recitation is obstructed, thereby improving the recitation effect.
[0004] In a first aspect, an embodiment of the present application provides an auxiliary recitation method applied to an extended reality device, the method comprising: Obtaining a first recitation text and collecting real-time recitation voice; By comparing the real-time recitation voice and the first recitation text, detecting whether the current recitation is blocked; When the current recitation is blocked, determining the blocked text from the first recitation text, and generating an associative image according to the blocked text; The associative image is displayed using a virtual display screen of the extended reality device.
[0005] In a second aspect, an embodiment of the present application provides an auxiliary recitation device, applied to an extended reality device, the device comprising: A voice collection module is used to obtain the first recitation text and collect real-time recitation voice; An obstruction detection module, configured to detect whether the current recitation is obstructed by comparing the real-time recitation voice with the first recitation text; an image generating module, configured to, when the current recitation is blocked, determine the blocked text from the first recitation text and generate an associative image based on the blocked text; An image display module is used to display the associative image using a virtual display screen of the extended reality device.
[0006] In a third aspect, an embodiment of the present application further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps in the above-mentioned assisted recitation method are implemented.
[0007] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above-mentioned auxiliary recitation method are implemented.
[0008] In a fifth aspect, embodiments of the present application further provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in the embodiments of the present application.
[0009] The embodiments of the present application have the following beneficial effects: The extended reality device can collect the user's real-time recitation voice and compare the real-time recitation voice with the first recitation text, so as to automatically determine whether the current recitation is blocked. When the current recitation is blocked, the blocked text can be determined, and an associative image can be generated based on the blocked text, and the associative image can be displayed using the virtual display screen of the extended reality device. In this way, it can automatically identify whether the current recitation is blocked without the user's active operation, which can avoid interrupting the user's recitation rhythm; the associative image corresponding to the blocked text can be displayed, so as to help the user continue to recite through the associative image; vivid associative images can attract the user's attention more than boring text, can reduce fatigue during the recitation process, increase the fun of recitation, and can also inspire users to actively associate images with recitation texts, and transform memory from passive reception to active construction; and associative images as visual information can help users form multi-sensory linkage, which is easier to leave a deep impression in the brain than single text stimulation, thereby effectively improving the user's memory efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in this application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0011] Figure 1 This is a schematic diagram of the steps of the auxiliary recitation method provided by an embodiment of the present application; Figure 2 This is a module diagram of an auxiliary recitation system provided by an embodiment of the present application; Figure 3 1 is a schematic structural diagram of a recitation assistance device provided in one embodiment of the present application; Figure 4It is a structural diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0012] The following will be combined with the drawings in this application to clearly and completely describe the technical solutions in this application. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0013] In one embodiment, Figure 1 As shown, a method for assisting recitation is provided. Although a logical order is shown in the step diagram, in some cases, the steps shown or described can be performed in a different order than shown in the figure. Specifically, the method for assisting recitation can be applied to an extended reality device.
[0014] Extended Reality (XR) technology uses computers to combine the real and the virtual to create a virtual environment that allows human-computer interaction. XR includes Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). XR devices include mobile handheld terminals such as mobile phones and tablets, heads-up displays such as heads-up displays in cars, and near-eye display terminals worn on the human head, such as glasses and helmet-shaped displays. In the embodiments of this application, the hardware structure of the XR device is explained using a display in the form of optical see-through glasses, namely XR glasses.
[0015] The extended reality glasses may include a frame body, namely the temples and the frame, to support the extended reality glasses to be worn on the human head; may include an optical display module mainly composed of a micro display screen, optical lenses and optical waveguide sheets to display virtual content to the user; may include an audio module mainly composed of a microphone and a speaker to collect and play sounds; may include sensors mainly composed of cameras, gyroscopes, barometers and infrared light transmitters and receivers to collect information data related to the human body, the glasses body and the external environment; may include an integrated processor with a microcontroller unit (MCU) or a central processing unit (CPU) as the core to perform data processing and calculation; may include a circuit board, flexible or rigid, to connect other electronic components to form an electronic circuit system, which is generally placed in the inner cavity of the glasses frame and the temples; may also include a battery power supply.
[0016] The extended reality device may also include an eye tracker that can detect the direction of the user's gaze. The extended reality device may also have a camera that captures the user's eyes, analyzing the user's gaze through the image of the user's eyes to determine the user's gaze direction.
[0017] The extended reality device may also include a virtual display screen. This virtual display screen is not a physical screen in the traditional sense, but rather a simulated display generated through technological means. It utilizes light, images, and other elements to create a user-perceived display area that can display various virtual information, images, videos, and interactive interfaces. The virtual display screen of the extended reality device can have a sufficiently large field of view (FOV). The FOV determines the area covered by the virtual screen in the user's field of view. A larger FOV allows the user to experience a wider virtual space.
[0018] It should be noted that the order of description of the following embodiments does not limit the priority order of the embodiments.
[0019] according to Figure 1 The auxiliary recitation method shown in the figure includes at least steps S110 to S140, which are described in detail as follows: In step S110 , a first recitation text is obtained, and real-time recitation speech is collected.
[0020] The first recitation text may be any text to be recited, such as a word, a poem, or an article.
[0021] In one embodiment, the first recitation text may be manually input by the user, or may be transmitted by the user to the augmented reality device through other electronic devices, or may be downloaded from a server.
[0022] In one embodiment, the first recitation text may be text on a paper book or other electronic device. The extended reality device may capture an image of the first recitation text using a camera and extract the first recitation text from the image using optical character recognition (OCR) technology.
[0023] When the user recites the first recitation text, the microphone of the extended reality device can collect the user's real-time recitation voice. After collecting the user's real-time recitation voice, the real-time recitation voice can be processed such as noise reduction to obtain a clear human voice.
[0024] In step S120, whether the current recitation is blocked is detected by comparing the real-time recitation voice with the first recitation text.
[0025] After the real-time recitation voice is collected, the real-time recitation voice can be converted into a target text, and by comparing the target text with the first recitation text, it can be determined whether the current recitation is blocked.
[0026] In one embodiment, it may be that when the pause duration of the real-time recitation voice characterizing the user reciting the first recitation text reaches a pause duration threshold, it is determined that the current recitation is blocked. The pause duration threshold can be set according to actual needs.
[0027] In one embodiment, it may be determined that the current recitation is obstructed when the real-time recitation voice representation shows that the user makes an error or omission in reciting the first recitation text.
[0028] In step S130 , when the current recitation is blocked, the blocked text is determined from the first recitation text, and an associative image is generated according to the blocked text.
[0029] When the current recitation is blocked, the blocked text may be determined from the first recitation text.
[0030] When the pause duration of the first recitation text represented by the real-time recitation voice characterization user reaches the pause duration threshold, the blocked text may be the text that should be recited after the pause. For example, if the pause duration of the user reciting "Bright moonlight in front of the bed" reaches the pause duration threshold, the "I doubt it is frost on the ground" following "Bright moonlight in front of the bed" may be determined as the blocked text.
[0031] When the user recites the first text to be recited in real time and an error or omission occurs, the text of the error or omission can be determined as an obstructed text. For example, when reciting "Quiet Night Thoughts", the user mistakenly recites "the moonlight before the bed" as "the moon and frost before the bed", then "the moonlight before the bed" or "the moonlight" can be determined as an obstructed text. For example, when reciting "Quiet Night Thoughts", the user directly recites "looking up at the bright moon" after reciting "the moonlight before the bed", and omits "I suspect it is frost on the ground", then the omitted "I suspect it is frost on the ground" can be determined as an obstructed text.
[0032] In one embodiment, the extended reality device can have a built-in text-based image model. When blocked text is input into the text-based image model, the model can output an associative image corresponding to the blocked text. The associative image is a visual representation of the blocked text. For example, if the blocked text is "the moon shines brightly in front of the bed," the associative image could be an image of the moon outside the window casting its light onto the bed inside.
[0033] In one embodiment, an extended reality device can communicate with another electronic device that has a built-in text-based graph model. The extended reality device can transmit the obstructed text to the other electronic device, which can then generate an associative image corresponding to the obstructed text based on the text-based graph model and return the associative image to the extended reality device.
[0034] A text-based graph model is an AI model that generates images based on input text. This model can be developed through adversarial training, supervised training, or deep learning. Examples include the Hunyuan-DiT model, UniT2IXL, Kolors, or other models.
[0035] In one embodiment, the same blocked text may have multiple associated images, and the multiple associated images may form a video, an animated image, or a short video, etc. The multiple associated images may form a story segment, a plot, etc.
[0036] For example, if a user is blocked from reciting the word "apple", the associated image can be an apple tree. For example, if a user is blocked from reciting "a volcano is erupting", the associated image can be multiple images, including a panoramic view of the eruption, a close-up of the lava jet, and an image of the lava flow. For example, if a user is blocked from reciting historical content, multiple images corresponding to a time chain scene can be generated. The multiple images can represent the various historical events that occurred in chronological order according to the historical content. For example, if a user is blocked from reciting text content, multiple continuous images can be generated according to the chapters of the text to form a "plot chain" learning experience.
[0037] In step S140, the associative image is displayed using the virtual display screen of the extended reality device.
[0038] The extended reality device includes a virtual display screen, through which associative images can be displayed.
[0039] Optionally, when displaying the associative image, the blocked text can also be displayed, and the speaker of the extended reality device can output the audio corresponding to the blocked text, thereby realizing multi-sensory linkage of images, text and audio to deepen the user's memory of the blocked text.
[0040] By adopting the technical solution of the embodiment of the present application, the extended reality device can collect the user's real-time recitation voice and compare the real-time recitation voice with the first recitation text, so as to automatically determine whether the current recitation is blocked. When the current recitation is blocked, the blocked text can be determined, and an associative image can be generated based on the blocked text, and the associative image can be displayed using the virtual display screen of the extended reality device; in this way, it is possible to automatically identify whether the current recitation is blocked without the user's active operation, which can avoid interrupting the user's recitation rhythm; the associative image corresponding to the blocked text can be displayed, so as to help the user continue to recite through the associative image; vivid associative images can attract the user's attention more than boring text, can reduce fatigue during the recitation process, increase the fun of recitation, and can also inspire the user to actively associate the image with the recitation text, and transform memory from passive reception to active construction; and associative images as visual information can help users form multi-sensory linkage, which is easier to leave a deep impression in the brain than single text stimulation, thereby effectively improving the user's memory efficiency.
[0041] On the basis of the above technical solution, as an embodiment, the method may further include: when an error is detected in the current recitation, giving an error prompt; when an omission is detected in the current recitation, giving an omission prompt.
[0042] If an error is detected in the user's current recitation based on the real-time recitation voice, an error prompt can be provided. The error prompt can be a voice prompt, a visual prompt, or a vibration prompt. For example, if the user mistakenly recites "bedside moonlight" instead of "bedside moon frost" when reciting "Quiet Night Thoughts", the user can be notified of the error by a voice prompt, or by a virtual display screen displaying the text "Current recitation error". Alternatively, the user can be notified of the error by a preset vibration.
[0043] If the user's current recitation is detected to have missed a part based on the real-time recitation voice, a reminder of the missed part can be provided. The error reminder can be a voice prompt, a visual prompt, or a vibration prompt. For example, if the user misses the line "I suspect it is frost on the ground" while reciting "Quiet Night Thoughts", a voice prompt "The current recitation has missed a part" can be given, or the text "The current recitation has missed a part" can be displayed on the virtual display screen, or the user can be vibrated in a preset manner to inform the user of the missed part.
[0044] Optionally, when an error prompt is given, an associative image corresponding to the erroneous text may be generated and displayed. Optionally, when an omission prompt is given, an associative image corresponding to the omitted text may be generated and displayed. Optionally, when an error prompt or an omission prompt is given, the associative image may not be displayed.
[0045] By adopting the technical solution of the embodiment of the present application, it is possible to detect in real time whether there are errors or omissions in the user's current recitation, and to provide corresponding prompts so that the user can get timely feedback and then make corrections.
[0046] Based on the above technical solution, as an embodiment, the detection of whether the current recitation is obstructed by comparing the real-time recitation voice and the first recitation text may include: real-time collection of physiological data of the user during the recitation process; obtaining reference physiological data when recitation is obstructed; determining a physiological judgment result based on whether the physiological data and the reference physiological data match; determining a voice judgment result based on whether the real-time recitation voice and the first recitation text match; and determining whether the current recitation is obstructed by combining the physiological judgment result and the voice judgment result.
[0047] The physiological data may include, but is not limited to, one or more of eye movement data, electroencephalography (EEG) data, and skin conduction data.
[0048] Eye movement data refers to characteristic information such as the position, direction, speed, and duration of the human eye during gaze, saccade, and following movements. Extended reality devices may include an eye tracker, which can collect real-time eye movement data from the user. Based on this data, the user's gaze direction can be determined.
[0049] EEG data are weak electrical signals generated by brain neuronal activity. Optionally, the extended reality device can include electrodes that can be attached to the user's scalp to detect the user's EEG data. Alternatively, the user's EEG data can be acquired using an electrode cap, which transmits the acquired EEG data to the extended reality device. Because raw EEG data contains significant interference (such as eye movements, myoelectricity, and power-frequency noise), after acquisition, the EEG data can be preprocessed through methods such as artifact removal, filtering, and re-referencing to produce preprocessed EEG data.
[0050] Electrodermal data is a physiological signal generated by fluctuations in the skin's electrical conductivity (skin resistance or conductance) due to changes in sweat gland activity. Extended reality devices can include galvanic skin sensors (consisting of two electrodes) that can detect and record this data in real time.
[0051] When a user encounters difficulty reciting, they may experience increased stress, tension, and irritability. Increased stress, tension, and irritability may manifest as follows: the user's gaze may remain fixed on one location for an extended period, or they may rapidly scan the area; theta waves (4-7Hz) in EEG data may increase significantly; and skin conductivity may rise. Theta waves are often associated with stress and anxiety, and elevated theta waves reflect a decrease in the brain's cognitive control (e.g., difficulty concentrating and confusion). Stress and tension activate the sympathetic nervous system, leading to increased sweat gland secretion and an increase in electrolytes in sweat on the skin's surface. This significantly enhances the skin's electrical conductivity, manifesting as a decrease in skin resistance and an increase in conductance.
[0052] Physiological data samples can be collected from multiple subjects while they recite a text, and each physiological data sample can be labeled as obstructed. Statistical analysis can be performed on the multiple physiological data samples labeled as obstructed to obtain reference physiological data for when recitation is obstructed. The reference physiological data can be a range of values corresponding to the multiple physiological data. The reference physiological data can include reference eye movement data, reference EEG data, and / or reference skin conductance data.
[0053] Therefore, the physiological determination result can be determined based on whether the physiological data of the user during the recitation process matches the reference physiological data when the recitation is blocked. The matching of the physiological data of the user during the recitation process and the reference physiological data when the recitation is blocked can mean that the physiological data of the user during the recitation process falls within the value range represented by the reference physiological data when the recitation is blocked. The physiological determination result can include a binary classification result of whether the current recitation is blocked or not; the physiological determination result can also be the probability of the current recitation being blocked and the probability of the current recitation being not blocked, where the sum of the probability of the current recitation being blocked and the probability of the current recitation being not blocked is 1.
[0054] Optionally, the physiological determination sub-result corresponding to the physiological data can be determined based on multiple physiological data and reference physiological data corresponding to the physiological data. When the number of physiological determination sub-results indicating that the current recitation is blocked exceeds the number of physiological determination sub-results indicating that the current recitation is not blocked, the final physiological determination result is determined to be that the current recitation is blocked. When the number of physiological determination sub-results indicating that the current recitation is blocked does not exceed the number of physiological determination sub-results indicating that the current recitation is not blocked, the final physiological determination result is determined to be that the current recitation is not blocked.
[0055] Optionally, the physiological determination sub-result corresponding to the physiological data can be determined based on multiple physiological data and reference physiological data corresponding to the physiological data. When any physiological determination sub-result indicates that the current recitation is obstructed, the final physiological determination result is determined to be that the current recitation is obstructed.
[0056] Optionally, the decision tree model may be supervised and trained in advance using multiple physiological data samples to obtain a trained decision tree model. The multiple physiological data may be input into the decision tree model to obtain a physiological determination result.
[0057] The speech determination result can be determined based on whether the collected real-time recitation speech of the user matches the first recitation text. The speech determination result can include a binary classification result of whether the current recitation is blocked or not blocked; the speech determination result can also be a probability of the current recitation being blocked or a probability of the current recitation being not blocked, where the sum of the probability of the current recitation being blocked and the probability of the current recitation being not blocked is 1.
[0058] After the real-time recitation voice is collected, the real-time recitation voice can be converted into a target text, and by comparing the target text with the first recitation text, it can be determined whether the current recitation is blocked.
[0059] In one embodiment, the extended reality device can capture a picture of the user's mouth through a camera, and based on the user's mouth image and the user's real-time recitation voice, when the user frequently opens his mouth but does not make any sound, or the pause time of the real-time recitation voice exceeds a pause time threshold, determine that the voice judgment result is that the current recitation is obstructed.
[0060] By combining the physiological determination results and the voice determination results, it can be determined whether the current recitation is obstructed.
[0061] In one embodiment, it may be determined that the current recitation is blocked when both the physiological determination result and the voice determination result determine that the current recitation is blocked; it may be determined that the current recitation is blocked when either the physiological determination result or the voice determination result determines that the current recitation is blocked.
[0062] In one embodiment, the probabilities of the current recitation being obstructed corresponding to the physiological determination result and the voice determination result can be combined to determine the total probability of the current recitation being obstructed, and then determine whether the current recitation is obstructed based on the total probability of the current recitation being obstructed.
[0063] By adopting the technical solution of the embodiment of the present application, it is considered that when the user's current recitation is obstructed, it will cause changes in physiological data, so as to judge whether the user's current recitation is obstructed through physiological data, and then combine the physiological judgment results and voice judgment results to determine whether the current recitation is obstructed. In this way, the accuracy of the judgment of whether the current recitation is obstructed can be improved.
[0064] On the basis of the above technical solution, as an embodiment, the determining of the physiological judgment result according to whether the physiological data and the reference physiological data match may include one or more of the following steps: when the eye movement data represents that the duration of the user's gaze at a position reaches a duration threshold, or there are multiple scans, determining that the physiological judgment sub-result corresponding to the eye movement data is that the current recitation is obstructed; when the EEG data represents that the rising amplitude of the target wave in the EEG wave exceeds the amplitude threshold, determining that the physiological judgment sub-result corresponding to the EEG data is that the current recitation is obstructed; when the skin electricity data represents that the skin conductivity rises by more than a threshold, determining that the physiological judgment sub-result corresponding to the skin electricity data is that the current recitation is obstructed; and determining the physiological judgment result based on multiple physiological judgment sub-results.
[0065] The reference physiological data corresponding to the eye movement data can be the duration of the user's gaze at a single location reaching a duration threshold, or multiple saccades. Therefore, when the user's eye movement data indicates that the user's gaze at a single location has lasted for a duration threshold, or multiple saccades, it can be determined that the user's eye movement data matches the reference eye movement data, and the physiological determination sub-result corresponding to the eye movement data is determined to be an obstruction in the current recitation.
[0066] The reference physiological data corresponding to the EEG data can be the increase in amplitude of the target wave (theta wave) in the EEG exceeding an amplitude threshold. Therefore, when the increase in amplitude of the theta wave in the user's EEG data exceeds the amplitude threshold, it can be determined that the user's EEG data matches the reference EEG data, and the physiological determination sub-result corresponding to the EEG data is that recitation is currently blocked.
[0067] The reference physiological data corresponding to the galvanic skin data may be skin conductivity exceeding a threshold. Therefore, when the user's skin conductivity exceeds the threshold, it can be determined that the user's galvanic skin data matches the reference galvanic skin data, thereby determining that the physiological determination sub-result corresponding to the galvanic skin data indicates that recitation is currently blocked.
[0068] The physiological determination result can be determined based on the physiological determination sub-results corresponding to the multiple physiological data. Optionally, when the number of physiological determination sub-results indicating that the current recitation is blocked exceeds the number of physiological determination sub-results indicating that the current recitation is not blocked, the final physiological determination result can be determined to be that the current recitation is blocked; when the number of physiological determination sub-results indicating that the current recitation is blocked does not exceed the number of physiological determination sub-results indicating that the current recitation is not blocked, the final physiological determination result can be determined to be that the current recitation is not blocked. Optionally, when any physiological determination sub-result indicates that the current recitation is blocked, the final physiological determination result can be determined to be that the current recitation is blocked.
[0069] By adopting the technical solution of the embodiment of the present application, the physiological determination sub-result corresponding to each physiological data can be determined separately based on each physiological data, thereby determining the final physiological determination result based on multiple physiological determination sub-results. By making judgments from multiple dimensions, the accuracy of the obtained physiological determination results can be improved.
[0070] Based on the above technical solution, as an embodiment, the method may further include: obtaining a second recitation text; collecting an image of the environment; analyzing the image of the environment to determine whether the environment includes a target object related to the second recitation text; when the environment includes the target object, using the virtual display screen to spatially display the second recitation text on the target object.
[0071] The extended reality device can have an assisted recitation system, Figure 2 This is a schematic diagram of a module of an auxiliary recitation system provided by an embodiment of the present application. Figure 2 As shown, the auxiliary recitation system can have an environmental perception module, and the environmental perception module can include a camera. The user can choose to turn on or off the environmental perception module. When turning on the environmental perception module, the second recitation text can be determined, thereby whether there is a target object relevant to the second recitation text in the environment based on the environmental perception module judgment. Wherein, the second recitation text can be any text, and the method for obtaining the second recitation text can refer to the method for obtaining the first recitation text, which will not be repeated here. The second recitation text and the first recitation text can be identical or different.
[0072] In one embodiment, after acquiring the second recited text, the extended reality device may acquire images of one or more target objects related to the second recited text from the second recited text. The environment perception module may acquire images of various objects in the environment in real time and compare the acquired images of each object with pre-acquired images of the target object (calculating similarity) to determine whether the environment includes target objects related to the second recited text.
[0073] For example, when the second recitation text is "Peach tree, apricot tree, pear tree, you don't let me, I don't let you," the target objects may include peach trees, apricot trees, and pear trees. Images of the target objects (peach tree images, apricot tree images, and pear tree images) can be acquired in advance. The environment perception module can collect images of each object in the environment in real time and determine whether the similarity between the collected images of each object and the image of the target object reaches a similarity threshold. When the similarity between any collected image of the object and any image of the target object reaches the similarity threshold, it is determined that the environment includes a target object related to the second recitation text. The similarity threshold can be set according to actual needs.
[0074] In one embodiment, after obtaining the second recitation text, the names of the target objects can be directly determined from the second recitation text. The environment perception module can collect images of the objects in the environment in real time and perform semantic analysis on the collected images to determine the names of the objects. When the name of any object in the environment matches the name of any target object, it can be determined that the environment includes a target object related to the second recitation text.
[0075] For example, when the second recitation text is "peach tree, apricot tree, pear tree, you don't let me, I don't let you", the names of the target objects can include peach tree, apricot tree and pear tree. The environmental perception module can collect images of various objects in the environment in real time, analyze each image, and thus determine the name of each object; judge whether the name of each object is peach tree, apricot tree or pear tree, so as to determine whether the environment includes target objects related to the second recitation text.
[0076] When the target object is determined to be included in the environment, the full or partial second recitation text can be spatially displayed on the target object using a virtual reality screen. Spatial display refers to the use of extended reality technology to display images in space. Extended reality technology can combine the real and virtual, and can display the second recitation text as a virtual element in a real environment.
[0077] Optionally, the full or partial text of the second recitation text is spatially displayed on the target object, and a lightweight prompt can be generated next to the target object in a non-interfering manner (such as a small icon, vocabulary card and / or voice prompt, etc.) to guide the user to recall or repeat the second recitation text.
[0078] For example, when the second recitation text is "peach tree, apricot tree, pear tree, you don't let me, I don't let you", if it is detected that the environment includes a peach tree, the text "peach tree, apricot tree, pear tree, you don't let me, I don't let you" can be displayed on (or around) the peach blossoms.
[0079] For example, when the second recitation text is "Quiet Night Thoughts", the target object is a bed. If a bed is detected in the current environment, the text "The moon shines brightly in front of the bed" can be displayed at the location of the bed, or the full text of "Quiet Night Thoughts" can be displayed at the location of the bed.
[0080] By adopting the technical solution of the embodiment of the present application, the real environment can be perceived in real time, and surrounding objects or scenes can be identified and labeled, thereby helping users to associate and deepen their memory; in addition, because the appearance of target objects in the environment is random, this method of breaking the fixed rhythm to memorize can avoid the slack of attention caused by mechanical repetition, strengthen memory at the critical point of memory forgetting, thereby reducing the forgetting rate, making it easier for the second recitation text to enter long-term memory, thereby improving the effect; and, it can realize the utilization of the user's fragmented time, without spending a long time on memorization separately.
[0081] Based on the above technical solution, as an embodiment, the method may further include: obtaining a third recitation text; using the virtual display screen to fully display the third recitation text and output a guided reading voice of the third recitation text; in the initial recitation stage, hiding the key text in the displayed third recitation text; in the non-initial recitation stage, hiding the entire text of the displayed third recitation text.
[0082] The third recitation text can be any text, and the method for obtaining the third recitation text can refer to the method for obtaining the first recitation text, which will not be described in detail here. The third recitation text and the first recitation text or the second recitation text can be the same or different.
[0083] like Figure 2 As shown, the auxiliary memory system can have a memory training engine, which can divide the recitation process into multiple stages in real time. In the first stage, the third recitation text can be fully displayed on a virtual display screen, and the guidance reading voice of the third recitation text can be output, so that the user can read the third recitation text in full and follow along. When the user follows along, the pronunciation quality and speaking speed of the user can be recognized and recorded. If the follow-up is correct, the next sentence of the guidance reading voice will not be played; if the follow-up is incorrect, a brief prompt or the original text playback can be provided.
[0084] The second stage (i.e., the initial recitation stage) may be entered after the voice guidance reading of the third recitation text reaches a preset number of times. During the initial recitation stage, only a portion of the third recitation text may be displayed, and key text within the third recitation text may be hidden, allowing the user to recite the hidden key text. The key text may be determined by the extended reality device based on big data or selected by the user. By hiding the key text, the user can be guided to recall the key text.
[0085] Optionally, the key text may be determined based on the user's performance in the previous round, or the blocked text that the user had difficulty reciting in the previous round may be determined as the key text.
[0086] Can be after the user recites the hidden key text, enter the third stage (i.e. non-initial recitation stage).In the non-initial recitation stage, can carry out full text hiding to the third recitation text, i.e. do not show the third recitation text, thereby allow the user to recite the third recitation text in full.
[0087] In one embodiment, if the user makes an error when reciting the key text in the second stage, an error prompt can be given; if the user omits the key text in the second stage, an omission prompt can be given; if the user is blocked from reciting the key text in the second stage, an associative image can be generated based on the key text and displayed. The specific method of generating and displaying the associative image can be referred to the previous text.
[0088] In one of the embodiments, if the user makes an error when reciting the third recitation text in the third stage, an error prompt can be given; if the user omits the third recitation text in the third stage, an omission prompt can be given; if the user is blocked from reciting the third recitation text in the third stage, the blocked text can be determined from the third recitation text, and an associative image can be generated based on the blocked text, and then the associative image can be displayed. The specific method of generating and displaying the associative image can be referred to the previous text.
[0089] By adopting the technical solution of the embodiment of the present application, the user's recitation process can be divided into multiple stages, so that different guidance can be provided at different stages to avoid the user directly reciting the entire text, which may cause recitation difficulties. In this way, the user's recitation effect can be improved through scientific stage guidance.
[0090] Based on the above technical solution, as an embodiment, the method may further include: recording statistical data from multiple recitations of the same recitation text; the statistical data include error frequency, pause time and / or accuracy rate; based on the statistical data, dynamically adjusting the hiding strategy and recitation path of the recitation text; wherein, the hiding strategy is used to determine the hiding method of the recitation text; the recitation path is used to determine the recitation time and recitation frequency.
[0091] The user performance evaluation engine of the auxiliary memory system can analyze indicators such as the frequency of incorrect words, pause time and recitation accuracy of each recitation by the user. Multiple indicators of multiple recitations can be statistically analyzed to obtain statistical data.
[0092] Based on the statistical data, the hiding strategy for the next recitation can be adjusted. For example, during the next recitation, the phrases that the user is prone to misreciting can be hidden, thereby guiding the user to deepen their memory of the phrases.
[0093] Based on statistical data, a user learning curve can be established and a personalized recitation path can be generated. The recitation path can represent the recitation time and recitation frequency.
[0094] For example, to deepen memory, a memorization path can be generated based on the Ebbinghaus Forgetting Curve. Based on this memorization path, the text can be reviewed (consolidated in short-term memory) within 30 minutes of initial learning, recited again within 12 hours after learning (such as before bed or the next morning), and then repeated at intervals of 1, 2, 4, 7, and 15 days, gradually transferring the information to long-term memory. If statistical data indicates that the user forgets quickly, the memorization path can be adjusted to shorten the time it takes to recite the text again, thus generating a personalized memorization path.
[0095] By adopting the technical solution of the embodiment of the present application, the user's performance can be evaluated in a timely manner, thereby generating a personalized hiding strategy and recitation path corresponding to the user, so as to flexibly adjust it for the user and improve the user's recitation efficiency.
[0096] In one example, when a user is learning the English text "The Rainy Day," the system first displays the entire text and guides the user through a repeat. It then hides key words, such as "dripping" and "window," to guide the user through recall. If the user pauses for a few seconds while reciting the word "dripping," the system, using eye movement data, identifies this as a delay in recitation and automatically triggers a three-dimensional (3D) scene of raindrops falling outside a window, accompanied by the voice prompt "dripping." Later, when the user walks into the office and the glasses recognize a window and associate it with the word, a lightweight word card automatically displays at the window to consolidate the user's memory.
[0097] Based on the above technical solution, as an embodiment, Figure 2As shown, the assisted recitation system can include a user state detection module that can obtain the user's physiological data and real-time recitation speech. The assisted recitation system can also include an environmental perception module that can determine whether there are target objects in the environment related to the recitation text and mark the recitation text on the target objects in the environment. The assisted memory system can also include a memory training engine that can divide the recitation process into multiple stages in real time and provide different guidance for the user's recitation in each stage. The assisted recitation system can also include an immersive feedback module that can interface with a semantic knowledge graph to extract the specific objects, scenes, and / or contexts represented by the blocked text. This module can then retrieve corresponding 3D scenes or dynamic animation resources from a preset resource library, generate an associative image based on the 3D scene or dynamic animation resource, or directly use the 3D scene or dynamic animation resource as an associative image and display the associative image. The assisted memory system can also include a user performance evaluation engine that can analyze indicators such as the frequency of incorrect words, pause time, and recitation accuracy of each recitation. The engine can also statistically analyze multiple indicators from multiple recitations to obtain statistical data, and then adjust the hiding strategy and recitation path based on the statistical data. The auxiliary memory system may have a display module, which may display associative images and recitation texts.
[0098] To facilitate better implementation of the auxiliary recitation method of the present application, the present application also provides an auxiliary recitation device based on the above auxiliary recitation method. The meanings of the terms are the same as those in the above auxiliary recitation method, and the specific implementation details can be referred to the description in the method embodiment.
[0099] See also Figure 3 , Figure 3 : is a schematic diagram of the structure of an auxiliary recitation device provided in an embodiment of the present application, wherein the auxiliary recitation device is applied to an extended reality device, and the auxiliary recitation device includes: The voice collection module 301 is used to obtain the first recitation text and collect the real-time recitation voice; An obstruction detection module 302 is configured to detect whether the current recitation is obstructed by comparing the real-time recitation voice with the first recitation text; An image generating module 303 is configured to determine a blocked text from the first recitation text when the current recitation is blocked, and generate an associative image based on the blocked text; The image display module 304 is configured to display the associative image using the virtual display screen of the extended reality device.
[0100] In one embodiment, the apparatus further comprises: A text acquisition module, used for acquiring a second recitation text; An image acquisition module, used to acquire images of the environment; an image analysis module, configured to analyze the image of the environment to determine whether the environment includes a target object related to the second recitation text; A spatial display module is used to spatially display the second recitation text on the target object using the virtual display screen when the target object is included in the environment.
[0101] In one embodiment, the apparatus further comprises: A data recording module is used to record statistical data of multiple recitations of the same recitation text; the statistical data includes error frequency, pause time and / or accuracy rate; A dynamic adjustment module, configured to dynamically adjust the hiding strategy and recitation path of the recitation text according to the statistical data; The hiding strategy is used to determine the hiding method of the recitation text; and the recitation path is used to determine the recitation time and recitation frequency.
[0102] In one embodiment, the obstruction detection module 302 is specifically configured to perform: Real-time collection of users’ physiological data during recitation; Obtain reference physiological data when recitation is blocked; determining a physiological determination result according to whether the physiological data matches the reference physiological data; Determining a voice determination result according to whether the real-time recitation voice matches the first recitation text; The physiological determination result and the voice determination result are combined to determine whether the current recitation is obstructed.
[0103] In one embodiment, the physiological data includes eye movement data, electroencephalogram (EEG) data, and / or skin conduction data; and determining the physiological determination result based on whether the physiological data matches the reference physiological data includes one or more of the following steps: When the eye movement data indicates that the duration of the user's gaze at a position reaches a duration threshold, or when the user makes multiple saccades, determining that the physiological determination sub-result corresponding to the eye movement data is that the current recitation is obstructed; When the rising amplitude of the target wave in the brain wave represented by the brain electrical data exceeds the amplitude threshold, determining that the physiological determination sub-result corresponding to the brain electrical data is that the current recitation is obstructed; When the skin electrical data indicates that the skin conductivity increases beyond a threshold, determining that the physiological determination sub-result corresponding to the skin electrical data is that the current recitation is obstructed; The physiological determination result is determined based on the plurality of physiological determination sub-results.
[0104] In one embodiment, the apparatus further comprises: A recitation text acquisition module, used for acquiring a third recitation text; A voice output module, configured to fully display the third recitation text using the virtual display screen and output a voice guidance for reading the third recitation text; A key hiding module, used for hiding key text in the displayed third recitation text during the initial recitation stage; The full-text hiding module is used to hide the full text of the displayed third recitation text in the non-initial recitation stage.
[0105] In one embodiment, the apparatus further comprises: An error prompt module is used to give an error prompt when an error is detected in the current recitation; The omission prompt module is used to provide omission prompts when omissions are detected in the current recitation.
[0106] By adopting the technical solution of the embodiment of the present application, the extended reality device can collect the user's real-time recitation voice and compare the real-time recitation voice with the first recitation text, so as to automatically determine whether the current recitation is blocked. When the current recitation is blocked, the blocked text can be determined, and an associative image can be generated based on the blocked text, and the associative image can be displayed using the virtual display screen of the extended reality device; in this way, it is possible to automatically identify whether the current recitation is blocked without the user's active operation, which can avoid interrupting the user's recitation rhythm; the associative image corresponding to the blocked text can be displayed, so as to help the user continue to recite through the associative image; vivid associative images can attract the user's attention more than boring text, can reduce fatigue during the recitation process, increase the fun of recitation, and can also inspire the user to actively associate the image with the recitation text, and transform memory from passive reception to active construction; and associative images as visual information can help users form multi-sensory linkage, which is easier to leave a deep impression in the brain than single text stimulation, thereby effectively improving the user's memory efficiency.
[0107] The specific definition of the auxiliary recitation device can be found in the definition of the auxiliary recitation method above, which will not be repeated here. The various modules in the above-mentioned auxiliary recitation device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0108] In addition, the present application also provides an electronic device, which can be an extended reality device. Figure 4 As shown, it shows a schematic diagram of the structure of the electronic device involved in this application, specifically: The electronic device may include one or more processing core processors 401 and one or more computer readable storage media memories 402 and other components. It will be understood by those skilled in the art that Figure 4 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently. Processor 401 is the control center of the electronic device, connecting the various parts of the entire electronic device using various interfaces and lines. By running or executing software programs and / or modules stored in memory 402 and accessing data stored in memory 402, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, processor 401 may include one or more processing cores; preferably, processor 401 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 401.
[0109] Memory 402 can be used to store software programs and modules. Processor 401 executes various functional applications and data processing by running the software programs and modules stored in memory 402. Memory 402 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as sound playback or image playback); the data storage area may store data generated based on the use of the electronic device. Memory 402 may also include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory 402 may also include a memory controller to provide processor 401 with access to memory 402.
[0110] In one embodiment, the electronic device further includes a power supply 403 for supplying power to various components. Preferably, the power supply 403 can be logically connected to the processor 401 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 403 can also include any of one or more DC or AC power supplies, a recharging system, a power supply device debugging circuit, a power converter or inverter, a power status indicator, and other components.
[0111] In one embodiment, the electronic device may further include an input unit 404, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0112] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 runs the application programs stored in the memory 402, thereby implementing the steps of any of the assisted recitation methods provided in the embodiments of the present application.
[0113] When the specific electronic device is wearable optical see-through smart glasses, in addition to the above structure, it also includes at least a glasses body frame, an optical display component, an electronic circuit component, a sensor, etc. The sensors built into the glasses include a heart rate monitor, a blood glucose meter, a microphone, a camera and / or an eye tracker.
[0114] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0115] In one embodiment, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method described in any embodiment of the present application is implemented.
[0116] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment of the present application is implemented.
[0117] In some embodiments, a computer program product is also proposed, including a computer program or instructions, which implements the method described in any embodiment of the present application when executed by a processor.
[0118] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0119] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0120] To this end, the present application provides a computer-readable storage medium having a computer program stored thereon. The computer program can be loaded by a processor to execute the steps of any one of the auxiliary recitation methods provided in the present application.
[0121] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0122] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0123] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the auxiliary recitation methods provided in this application, the beneficial effects that can be achieved by any of the auxiliary recitation methods provided in this application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0124] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements that are inherent to such process, method, article, or terminal device. In the absence of further restrictions, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0125] The above is a detailed introduction to the auxiliary recitation method, device, electronic device and computer-readable storage medium provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A method for assisting recitation, characterized in that: Applied to an extended reality device, the method includes: Obtaining a first recitation text and collecting real-time recitation voice; By comparing the real-time recitation voice and the first recitation text, detecting whether the current recitation is blocked; When the current recitation is blocked, determining the blocked text from the first recitation text, and generating an associative image according to the blocked text; The associative image is displayed using a virtual display screen of the extended reality device.
2. The method according to claim 1, characterized in that The method further comprises: Obtain the second recitation text; Collect images of the environment; analyzing the image of the environment to determine whether the environment includes a target object related to the second recitation text; When the target object is included in the environment, the second recitation text is spatially displayed on the target object using the virtual display screen.
3. The method according to claim 1, characterized in that The method further comprises: Recording statistical data from multiple recitations of the same recitation text; the statistical data including error frequency, pause time and / or accuracy rate; Dynamically adjusting the hiding strategy and recitation path of the recitation text according to the statistical data; The hiding strategy is used to determine the hiding method of the recitation text; and the recitation path is used to determine the recitation time and recitation frequency.
4. The method according to claim 1, wherein The detecting whether the current recitation is blocked by comparing the real-time recitation voice and the first recitation text includes: Real-time collection of users’ physiological data during recitation; Obtain reference physiological data when recitation is blocked; determining a physiological determination result according to whether the physiological data matches the reference physiological data; Determining a voice determination result according to whether the real-time recitation voice matches the first recitation text; The physiological determination result and the voice determination result are combined to determine whether the current recitation is obstructed.
5. The method according to claim 4, characterized in that The physiological data includes: eye movement data, EEG data and / or skin electrical data; and determining the physiological determination result based on whether the physiological data matches the reference physiological data includes one or more of the following steps: When the eye movement data indicates that the duration of the user's gaze at a position reaches a duration threshold, or when the user makes multiple saccades, determining that the physiological determination sub-result corresponding to the eye movement data is that the current recitation is obstructed; When the rising amplitude of the target wave in the brain wave represented by the brain electrical data exceeds the amplitude threshold, determining that the physiological determination sub-result corresponding to the brain electrical data is that the current recitation is obstructed; When the skin electrical data indicates that the skin conductivity increases beyond a threshold, determining that the physiological determination sub-result corresponding to the skin electrical data is that the current recitation is obstructed; The physiological determination result is determined based on the plurality of physiological determination sub-results.
6. The method according to claim 1, characterized in that The method further comprises: Get the third recitation text; Utilizing the virtual display screen to completely display the third recitation text, and outputting a guided reading voice of the third recitation text; During the initial recitation phase, the key text in the displayed third recitation text is hidden; In the non-initial recitation stage, the displayed third recitation text is fully hidden.
7. The method according to claim 1, characterized in that The method further comprises: When an error is detected in the current recitation, an error prompt will be given; When omissions are detected in the current recitation, omission prompts will be given.
8. A device for assisting recitation, characterized in that: Applied to an extended reality device, the apparatus comprises: A voice collection module is used to obtain the first recitation text and collect real-time recitation voice; An obstruction detection module, configured to detect whether the current recitation is obstructed by comparing the real-time recitation voice with the first recitation text; an image generating module, configured to, when the current recitation is blocked, determine the blocked text from the first recitation text and generate an associative image based on the blocked text; An image display module is used to display the associative image using a virtual display screen of the extended reality device.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the auxiliary recitation method as claimed in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the auxiliary recitation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Auxiliary recitation method and terminal device
CN109326178A
Auxiliary recitation reminding method and device and storage medium
CN110310086A
System and method for assisting students in memorizing text content
CN113793536A
Wearable device and operation method thereof, wearable intelligent system and storage medium
CN120319246A
System for generating background of augmented reality according to element in real image and method thereof
TW202146980A