Auxiliary recitation method and device, electronic equipment and computer readable storage medium

By automatically identifying obstacles in memorization and generating associative images using augmented reality devices, the problem of dull experience and insufficient feedback in traditional memorization is solved, achieving a vivid memorization experience and efficient memory.

CN120632141BActive Publication Date: 2026-01-02FALCON INNOVATIONS TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511115597.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2026-01-02
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Traditional paper-based memorization and electronic device-assisted memorization suffer from a dull experience and an inability to obtain timely feedback. In particular, when users forget words, they need to actively operate to get prompts, which interrupts the memorization rhythm and is not efficient.

Method used

Extended reality devices collect users' real-time recitation audio, automatically identify whether recitation is hindered, and generate associative images for display. By using virtual displays to provide contextual stimulation, they can improve recitation effectiveness.

Benefits of technology

It enables the identification of memorization obstacles without user intervention, displays vivid associative images, reduces fatigue, increases the enjoyment of memorization, improves memory efficiency, forms multi-sensory linkage, and enhances memory effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632141B_ABST
    Figure CN120632141B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a kind of auxiliary recitation method, device, electronic equipment and computer readable storage medium, it is related to the field of extended reality technology;The method is applied to extended reality device.The method comprises: obtaining first recitation text, and collecting real-time recitation voice;By comparing the real-time recitation voice and the first recitation text, it is detected whether current recitation is blocked;When current recitation is blocked, blocked text is determined from the first recitation text, and an associative image is generated according to the blocked text;Using the virtual display screen of the extended reality device, the associative image is displayed.Such, the present scheme can automatically identify whether user recitation is blocked, and when recitation is blocked, associative image is given to user situational stimulation, to improve recitation effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of extended reality, and in particular to an auxiliary recitation method and device, an electronic device, and a computer readable storage medium. BACKGROUND

[0002] When a user recites a traditional paper text or words, there are problems such as a dull experience and an inability to obtain feedback in a timely manner. Related technologies assist users in reciting based on learning machines and other electronic devices. However, when a user forgets a word, the user needs to actively operate the electronic device, which will then prompt the correct recitation content. This easily disrupts the recitation rhythm and lacks situational stimulation, so the recitation efficiency is not high. SUMMARY

[0003] Embodiments of the present application provide an auxiliary recitation method and device, an electronic device, and a computer readable storage medium, which can automatically identify whether a user's recitation is blocked and provide situational stimulation to the user through an associative image when the recitation is blocked, thereby improving the recitation effect.

[0004] In a first aspect, embodiments of the present application provide an auxiliary recitation method applied to an extended reality device, the method comprising:

[0005] obtaining a first recitation text and collecting real-time recitation speech;

[0006] detecting whether the current recitation is blocked by comparing the real-time recitation speech and the first recitation text;

[0007] when the current recitation is blocked, determining a blocked text from the first recitation text and generating an associative image according to the blocked text;

[0008] displaying the associative image using a virtual display screen of the extended reality device.

[0009] In a second aspect, embodiments of the present application provide an auxiliary recitation device applied to an extended reality device, the device comprising:

[0010] a speech collection module configured to obtain a first recitation text and collect real-time recitation speech;

[0011] a block detection module configured to detect whether the current recitation is blocked by comparing the real-time recitation speech and the first recitation text;

[0012] an image generation module configured to, when the current recitation is blocked, determine a blocked text from the first recitation text and generate an associative image according to the blocked text;

[0013] an image display module configured to display the associative image using a virtual display screen of the extended reality device.

[0014] In a third aspect, the embodiments of the present application further provide an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program, when executed by the processor, implements the steps in the recitation-assisting method.

[0015] In a fourth aspect, the embodiments of the present application further provide a computer-readable storage medium, which stores a computer program, and the computer program, when executed by a processor, implements the steps in the recitation-assisting method.

[0016] In a fifth aspect, the embodiments of the present application further provide a computer program product or a computer program, which comprises computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in various optional implementation manners of the embodiments of the present application.

[0017] The embodiments of the present application have the following beneficial effects:

[0018] The extended reality device can collect real-time recitation speech of the user, and compare the real-time recitation speech with the first recitation text, so as to automatically determine whether the current recitation is blocked. When the current recitation is blocked, the blocked text can be determined, an association image can be generated based on the blocked text, and the association image can be displayed on the virtual display screen of the extended reality device. In this way, it can automatically identify whether the current recitation is blocked without the user actively operating, and can avoid interrupting the recitation rhythm of the user. The association image corresponding to the blocked text can be displayed, so as to help the user continue the recitation through the association image. The vivid association image can attract the attention of the user more than the dull text, can reduce the fatigue in the recitation process, increase the recitation pleasure, and also can stimulate the user to actively associate the image with the recitation text, and convert the memory from passive reception to active construction. In addition, the association image as visual information can help the user form multi-sensory linkage, and compared with single text stimulation, it is easier to leave a deep impression in the brain, so as to effectively improve the memory efficiency of the user. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0020] Figure 1 is a step schematic diagram of the recitation-assisting method provided by an embodiment of the present application;

[0021] Figure 2 is a module schematic diagram of an auxiliary recitation system provided by an embodiment of the present application;

[0022] Figure 3 is a structural schematic diagram of an auxiliary recitation device provided by an embodiment of the present application;

[0023] Figure 4 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0024] The technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.

[0025] In one embodiment, as shown in Figure 1 , an auxiliary recitation method is provided, although a logical order is shown in the step schematic diagram, in some cases, the steps shown or described can be performed in an order different from that shown in the drawing. Specifically, the auxiliary recitation method can be applied in an extended reality device.

[0026] Extended reality (XR) technology combines reality and virtuality through a computer to create a virtual environment that can be interacted with by a human-computer. Extended reality includes virtual reality (VR), augmented reality (AR), and mixed reality (MR), and extended reality devices include mobile handheld terminals such as mobile phones and tablets, heads-up displays such as car head-up displays, and near-eye display terminals worn on the head such as glasses-shaped displays and helmet-shaped displays. In the embodiments of the present application, the hardware structure of the extended reality device is illustrated by taking the optical see-through glasses-shaped display, i.e., the extended reality glasses, as an example.

[0027] The extended reality glasses can include a frame body, i.e., a temple and a frame, to support the extended reality glasses to be worn on a human head; can include an optical display module mainly composed of a micro display screen, an optical lens and an optical waveguide sheet to display virtual content to a user; can include an audio module mainly composed of a microphone and a speaker to collect and play sound; can include a sensor mainly composed of a camera, a gyroscope, a barometer and an infrared light emitter-receiver to collect information data related to the human body, the glasses body and the external environment; can include an integrated processor mainly composed of a microcontroller unit (MCU) or a central processing unit (CPU) to perform data processing and calculation; can include a circuit board, which is flexible or rigid, to connect other electronic components to form an electronic circuit system, which is generally placed in the inner cavity of the frame and the temple; and can further include a battery power supply.

[0028] The extended reality device can also include an eye tracker that can detect the user's gaze direction. The extended reality device can also have a camera that captures the user's eyes, and the user's gaze is analyzed through the image of the user's eyes to determine the user's gaze direction.

[0029] The extended reality device can also include a virtual display screen, which is not a physical screen in the traditional sense, but is generated by technical means, and uses light, image and other elements to make the user perceive a display area that can display various virtual information, images, videos and interactive interfaces. The virtual display screen of the extended reality device can have a large enough field of view (FOV), which determines the size of the virtual screen in the user's field of view. A larger field of view can make the user feel a wider virtual space.

[0030] The following will be described in detail. It should be noted that the order of the following embodiments is not limited as the priority order of the embodiments.

[0031] According to Figure 1 The auxiliary recitation method shown in the figure includes at least steps S110 to S140, which are described in detail as follows:

[0032] In step S110, a first recitation text is obtained, and real-time recitation speech is collected.

[0033] The first recitation text can be any text to be recited, and the first recitation text can be a word, a poem or an article.

[0034] In one embodiment, the first recitation text can be manually input by the user, transmitted to the extended reality device by the user through other electronic devices, or downloaded from a server.

[0035] In one embodiment, the first recitation text can be a text on a paper book or other electronic devices, and the extended reality device can capture an image of the first recitation text based on the camera, and extract the first recitation text from the image of the first recitation text by using optical character recognition (OCR) technology.

[0036] When the user recites the first recitation text, the microphone of the extended reality device can collect the real-time recitation voice of the user. After collecting the real-time recitation voice of the user, the real-time recitation voice can be processed, such as noise reduction, to obtain clear human voice.

[0037] In step S120, it is detected whether the current recitation is blocked by comparing the real-time recitation voice and the first recitation text.

[0038] After collecting the real-time recitation voice, the real-time recitation voice can be converted into a target text, and by comparing the target text with the first recitation text, it is determined whether the current recitation is blocked.

[0039] In one embodiment, the current recitation can be determined to be blocked when the real-time recitation voice represents that the user's recitation of the first recitation text has a pause duration reaching a pause duration threshold. The pause duration threshold can be set according to actual needs.

[0040] In one embodiment, the current recitation can be determined to be blocked when the real-time recitation voice represents that the user's recitation of the first recitation text has errors or omissions.

[0041] In step S130, when the current recitation is blocked, a blocked text is determined from the first recitation text, and an association image is generated according to the blocked text.

[0042] When the current recitation is blocked, the blocked text can be determined from the first recitation text.

[0043] When the real-time recitation voice represents that the user's recitation of the first recitation text has a pause duration reaching a pause duration threshold, the blocked text can be the text that should be recited after the pause. For example, when the user's pause duration in reciting "bedside moonlight" reaches the pause duration threshold, "doubt is ground frost" after "bedside moonlight" can be determined as the blocked text.

[0044] When the user recites the first recitation text incorrectly or omits some text during real-time recitation, the incorrect or omitted text can be determined as the blocked text. For example, when the user recites “The Midnight” incorrectly as “The Moon Frost” instead of “The Moonlight”, “The Moonlight” or “Moonlight” can be determined as the blocked text. For example, when the user recites “The Midnight” and directly recites “Lift your head and look at the bright moon” after reciting “The Moonlight in front of the bed”, the user omits “Is it frost on the ground”, and the omitted “Is it frost on the ground” can be determined as the blocked text.

[0045] In one embodiment, the extended reality device can input the blocked text into a text-to-image model, and the text-to-image model can output an associated image corresponding to the blocked text. The associated image is a visual representation of the blocked text. For example, when the blocked text is “The Moonlight in front of the bed”, the associated image can be an image of the moon outside the window shining moonlight through the window onto the bed inside the house.

[0046] In one embodiment, the extended reality device can communicate with another electronic device, and the other electronic device can have a text-to-image model. The extended reality device can transmit the blocked text to the other electronic device, and the other electronic device can generate an associated image corresponding to the blocked text based on the text-to-image model, and return the associated image to the extended reality device.

[0047] The text-to-image model is an artificial intelligence model that can generate an image corresponding to the input text. The text-to-image model can be obtained through adversarial training, supervised training or deep learning. The text-to-image model can be Hunyuan-DiT, UniT2IXL, Kolors or other models.

[0048] In one embodiment, the same blocked text can have multiple associated images, and the multiple associated images can form a video, an animated image or a short video, etc. The multiple associated images can form a story segment, a plot, etc.

[0049] For example, when the user recites “apple” incorrectly, the associated image can be an apple tree. For example, when the user recites “The volcano is erupting” incorrectly, the associated image can be multiple images, which can include a panoramic view of the volcano eruption, a close-up of the magma eruption, and an image of the magma flowing. For example, when the user recites historical content incorrectly, multiple images corresponding to a time chain scene can be generated, and the multiple images can represent each historical event of the historical content in chronological order. For example, when the user recites a textbook content incorrectly, multiple images can be generated according to the chapters of the textbook, forming a “plot chain” learning experience.

[0050] In step S140, the association image is displayed by using a virtual display screen of the extended reality device.

[0051] The extended reality device comprises a virtual display screen through which the association image can be displayed.

[0052] Optionally, when the association image is displayed, the blocked text can also be displayed, and a speaker of the extended reality device can output audio corresponding to the blocked text, so as to realize multi-sensory linkage of images, texts and audio, so as to deepen the user's memory of the blocked text.

[0053] By adopting the technical solutions of the embodiments of the present application, the extended reality device can collect real-time recitation speech of the user, and compare the real-time recitation speech with the first recitation text, so as to automatically determine whether the current recitation is blocked. When the current recitation is blocked, the blocked text can be determined, and an association image based on the blocked text can be generated, and the association image can be displayed by using the virtual display screen of the extended reality device. In this way, it can be automatically determined whether the current recitation is blocked without the user's active operation, and the recitation rhythm of the user can be avoided from being interrupted. The blocked text corresponding to the association image can be displayed, so as to help the user continue to recite through the association image. The lively association image can attract the user's attention more than the dull text, can reduce the fatigue in the recitation process, increase the recitation pleasure, and also can stimulate the user to actively associate the image with the recitation text, and convert the memory from passive reception to active construction. The association image as visual information can help the user form multi-sensory linkage, and compared with single text stimulation, it is easier to leave a deep impression in the brain, so as to effectively improve the user's memory efficiency.

[0054] On the basis of the above technical solutions, as an embodiment, the method can further comprise: when it is detected that the current recitation is wrong, an error prompt is given; and when it is detected that the current recitation is missed, a missing prompt is given.

[0055] When it is detected based on the real-time recitation speech that the current recitation of the user is wrong, an error prompt can also be given. The error prompt can be a voice prompt, a visual prompt or a vibration prompt. For example, when the user recites “The Midnight Thought” incorrectly as “bedside moon frost” instead of “bedside moonlight”, a voice prompt “current recitation error” can be given, or a text “current recitation error” can be displayed through the virtual display screen, or a vibration can be performed in a preset manner, so that the user knows that the current recitation is wrong.

[0056] When the real-time recitation voice is detected to have a missing in the current recitation of the user, a missing prompt can also be performed. The error prompt can be a voice prompt, a visual prompt or a vibration prompt. For example, when the user misses “doubt is ground frost” in reciting “Quiet Night Thoughts”, a voice prompt “current recitation has a missing” can be performed, or a text “current recitation has a missing” can be displayed on the virtual display screen, or a vibration can be performed in a preset manner to make the user know that the current recitation has a missing.

[0057] Optionally, when the error prompt is performed, an associated image corresponding to the error text can also be generated and displayed. Optionally, when the missing prompt is performed, an associated image corresponding to the missing text can also be generated and displayed. Optionally, when the error prompt or the missing prompt is performed, the associated image can not be displayed.

[0058] By adopting the technical solutions of the embodiments of the present application, it can be detected in real time whether the current recitation of the user has an error or a missing, and corresponding prompts can be performed to make the user get feedback in time and then get correction.

[0059] On the basis of the above technical solutions, as an embodiment, the detecting whether the current recitation is blocked by comparing the real-time recitation voice and the first recitation text can include: collecting physiological data of the user in the recitation process in real time; obtaining reference physiological data when the recitation is blocked; determining a physiological determination result according to whether the physiological data and the reference physiological data match; determining a voice determination result according to whether the real-time recitation voice and the first recitation text match; and determining whether the current recitation is blocked by integrating the physiological determination result and the voice determination result.

[0060] The physiological data can include but is not limited to one or more of eye movement data, electroencephalography (EEG) data and electrodermal data.

[0061] The eye movement data refers to position, direction, speed, duration and other characteristic information of the human eye in the movement process of fixation, saccade, pursuit and the like. The extended reality device can have an eye tracker, and the eye tracker can collect eye movement data of the user in real time. According to the eye movement data of the user, the line-of-sight direction of the user can be determined.

[0062] Electroencephalogram (EEG) data is a weak electrical signal generated by the activity of neurons in the brain. Optionally, the extended reality device can have an electrode patch that can be attached to the user's scalp, so as to detect the user's EEG data. Optionally, the user's EEG data can be obtained through an electrode cap, which transmits the obtained EEG data to the extended reality device. Because the original EEG data contains a large amount of interference (such as eye movement, electromyogram, and power frequency noise), after the EEG data is collected, the EEG data can be preprocessed through processing methods such as artifact removal, filtering, and re-reference to obtain preprocessed EEG data.

[0063] Electrodermal data is a physiological signal of fluctuation of conductive ability (skin resistance or conductance) of the surface of human skin due to sweat gland activity changes. The extended reality device can have an electrodermal sensor (consisting of two electrodes), and the extended reality device can detect and record electrodermal data in real time based on the electrodermal sensor.

[0064] When the user is reciting a text and is blocked, the user's stress increases, and emotions such as tension and irritability appear. When the stress increases, tension and irritability appear, the following situations can occur: the user's gaze can be fixed on a position for a long time, or rapid saccades can occur; the theta wave (4-7 Hz) in the EEG data can be significantly increased; the skin conductivity can be increased. Among them, the theta wave is usually related to stress and anxiety, and the increase of the theta wave reflects the decrease of the cognitive control ability of the brain (such as difficulty in concentrating, disordered thinking). Stress and tension can activate the sympathetic nervous system, increase the secretion of sweat glands, and increase the electrolytes in the sweat on the surface of the skin, thereby significantly enhancing the conductive ability of the skin, which is manifested as a decrease in skin resistance and an increase in conductance value.

[0065] A plurality of physiological data samples of subjects reciting a text can be collected, and each physiological data sample can be labeled as blocked or not blocked. Statistical analysis can be performed on the plurality of physiological data samples labeled as blocked, so as to obtain reference physiological data when reciting is blocked. The reference physiological data can be a value range corresponding to a plurality of physiological data. The reference physiological data can include reference eye movement data, reference EEG data, and / or reference electrodermal data.

[0066] Therefore, whether the physiological data of the user during recitation and the reference physiological data when recitation is blocked match can be determined based on the physiological data of the user during recitation and the reference physiological data when recitation is blocked, so as to determine a physiological determination result. The physiological data of the user during recitation and the reference physiological data when recitation is blocked can mean that the physiological data of the user during recitation falls within the value range represented by the reference physiological data when recitation is blocked. The physiological determination result can include a binary classification result of current recitation being blocked and current recitation not being blocked; the physiological determination result can also be a probability of current recitation being blocked and a probability of current recitation not being blocked, wherein the sum of the probability of current recitation being blocked and the probability of current recitation not being blocked is 1.

[0067] Optionally, the physiological data corresponding physiological determination sub-results can be determined based on the plurality of physiological data and the reference physiological data corresponding to the physiological data, and when the number of physiological determination sub-results representing that the current recitation is blocked is more than the number of physiological determination sub-results representing that the current recitation is not blocked, the final physiological determination result is determined to be that the current recitation is blocked, and when the number of physiological determination sub-results representing that the current recitation is blocked is not more than the number of physiological determination sub-results representing that the current recitation is not blocked, the final physiological determination result is determined to be that the current recitation is not blocked.

[0068] Optionally, the physiological data corresponding physiological determination sub-results can be determined based on the plurality of physiological data and the reference physiological data corresponding to the physiological data, and when any of the physiological determination sub-results represents that the current recitation is blocked, the final physiological determination result is determined to be that the current recitation is blocked.

[0069] Optionally, the decision tree model can be supervised trained in advance by a plurality of physiological data samples to obtain a trained decision tree model. The plurality of physiological data can be input into the decision tree model to obtain the physiological determination result.

[0070] The speech determination result can be determined according to whether the collected real-time recitation speech matches the first recitation text. The speech determination result can include a binary classification result of current recitation being blocked and current recitation not being blocked; or the speech determination result can be a probability of current recitation being blocked and a probability of current recitation not being blocked, wherein the sum of the probability of current recitation being blocked and the probability of current recitation not being blocked is 1.

[0071] In one of the embodiments, the extended reality device can collect the user's mouth picture through the camera, and determine the speech determination result to be that the current recitation is blocked when the user frequently opens the mouth without making a sound or the pause time of the real-time recitation speech exceeds the pause time threshold.

[0072] In one of the embodiments, the extended reality device can collect the user's mouth picture through the camera, and determine the speech determination result to be that the current recitation is blocked when the user frequently opens the mouth without making a sound or the pause time of the real-time recitation speech exceeds the pause time threshold.

[0073] The physiological determination result and the speech determination result can be combined to determine whether the current recitation is blocked.

[0074] In one of the embodiments, the current recitation is determined to be blocked when both the physiological determination result and the speech determination result determine that the current recitation is blocked; or the current recitation is determined to be blocked when either the physiological determination result or the speech determination result determines that the current recitation is blocked.

[0075] In one of the embodiments, the total probability of the current recitation being blocked can be determined by integrating the probability of the current recitation being blocked corresponding to the physiological determination result and the speech determination result respectively, so as to determine whether the current recitation is blocked according to the total probability of the current recitation being blocked.

[0076] The technical scheme of the embodiments of the present application considers that the physiological data will change when the user is currently blocked in recitation, so as to determine whether the user is currently blocked in recitation by the physiological data, and then determine whether the current recitation is blocked by integrating the physiological determination result and the speech determination result, so as to improve the accuracy of the determination of whether the current recitation is blocked.

[0077] On the basis of the above technical scheme, as one of the embodiments, the physiological determination result can be determined according to whether the physiological data and the reference physiological data match, which can include one or more of the following steps: when the duration of the user's gaze at a position represented by the eye movement data reaches a duration threshold or the user's saccades are multiple, the physiological determination sub-result corresponding to the eye movement data is determined to be that the current recitation is blocked; when the rising amplitude of the target wave (theta wave) in the brain electrical wave represented by the electroencephalogram data exceeds an amplitude threshold, the physiological determination sub-result corresponding to the electroencephalogram data is determined to be that the current recitation is blocked; when the skin conductivity represented by the skin electricity data rises above a threshold, the physiological determination sub-result corresponding to the skin electricity data is determined to be that the current recitation is blocked; and the physiological determination result is determined according to the plurality of physiological determination sub-results.

[0078] The reference physiological data corresponding to the eye movement data can be that the duration of the user's gaze at a position reaches a duration threshold or the user's saccades are multiple. Therefore, when the duration of the user's gaze at a position represented by the user's eye movement data reaches a duration threshold or the user's saccades are multiple, it can be determined that the user's eye movement data and the reference eye movement data match, so as to determine that the physiological determination sub-result corresponding to the eye movement data is that the current recitation is blocked.

[0079] The reference physiological data corresponding to the electroencephalogram data can be that the rising amplitude of the target wave (theta wave) in the brain electrical wave exceeds an amplitude threshold. Therefore, when the rising amplitude of the theta wave in the brain electrical wave represented by the user's electroencephalogram data exceeds the amplitude threshold, it can be determined that the user's electroencephalogram data and the reference electroencephalogram data match, so as to determine that the physiological determination sub-result corresponding to the electroencephalogram data is that the current recitation is blocked.

[0080] The reference physiological data corresponding to the skin electricity data can be that the skin conductivity rises above a threshold. Therefore, when the user's skin conductivity rises above a threshold, it can be determined that the user's skin electricity data and the reference skin electricity data match, so as to determine that the physiological determination sub-result corresponding to the skin electricity data is that the current recitation is blocked.

[0081] The physiological determination result can be determined according to the physiological determination sub-results corresponding to the plurality of physiological data. Optionally, when the number of physiological determination sub-results representing that the current recitation is blocked is more than the number of physiological determination sub-results representing that the current recitation is not blocked, the final physiological determination result is determined to be that the current recitation is blocked; when the number of physiological determination sub-results representing that the current recitation is blocked is not more than the number of physiological determination sub-results representing that the current recitation is not blocked, the final physiological determination result is determined to be that the current recitation is not blocked. Optionally, when any physiological determination sub-result represents that the current recitation is blocked, the final physiological determination result is determined to be that the current recitation is blocked.

[0082] By adopting the technical solutions of the embodiments of the present application, the physiological determination sub-results corresponding to the respective physiological data can be determined based on the respective physiological data, the final physiological determination result can be determined based on the plurality of physiological determination sub-results, and the accuracy of the obtained physiological determination result can be improved by determining from multiple dimensions.

[0083] Based on the above technical solutions, as an embodiment, the method can further include: obtaining a second recitation text; collecting an image of an environment; analyzing the image of the environment to determine whether the environment includes a target object related to the second recitation text; and when the environment includes the target object, using the virtual display screen to spatially display the second recitation text on the target object.

[0084] The extended reality device can have an auxiliary recitation system, Figure 2 is a module schematic diagram of the auxiliary recitation system provided by an embodiment of the present application, as Figure 2 The auxiliary recitation system can have an environment perception module, and the environment perception module can include a camera. The user can choose to turn on or turn off the environment perception module. When the environment perception module is turned on, the second recitation text can be determined, so as to determine whether there is a target object related to the second recitation text in the environment based on the environment perception module. The second recitation text can be any text, and the method of obtaining the second recitation text can refer to the method of obtaining the first recitation text, which will not be described here. The second recitation text and the first recitation text can be the same or different.

[0085] In one of the embodiments, after obtaining the second recitation text, the extended reality device can obtain the image of one or more target objects related to the second recitation text from the second recitation text. The environment perception module can collect the images of the respective objects in the environment in real time, and compare (calculate the similarity) the collected images of the respective objects with the pre-obtained images of the target objects, so as to determine whether the environment includes the target object related to the second recitation text.

[0086] For example, when the second recited text is "peach tree, apricot tree, pear tree, you don't let me, I don't let you", the target objects can include peach tree, apricot tree and pear tree, and the images of the target objects (peach tree image, apricot tree image and pear tree image) can be obtained in advance. The environment perception module can collect images of objects in the environment in real time, and determine whether the similarity between the collected images of the objects and the images of the target objects reaches a similarity threshold. When the similarity between the image of any object collected and the image of any target object reaches the similarity threshold, it is determined that the environment includes the target object related to the second recited text. The similarity threshold can be set according to actual needs.

[0087] In one embodiment, after obtaining the second recited text, the names of the target objects can be directly determined from the second recited text. The environment perception module can collect images of objects in the environment in real time, and perform semantic analysis on the collected images of the objects to determine the names of the objects. When the name of any object in the environment is consistent with the name of any target object, it is determined that the environment includes the target object related to the second recited text.

[0088] For example, when the second recited text is "peach tree, apricot tree, pear tree, you don't let me, I don't let you", the names of the target objects can include peach tree, apricot tree and pear tree. The environment perception module can collect images of objects in the environment in real time, analyze the images to determine the names of the objects, and determine whether the names of the objects are peach tree, apricot tree or pear tree, so as to determine whether the environment includes the target object related to the second recited text.

[0089] When it is determined that the environment includes the target object, the virtual reality screen can be used to spatially display the full text or part of the second recited text on the target object. Spatial display refers to displaying a picture in space by using extended reality technology. The extended reality technology can combine reality and virtuality, and the extended reality technology can display the second recited text as a virtual element in the real environment.

[0090] Alternatively, spatially displaying the full text or part of the second recited text on the target object can generate a light prompt in a non-interfering manner (such as a small icon, a word card and / or a voice prompt) beside the target object, guiding the user to recall or repeat the second recited text.

[0091] For example, when the second recited text is "peach tree, apricot tree, pear tree, you don't let me, I don't let you", if it is detected that the environment includes a peach tree, the text "peach tree, apricot tree, pear tree, you don't let me, I don't let you" can be displayed on (or around) the peach tree.

[0092] For example, when the second recited text is “Quiet Night Thoughts”, the target object is a bed, and it is detected that there is a bed in the current environment, the text “moonlight before the bed” can be displayed at the location of the bed, or the full text of “Quiet Night Thoughts” can be displayed at the location of the bed.

[0093] The technical solution of the embodiment of the present application can perceive the real environment in real time, identify and label the surrounding objects or scenes, thereby helping the user to associate and deepen the memory. In addition, because the appearance of the target object in the environment is random, this method of breaking the fixed rhythm for memory can avoid the attention slack caused by mechanical repetition, strengthen the memory at the memory forgetting critical point, thereby reducing the forgetting rate, making the second recited text more easily enter the long-term memory, and thereby improving the effect. Moreover, the fragmented time of the user can be utilized, and there is no need to spend a long time for memory alone.

[0094] On the basis of the above technical solution, as an embodiment, the method can further include: obtaining a third recited text; using the virtual display screen to completely display the third recited text and output a guided reading voice of the third recited text; in an initial recitation stage, hiding key texts in the displayed third recited text; and in a non-initial recitation stage, hiding the full text of the displayed third recited text.

[0095] The third recited text can be any text, and the method of obtaining the third recited text can refer to the method of obtaining the first recited text, which will not be described here. The third recited text can be the same as or different from the first recited text or the second recited text.

[0096] As shown in Figure 2 The memory training engine can divide the recitation process into multiple stages in real time. In the first stage, the virtual display screen can be used to completely display the third recited text and output a guided reading voice of the third recited text, so that the user can read the third recited text completely and follow the reading. When the user follows the reading, the pronunciation quality and the speed of the user's reading can be recognized and recorded. If the reading is correct, the next sentence of the guided reading voice cannot be played; if the reading is incorrect, a brief prompt or a playback of the original text can be provided.

[0097] The second stage (i.e., the initial recitation stage) can be entered after the number of times of outputting the guided reading voice of the third recited text reaches a preset number of times. In the initial recitation stage, only part of the third recited text can be displayed, and key texts in the third recited text can be hidden, so that the user recites the hidden key texts. The key texts can be determined by the extended reality device based on big data, or can be selected by the user. By hiding the key texts, the user can be guided to recall the key texts.

[0098] Optionally, the key text can be determined according to the user's last round performance. The user's recitation of the blocked text in the last round can be determined as the key text.

[0099] The third recitation text can be fully hidden in the third stage, i.e., the non-initial recitation stage. In the non-initial recitation stage, the third recitation text can be fully hidden, i.e., not displayed, so that the user recites the third recitation text in full.

[0100] In one of the embodiments, if the user makes a mistake in reciting the key text in the second stage, an error prompt can be given; if the user omits the key text in the second stage, an omission prompt can be given; and if the user is blocked in reciting the key text in the second stage, an association image can be generated according to the key text, and the association image can be displayed. The specific method of generating and displaying the association image can be referred to the foregoing.

[0101] In one of the embodiments, if the user makes a mistake in reciting the third recitation text in the third stage, an error prompt can be given; if the user omits the third recitation text in the third stage, an omission prompt can be given; and if the user is blocked in reciting the third recitation text in the third stage, a blocked text can be determined from the third recitation text, an association image can be generated according to the blocked text, and the association image can be displayed. The specific method of generating and displaying the association image can be referred to the foregoing.

[0102] The technical solution of the embodiments of the present application can divide the user's recitation process into multiple stages, so that different guidance can be given in different stages, avoiding the difficulty of direct full-text recitation by the user. In this way, the recitation effect of the user can be improved through scientific stage guidance.

[0103] On the basis of the above technical solution, as an embodiment, the method can further include: recording statistical data in multiple recitations of the same recitation text; the statistical data includes error frequency, pause time, and / or accuracy rate; and dynamically adjusting the hiding strategy and recitation path of the recitation text according to the statistical data; wherein the hiding strategy is used to determine the hiding manner of the recitation text; and the recitation path is used to determine the recitation time and recitation frequency.

[0104] The user performance evaluation engine of the auxiliary memory system can analyze the error word frequency, pause time, and recitation accuracy rate of each recitation of the user, and can statistically analyze multiple indicators of multiple recitations to obtain statistical data.

[0105] According to the statistical data, the hiding strategy of the next recitation can be adjusted. For example, in the next recitation process, the user can be guided to deepen the memory of the word group that the user is prone to recite incorrectly by hiding the word group.

[0106] According to the statistical data, a user learning curve can be established, and a recitation path can be generated individually. The recitation path can represent recitation time and recitation frequency.

[0107] For example, in order to deepen memory, a recitation path can be generated according to an Ebbinghaus Forgetting Curve, based on which the recitation text can be reviewed (to consolidate short-term memory) within 30 minutes after first learning, and the recitation text can be recited again within 12 hours after learning (e.g., before going to bed or in the morning of the next day), and then the recitation is repeated at intervals of 1 day, 2 days, 4 days, 7 days, and 15 days, so as to gradually transfer information into long-term memory. When the statistical data represent that the user's forgetting speed is fast, the recitation path can be adjusted to shorten the time for reciting again, and an individualized recitation path can be generated.

[0108] By using the technical solution of the embodiments of the present application, the performance of the user can be evaluated in a timely manner, so as to generate an individualized hiding strategy and recitation path corresponding to the user, to flexibly adjust for the user, and to improve the recitation efficiency of the user.

[0109] In one embodiment, when a user learns an English text "The Rainy Day", the system first displays the text completely, and guides the user to read it once. Then, some key texts such as "dripping" and "window" are hidden, and the user is guided to recall the key texts. The user pauses for a few seconds at the recitation of "dripping", and the system can determine that the recitation is blocked through eye movement data, and automatically trigger a three-dimensional (3D) scene of raindrops falling outside the window, with a voice prompt of "dripping". Then, when the user walks into an office and the glasses recognize the window and associate the word, a light word card is automatically displayed at the window, to form a consolidated memory.

[0110] On the basis of the above technical solution, as an embodiment, as Figure 2As shown, the auxiliary recitation system can have a user state detection module that can obtain physiological data of the user and real-time recitation speech. The auxiliary recitation system can have an environment perception module that can determine whether there is a target object related to the recitation text in the environment and mark the recitation text on the target object in the environment. The auxiliary memory system can have a memory training engine that can divide the recitation process into multiple stages in real time and guide the user's recitation in different stages. The auxiliary recitation system can have an immersion feedback module that can interface with a semantic knowledge graph, extract specific objects, scenes and / or contexts represented by the blocked text, so as to call corresponding 3D scene or dynamic animation resources from a preset resource library, and then generate an association image based on the 3D scene or dynamic animation resources, or directly use the 3D scene or dynamic animation resources as the association image, and display the association image. The auxiliary memory system can have a user performance evaluation engine that can analyze the error word frequency, pause time and recitation accuracy of each recitation of the user, and statistically analyze multiple indicators of multiple recitations to obtain statistical data, and then adjust the hiding strategy and recitation path according to the statistical data. The auxiliary memory system can have a display module that can display the association image and the recitation text.

[0111] To better implement the auxiliary recitation method of the present application, the present application also provides an auxiliary recitation device based on the above-mentioned auxiliary recitation method. The meanings of the terms are the same as in the above-mentioned auxiliary recitation method, and the specific implementation details can be referred to the description in the method embodiment.

[0112] Please refer to Figure 3 , Figure 3 is a structural schematic diagram of an auxiliary recitation device provided by an embodiment of the present application, wherein the auxiliary recitation device is applied to an extended reality device, and the auxiliary recitation device comprises:

[0113] The voice collection module 301 is configured to obtain a first recitation text and collect real-time recitation speech.

[0114] The blocked detection module 302 is configured to detect whether the current recitation is blocked by comparing the real-time recitation speech with the first recitation text.

[0115] The image generation module 303 is configured to determine blocked text from the first recitation text when the current recitation is blocked, and generate an association image according to the blocked text.

[0116] The image display module 304 is configured to display the association image by using a virtual display screen of the extended reality device.

[0117] In one embodiment, the device further comprises:

[0118] a text acquisition module, configured to acquire a second recitation text;

[0119] an image acquisition module, configured to acquire an image of an environment;

[0120] an image analysis module, configured to analyze the image of the environment to determine whether the environment includes a target object related to the second recitation text;

[0121] a spatial display module, configured to, when the environment includes the target object, display the second recitation text on the target object by using the virtual display screen.

[0122] In one of the embodiments, the device further includes:

[0123] a data recording module, configured to record statistical data in multiple recitations of the same recitation text; the statistical data includes error frequency, pause time and / or accuracy rate;

[0124] a dynamic adjustment module, configured to dynamically adjust a hiding strategy and a recitation path of the recitation text according to the statistical data;

[0125] The hiding strategy is used to determine the hiding manner of the recitation text; and the recitation path is used to determine the recitation time and the recitation frequency.

[0126] In one of the embodiments, the obstruction detection module 302 is specifically configured to perform:

[0127] acquire physiological data of a user in a recitation process in real time;

[0128] acquire reference physiological data when the recitation is obstructed;

[0129] determine a physiological determination result according to whether the physiological data and the reference physiological data match;

[0130] determine a speech determination result according to whether the real-time recitation speech and the first recitation text match;

[0131] integrate the physiological determination result and the speech determination result to determine whether the current recitation is obstructed.

[0132] In one of the embodiments, the physiological data includes eye movement data, electroencephalogram data and / or electrodermal data; and the determination of the physiological determination result according to whether the physiological data and the reference physiological data match includes one or more of the following steps:

[0133] when the eye movement data represents that the user gazes at a position for a duration reaching a duration threshold, or a plurality of saccades, determining that a physiological judgment sub-result corresponding to the eye movement data is that the current recitation is blocked;

[0134] when the electroencephalogram data represents that an amplitude of a target wave in the electroencephalogram wave exceeds an amplitude threshold, determining that a physiological judgment sub-result corresponding to the electroencephalogram data is that the current recitation is blocked;

[0135] when the electrodermal data represents that the skin conductivity increases by more than a threshold, determining that a physiological judgment sub-result corresponding to the electrodermal data is that the current recitation is blocked;

[0136] determining the physiological judgment result according to a plurality of the physiological judgment sub-results.

[0137] In one of the embodiments, the device further comprises:

[0138] a recitation text obtaining module, configured to obtain a third recitation text;

[0139] a voice output module, configured to display the third recitation text completely by using the virtual display screen, and output a guided reading voice of the third recitation text;

[0140] a key hiding module, configured to hide a key text in the displayed third recitation text in an initial recitation stage;

[0141] a full-text hiding module, configured to hide the displayed third recitation text in a non-initial recitation stage.

[0142] In one of the embodiments, the device further comprises:

[0143] an error prompt module, configured to perform error prompt when detecting that the current recitation is wrong;

[0144] an omission prompt module, configured to perform omission prompt when detecting that the current recitation is omitted.

[0145] Using the technical solution of this application embodiment, the extended reality device can collect the user's real-time recitation voice and compare the real-time recitation voice with the first recitation text to automatically determine whether the current recitation is blocked. When the current recitation is blocked, the blocked text can be identified, and an associative image can be generated based on the blocked text. The associative image is then displayed on the virtual display screen of the extended reality device. In this way, the current recitation can be automatically identified without the user's active operation, which can avoid interrupting the user's recitation rhythm. The associative image corresponding to the blocked text can be displayed, thereby helping the user continue to recite. Vivid associative images are more attractive to users than dry text, which can reduce fatigue during the recitation process, increase the fun of recitation, and also stimulate users to actively associate images with the recitation text, turning memory from passive reception to active construction. Moreover, as visual information, associative images can help users form multi-sensory linkage, which is more likely to leave a deep impression on the brain than single text stimulation, thereby effectively improving the user's memory efficiency.

[0146] Specific limitations regarding the aided memorization device can be found in the limitations of the aided memorization method above, and will not be repeated here. Each module in the aforementioned aided memorization device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0147] Furthermore, this application also provides an electronic device that can be an augmented reality device. For example... Figure 4 As shown, it illustrates the structural diagram of the electronic device involved in this application, specifically:

[0148] The electronic device may include components such as a processor 401 with one or more processing cores and a memory 402 with one or more computer-readable storage media. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0149] The processor 401 is the control center of the electronic device, connects each part of the entire electronic device through various interfaces and lines, executes various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 402 and calling data stored in the memory 402, thereby overall monitoring the electronic device. Optionally, the processor 401 can include one or more processing cores; preferably, the processor 401 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 401.

[0150] The memory 402 can be used to store software programs and modules, and the processor 401 executes various functions and data processing by running the software programs and modules stored in the memory 402. The memory 402 can mainly include a program storage area and a data storage area, wherein the program storage area can store the operating system, at least one application program required by the function (such as sound playing function, image playing function, etc.), etc.; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 402 can include a high-speed random access memory, and can also include a non-volatile memory, for example, at least one magnetic disk storage device, flash memory device, or other volatile solid-state memory device. Accordingly, the memory 402 can also include a memory controller to provide the processor 401 with access to the memory 402.

[0151] In one embodiment, the electronic device further includes a power supply 403 for supplying power to each component, and preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management, etc. through the power management system. The power supply 403 can also include one or more than one direct current or alternating current power supply, a recharging system, a power supply device debugging circuit, a power supply converter or inverter, a power supply state indicator, etc. any component.

[0152] In one embodiment, the electronic device can further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0153] Although not shown, the electronic device can further include a display unit, etc., which will not be described herein. Specifically, in the present embodiment, the processor 401 in the electronic device will load the executable file corresponding to the process of one or more application programs into the memory 402 according to the following instructions, and run the application program stored in the memory 402 by the processor 401, thereby implementing the steps in any of the auxiliary reciting methods provided in the present application.

[0154] Specifically, when the electronic device is wearable optical see-through smart glasses, in addition to the above structure, it at least includes a glasses body frame, an optical display component, an electronic circuit component, a sensor, etc., and the sensors built-in the glasses include a heart rate monitor, a blood glucose detector, a microphone, a camera and / or an eye tracker.

[0155] Those skilled in the art can understand that, Figure 4 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. Specifically, the electronic device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0156] In one embodiment, an electronic device is provided, including a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the method described in any embodiment of the present application.

[0157] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the method described in any embodiment of the present application.

[0158] In some embodiments, a computer program product is also provided, which includes a computer program or instructions, and the computer program or instructions are executed by a processor to implement the method described in any embodiment of the present application.

[0159] The specific implementation of each operation can refer to the previous embodiments, which will not be described herein.

[0160] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by related hardware controlled by the instructions, which can be stored in a computer readable storage medium and loaded and executed by a processor.

[0161] To this end, the present application provides a computer readable storage medium, which stores a computer program, and the computer program can be loaded by a processor to execute the steps in any of the auxiliary reciting methods provided in the present application.

[0162] The specific implementation of the above operations can refer to the foregoing embodiments, which will not be repeated here.

[0163] The computer readable storage medium can include a read only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0164] Due to the instructions stored in the computer readable storage medium, the steps in any of the recitation assisting methods provided in the present application can be executed, and thus the beneficial effects of any of the recitation assisting methods provided in the present application can be achieved. Details are shown in the foregoing embodiments, which will not be repeated here.

[0165] Finally, it should be noted that in this document, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between the entities or operations. Moreover, the terms “include”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or terminal device. Without more limitations, the element defined by the statement “including a…” does not exclude the presence of other identical elements in the process, method, article or terminal device including the element.

[0166] The above provides a detailed description of the recitation assisting method, device, electronic equipment and computer readable storage medium provided in the present application. The principles and implementation modes of the present application are described by applying specific examples in this document. The above example is only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed; in conclusion, the content of the specification should not be understood as a limitation of the present application.

Claims

1. An assisted recitation method, characterized by, The method is applied to an extended reality device, and the method comprises: obtaining a first recited text and collecting real-time recited speech; detecting whether the current recitation is blocked by comparing the real-time recited speech and the first recited text; when the current recitation is blocked, determining blocked text from the first recited text, and generating an associated image according to the blocked text, comprising: inputting the blocked text into a text-to-image model, and the text-to-image model outputs an associated image corresponding to the blocked text; the associated image is a visual representation of the blocked text; the same blocked text corresponds to multiple associated images, and the multiple associated images form a story segment or a plot; when the user recites historical content, the multiple associated images represent each historical event of the historical content in chronological order; displaying the associated image on a virtual display screen of the extended reality device; when an environment sensing module is turned on, obtaining a second recited text, and obtaining an image and a name of a target object related to the second recited text; collecting images of each object in the environment in real time and determining the name of the object; analyzing the image of the environment to determine whether the target object related to the second recited text is included in the environment, comprising: when the object is similar to the image of the target object or the name is consistent, determining that the target object is included in the environment; when the target object is included in the environment, using the virtual display screen and the extended reality technology to display the second recited text as a virtual element on the target object in the real environment.

2. The method of claim 1, wherein, The method further comprises: recording statistical data in multiple recitations of the same recited text; the statistical data includes error frequency, pause time and / or accuracy; according to the statistical data, dynamically adjusting the hiding strategy and recitation path of the recited text; wherein the hiding strategy is used to determine the hiding mode of the recited text; the recitation path is used to determine the recitation time and recitation frequency.

3. The method of claim 1, wherein, The method further comprises: collecting physiological data of the user during the recitation process in real time; obtaining reference physiological data when the recitation is blocked; determining a physiological determination result according to whether the physiological data and the reference physiological data match; determining a speech determination result according to whether the real-time recited speech and the first recited text match; combining the physiological determination result and the speech determination result to determine whether the current recitation is blocked.

4. The method of claim 3, wherein, The physiological data includes eye movement data, electroencephalogram data and / or electrodermal data; and determining a physiological determination result according to whether the physiological data and the reference physiological data match comprises one or more of the following steps: when the duration of the user's gaze at a position represented by the eye movement data reaches a duration threshold, or when the user's gaze is scanned multiple times, determining that the physiological determination sub-result corresponding to the eye movement data is that the current recitation is blocked; when the rising amplitude of a target wave in the electroencephalogram represented by the electroencephalogram data exceeds an amplitude threshold, determining that the physiological determination sub-result corresponding to the electroencephalogram data is that the current recitation is blocked. determine that the skin conductance data corresponds to a physiological judgment sub-result of current recitation blockage when the skin conductance data represents a skin conductivity rise exceeding a threshold value; determine the physiological judgment result according to a plurality of the physiological judgment sub-results.

5. The method of claim 1, wherein, The method further comprises: acquiring third recitation text; displaying the third recitation text completely on the virtual display screen and outputting guided reading voice of the third recitation text; hiding key text in the displayed third recitation text in an initial recitation stage; completely hiding the displayed third recitation text in a non-initial recitation stage.

6. The method of claim 1, wherein, The method further comprises: performing an error prompt when detecting that the current recitation is wrong; performing a missing prompt when detecting that the current recitation is missing.

7. An apparatus for aiding recitation, characterized by The device is applied to an extended reality device, and the device comprises: a voice acquisition module configured to acquire first recitation text and acquire real-time recitation voice; a blockage detection module configured to detect whether the current recitation is blocked by comparing the real-time recitation voice and the first recitation text; an image generation module configured to determine blocked text from the first recitation text when the current recitation is blocked, and generate an association image according to the blocked text, including: inputting the blocked text into a text-to-image model, and the text-to-image model outputs an association image corresponding to the blocked text; the association image is a visual expression of the blocked text; a plurality of association images corresponding to the same blocked text form a story fragment or a plot; when a user recites historical content and the historical content is blocked, a plurality of association images represent each historical event of the historical content in chronological order; an image display module configured to display the association image on a virtual display screen of the extended reality device; a text acquisition module configured to acquire second recitation text when an environment perception module is started, and acquire an image and a name of a target object related to the second recitation text; an image acquisition module configured to acquire images of each object in the environment in real time and determine the name of the object; an image analysis module configured to analyze the images of the environment to determine whether the target object related to the second recitation text is included in the environment, including: determining that the target object is included in the environment when the object is similar to the image of the target object or the name is consistent; a spatial display module configured to display the second recitation text as a virtual element on the target object in the real environment using the virtual display screen and the extended reality technology when the target object is included in the environment.

8. An electronic device, comprising: The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps in the auxiliary recitation method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps in the auxiliary recitation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Auxiliary recitation method and terminal device

    CN109326178A

  • Auxiliary recitation reminding method and device and storage medium

    CN110310086A

  • Wearable device and operation method thereof, wearable intelligent system and storage medium

    CN120319246A