Extended reality interaction method and apparatus, electronic device, and storage medium

By using augmented reality devices to collect audience status and generate speaking assistance information, the problem of insufficient attention to audience status in speaking scenarios is solved, thus improving the speaking experience and interactive effects.

CN121326155BActive Publication Date: 2026-06-16FALCON INNOVATIONS TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511864071.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-06-16
Estimated Expiration
2045-12-11

AI Technical Summary

Technical Problem

Current extended reality technology does not pay enough attention to the audience's state in speaking scenarios, resulting in a speaking experience that needs improvement.

Method used

By collecting audience status information through augmented reality devices, speech aids such as attention markers, suggestion markers, and related interest information are generated and displayed to help users optimize their speaking performance.

Benefits of technology

It enhances user interaction and experience during the speaking process, and improves audience attention and engagement by adjusting the content and style of the speech in real time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121326155B_ABST
    Figure CN121326155B_ABST
Patent Text Reader

Abstract

The application discloses an extended reality interaction method and device, electronic equipment and a storage medium, and relates to the technical field of extended reality, and the method comprises the following steps: in the process that an extended reality device is worn by a user to speak, collecting state information of a listener, and generating speech auxiliary information according to the state information, and displaying the speech auxiliary information. So that the user can collect the state of the listener when speaking to the listener by using the extended reality device, and generate auxiliary information in the speech process based on the state of the listener, so as to help the user to better speak to the listener, and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of extended reality technology, specifically to an extended reality interaction method, device, electronic device, and storage medium. Background Technology

[0002] Extended Reality (XR) is a comprehensive term that encompasses Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). Extended Reality technologies expand the ways humans perceive and interact with the real world, creating virtual, augmented, and mixed reality experiences through digital technology.

[0003] Currently, extended reality technology can be applied to speaking scenarios, such as translating during conversations, navigating dialogues, or giving speeches, enhancing the convenience and enjoyment of speaking to an audience.

[0004] However, current speech application scenarios do not pay enough attention to the audience's state, resulting in a speech experience that needs improvement. Summary of the Invention

[0005] This application provides an extended reality interaction method, device, electronic device, and storage medium that can enhance the speaking experience.

[0006] In a first aspect, embodiments of this application provide an extended reality interaction method applied to an extended reality device, the method comprising:

[0007] During the user's speech, the system collects the listener's status information;

[0008] Based on the status information, speech assistance information is generated;

[0009] Display the speech support information.

[0010] Secondly, embodiments of this application also provide an extended reality interactive device, applied to an extended reality device, the device comprising:

[0011] The data acquisition module is used to collect the listener's status information during the user's speech.

[0012] The generation module is used to generate speech assistance information based on the status information;

[0013] The display module is used to display the speech assistance information.

[0014] Optionally, in some embodiments of this application, the speech assistance information includes attention marker information; displaying the speech assistance information includes:

[0015] Locate the first target location information for the listener;

[0016] The attention marker information is displayed based on the first target location information.

[0017] Optionally, in some embodiments of this application, the speaking assistance information includes suggestion identification information; displaying the speaking assistance information includes:

[0018] Identify the reference object from which the wearer speaks;

[0019] The suggestion identification information is displayed on the reference object, so that the suggestion identification information is superimposed on the reference object;

[0020] The reference object includes the audience or the prompting interface. If the reference object is the prompting interface, the user speaking is based on the content displayed on the prompting interface.

[0021] Optionally, in some embodiments of this application, the speech assistance information includes associated interest information; generating speech assistance information based on the state information includes:

[0022] If the user speaks based on the prompting interface, then the associated interest information of the current speech content in the prompting interface is generated according to the status information;

[0023] The display of the speech assistance information includes:

[0024] Locate the second target location information for the listener, and display the associated interest information based on the second target location information.

[0025] Optionally, in some embodiments of this application, the apparatus further includes:

[0026] Collect the audience's voice information;

[0027] If the duration of the pause in the user's speech is detected to reach a preset threshold, then response information corresponding to the voice information is generated;

[0028] Display the response information.

[0029] Optionally, in some embodiments of this application, the state information includes distraction state information or focus state information;

[0030] The process of collecting listener status information during the user's speech includes:

[0031] During the user's speech, information about the listener's head orientation or eye gaze is collected;

[0032] If the head orientation information or the eye gaze information meets the first preset condition, then the listener's state information is determined to be focused state information;

[0033] If neither the head orientation information nor the eye gaze information meets the first preset condition, then the listener's state information is determined to be a distracted state.

[0034] Optionally, in some embodiments of this application, the user speaks based on the content displayed in the prompting interface, and the device further includes:

[0035] Collect the speaking status information of the wearer;

[0036] If the speech status information meets the second preset condition, the target content corresponding to the current speech is identified from the prompting interface, and annotation information for the target content is generated.

[0037] If the speaking status information meets the second preset condition and the ambient light intensity information meets the third preset condition, then the target content is enhanced based on the ambient light intensity information to obtain the target display result.

[0038] Thirdly, embodiments of this application also provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the extended reality interaction method described above.

[0039] Fourthly, embodiments of this application also provide a storage medium, which includes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the steps in the extended reality interaction method described above.

[0040] Fifthly, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in embodiments of this application.

[0041] In summary, the extended reality device of this application collects the listener's status information during the user's speech, generates and displays speech assistance information based on the status information.

[0042] In this application embodiment, when a user is speaking to an audience, they can use an extended reality device to collect the audience's state and generate auxiliary information during the speaking process based on the audience's state, so as to help the user speak to the audience better and improve the user experience. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a schematic diagram of a scenario in which the extended reality device provided in this application executes the extended reality interaction method;

[0045] Figure 2 This is a flowchart illustrating the extended reality interaction method provided in the embodiments of this application;

[0046] Figure 3 This is a schematic diagram of the structure of the extended reality interactive device provided in the embodiments of this application;

[0047] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.

[0048] Explanation of icon numbers:

[0049] 101-Extended Reality Device; 301-Acquisition Module; 302-Generation Module; 303-Display Module; 401-Processor; 402-Memory; 403-Power Supply; 404-Input Unit. Detailed Implementation

[0050] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] In the description of the embodiments of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. Furthermore, the terms "first," "second," "third," and "fourth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first," "second," "third," and "fourth" may explicitly or implicitly include one or more features. In the description of the present invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0052] In this application, the term "exemplary" is used to mean "serving as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0053] This application provides an extended reality interaction method, apparatus, electronic device, and computer-readable storage medium. Specifically, this application provides an extended reality interaction apparatus suitable for electronic devices (e.g., extended reality devices). The electronic device includes an extended reality device, which includes, but is not limited to, head-mounted displays, wearable glasses, and can be an integrated extended reality device built into a computing processing unit, or a separate extended reality device external to the computing processing unit. The extended reality device includes, but is not limited to, airborne optical display systems (i.e., head-up display systems) used in vehicles such as aircraft, automobiles, and ships, such as AR-HUD (Augmented Reality Head-Up Display) mounted on intelligent connected vehicles; extended reality applications (such as extended reality games, virtual tourism, telemedicine, or virtual experiments) used in handheld mobile devices such as mobile phones, laptops, and tablets; and near-eye display systems used in wearable devices such as head-mounted displays and smart glasses.

[0054] For example, please see Figure 1 , Figure 1 This is a schematic diagram of a scenario in which an extended reality device, according to an embodiment of this application, executes the extended reality interaction method. Specifically, the execution process of the extended reality device executing the extended reality interaction method is as follows:

[0055] The extended reality device 101 collects the listener's status information during the user's speech, and generates and displays speech assistance information based on the status information.

[0056] For example, when a user is speaking to an audience while wearing the extended reality device, the device can collect the audience's state, such as whether the audience is distracted. If the audience is distracted, it means that the speech is not effective. In this case, some auxiliary information can be generated for the speech, so that the user can use this auxiliary information to improve the speech effect for the audience.

[0057] In summary, the embodiments of this application enable users to use extended reality devices to collect the state of the audience when speaking to them, and generate auxiliary information during the speaking process based on the state of the audience, so as to help users speak to the audience better and improve the user experience.

[0058] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the priority of the embodiments.

[0059] Please see Figure 2 , Figure 2This is a flowchart illustrating an extended reality interaction method provided in an embodiment of this application. Although the flowchart shows a logical order, in some cases, the steps shown or described can be performed in a different order than that shown in the flowchart. Specifically, this extended reality interaction method is applied to an extended reality device, and the specific flow of the extended reality interaction method is as follows:

[0060] S201. Collect the listener's status information during the user's speech.

[0061] In this context, "wearing user" refers to the user who wears the augmented reality device. This can be understood as the user wearing the augmented reality device to speak to an audience. For example, in a speech scenario, the speaker is speaking to the audience, and during the speech, the speaker wears the augmented reality device (such as AR glasses).

[0062] It should be noted that status information refers to the audience's state while listening to the user speaking, reflecting their level of interest and focus on the content, and indirectly reflecting the effectiveness of the user's speech. For example, this status information includes focused state information or distracted state information. Distraction refers to the audience's attention not being on the current content of the speech. For example, if the audience is looking down at their phone or their gaze is not in the direction of the user, they are considered to be in a distracted state.

[0063] S202. Generate speech assistance information based on the status information.

[0064] It should be noted that speech assistance information is designed to help the wearer speak, for example, by improving the effectiveness of the speech and getting the audience back into a focused state. This speech assistance information could include points of interest for the audience, prompts for increasing the volume of the speech, or indicators for distracted listeners.

[0065] S203. Display the speech assistance information.

[0066] It is understood that the embodiments of this application display speech assistance information on the virtual screen of the extended reality device, allowing users to view this information on the virtual screen while speaking using the extended reality device. Users can then select or use this speech assistance information as a reference to optimize their speaking performance. For example, increasing the volume of their speech to attract the audience's attention, or presenting topics of interest to the audience to engage them in discussion.

[0067] In summary, the embodiments of this application enable users to use extended reality devices to collect the state of the audience when speaking to them, and generate auxiliary information during the speaking process based on the state of the audience, so as to help users speak to the audience better and improve the user experience.

[0068] Optionally, in some embodiments of this application, the speaking assistance information may be attentional alerts, which can be displayed at the location of the corresponding audience to help the wearer notice the audience's state. That is, optionally, in some embodiments of this application, the speaking assistance information includes attentional marker information, and the step "displaying the speaking assistance information" includes:

[0069] Locate the first target location information for the listener;

[0070] The attention marker information is displayed based on the first target location information.

[0071] The first target location information is the location information of the audience; that is, the first target location information is the location reference for displaying the attention marker information. For example, the first target location information could be the audience member's shoulder. Correspondingly, the attention marker information is displayed at the corresponding location on the virtual screen of the extended reality device, using the shoulder as the display reference. At this time, the user wearing the device can see a virtual attention marker hovering over the audience member's shoulder through the virtual screen of the extended reality device.

[0072] It is understood that, in the embodiments of this application, a three-dimensional model of the venue can be constructed using SLAM to locate the primary target location information of the audience.

[0073] It is understood that the first target location information is primarily intended to associate the attention marker information with the audience, enabling the wearer to locate the associated audience through the location of the attention marker information. Therefore, the first target location information can be not only the audience's shoulder, but also the audience's head, chest, side of the ear, etc. The specific location of the first target location information is not limited in this application embodiment.

[0074] In this embodiment of the application, the purpose of the attention marking information is to attract the wearer's attention. Therefore, the attention marking information can be the text "attention" or a graphic symbol of attention or warning. The specific type and content of the attention marking information are not limited here.

[0075] Optionally, for better display and to fully attract user attention, the attention marker can also be displayed using flashing, pulsed, or vibrating methods. For example, the attention marker could be displayed at a frequency of 3Hz for 3 seconds. Understandably, short-duration flashing not only attracts the wearer's attention but also avoids disrupting the user's speaking rhythm due to prolonged display, thus improving the speaking experience. This is especially important in speech scenarios, where short-duration flashing minimizes disruption to the delivery of information.

[0076] In this embodiment of the application, if the same listener is in a state that requires attention for a long time, in order to avoid continuously detecting the listener and displaying attention marker information, a display cycle can be set. For example, a display cycle of five or ten minutes can be set, and the listener's state information can be analyzed every five or ten minutes to determine whether to generate and display the speech assistance information.

[0077] Optionally, in embodiments of this application, in addition to displaying attention marker information to attract the wearer's attention, suggestions can also be generated to help optimize the speaking effect. By displaying these suggestions, users can quickly identify or choose reasonable measures in the speaking situation. That is, optionally, in some embodiments of this application, the speaking assistance information includes suggestion marker information, and the step of "displaying the speaking assistance information" includes:

[0078] Identify the reference object from which the wearer speaks;

[0079] The suggestion identification information is displayed on the reference object, so that the suggestion identification information is superimposed on the reference object;

[0080] The reference object includes the audience or the prompting interface. If the reference object is the prompting interface, the user speaking is based on the content displayed on the prompting interface.

[0081] The reference object is an object closely related to the speech, associated with the user or the content of the speech. For example, the reference object could be an audience member to whom the user is currently speaking, or it could be the teleprompter on which the speech is based (e.g., including the teleprompter interface). A teleprompter is an electronic device used to display text content, helping speakers, presenters, actors, or news anchors read the script fluently while maintaining natural eye contact (avoiding frequent looking down at the script). Understandably, the interface displaying the text content is the teleprompter interface.

[0082] Overlaying suggestion information onto a reference object helps associate the suggestion with the reference object, allowing the user to quickly perceive the target of the suggestion. For example, the suggestion could include increasing volume; when overlaid on the teleprompter interface, the user can understand that they need to increase their speaking volume when delivering content via the teleprompter. Another example is slowing down the speaking speed; when overlaid on the listener's body outline, the user can understand that they need to slow down their speaking speed to avoid making it difficult for the listener to understand. Furthermore, this suggestion can be generated based on the listener's facial expressions. For instance, if the listener's face indicates a thoughtful or deep contemplative state (e.g., blank stare, still facial expression, slightly tilted head, frowning), a suggestion to slow down the speaking speed is generated.

[0083] In this embodiment, directional audio enhancement can be used to alert or attract the attention of distracted listeners. For example, after identifying a specific distracted listener, the voice of the user (e.g., the speaker) can be captured through the microphone of an augmented reality device. Then, using directional beamforming audio processing technology, the volume is locally amplified only in the direction of the listener via the venue's sound system, effectively attracting their attention without affecting other listeners. Specifically, the user's voice is captured, and beamforming strategy and the venue's sound system are used to directionally enhance the voice, resulting in enhanced voice. This enhanced voice is then output to the sound system, causing the system to output directional enhanced voice (i.e., enhanced voice) specifically for the distracted listener.

[0084] Optionally, in this embodiment, the speaking assistance information can also be content that the audience is interested in. That is, optionally, in some embodiments of this application, the speaking assistance information includes related interest information, and the step of "generating speaking assistance information based on the state information" includes:

[0085] If the user speaks based on the prompting interface, then the associated interest information of the current speech content in the prompting interface is generated according to the status information;

[0086] The display of the speech assistance information includes:

[0087] Locate the second target location information for the listener, and display the associated interest information based on the second target location information.

[0088] For example, if the current speech content is an academic formula, the relevant interest information could be related knowledge points, background figures, stories, etc. If the current speech content is product performance parameters, the relevant interest information could be industry ratings, performance parameter comparison data of related products, etc.

[0089] Understandably, this extended reality interaction can also be applied in classroom scenarios. Teachers wearing the device can lecture students using content on a blackboard. If the device detects that students are distracted, it generates relevant interest information about the current lesson content. Teachers can then use this information to re-engage students' attention and adjust the classroom atmosphere. Furthermore, the device can display prompts, such as suggesting that the wearer guide the audience back to focus through voice or interactive questions.

[0090] In this embodiment, the second target location information is the location where the associated interest information is displayed. The second target location information can be the prompting interface, the black and white screen, or the side of the audience.

[0091] In this embodiment of the application, the related interest information can also be projected to the distracted audience member through targeted projection. For example, after identifying the distracted audience member and analyzing their points of interest, the generated related interest content can be projected directly into the distracted audience member's field of vision from a specific perspective (e.g., towards the user's perspective) through the display system of an extended reality device, achieving personalized content intervention without interrupting the speech process and guiding their attention back.

[0092] In addition, the relevant interest information can be sent to the audience's display device / mobile phone, so that when the user is distracted by browsing other display devices or mobile phones, they can see the relevant interest information through the display device or mobile phone.

[0093] Additionally, if the audience member is also wearing an augmented reality device, the relevant interest information can be sent to that device, allowing the audience member to directly receive the relevant interest information sent by the speaker's augmented reality device through their own device.

[0094] In this application embodiment, to enhance the interactive effect during the speech process, the audience's voice can be collected and analyzed, and responses generated to the voice can be generated at the end of a speech segment or during a pause, to help the wearer interact with the audience. Optionally, in some embodiments of this application, the method further includes:

[0095] Collect the audience's voice information;

[0096] If the duration of the pause in the user's speech is detected to reach a preset threshold, then response information corresponding to the voice information is generated;

[0097] Display the response information.

[0098] For example, after a user finishes a segment of a presentation and before starting another, there's a brief pause. During this interval, the audience often offers their insights, comments, or questions. The augmented reality device captures speech in real-time from all directions and analyzes it to determine its meaning. Based on this meaning, the device automatically generates feedback and guides the speaker's gaze (e.g., displaying arrows) to that direction. The feedback is then displayed in an augmented reality manner to encourage interaction between the speaker and the audience.

[0099] In this embodiment of the application, the voice of the listener can be collected by a microphone array set at the listener's location in a speaking scenario, and the voice of the listener can be obtained by receiving the transmission from the microphone array.

[0100] In this embodiment of the application, the listener's state information can be analyzed based on the listener's head orientation or eye gaze. Optionally, in some embodiments of this application, the state information includes distraction information or focused state information. The step "collecting listener state information during the user's speech" includes:

[0101] During the user's speech, information about the listener's head orientation or eye gaze is collected;

[0102] If the head orientation information or the eye gaze information meets the first preset condition, then the listener's state information is determined to be focused state information;

[0103] If neither the head orientation information nor the eye gaze information meets the first preset condition, then the listener's state information is determined to be a distracted state.

[0104] The first preset condition refers to the direction of the head or the line of sight of the eyes being directed towards the augmented reality device or the user wearing it. For example, when the listener's head or the line of sight of the eyes is directed towards the augmented reality device or the user wearing it, it is considered that the listener is in a focused state, and otherwise it is considered that the user is in a distracted state.

[0105] In this embodiment, the display of the extended reality device can also be adjusted according to the user's own state to better assist the user in speaking. That is, optionally, in some embodiments of this application, the user speaks based on the content displayed in the prompting interface. The method further includes:

[0106] Collect the speaking status information of the wearer;

[0107] If the speech status information meets the second preset condition, the target content corresponding to the current speech is identified from the prompting interface, and annotation information for the target content is generated.

[0108] If the speaking status information meets the second preset condition and the ambient light intensity information meets the third preset condition, then the target content is enhanced based on the ambient light intensity information to obtain the target display result.

[0109] The speech status information refers to the user's own state information, including eye movement and speech pauses. This can be determined by eye tracking; for example, if the user deviates from the anchor point for 0.5 seconds, it is considered eye movement. Alternatively, microphone analysis can be used to determine if the user is experiencing speech pauses; for example, silence exceeding 2 seconds or a sudden drop in speech rate of 50% is considered speech pauses.

[0110] The second preset condition refers to a negative state where the user's gaze wanders or their speech is interrupted.

[0111] Ambient light intensity information refers to the intensity of the ambient light at the speaking location. Meeting the third preset condition means that the ambient light intensity meets the conditions for enhanced content display. For example, when the ambient light suddenly dims (e.g., ≤50 lux), the user's speaking speed suddenly decreases, or the speech is interrupted, the target content corresponding to the current speech is enhanced, for example, by displaying the target content with a semi-transparent highlight layer. For instance, for the target content displayed in the prompting interface, it is enhanced in the virtual screen of the extended reality device using a semi-transparent highlight layer so that the user can still view the content of the prompting interface when the ambient light intensity changes. For example, when the ambient light suddenly dims, the prompting screen becomes brighter and more glaring than the surrounding environment. Displaying the target content with a semi-transparent highlight layer at this time helps the user clearly view the target content. In addition, other enhanced display methods can be designed, such as opaque highlight layer display or color display, which must ensure a good viewing effect when the ambient light intensity changes.

[0112] In addition, if a key word is detected during the user's speech, preset auxiliary phrases can automatically appear. For example, if the word "innovation" is detected, related cases will be displayed, and if the word "price" is detected, the online store price will automatically appear.

[0113] In this embodiment of the application, when speaking through a teleprompter, the scrolling speed of the text in the teleprompter can be adjusted according to the user's speaking speed, avoiding the problem of mismatch between the text scrolling progress and the speaking progress caused by mechanical scrolling for a specific duration.

[0114] In summary, the embodiments of this application enable users to use extended reality devices to collect the state of the audience when speaking to them, and generate auxiliary information during the speaking process based on the state of the audience, so as to help users speak to the audience better and improve the user experience.

[0115] During the speech, when specific events or behaviors of the audience are detected (such as users dozing off, users talking, users looking away), the system analyzes the topics that the user is interested in based on user profiles or speech recognition results, generates content, and presents it to the speaker in an augmented reality manner to prompt the speaker to begin the speech.

[0116] To facilitate better implementation of the extended reality interaction method of this application, this application also provides an extended reality interaction device based on the above-described extended reality interaction method. The meanings of the terms used are the same as in the extended reality interaction method described above, and specific implementation details can be found in the descriptions of the method embodiments.

[0117] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of the extended reality interaction device provided in the embodiments of this application, wherein the extended reality interaction device is applied to an extended reality device, and the extended reality interaction device can be specifically as follows:

[0118] The data acquisition module 301 is used to collect the listener's status information during the user's speech.

[0119] The generation module 302 is used to generate speech assistance information based on the status information;

[0120] Display module 303 is used to display the speech assistance information.

[0121] Optionally, in some embodiments of this application, the speech assistance information includes attention marker information; displaying the speech assistance information includes:

[0122] Locate the first target location information for the listener;

[0123] The attention marker information is displayed based on the first target location information.

[0124] Optionally, in some embodiments of this application, the speaking assistance information includes suggestion identification information; displaying the speaking assistance information includes:

[0125] Identify the reference object from which the wearer speaks;

[0126] The suggestion identification information is displayed on the reference object, so that the suggestion identification information is superimposed on the reference object;

[0127] The reference object includes the audience or the prompting interface. If the reference object is the prompting interface, the user speaking is based on the content displayed on the prompting interface.

[0128] Optionally, in some embodiments of this application, the speech assistance information includes associated interest information; generating speech assistance information based on the state information includes:

[0129] If the user speaks based on the prompting interface, then the associated interest information of the current speech content in the prompting interface is generated according to the status information;

[0130] The display of the speech assistance information includes:

[0131] Locate the second target location information for the listener, and display the associated interest information based on the second target location information.

[0132] Optionally, in some embodiments of this application, the apparatus further includes:

[0133] Collect the audience's voice information;

[0134] If the duration of the pause in the user's speech is detected to reach a preset threshold, then response information corresponding to the voice information is generated;

[0135] Display the response information.

[0136] Optionally, in some embodiments of this application, the state information includes distraction state information or focus state information;

[0137] The process of collecting listener status information during the user's speech includes:

[0138] During the user's speech, information about the listener's head orientation or eye gaze is collected;

[0139] If the head orientation information or the eye gaze information meets the first preset condition, then the listener's state information is determined to be focused state information;

[0140] If neither the head orientation information nor the eye gaze information meets the first preset condition, then the listener's state information is determined to be a distracted state.

[0141] Optionally, in some embodiments of this application, the user speaks based on the content displayed in the prompting interface, and the device further includes:

[0142] Collect the speaking status information of the wearer;

[0143] If the speech status information meets the second preset condition, the target content corresponding to the current speech is identified from the prompting interface, and annotation information for the target content is generated.

[0144] If the speaking status information meets the second preset condition and the ambient light intensity information meets the third preset condition, then the target content is enhanced based on the ambient light intensity information to obtain the target display result.

[0145] In this embodiment, the acquisition module 301 first acquires the listener's status information during the user's speech, the generation module 302 generates speech assistance information based on the status information, and the display module 303 displays the speech assistance information.

[0146] In summary, the embodiments of this application enable users to use extended reality devices to collect the state of the audience when speaking to them, and generate auxiliary information during the speaking process based on the state of the audience, so as to help users speak to the audience better and improve the user experience.

[0147] In addition, this application also provides an electronic device, such as Figure 4 As shown, it illustrates a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Specifically:

[0148] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0149] The processor 401 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0150] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0151] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power equipment debugging circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0152] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0153] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402, thereby implementing the steps in any of the extended reality interaction methods provided in the embodiments of this application.

[0154] The extended reality device of this application collects the listener's status information during the user's speech, generates and displays speech assistance information based on the status information.

[0155] In this application embodiment, when a user is speaking to an audience, they can use an extended reality device to collect the audience's state and generate auxiliary information during the speaking process based on the audience's state, so as to help the user speak to the audience better and improve the user experience.

[0156] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0157] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0158] To this end, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor to perform the steps in any of the extended reality interaction methods provided in this application.

[0159] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0160] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0161] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the extended reality interaction methods provided in this application, the beneficial effects that any of the extended reality interaction methods provided in this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0162] The above provides a detailed description of an extended reality interaction method, apparatus, electronic device, and computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. An extended reality interaction method, characterized in that, Applied to extended reality devices, the method includes: During the user's speech, the system collects the listener's status information; If the state information indicates a distracted state, then speech assistance information is generated based on the state information; The speech assistance information is displayed to assist the wearer in speaking, so that the distracted listener can return to a state of focused listening to the speech. The display of the speech assistance information includes: Locate the first target location information of the listener corresponding to the distracted state from among multiple listeners; The speech assistance information is displayed based on the first target location information; The display of the speech assistance information includes: overlaying the speech assistance information onto the extended reality device; The method further includes: The user speaks based on the content displayed in the prompting interface. Collect the speaking status information of the wearer; If the speech status information meets the second preset condition, the target content corresponding to the current speech is identified from the prompting interface, and annotation information for the target content is generated. If the speaking status information meets the second preset condition and the ambient light intensity information meets the third preset condition, then the target content is enhanced based on the ambient light intensity information to obtain the target display result.

2. The extended reality interaction method according to claim 1, characterized in that, The speech assistance information includes attention marker information; displaying the speech assistance information includes: Locate the first target location information for the listener; The attention marker information is displayed based on the first target location information.

3. The extended reality interaction method according to claim 1, characterized in that, The speech assistance information includes suggestion identification information; displaying the speech assistance information includes: Identify the reference object from which the wearer speaks; The suggestion identification information is displayed on the reference object, so that the suggestion identification information is superimposed on the reference object; The reference object includes the audience or the prompting interface. If the reference object is the prompting interface, the user speaking is based on the content displayed on the prompting interface.

4. The extended reality interaction method according to claim 1, characterized in that, The speech assistance information includes related interest information; the step of generating speech assistance information based on the status information includes: If the user speaks based on the prompting interface, then the associated interest information of the current speech content in the prompting interface is generated according to the status information; The display of the speech assistance information includes: Locate the second target location information for the listener, and display the associated interest information based on the second target location information.

5. The extended reality interaction method according to claim 1, characterized in that, The method further includes: Collect the audience's voice information; If the duration of the pause in the user's speech is detected to reach a preset threshold, then response information corresponding to the voice information is generated; Display the response information.

6. The extended reality interaction method according to claim 1, characterized in that, The state information includes distraction state information or focus state information; The process of collecting listener status information during the user's speech includes: During the user's speech, information about the listener's head orientation or eye gaze is collected; If the head orientation information or the eye gaze information meets the first preset condition, then the listener's state information is determined to be focused state information; If neither the head orientation information nor the eye gaze information meets the first preset condition, then the listener's state information is determined to be a distracted state.

7. An extended reality interactive device, characterized in that, Applied to an extended reality device, the device includes: The data acquisition module is used to collect the listener's status information during the user's speech. The generation module is used to generate speech assistance information based on the state information if the state information indicates a distracted state. The display module is used to display the speech assistance information, which is used to assist the wearer in speaking and to help the distracted listener return to a state of focused listening to the speech. The display of the speech assistance information includes: Locate the first target location information of the listener corresponding to the distracted state from among multiple listeners; The speech assistance information is displayed based on the first target location information; The display of the speech assistance information includes: overlaying the speech assistance information onto the extended reality device; The device further includes: The user speaks based on the content displayed in the prompting interface. Collect the speaking status information of the wearer; If the speech status information meets the second preset condition, the target content corresponding to the current speech is identified from the prompting interface, and annotation information for the target content is generated. If the speaking status information meets the second preset condition and the ambient light intensity information meets the third preset condition, then the target content is enhanced based on the ambient light intensity information to obtain the target display result.

8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the extended reality interaction method as described in any one of claims 1-6.

9. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the extended reality interaction method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Conference system-based conference participant monitoring processing method and device, and intelligent terminal

    CN113783709A

  • Analysis method based on visual scene analysis algorithm model and AI glasses

    CN120412048A

  • Intelligent dialogue auxiliary method and device based on AR glasses and electronic equipment

    CN120748009A