Personalized and careful transcription for auditory experience to improve user engagement

Through personalized speech-to-text transcription systems, using biomarkers and machine learning models to selectively transcribe keywords, the problems of distraction and information omission in traditional systems are solved, and the effectiveness of user engagement and auditory experience is improved.

CN120418756APending Publication Date: 2025-08-01CTRL-LABS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380088439.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-26
Filing Date
2023-12-18
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Traditional pronunciation-to-text systems, when providing verbatim transcription, cause user attention distraction, reduce engagement, and fail to effectively evaluate user attention status and dialogue context, resulting in information omission or visual confusion.

Method used

By calculating user attention status, word comprehensibility, and contextual importance, using biomarkers and machine learning models, keywords are individually selectively transcribed, reducing unnecessary text output.

Benefits of technology

Increase user engagement, reduce visual confusion, ensure users maintain attention on key information, and enhance the effectiveness of auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120418756A_ABST
    Figure CN120418756A_ABST
Patent Text Reader

Abstract

One embodiment of the present invention sets forth a technique for developing a careful transcription of an acoustic experience including spoken words. The technique includes calculating a first metric for a word based on a biomarker associated with a user, wherein the metric indicates a state of attention of the user perceiving the word during an auditory experience. The technique further includes calculating a second metric corresponding to the intelligibility of the word during the auditory experience. The technique further includes calculating a third measure of importance of the word in the context of other close words during the auditory experience. Based on a weight assigned to each of the first metric, the second metric, and the third metric, the technique includes determining whether to transcribe the word on a display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various embodiments generally relate to speech-to-text transcription, and more specifically, to techniques for providing personalized and curated speech-to-text transcription for an auditory experience to increase user engagement. Background Art

[0002] Speech-to-text systems typically identify spoken words in a speech-based auditory experience (e.g., speech, conversations, audiobooks, music lyrics, movies, etc.) and produce a text output that includes all the spoken words that the system is able to identify. However, users participating in and listening to a speech-based auditory experience may not need a transcription of all the spoken words to keep up. For users who may be temporarily distracted or do not understand parts of a conversation, accessing just a few missed keywords may be sufficient to follow what is being discussed. In fact, for some users, providing a transcription of the entire conversation may cause them to disengage from the auditory experience. Additionally, providing a verbatim transcription of a conversation can be frustrating for users because inundating the user with text limits cognitive access to body language or lip-reading cues, which has the effect of reducing auditory presence and social connection.

[0003] In the case of dual-tasking, which is becoming increasingly common, providing a verbatim transcription may also be excessive. For example, in the case where a transcription of an audiobook is presented on a near-eye display (NED) system (e.g., smart glasses) while a user is out walking or driving, the verbatim transcription of the audiobook may clutter the display and distract the user. Even in a single-tasking scenario, a verbatim transcription may be unappealing. For example, in the case where a user is trying to participate in a classroom lecture but having difficulty concentrating, the information clutter on the NED may cause a further loss of concentration.

[0004] The speech-to-text system is further limited by the fact that the cognitive allocation of visual attention to the transcription of all spoken words produced by such a system would reduce the remaining bandwidth available for any other sensory cues. As a result, the transcription produced by such a system may have the unintended consequence of disengaging certain users from the experience (e.g., a user distracted by the transcription of the dialogue in a movie presented on the display of an NED may be disengaged from the visual experience of the movie). Additionally, there may be significant differences in users' sensitivity to visual clutter, and thus, any standardized method of providing a transcription auditory experience by a traditional speech-to-text system that does not take into account the subjective experience of the user is insufficient to increase user engagement. Furthermore, selective transcription has not been successful because users vary significantly in their ability to concentrate their attention, and it is challenging for any speech-to-text system to determine what information to present. Additionally, general solutions do not evaluate the context of the speech, and thus, there is a risk that critical information may be omitted.

[0005] As described above, there is a need in the art for a more effective method of transcribing an auditory experience. Summary of the Invention

[0006] According to a first aspect, there is provided a computer-implemented method, the method comprising: calculating a first metric for a word, the first metric indicating an attention state of a user perceiving the word during an auditory experience, wherein the metric is calculated based on a biomarker associated with the user; calculating a second metric corresponding to the intelligibility of the word during the auditory experience; calculating the importance of the word in the context of one or more words spoken in proximity to the word during the auditory experience; and calculating a fourth metric based on weights assigned to each of the first metric, the second metric, and the third metric; and determining whether to transcribe the word for display based on the fourth metric.

[0007] The fourth metric can be used to determine whether to transcribe the word for display by comparing the fourth metric to a predetermined threshold. If the fourth metric is higher than the predetermined threshold, the word can be displayed on a display.

[0008] The display can be included within a near-eye display (NED) system operating in an artificial reality environment.

[0009] The biomarker can include one or more of the following: blink rate; pupil dilation; fixation stability; or fixation acceleration.

[0010] A biomarker can be calculated using signals from one or more sensors, which include electroencephalography (EEG) electrodes, electrooculography (EOG) electrodes, functional near-infrared spectroscopy (fNIRS) optodes, or multi-wavelength photoplethysmography (MW-PPG) sensors.

[0011] The first metric, the second metric, and the third metric can each include a numerical value ranging between 0 and 1.

[0012] Calculating a second metric corresponding to the comprehensibility of the word can include: estimating the audibility of the word using the signal-to-noise ratio (SNR) associated with the word.

[0013] Calculating a third metric corresponding to the importance of the word can include: using a machine learning model to determine the importance of the word. The machine learning model can include one of the following: recurrent neural network (RNN), convolutional neural network (CNN), deep neural network (DNN), deep convolutional network (DCN), residual neural network (ResNet), graph neural network, autoencoder, transformer neural network, deep stereo geometry network (DSGN), or region-based convolutional neural network (R-CNN).

[0014] The assignment of the weights can include using a machine learning model to determine the optimal weights. The machine learning model can include one of the following: recurrent neural network (RNN), convolutional neural network (CNN), deep neural network (DNN), deep convolutional network (DCN), residual neural network (ResNet), graph neural network, autoencoder, transformer neural network, deep stereo geometry network (DSGN), or stereo R-CNN.

[0015] The assignment of the weights can include performing a regression analysis.

[0016] The weights can be selectively assigned by a user.

[0017] The assignment of the weights may be based on the following combination: automatic selection; and user-based selection.

[0018] According to a second aspect, there is provided one or more computer-readable media storing multiple instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to the first aspect. The media may be non-transitory.

[0019] According to a third aspect, there is provided a wearable device comprising: one or more memories storing instructions; and one or more processors coupled to the one or more memories and configured to perform the method according to the first aspect when executing the instructions.

[0020] According to a fourth aspect, there is provided a computer program product comprising multiple instructions which, when the program is executed by one or more processors of a computer, cause the one or more processors to perform the method according to the first aspect.

[0021] At least one advantage of the disclosed technology is that users of a personalized speech-to-text transcription system can rely on curated transcriptions to grasp certain spoken words that the user did not auditorily perceive during the auditory experience. The transcription system personalizes and curates the transcription based on cues and biomarkers collected from both the spoken words and the user, which not only increases user engagement but also prevents the user from being overwhelmed by a verbatim transcription of the auditory experience and the accompanying visual clutter. By relying on cues from the user (e.g., physiological cues that determine the user's attention state), the transcription can also be selectively curated for each individual user, which further increases user engagement. User engagement depends on rapid task switching between real-time audio perception and visual perception. The display of keyword selection can help minimize the user's reading effort as much as possible while still providing enough information to generate context. For these reasons, the disclosed technology represents a technological advancement compared to previous methods that verbatim transcribed speech and resulted in low user engagement. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To understand the above features of various embodiments in detail, reference may be made to a more specific description of the inventive concept briefly outlined above in terms of various embodiments, some of which are illustrated in the drawings. However, it should be noted that the drawings only show typical embodiments of the inventive concept and should not be considered to limit the scope in any way, and there are other equivalent embodiments.

[0023] Figure 1Block diagram of an embodiment of a near-eye display (NED) system in which a console operates according to various embodiments.

[0024] Figure 2A Schematic diagram of a NED according to various embodiments.

[0025] Figure 2B Schematic diagram of a NED according to various embodiments. In various embodiments, the NED presents media to a user.

[0026] Figure 3 Illustrates techniques for determining personalized and curated transcripts of an auditory experience according to various embodiments.

[0027] Figure 4 Flowchart of the multiple steps of a method for developing a curated transcript of an acoustic experience according to various embodiments. Detailed Description

[0028] In the following description, numerous specific details are set forth to provide a thorough understanding of the various embodiments. However, it will be apparent to one of ordinary skill in the art that the concepts of the present invention may be practiced without one or more of these specific details.

[0029] As mentioned above, traditional speech-to-text systems typically produce a transcript of all spoken words of a conversation, which may not be ideal for a user who is participating in the conversation but misses some words either due to being temporarily distracted or due to some words in the conversation being difficult to understand. Providing a complete transcript of the conversation on a display (especially in the case of a near-eye display (NED) system such as smart glasses) is not ideal because it may increase the user's cognitive load and may further distract the user from the auditory experience. Thus, traditional speech-to-text systems do not enable the user to effectively participate.

[0030] To address these issues, various embodiments include a transcription engine configured to intelligently distribute cognitive load between the visual and auditory domains by providing a personalized and curated transcription of an auditory experience (e.g., a conversation) to a user on a display screen (e.g., on a display of the NED). The transcription displays on the display only those keywords determined to be relevant to the user, while omitting or de-emphasizing words not determined to be relevant. Relevance can be determined based on a combination of one or more of the following factors: a) a determination of how intelligible a particular spoken word is to the user (e.g., using an audibility estimation method such as the Speech Intelligibility Index); b) an estimation of the importance of a word in the context of a conversation (e.g., using a machine learning model); and c) using various biomarkers (e.g., (i) an estimation of blink rate, pupil dilation, gaze stability, and gaze acceleration obtained from sensors located on the NED or on the user's face); (ii) an EEG signal that estimates the intensity and / or fluctuations of the user's alpha (α) and theta (θ) waves, the EEG signal being obtained using sensors located on the NED or placed on the user's head or in the user's ears) to determine the user's attention state. In some embodiments, the user is allowed to control the weight assigned to each of factors (a), (b), and (c).

[0031] At least one advantage of the disclosed technology is that users of a personalized speech-to-text transcription system can rely on the curated transcription to grasp certain spoken words that the user did not auditorily perceive during an auditory experience. The transcription system personalizes and curates the transcription based on physiological cues and other biomarkers collected from both the spoken words and the user, which not only increases user engagement but also prevents the user from being overwhelmed by a verbatim transcription of the auditory experience. By relying on cues or biomarkers from the user (e.g., physiological cues that determine the user's attention state), the transcription is also selectively curated for each individual user, which further increases user engagement. Thus, compared to previous methods of verbatim transcription of speech that resulted in low user engagement, the disclosed technology represents a technological advancement.

[0032] Embodiments of the present disclosure may include an artificial reality system or be implemented in combination with an artificial reality system. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user. Artificial reality may include, for example, a virtual reality (VR) system, an augmented reality (AR) system, a mixed reality (MR) system, a hybrid reality system, or some combination and / or derivative thereof. Artificial reality content may include, but is not limited to, fully generated content or content generated in combination with captured (e.g., real-world) content. Artificial reality content may include, but is not limited to, video, audio, haptic feedback, or some combination thereof. Artificial reality content may be presented in a single channel or multiple channels (e.g., a stereoscopic video that gives a three-dimensional effect to a viewer). Additionally, in some embodiments, the artificial reality system may also be associated with an application, a product, an accessory, a service, or some combination thereof, which is used, for example, to create content in the artificial reality system and / or otherwise used in the artificial reality system (e.g., perform an activity in the artificial reality system). The artificial reality system may be implemented on various platforms, including a wearable head-mounted display (HMD) connected to a host computer system, a stand-alone HMD, a mobile device or computing system, or any other hardware platform capable of providing artificial reality content to one or more viewers.

[0033] Note that although NED and head-mounted display (HMD) are disclosed herein as reference examples, the transcription engine disclosed herein may also operate on other types of wearable and non-wearable display elements and display devices, including, for example, display elements and devices that can be configured to be placed at a fixed position near one or both eyes of a user without being head-mounted (e.g., a display device can be mounted in a vehicle (e.g., a car or an airplane) to be placed in front of one or both eyes of a user). Additionally, embodiments of the present disclosure are not limited to being implemented in combination with an artificial reality system and may also be implemented using other types of audiovisual systems.

[0034] Figure 1 is a block diagram of an embodiment of a near-eye display (NED) system 100 in which a console operates according to various embodiments. The NED system 100 may operate in a virtual reality (VR) system environment, an augmented reality (AR) system environment, a mixed reality (MR) system environment, or some combination thereof. Figure 1The NED system 100 shown in [figure] includes a NED 105 and an input / output (I / O) interface 175 coupled to a console 170. In various embodiments, the composite display system 100 is included in or operates in conjunction with the NED 105. For example, the composite display system 100 may be included within the NED 105 or may be coupled to the console 170 and / or the NED 105.

[0035] Although Figure 1 the example NED system 100 is shown including one NED 105 and one I / O interface 175, in other embodiments, the NED system 100 may include any number of these components. For example, there may be multiple NEDs 105, and each NED 105 has an associated I / O interface 175. Each NED 105 and I / O interface 175 communicate with the console 170. In an alternative configuration, the NED system 100 may include different and / or additional components. Additionally, in some embodiments, the various components included within the NED 105, the console 170, and the I / O interface 175 may be distributed in a manner different from that described in connection with Figures 1 to 2B For example, some or all of the functions of the console 170 may be provided by the NED 105 and vice versa.

[0036] The NED 105 may be a head-mounted display that presents content to a user. The content may include: virtual and / or augmented views of the physical real-world environment, which views include computer-generated elements (e.g., two-dimensional or three-dimensional images, two-dimensional or three-dimensional video, sound, etc.). In some embodiments, the NED 105 may also present audio content to the user. The NED 105 and / or the console 170 may send the audio content to an external device via the I / O interface 175. The external device may include various forms of speaker systems and / or headphones. In various embodiments, the audio content is synchronized with the visual content being displayed by the NED 105. In some embodiments, the visual content includes a transcription of the audio content captured using a microphone 176 included in the NED 105, which transcription is used to assist the user in understanding the audio content.

[0037] The NED 105 may include one or more rigid bodies that may be rigidly or non-rigidly coupled to each other. A rigid coupling between multiple rigid bodies may cause the coupled rigid bodies to act as a single rigid entity. In contrast, a non-rigid coupling between multiple rigid bodies may allow the rigid bodies to move relative to each other.

[0038] As Figure 1As shown, the NED 105 may include EOG electrodes 110, a depth camera assembly (DCA) 155, one or more locators 120, a display 125, an optical component 130, one or more position sensors 135, an inertial measurement unit (IMU) 140, an eye tracking system 145, EEG electrodes 190, an optical sensor 195, and a zoom module 150. In some embodiments, the display 125 and the optical component 130 may be integrated together into a projection component. Compared with the components listed above, various embodiments of the NED 105 may have more, fewer, or different components. Additionally, in various embodiments, the function of each component may be partially or fully subsumed by the function of one or more other components.

[0039] The DCA 155 acquires sensor data as follows: the sensor data describes depth information of an area around the NED 105. The sensor data may be generated by one or a combination of multiple depth imaging techniques, such as triangulation, structured light imaging, time-of-flight imaging, stereoscopic imaging, and laser scanning. The DCA 155 may use the sensor data to calculate various depth attributes of the area around the NED 105. Additionally or alternatively, the DCA 155 may send the sensor data to the console 170 for processing. Further, in various embodiments, the DCA 155 acquires or samples the sensor data at different times. For example, the DCA 155 may sample the sensor data at different times within a time window to obtain the sensor data along the time dimension.

[0040] The DCA 155 includes a light source, an imaging device, and a controller. The light source emits light onto the area around the NED 105. In one embodiment, the emitted light is structured light. The light source includes a plurality of emitters, each emitter emitting light having certain characteristics (e.g., wavelength, polarization, coherence, temporal behavior, etc.). These characteristics can be the same or different among the emitters, and the emitters can operate simultaneously or individually. In one embodiment, the plurality of emitters can be, for example, laser diodes (e.g., edge emitters), inorganic or organic light-emitting diodes (LEDs), vertical-cavity surface-emitting lasers (VCSELs), or some other source. In some embodiments, a single emitter or a plurality of emitters in the light source can emit light having a structured light pattern. The imaging device collects not only the light generated by the plurality of emitters and reflected by objects in the environment but also the ambient light in the environment around the NED 105. In various embodiments, the imaging device can be an infrared camera or can be a camera configured to operate in the visible light spectrum. The controller coordinates how the light source emits light and how the imaging device collects light. For example, the controller can determine the brightness of the emitted light. In some embodiments, the controller also analyzes the detected light to detect objects in the environment and the position information associated with these objects.

[0041] Each locator 120 is an object located at a specific position on the NED 105 relative to each other and relative to a specific reference point on the NED 105. The locator 120 can be a light-emitting diode (LED), a corner cube reflector, a reflective marker, a type of light source that contrasts with the environment in which the NED 105 operates, or some combination thereof. In embodiments where the locator 120 is an active component (e.g., an LED or other type of light-emitting device), these locators 120 can emit light in the visible light band (e.g., from about 380 nanometers (nm) to 950 nm), in the infrared (IR) band (from about 950 nm to 9700 millimeters (mm)), in the ultraviolet band (70 nm to 380 nm), in another part of the electromagnetic spectrum, or some combination of the above lights.

[0042] In some embodiments, the locator 120 is located beneath the outer surface of the NED 105, which is transparent to the wavelength of the light emitted by the locator 120 or the wavelength of the light reflected by the locator 120, or the outer surface is thin enough so as to substantially not attenuate the wavelength of the light emitted by the locator 120 or the wavelength of the light reflected by the locator 120. Additionally, in some embodiments, the outer surface or other portions of the NED 105 are opaque in the visible light band of light. Accordingly, the locator 120 can emit light in the IR band beneath an outer surface that is transparent in the IR band but opaque in the visible light band.

[0043] The display 125 displays two-dimensional or three-dimensional images to the user based on pixel data received from the console 170 and / or one or more other sources. In various embodiments, the display 125 includes a single display or multiple displays (e.g., separate displays for each eye of the user). In some embodiments, the display 125 includes a single or multiple waveguide displays. Light can be coupled into the single or multiple waveguide displays through a display such as a liquid crystal display (LCD), an organic light emitting diode (OLED) display, an inorganic light emitting diode (ILED) display, an active-matrix organic light-emitting diode (AMOLED) display, a transparent organic light emitting diode (TOLED) display, a laser-based display, one or more waveguides, other types of displays, a scanner, a one-dimensional array, etc. Additionally, combinations of display types can be incorporated into the display 125 and used individually, in parallel, and / or in combination.

[0044] The optical assembly 130 amplifies the received image light from the display 125, corrects the optical errors associated with the image light, and presents the corrected image light to the user of the NED 105. The optical assembly 130 includes a plurality of optical elements. For example, the optical assembly 130 may include one or more of the following optical elements: an aperture; a Fresnel lens; a convex lens; a concave lens; a filter; a reflective surface; or any other suitable optical element that deflects, reflects, refracts, and / or otherwise modifies the image light. Additionally, the optical assembly 130 may include a combination of different optical elements. In some embodiments, one or more of the plurality of optical elements of the optical assembly 130 may have one or more coatings, such as a partially reflective coating or an anti-reflective coating.

[0045] In some embodiments, the optical assembly 130 may be designed to correct one or more types of optical errors. Examples of optical errors include barrel distortion or pincushion distortion, longitudinal chromatic aberration or lateral chromatic aberration. Other types of optical errors may also include spherical aberration; chromatic aberration; or errors caused by lens field curvature, astigmatism; and other types of optical errors. In some embodiments, the visual content transmitted to the display 125 is pre-distorted, and when the image light from the display 125 passes through the various optical elements of the optical assembly 130, the optical assembly 130 corrects the distortion. In some embodiments, the optical elements of the optical assembly 130 are integrated into the display 125 as a projection assembly that includes at least one waveguide coupled to one or more optical elements.

[0046] The microphone 176 captures audio content. In some embodiments, an array of microphones 176 may also be used, where various microphones are required to implement beamforming and direction of arrivals (DOA). The microphone 176 is capable of receiving spoken voice from the user or any other person. The microphone may be connected to the console 170 using or any other type of wireless or wired technology. In some embodiments, the user may speak commands into the microphone, and these commands are executed by the console 170 to perform an action. In some embodiments, a transcription engine 185 in the console 170 may transcribe the voice or external sound input to the microphone from the user and display it on the display 125.

[0047] The IMU 140 is an electronic device that generates data indicating the position of the NED 105 based on multiple measurement signals received from one or more of the plurality of position sensors 135 and depth information received from the DCA 155. In some embodiments of the NED 105, the IMU 140 may be a dedicated hardware component. In other embodiments, the IMU 140 may be a software component implemented in one or more processors.

[0048] In operation, the position sensors 135 generate one or more measurement signals in response to the movement of the NED 105. Examples of the position sensors 135 include one or more accelerometers, one or more gyroscopes, one or more magnetometers, one or more altimeters, one or more inclinometers, and / or various types of sensors for motion detection, drift detection, and / or error detection. The position sensors 535 may be located outside the IMU 140, inside the IMU 140, or some combination thereof.

[0049] The IMU 140 generates data indicating an estimated position of the near-eye display 105 relative to an initial position of the near-eye display 105 based on one or more measurement signals from one or more of the position sensors 135. For example, the position sensors 135 include multiple accelerometers for measuring translational motion (forward / backward, up / down, left / right) and multiple gyroscopes for measuring rotational motion (e.g., pitch, yaw, and roll). In some embodiments, the IMU 140 samples the measurement signals rapidly and calculates the estimated current position of the NED 105 based on the sampled data. For example, the IMU 140 integrates the measurement signals received from the accelerometers over time to estimate the velocity vector and integrates the velocity vector over time to determine the estimated current position of a reference point on the NED 105. Alternatively, the IMU 140 provides the sampled measurement signals to the console 170, which analyzes the sampled data to determine one or more measurement errors. The console 170 may also send one or more control signals and / or measurement errors to the IMU 140 to configure the IMU 140 to correct and / or reduce one or more measurement errors (e.g., drift errors). The reference point is a point that can be used to describe the position of the NED 105. The reference point can generally be defined as a point in space or a position related to the position and / or orientation of the NED 105.

[0050] In various embodiments, the IMU 140 receives one or more parameters from the console 170. The one or more parameters are used to maintain tracking of the NED 105. The IMU 140 may adjust one or more IMU parameters (e.g., sampling rate) based on the received parameters. In some embodiments, certain parameters cause the IMU 140 to update the initial position of the reference point so that the IMU corresponds to the next position of the reference point. Updating the initial position of the reference point to the next calibrated position of the reference point helps reduce drift error in detecting the current position estimate of the IMU 140.

[0051] In various embodiments, the eye tracking system 145 is integrated into the NED 105. The eye tracking system 145 may include one or more illumination sources (e.g., infrared illumination sources, visible light illumination sources) and one or more imaging devices (e.g., one or more cameras). In operation, when the user wears the NED 105, the eye tracking system 145 generates tracking data related to the user's eyes and analyzes the tracking data. In various embodiments, the eye tracking system 145 estimates the angular orientation of the user's eyes. This orientation of the eyes corresponds to the user's gaze direction within the NED 105. The orientation of the user's eyes is defined herein as the direction of the foveal axis, which is the axis between the fovea (the area of the retina of the eye with the highest concentration of photoreceptors) and the pupil center of the eye. Generally, when the user's eyes are fixed on a point, the foveal axis of the user's eyes will intersect that point. The pupil axis is another axis of the eye, which is defined as the axis passing through the pupil center and perpendicular to the corneal surface. Generally, the pupil axis does not align directly with the foveal axis. These two axes intersect at the pupil center, but the orientation of the foveal axis is laterally offset from the pupil axis by approximately -1° to 8° and longitudinally offset by approximately ±4°. Since the foveal axis is defined based on the fovea located at the back of the eye, it may be difficult or impossible to directly detect the foveal axis in some eye tracking embodiments. Therefore, in some embodiments, the orientation of the pupil axis is detected and the foveal axis is estimated based on the detected pupil axis.

[0052] Typically, eye movements not only correspond to angular rotations of the eyes, but also to translations of the eyes, changes in eye torsion, and / or changes in eye shape. The eye tracking system 145 may also detect translations of the eyes, i.e., changes in the position of the eyes relative to the eye socket. In some embodiments, the translations of the eyes are not directly detected, but are approximated based on a mapping from the detected angular orientations. Eye translations corresponding to changes in the position of the eyes relative to the detection components of the eye tracking unit may also be detected. This type of translation may occur, for example, due to movement of the NED 105 in the position of the user's head. The eye tracking system 145 may also detect eye torsion, i.e., rotation of the eyes about the pupil axis. The eye tracking system 145 may use the detected eye torsion to estimate the orientation of the fovea axis relative to the pupil axis. The eye tracking system 145 may also track changes in eye shape, which may be approximated as a skew or scaled linear transformation or a distortion (e.g., due to a torsional distortion). The eye tracking system 145 may estimate the fovea axis based on some combination of the angular orientation of the pupil axis, the translation of the eyes, the torsion of the eyes, and the current shape of the eyes.

[0053] Since the orientation of the user's binocular eyes can be determined, the eye tracking system 145 is able to determine where the user is looking. The NED 105 may use the orientation of the eyes to, for example, determine the user's inter-pupillary distance (IPD), determine the gaze direction, introduce depth cues (e.g., a blurred image outside the user's primary line of sight), collect heuristics regarding user interaction in VR media (e.g., the time spent on any particular topic, object, or frame as a function of the stimuli encountered), some other function based at least in part on the orientation of at least one of the user's eyes, or some combination thereof. Determining the user's gaze direction may include: determining the convergence point based on the determined orientations of the user's left and right eyes. The convergence point may be the point at which the two fovea axes of the user's binocular eyes intersect (or the closest point between the two axes). The user's gaze direction may be the direction of the line passing through the convergence point and the midpoint between the pupils of the user's binocular eyes.

[0054] In some embodiments, in addition to other eye-tracking cues or biomarkers, the eye-tracking system may also be configured to estimate (measure) the user's blink rate, pupil dilation, fixation stability, and fixation acceleration. In some embodiments, the calculations for determining the estimated values may be performed in real time. In other embodiments, the recorded user fixations may be used to calculate the estimated values. In some embodiments, electrooculogram (EOG) electrodes 110 may be used to calculate estimated values regarding the user's blink rate and other eye-tracking biomarkers, which will be further explained below. In some embodiments, an optical sensor 195 may be used to additionally determine or confirm the estimated values regarding the user's blink rate and other eye-tracking biomarkers, which will be further explained below. In some embodiments, both the EOG electrodes 110 and the optical sensor 195 may be included within the eye-tracking system 145.

[0055] In some embodiments, a zoom module 150 is integrated into the NED 105. The zoom module 150 may be communicatively coupled to the eye-tracking system 145 so that the zoom module 150 can receive eye-tracking information from the eye-tracking system 145. The zoom module 150 may also modify the focus of the image light emitted from the display 125 based on the received eye-tracking information from the eye-tracking system 145. Thus, the zoom module 150 may reduce the vergence-accommodation conflict that may occur when the user's eyes resolve the image light. In various embodiments, the zoom module 150 may engage at least one optical element in the optical assembly 130 (e.g., mechanically or electrically).

[0056] In operation, the zoom module 150 may adjust the position and / or orientation of one or more optical elements in the optical assembly 130 in order to adjust the focus of the image light propagating through the optical assembly 130. In various embodiments, the zoom module 150 may use the eye-tracking information obtained from the eye-tracking system 145 to determine how to adjust one or more optical elements in the optical assembly 130. In some embodiments, the zoom module 150 may perform foveated rendering of the image light based on the eye-tracking information obtained from the eye-tracking system 145 in order to adjust the resolution of the image light emitted by the display 125. In such a case, the zoom module 150 configures the display 125 to display a high pixel density in the foveal region where the user's eyes are fixated and a low pixel density in other regions where the user's eyes are fixated.

[0057] In some embodiments, physiological sensors such as electroencephalogram (EEG) electrodes 190 and / or electrooculogram (EOG) electrodes 110 can be used to determine a user's engagement in different contexts and enhance learning activities. One or more EEG electrodes 190 collect the electric charges generated by the activity of brain cells of the user. One or more EEG electrodes 190 can use the principle of differential amplification by recording the voltage difference between different points, which compares an active detection electrode site with another adjacent or distant reference electrode. The electrical signals collected by the EEG electrodes 190 can be used to generate EEG signal data, which defines a waveform that varies over time and represents the electrical activity occurring within the user's brain. In some embodiments, the EEG electrodes 190 can also be part of a set of electrodes that can be used to generate different types of electrograms of the brain, eyes, heart, etc. (e.g., electroencephalogram (EEG), electrocorticography (ECoG or iEEG), electrooculogram (EOG), electroretinogram (ERG), electrocardiogram (ECG)).

[0058] It should be noted that the same set of electrodes can be used to collect and acquire both EOG signals and EEG signals. The sources of the electrical activities are from different locations (e.g., EOG signals are generated by the electrical activity of the corneal-retinal resting potential, and EEG signals are generated by the electrical activity of the brain). However, EOG signals and EEG signals are typically acquired using the same set of electrodes.

[0059] In some embodiments, the NED 105 includes EEG electrodes 190. In some embodiments, these electrodes are directly disposed on the NED 105 (e.g., on the top of the auricle, the contact point where the nose pad touches the nose, on the temple or the bridge of a pair of smart glasses as shown), and when the user wears the NED, the electrodes contact the user's anatomy. In other embodiments, the electrodes 190 are placed on the user's scalp, face, or ear (e.g., located in the user's ear using an in-ear device such that the electrodes contact the inner surface of the user's ear canal) and are communicatively coupled to the NED 105 via a wired or wireless medium. Figure 2B In some embodiments, the NED 105 includes EEG electrodes 190. In some embodiments, these electrodes are directly disposed on the NED 105 (e.g., on the top of the auricle, the contact point where the nose pad touches the nose, on the temple or the bridge of a pair of smart glasses as shown), and when the user wears the NED, the electrodes contact the user's anatomy. In other embodiments, the electrodes 190 are placed on the user's scalp, face, or ear (e.g., located in the user's ear using an in-ear device such that the electrodes contact the inner surface of the user's ear canal) and are communicatively coupled to the NED 105 via a wired or wireless medium.

[0060] Brain cells communicate via electrical impulses and are constantly active, even during sleep. EEG signals measure brain waves of different frequencies within the brain. Brain waves are oscillating voltages measured in the brain that are only a few millionths of a volt. There are five widely recognized brain waves, and the main frequencies of human EEG waves are gamma (γ), beta (β), alpha, theta, and delta (Δ). Fluctuations in a user's brain waves typically provide information about the user's attention state. In some embodiments, EEG recordings are used to calculate the amplitude of the fluctuations in the user's brain waves, and conclusions about the user's attention state are drawn based on these calculations. In particular, the intensity or dynamic range of the user's alpha, beta, and theta brain waves, as well as the fluctuations in these waves, provide information about the user's attention state and working memory capacity. In some embodiments, the EEG signal is used in combination with the signal from the EOG electrode 110 or other biomarkers from the eye tracking system 145 to filter for the user's attention state.

[0061] EOG is a technique for measuring the corneal-retinal resting potential that exists between the front and back of the human eye. EOG is mainly used to detect eye movements. Using an EOG sensor can help correlate the eye movement artifacts that appear in the EEG signal, thus helping to eliminate such eye movement artifacts. For example, blinking produces certain electrical activities that can distort the EEG signal. The artifacts in the EEG signal generated by blinking may mask the information related to the user's attention state. Therefore, the EOG signal can be used to remove any distortion from the EEG signal. The blink rate determined by the EOG sensor can be used to remove the distortion from the EEG signal and is also one of the multiple biomarkers considered when determining the user's attention state.

[0062] As described above, in some embodiments, the same set of electrodes will be used to collect and acquire both the EOG signal and the EEG signal. It should also be noted that using a sensor based on eye tracking imaging (embedded in the NED system), in combination with the EEG electrodes and the EOG electrodes, trends corresponding to the eye artifacts can be obtained from the EEG data. Eventually, a neural network can be trained to collect the correspondence between the eye tracking sensor and the EOG signal acquired from the electrodes within the glasses.

[0063] In some embodiments, the EOG electrode 110 can be included within the eye tracking system 145. In some embodiments, the EOG readings can be obtained from the same sensor as the EEG electrode 190. In other embodiments, the EOG electrode 110 can be different from the eye tracking system 145 and the EEG electrode 190.

[0064] Similar to EEG electrodes, in some embodiments, the EOG electrodes 110 are disposed directly on the NED 105 (e.g., on the temple or nose bridge of a pair of smart glasses), and when the user wears the NED, the electrodes contact the user's anatomical structure. In other embodiments, the EOG electrodes 110 are placed directly on the user's anatomical structure (e.g., on the user's face) and are communicatively coupled to the NED 105 via a wired or wireless medium. In some embodiments, the EOG electrodes 110 are used to determine the user's blink rate.

[0065] In some embodiments, other optical sensors 195 are used to collect more information for determining the user's attention state. Information from the optical sensors can be used to provide additional information about the user's attention state, particularly in cases where EEG signals are severely distorted, e.g., due to a high blink rate of the user. In some embodiments, the optical sensors 195 can be disposed on the temple or nose bridge of the NED 105, or alternatively can be included in an in-ear device communicatively coupled to the NED 105.

[0066] In some embodiments, the optical sensor 195 can include a functional near-infrared spectroscopy (fNIRS) system, which can include one or more light-emitting diodes (light sources) and one or more sensors (detectors). fNIRS is a non-invasive and non-ionizing optical brain imaging technique that estimates hemodynamic changes in the cerebral cortex by shining light (from, e.g., light-emitting diodes (LEDs), lasers, etc.) onto the user's head and comparing the absorption of light at different wavelengths based on the Beer-Lambert law principle. Different from other tissues in the head, in neural tissue, the hemodynamic changes of hemoglobin oxygenation (HbO) and hemoglobin deoxygenation (HbR) are restricted to be anti-correlated over time. Thus, since the oxygenation level changes as brain regions become more active, it is possible to detect the blood oxygenation changes represented by HbO and / or HbR trajectories in real time to identify and monitor brain activity by using an fNIRS device including a set of fNIRS optodes (e.g., sources, detectors).

[0067] When a user listens to sounds in a crowded environment, background sounds make it difficult for the user to understand what the people around the user are saying. The brain activity identified and monitored in real time by the fNIRS device can be used to estimate what the user is trying to hear and the degree of stress the user is experiencing when trying to hear what the user is focusing on (e.g., estimating the level of difficulty the person is experiencing, estimating cognitive load, estimating listening effort, and estimating the listener's intention, etc.).

[0068] The fNIRS device can be applied to the cortical regions of the brain, where more cortical regions of the brain are involved in active listening compared to when a person hears sounds passively. More specifically, the fNIRS device can be applied to the part of the temporal lobe of the brain called the superior temporal gyrus (STG) and the part of the frontal cortex recruited for active listening. That is, the cognitive load is considered to be related to the amount of blood oxygen change in the STG. One way to obtain such signals is through fNIRS, which detects the degree of blood oxygen change when light passes through the skull. Therefore, the fNIRS device can be applied to measure the activation near the STG, which is related to the vulnerability of the listener to background sounds and may be attributed to cognitive load (e.g., listening effort, listening fatigue), rather than just the percentage of words that the listener can correctly understand (or not understand) when they can clearly hear the words.

[0069] The data obtained from the fNIRS sensor can be used to estimate the user's cognitive load (e.g., listening effort and the listener's intention, etc.). As described above, in some embodiments, the fNIRS optodes can be disposed on the temple or the bridge of the NED 105, or alternatively can be included in an in-ear device communicatively coupled to the NED 105.

[0070] In some embodiments, the optical sensor 195 can include a multi-wavelength photoplethysmography (MW-PPG) sensor. As is well known to those skilled in the art, PPG is a commonly used optical sensing method that collects the light reflected or transmitted through the skin to non-invasively monitor the blood flow pulsation in the subcutaneous blood vessels. Since the blood flow pulsation can reflect the operating conditions of the human circulatory system and respiratory system, the PPG signal can be used as an indicator of the user's attention state, etc.

[0071] For example, the PPG sensor can be used to continuously measure the user's heart rate, respiratory rate, maximum oxygen uptake (VO2), energy consumption, blood oxygen saturation (SpO2), blood pressure, etc. Using the user's heart rate, respiratory rate, and oxygenation level, additional estimates of the user's cognitive load and effort can be calculated. In some embodiments, the optical sensor 195 is directly disposed on the NED 105 (e.g., on the temple of the smart glasses, on the nose pad or the bridge of the smart glasses), or communicatively coupled to the NED 105 using a wired or wireless medium (e.g., an in-ear device is required to place the sensor in the user's ear).

[0072] The I / O interface 175 facilitates the transmission of action requests from a user to the console 170. Additionally, the I / O interface 175 facilitates the transmission of device feedback from the console 170 to the user. An action request is a request to perform a specific action. For example, an action request can be an instruction to start or end the acquisition of image or video data, or an instruction to perform a specific action within an application, such as pausing video playback, increasing or decreasing the volume of audio playback, and initiating and pausing the transcription of audio. In various embodiments, the I / O interface 175 can include one or more input devices. Example input devices include: a keyboard, a mouse, a game controller, a joystick, and / or any other suitable device for receiving action requests and transmitting the action requests to the console 170. In some embodiments, the I / O interface 175 includes an IMU 140 that acquires calibration data indicating an estimated current position of the I / O interface 175 relative to an initial position of the I / O interface 175.

[0073] During operation, the I / O interface 175 receives action requests from the user and sends these action requests to the console 170. The console 170 performs a corresponding action in response to receiving the action request. For example, the console 170 can configure the I / O interface 175 to emit haptic feedback to the user's arm in response to receiving the action request. For example, the console 170 can configure the I / O interface 175 to deliver haptic feedback to the user when the action request is received. Additionally or alternatively, the console 170 can configure the I / O interface 175 to generate haptic feedback when the console 170 performs an action in response to receiving the action request.

[0074] The console 170 provides content for processing to the NED 105 based on information received from one or more of the following: the DCA 155; the eye tracking system 145; one or more other components of the NED 105; and the I / O interface 175. In Figure 1 the illustrated embodiment, the console 170 includes an application repository 160 and an engine 165. In Figure 1 the illustrated embodiment, the engine 165 includes a transcription engine 185. Various embodiments include a transcription engine 185 that is configured to intelligently allocate cognitive load across the visual and auditory domains by determining a personalized and curated transcription of a conversation to be displayed to the user on the display 125. The transcription generated by the transcription engine 185 and provided to the NED 105 includes a plurality of words determined to be relevant, where the relevance of each word is determined based on a combination of multiple factors, which will be discussed further below.

[0075] In some embodiments, related to Figure 1Compared with the modules and / or components described, the console 170 may have more, fewer, or different modules and / or components. Similarly, the functions further described below may be distributed among the components of the console 170 in a manner different from that described in connection with Figure 1 and may be distributed among the components of the console 170 in a manner different from that described in connection with

[0076] The application repository 160 stores one or more applications for execution by the console 170. An application is a set of instructions that, when executed by a processor, perform a set of specific functions, such as generating content for presentation to a user. For example, an application may generate content in response to input received from a user (e.g., movement of the user's head as detected by the NED 105, via the I / O interface 175, etc.). Examples of applications include: game applications, conferencing applications, video playback applications, or other suitable applications.

[0077] In some embodiments, the engine 165 generates a three-dimensional map of the area around the NED 105 (i.e., the "local area") based on information received from the NED 105. In some embodiments, the engine 165 determines depth information for the three-dimensional map of the local area based on depth data received from the NED 105. In various embodiments, the engine 165 uses the depth data received from the NED 105 to update the model of the local area and, in part, generates and / or modifies media content based on the updated model of the local area.

[0078] The engine 165 may also execute applications within the NED system 100 and receive position information, acceleration information, velocity information, predicted future positions, eye gaze information, EEG / EOG information, information from fNIRS or PPG sensors, or some combination thereof, of the NED 105. Based on the information received, the engine 165 determines the various forms of media content to be sent to the NED 105 for presentation to the user. For example, if the information received indicates that the user has looked to the left, the engine 165 generates media content for the NED 105 that reflects the user's movement in a virtual environment or in an environment enhanced with additional content in the local area. Thus, the engine 165 may generate and / or modify media content (e.g., visual content and / or audio content) for presentation to the user. The engine 165 may also send the media content to the NED 105. Additionally, the engine 165 may execute an action within an application executing on the console 170 in response to an action request received from the I / O interface 175. The engine 165 may also provide feedback when executing the action. For example, the engine 165 may configure the NED 105 to generate visual feedback and / or audio feedback, and / or may configure the I / O interface 175 to generate tactile feedback to the user.

[0079] In some embodiments, the engine 165 determines the resolution of media content provided to the NED 105 for presentation to the user on the display 125 based on the received eye-tracking information from the eye-tracking system 145 (e.g., the orientation of the user's eyes). The engine 165 can adjust the resolution of the visual content provided to the NED 105 by configuring the display 125 to perform fixation rendering of visual content based at least in part on the direction of the user's gaze received from the eye-tracking system 145. The engine 165 provides the NED 105 with content that has a high resolution in the foveal region of the user's gaze on the display 125 and a low resolution in other regions, thereby reducing the power consumption of the NED 105. Additionally, using fixation rendering reduces the number of computational cycles used to render visual content without degrading the quality of the user's visual experience. In some embodiments, the engine 165 can also use the eye-tracking information to adjust the focus of the image light emitted from the display 125 in order to reduce vergence-accommodation conflict.

[0080] In some embodiments, the transcription engine 185 determines a personalized and curated transcription of a conversation to be provided to the NED 105 and displayed to the user on the display 125 based on EEG signals from the electrodes 190, blink rate, pupil dilation, fixation stability, and fixation acceleration information obtained from the eye-tracking system 145, and additional attention information obtained from various optical sensors 195. Additionally, the transcription engine can perform actions within an application executing on the console 170 to generate a curated transcription or in response to an action request received from the I / O interface 175.

[0081] Figure 2A is a schematic diagram of the NED 200 according to various embodiments. In various embodiments, the NED 200 presents media to the user. The media can include visual content, auditory content, and tactile content. In some embodiments, the NED 200 provides artificial reality (e.g., virtual reality) content by providing a real-world environment and / or computer-generated content. In some embodiments, the computer-generated content can include visual information, auditory information, and tactile information.

[0082] Those of ordinary skill in the art will understand that the NED 200 may include a see-through NED. A see-through NED keeps the user's view of the real world open and creates a transparent image or a small opaque image that blocks only a small portion of the user's peripheral vision. The see-through category generally includes augmented reality headsets and smart glasses. Augmented reality headsets typically have a field of view of 20 degrees to 60 degrees and overlay information and graphics on the user's view of the real world. Smart glasses typically have a smaller field of view and a display that the user periodically scans rather than continuously views.

[0083] The NED 200 is an embodiment of the NED 105 and includes a front rigid body 205 and a strap 210. The front rigid body 205 includes: electronic display elements of the electronic display 225 ( Figure 2A not shown in the figure), an optical component 130 ( Figure 2A not shown in the figure), an IMU 240, one or more position sensors 135, an eye tracking system 145, and a plurality of locators 120. In Figure 2A the illustrated embodiment, the position sensor 235 is located within the IMU 140, and both the IMU 140 and the position sensor 235 are invisible to the user. EEG electrodes 190 and EOG electrodes ( Figure 2A not shown in the figure) and optical sensors 195 ( Figure 2A not shown in the figure) may be disposed at different positions on the front rigid body 205 or

[0084] A plurality of locators 222 are located at fixed positions on the front rigid body 205 relative to each other and relative to a reference point 215. In Figure 2A the example, the reference point 215 is located at the center of the IMU 140. Each of the plurality of locators 222 emits light that can be detected by the imaging device in the DCA 155. In Figure 2A the example, a plurality of locators 222 or multiple portions of these locators 222 are located on the front side 220A, top side 220B, bottom side 220C, right side 220D, and left side 220E of the front rigid body 205. In some embodiments, the EEG electrodes 190 and EOG electrodes 110 ( Figure 2A not shown in the figure) and optical sensors 195 ( Figure 2A not shown in the figure) may be arranged on the front rigid body 205 or on the front side 220A, top side 220B, bottom side 220C, right side 220D, and left side 220E of the front rigid body 205. In some embodiments, the EEG electrodes 190 and EOG electrodes 110 ( Figure 2A not shown in the figure) and optical sensors 195 ( Figure 2AVarious positions that can be set on the front rigid body 205 (not shown in the figure) or on the front side 220A, top side 220B, bottom side 220C, right side 220D, and left side 220E of the front rigid body 205. Alternatively, in some embodiments, the sensor can be communicatively coupled to the NED 200 using wired or wireless technology.

[0085] The NED 200 includes an eye tracking system 245. As described above, the eye tracking system 245 can include a structured light generator that projects an interference structured light pattern onto the user's eyes and a camera for detecting the illuminated portion of the eyes. The structured light generator and the camera can be located outside the axis of the user's gaze. In various embodiments, the eye tracking system 245 can additionally or alternatively include one or more time-of-flight sensors and / or one or more stereo depth sensors. In Figure 2A one embodiment, the eye tracking system 245 is located below the axis of the user's gaze, but the eye tracking system 245 can alternatively be placed in other positions. Additionally, in some embodiments, there is at least one eye tracking unit for the user's left eye and at least one tracking unit for the user's right eye.

[0086] In various embodiments, the eye tracking system 245 includes one or more cameras located inside the NED 200. When the user is wearing the NED 200, the one or more cameras of the eye tracking system 245 can point inwardly towards one or both of the user's eyes such that the one or more cameras can image one or both of the user's eyes and one or more eye regions while the user is wearing the NED 200. The one or more cameras can be located outside the axis of the user's gaze. In some embodiments, the eye tracking system 245 includes separate cameras for the left and right eyes (e.g., one or more cameras pointing at the user's left eye and one or more cameras separately pointing at the user's right eye).

[0087] Figure 2B is a schematic diagram of the NED 250 according to various embodiments. In various embodiments, the NED 250 presents media to the user. The media can include visual content, auditory content, and tactile content. In some embodiments, the NED 250 provides artificial reality (e.g., augmented reality) content by providing a real-world environment and / or computer-generated content. In some embodiments, the computer-generated content can include visual information, auditory information, and tactile information. The NED 250 is an embodiment of the NED 105. In one embodiment, the NED 250 includes perspective smart glasses.

[0088] The NED 250 includes a frame 252 and a display 254. In various embodiments, the NED 250 may include one or more additional elements. The display 254 may be located at a position on the NED 250 different from the Figure 2B position shown. The display 254 is configured to provide content to a user, the content including audiovisual content. In some embodiments, one or more displays 254 may be located within the frame 252.

[0089] The NED 250 further includes an eye tracking system 245 and one or more corresponding modules 256. These modules 256 may include transmitters (e.g., light transmitters) and / or sensors (e.g., image sensors, cameras). In various embodiments, a plurality of modules 256 are arranged at different positions along the inner surface of the frame 252 such that these modules 256 face the eyes of a user wearing the NED 250. For example, the plurality of modules 256 may include a transmitter that emits a structured light pattern onto the eyes and an image sensor for acquiring an image of the structured light pattern on the eyes. As another example, the plurality of modules 256 may include a plurality of time-of-flight sensors that are configured to direct light onto the eyes and measure the propagation time of the light at each pixel of these sensors. As yet another example, the plurality of modules 256 may include a plurality of stereo depth sensors that are configured to acquire images of the eyes from different advantageous positions. In various embodiments, the plurality of modules 256 further includes an image sensor for acquiring a 2D image of the eyes.

[0090] In some embodiments, EEG and / or EOG electrodes 293 may be provided on the temple 291 of the NED 250 and / or on the nose pad 292 of the NED 250. The EEG and / or EOG electrodes 293 perform functions Figure 1 substantially the same as the EEG electrodes 190 and EOG electrodes 110 shown. In some embodiments, other optical sensors 294 (e.g., fNIRS optodes) may be provided on the temple 291 of the NED 250. The optical sensors 294 perform functions Figure 1 substantially the same as the optical sensors 195 in

[0091] Figure 3 Techniques for determining personalized and curated transcripts of an auditory experience are shown in accordance with various embodiments. Figure 3 The techniques shown in

[0092] may be implemented at least in part using a transcription engine 185. Multiple portions of the techniques shown may also be implemented in combination with an application repository 160 and various components included within the NED 105. Figure 1The described EOG electrode 110, eye tracking system 145, and optical sensor 195 generate multimodal signals that are input to the blink rate module 308. It should be noted that the EOG electrode 110 and the optical sensor 195 may partially or fully overlap with various components of the eye tracking system 145. The blink rate module 308 estimates (measures) the user's blink rate, pupil dilation, fixation stability, fixation acceleration, and other eye tracking biomarkers. A biomarker is a descriptor or measure of a biological system. Biomarkers include obvious measures such as blood pressure, body temperature, or heart rate, and also include less obvious measures such as hair color, brain wave activity, eye fixation, or blink rate.

[0093] The EEG electrode 190 acquires EEG signal data that defines a waveform that varies over time, and this waveform represents the electrical activity occurring within the user's brain. As described above, the intensity of the user's alpha and theta brain waves and the fluctuations of these waves provide information about the user's attention state. However, before using these signals to estimate the user's attention state, the signals from the EEG electrode 190 need to be filtered to generate the filtered EEG signal 310. As previously mentioned, the EOG signals from the EOG electrode 110 can be used to filter the signals from the EEG electrode 190 and remove any distortion.

[0094] In some embodiments, the attention state estimation module 312 uses multimodal inputs (specifically, the filtered EEG signal 310 and the biomarkers from the blink rate module 308 as inputs) to estimate the user's attention state. The user's attention state can be estimated using a calculated quantification or numerical value (based on various biomarkers) that measures the user's ability to focus on spoken words. In some embodiments, the signals from the optical sensor 195 can be directly input into the attention state estimation module 312. As described above, the optical sensor 195 can include an fNIRS system, and the signals from the fNIRS device can be used to estimate the user's cognitive load (e.g., listening effort and the listener's intent, etc.). Therefore, the signals from the optical sensor 195 are input into the attention state estimation module 312, and the attention state estimation module 312 can adjust the signals (e.g., using the EEG signal to filter the fNIRS signal data to separate the neural signals representing brain activity from the noise) to determine the user's cognitive load. The user's cognitive load determined by the attention state estimation module 312 can be one of the multiple evaluation factors for determining the user's attention state. In one embodiment, the attention estimation module 312 assigns a metric or numerical value to the estimated value of the cognitive load experienced by the user.

[0095] In some embodiments, a value between "0" and "1.0" corresponding to the attention state of the user (and / or the cognitive load experienced by the user) is calculated, where "0" indicates that the subject is completely inattentive, and "1.0" indicates that the subject is completely focused. In some embodiments, this value can be assigned to each word in a particular segment of the recorded data. Since the attention state of the user may remain constant over a given period of time, in other embodiments, the calculated value can be assigned to a segment of the recording as a whole. It should be noted that the value is not limited to any specific range of values. Additionally, any other metric can be calculated by the attention state estimation module 312 to track the attention state of the user.

[0096] In some embodiments, the attention state estimation module 312 includes a machine learning model ( Figure 3 not shown), which includes a pre-trained model that is used to quantify the attention state of the user and assign a value within a given range. For example, the machine learning model is trained to assign a value between "0" and "1.0" to the attention state of the subject, where "0" indicates that the subject is completely inattentive, and "1.0" indicates that the subject is completely focused. For example, the machine learning model can include one or more recurrent neural networks (RNNs), one or more convolutional neural networks (CNNs), one or more deep neural networks (DNNs), one or more deep convolutional networks (DCNs), one or more residual neural networks (ResNets), one or more graph neural networks, one or more autoencoders, one or more transformer neural networks, one or more deep stereo geometric networks (DSGNs), one or more stereo R-CNNs, and / or one or more other types of artificial neural networks or components of artificial neural networks. The machine learning model can also or alternatively include a regression model, a support vector machine, a decision tree, a random forest, a gradient boosting tree, a naive Bayes classifier, a Bayesian network, a Hidden Markov model (HMM), a hierarchical model, an ensemble model, a clustering technique, and / or another type of machine learning model that does not utilize components of artificial neural networks.

[0097] In some embodiments, sensor readings (e.g., EOG, EEG, optical sensor readings) from a large number of users in various attention states can be used to train the machine learning model. The model can use machine learning techniques to learn which of the multiple signals from various sensors is most directly related to the attention of the user and needs to be weighted more heavily relative to other signals when determining the attention state. The model can also determine the ideal weights between various sensor inputs for calculating the value of the attention state of the user (e.g., between "0" and "1.0").

[0098] The microphone 176 captures the acoustic experience near the NED 105, particularly any conversations or spoken words that are part of the acoustic experience. The signal of the captured acoustic experience from the microphone 176 is directed to the word intelligibility estimation module 318. In some embodiments, the word intelligibility estimation module 318 includes a classification system that is used to determine the degree of intelligibility of the spoken words perceived by the user of the NED 105. In some embodiments, this determination can be performed based on an audibility estimation method (e.g., the speech intelligibility index). The Speech Intelligibility Index (SII) is a standardized measurement that ranges between "0.0" and "1.0" and is highly correlated with the intelligibility of speech. An SII of "0" means that no speech information is available (audible and / or usable) in the given environment to improve speech understanding. An SII of "1.0" means that all speech information in the given situation is audible and usable to the listener. The word intelligibility estimation module 318 is not limited to using the SII and can use any other technique for calculating the intelligibility of the words spoken in the conversation captured by the microphone 176.

[0099] In some embodiments, any system capable of performing audibility estimation based on the acoustic signal-to-noise ratio (SNR) (e.g., the ratio of the acoustic energy of the spoken word to the energy of the background sound) can be used. The intelligibility of a given spoken word can be determined based on an audibility estimation method that depends on the SNR associated with the word.

[0100] Similar to the attention state estimation module 312, in some embodiments, a numerical value between "0" and "1.0" can be assigned to each word in the record based on the intelligibility of the word. Since the intelligibility of speech may remain constant over a given period of time, in other embodiments, the calculated numerical value can be assigned to a segment of the record as a whole. It should be noted that the numerical value is not limited to any specific value range. In addition, any other metric can also be calculated by the word intelligibility estimation module 318 to track speech intelligibility.

[0101] The signal for the auditory experience collected from the microphone 176 is also directed to the keyword detection module 334. The keyword detection module 334 estimates the importance of a given word in the context of the collected conversation. In other words, the keyword detection module 334 determines whether a given word is prominent enough in the context of the conversation to be considered a "keyword". In some embodiments, a score between 0 and 1 can be attributed to each word, where "0" indicates that the word is irrelevant in the context of the conversation, and "1" indicates that the word is crucial in the context of the conversation. In some embodiments, each word is assigned a metric or numerical value that indicates the importance of the word in the context of the conversation. In some embodiments, the context of the conversation can be determined by analyzing the words close to the word being evaluated. It should be noted that the numerical values are not limited to any specific value range. Additionally, the keyword detection module 334 can also calculate any other metric to determine whether a word is critical in the context of the conversation.

[0102] In some embodiments, the keyword detection module 334 includes a machine learning model ( Figure 3 not shown), which includes a pre-trained model for quantifying the importance of a given word in the context of the collected speech and assigning it a numerical value within a given range (e.g., between "0" and "1.0" as described above). For example, the machine learning model can include one or more recurrent neural networks (RNNs), one or more convolutional neural networks (CNNs), one or more deep neural networks (DNNs), one or more deep convolutional networks (DCNs), one or more residual neural networks (ResNets), one or more graph neural networks, one or more autoencoders, one or more transformer neural networks, one or more deep stereometric networks (DSGNs), one or more stereo R-CNNs, and / or one or more other types of artificial neural networks or components of artificial neural networks. The machine learning model can also or alternatively include regression models, support vector machines, decision trees, random forests, gradient boosting trees, naive Bayes classifiers, Bayesian networks, hidden Markov models (HMMs), hierarchical models, ensemble models, clustering techniques, and / or other types of machine learning models that do not utilize components of artificial neural networks.

[0103] In some embodiments, the keyword detection module 334 includes a machine learning model that uses a tool similar to ChatGPT, where ChatGPT selects keywords from the conversation or evaluates each word individually and assigns a score between "0" and "1.0" to each word. Alternatively, the ChatGPT tool can use any other metric to evaluate the importance of a word in the context.

[0104] The weighting module 332 analyzes inputs (e.g., numerical values or scores for each word or word segment) from the attention estimation module 312, the keyword detection module 334, and the word comprehensibility estimation module 318 to determine relevant keywords to be displayed on the display 125 as part of the personalized and curated transcription. When the attention estimation module 318 makes a judgment based on biomarkers collected from the user's anatomy, the word comprehensibility estimation module and the keyword detection module 334 analyze signals received from outside the user (e.g., acoustic microphone recordings of speech near the user). The relevance of each word is determined based on the analysis and weighting of all three inputs from modules 312, 334, and 318.

[0105] In cases where a word may be acoustically incomprehensible (e.g., as determined by the output of the word comprehensibility estimation module 318) and / or the user exhibits inattentiveness (e.g., based on the output of the attention state estimation module 312) and / or a keyword is spoken (e.g., based on the output of the keyword detection module 334), the weighting module 332 selects the words to be selectively rendered on the display 125 as part of the curated transcription 330. The words selected for display will depend on the weights assigned to each of the modules 312, 318, and 334.

[0106] In some embodiments, weighting includes calculating a metric or numerical value for each word based on digital readings extracted from each of the attention estimation module 312, the keyword detection module 334, and the word comprehensibility estimation module 318. For example, a weighting value can be calculated based on digital readings (e.g., digital readings between "0" and "1.0" calculated for each of modules 312, 334, and 318) and compared to a predetermined threshold. If the weighting value is above the threshold, the word is displayed on the display 125. However, if the weighting value is below the threshold, the word will not be emphasized or omitted from the curated transcription. It should be noted that the numerical values assigned by the weighting module 332 are not limited to any specific value range. Additionally, the weighting module 332 can use any other metric to weight between various inputs.

[0107] The weighting module 332 can use one of several different methods to determine the weights assigned to each of the multiple inputs. In one embodiment, regression analysis or a similar statistical method can be used to determine how to weight each of the multiple inputs in a way that optimizes a particular user's engagement. For example, regression analysis can be used to determine the weighting value to be compared to the predetermined threshold.

[0108] The weighting module 332 can be configured to change the weights according to the user, rather than prescribing a standard weighting for all users. For example, if the biomarker or EEG signal of a particular user indicates that the user is distracted, the weighting module 332 can assign a higher weight to the signals from the attention state estimation module 312, which may have the effect of over-including words (and including words such as: even if these words are not determined to be keywords by the keyword detection module 334, and even if the SNR readings of these words are high enough to make the words understandable) when generating the curated transcription. In other words, the weighting module 332 can be configured to give precedence to the readings from one of the modules 318, 332, and 312 based on the subjective characteristics of the user or various other characteristics of the acoustic environment.

[0109] In some embodiments, the weighting module 332 includes a machine learning model ( Figure 3 not shown), the machine learning model includes a pre-trained model, and the pre-trained model is trained to determine the optimal balance between various signal inputs. For example, the machine learning model can include one or more recurrent neural networks (RNNs), one or more convolutional neural networks (CNNs), one or more deep neural networks (DNNs), one or more deep convolutional networks (DCNs), one or more residual neural networks (ResNets), one or more graph neural networks, one or more autoencoders, one or more transformer neural networks, one or more deep stereometric networks (DSGNs), one or more stereo R-CNNs, and / or one or more other types of artificial neural networks or components of artificial neural networks. The machine learning model can also or alternatively include a regression model, a support vector machine, a decision tree, a random forest, a gradient boosting tree, a naive Bayes classifier, a Bayesian network, a hidden Markov model (HMM), a hierarchical model, an ensemble model, a clustering technique, and / or other types of machine learning models that do not utilize components of artificial neural networks. The weighting module 332 uses the machine learning model to determine a personalized and unique weight for each user.

[0110] In some embodiments, the user can selectively control the weights assigned to each of the attention state estimation module 312, the word intelligibility estimation module 318, and the keyword detection module 334 through the user input 320. The user input 320 to the sensitivity setting module 324 allows the user to override the default parameters used by the weighting module 332. For example, if the user is in a noisy environment, the user can manually increase the weight attributed to the word intelligibility estimation module 318. Or, if the user is suffering from attention deficit, the user can lower the attention threshold of the attention state estimation module 312 so that more words are included in the transcription.

[0111] The sensitivity setting module 324 allows for the control of the weighting module 332. The sensitivity setting module 324 can allow a user to increase the sensitivity of the weighting module 332 to certain inputs over other inputs. For example, if a user is listening to a lecture full of unfamiliar terms in a classroom, the sensitivity of the keyword detection module 334 can be increased so that the weighting module 332 over-includes words determined to be keywords.

[0112] In some embodiments, the sensitivity setting module 324 can be configured to automatically prompt the weighting module 332 to adjust the weights. As Figure 3 shown, the background noise 322 is a separate input to the sensitivity setting module 324. For example, when the intensity of the background noise 322 exceeds a given threshold, the sensitivity setting module 324 can be configured to automatically increase the weight assigned to the word intelligibility estimation module 318. For example, in such a case, the SNR threshold used by the module 318 can be increased so that any word below the given SNR threshold can be designated as part of the curated transcription and displayed on the display 125.

[0113] In some embodiments, the curated transcription 330 includes only those words determined by the weighting module 332 to meet the criteria established by the weighting module 332, while the remaining words not determined to be relevant are omitted. In some embodiments, the curated transcription 330 includes a verbatim transcription, where the keywords are highlighted or bolded in some way to attract the user's attention, while the words not determined to be keywords are not emphasized.

[0114] Figure 4 is a flowchart of method steps for developing a curated transcription of an acoustic experience according to various embodiments. Although the method steps are described with reference to Figures 1 to 3 the system, those skilled in the art will understand that in other embodiments, any system can be configured to implement these method steps in any order.

[0115] As shown, the method 400 begins at step 402, where the attention state estimation module 312 calculates a first metric or numerical value of the spoken words in an auditory experience (e.g., a conversation) based on biometric markers associated with the user, and this first metric or numerical value indicates the user's attention state. As previously described, in some embodiments, EEG / EOG signals and other signals collected from various sensors connected to or disposed on the NED system are used to determine the user's attention state. The biometric markers can include the user's blink rate, pupil dilation, gaze stability, gaze acceleration, the intensity of the user's alpha and theta brain waves, or other biometric markers collected from one or more of the EEG electrodes 190, EOG electrodes 110, eye tracking system 145, and optical sensor 195.

[0116] In step 404, the word intelligibility estimation module 318 calculates a second measure or value corresponding to the intelligibility of the spoken word. In some embodiments, this determination may be performed based on an audibility estimation method (e.g., the Speech Intelligibility Index), but other SNR-based methods may also be used.

[0117] In step 406, the keyword detection module 318 calculates a third measure or value corresponding to the importance of the word in the context of one or more other nearby words in the conversation. This third measure or value can be used to determine whether the word is a keyword in the context of the conversation. In some embodiments, a machine learning model is used to determine whether the word is a keyword and assign a numerical score to the word.

[0118] In step 408, the weighting module 332 weights among the first measure or value, the second measure or value, and the third measure or value to determine a fourth measure or value. As described above, Figure 3 the weighting module 332 of analyzes the inputs from the attention estimation module 312, the keyword detection module 334, and the word intelligibility estimation module 318 to determine relevant keywords to be displayed on the display 125 as part of a personalized and curated transcription. In some embodiments, regression analysis or a similar statistical method may be used to perform the weighting. In some embodiments, a trained machine learning model is used to perform the weighting.

[0119] In step 410, the fourth measure or value is used to determine whether to transcribe the word for display on a display device (e.g., on a NED display) as part of a curated and personalized transcription presented to the user, the curated and personalized transcription taking into account the user's attention state, the intelligibility of the word, and the importance of the word in the context of the auditory experience. For example, the fourth measure or value may be compared to a threshold to determine whether a particular word should be displayed.

[0120] In summary, the transcription engine is configured to intelligently distribute cognitive load between the visual and auditory domains by providing the user with a personalized and curated transcription of the conversation on a display screen (e.g., on the display of the NED). The transcription only displays on the display those keywords that are determined to be relevant to the user, while omitting or de-emphasizing words that are not determined to be relevant to the user. Relevance can be determined based on a combination of one or more of the following factors: a) determination of the user's level of intelligibility of a particular spoken word (e.g., using an audibility estimation method such as the Speech Intelligibility Index); b) estimation of the importance of a word in the context of the conversation (e.g., using a machine learning model); and c) determination of the user's attention state using various biomarkers (e.g., (i) estimation of blink rate, pupil dilation, fixation stability, and fixation acceleration obtained from sensors located on the NED or on the user's face; (ii) EEG signals that estimate the intensity and / or fluctuations of the user's alpha and theta waves, the EEG signals being obtained using sensors placed on the NED or on the user's anatomy). In some embodiments, the user is allowed to control the weights or sensitivities assigned to each of factors (a), (b), and (c).

[0121] At least one advantage of the disclosed technology is that users of a personalized speech-to-text transcription system can rely on the curated transcription to grasp certain spoken words that the user did not perceive auditorily during the auditory experience. The transcription system personalizes and curates the transcription based on cues collected from both the spoken words and the user, which not only increases user engagement but also prevents the user from being overwhelmed by a verbatim transcription of the auditory experience. By relying on cues from the user (e.g., physiological cues that determine the user's attention state), the transcription is also selectively curated for each individual user, which further increases user engagement. Thus, compared to previous methods that resulted in low user engagement by transcribing the entire conversation verbatim, the disclosed technology represents a technological advancement.

[0122] Any combination and all combinations of any claim element stated in any claim and / or any element described in this application fall within the intended scope of this embodiment and protection in any way.

[0123] The foregoing description of the embodiments has been presented for purposes of illustration; the foregoing description is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. It will be appreciated by those skilled in the relevant art that many modifications and variations are possible in light of the above disclosure.

[0124] Some parts of this specification describe embodiments of the present disclosure in terms of algorithms and symbolic representations of operations on information. Those skilled in the art of data processing typically use these algorithmic descriptions and representations to effectively convey the substance of their work to other skilled persons in the art. Although these operations are described functionally, computationally, or logically, these operations are understood to be implemented by a computer program, an equivalent circuit, microcode, or the like. Additionally, it has been found that, without loss of generality, the arrangement of these operations is sometimes conveniently referred to as a module for convenience. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combination thereof.

[0125] Any of the steps, operations, or processes described herein may be performed or implemented singly or in combination with other devices using one or more hardware or software modules. In one embodiment, a software module is implemented using a computer program product that includes a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the steps, operations, or processes described.

[0126] Embodiments of the present disclosure may also relate to an apparatus for performing the operations herein. The apparatus may be specially constructed for the required purposes, and / or the apparatus may include a general-purpose computing device selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory tangible computer-readable storage medium coupled to a computer system bus, or in any type of medium suitable for storing electronic instructions. Additionally, any computing system mentioned in this specification may include a single processor or may be an architecture employing a multi-processor design for increased computing power.

[0127] Embodiments of the present disclosure may also relate to a product generated by the computing processes described herein. Such a product may include information generated from the computing process, where the information is stored on a non-transitory tangible computer-readable storage medium and may include any embodiment of a computer program product or other data combination described herein.

[0128] Finally, the language used in the specification has been principally selected for readability and guidance, and it may not have been selected to delimit or circumscribe the subject matter of the invention. Accordingly, the scope of the present disclosure is intended not to be limited by this specific embodiment, but rather by any claims that may issue from this application based on this disclosure. Thus, the disclosure of the embodiments is intended to illustrate rather than limit the scope of the present disclosure, which is set forth in the appended claims.

[0129] For illustrative purposes, descriptions of various embodiments have been given, but these descriptions are not intended to be exhaustive or to limit the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the embodiments.

[0130] Aspects of the present embodiments may be embodied in a system, a method, or a computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects that are generally referred to herein as a "module" or "system". In addition, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code thereon.

[0131] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above storage media. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0132] As described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine. When executed by the processor of a computer or other programmable data apparatus, these instructions are capable of implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, an application-specific processor, or a field-programmable gate array.

[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architectures, functions, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code that includes one or more executable instructions for implementing one or more specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks may not be performed in the order noted in the accompanying drawings. For example, two blocks shown in succession may, depending on the functions involved, actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order. It will be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a system based on dedicated hardware that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0134] Although the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure can be designed without departing from the basic scope of the present disclosure, and the scope of the present disclosure is determined by the appended claims.

Claims

1. A computer-implemented method, comprising: Calculating a first metric for a word, the first metric indicating an attentional state of a user perceiving the word during an auditory experience, wherein the metric is calculated based on a biomarker associated with the user; Calculating a second metric corresponding to the intelligibility of the word during the auditory experience; Calculating a third metric corresponding to the importance of the word in the context of one or more words spoken in proximity to the word during the auditory experience; and Calculating a fourth metric based on weights assigned to each of the first metric, the second metric, and the third metric; and Determining whether to transcribe the word for display based on the fourth metric.

2. The computer-implemented method according to claim 1, wherein, The fourth metric is used to determine whether to transcribe the word for display by comparing the fourth metric with a predetermined threshold, wherein if the fourth metric is higher than the predetermined threshold, the word is displayed on a display.

3. The computer-implemented method according to claim 1 or 2, wherein, The display is included within a near-eye display (NED) system operating in an artificial reality environment.

4. The computer-implemented method according to any one of the preceding claims, wherein, The biomarker includes one or more of the following: blink rate; pupil dilation; fixation stability; or fixation acceleration.

5. The computer-implemented method according to any one of the preceding claims, wherein, The biomarker is calculated using signals from one or more sensors, the one or more sensors including electroencephalogram (EEG) electrodes, electrooculogram (EOG) electrodes, functional near-infrared spectroscopy (fNIRS) optodes, or multi-wavelength photoplethysmography (MW-PPG) sensors.

6. The computer-implemented method according to any one of the preceding claims, wherein, Each of the first metric, the second metric, and the third metric includes a numerical value ranging between 0 and 1.

7. The computer-implemented method according to any one of the preceding claims, wherein, Calculating the second metric corresponding to the intelligibility of the word includes: estimating the audibility of the word using a signal-to-noise ratio (SNR) associated with the word.

8. The computer-implemented method according to any one of the preceding claims, wherein, Calculating the third metric corresponding to the importance of the word includes: determining the importance of the word using a machine learning model, wherein the machine learning model includes one of the following: recurrent neural network (RNN), convolutional neural network (CNN), deep neural network (DNN), deep convolutional network (DCN), residual neural network (ResNet), graph neural network, autoencoder, transformer neural network, deep stereometric network (DSGN), or region-based convolutional neural network for stereoscopic images R-CNN.

9. The computer-implemented method according to any one of the preceding claims, wherein, The assignment of the weights includes using a machine learning model to determine optimal weights, wherein the machine learning model includes one of the following: recurrent neural network (RNN), convolutional neural network (CNN), deep neural network (DNN), deep convolutional network (DCN), residual neural network (ResNet), graph neural network, autoencoder, transformer neural network, deep stereometric network (DSGN), or region-based convolutional neural network for stereoscopic images R-CNN.

10. The computer-implemented method according to any one of the preceding claims, wherein, The assignment of the weights includes performing a regression analysis.

11. The computer-implemented method according to any one of the preceding claims, wherein, The weights are selectively assigned by the user.

12. The computer-implemented method according to any one of the preceding claims, wherein, The assignment of the weights is based on a combination of: automatic selection; and user-based selection.

13. One or more computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of the preceding claims.

14. A wearable device comprising: one or more memories storing instructions, and one or more processors coupled to the one or more memories and configured to perform the method according to any one of claims 1 to 12 when executing the instructions.

15. A computer program product comprising instructions that, when the program is executed by one or more processors of a computer, cause the one or more processors to perform the method according to any one of claims 1 to 12.