Voice text information display method, head-mounted display device and readable medium

By using a head-mounted display device to determine the direction of the sound source and display the voice and text in three-dimensional space, the problem of users having difficulty distinguishing between voice and text when communicating with multiple people is solved, thus improving the user experience and the convenience of interaction.

CN121807190APending Publication Date: 2026-04-07HANGZHOU LINGBAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

When hearing aids help users with communication difficulties communicate, they only provide translation and display functions. This requires users to distinguish the corresponding communication object based on semantics, resulting in a poor user experience.

Method used

The system uses a head-mounted display device to determine the direction of the voice source, performs text conversion processing, and displays the voice and text information in a three-dimensional display space with preset degrees of freedom based on the direction of the voice source. This allows users to distinguish the communication partner by the text display location.

Benefits of technology

It improves the user experience, especially in situations involving multiple communication partners. Users can more easily identify the corresponding communication partner for the voice message, and the text display is anchored according to the user's head posture, enhancing the convenience of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121807190A_ABST
    Figure CN121807190A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice text information display method, head-mounted display equipment and a readable medium. A specific embodiment of the method comprises the steps of determining sound source direction information corresponding to current voice information in response to determining that the head-mounted display device receives the current voice information; performing text conversion processing on the current voice information to obtain voice text information; determining text display position information according to the sound source direction information; the voice text information is displayed at the position, corresponding to the text display position information, in the three-dimensional display space of the head-mounted display device, and the voice text information displayed in the three-dimensional display space is displayed with the preset freedom degree. According to the embodiment, the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to a method for displaying voice and text information, a head-mounted display device, and a readable medium. Background Technology

[0002] Hearing assistive devices can help users with communication difficulties understand what others are saying in real time. Currently, when using hearing assistive devices to help users with communication difficulties communicate, the common method is that the hearing assistive device uses speech recognition technology to recognize spoken text, and then displays the recognized spoken text on the screen of the hearing assistive device for the user to view.

[0003] However, the inventors discovered that when using the above methods to help users with communication difficulties communicate, the following technical problems often arise: when hearing aids help users with communication difficulties communicate, they only realize the functions of translation and display. When a user communicates with multiple people, the user needs to distinguish the corresponding person for the voice text based on the semantics of the communication, resulting in a poor user experience.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure provide methods for displaying voice and text information, head-mounted display devices, and computer-readable media to address the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a method for displaying voice-to-text information. The method includes: in response to determining that the head-mounted display device has received current voice information, determining sound source direction information corresponding to the current voice information; performing text conversion processing on the current voice information to obtain voice-to-text information; determining text display position information based on the sound source direction information; and displaying the voice-to-text information at a position corresponding to the text display position information in the three-dimensional display space of the head-mounted display device, wherein the voice-to-text information displayed in the three-dimensional display space is displayed with preset degrees of freedom.

[0008] Optionally, determining the sound source direction information corresponding to the current speech information includes: performing voiceprint recognition processing on the current speech information to obtain voiceprint category information; and determining the sound source direction information based on the voiceprint category information in response to determining that the voiceprint category information meets preset category conditions.

[0009] Optionally, the head-mounted display device includes a camera; and determining the sound source direction information corresponding to the current voice information includes: determining the orientation of the camera as the sound source direction information of the current voice information.

[0010] Optionally, the above-mentioned text conversion processing of the current speech information to obtain speech-text information includes: acquiring a sequence of facial images of the communication subject captured by the camera corresponding to the current time period, wherein the current time period is the time period corresponding to the current speech information; performing speech detection processing on the facial image sequence of the communication subject to obtain detection result information; and, in response to determining that the detection result information meets preset result conditions, performing text conversion processing on the current speech information to obtain speech-text information.

[0011] Optionally, the above method further includes: in response to determining that the above detection result information does not meet the above preset result conditions, displaying preset sound prompt information in the three-dimensional display space of the above head-mounted display device; performing sound source localization processing on the above current voice information to obtain updated sound source direction information; and in response to determining that the above camera orientation corresponds to the sound source direction corresponding to the above updated sound source direction information, performing text conversion processing on the above current voice information to obtain voice text information.

[0012] Optionally, determining the text display position information based on the sound source direction information includes: determining the communication object information based on the detection result information; obtaining the virtual mouth position information corresponding to the communication object information, wherein the virtual mouth position information represents the position of the mouth of the communication object corresponding to the communication object information in the three-dimensional display space; and determining the text display position information based on the virtual mouth position information.

[0013] Optionally, displaying the voice text information at the position corresponding to the text display position information in the three-dimensional display space of the head-mounted display device includes: determining the text display style based on the voiceprint category information; and displaying the voice text information at the position corresponding to the text display position information in the three-dimensional display space of the head-mounted display device according to the text display style.

[0014] Optionally, the above method further includes: adjusting the position of the voice text information in response to detecting a position adjustment operation corresponding to the voice text information.

[0015] In a second aspect, some embodiments of this disclosure provide a head-mounted display device, including: one or more processors; a storage device storing one or more programs thereon; a display screen for displaying voice and text information; and when one or more programs are executed by one or more processors, causing the one or more processors to implement the method described in any implementation of the first aspect above.

[0016] Thirdly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in any implementation of the first aspect.

[0017] The above-described embodiments of this disclosure have the following beneficial effects: the voice-to-text information display method of some embodiments of this disclosure can improve the user experience. Specifically, the reason for the poor user experience is that when hearing aids help users with communication difficulties communicate, they only realize the functions of translation and display. When a user communicates with multiple communication partners, the user needs to distinguish the corresponding communication partner based on the semantics of the communication, resulting in a poor user experience. Based on this, the voice-to-text information display method of some embodiments of this disclosure firstly determines the sound source direction information corresponding to the current voice information in response to determining that the head-mounted display device has received current voice information. Thus, the direction of the sound source corresponding to the voice can be obtained, which can be used to determine the display position of the text in the three-dimensional display space. Secondly, the current voice information is processed by text conversion to obtain voice-to-text information. Thus, the text corresponding to the received voice can be obtained, which can be used to help the user understand the voice content of the communication partner. Then, based on the sound source direction information, the text display position information is determined. Thus, the display position of the text in the three-dimensional display space can be obtained, which can be used for the user to view the text. Finally, the aforementioned voice-text information is displayed at the position corresponding to the text display location information in the three-dimensional display space of the aforementioned head-mounted display device. The voice-text information displayed in the three-dimensional display space is displayed with preset degrees of freedom. Therefore, the text corresponding to the voice can be displayed with preset degrees of freedom in the three-dimensional display space of the head-mounted display device, allowing the user to view the text from any angle. Because the head-mounted display device first determines the location of the sound source when receiving external voice information, and determines different text display positions for different sound sources, when there are multiple communication partners, the user can distinguish the communication partners through the text display position, thereby improving the user experience. Furthermore, because the text can be displayed using the preset degrees of freedom of the head-mounted display device, the text can be anchored in the three-dimensional display space when the user's head posture changes, further enhancing the user experience. Attached Figure Description

[0018] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0019] Figure 1 This is a schematic diagram of an application scenario of a voice text information display method according to some embodiments of the present disclosure;

[0020] Figure 2 This is a flowchart of some embodiments of the voice and text information display method according to the present disclosure;

[0021] Figure 3 This is a schematic diagram of an application scenario of a voice text information display method according to some other embodiments of the present disclosure;

[0022] Figure 4 This is a schematic diagram illustrating an application scenario of the voice text information display method according to other embodiments of this disclosure;

[0023] Figure 5 These are flowcharts of other embodiments of the voice and text information display method according to this disclosure;

[0024] Figure 6 This is a schematic diagram of a head-mounted display device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0028] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0030] The collection, storage, and use of user personal information (such as facial images of communication partners) involved in this disclosure shall be carried out in accordance with relevant laws and regulations, provided that the relevant organizations or individuals have fulfilled their obligations, including conducting personal information security impact assessments, informing personal information subjects, obtaining prior authorization and consent from personal information subjects, and other obligations before performing such operations.

[0031] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0032] Figure 1 This is a schematic diagram illustrating an application scenario of the voice text information display method according to some embodiments of this disclosure.

[0033] exist Figure 1 In the application scenario, firstly, the head-mounted display device 101 can determine the sound source direction information 103 corresponding to the received current voice information 102 in response to the determination that the current voice information 102 has been received. Secondly, the head-mounted display device 101 can perform text conversion processing on the current voice information 102 to obtain voice-text information 104. Then, the head-mounted display device 101 can determine the text display position information 105 based on the sound source direction information 103. Finally, the head-mounted display device 101 can display the voice-text information 104 at the position corresponding to the text display position information 105 in the three-dimensional display space 106. The voice-text information displayed in the three-dimensional display space 106 is displayed with preset degrees of freedom.

[0034] It should be understood that Figure 1 The number of head-mounted display devices shown is merely illustrative. Any number of head-mounted display devices can be used depending on the implementation requirements.

[0035] Continue to refer to Figure 2 The diagram illustrates a flow 200 of some embodiments of a voice-to-text information display method according to the present disclosure. This voice-to-text information display method, applied to a head-mounted display device, includes the following steps:

[0036] Step 201: In response to determining that the head-mounted display device has received the current voice information, determine the direction information of the sound source corresponding to the current voice information.

[0037] In some embodiments, the execution subject of the voice text information display method (e.g. Figure 1 The head-mounted display device (as shown) can determine the sound source direction information corresponding to the current voice information in response to determining that the head-mounted display device has received current voice information. The head-mounted display device may be equipped with a microphone array. The head-mounted display device can be used to display mixed reality scenes or virtual scenes. For example, the head-mounted display device can be MR glasses, VR glasses, or AR glasses. The microphone array can be used to collect sound. The current voice information can be the currently received voice signal. The sound source direction information can represent the direction of the position of the sound source emitting the current voice information relative to the position of the head-mounted display device. In practice, the execution entity can, in response to determining that the head-mounted display device has received current voice information, perform sound source localization processing on the current voice information using a preset sound source localization algorithm to obtain the azimuth and elevation angles corresponding to the current voice information as sound source direction information. The preset sound source localization algorithm can be a pre-set algorithm for sound source localization. For example, the aforementioned preset sound source localization algorithm can be an MVDR (Minimum Variance Distortionless Response) beamforming algorithm, or an SRP-PHAT (Steered Response Power-Phase Transform) sound source localization algorithm.

[0038] Optionally, the aforementioned execution entity may also, in response to determining that the head-mounted display device has received current voice information, perform sound source localization processing on the current voice information using a preset sound source localization algorithm to obtain the azimuth, elevation, and distance corresponding to the current voice information as sound source direction information. Thus, the depth of the text in the three-dimensional display space can be determined using the distance included in the sound source direction information.

[0039] In some optional implementations of certain embodiments, the execution entity may determine the sound source direction information corresponding to the current speech information through the following steps:

[0040] The first step is to perform voiceprint recognition processing on the current voice information to obtain voiceprint category information. This voiceprint category information indicates whether the voiceprint corresponding to the current voice information belongs to the communication target and the type of the communication target. The communication target can be the person with whom the user is currently communicating via voice with the user wearing the head-mounted display device. For example, the voiceprint category information can be either a match with the communication target or a non-match. It should be noted that when the voiceprint category information indicates that the voiceprint corresponding to the current voice information belongs to the communication target, the voiceprint category information may also include the communication target's sound source direction information and the communication target's identifier. The communication target's sound source direction information indicates the direction of the communication target's sound relative to the head-mounted display device. This information may include, but is not limited to, the communication target's azimuth angle and elevation angle. The communication target's azimuth angle can be the azimuth angle of the communication target's throat relative to the head-mounted display device. The communication target's elevation angle can be the elevation angle of the communication target's throat relative to the head-mounted display device. The communication target's identifier can be the communication target's identifier. In practice, the aforementioned executing entity can use a preset voiceprint recognition algorithm to process the current speech information and obtain voiceprint category information. The preset voiceprint recognition algorithm can be a pre-defined algorithm for voiceprint recognition. For example, it could be a voiceprint recognition algorithm based on a Gaussian mixture model and Mel-frequency cepstral coefficients, or it could be an ECAPA-TDNN (Emphasized Channel Attention, Propagation and Aggregation in timedelay neural network Based Speaker Verification) voiceprint recognition model.

[0041] The second step involves determining the sound source direction information based on the voiceprint category information, in response to the determination that the voiceprint category information meets the preset category conditions. The preset category conditions can represent that the voiceprint corresponding to the current speech information belongs to the communication target. In practice, the executing entity can determine the communication target sound source direction information included in the voiceprint category information as the sound source direction information in response to the determination that the voiceprint category information meets the preset category conditions.

[0042] Therefore, by identifying whether the voice belongs to the person the user is communicating with, and then determining the direction of the sound source, it is easier for the user to identify the type of person to whom the voice belongs.

[0043] Step 202: Perform text conversion processing on the current speech information to obtain speech-text information.

[0044] In some embodiments, the executing entity can perform text conversion processing on the current speech information to obtain speech-text information. The speech-text information can be text representing speech. In practice, the executing entity can use a preset speech recognition algorithm to perform text conversion processing on the current speech information to obtain speech-text information. The preset speech recognition algorithm can be a pre-defined algorithm for converting speech into text. For example, the preset speech recognition algorithm can be a speech recognition algorithm based on a deep fully convolutional neural network, or it can be a speech recognition algorithm based on a streaming multi-level truncated attention model.

[0045] Step 203: Determine the text display position information based on the sound source direction information.

[0046] In some embodiments, the execution entity can determine the text display position information based on the sound source direction information. The text display position information can represent the position of the text box displaying the text in the three-dimensional display space. In practice, the execution entity can determine the text display position information as the preset text display position information corresponding to the sound source direction information from a preset text display position information set. The preset text display position information in the preset text display position information set can be a pre-set position representing the display of text in the three-dimensional display space. The preset text display position information can include, but is not limited to, the coordinates of the text display box position. The text display box position coordinates can be the three-dimensional coordinates of the text box used to display the corresponding text in the three-dimensional display space. Each preset text display position information can correspond to a preset sound source direction range information. The preset sound source direction range information can represent a pre-set sound source direction range. The preset sound source direction range information can include, but is not limited to, a preset azimuth angle range and a preset elevation angle range. The preset azimuth angle range can be a pre-set range of the azimuth angle of the sound source relative to the head-mounted display device. The preset elevation angle range can be a pre-set range of the elevation angle of the sound source relative to the head-mounted display device. The preset text display position information corresponding to the above sound source direction information can be: the azimuth angle included in the above sound source direction information is within the preset azimuth angle range included in the preset sound source direction range information corresponding to the preset text display position information, and the elevation angle included in the above sound source direction information is within the preset elevation angle range included in the preset sound source direction range information corresponding to the preset text display position information.

[0047] Optionally, the aforementioned preset sound source direction range information may further include a preset distance range. The preset distance range can be a pre-defined range of distances between the sound source and the head-mounted display device. Correspondingly, the preset text display position information can be: the azimuth angle included in the sound source direction information is within the preset azimuth angle range included in the preset sound source direction range information corresponding to the preset text display position information; the elevation angle included in the sound source direction information is within the preset elevation angle range included in the preset sound source direction range information corresponding to the preset text display position information; and the distance included in the sound source direction information is within the preset distance range included in the preset sound source direction range information corresponding to the preset text display position information. This facilitates user differentiation between voice and text messages from different communicators within the same direction.

[0048] Step 204: Display the voice text information at the location corresponding to the text display position information in the three-dimensional display space of the head-mounted display device.

[0049] In some embodiments, the execution entity may display the speech-text information at a position corresponding to the text display position information in the three-dimensional display space of the head-mounted display device. The speech-text information displayed in the three-dimensional display space is displayed with preset degrees of freedom. These preset degrees of freedom can be pre-defined degrees of freedom of the head-mounted display device. These preset degrees of freedom can be 3DOF (Degrees of Freedom) or 6DOF.

[0050] In some optional implementations of certain embodiments, the aforementioned execution entity may display the aforementioned voice-text information at the location corresponding to the aforementioned text display position information in the three-dimensional display space of the aforementioned head-mounted display device through the following steps:

[0051] The first step is to determine the text display style based on the aforementioned voiceprint category information. This text display style can be the style of the text box displaying the text. In practice, firstly, the executing entity can, in response to determining that the aforementioned voiceprint category information meets preset category conditions, determine the text display style of the communication object corresponding to the aforementioned communication object identifier as the text display style. This communication object text display style can be the text display style of the communication object itself. The communication object text display style can be pre-set or determined when the corresponding communication object first communicates with the user wearing the aforementioned head-mounted display device. Secondly, in response to determining that the aforementioned voiceprint category information does not meet preset category conditions, a preset set of text display styles to be applied is retrieved from the database via a wired or wireless connection. The preset text display styles to be applied in the preset set of text display styles to be applied can be pre-set text display styles waiting to be applied. It should be noted that the aforementioned wireless connection methods can include, but are not limited to, 3G / 4G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future wireless connection methods. Then, any preset text display style from the set of preset text display styles to be applied is selected as the text display style.

[0052] The second step is to display the voice text information in the three-dimensional display space of the head-mounted display device at the position corresponding to the text display position information, according to the text display style.

[0053] As an example, such as Figure 3 As shown, the text display style of the communication object 301 is elliptical, and the text display style of the communication object 302 is rectangular. When the head-mounted display device determines that the above voiceprint category information meets the preset category conditions and the matched communication object is the communication object 301, the voice text information is displayed in the three-dimensional display space 303 of the head-mounted display device at the position corresponding to the above text display position information, according to the above text display style 304.

[0054] Therefore, different communication partners can correspond to different text display styles, making it easier for users to distinguish communication partners.

[0055] Optionally, the aforementioned executing entity may also adjust the position of the aforementioned voice-text information in response to detecting a position adjustment operation corresponding to the voice-text information. The position adjustment operation may be an operation to adjust the display position of the voice-text information. The position adjustment operation may be, but is not limited to, at least one of the following: clicking, dragging, hovering, voice control, gesture control, and head posture control. Therefore, users can adjust the display position of the voice-text themselves, thereby improving the user experience.

[0056] The above-described embodiments of this disclosure have the following beneficial effects: the voice-to-text information display method of some embodiments of this disclosure can improve the user experience. Specifically, the reason for the poor user experience is that when hearing aids help users with communication difficulties communicate, they only realize the functions of translation and display. When a user communicates with multiple communication partners, the user needs to distinguish the corresponding communication partner based on the semantics of the communication, resulting in a poor user experience. Based on this, the voice-to-text information display method of some embodiments of this disclosure firstly determines the sound source direction information corresponding to the current voice information in response to determining that the head-mounted display device has received current voice information. Thus, the direction of the sound source corresponding to the voice can be obtained, which can be used to determine the display position of the text in the three-dimensional display space. Secondly, the current voice information is processed by text conversion to obtain voice-to-text information. Thus, the text corresponding to the received voice can be obtained, which can be used to help the user understand the voice content of the communication partner. Then, based on the sound source direction information, the text display position information is determined. Thus, the display position of the text in the three-dimensional display space can be obtained, which can be used for the user to view the text. Finally, the aforementioned voice-text information is displayed at the position corresponding to the text display location information in the three-dimensional display space of the aforementioned head-mounted display device. The voice-text information displayed in the three-dimensional display space is displayed with preset degrees of freedom. Therefore, the text corresponding to the voice can be displayed with preset degrees of freedom in the three-dimensional display space of the head-mounted display device, allowing the user to view the text from any angle. Because the head-mounted display device first determines the location of the sound source when receiving external voice information, and determines different text display positions for different sound sources, when there are multiple communication partners, the user can distinguish the communication partners through the text display position, thereby improving the user experience. Furthermore, because the text can be displayed using the preset degrees of freedom of the head-mounted display device, the text can be anchored in the three-dimensional display space when the user's head posture changes, further enhancing the user experience.

[0057] Figure 4This is a schematic diagram illustrating an application scenario of the voice text information display method according to other embodiments of this disclosure.

[0058] exist Figure 4 In the application scenario, firstly, the head-mounted display device 401, in response to determining that it has received current voice information 402, determines the orientation of the camera as the sound source direction information 403 of the current voice information 402. Secondly, the head-mounted display device 401 acquires a sequence of facial images of the communication subject corresponding to the current time period captured by the camera. Then, the head-mounted display device 401 performs speech detection processing on the facial image sequence 404 to obtain detection result information 405. Afterwards, in response to determining that the detection result information 405 meets preset result conditions, the current voice information 402 is processed into text to obtain voice-text information 406. Then, based on the sound source direction information 403, the text display position information 407 is determined. Finally, the head-mounted display device 401 can display the voice-text information 406 at the position corresponding to the text display position information 407 in the three-dimensional display space 408.

[0059] It should be understood that Figure 4 The number of head-mounted display devices shown is merely illustrative. Any number of head-mounted display devices can be used depending on the implementation requirements.

[0060] Further reference Figure 5 The document illustrates a flow 500 of another embodiment of the voice-to-text information display method. Flow 500 of this voice-to-text information display method includes the following steps:

[0061] Step 501: In response to determining that the head-mounted display device has received the current voice information, the orientation of the camera is determined as the sound source direction information of the current voice information.

[0062] In some embodiments, the head-mounted display device described above may include a camera. The camera may be used to capture an image of the face of the person being communicated with. This facial image may be a picture of the face of the person being communicated with. The entity executing the voice-text information display method (e.g., Figure 4 The head-mounted display device 401 shown can, in response to determining that the head-mounted display device has received current voice information, determine the orientation of the camera as the sound source direction information of the current voice information. The orientation of the camera can be the direction in which the camera is shooting relative to the head-mounted display device.

[0063] Step 502: Obtain the sequence of facial images of the people being interacted with, captured by the camera, corresponding to the current time period.

[0064] In some embodiments, the executing entity can acquire a sequence of facial images of the interacting parties corresponding to the current time period captured by the camera. The current time period is the time period corresponding to the current voice information. The time period corresponding to the current voice information can be the period from the start time of the current voice information to the end time of the current voice information. The sequence of facial images of the interacting parties can be a sequence of facial images of the interacting parties arranged in ascending chronological order. In practice, the executing entity can acquire the sequence of facial images of the interacting parties corresponding to the current time period captured by the camera from a database via a wired or wireless connection.

[0065] Step 503: Perform speech detection processing on the facial image sequence of the communication subject to obtain detection result information.

[0066] In some embodiments, the execution entity can perform speech detection processing on the facial image sequence of the communication subject to obtain detection result information. In practice, the execution entity can input the facial image sequence of the communication subject into a pre-trained detection result information generation model to obtain detection result information. The detection result information generation model can be a classification model that takes the facial image sequence of the communication subject as input and the detection result information as output. The detection result information can characterize whether the communication subject is speaking in the current time period. For example, the detection result information can be "the communication subject is speaking" or "the communication subject is not speaking." It should be noted that when the detection result information characterizes that the communication subject is speaking in the current time period, the detection result information may also include a communication subject identifier. The classification model can be a classification model based on a convolutional neural network or a support vector machine.

[0067] Step 504: In response to determining that the detection result information meets the preset result conditions, the current speech information is processed into text to obtain speech text information.

[0068] In some embodiments, in response to determining that the detection result information meets preset result conditions, the execution entity performs text conversion processing on the current speech information to obtain speech-text information. The preset result conditions may indicate that the detection result information represents that the communication subject is speaking during the current time period. In practice, in response to determining that the detection result information meets preset result conditions, the execution entity may perform text conversion processing on the current speech information using a preset speech recognition algorithm to obtain speech-text information.

[0069] Optionally, the aforementioned implementing entity may also perform the following steps:

[0070] The first step involves, in response to the determination that the aforementioned detection result information does not meet the aforementioned preset result conditions, displaying preset sound prompt information in the three-dimensional display space of the aforementioned head-mounted display device. This preset sound prompt information can be a pre-set message used to alert the user that other sounds are being heard. These other sounds can be voices emitted by non-interaction parties. For example, the preset sound prompt information could be "A new sound has entered."

[0071] The second step is to perform sound source localization processing on the current speech information to obtain updated sound source direction information. In practice, the executing entity can use the preset sound source localization algorithm to perform sound source localization processing on the current speech information to obtain the azimuth and elevation angles corresponding to the current speech information as updated sound source direction information.

[0072] The third step involves, in response to determining that the camera orientation corresponds to the sound source direction in the updated sound source direction information, performing text conversion processing on the current speech information to obtain speech-text information. Specifically, the correspondence between the camera orientation and the sound source direction in the updated sound source direction information can be defined as follows: the difference between the camera azimuth angle of the camera orientation and the azimuth angle included in the updated sound source direction information is less than or equal to a preset azimuth angle difference value, and the difference between the camera elevation angle of the camera orientation and the elevation angle included in the updated sound source direction information is less than or equal to a preset elevation angle difference value. The preset azimuth angle difference value can be a pre-set difference in azimuth angles. The preset elevation angle difference value can be a pre-set difference in elevation angles. In practice, the executing entity can, in response to determining that the camera orientation corresponds to the sound source direction in the updated sound source direction information, perform text conversion processing on the current speech information using the preset speech recognition algorithm to obtain speech-text information.

[0073] Therefore, when the voice of the currently received non-face-to-face communication partner is detected, the user can be notified that there is a new voice source. The voice information will only be converted into text when the user is face-to-face with the corresponding communication partner, thus facilitating one-on-one communication.

[0074] Step 505: Determine the text display position information based on the sound source direction information.

[0075] In some embodiments, the specific implementation of step 505 and its resulting technical effects can be found in [reference needed]. Figure 2 Step 203 in the corresponding embodiments will not be repeated here.

[0076] Optionally, the aforementioned execution entity can determine the text display location information based on the aforementioned sound source direction information through the following steps:

[0077] The first step is to identify any face image of the target communication object from the above sequence of face images of the communication object.

[0078] The second step involves performing distance measurement processing on the facial image of the target communication subject to obtain distance information. This distance information can be the distance from the communication subject to the camera. In practice, the executing entity can use a preset image ranging algorithm to perform distance measurement processing on the facial image of the target communication subject to obtain the distance information. This preset image ranging algorithm can be a pre-defined algorithm used to estimate the distance from the face of the communication subject in the image to the camera. For example, the preset image ranging algorithm can be a monocular ranging algorithm.

[0079] The third step is to determine the target sound source direction information based on the aforementioned sound source direction information and the distance information of the communication object. In practice, firstly, the executing entity can determine the distance information of the communication object as the sound source distance. Then, the aforementioned sound source direction information and the aforementioned sound source distance are determined as the target sound source direction information.

[0080] The fourth step is to determine the preset text display position information that corresponds to the target sound source direction information in the preset text display position information set as the text display position information.

[0081] In some optional implementations of certain embodiments, the execution entity may determine the text display position information based on the sound source direction information through the following steps:

[0082] The first step is to determine the communication partner information based on the aforementioned test results. In practice, the implementing entity can identify the communication partner identifier included in the test results as the communication partner information.

[0083] The second step is to obtain the virtual mouth position information corresponding to the aforementioned communication object information. This virtual mouth position information represents the position of the mouth of the communication object in the aforementioned three-dimensional display space. This virtual mouth position information may include, but is not limited to, the coordinates of the mouth's center position. Specifically, the mouth's center position coordinates can be the coordinates of the center point of the mouth of the communication object in the aforementioned three-dimensional display space. In practice, the executing entity can obtain the virtual mouth position information corresponding to the aforementioned communication object information using SLAM (simultaneous localization and mapping) technology.

[0084] The third step is to determine the text display position information based on the aforementioned virtual mouth position information. In practice, the executing entity can use a preset position offset algorithm to offset the mouth center position coordinates included in the virtual mouth position information, obtaining the offset position coordinates as the text display position information. The preset position offset algorithm can be any algorithm used for offsetting the position. For example, the preset position offset algorithm can be the OFFSET function. Therefore, the text can be displayed next to the mouth of the person being communicated with, thus displaying the text without obscuring the image of the mouth of the person being communicated with.

[0085] Therefore, the text can be displayed near the mouth of the person communicating with the voice message, which can improve the realism of the text when users view the voice message and thus enhance the user experience.

[0086] Step 506: Display the voice text information at the location corresponding to the text display position information in the three-dimensional display space of the head-mounted display device.

[0087] In some embodiments, the specific implementation of step 506 and its resulting technical effects can be found in [reference needed]. Figure 2 Step 204 in the corresponding embodiments will not be repeated here.

[0088] from Figure 5 It can be seen from this that, with Figure 2 Compared to the description of some corresponding embodiments, Figure 5 The flow 500 of the voice-to-text information display method in some corresponding embodiments embodies the step of extending the current voice information by performing text conversion processing. Therefore, the solutions described in these embodiments can achieve voice-to-text conversion only when the user wearing the head-mounted display device speaks face-to-face with the communication partner, thereby facilitating face-to-face communication and improving the user experience.

[0089] The following is for reference. Figure 6 It illustrates a head-mounted display device 600 suitable for implementing some embodiments of the present disclosure (e.g., Figure 1 A schematic diagram of the hardware structure of the head-mounted display device 101 in the image. Figure 6 The head-mounted display device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0090] like Figure 6As shown, the head-mounted display device 600 may include a processing unit 601 (e.g., a central processing unit, a graphics processing unit, etc.), a memory 602, an input unit 603, and an output unit 604. The processing unit 601, memory 602, input unit 603, and output unit 604 are interconnected via a bus 605. Here, the method according to embodiments of this disclosure can be implemented as a computer program and stored in the memory 602. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the method shown in the flowchart. The processing unit 601 in the head-mounted display device implements the voice-text information display method of this disclosure by calling the aforementioned computer program stored in the memory 602. In some implementations, the output unit 604 may include a display screen for displaying voice-text information.

[0091] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0092] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0093] The aforementioned computer-readable medium may be included in the aforementioned head-mounted display device; or it may exist independently and not assembled into the head-mounted display device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the head-mounted display device, cause the head-mounted display device to: determine the direction information of the sound source corresponding to the current speech information in response to determining that the head-mounted display device has received current speech information; perform text conversion processing on the current speech information to obtain speech-text information; determine the text display position information based on the sound source direction information; and display the speech-text information at the position corresponding to the text display position information in the three-dimensional display space of the head-mounted display device, wherein the speech-text information displayed in the three-dimensional display space is displayed with preset degrees of freedom.

[0094] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0096] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0097] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for displaying voice and text information, applied to a head-mounted display device, comprising: In response to determining that the head-mounted display device has received current voice information, the direction information of the sound source corresponding to the current voice information is determined; The current speech information is converted into text to obtain speech-text information; Based on the sound source direction information, determine the text display position information; The voice text information is displayed at the position corresponding to the text display position information in the three-dimensional display space of the head-mounted display device, wherein the voice text information displayed in the three-dimensional display space is displayed with a preset degree of freedom.

2. The method according to claim 1, wherein, Determining the sound source direction information corresponding to the current speech information includes: The current voice information is processed by voiceprint recognition to obtain voiceprint category information; In response to determining that the voiceprint category information meets preset category conditions, the sound source direction information is determined based on the voiceprint category information.

3. The method according to claim 1, wherein, The head-mounted display device includes a camera; and Determining the sound source direction information corresponding to the current speech information includes: The orientation of the camera is determined as the sound source direction information of the current voice information.

4. The method according to claim 3, wherein, The step of performing text conversion processing on the current speech information to obtain speech-text information includes: Obtain a sequence of facial images of the person being interacted with, captured by the camera, corresponding to the current time period, wherein the current time period is the time period corresponding to the current voice information; Speech detection processing is performed on the facial image sequence of the person being communicated with to obtain detection result information; In response to determining that the detection result information meets the preset result conditions, the current voice information is processed into text to obtain voice text information.

5. The method according to claim 4, wherein, The method further includes: In response to determining that the detection result information does not meet the preset result condition, a preset sound prompt is displayed in the three-dimensional display space of the head-mounted display device; The current speech information is processed to locate the sound source, and the updated sound source direction information is obtained. In response to determining that the camera orientation corresponds to the sound source direction of the updated sound source direction information, the current speech information is processed into text to obtain speech-text information.

6. The method according to claim 4, wherein, The step of determining the text display position information based on the sound source direction information includes: Based on the detection results, the information of the communication partner is determined; Obtain virtual mouth position information corresponding to the communication object information, wherein the virtual mouth position information represents the position of the mouth of the communication object corresponding to the communication object information in the three-dimensional display space; Based on the virtual mouth position information, the text display position information is determined.

7. The method according to claim 2, wherein, Displaying the voice text information at the position corresponding to the text display position information in the three-dimensional display space of the head-mounted display device includes: The text display style is determined based on the voiceprint category information; The voice text information is displayed in accordance with the text display style at the position corresponding to the text display location information in the three-dimensional display space of the head-mounted display device.

8. The method according to claim 1, wherein, The method further includes: In response to detecting a position adjustment operation corresponding to the voice text information, the position of the voice text information is adjusted.

9. A head-mounted display device, comprising: One or more processors; A storage device on which one or more programs are stored; A display screen for showing voice and text information; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-8.

10. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.