Translation system, smart glasses, computer readable medium, and computer program product

Through the collaborative work of smart glasses and mobile terminals, directional audio acquisition and low power transmission are realized, solving the problems of poor user experience, low speech recognition accuracy and long network delay time in the prior art, and improving the user experience and translation quality of cross-language communication.

CN120452418APending Publication Date: 2025-08-08HANGZHOU LINGBAN TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510527602.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In cross-language communication, the reliance on mobile phone translation software leads to poor user experience, low speech recognition accuracy, smart glasses are limited by local computing power and energy consumption, and insufficient real-time translation quality and network delay time.

Method used

Audio data is collected in a directional manner through the built-in audio acquisition unit of the smart glasses, and sent it to the mobile terminal for translation using a low-power transmission channel. The translation results are displayed on the smart glasses. Combined with microphone array and signal-to-noise ratio processing to improve the audio data quality, the mobile terminal translates and returns to the smart glasses to display.

Benefits of technology

It improves the user experience of cross-language communication, enhances the accuracy of speech recognition, shortens network latency, and improves the quality of real-time translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452418A_ABST
    Figure CN120452418A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a translation system, smart glasses, a computer readable medium and a computer program product. According to one specific embodiment, the translation system comprises an intelligent glasses terminal which is configured to directionally collect audio data of a first sound source through a built-in audio collection unit, send the audio data of the first sound source to a mobile terminal in communication connection through a low-power-consumption transmission channel, respond to received translation information sent by the mobile terminal and send the translation information to the intelligent glasses terminal; the translation information is displayed; and the mobile terminal is configured to translate the received audio data of the first sound source to obtain translation information corresponding to the first sound source, and send the translation information corresponding to the first sound source to the intelligent glasses terminal. According to the embodiment, the user experience during cross-language communication is improved, the accuracy of voice recognition is improved, the real-time translation quality is improved, and the network delay time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a translation system, smart glasses, a computer-readable medium, and a computer program product. Background Art

[0002] Cross-language communication requires language translation. Currently, this mainly relies on mobile phone translation software, where users point their phone's microphone at the speaker to pick up the audio, and the translated result is displayed on the phone screen. Alternatively, they can use smart glasses with voice interaction capabilities.

[0003] However, these approaches often present the following technical challenges: Relying on mobile phone translation software requires both parties to frequently look down at their screens, disrupting eye contact and body language. Handheld devices can cause stiff movements, disrupting natural conversation and resulting in a poor user experience. Mobile phone microphones are non-directional, so ambient noise can interfere with recognition accuracy, leading to poor recognition accuracy. While smart glasses offer voice interaction capabilities, they are limited by local computing power and energy consumption, resulting in poor real-time translation quality. Cloud-based processing also presents significant network latency.

[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention

[0005] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] Some embodiments of the present disclosure propose a translation system, smart glasses, a computer-readable medium, and a computer program product to solve one or more of the technical problems mentioned in the above background technology section.

[0007] In a first aspect, some embodiments of the present disclosure provide a translation system, including: a smart glasses end, configured to directionally collect audio data of a first sound source through a built-in audio collection unit, send the audio data of the first sound source to a communication-connected mobile terminal through a low-power transmission channel, and display the translation information in response to receiving the translation information sent by the mobile terminal; the mobile terminal, configured to translate the received audio data of the first sound source, obtain translation information corresponding to the first sound source, and send the translation information corresponding to the first sound source to the smart glasses end.

[0008] Optionally, the smart glasses are further configured to, in response to detecting a selection operation corresponding to the first translation mode, determine a direction range of a first preset angle in front of the smart glasses as the source direction of the first sound source.

[0009] Optionally, the smart glasses are further configured to send the audio data of the first sound source to the communication-connected mobile terminal through the low-power transmission channel through the following steps: compressing the audio data of the first sound source to obtain compressed audio data; and sending the compressed audio data to the mobile terminal through the low-power transmission channel.

[0010] Optionally, the audio acquisition unit includes a microphone ring array consisting of at least two microphones, and the smart glasses end is further configured to directionally collect audio data of the first sound source through the following steps: synchronously collecting the sound signals received by each microphone in the microphone ring array to obtain each sound signal; for each of the above sound signals, generating sound source direction information of the above sound signal; generating a directional enhanced sound signal for each sound signal that meets the preset direction conditions corresponding to the above first sound source according to the corresponding sound source direction information; and generating audio data of the first sound source based on the above directional enhanced sound signal.

[0011] Optionally, the smart glasses include an ambient light sensor, and the smart glasses are further configured to display the translation information through the following steps: collecting ambient brightness information through the ambient light sensor included in the smart glasses; generating subtitle text configuration information corresponding to the translation information based on the ambient brightness information; and displaying the translation information based on the subtitle text configuration information.

[0012] Optionally, the smart glasses are further configured to: in response to detecting a selection operation corresponding to the second translation mode, determine a direction range of a second preset angle below the smart glasses as the source direction of a second sound source; directionally collect audio data of the second sound source through the audio collection unit; and send the audio data of the second sound source to the mobile terminal through a low-power transmission channel.

[0013] Optionally, the mobile terminal is further configured to: translate the received audio data of the second sound source to obtain translation information corresponding to the second sound source; display the translation information corresponding to the second sound source; collect audio data as mobile-end audio data through the microphone included in the mobile terminal; translate the mobile-end audio data to obtain translation information corresponding to the mobile-end audio data; and send the translation information corresponding to the mobile-end audio data to the smart glasses.

[0014] Optionally, the smart glasses are further configured to generate a directional enhanced sound signal according to the sound signal corresponding to the first sound source that satisfies the preset direction condition corresponding to the corresponding sound source direction information through the following steps: determining the microphones corresponding to the sound signals that meet the preset direction condition as a microphone set; screening the target microphone from the microphone set; for each microphone in the microphone set except the target microphone, performing the following steps: determining the distance between the microphone and the target microphone; generating a time difference between the microphone and the target microphone according to the preset focal length direction and the distance; and delaying the sound signal corresponding to the microphone according to the time difference. compensation processing to obtain an updated sound signal; the sound signal of the above-mentioned target microphone and the updated sound signals are determined as a sound signal set; for each sound signal in the above-mentioned sound signal set, the following steps are performed: generating a signal-to-noise ratio corresponding to the above-mentioned sound signal; generating a first weight corresponding to the above-mentioned sound signal based on the above-mentioned signal-to-noise ratio; generating a second weight corresponding to the above-mentioned sound signal based on the sound source direction information corresponding to the above-mentioned sound signal; generating an enhanced weight corresponding to the above-mentioned sound signal based on the above-mentioned first weight and the above-mentioned second weight; weighting each sound signal in the above-mentioned sound signal set according to the enhanced weight corresponding to each sound signal in the above-mentioned sound signal set to obtain a directionally enhanced sound signal.

[0015] In a second aspect, some embodiments of the present disclosure provide a pair of smart glasses, comprising: at least one display unit for forming an image in front of the user's eyes; one or more processors; and a storage device storing one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement the method executed by the smart glasses described in any implementation of the first aspect above.

[0016] In a third aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method executed by any end described in any implementation manner of the first aspect above is implemented.

[0017] In a fourth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which, when executed by a processor, implements the method performed by any end described in any implementation manner of the first aspect above.

[0018] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the translation system of some embodiments of the present disclosure, the user experience in cross-language communication is improved, the accuracy of voice recognition is improved, the quality of real-time translation is improved, and the network delay time is shortened. Specifically, the reasons for the poor user experience, recognition accuracy, real-time translation quality, and long network delay time are: relying on mobile phone translation software, the communicating parties need to frequently look down at the screen, interrupting eye contact and body language expression, handheld devices are prone to cause stiff movements, destroying natural conversation scenes, resulting in poor user experience, mobile phone microphones are non-directionally designed, and environmental noise is easy to interfere with recognition accuracy, resulting in poor recognition accuracy. Although smart glasses products have voice interaction functions, they are limited by local computing power and energy consumption, and the quality of real-time translation is poor. The cloud-based processing solution has the problem of long network delay time. Based on this, some embodiments of the present disclosure include a translation system comprising: a smart glasses device configured to use a built-in audio acquisition unit to directionally acquire audio data from a first sound source, transmit the audio data from the first sound source to a mobile terminal connected thereto via a low-power transmission channel, and display the translation information in response to receiving translation information from the mobile terminal; and a mobile terminal configured to translate the received audio data from the first sound source to obtain translation information corresponding to the first sound source, and transmit the translation information corresponding to the first sound source to the smart glasses device. Because the smart glasses device can directionally acquire audio data from a specific sound source, the impact of ambient noise can be reduced, the quality of the acquired audio data can be improved, and thus the accuracy of speech recognition can be improved. Furthermore, because the translation processing is performed by the mobile terminal, the smart glasses device is not limited by computing power and energy consumption, thereby improving translation quality. Furthermore, the smart glasses device and the mobile terminal are connected via a low-power transmission channel, which can reduce network latency during data transmission. Furthermore, the translation information obtained by the mobile terminal is directly displayed on the smart glasses device, and the smart glasses device directly acquires audio, eliminating the need for the user to frequently look down at the screen or focus on the speaker, thereby improving the user experience during cross-language communication. This improves the user experience during cross-language communication, improves the accuracy of speech recognition, improves the quality of real-time translation, and shortens network latency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0020] Figure 1 is a schematic diagram of an application scenario of a translation system according to some embodiments of the present disclosure;

[0021] Figure 2 is an interaction flow chart of a translation system according to some embodiments of the present disclosure;

[0022] Figure 3 is a schematic diagram of the structure of some embodiments of the translation system according to the present disclosure;

[0023] Figure 4 Schematic diagram of the structure of smart glasses suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0024] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0025] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0026] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0027] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0028] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0029] Before performing any operations involving the collection, storage, and use of user personal information (such as audio data) in this disclosure, relevant organizations or individuals must fulfill their obligations, including conducting personal information security impact assessments, fulfilling their obligation to inform the personal information subject, and obtaining the prior authorization and consent of the personal information subject.

[0030] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0031] Figure 1 is a schematic diagram of an application scenario of a translation system according to some embodiments of the present disclosure.

[0032] like Figure 1 As shown, user 101 can wear smart glasses 102 and communicate with user 103 across languages. At the same time, smart glasses 102 and mobile terminal 104 are connected via a low-power transmission channel. User 101 and user 103 communicate face to face, and smart glasses 102 can directionally collect the sound in front of them. When user 103 speaks, smart glasses 102 can directionally collect the audio data of user 103 speaking. Smart glasses 102 can then transmit audio data 105 to mobile terminal 104 via the low-power transmission channel. After mobile terminal 104 translates and obtains translation information 105, it transmits translation information 105 to smart glasses 102 via the low-power transmission channel. Finally, smart glasses 102 can display translation information 105 for user 101 to view directly.

[0033] It should be understood that Figure 1 The number of smart glasses and mobile terminals in the embodiment is merely illustrative. Any number of smart glasses and mobile terminals may be provided as required.

[0034] Continue to refer Figure 2 , shows an interaction process 200 according to some embodiments of the translation system of the present disclosure. The interaction process 200 includes the following steps:

[0035] In step 201, the smart glasses directionally collect audio data of a first sound source through a built-in audio collection unit.

[0036] In some embodiments, the smart glasses can use a built-in audio acquisition unit to directionally acquire audio data from a first sound source. The smart glasses may include smart glasses. Smart glasses can be used to create images in front of the user's eyes. Smart glasses may include, but are not limited to, AR glasses, MR glasses, and VR glasses. For example, the smart glasses may be binocular diffraction waveguide AR glasses. The smart glasses may be provided with an audio acquisition unit. The audio acquisition unit may include at least one microphone. The at least one microphone may be used to directionally acquire audio data. The first sound source may represent a spatial range at a preset angle in front of the user when wearing the smart glasses. For example, the preset angle may be 90°. The first sound source may correspond to the speaker on the other side during cross-language communication. In practice, the smart glasses may use a beamforming algorithm to directionally acquire audio data from the first sound source using a built-in audio acquisition unit.

[0037] Optionally, the smart glasses may further, in response to detecting a selection operation corresponding to the first translation mode, determine a directional range within a first preset angle in front of the smart glasses as the source direction of the first sound source. The first translation mode may represent a one-way translation mode for the wearer of the smart glasses. For example, the wearer may select a control representing the first translation mode in the display interface of the smart glasses. For example, the first preset angle may be 90°. The directional range within the first preset angle in front of the smart glasses may be a directional range within the first preset angle centered directly in front of the smart glasses. Thus, in the first translation mode, the source direction of the first sound source may represent the range of the sound source of the opposite speaker.

[0038] Optionally, the audio collection unit includes a microphone ring array consisting of at least two microphones. The microphone ring array can be set on the glasses frame of the smart glasses.

[0039] In some optional implementations of some embodiments, the smart glasses may be further configured to directionally collect audio data of the first sound source through the following steps:

[0040] The first step is to synchronously collect the sound signals received by each microphone in the microphone ring array to obtain individual sound signals. In practice, a multi-channel audio interface can be used to synchronously collect the sound signals received by each microphone. This ensures that all microphone signals are time-aligned.

[0041] The second step is to generate the direction of sound source information for each of the above-mentioned sound signals. In practice, a DOA (Direction of Arrival) algorithm, such as MUSIC, ESPRIT, or an energy-based method, can be used to generate the direction of sound source information for the above-mentioned sound signals.

[0042] In the third step, a directional enhanced sound signal is generated based on each sound signal whose corresponding sound source direction information satisfies the preset direction condition corresponding to the first sound source. The preset direction condition may be that the sound source direction information is within the direction range corresponding to the first sound source.

[0043] Step 4: Generate audio data for the first sound source based on the directionally enhanced sound signal. In practice, the directionally enhanced sound signal can be determined as the audio data for the first sound source. This allows for directionally enhancing the sound signal corresponding to the first sound source while automatically discarding sound sources in other ranges.

[0044] In some optional implementations of some embodiments, the smart glasses may be further configured to generate audio data of the first sound source based on the directionally enhanced sound signal by performing the following steps:

[0045] The first step is to filter the directional enhanced sound signal through a pre-set bandpass filter to obtain a first sound signal. The bandpass filter can be used to retain signals within a predetermined range. The predetermined range can be the distribution range of human voice frequencies. For example, the preset range can be between 300Hz and 3400Hz. In practice, the bandpass filter can be used to retain signals within a predetermined range in the directional enhanced sound signal to obtain the first sound signal. Thus, low-frequency noise and high-frequency noise can be filtered out by the bandpass filter.

[0046] The second step is to perform periodic noise identification on the first sound signal to obtain periodic noise frequency information. In practice, a pre-modeled adaptive filter can be used to perform periodic noise identification on the first sound signal to obtain periodic noise frequency information. The periodic noise frequency information may include the frequency of the periodic noise.

[0047] The third step is to filter the first sound signal based on the periodic noise frequency information to obtain a second sound signal. For example, if there is a fan in the room emitting a steady 500Hz noise, the adaptive filter will gradually learn this frequency and filter it.

[0048] The fourth step is to generate noise spectrum information based on the pre-collected noise signal. The noise signal is a sound signal collected by the microphone ring array when the first sound source is inactive. The inactive state of the first sound source can indicate that the user corresponding to the first sound source is not making any sound. In practice, the noise spectrum information can be determined as the power spectral density of the noise signal.

[0049] The fifth step is to transform the second sound signal to obtain sound signal spectrum information. In practice, the second sound signal can be transformed by performing a short-time Fourier transform to obtain the spectrum of the second sound signal as the sound signal spectrum information.

[0050] In step 6, the noise spectrum information is removed from the sound signal spectrum information to obtain enhanced sound spectrum information. In practice, the noise spectrum information can be subtracted from the sound signal spectrum information to obtain enhanced sound spectrum information.

[0051] Step 7: Inverse transform the enhanced sound spectrum information to obtain the third sound signal. In practice, the enhanced sound spectrum information can be inversely transformed to obtain the third sound signal.

[0052] In step 8, echo cancellation is performed on the third sound signal to obtain a fourth sound signal. In practice, an adaptive filter method can be used to perform echo cancellation on the third sound signal to obtain the fourth sound signal. The adaptive filter method can predict and cancel echoes by modeling the transfer function between the speaker signal and the microphone signal.

[0053] In step nine, the fourth sound signal is subjected to speech enhancement processing to obtain audio data of the first sound source. In practice, the amplitude of the fourth sound signal can be nonlinearly compressed or limited to reduce the noise component in the weak signal. For example, a threshold can be set, and signals below the threshold are directly reset to zero. In this way, the quality of the audio data of the first sound source can be significantly improved through filtering, noise reduction, echo cancellation, and speech enhancement.

[0054] In some optional implementations of some embodiments, the smart glasses may be further configured to generate a directional enhanced sound signal based on the sound signals whose corresponding sound source direction information satisfies the preset direction condition corresponding to the first sound source through the following steps:

[0055] In the first step, each microphone corresponding to each sound signal that meets the above-mentioned preset direction condition is determined as a microphone set.

[0056] The second step is to select a target microphone from the microphone set. For example, the microphone in the microphone set that is closest to the first sound source can be determined as the target microphone.

[0057] In the third step, for each microphone in the microphone set except the target microphone, perform the following steps:

[0058] The first sub-step is to determine the distance between the microphone and the target microphone. In practice, the distance between the microphones can be pre-stored.

[0059] The second sub-step is to generate the time difference between the microphone and the target microphone based on the preset focal length direction and the above-mentioned spacing. For example, in the first translation mode, the preset focal length direction can be the direction corresponding to the front, that is, the center direction of the direction range corresponding to the first sound source. In the second translation mode, the preset focal length direction can be the direction corresponding to the bottom, that is, the center direction of the direction range corresponding to the second sound source. In practice, the sine value of the preset focal length direction can be determined first. Then, the product of the above-mentioned sine value and the above-mentioned spacing can be determined. Finally, the ratio of the above-mentioned product to the speed of sound can be determined as the time difference between the microphone and the target microphone. For example, the speed of sound can be 340m / s.

[0060] The third sub-step is to perform delay compensation processing on the sound signal corresponding to the microphone according to the time difference to obtain an updated sound signal. In practice, the sound signal corresponding to the microphone can be delayed by the time difference to obtain an updated sound signal.

[0061] In the fourth step, the sound signal of the target microphone and the updated sound signals are determined as a sound signal set.

[0062] Step 5: For each sound signal in the sound signal set, perform the following steps:

[0063] The first sub-step is to generate a signal-to-noise ratio corresponding to the sound signal. The signal-to-noise ratio can represent the contribution of the microphone corresponding to the sound signal.

[0064] The second sub-step is to generate a first weight corresponding to the sound signal based on the signal-to-noise ratio. In practice, the first weight corresponding to the sound signal can be determined as the ratio of the signal-to-noise ratio to the sum of the signal-to-noise ratios of the sound signals in the sound signal set.

[0065] The third sub-step involves generating a second weight corresponding to the sound signal based on the sound source direction information corresponding to the sound signal. In practice, the difference between the sound source direction information and the preset focal length direction can be first determined. The second weight can then be determined as the ratio of the difference to a preset angle corresponding to the directional range of the first sound source.

[0066] The fourth sub-step is to generate an enhancement weight corresponding to the sound signal based on the first weight and the second weight. In practice, the first weight and the second weight can be weighted and summed to obtain the enhancement weight of the sound signal. Here, the weighting coefficient can be pre-set. For example, the weighting coefficient corresponding to the first weight and the second weight can both be 0.5.

[0067] In the sixth step, the sound signals in the sound signal set are weighted according to the enhancement weights corresponding to the sound signals in the sound signal set to obtain a directionally enhanced sound signal. In practice, the sound signals in the sound signal set can be weighted and summed according to the enhancement weights corresponding to the sound signals in the sound signal set to obtain a directionally enhanced sound signal. Thus, the enhancement weight of each sound signal can be dynamically determined from the two dimensions of signal-to-noise ratio and sound source direction, and the enhancement weight can be applied to signal superposition, so that the high-contribution microphone can obtain a larger weight, thereby improving the directional enhancement effect of the audio data.

[0068] In step 202, the smart glasses transmit the audio data of the first sound source to the mobile terminal connected thereto via a low-power transmission channel.

[0069] In some embodiments, the smart glasses can send the audio data of the first sound source to a mobile terminal connected to the communication channel through a low-power transmission channel. The low-power transmission channel can be a technology and protocol designed to minimize energy consumption while ensuring data transmission. The low-power transmission channel can include but is not limited to: Bluetooth low energy, zigbee, Z-Wave, Wi-Fi HaLow. Preferably, the low-power transmission channel can be Bluetooth low energy. For example, the low-power transmission channel can be a two-way communication established through the BLE 5.2 protocol. Mobile terminals can include but are not limited to: mobile phones and tablets.

[0070] In some optional implementations of some embodiments, the smart glasses may be further configured to send the audio data of the first sound source to the mobile terminal in communication connection via a low-power transmission channel through the following steps:

[0071] The first step is to compress the audio data of the first sound source to obtain compressed audio data. In practice, the smart glasses can encode and compress the audio data of the first sound source to obtain compressed audio data. For example, the encoding and compression process can be AAC-LC encoding and compression.

[0072] The second step is to send the compressed audio data to the mobile terminal through the low-power transmission channel.

[0073] Step 203: The mobile terminal translates the received audio data of the first sound source to obtain translation information corresponding to the first sound source.

[0074] In some embodiments, the mobile terminal can translate the audio data received from the above-mentioned first sound source to obtain translation information corresponding to the above-mentioned first sound source. The mobile terminal can be pre-configured with a translation engine. In practice, the mobile terminal can translate the audio data through the translation engine to obtain translation information corresponding to the above-mentioned first sound source. The translation information can be a translation result from the source language to the target language. The source language can be the language type of the audio data. The source language can be automatically identified. The target language can be the language type that needs to be converted. The target language can be set by itself. For example, the source language can be English. The target language can be Chinese. The translation result can be the text obtained by translation.

[0075] In step 204 , the mobile terminal sends the translation information corresponding to the first sound source to the smart glasses.

[0076] In some embodiments, the mobile terminal may send translation information corresponding to the first sound source to the smart glasses.

[0077] In step 205 , the smart glasses display the translated information in response to receiving the translated information sent by the mobile terminal.

[0078] In some embodiments, the smart glasses can display the translation information in response to receiving the translation information sent by the mobile terminal. In practice, the smart glasses can display the translation information at a preset depth of the display space. For example, the preset depth can be 6 meters. The display space can be a three-dimensional space used by the smart glasses to display content. Optionally, in response to detecting a font size adjustment operation for the corresponding translation information, the font size of the target translation information is adjusted. The font size adjustment operation can be an operation for adjusting the font size of the translation information. For example, the font size adjustment operation can include but is not limited to at least one of the following: touch control, gesture control, voice control, and head control. The target translation information can include the translation information to be displayed. The target translation information can also include the translation information that has been displayed, without specific limitation. In this way, the user can adjust the font size of the translation information displayed on the smart glasses by himself.

[0079] Optionally, the smart glasses may include an ambient light sensor. The ambient light sensor may be provided on the frames of the smart glasses. The ambient light sensor may be used to detect the current ambient light intensity, and the output value may reflect the ambient brightness level.

[0080] In some optional implementations of some embodiments, the smart glasses may be further configured to display the translation information through the following steps:

[0081] The first step is to collect ambient brightness information through the ambient light sensor included in the above-mentioned smart glasses.

[0082] The second step is to generate subtitle text configuration information corresponding to the above translation information based on the above environmental brightness information.

[0083] The third step is to display the translation information according to the subtitle text configuration information. In practice, the subtitle text configuration information can be updated regularly at least according to the ambient brightness information, and the display effect of the translation information can also be updated accordingly.

[0084] Optionally, the smart glasses may further include a camera, which may be disposed in front of the smart glasses frame.

[0085] In some optional implementations of some embodiments, the smart glasses are further configured to generate subtitle text configuration information corresponding to the translation information based on the ambient brightness information through the following steps:

[0086] The first step is to perform a text size mapping on the ambient brightness information to obtain a text size type. In practice, a nonlinear text size mapping can be performed on the ambient brightness information to obtain a text size type. For example, the mapping function corresponding to the nonlinear text size mapping can be: ambient brightness information <50 lux → text size type = large; ambient brightness information >1000 lux → text size type = small; else → text size type = medium.

[0087] In the second step, the text size corresponding to the above text size type is determined as the subtitle text size information. The text size corresponding to each text size type can be pre-set. For example, the text size corresponding to "Large" can be 24 pt. The text size corresponding to "Medium" can be 18 pt. The text size corresponding to "Small" can be 12 pt.

[0088] The third step is to perform a light / dark type mapping on the ambient brightness information to obtain a light / dark type. In practice, a nonlinear light / dark type mapping can be performed on the ambient brightness information to obtain a text size type. For example, the mapping function corresponding to the nonlinear light / dark type mapping can be: ambient brightness information < 50 lux → light / dark type = dark environment; ambient brightness information ≥ 50 lux → light / dark type = bright environment.

[0089] Step 4: Determine subtitle text color information based on the aforementioned light / dark type. For example, in response to determining that the aforementioned light / dark type represents a dark environment, preset light color information may be determined as the subtitle text color information. In response to determining that the aforementioned light / dark type represents a bright environment, preset dark color information may be determined as the subtitle text color information. The preset light color information and the preset dark color information may both include color values. For example, the preset light color information may include a color value of white. The preset dark color information may include a color value of black.

[0090] Step 5: Determine the subtitle text transparency information based on the light and dark type. For example, in response to determining that the light and dark type represents a dark environment, a first preset value may be added to the preset transparency to obtain an updated transparency as the subtitle text transparency information. In response to determining that the light and dark type represents a bright environment, a second preset value may be subtracted from the preset transparency to obtain an updated transparency as the subtitle text transparency information. The first preset value and the second preset value may be pre-set.

[0091] The sixth step is to collect an environmental image through the above-mentioned camera. The environmental image can represent the image of the real scene in front of the wearer's field of vision.

[0092] Step 7: Identify the direction of the strong light from the above environment image. In practice, the main source direction of the light in the environment image can be identified as the strong light direction through image processing algorithms.

[0093] In step 8, first subtitle text position information is generated based on the aforementioned strong light direction. In practice, the image region in which the aforementioned strong light direction is located in the aforementioned environmental image can be determined. The determined image region can then be used as the first subtitle text position information. The aforementioned environmental image can be pre-divided into image regions. For example, the environmental image can be evenly divided into four image regions.

[0094] In the ninth step, semantic segmentation is performed on the environmental image to obtain a set of segmented regions. In practice, a pre-trained semantic segmentation model can be used to perform semantic segmentation on the environmental image to obtain a set of segmented regions. For example, the semantic segmentation model can be U-Net or Mask R-CNN. Through semantic segmentation, different objects (corresponding to different segmented regions) in the environmental image can be identified, such as faces, screens, whiteboards, blank walls, etc.

[0095] In step 10, each segmented region that meets a preset placement condition is selected from the aforementioned segmented region set as a target segmented region set. The preset placement condition may be that the object corresponding to the segmented region exists in a set of placeable objects. The placeable object set may be a set of pre-defined objects that can be used to place subtitle text. For example, the placeable object set may include, but is not limited to, a blank wall, a whiteboard, and an open space.

[0096] In step 11, each spatial position corresponding to the target segmented region set in the display space of the smart glasses is determined as the position information of each second subtitle text. The spatial position may be a three-dimensional position in the display space of the smart glasses. The spatial position corresponding to the target segmented region may be a spatial range or a spatial position point representing a center position.

[0097] Step 12: Select from the above-mentioned second subtitle text position information the second subtitle text position information that matches the above-mentioned first subtitle text position information as the third subtitle text position information. In practice, the second subtitle text position information whose corresponding target segmented area is outside the range of the first subtitle text position information can be determined as the third subtitle text position information.

[0098] Step 13: Select the third subtitle text position information that meets the preset area condition from the above third subtitle text position information as the subtitle text position information. In practice, the third subtitle text position information with the largest area of the corresponding segmented area can be determined as the subtitle text position information.

[0099] In the fourteenth step, the subtitle text size information, the subtitle text color information, the subtitle text transparency information and the subtitle text position information are determined as target subtitle text information.

[0100] In step 15, based on the initial subtitle text information and the target subtitle text information, a smooth subtitle text information sequence is generated as the subtitle text configuration information corresponding to the translation information. The last smooth subtitle text information in the smooth subtitle text information sequence is the target subtitle text information. When the program starts, the initial subtitle text information may be the default subtitle text information. Thereafter, the initial subtitle text information may be the current subtitle text information. The target subtitle text information may be the subtitle text information to be converted. The initial subtitle text information may include initial subtitle text size information, initial subtitle text color information, initial subtitle text transparency information, and initial subtitle text position information. In practice, the initial subtitle text size information may be used as the starting value and the target value, and linear interpolation may be performed according to a preset interpolation step number to obtain the interpolated subtitle text size information corresponding to each interpolation step number. The initial subtitle text color information may be used as the starting value and the target value, and linear interpolation may be performed according to a preset interpolation step number to obtain the interpolated subtitle text color information corresponding to each interpolation step number. The initial subtitle transparency information can be used as a starting value, the subtitle transparency information can be used as a target value, and linear interpolation can be performed according to a preset interpolation step number to obtain interpolated subtitle transparency information corresponding to each interpolation step. The center point corresponding to the initial subtitle position information can be used as a starting value, the center point corresponding to the subtitle position information can be used as a target value, and linear interpolation can be performed according to a preset interpolation step number to obtain interpolation center points corresponding to each interpolation step. Each interpolation center point can then be determined as each interpolated subtitle position information, or the spatial range corresponding to each interpolation center point can be determined as each interpolated subtitle position information. The size of the spatial range corresponding to the interpolation center point can be the same as the size of the spatial range corresponding to the subtitle position information. Finally, for each interpolation step number, the interpolated subtitle size information, interpolated subtitle color information, interpolated subtitle transparency information, and interpolated subtitle position information corresponding to the interpolation step number can be determined as smoothed subtitle information. Finally, the initial subtitle information, the smoothed subtitle information corresponding to each interpolation step, and the target subtitle information can be sequentially combined to form a smoothed subtitle information sequence. Thus, by dynamically adjusting the font size, color, transparency, and position, subtitles can be made more clearly visible under different lighting conditions. At the same time, through semantic segmentation and dynamic region selection, subtitles can be prevented from obscuring important visual information (such as faces and screens), thereby increasing the readability of subtitle text. In addition, by smoothing the initial and target subtitle text information, the display of subtitles can smoothly transition when environmental conditions change (such as from indoors to outdoors), avoiding visual discomfort caused by sudden switching.

[0101] Optionally, the smart glasses may be further configured to perform the following steps:

[0102] In the first step, in response to detecting a selection operation corresponding to a second translation mode, a directional range within a second preset angle below the end of the smart glasses is determined as the source direction of a second sound source. The second translation mode may be a mode representing face-to-face, two-way translation. In this second translation mode, both the wearer of the smart glasses and the user facing the wearer can view the corresponding translation information. The wearer of the smart glasses can view the translation information directly on the smart glasses, while the user facing the wearer can view the translation information directly on a mobile terminal. For example, the wearer can select a control representing the second translation mode in the display interface of the smart glasses. For example, the second preset angle may be 90°. The directional range within the second preset angle below the end of the smart glasses may be a directional range centered directly below the end of the smart glasses within the second preset angle. The second sound source may represent the spatial range below the end of the smart glasses at the second preset angle when the smart glasses are worn. The second sound source may correspond to the wearer of the smart glasses during cross-language communication. Therefore, in the second translation mode, the source direction of the second sound source may represent the range of the sound source of the wearer of the smart glasses, allowing the smart glasses to capture the wearer's speech.

[0103] In the second step, the audio data of the second sound source is directionally collected by the audio collection unit. Here, the specific implementation of directionally collecting the audio data of the second sound source can refer to any implementation method of directionally collecting the audio data of the first sound source, which will not be repeated here.

[0104] The third step is to transmit the audio data of the second sound source to the mobile terminal via a low-power transmission channel. In this way, in the second translation mode, the smart glasses can directionally collect the user's voice and transmit the collected audio data to the mobile terminal.

[0105] Optionally, the mobile terminal may be further configured to perform the following steps:

[0106] In the first step, the received audio data of the second sound source is translated to obtain translation information corresponding to the second sound source. In practice, the mobile terminal can translate the audio data of the second sound source through a translation engine to obtain translation information corresponding to the second sound source.

[0107] The second step is to display the translation information corresponding to the second sound source. In practice, the mobile terminal can directly display the translation information corresponding to the second sound source on the display screen.

[0108] In the third step, the microphone included in the mobile terminal collects audio data as mobile audio data. In the second translation mode, the user opposite the wearer of the smart glasses can carry a mobile terminal and speak into the microphone included in the mobile terminal, so that the mobile terminal can collect the voice of the user opposite.

[0109] The fourth step is to translate the mobile terminal audio data to obtain translation information corresponding to the mobile terminal audio data. In practice, the mobile terminal can translate the mobile terminal audio data through a translation engine to obtain translation information corresponding to the mobile terminal audio data.

[0110] The fifth step is to send the translation information corresponding to the audio data on the mobile terminal to the smart glasses, so that the smart glasses can display the translation information corresponding to the audio data on the mobile terminal.

[0111] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the translation system of some embodiments of the present disclosure, the user experience in cross-language communication is improved, the accuracy of voice recognition is improved, the quality of real-time translation is improved, and the network delay time is shortened. Specifically, the reasons for the poor user experience, recognition accuracy, real-time translation quality, and long network delay time are: relying on mobile phone translation software, the communicating parties need to frequently look down at the screen, interrupting eye contact and body language expression, handheld devices are prone to cause stiff movements, destroying natural conversation scenes, resulting in poor user experience, mobile phone microphones are non-directionally designed, and environmental noise is easy to interfere with recognition accuracy, resulting in poor recognition accuracy. Although smart glasses products have voice interaction functions, they are limited by local computing power and energy consumption, and the quality of real-time translation is poor. The cloud-based processing solution has the problem of long network delay time. Based on this, some embodiments of the present disclosure include a translation system comprising: a smart glasses device configured to use a built-in audio acquisition unit to directionally acquire audio data from a first sound source, transmit the audio data from the first sound source to a mobile terminal connected thereto via a low-power transmission channel, and display the translation information in response to receiving translation information from the mobile terminal; and a mobile terminal configured to translate the received audio data from the first sound source to obtain translation information corresponding to the first sound source, and transmit the translation information corresponding to the first sound source to the smart glasses device. Because the smart glasses device can directionally acquire audio data from a specific sound source, the impact of ambient noise can be reduced, the quality of the acquired audio data can be improved, and thus the accuracy of speech recognition can be improved. Furthermore, because the translation processing is performed by the mobile terminal, the smart glasses device is not limited by computing power and energy consumption, thereby improving translation quality. Furthermore, the smart glasses device and the mobile terminal are connected via a low-power transmission channel, which can reduce network latency during data transmission. Furthermore, the translation information obtained by the mobile terminal is directly displayed on the smart glasses device, and the smart glasses device directly acquires audio, eliminating the need for the user to frequently look down at the screen or focus on the speaker, thereby improving the user experience during cross-language communication. This improves the user experience during cross-language communication, improves the accuracy of speech recognition, improves the quality of real-time translation, and shortens network latency.

[0112] Further references Figure 3 The present disclosure provides some embodiments of a translation system. These system embodiments are related to Figure 2 The interactive method embodiments correspond to those shown.

[0113] In some embodiments, the translation system 300 may include:

[0114] The smart glasses end 301 is configured to directionally collect audio data of the first sound source through a built-in audio collection unit, send the audio data of the first sound source to a communication-connected mobile terminal through a low-power transmission channel, and display the translation information in response to receiving the translation information sent by the mobile terminal.

[0115] The mobile terminal 302 is configured to translate the received audio data of the first sound source, obtain translation information corresponding to the first sound source, and send the translation information corresponding to the first sound source to the smart glasses.

[0116] It is understandable that the smart glasses and mobile terminal described in the translation system 300 are similar to the reference Figure 2 The steps described above for the smart glasses and mobile terminal correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method are also applicable to the translation system 300 and the smart glasses and mobile terminal included therein, and will not be repeated here.

[0117] Reference below Figure 4 , which shows smart glasses 400 (e.g., Figure 1 Schematic diagram of the structure of smart glasses). Figure 4 The smart glasses shown are merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0118] like Figure 4 As shown, smart glasses 400 may include a processing device 401 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of smart glasses 400 are also stored in RAM 403. Processing device 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to bus 404.

[0119] Typically, the following devices can be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a microdisplay, a speaker, a vibrator, etc.; and a communication device 409. The communication device 409 can allow the smart glasses 400 to communicate with other devices wirelessly or by wire to exchange data. The above-mentioned smart glasses may include at least one display unit for imaging in front of the user's eyes. The display unit may include a microdisplay and an optical element. For example, the optical element may be an optical lens or an optical waveguide. Although Figure 4The smart glasses 400 are shown with various devices, but it should be understood that it is not required to implement or have all the devices shown. More or fewer devices may be implemented or have instead. Figure 4 Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0120] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.

[0121] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0122] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0123] The computer-readable medium may be included in the smart glasses, or may exist independently and not incorporated into the smart glasses. The computer-readable medium carries one or more programs. When executed by the smart glasses, the one or more programs cause the smart glasses to: directionally collect audio data from a first sound source via a built-in audio collection unit; transmit the audio data from the first sound source to a mobile terminal connected to the smart glasses via a low-power transmission channel; and, in response to receiving translation information from the mobile terminal, display the translation information.

[0124] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0126] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0127] Some embodiments of the present disclosure further provide a computer program product, including a computer program, which, when executed by a processor, implements the method performed by the above-mentioned smart glasses or mobile terminal.

[0128] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A translation system comprising: The smart glasses are configured to directionally collect audio data from a first sound source through a built-in audio collection unit, transmit the audio data from the first sound source to a mobile terminal connected thereto through a low-power transmission channel, and display the translation information in response to receiving the translation information transmitted by the mobile terminal; The mobile terminal is configured to translate the received audio data of the first sound source, obtain translation information corresponding to the first sound source, and send the translation information corresponding to the first sound source to the smart glasses.

2. The translation system according to claim 1, wherein: The smart glasses are further configured to, in response to detecting a selection operation corresponding to the first translation mode, determine a direction range of a first preset angle in front of the smart glasses as a source direction of the first sound source.

3. The translation system according to claim 1, wherein: The smart glasses are further configured to send the audio data of the first sound source to the mobile terminal in communication connection through the low-power transmission channel by the following steps: compressing the audio data of the first sound source to obtain compressed audio data; The compressed audio data is sent to the mobile terminal through the low-power transmission channel.

4. The translation system according to claim 1, wherein: The audio collection unit includes a microphone ring array consisting of at least two microphones, and the smart glasses are further configured to directionally collect audio data of the first sound source through the following steps: Synchronously collecting sound signals received by each microphone in the microphone ring array to obtain each sound signal; For each of the sound signals, generating sound source direction information of the sound signal; generating a directional enhanced sound signal according to each sound signal whose corresponding sound source direction information satisfies a preset direction condition corresponding to the first sound source; Audio data of a first sound source is generated according to the directional enhanced sound signal.

5. The translation system according to claim 1, wherein: The smart glasses include an ambient light sensor, and the smart glasses are further configured to display the translation information through the following steps: collecting ambient brightness information through an ambient light sensor included in the smart glasses; generating subtitle text configuration information corresponding to the translation information according to the ambient brightness information; The translation information is displayed according to the subtitle text configuration information.

6. The translation system according to claim 1, wherein: The smart glasses are further configured to: In response to detecting a selection operation corresponding to a second translation mode, determining a direction range of a second preset angle below the end of the smart glasses as a source direction of a second sound source; Directedly collecting audio data of the second sound source through the audio collection unit; The audio data of the second sound source is sent to the mobile terminal through a low-power transmission channel.

7. The translation system according to claim 6, wherein: The mobile terminal is further configured to: translating the received audio data of the second sound source to obtain translation information corresponding to the second sound source; displaying the translation information corresponding to the second sound source; collecting audio data as mobile terminal audio data through a microphone included in the mobile terminal; Translating the mobile terminal audio data to obtain translation information corresponding to the mobile terminal audio data; The translation information corresponding to the audio data of the mobile terminal is sent to the smart glasses terminal.

8. The translation system according to claim 4, wherein: The smart glasses are further configured to generate a directional enhanced sound signal according to each sound signal whose corresponding sound source direction information satisfies a preset direction condition corresponding to the first sound source through the following steps: Determining each microphone corresponding to each sound signal that meets the preset direction condition as a microphone set; Filtering a target microphone from the microphone set; For each microphone in the microphone set except the target microphone, perform the following steps: determining a distance between the microphone and the target microphone; generating a time difference between the microphone and the target microphone according to a preset focal length direction and the distance; performing delay compensation processing on the sound signal corresponding to the microphone according to the time difference to obtain an updated sound signal; Determine the sound signal of the target microphone and the updated sound signals as a sound signal set; For each sound signal in the sound signal set, performing the following steps: generating a signal-to-noise ratio corresponding to the sound signal; generating a first weight corresponding to the sound signal according to the signal-to-noise ratio; generating a second weight corresponding to the sound signal according to sound source direction information corresponding to the sound signal; generating an enhancement weight corresponding to the sound signal according to the first weight and the second weight; Each sound signal in the sound signal set is weighted according to the enhancement weight corresponding to each sound signal in the sound signal set to obtain a directionally enhanced sound signal.

9. Smart glasses, comprising: at least one display unit for forming an image in front of a user's eyes; one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method executed by the smart glasses end as claimed in one of claims 1-8.

10. A computer-readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

11. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Intelligent voice-to-text and simultaneous translation system based on microphone array

    CN111276150A

  • Translation method and device, wearable equipment and computer readable storage medium

    CN111814497A

  • Voice information processing method and system for AR glasses

    CN113763940A

  • Translation system based on artificial intelligence, intelligent glasses for translation and translation method

    CN119670767A

  • Rail system for automatic transport of objects to be coated

    KR102637805B1