Interaction method and system of intelligent glasses and translation machine

Voice is collected through the smart glasses multi-microphone matrix and transmitted to the translator in real time. It combines the preset acoustic model and the local environment vocabulary selection algorithm for translation processing, and real-time feedback of the translation results is achieved through the incremental feedback algorithm, which solves the problems of low speech recognition accuracy and insufficient translation real-time performance in the interaction between the smart glasses and the translator, improving the accuracy and user experience of translation.

CN120199236APending Publication Date: 2025-06-24深圳目渡科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510359125.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When interacting with the translator, existing smart glasses face the problems of low speech recognition accuracy and insufficient translation real-time performance, resulting in delay or interruption of translation results feedback.

Method used

Voice information is collected synchronously through multiple microphone matrices of smart glasses, and transmitted it to the translator in real time through wireless communication protocols. The translator uses a segmented processing algorithm with a preset acoustic model and a vocabulary selection algorithm based on the locale for processing, and feeds the translation results to the smart glasses through an incremental feedback algorithm.

Benefits of technology

It realizes high-precision speech recognition and fast translation, ensuring low latency and high accuracy feedback on translation results, significantly improving the fluency and immediacy of interaction, and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199236A_ABST
    Figure CN120199236A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the field of intelligent interaction, and relates to an interaction method of intelligent glasses and a translator, which comprises the following steps: the intelligent glasses synchronously acquire voice information of a user through a plurality of microphone matrixes; the voice information is transmitted to a translator in real time through a wireless communication protocol, a timestamp is added to each frame of voice information, and the translator detects transmission delay according to the timestamps and dynamically adjusts the transmission rate of the voice information; the translation machine divides the voice information into different subunits for processing according to syllables, intonations and grammar by adopting a preset acoustic model segmentation processing algorithm; the translation machine translates the voice information by using a vocabulary selection algorithm based on a language environment; and the translation machine feeds back a translation result to the intelligent glasses through an incremental feedback algorithm. The invention further provides an interaction system of the intelligent glasses and the translation machine. The objective of the invention is to realize low-delay and high-accuracy translation result feedback while ensuring high-precision speech recognition and rapid translation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent interaction technologies, and in particular, to an interaction method and system between a smart glasses and a translator. Background Art

[0002] With the rapid development of intelligent devices and artificial intelligence technologies, smart glasses, as a new type of wearable device, have been widely used in multiple fields, especially showing great potential in augmented reality (AR), health monitoring, voice interaction, etc. Smart glasses can display information in real time through built-in displays and sensors, and receive and process external voice, image and other data. However, when existing smart glasses interact with a translator for functions such as language translation, they face problems of low speech recognition accuracy and insufficient translation real-time performance. Summary of the Invention

[0003] The purpose of the embodiments of this application is to propose an interaction method and system between a smart glasses and a translator, aiming to ensure high-precision speech recognition and fast translation while achieving low-latency and high-accuracy translation result feedback.

[0004] To solve the above technical problems, the embodiments of this application provide an interaction method between a smart glasses and a translator, which adopts the following technical solutions:

[0005] An interaction method between a smart glasses and a translator includes the following steps:

[0006] The smart glasses synchronously collect the user's speech information through multiple microphone arrays;

[0007] The speech information is transmitted to the translator in real time through a wireless communication protocol, and a timestamp is added to each frame of speech information. The translator detects the transmission delay according to the timestamp and dynamically adjusts the transmission rate of the speech information;

[0008] The translator uses a segmentation processing algorithm of a preset acoustic model to divide the speech information into different sub-units according to syllables, intonations and grammar for processing;

[0009] The translator uses a vocabulary selection algorithm based on the language environment to translate the speech information;

[0010] The translator feeds back the translation result to the smart glasses through an incremental feedback algorithm.

[0011] In a possible implementation manner, after the step of the translator feeding back the translation result to the smart glasses through an incremental feedback algorithm, the method further includes:

[0012] The smart glasses monitor the user's fixation points in real time through an eye-tracking module, and automatically adjust the display position of the translation result or the voice playback order.

[0013] In a possible implementation, after the step of the smart glasses monitoring the user's fixation points in real time through an eye-tracking module and automatically adjusting the display position of the translation result or the voice playback order, the method further includes:

[0014] If the user's conversation needs change or are interrupted, the smart glasses automatically exit the current translation process and adjust the display content according to the voice commands or actions input by the user.

[0015] In a possible implementation, the timestamp is generated based on the system clock and is used to record the acquisition time of the voice information. The step of transmitting the language information to the translator in real time through a wireless communication protocol and adding a timestamp to each frame of voice information, and the translator detecting the transmission delay according to the timestamp and dynamically adjusting the transmission rate of the voice information specifically includes:

[0016] After receiving the voice information, the translator extracts the timestamp information therein and compares it with the local system clock to calculate the transmission delay of the voice data from the smart glasses to the translator;

[0017] Compare the transmission delay with a preset threshold. If the transmission delay is greater than the preset threshold, correction is performed.

[0018] In a possible implementation, the step of the translator using a segmentation processing algorithm of a preset acoustic model to divide the voice information into different sub-units according to syllables, intonations, and grammar specifically includes:

[0019] Perform preprocessing on the voice information, including noise removal, echo suppression, and gain adjustment;

[0020] Analyze the temporal characteristics of the voice information based on a hidden Markov model to identify the syllables in the voice information;

[0021] Extract the pitch, speech rate, and intonation change characteristics in the voice information based on a deep neural network;

[0022] Identify the grammatical structure of the voice information according to a preset language model and vocabulary library;

[0023] Divide the voice information into different sub-units according to syllables, pitch, speech rate, intonation, and grammatical structure;

[0024] Perform independent translation processing on each sub-unit.

[0025] In a possible implementation, the step of the translator using a language environment-based vocabulary selection algorithm to translate the speech information specifically includes:

[0026] Obtain the current language environment information, where the language environment information includes geographical location, time period, and historical translation records;

[0027] Analyze the current language environment through a context awareness module and establish a context model according to the user's language preference;

[0028] Adjust the vocabulary used in translation according to the context model.

[0029] To solve the above technical problems, an interaction system between a smart glasses and a translator is further provided in an embodiment of the present application, adopting the following technical solutions:

[0030] An interaction system between a smart glasses and a translator includes:

[0031] An acquisition module, configured to enable the smart glasses to synchronously acquire the user's speech information through a plurality of microphone matrices;

[0032] A communication module, configured to transmit the language information to the translator in real time through a wireless communication protocol and add a timestamp to each frame of speech information, and the translator detects the transmission delay according to the timestamp and dynamically adjusts the transmission rate of the speech information;

[0033] A division module, configured to enable the translator to divide the speech information into different sub-units for processing according to syllables, intonations, and grammar by using a segmentation processing algorithm of a preset acoustic model;

[0034] A translation module, configured to enable the translator to translate the speech information by using a language environment-based vocabulary selection algorithm;

[0035] A feedback module, configured to enable the translator to feedback the translation result to the smart glasses through an incremental feedback algorithm.

[0036] To solve the above technical problems, an embodiment of the present application further provides a computer device, adopting the following technical solutions:

[0037] A computer device includes a memory and a processor, where computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of an interaction method between a smart glasses and a translator as described above are implemented.

[0038] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, adopting the following technical solutions:

[0039] A computer-readable storage medium stores computer-readable instructions thereon, and when the computer-readable instructions are executed by a processor, the steps of an interaction method between a smart glasses and a translator as described above are implemented.

[0040] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:

[0041] An interaction method between a smart glasses and a translator disclosed in the present application, the smart glasses synchronously collect voice information of a user through multiple microphone matrices; transmit the language information to the translator in real time through a wireless communication protocol, and add a timestamp to each frame of voice information, and the translator detects transmission delay according to the timestamp and dynamically adjusts the transmission rate of the voice information; the translator uses a segmentation processing algorithm of a preset acoustic model to divide the voice information into different sub-units according to syllables, intonations and grammar for processing; the translator uses a vocabulary selection algorithm based on the language environment to translate the voice information; the translator feeds back the translation result to the smart glasses through an incremental feedback algorithm. The present application realizes low-latency real-time communication through a wireless communication protocol and a timestamp mechanism, ensures fast transmission and efficient processing of voice information during the translation process, the translator can detect and correct transmission delay in real time according to the timestamp, improves the accuracy and stability of translation; by precisely synchronizing the transmission of voice information and translation results, it avoids the user encountering translation delay or interruption in real-time conversations, significantly improves the fluency and immediacy of interaction, and enhances the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0043] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0044] Figure 2 is a flowchart of an embodiment of an interaction method between a smart glasses and a translator according to the present application;

[0045] Figure 3 is a schematic structural diagram of an embodiment of an interaction system between a smart glasses and a translator according to the present application;

[0046] Figure 4 is a schematic structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0048] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0049] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0050] Terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, etc.

[0051] The server 105 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal devices 101 , 102 , and 103 .

[0052] It should be noted that the interaction method between smart glasses and a translation machine provided in the embodiment of the present application is generally executed by a server. Correspondingly, an interaction system between smart glasses and a translation machine is generally set in the server.

[0053] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0054] Continue to refer Figure 2 , shows a flow chart of an embodiment of a method for interaction between smart glasses and a translation machine according to the present application. The method for interaction between smart glasses and a translation machine comprises the following steps:

[0055] Step S201, the smart glasses synchronously collect the user's speech information through multiple microphone arrays.

[0056] In this embodiment, an electronic device (such as Figure 1 the server shown) on which an interaction method between a smart glasses and a translator runs can send or receive data through a wired connection or a wireless connection. It should be noted that the above wireless connection methods can include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.

[0057] In this embodiment, the smart glasses are equipped with multiple microphone arrays for capturing the user's speech information. The simultaneous operation of multiple microphones can effectively improve the quality of speech collection and avoid interference with speech recognition due to noise or echo. By synchronously collecting through multiple microphone arrays, the clarity of the conversation content can be enhanced. Especially in a noisy environment, the layout of multiple microphone arrays also helps to improve the directivity of the speech source and the accuracy of the speech recognition system.

[0058] Step S202, transmit the speech information to the translator in real time through a wireless communication protocol, and add a timestamp to each frame of speech information. The translator detects the transmission delay according to the timestamp and dynamically adjusts the transmission rate of the speech information.

[0059] In this embodiment, the smart glasses transmit the collected speech information to the translator in real time through a wireless communication protocol (such as Wi-Fi, Bluetooth, 5G, etc.). The wireless communication protocol ensures that data can be transmitted quickly and stably. In order to monitor the transmission status of the speech information in real time, the smart glasses add a timestamp to each frame of speech data. The timestamp is a mark of the speech information, indicating the collection and transmission time of each frame of speech information. The translator detects the transmission delay by receiving the timestamp and comparing the transmission time of each frame of data. The detected transmission delay will be used to dynamically adjust the transmission rate of the speech data. In the case of an unstable network environment or a large delay, the translator will adjust the transmission size or priority of the data packet according to the delay situation to ensure the real-time and continuous transmission of the speech data. By marking each frame of speech information with a timestamp, the translator can accurately master the transmission situation of the speech data, automatically adjust the transmission strategy of the data stream, ensure that the speech information reaches the translator in time, and improve the real-time and accuracy of the translation process.

[0060] Step S203, the translator uses a segmentation processing algorithm of a preset acoustic model to divide the speech information into different sub-units according to syllables, intonations, and grammar for processing.

[0061] In this embodiment, the translator is built-in with a segmentation processing algorithm based on an acoustic model. This algorithm can segment the received speech signal, usually divided by syllables, intonation, and grammatical structure. These different subunits may include:

[0062] Syllable: The smallest pronunciation unit in speech.

[0063] Intonation: Refers to the changes in tone and intonation in speech (such as rising tone, falling tone, etc.).

[0064] Grammar: The grammatical structure in a sentence, such as subject-predicate-object, clause, interrogative sentence, etc.

[0065] The purpose of the segmentation processing algorithm is to convert the speech signal into structured language units, making subsequent speech recognition, grammatical analysis, and translation processing more accurate. The use of the acoustic model combines the physical characteristics of the speech signal and the audio features of speech, thereby improving the accuracy of speech recognition.

[0066] Step S204, the translator uses a vocabulary selection algorithm based on the language environment to translate the speech information.

[0067] In this embodiment, the translator uses a vocabulary selection algorithm based on the language environment to translate the speech according to the context and language environment of the speech information. This algorithm will dynamically select the vocabulary that best matches the current conversation environment. The translator adjusts the vocabulary selection according to information such as the user's geographical location (such as a restaurant, business meeting, etc.), the current time period (such as morning, night, etc.), the historical conversation content, and the user's translation preferences. For example, if the user is having a business conversation with others, the translator may give priority to using business-related vocabulary; if it is in a daily conversation, the system will select common life terms. The vocabulary selection algorithm based on the language environment makes the translation more in line with the actual communication scenario by analyzing the context, avoiding inappropriate or irrelevant translation content, and the analysis of the language environment helps to improve the accuracy and naturalness of the translation.

[0068] Step S205, the translator feeds back the translation result to the smart glasses through an incremental feedback algorithm.

[0069] In this embodiment, the translator divides the translation result into multiple data packets and feeds them back to the smart glasses device in sequence. The key to incremental feedback is that the translator not only transmits the complete translation result at one time, but gradually transmits the translation content frame by frame to ensure that the translation result is real-time, continuous and displayed in time. The translator gradually generates and feeds back the translation result during the translation process. The translation result can be text or voice. The incremental feedback algorithm ensures that the translation content is displayed on the display screen of the smart glasses in time, or played out through voice. The incremental feedback algorithm solves the problem of delay or discontinuity in the translation process, ensures the efficiency and real-time nature of the translation process, and through real-time feedback, users can continuously obtain translation results when communicating with others to avoid lag or interruption of the translation content.

[0070] This application realizes low-latency real-time communication through wireless communication protocols and timestamp mechanisms, ensuring the rapid transmission and efficient processing of voice information during the translation process. The translation machine can detect and correct transmission delays in real time according to timestamps, thereby improving the accuracy and stability of translation. By accurately synchronizing the transmission of voice information and translation results, users are prevented from encountering translation delays or interruptions in real-time conversations, significantly improving the fluency and immediacy of interaction and enhancing the user experience.

[0071] In some optional implementations of this embodiment, after the translation machine feeds back the translation result to the smart glasses through an incremental feedback algorithm, the method further includes:

[0072] The smart glasses monitor the user's gaze point in real time through an eye tracking module and automatically adjust the display position of the translation result or the voice playback order.

[0073] In this embodiment, the smart glasses are equipped with eye tracking sensors (such as infrared sensors, eye trackers, etc.) to monitor the user's eye movements in real time. Through eye tracking technology, the smart glasses can accurately sense the user's gaze point, that is, the specific location where the user's eyes are currently focused. The eye tracking sensor captures the movement trajectory of the eyes through infrared emission and reflection technology, and then analyzes the user's line of sight. These data are transmitted to the processing unit of the smart glasses for analysis and processing. Eye tracking technology can help the device perceive the user's attention and focus position in real time, and is widely used in smart glasses, virtual reality (VR) and augmented reality (AR) and other devices. Through eye tracking, the device can provide a more intelligent and personalized interactive experience in the user's vision.

[0074] Based on the real-time monitored gaze point, the smart glasses automatically adjust the display position of the translation results. If the user's gaze is focused on a specific area (such as the upper right corner), the translated text will automatically appear in that area instead of being fixed at a certain position on the screen. This ensures that the translated content is always in the position that is easiest for the user to see, thereby improving the user experience. For voice translation scenarios, smart glasses will also adjust the order of voice playback according to the user's gaze point. For example, if the user focuses on the previous sentence of the translation when talking to others, the smart glasses will give priority to playing the translation content of the sentence, and slightly delay the translation of the next sentence. This method can avoid the situation where the voice playback is inconsistent with the user's sight. Dynamic adjustment of the displayed content through eye tracking is the key to improving the user's interactive experience. By sensing the position of the user's gaze point in real time, smart glasses can adaptively adjust the position or order of the translated content to avoid the situation where the translated content is blocked or the playback order is inconsistent, ensuring that the translated content is always consistent with the user's focus.

[0075] This application uses an eye tracking module to monitor the user's gaze point in real time, and the smart glasses can dynamically adjust the display position of the translation result or the order of voice playback according to the user's line of sight. It can ensure that the translation result is always in the best position of the user's line of sight, avoiding the misplacement or inconvenient display of the translation information, thereby improving the user's interactive experience and visual comfort, especially in dynamic dialogue scenarios.

[0076] In some optional implementations of this embodiment, after the step of the smart glasses monitoring the user's gaze point in real time through the eye tracking module and automatically adjusting the display position of the translation result or the voice playback order, the smart glasses further include:

[0077] If the user's conversation needs change or are interrupted, the smart glasses automatically exit the current translation process and adjust the displayed content according to the voice instructions or actions input by the user.

[0078] In this embodiment, the smart glasses determine whether to exit the current translation process by detecting changes in the user's conversation needs. For example, the user may pause the conversation, switch topics, or stop interacting with the translator during the conversation, and the smart glasses automatically detect the interruption or change of the conversation based on these behaviors. The smart glasses can receive the user's voice commands (such as: "Pause translation", "Stop translation", or "Switch language"), and automatically adjust the translation process according to these commands. For example, if the user issues the voice command "Stop translation", the smart glasses will immediately abort the current translation task. In addition to voice commands, the smart glasses can also adjust the translation process through the user's action inputs. For example, the user may pause or stop the translation process by gestures (such as finger swiping, gesture recognition, etc.) or by touching the screen. In this case, the smart glasses can recognize these actions and respond. Once the conversation needs change or are interrupted, the smart glasses will automatically exit the current translation process and adjust the display content according to the new needs. For example, the smart glasses can display "Translation paused" or switch to a different display mode according to the user's instructions, such as entering the main interface, displaying other information, or performing other operations. Automatically recognizing changes or interruptions in conversation needs is the key to enhancing the interaction experience of smart devices. By integrating speech recognition and action recognition technologies, the smart glasses can quickly respond to changes in user needs and adapt to new conversation scenarios. This function is particularly suitable for multi-tasking environments and can automatically adjust the device's behavior when the user interrupts or changes topics to ensure that the translation function always matches the user's needs.

[0079] According to the voice commands or actions input by the user, the smart glasses of the present application can automatically recognize changes or interruptions in conversation needs, and timely adjust the translation process or exit the current translation process, enabling the user to flexibly control the translation process, avoiding unnecessary information display or mistranslation, and providing a more personalized and convenient translation experience, especially improving the adaptability and flexibility of the system in a changing conversation environment.

[0080] In some optional implementation manners of this embodiment, the above timestamp is generated based on the system clock and is used to record the acquisition time of the voice information. The step of transmitting the language information to the translator in real time through a wireless communication protocol and adding a timestamp to each frame of voice information, and the translator detecting the transmission delay according to the timestamp and dynamically adjusting the transmission rate of the voice information specifically includes:

[0081] After receiving the voice information, the translator extracts the timestamp information therein and compares it with the local system clock to calculate the transmission delay of the voice data from the smart glasses to the translator;

[0082] Compare the transmission delay with a preset threshold. If the transmission delay is greater than the preset threshold, make a correction.

[0083] In this embodiment, when the smart glasses collect each frame of voice information, they will use the internal system clock to generate a timestamp and attach this timestamp to the voice data frame. This timestamp records the collection time of each frame of voice data, that is, the exact time point when the user emits a voice signal. The system clock is the hardware clock inside the smart glasses, usually with high precision, and can accurately record the time of each frame of voice collection. This timestamp is crucial for subsequent delay detection and data synchronization. The timestamp is the key to synchronization and delay processing, and can ensure the timing consistency of voice information during transmission and processing. Through the timestamp, the transmission process of voice information can be accurately traced, helping to optimize the real-time performance of voice translation.

[0084] The smart glasses transmit the voice information to the translator through a wireless communication protocol (such as Wi-Fi, Bluetooth, 5G, etc.). Each frame of voice data will carry a timestamp during the transmission process, recording the exact time of collection. The translator dynamically adjusts the transmission rate of the voice information according to the received timestamp information and the current transmission status. If the network delay is large, the translator will slow down the transmission rate of the voice data to ensure that the data transmission will not be lost and can reach smoothly.

[0085] After receiving the voice data, the translator extracts the attached timestamp information from each frame of data. These timestamps represent the collection time of each frame of voice data and accurately reflect the sending time of the voice signal. The translator will compare the received timestamp with its own local system clock (or clock synchronization system). By calculating the difference between the current time of the received data and the timestamp, the translator can accurately calculate the transmission delay from the smart glasses to the translator. Transmission delay = current time - timestamp. This calculation result helps to evaluate the total delay duration from the collection of voice data to the reception by the translator. By comparing the timestamp with the local clock, the translator can monitor the delay situation during the data transmission process in real time. The purpose of calculating the transmission delay is to provide a basis for subsequent delay compensation and adjustment.

[0086] The translator presets a transmission delay threshold, which is the maximum allowable delay value. If the calculated transmission delay exceeds this threshold, the translator will take corrective measures to ensure that the delay during the translation process will not affect the real-time performance. When the transmission delay exceeds the preset threshold, the translator will correct it in the following ways:

[0087] Adjust the packet size: If the network bandwidth is low, the translator can reduce the size of each packet to reduce the delay of a single transmission.

[0088] Adjust the data transmission rate: The translator can also reduce the data transmission rate to ensure that the voice information can arrive and be processed on time, avoiding packet loss caused by too fast data transmission.

[0089] Retransmission mechanism: If some data packets are lost due to network fluctuations, the translator can trigger the retransmission mechanism to resend the unreceived data packets.

[0090] Delay correction is an important part of ensuring the real-time performance of speech translation. By dynamically adjusting the transmission rate and packet size, the translator can ensure the smoothness of the speech translation process. Especially in an unstable network environment, this correction mechanism can minimize the impact of delay on translation quality.

[0091] Through the timestamp-based delay correction mechanism in this application, the translator can accurately detect and correct the transmission delay of speech information, ensuring the real-time performance of information during the translation process. The introduction of timestamps enables the translator to effectively identify delays and make corrections when the delay exceeds a preset threshold, guaranteeing the data smoothness during the translation process and improving the stability of translation under changing network environments.

[0092] In some optional implementation manners of this embodiment, the step of the above-mentioned translator using a segmentation processing algorithm of a preset acoustic model to divide the speech information into different sub-units according to syllables, intonations, and grammar specifically includes:

[0093] Preprocess the speech information, including noise removal, echo suppression, and gain adjustment;

[0094] Analyze the temporal characteristics of the speech information based on the hidden Markov model to identify the syllables in the speech information;

[0095] Extract the pitch, speech rate, and intonation change characteristics in the speech information based on a deep neural network;

[0096] Identify the grammatical structure of the speech information according to a preset language model and vocabulary library;

[0097] Divide the speech information into different sub-units according to syllables, pitch, speech rate, intonation, and grammatical structure;

[0098] Perform independent translation processing on each sub-unit.

[0099] In this embodiment, when the translator preprocesses the voice information, it first applies a noise removal algorithm to filter out ambient noise. This step is to improve the clarity of the voice signal, especially in a noisy environment. The echo cancellation technology is used to eliminate the echo in the voice signal, ensuring that the voice signal is clear and non-overlapping, thereby improving the accuracy of speech recognition. Gain adjustment optimizes the volume of the voice signal to ensure that the amplitude of the voice signal is suitable for subsequent processing steps, avoiding the loss or distortion of voice information caused by too low or too high volume. Preprocessing is the basic step of voice signal processing. By removing background noise and echo and enhancing the voice signal, the translator can better extract effective information and improve the accuracy of subsequent speech recognition and translation.

[0100] Hidden Markov Model (HMM): HMM is a statistical model used to analyze the temporal characteristics of voice signals. In this step, the translator uses the HMM model to analyze the temporal characteristics in the voice signal and identify the syllables in the speech. HMM can handle the temporal correlation between individual sound units in the voice signal, helping the translator understand the structure of the speech.

[0101] Syllable recognition: Syllables are the basic units in speech recognition. Through HMM analysis, the translator can accurately identify the pronunciation, duration, and position of each syllable in the speech.

[0102] HMM is a widely used model in traditional speech recognition and is suitable for processing speech data with time series. By analyzing the temporal characteristics of syllables, the translator can extract meaningful syllable information from complex voice signals.

[0103] Deep Neural Network (DNN): The translator uses the DNN model to extract the pitch, speech rate, and intonation change characteristics in the voice signal. DNN can automatically learn and identify these characteristics in the voice signal, which are crucial for accurately understanding the meaning of sentences in speech recognition.

[0104] Pitch, speech rate, and intonation: These characteristics help better understand the emotional color, speaking speed, and speaking style of the speech. For example, too fast a speech rate or too high a pitch may indicate a tense or urgent mood in the conversation.

[0105] Pitch, speech rate, and intonation are important dimensions in language understanding, affecting the emotional expression and grammatical understanding of sentences. Through multi-level feature learning, DNN can extract these complex language features from the voice signal, thereby improving the accuracy of translation.

[0106] The translation machine uses a preset language model to perform grammatical analysis on the voice information. The language model contains an understanding of the grammar of a specific language, such as common sentence patterns, common vocabulary collocations, etc. Through the language model, the translation machine can understand the structure and grammar rules of sentences. The translation machine also performs lexical-level analysis on the voice information by means of a preset vocabulary database. The vocabulary database contains a large number of words and their contextual meanings. The translation machine uses the vocabulary database to help identify the meaning of each word and its role in the sentence. The combination of the language model and the vocabulary database enables the translation machine to not only stay at the word level during the translation process, but also understand the grammatical structure, thereby improving the translation quality and naturalness.

[0107] According to syllables, pitch, speech rate, intonation, and grammatical structure, the translation machine divides the entire voice information into multiple sub-units. These sub-units may be single syllables, words, phrases, or parts of sentences, and each sub-unit has its unique voice characteristics. The basis for dividing the voice information is usually the pitch, speech rate changes, grammatical structure, and other language characteristics of the voice signal, ensuring that the content within each sub-unit has strong relevance and independence. By dividing the voice information into smaller and more easily processed sub-units, the translation machine can perform more accurate translation tasks on each sub-unit, improving the quality of voice translation.

[0108] Each voice sub-unit will be processed separately. The translation machine performs independent translation on each sub-unit to ensure that each part of the translation process can accurately correspond to the meaning in the source language. Processing each sub-unit independently helps improve the accuracy of translation because the semantics of each sub-unit can be analyzed independently without being interfered by other parts. Performing independent translation processing on each sub-unit helps reduce mistranslation. Especially in the case of complex sentences, through fine-grained translation processing, the translation machine can better grasp the accurate meaning of each sub-unit and avoid misunderstandings in the overall translation.

[0109] This application processes the voice information in segments, accurately analyzes the characteristics of the voice such as syllables, pitch, and speech rate based on an acoustic model, enabling the translation machine to independently identify and process each sub-unit in the voice information, thereby improving the accuracy of voice recognition and translation; through preprocessing measures such as noise removal and echo suppression, the accuracy of translation and the system's processing ability for complex voices are further improved, and high-quality translation effects can still be ensured especially in an environment with high noise.

[0110] In some optional implementation manners of this embodiment, the step of the translation machine using a vocabulary selection algorithm based on the language environment to translate the voice information specifically includes:

[0111] Obtain the current language environment information, where the language environment information includes geographical location, time period, and historical translation records;

[0112] The context awareness module analyzes the current language environment and establishes a context model based on the user's language preferences;

[0113] Adjust the vocabulary used in translation according to the context model.

[0114] In this embodiment, this step involves the translator selecting appropriate vocabulary according to the language environment (such as geographical location, time period, and historical translation records) when translating voice information, so that the translation result is more in line with the current conversation context and user needs. Different language environments will affect the accuracy and naturalness of translation. For example, in a restaurant setting, users are more likely to need food and beverage related vocabulary, while in a business setting, users may need more business related terms. By analyzing this language environment information, the translator can dynamically adjust the translation content to ensure that the translation is more in line with the context.

[0115] The translator can obtain the current geographical location information through a built-in positioning module (such as GPS). Geographical location can help the translator identify the current city, country, or region, so as to decide whether to use localized vocabulary or specific dialects. The translator can also obtain the current time information. Different time periods (such as morning, afternoon, or evening) may be related to certain specific occasions or activities. For example, meetings or dining activities carried out at different times may have different language requirements. The translator can extract information from the user's past translation history, and this information can reflect the user's language preferences and commonly used vocabulary. For example, if a user often translates certain specific vocabulary in a specific field (such as tourism, business), the translator can infer the current translation needs based on these historical records. The acquisition of language environment information provides the translator with more context data to help it make more accurate translation decisions. By combining these environmental data, the translator can effectively adapt to different translation scenarios.

[0116] The context awareness module is an intelligent analysis module of the translator, which can combine various factors such as real-time obtained geographical location, time period, and historical translation records to evaluate the current conversation context. This module can dynamically sense the user's language needs and convert them into a translation model.

[0117] The context model is a dynamic adjustment mechanism established based on data of the user's language environment and preferences. This model can reflect the current conversation scenario, the user's personalized language needs, and the possible vocabulary categories. For example, if the user is in a travel scenario, the context model will automatically give priority to selecting travel related vocabulary, such as "flight ticket", "hotel", "scenic spot", etc.

[0118] The combination of the context awareness module and the context model enables the translator to not only understand the literal meaning of each word, but also adjust its translation content according to the current situation, thereby improving the accuracy and naturalness of translation. This dynamic adjustment makes the translation process more personalized and adaptable.

[0119] Based on the established context model, the translator automatically adjusts the vocabulary used during translation. For example, if the user is having a conversation in a restaurant, the translator will preferentially select catering-related words such as "menu", "order dishes", "waiter", rather than words unrelated to travel or business.

[0120] As the conversation scenario changes, the translator can adjust the translation vocabulary in real time. For example, when the user moves from a restaurant to a business meeting, the translator will quickly switch to business-related words such as "contract", "meeting", "sign a contract", etc.

[0121] The adjustment mechanism of the context model can ensure that the translator automatically adapts the correct vocabulary in different language environments to suit the current conversation content and requirements. This method improves the accuracy and context relevance of translation, avoiding mistranslation or translations that do not fit the occasion. Through the vocabulary selection algorithm based on the language environment, the translator can intelligently adjust the translation vocabulary according to the current geographical location, time period, and the user's historical translation records, ensuring that the translation result is more in line with the context of the current conversation, making the translation more natural, accurate, and contextually relevant, thereby enhancing the adaptability and user experience of translation and avoiding translation content that does not match the scenario.

[0122] In some alternative implementation manners of this embodiment, the step of the translator feeding back the translation result to the smart glasses through the incremental feedback algorithm specifically includes:

[0123] The translator divides the translation result into multiple incremental data packets, and each data packet represents a sub-unit of the translation;

[0124] Feed back the translation content to the smart glasses according to the priority of the translation content.

[0125] In this embodiment, the translator splits the translation result into multiple small data packets according to logical units or time periods. Each data packet contains a "sub-unit" of the translation result, which can be a word, a phrase, or even a part of a sentence. This method enables the translation result to be fed back to the smart glasses more quickly, avoiding delays for the user while waiting for the complete translation result. Each incremental data packet can represent a syllable in the voice signal, a partial translation of a sentence, or a translation segment. These sub-units are gradually constructed during the translation process to form the complete translation result. Splitting the translation result into multiple incremental data packets helps to provide real-time feedback during the translation process. Especially in complex sentences or contexts, by gradually feeding back the translation result, the smart glasses can display the translated part that has been completed in the shortest possible time, greatly improving the translation response speed and user experience.

[0126] The translator determines the feedback order according to the priority of the translation content. For example, if certain translation content (such as key sentences or important information) is more urgent or important than other content, the translator will send this content first. This priority setting can be determined based on various factors, such as the importance of the content, the length of the sentence, the urgency of the context, etc.

[0127] The translator gradually feeds back the data packets according to the priority, so that important translation information is displayed on the smart glasses as soon as possible, ensuring that the user can obtain the most critical information as soon as possible during the conversation, rather than waiting for the completion of the complete translation result.

[0128] If there are information changes during the translation process (such as new keywords or sentences appearing in the conversation), the translator will dynamically adjust the feedback order and priority to ensure that the most relevant content is presented to the user first.

[0129] The priority feedback mechanism is an intelligent algorithm that helps the translator optimize data transmission in cases where resources are limited or latency is high. By prioritizing the processing of important or urgent information, the translator can ensure that the user receives immediate and valuable feedback without having to wait for the complete translation result.

[0130] Through the incremental feedback algorithm in this application, the translator splits the translation result into multiple incremental data packets and gradually feeds them back to the smart glasses according to the priority order of the content, so that more efficient and real-time feedback can be achieved during the translation process. Priority control ensures that important content can be transmitted first, especially ensuring the rapid transmission of key translation information even when the network condition is poor, thus improving the real-time performance of translation and user experience.

[0131] Embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0132] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0133] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, read-only memory (ROM), or a random access memory (RAM), etc.

[0134] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0135] Further reference Figure 3 to Figure 2 As an implementation of the method shown above, the present application provides an embodiment of an interaction system between a smart glasses and a translator. This system embodiment corresponds to the method embodiment shown in Figure 2 and this system can be specifically applied to various electronic devices.

[0136] Such as Figure 3As shown in the figure, an interaction system between a smart glasses and a translator 300 in this embodiment includes: an acquisition module 301, an identification module 302, a calculation module 303, a training module 304, and a processing module 305. Among them:

[0137] The acquisition module 301 is configured to enable the smart glasses to synchronously acquire the user's voice information through a plurality of microphone matrices;

[0138] The communication module 302 is configured to transmit the language information to the translator in real time through a wireless communication protocol, and add a timestamp to each frame of voice information. The translator detects the transmission delay according to the timestamp and dynamically adjusts the transmission rate of the voice information;

[0139] The division module 303 is configured to enable the translator to divide the voice information into different sub-units for processing according to syllables, intonations, and grammar by using a segmentation processing algorithm of a preset acoustic model;

[0140] The translation module 304 is configured to enable the translator to translate the voice information by using a vocabulary selection algorithm based on the language environment;

[0141] The feedback module 305 is configured to enable the translator to feedback the translation result to the smart glasses through an incremental feedback algorithm.

[0142] The interaction system between a smart glasses and a translator provided by this application realizes low-latency real-time communication through a wireless communication protocol and a timestamp mechanism, ensures the fast transmission and efficient processing of voice information during the translation process. The translator can detect and correct the transmission delay in real time according to the timestamp, improving the accuracy and stability of translation; by precisely synchronizing the transmission of voice information and translation results, it avoids translation delays or interruptions that users may encounter in real-time conversations, significantly improving the fluency and immediacy of interaction and enhancing the user experience.

[0143] To solve the above technical problems, the embodiments of this application also provide computer devices. For details, please refer to Figure 4 , Figure 4 This is the basic structural block diagram of the computer device in this embodiment.

[0144] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0145] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.

[0146] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 4. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for an interaction method between a smart glasses and a translator. In addition, the memory 41 can also be used to temporarily store various data that have been output or will be output.

[0147] In some embodiments, the processor 42 may be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run the computer-readable instructions stored in the memory 41 or process data, such as running the computer-readable instructions of the interaction method between an intelligent glasses and a translator.

[0148] The network interface 43 may include a wireless network interface or a wired network interface, which is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0149] The computer device provided by this application realizes low-latency real-time communication through a wireless communication protocol and a timestamp mechanism, ensuring the fast transmission and efficient processing of voice information during the translation process. The translator can detect and correct transmission delays in real time according to the timestamp, improving the accuracy and stability of translation; by precisely synchronizing the transmission of voice information and translation results, it avoids translation delays or interruptions that users may encounter during real-time conversations, significantly improving the fluency and immediacy of interaction and enhancing the user experience.

[0150] This application also provides another implementation, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that the at least one processor executes the steps of the interaction method between an intelligent glasses and a translator as described above.

[0151] The computer-readable storage medium provided by this application realizes low-latency real-time communication through a wireless communication protocol and a timestamp mechanism, ensuring the fast transmission and efficient processing of voice information during the translation process. The translator can detect and correct transmission delays in real time according to the timestamp, improving the accuracy and stability of translation; by precisely synchronizing the transmission of voice information and translation results, it avoids translation delays or interruptions that users may encounter during real-time conversations, significantly improving the fluency and immediacy of interaction and enhancing the user experience.

[0152] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0153] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements for some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is equally within the scope of the patent protection of the present application.

Claims

1. A method for interaction between smart glasses and a translation machine, characterized in that: The steps include: Smart glasses collect user’s voice information synchronously through multiple microphone matrices; The language information is transmitted to the translator in real time through a wireless communication protocol, and a timestamp is added to each frame of voice information. The translator detects transmission delay according to the timestamp and dynamically adjusts the transmission rate of the voice information; The translation machine uses a segmentation processing algorithm of a preset acoustic model to divide the speech information into different sub-units according to syllables, intonation and grammar for processing; The translation machine translates the speech information using a vocabulary selection algorithm based on the language environment; The translation machine feeds back the translation result to the smart glasses through an incremental feedback algorithm.

2. The method for interaction between smart glasses and a translation machine according to claim 1, characterized in that: After the translation machine feeds back the translation result to the smart glasses through an incremental feedback algorithm, the step further includes: The smart glasses monitor the user's gaze point in real time through an eye tracking module and automatically adjust the display position of the translation result or the voice playback order.

3. The method for interaction between smart glasses and a translation machine according to claim 2, characterized in that: After the step of the smart glasses monitoring the user's gaze point in real time through the eye tracking module and automatically adjusting the display position of the translation result or the voice playback order, the smart glasses further include: If the user's conversation needs change or are interrupted, the smart glasses automatically exit the current translation process and adjust the displayed content according to the voice instructions or actions input by the user.

4. The method for interaction between smart glasses and a translation machine according to claim 1, characterized in that: The timestamp is generated based on the system clock and is used to record the acquisition time of the voice information. The steps of transmitting the language information to the translator in real time through the wireless communication protocol and adding a timestamp to each frame of the voice information, and the translator detecting the transmission delay according to the timestamp and dynamically adjusting the transmission rate of the voice information specifically include: After receiving the voice information, the translation machine extracts the timestamp information therein and compares it with the local system clock to calculate the transmission delay of the voice data from the smart glasses to the translation machine; The transmission delay is compared with a preset threshold, and if the transmission delay is greater than the preset threshold, correction is performed.

5. The method for interaction between smart glasses and a translation machine according to claim 1, characterized in that: The translation machine uses a segmentation processing algorithm of a preset acoustic model to divide the speech information into different sub-units according to syllables, intonation and grammar for processing, which specifically includes: Preprocessing the speech information, including noise removal, echo suppression and gain adjustment; Analyzing the temporal characteristics of the speech information based on a hidden Markov model to identify syllables in the speech information; Extract pitch, speaking speed, and intonation change features from speech information based on deep neural networks; Recognize the grammatical structure of the speech information according to a preset language model and vocabulary library; Divide speech information into different sub-units according to syllables, pitch, speaking rate, intonation and grammatical structure; Each subunit is processed for translation independently.

6. The method for interaction between smart glasses and a translation machine according to claim 1, characterized in that: The step of the translation machine translating the voice information using a vocabulary selection algorithm based on the language environment specifically includes: Obtaining current language environment information, the language environment information including geographic location, time period and historical translation records; Analyze the current language environment through the context-aware module and build a context model based on the user's language preference; The vocabulary used in the translation is adapted according to the context model.

7. The method for interaction between smart glasses and a translation machine according to claim 1, characterized in that: The step of feeding back the translation result to the smart glasses by the translation machine through an incremental feedback algorithm specifically includes: The translator divides the translation result into multiple incremental data packets, each of which represents a sub-unit of the translation; The translated content is fed back to the smart glasses according to the priority of the translated content.

8. An interactive system between smart glasses and a translation machine, characterized in that: include: A collection module, used to enable the smart glasses to synchronously collect the user's voice information through multiple microphone matrices; A communication module, used to transmit the language information to the translator in real time through a wireless communication protocol, and add a timestamp to each frame of voice information, wherein the translator detects transmission delay according to the timestamp and dynamically adjusts the transmission rate of the voice information; A division module, used to enable the translation machine to use a segmentation processing algorithm of a preset acoustic model to divide the speech information into different sub-units according to syllables, intonation and grammar for processing; A translation module, used to enable the translation machine to translate the voice information using a vocabulary selection algorithm based on a language environment; The feedback module is used to enable the translation machine to feed back the translation result to the smart glasses through an incremental feedback algorithm.

9. A computer device, characterized in that: It comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the method for interaction between smart glasses and a translation machine as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps of the method for interacting between the smart glasses and the translation machine as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Intelligent blind-assisting glasses system and method capable of realizing stereoscopic perception of environment

    CN120918922A