Voice information processing method and device for intelligent glasses and intelligent glasses
Through the combination of the multimodal conversion model and the tone library, smart glasses realize accurate recognition and translation of multispeaker voice, solving the dual challenges of speaker identity confusion and multilingual translation, ensuring the consistency and timing accuracy of translation results, and improving user interaction experience.
Patent Information
- Application Number
- CN202510883395.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-29
AI Technical Summary
Currently, smart glasses are difficult to accurately distinguish the identities of different speakers in multi-language translation scenarios where multiple people are talking in real time, resulting in confusion of voice content attributes, affecting the accuracy of personalized processing and translation results. At the same time, they cannot identify and translate different languages used by multiple speakers in real time, limiting their practicality in complex multilingual interaction scenarios.
By receiving the environmental audio marked with the recording time period collected by the smart glasses, the pre-constructed multi-modal conversion model is used to perform multi-modal analysis and translation processing, and an output item containing language word elements and tone vectors is generated, and the speech user is retrieved and confirmed in the tone library based on the tone vector, timing splicing is performed for the output items of the same user, target data is generated and output through the smart glasses.
It realizes accurate identification and translation of voice content for many speakers, solves the problem of confusion of speaker identity, ensures the consistency and timing accuracy of translation results, and improves user interaction experience and communication efficiency.
Smart Images

Figure CN120388567A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method and apparatus for processing voice information for smart glasses, and smart glasses. Background Art
[0002] Current smart glasses have obvious technical bottlenecks in the multi-language translation scenario of multi-person real-time conversations. First, when these devices perform speech recognition, it is usually difficult to accurately distinguish the identities of different speakers. When multiple users speak simultaneously or alternately, the speech recognition system relied on by smart glasses cannot effectively identify who each speaker is, resulting in confusion about the attribution of speech content and affecting the accuracy of subsequent personalized processing and translation results. Second, although the speech recognition system relied on by smart glasses has speech recognition and translation functions, when the conversation involves multiple languages, the speech recognition system relied on by smart glasses cannot real-time recognize and separately translate the different languages used by multiple speakers, limiting its practicality in complex, multi-language interaction scenarios. Therefore, current smart glasses face the dual challenges of being unable to achieve speaker identity recognition and multi-language synchronous translation. Summary of the Invention
[0003] Based on the above problems, this application provides a method and apparatus for processing voice information for smart glasses, and smart glasses, aiming to achieve accurate speaker identity recognition and multi-language synchronous translation to improve the practicality of smart glasses in multi-person real-time conversations.
[0004] The embodiments of this application disclose the following technical solutions: A method for processing voice information for smart glasses, characterized in that the method includes: Receiving ambient audio collected by smart glasses and marked with a recording time period; the ambient audio includes the speech of at least one speaking user; Using a pre-constructed multi-modal conversion model to perform multi-modal parsing and translation processing on the ambient audio to obtain a plurality of output items; each output item includes a language token and a timbre vector corresponding to the language token; the language token is marked with a time stamp, and the time stamp corresponds to the recording time period; For each output item, retrieving in a timbre library based on the included timbre vector to confirm the speaking user corresponding to the output item; For the language tokens in all output items corresponding to the same speaking user, performing sequential splicing based on the time stamps corresponding to each output item to obtain the target data of the speaking user; the target data is marked with a speaking time, and the speaking time corresponds to the time stamp of the first language token in the target data to indicate the start time of the target data; Output and display the target data corresponding to each speaking user through the smart glasses.
[0005] A voice information processing device for smart glasses, the device comprising: A receiving unit, configured to receive ambient audio marked with a recording time period collected by the smart glasses; the ambient audio includes the speech of at least one speaking user; An output item obtaining unit, configured to perform multimodal parsing and translation processing on the ambient audio by using a pre-constructed multimodal conversion model to obtain a plurality of output items; each output item includes a language token and a timbre vector corresponding to the language token; the language token is marked with a time stamp, and the time stamp corresponds to the recording time period; A second speaking user confirmation unit, configured to retrieve, based on the included timbre vector, in a timbre library for each output item to confirm the speaking user corresponding to the output item; A conversion content generation unit, configured to perform sequential splicing on the language tokens in all output items corresponding to the same speaking user based on the time stamps corresponding to each output item to obtain the target data of the speaking user; the target data is marked with a speaking time, and the speaking time corresponds to the time stamp of the first language token in the target data to indicate the start time of the target data; An output display unit, configured to output and display the target data corresponding to each speaking user through the smart glasses.
[0006] A smart glasses, the smart glasses comprising an information processing system for identifying user speech and user manual configuration information, the information processing system comprising: a display module, a transmission module, sensors and a control module; the display module includes smart display lenses; the sensors include a microphone and a speaker module; the control module includes a computing unit and a user interaction control unit of the smart glasses The transmission module is configured to send ambient audio, registered speech and user profiles, and receive converted audio and converted text; The smart display lenses are configured to present converted text and interaction information; The microphone is configured to collect ambient audio and user speech input; The speaker module is configured to play converted audio and prompt sounds; The computing unit is configured to process data collected by the sensors; The user interaction control unit is configured to receive user input instructions and user profiles.
[0007] Compared with the prior art, the present application has the following beneficial effects: In the embodiments of the present application, first, the environmental audio marked with the recording time period collected by the smart glasses is received. Then, the multi-modal parsing and translation processing is performed on the environmental audio by using the pre-constructed multi-modal conversion model to obtain a plurality of output items, and each output item includes a language token and a timbre vector. The time stamps of these output items correspond to the specific time points in the recording time period, ensuring the specific time position of the tokens in the environmental audio. For each output item, the corresponding speaking user is retrieved and confirmed in the timbre library based on its timbre vector. Then, for all the output items of the same speaking user, the time series splicing is performed based on the time stamps corresponding to each output item to generate the target data of the speaking user. Finally, the target data of each speaking user is output and displayed through the smart glasses.
[0008] By using the multi-modal conversion model, the present application converts the collected environmental audio into language tokens with time stamps and timbre vectors, realizing the accurate recognition and translation of the speech content of multiple speakers. Through the retrieval in the timbre library based on the timbre vector, the specific speaking user corresponding to each output item can be accurately confirmed, effectively solving the problem of speaker identity confusion in traditional devices. At the same time, based on the time stamps, the language token content of the same user is spliced in time series to generate a continuous and synchronous conversion result, ensuring the coherence and time series accuracy of the translation content. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Figure 1 FIG. 1 is a schematic diagram of an application scenario provided by an embodiment of the present application; Figure 2 FIG. 2 is a flowchart of a method for processing voice information for smart glasses provided by an embodiment of the present application; Figure 3 FIG. 3 is a schematic structural diagram of a multi-modal conversion model provided by an embodiment of the present application; Figure 4 FIG. 4 is a flowchart of a method for obtaining output items provided by an embodiment of the present application; Figure 5 FIG. 5 is a flowchart of a method for confirming a speaking user provided by an embodiment of the present application; Figure 6 FIG. 6 is a flowchart of a method for expanding a timbre library provided by an embodiment of the present application; Figure 7 FIG. 7 is a flowchart of another method for processing voice information for smart glasses provided by an embodiment of the present application; Figure 8 It is a flowchart of a method for updating a dynamic translation matrix provided by an embodiment of the present application; Figure 9 It is a schematic diagram of a voice information processing device for smart glasses provided by an embodiment of the present application; Figure 10 It is a schematic diagram of a smart glass provided by an embodiment of the present application. Detailed implementation manners
[0010] To facilitate understanding of the technical solutions provided by the embodiments of the present application, the background technologies related to the embodiments of the present application will be described first.
[0011] Current smart glasses have obvious technical bottlenecks in the multi-language translation scenario of multi-person real-time conversations. First, it is difficult for the device to accurately distinguish the identities of different speakers, resulting in confusion in the attribution of speech content, affecting personalized processing and the accuracy of translation results. Second, the speech recognition system of smart glasses cannot identify and translate different languages used by multiple speakers in real time, limiting its practicality in complex multi-language interaction scenarios. Therefore, smart glasses face the dual challenges of speaker identity recognition and multi-language synchronous translation.
[0012] Based on this, the embodiments of the present application provide a voice information processing method, device and smart glasses for smart glasses. After receiving the environmental audio marked with a recording time period collected by the smart glasses, this method first uses a pre-constructed multi-modal conversion model to perform multi-modal parsing and translation processing on the environmental audio to obtain multiple output items. Each output item includes a language token and a timbre vector, and the language token is marked with a time stamp, which corresponds to the original recording time period. For each output item, based on the timbre vector therein, a search is performed in the timbre library to confirm the identity of the corresponding speaking user. Then, for all output items of the same speaking user, temporal splicing is performed according to the time stamps to generate the target data of the speaking user. These target data are all marked with the speaking time, which indicates the start time of the target data and corresponds to the time stamp of the first language token in the target data. Finally, the target data of each speaking user are output and displayed through the smart glasses. The present application realizes the accurate recognition and translation of multi-speaker speech in environmental audio through a multi-modal conversion model combined with a preset language, uses timbre vector retrieval to accurately distinguish speaking users, solves the problem of speaker identity confusion, and ensures the coherence and temporal accuracy of the translation result by temporally splicing the speech and text of the same user through time stamps.
[0013] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0014] See also Figure 1 , Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of the present application. Figure 1 As shown, the application scenario includes three users A, B, and C. Each user wears smart glasses 110, and each user has pre-selected the expected receiving language in the glasses they wear. Each smart glasses is in communication with the server 120.
[0015] The smart glasses 110 are provided with a microphone and a computing unit. The microphone can receive ambient audio, and the computing unit can calculate the recording time period of the ambient audio, so that the server 120 can obtain the ambient audio marked with the recording time period from the smart glasses 110. The server 120 is used to perform the voice information processing for smart glasses provided in the embodiment of the present application.
[0016] In this embodiment, user A expresses "私は果物が好きです" (meaning "I like fruit" in Chinese) in Japanese, while user B expresses "我爱大模" in Chinese. These sounds interweave to create ambient audio. At this point, the microphones of the smart glasses 110 of users A, B, and C all receive this ambient audio, and the computing unit of the smart glasses 110 calculates the acquisition time period (i.e., the recording time period) for this ambient audio.
[0017] This embodiment uses English as an example. Specifically, server 120 first receives ambient audio captured by smart glasses and annotated with the recorded time period. The server then preprocesses the ambient audio, extracting Mel-Frequency Cepstral Coefficients (MFCC) features and converting the audio signal into a Mel-spectrogram representation that reflects the acoustic characteristics. This process effectively captures the time-frequency characteristics of speech, laying the foundation for subsequent speech recognition and translation tasks, and improving processing accuracy and efficiency in multilingual environments.
[0018] After that, the server 120 can use a pre-built multi-modal conversion model corresponding to the expected receiving language to perform multi-modal parsing and translation processing on the Mel spectrogram, obtaining multiple output items such as "Output item 1: [phonetic token: 'I', text token: 'I', timbre vector: 'v1'], Output item 2: [phonetic token: 'like', text token: 'like', timbre vector: 'v1'], Output item 3: [phonetic token: 'fruit', text token: 'fruit', timbre vector: 'v1'], Output item 4: [phonetic token: 'I', text token: 'I', timbre vector: 'v2'], Output item 5: [phonetic token: 'love', text token: 'love', timbre vector: 'v2'], Output item 6: [phonetic token: 'large', text token: 'large', timbre vector: 'v2']", Output item 7: "[phonetic token:'model', text token:'model', timbre vector: 'v2']". In this example, the timbre vector v1 represents the timbre characteristics of user A (who originally spoke Japanese), and v2 represents the timbre characteristics of user B (who speaks Chinese), reflecting the distinction between multiple speakers. Each output item contains the corresponding phonetic token, its transcribed text token, and the timbre vector representing the speaker's characteristics, and all outputs are in the target language of the expected receiving language. At the same time, each phonetic token and text token is marked with a time stamp, which is consistent with the environmental audio acquisition time period calculated by the smart glasses 110, ensuring the accurate alignment and synchronization of the recognition results in terms of time sequence.
[0019] Furthermore, the server 120 can retrieve and match the timbre vector included in each output item in the pre-established timbre library to confirm the specific speaking user corresponding to this output item. In this way, the server 120 can effectively associate the voice content with the speaker's identity. The timbre vectors of Output items 1, 2, and 3 all match the voiceprint characteristics of user A, so it is determined that these three output items belong to the speech of user A; while the timbre vectors of Output items 4, 5, 6, and 7 correspond to the voiceprint characteristics of user B, determining that these output items are from user B. This can not only distinguish the voice content of multiple speakers, but also provide a solid foundation for subsequent processing and analysis based on the speaker's identity.
[0020] Subsequently, for the voice tokens and text tokens in all output items corresponding to the same speaking user, the server 120 can perform sequential splicing based on their respective time stamps, orderly combine the discrete tokens in chronological order, and generate the continuous conversion audio and corresponding conversion text of the speaking user in the expected receiving language (such as English). For output items 1, 2, and 3 corresponding to user A, the server splices the voice tokens "I", "like", "fruit" and their corresponding text tokens in sequence to obtain the continuous conversion audio "audio_I_like_fruit.wav" and text "I like fruit" of user A; for output items 4, 5, 6, and 7 corresponding to user B, the server sequentially splices voice tokens such as "love", "large", "model", etc. and their corresponding texts to obtain the continuous conversion audio "audio_I_love_large_model.wav" and text "I love largemodel" of user B. The generated conversion audio and conversion text are both marked with the speaking time, which is used to indicate the starting moment of the entire spliced content and is consistent with the time stamp of the first voice token or text token in the splicing, so as to ensure the accurate positioning and coherence of the conversion result in time, and effectively support the synchronous display and analysis in a multi-speaker environment.
[0021] Finally, after the server 120 completes the speech recognition, translation, and sequential splicing of user A and user B, it can separately package the generated conversion audio and conversion text, and then send these personalized conversion results back to the corresponding smart glasses device through network transmission. In this way, the smart glasses of user A and user B can receive and present their own translation content in real time, realizing the accurate distribution and display of information in a multi-speaker environment, and improving the user's interaction experience and communication efficiency.
[0022] Those skilled in the art can understand that Figure 1 The shown framework schematic diagram is only an example in which the implementation manner of the present application can be realized. The scope of application of the implementation manner of the present application is not limited by any aspect of this framework.
[0023] To facilitate the understanding of the present application, a method for processing voice information for smart glasses provided in an embodiment of the present application will be described below with reference to the accompanying drawings.
[0024] See Figure 2As shown in the figure, this is a flowchart of a voice information processing method for smart glasses provided by an embodiment of the present application. For ease of description, hereinafter, the execution subject of this voice information processing method for smart glasses is taken as an example of a server for introduction. However, the present application does not specifically limit the execution subject of this method, and the execution subject of this method may also be other terminal devices or network edge computing nodes with corresponding processing capabilities, such as smart glasses.
[0025] As Figure 2 shown, this method may include S201 - S205: S201: Receive the ambient audio collected by the smart glasses and marked with a recording time period.
[0026] The smart glasses will capture the voice information in the surrounding environment in real time to obtain the ambient audio, and transmit the collected audio data (i.e., the ambient audio) to the server. Subsequently, the server receives this ambient audio from the smart glasses. This audio not only contains the voice information (i.e., the speech information) of at least one speaking user within a specific time period, but also is marked with an accurate recording time period to ensure that subsequent processing can accurately correspond to the occurrence moments of different voice segments.
[0027] In a possible implementation manner, since the received ambient audio often contains rich voice signals and background noise, it is necessary to perform effective feature extraction on it for subsequent recognition and translation. For this purpose, the original audio signal (i.e., the ambient audio) can be converted into a Mel spectrogram. Therefore, the method further includes: Convert the ambient audio into a Mel spectrogram.
[0028] Using the Mel scale to simulate the non - linear response of the human ear to different frequencies, through short - time Fourier transform and filter bank analysis, the audio signal is decomposed in the time - frequency domain into a feature representation with speech recognition significance. This representation method not only compresses redundant information, improves processing efficiency, but also enhances the speech model's ability to capture language and speaker features, providing a solid data foundation for subsequent multi - modal conversion models to perform high - precision speech recognition and translation. Through this method, it is possible to better understand and process multi - speaker speech inputs from complex environments and achieve more accurate speech conversion effects.
[0029] In a possible implementation manner, use the pre - constructed multi - modal conversion model to perform multi - modal parsing and translation processing on the ambient audio to obtain multiple output items, including: Use the pre - constructed multi - modal conversion model to perform multi - modal parsing and translation processing on the Mel spectrogram to obtain multiple output items.
[0030] S202: Use the pre - built multimodal conversion model to perform multimodal parsing and translation processing on the environmental audio, obtaining multiple output items.
[0031] After receiving the multimodal parsing and translation processing, the environmental audio can be processed based on the pre - built multimodal conversion model. This model is specifically trained and optimized for the expected receiving language, and can efficiently process the input environmental audio. Specifically, the multimodal conversion model first receives the environmental audio conversion as input to accurately identify and translate the speech signals in the audio. The processing results are presented in the form of multiple output items, and each output item contains not only the corresponding language token, but also the timbre vector corresponding to this language token.
[0032] In a possible implementation, the user has pre - selected the expected receiving language in the worn glasses. The smart glasses will send the collected environmental audio marked with the recording time period and the expected receiving language to the server together, so that the server can ensure that the output language of the multimodal conversion model used to process the current environmental audio is consistent with the expected receiving language.
[0033] In a possible implementation, the language token includes a speech token and the corresponding text token of this speech token.
[0034] It can be understood that the speech token represents the basic unit of sound, the text token is the text content corresponding to this sound, and the timbre vector reflects the unique voiceprint characteristics of the speaker, which helps to distinguish different speakers.
[0035] In addition, the language type of the output items output by the multimodal conversion model is consistent with the expected receiving language to ensure that the translated content meets the target language requirements. To ensure the temporal consistency of the speech recognition and translation results, all speech tokens and text tokens (i.e., language tokens) are marked with time stamps, and these time stamps correspond to the recording time period in the original environmental audio, so as to achieve precise positioning and synchronous processing of the speech content of multiple speakers on the time axis. This process greatly improves the accuracy of speech recognition and the coherence of translation, laying a solid foundation for subsequent content splicing and personalized processing based on the speaker's identity.
[0036] Exemplarily, assume that the smart glasses collect an environmental audio with a recording time period from 10:00:05 to 10:00:10. After processing this audio, the multimodal conversion model identifies several speech tokens and corresponding text tokens in a sentence: Speech token 1: "hey", text token: "hey", time stamp: 10:00:05.2, indicating that this token occurs at 10:00:05 seconds and 200 milliseconds. Speech token 2: "I", text token: "I", time marked at 10:00:05.5, indicating that this token occurred at 10 hours, 0 minutes, 5 seconds, and 500 milliseconds; Speech token 3: "want", text token: "want", time marked at 10:00:06.0, indicating that this token occurred at 10 hours, 0 minutes, 6 seconds.
[0037] These time marks are all within the time period from 10:00:05 to 10:00:10 of the original ambient audio recording, ensuring that the speech recognition and text content can be accurately corresponding to specific time nodes, achieving precise matching in terms of time sequence.
[0038] S203: For each output item, retrieve in the voiceprint library based on the included voiceprint vector to confirm the speaking user corresponding to this output item.
[0039] During the speech recognition and translation process, each output item not only contains speech tokens and corresponding text tokens, but also carries a voiceprint vector representing the voice characteristics of the speaker. For each such output item, this voiceprint vector can be used as an index to retrieve in the pre-established voiceprint library. The voiceprint library stores the voice characteristics of multiple known speaking users. By comparing the voiceprint vector of the current output item with the characteristics in the voiceprint library, the similarity or matching degree can be calculated, so as to accurately determine which specific speaking user this speech segment belongs to. This process effectively differentiates the voice characteristics of different speakers, avoiding recognition errors caused by the confusion of speaker identities in traditional speech processing systems. At the same time, the matching method based on the voiceprint vector has strong robustness and can adapt to complex scenarios such as environmental noise and multi-speaker crossover, realizing the accurate attribution of multi-speaker speech. Through this retrieval and confirmation mechanism, not only can the speech content be bound to the correct speaker, but also a reliable basis is provided for subsequent personalized processing, dialogue management, and multilingual translation, which helps to improve the accuracy of overall speech information processing and the user experience.
[0040] S204: For the language tokens in all output items corresponding to the same speaking user, perform time sequence splicing based on the time marks corresponding to each output item to obtain the target data of this speaking user.
[0041] For all output items corresponding to the same speaking user, based on these time marks, the speech tokens and text tokens (i.e., language tokens) can be spliced in order according to the time sequence, continuously integrating the scattered speech segments and text contents, so as to generate the target data of this speaking user.
[0042] In a possible implementation, the target data includes converted audio and converted text. Therefore, the sequential splicing in this application not only ensures the coherence of the speech content and text information, but also effectively avoids the situation of semantic breakage or confusion.
[0043] Meanwhile, for the convenience of subsequent processing and display, the generated converted audio and converted text (i.e., the target data) are both assigned a unified speaking time, which indicates the starting moment of the spliced content, and its value is consistent with the time stamp marked by the first speech token or text token in the splicing sequence. In this way, the speech and translation content of each speaking user can be accurately located and synchronized, ensuring the accurate timing and logical coherence of the speech information in a multi-speaker environment, providing clear, synchronized and personalized speech conversion results for terminal devices such as smart glasses, and greatly enhancing the user's auditory understanding and interaction experience.
[0044] S205: Output and display the target data corresponding to each speaking user through the smart glasses.
[0045] After completing the recognition, translation, and sequential splicing based on time stamps of the audio in a multi-speaker environment, the converted audio and converted text (i.e., the target data) corresponding to each speaking user can be generated. These contents not only accurately reflect the user's speech information, but also maintain the coherence of the language and the consistency of the timing. Next, these personalized conversion results need to be output and displayed through the smart glasses terminal so that users can obtain clear and structured speech information in real-time or near real-time. For this purpose, the converted audio and converted text corresponding to each speaking user are packaged and sent to the smart glasses through a stable and reliable data transmission channel. This process not only ensures the integrity of the data without loss, but also takes into account the transmission delay and bandwidth limitations to ensure that the final presented speech and text content can be updated synchronously, enhancing the interaction experience. At the same time, by separating and classifying the conversion results of multiple speakers and sending them, the smart glasses can perform differential display or processing for different speaking users, enabling the wearer to clearly distinguish the information sources of each speaker and achieving more accurate speech understanding and multilingual communication.
[0046] Based on the descriptions of S201 - S205, an embodiment of the present application provides a method for processing voice information for smart glasses. First, the smart glasses collect and label the ambient audio with time periods. Then, using a multi - modal conversion model, voice recognition and translation are performed on the ambient audio to generate multiple output items containing language tokens and timbre vectors, and all tokens are with time stamps corresponding to the recorded time periods. Next, the identity of the speaking user for each output item is determined by retrieving in the timbre library according to the timbre vector. Subsequently, for all output items of the same user, sequential splicing is performed according to the time stamps to generate continuous target data for this user and label the speaking time, which indicates the start time of the target data and is consistent with the time stamp of the first voice or text token in the target data. Finally, the converted audio and text of each speaking user are output and displayed through the smart glasses to achieve personalized and multi - language real - time voice processing. The present application uses a multi - modal conversion model to convert the ambient audio into language tokens with time stamps and timbre vectors, realizing accurate recognition and translation of the speech content of multiple speakers. By retrieving the timbre vector in the timbre library, the speaking user corresponding to each output item is accurately identified, effectively avoiding the problem of speaker identity confusion in traditional devices. At the same time, with the help of time stamps, sequential splicing of the language tokens of the same user is performed to generate continuous and synchronous conversion results, ensuring the coherence and sequential accuracy of the translation content.
[0047] In a possible implementation manner, an embodiment of the present application further provides a structural schematic diagram of a multi - modal conversion model, as Figure 3 shown. This multi - modal conversion model includes a multi - modal encoding unit 301 and a multi - modal decoding unit 302. The multi - modal decoding unit 302 is obtained by gradually performing speech synthesis training, language translation training, and cross - language conversion training through a basic decoding unit. The multi - modal encoding unit 303 is obtained by gradually performing text - audio alignment training, timbre alignment training, and language alignment training through a basic encoding unit.
[0048] Among them, the multi - modal decoding unit 302 includes a timbre decoding unit 3021, a speech decoding unit 3022, a text decoding unit 3023, and a language decoding unit 3024. The input ends of the speech decoding unit 3022 and the text decoding unit 3023 are both connected to the output end of the multi - modal encoding unit; the output end of the timbre decoding unit 3021 is respectively connected to the input ends of the speech decoding unit 3022 and the text decoding unit 3023; the input end of the timbre decoding unit 3021 is connected to the output end of the multi - modal encoding unit; the output end of the language decoding unit 3024 is respectively connected to the input ends of the speech decoding unit 3022 and the text decoding unit 3023.
[0049] The multi - modal encoding unit 301 is used to perform feature extraction and fusion on the ambient audio to obtain a vector representation; The timbre decoding unit 3021 is configured to provide converted timbre information for the speech decoding unit and the text decoding unit, and perform timbre feature decoding processing on the vector representation to obtain the timbre vector; The language decoding unit 3022 is configured to provide output language information for the speech decoding unit and the text decoding unit; The speech decoding unit 3023 is configured to perform speech decoding processing on the vector representation based on the converted timbre information and the output language information to generate a plurality of speech tokens; The text decoding unit 3024 is configured to perform text decoding processing on the vector representation based on the converted timbre information and the output language information to generate a plurality of text tokens.
[0050] In a possible implementation manner, the present application provides a method for obtaining output items. Refer to Figure 4 , Figure 4 which is a flowchart of a method for obtaining output items provided by an embodiment of the present application. Correspondingly, step S201 uses a pre-constructed multimodal conversion model to perform multimodal parsing and translation processing on the environmental audio to obtain a plurality of output items, which can be specifically implemented through steps S401 - S404: S401: Obtain the output language information in the language decoding unit and the converted timbre information in the timbre decoding unit, and use the multimodal encoding unit to perform encoding processing on the environmental audio to obtain the vector representation of the environmental audio.
[0051] In this step, first, obtain the target output language information from the language decoding unit and the converted timbre information from the timbre decoding unit. Then, use the multimodal encoding unit to perform encoding processing on the environmental audio and convert it into a vector representation form of the environmental audio. This process converts the original audio signal into a digital representation form that can be used in subsequent processing.
[0052] S402: Use the timbre decoding unit to perform timbre feature decoding processing on the vector representation to obtain a plurality of timbre vectors.
[0053] In this step, the encoded environmental audio vector is sent to the timbre decoding unit for processing. This unit is responsible for parsing the timbre features in the vector representation and generating a plurality of timbre vectors. These timbre vectors represent the timbre information of different parts of the audio signal and help to identify and distinguish different speakers.
[0054] S4031: Use the speech decoding unit to perform speech decoding processing on the vector representation based on the converted timbre information and the output language information to obtain the speech tokens in the language tokens.
[0055] Using the voice decoding unit, based on the previously obtained converted timbre information and output language information, further process the environmental audio vector to extract the voice tokens in the language tokens. This step will identify and separate the language components in the audio, including the specific content of the voice data.
[0056] S4032: Use the text decoding unit to perform text decoding processing on the vector representation based on the converted timbre information and the output language information to obtain the text tokens in the language tokens.
[0057] Similar to step S403, the text decoding unit here also processes the environmental audio vector based on the converted timbre information and the output language information, but this time it is to generate the text tokens in the language tokens. This is the process of converting voice data into corresponding text information.
[0058] S404: For each voice token, combine the voice token, the text token corresponding to the voice token, and the timbre vector corresponding to the voice token based on its time stamp to obtain an output item.
[0059] Finally, for each voice token, based on its time stamp, combine the voice token, the corresponding text token, and the timbre vector into an output item. This step integrates the results of the previous steps to ensure that each output item contains complete voice, text, and timbre information, facilitating subsequent analysis and application. In a possible implementation, the basic decoding unit includes a basic timbre decoding unit, a basic voice decoding unit, a basic text decoding unit, and a basic language decoding unit.
[0060] In a possible implementation, the voice synthesis training of the basic decoding unit includes: Dataset construction: Construct a first training dataset, which contains voice vector samples with text vector labels, and these samples have the same language as the labels.
[0061] Training objective: Minimize the text output error by minimizing the timbre prediction error and the language prediction error. Use the first loss function to measure these errors.
[0062] Result: Obtain the voice synthesis decoding unit.
[0063] In a possible implementation, the language translation training of the voice synthesis decoding unit includes: Dataset construction: Construct a second training dataset, which contains text vector samples with voice vector labels, and these samples have the same language as the labels.
[0064] Training objective: Minimize the voice output error by minimizing the timbre prediction error and the language prediction error. Use the second loss function to measure these errors.
[0065] Result: A language translation decoding unit is obtained.
[0066] In a possible implementation, the cross-lingual conversion training of the language translation decoding unit includes: Dataset construction: Construct a third training dataset, which contains text vector samples with speech vector labels and speech vector samples with text vector labels. The languages of the samples and the labels are different, but the output languages are the same.
[0067] Training objective: Minimize the multi-lingual translation text error and the speech output error by minimizing the timbre prediction error and the language prediction error. Use the third loss function to measure these errors.
[0068] Result: A multi-modal decoding unit is obtained.
[0069] In a possible implementation, the text-audio alignment training of the basic encoding unit includes: Dataset construction: Construct a fourth training dataset, which contains multiple pairs of first sample pairs. Each pair includes a text sample and an audio sample. The first sample pairs are divided into two categories: one category is text and audio samples with the same language but different semantics; the other category is text and audio samples with the same language and the same semantics.
[0070] Training objective: Use the fourth loss function to make the basic encoding unit maximize the text-audio alignment error in the first category of sample pairs and minimize the text-audio alignment error in the second category of sample pairs.
[0071] Result: An audio alignment encoding unit is obtained.
[0072] In a possible implementation, the timbre alignment training of the audio alignment encoding unit includes: Dataset construction: Construct a fifth training dataset, which contains multiple pairs of second sample pairs. Each pair includes two audio samples. The second sample pairs are divided into two categories: one category is audio sample pairs with the same semantics but different timbres; the other category is audio sample pairs with the same semantics and the same timbre.
[0073] Training objective: Use the fifth loss function to make the audio alignment encoding unit maximize the timbre alignment error within the first category of audio sample pairs and minimize the timbre alignment error within the second category of audio sample pairs.
[0074] Result: An audio timbre alignment encoding unit is obtained.
[0075] In a possible implementation, the language alignment training of the audio timbre alignment encoding unit includes: Dataset construction: Construct the sixth training dataset, which includes multiple groups of first sample groups. Each group contains text samples and audio samples in two languages under the same semantics.
[0076] Training objective: Use the sixth loss function to maximize the language alignment error between text and audio samples in different languages within the first sample group, while minimizing the language alignment error between text and audio samples in the same language.
[0077] Result: Obtain a multi-modal coding unit.
[0078] Through a hierarchical and modular design, this application integrates and decodes multi-modal (timbre, speech, text, language) features, realizing an efficient conversion from environmental audio to multi-language, multi-modal speech and text output. Its training process is refined step by step to ensure the accurate alignment and expression of various features.
[0079] In a possible implementation, the timbre library includes multiple user voiceprint record entries; each user voiceprint record entry includes a user identification ID, a user voiceprint feature, and a user profile. This structured management method helps to accurately distinguish and quickly identify different user identities.
[0080] In a possible implementation, the user profile includes, but is not limited to, user-related personal information (such as name, gender, language preference, etc.) and avatar setting information.
[0081] In a possible implementation, this application provides a method for confirming the speaking user. Refer to Figure 5 , Figure 5 which is a flowchart of a method for confirming the speaking user provided by an embodiment of this application. Correspondingly, for each output item in step S203, retrieve in the timbre library based on the included timbre vector to confirm the speaking user corresponding to this output item, which can be specifically implemented through steps S501 - S503: S501: For each output item, calculate the similarity between each user voiceprint feature in the timbre library and the timbre vector to obtain multiple similarities.
[0082] For each output item, the timbre vector included in this output item can be compared one by one with each user voiceprint feature stored in the timbre library, and multiple similarity values can be obtained through the similarity calculation method, so as to quantify the matching degree between this timbre vector and different user voiceprint features.
[0083] S502: Identify the similarities among the multiple similarities that are greater than the similarity threshold, and determine the user voiceprint feature corresponding to the similarity greater than the similarity threshold as the target voiceprint feature.
[0084] After obtaining multiple similarity degrees, those values exceeding a preset similarity threshold can be filtered out from these similarity values, and it is considered that the user voiceprint features with relatively high similarity have a strong matching relationship with the current tone vector. Subsequently, these user voiceprint features that meet the conditions are used as target voiceprint features for subsequent confirmation of the speaking user.
[0085] It should be noted that when more than one similarity value exceeds the preset similarity threshold, the magnitudes of these similarity values can be compared, and the user voiceprint feature corresponding to the maximum similarity value is selected as the final target voiceprint feature to ensure the accuracy and uniqueness of the speaking user confirmation.
[0086] S503: Determine the speaking user corresponding to the output item based on the user profile associated with the target voiceprint feature.
[0087] By searching for the user profile associated with the determined target voiceprint feature, the specific user information corresponding to this voiceprint feature is obtained, and then the identity of the speaking user of the current output item is accurately identified and determined.
[0088] Steps S501 - S503 can comprehensively evaluate the matching degree of potential speaking users by calculating the similarity between the tone vector of each output item and the user voiceprint features in the tone library. At the same time, by setting the similarity threshold, the target voiceprint features with high correlation are effectively filtered out.
[0089] In a possible implementation manner, if there is no similarity greater than the similarity threshold among the multiple similarities generated by step S501, then the tone library needs to be expanded. For this purpose, this application provides a method for expanding the tone library. See Figure 6 , Figure 6 is a flowchart of a method for expanding the tone library provided by an embodiment of this application, which can be specifically implemented through steps S601 - S604: S601: Prompt the users in the scene where the smart glasses are located and whose user voiceprint record entries are not entered to set up user profiles, and prompt the users whose user voiceprint record entries are not entered to send a registration voice that conforms to the text template.
[0090] When it is detected that there is a user in the current environment who has not been registered in the voiceprint library, the server will actively prompt this user through the smart glasses to complete the setting of the personal information profile. At the same time, guide this user to read a registration voice according to a predetermined text template through the smart glasses, so as to collect their unique voiceprint features and establish a complete and standardized user voiceprint record.
[0091] S602: Identify the user voiceprint features of the registration voice and assign a user ID to this user voiceprint feature.
[0092] When the server receives the user profile and the registered voice, it can extract the voiceprint features from the registered voice read by the user according to the text template, and generate a unique voiceprint feature representation for the user. Subsequently, a unique user ID is assigned to this voiceprint feature to achieve the unique identification and management of the user's identity, facilitating subsequent voiceprint matching and user identification.
[0093] S603: Associate the newly recognized user voiceprint features, the corresponding newly set user profile, and the newly assigned user ID to obtain a new user voiceprint record entry and store it in the tone color library.
[0094] Then, bind the newly recognized user voiceprint features with the newly set personal profile information and the assigned unique user ID of the user to form a complete user voiceprint record entry. This entry is then stored in the tone color library to achieve the unified management of the new user's identity and voiceprint data, providing a reliable data basis for subsequent user identification and voice processing.
[0095] S604: Execute again the "recognize the similarity greater than the similarity threshold among the multiple similarities" in step S502. If there is still no similarity greater than the similarity threshold, repeat steps S601 - S603 until a similarity greater than the similarity threshold is found.
[0096] After the new user's voiceprint features are entered, the similarity screening step (i.e., "recognize the similarity greater than the similarity threshold among the multiple similarities" in step S502) will be executed again to check whether there is a matching item in the tone color library whose similarity with the current voiceprint vector exceeds the preset threshold. If no matching similarity that meets the conditions is still found, repeat to guide the user to set the user profile and collect the registered voice through the smart glasses (steps S601 - S603) until a valid match with a similarity greater than the threshold is successfully obtained, thus ensuring the accurate entry and recognition of the new user's voiceprint data.
[0097] Steps S601 - S603 can effectively avoid the problems of misidentification and loss of new user information, achieve dynamic, continuous, and accurate user voiceprint update and entry, and significantly improve the adaptability of the server and the user coverage rate.
[0098] In a possible implementation manner, the method further includes: Once the user registers successfully, the server will send the information of successful registration to the smart glasses for display, and remind the user that they can directly log in through voiceprint recognition in future use. At the same time, the user's voiceprint model will be optimized, and the user's personalized settings (such as voice echo, translation preferences, etc.) will be customized according to the user profile.
[0099] Specifically, after the user completes the registration process, the server will immediately display a registration success notification through the smart glasses and inform the user that they can log in quickly through voiceprint recognition in the future. In addition, the smart glasses will further optimize the user's voiceprint data to improve the accuracy and efficiency of recognition. The user's personalized settings, such as voice echo and translation preferences, will also be adjusted and customized according to the user's personal profile, thus providing a more personalized user experience.
[0100] In a possible implementation manner, the method further includes: after registration is completed, the user can continue to use the smart glasses for daily operations. The server will automatically recognize the user in subsequent conversations and provide personalized voice services, translation services, voice echo and other functions according to their profile. This means that the user does not need to log in again, and the server will automatically recognize the user's identity and provide a customized service experience.
[0101] In a possible implementation manner, the method further includes: the user profile supports dynamic updates. When the user's voice or personal information changes, the user can perform simple voice or touch operations through the smart glasses to update their profile. The server will re-collect new voiceprint data through the smart glasses and update the user's voice model and personalized settings to ensure that the service always meets the user's latest needs.
[0102] Steps S601 - S604 actively prompt users who have not registered their voiceprints in the smart glasses scenario to register, ensuring that each new user can establish a complete user profile and voiceprint record in a timely manner, realizing dynamic and real-time user identity management. With the registration voice that conforms to the text template, the accuracy and consistency of voiceprint recognition are effectively improved, avoiding the influence of environmental noise or non-standard speech on the recognition effect. The newly recognized user voiceprint features are associated and stored with the user profile and ID, improving the data richness of the voice model and supporting more accurate multi-user differentiation in the future. By repeatedly executing the similarity comparison and registration process, the server's continuous learning and adaptation ability to unknown users is ensured, minimizing the risk of misrecognition and enhancing the user experience.
[0103] In a possible implementation manner, the present application further provides a method for processing voice information for smart glasses. Refer to Figure 7 , Figure 7 which is a flowchart of another method for processing voice information for smart glasses provided by the embodiments of the present application. Specifically, it can be implemented through steps S701 - S702: S701: For each user voiceprint feature in the voice model, extract and model the voice characteristics of the user voiceprint feature to obtain a voice matrix.
[0104] For each user voiceprint feature stored in the voiceprint library, by extracting its unique timbre feature parameters and using modeling techniques to structurally represent these parameters, a corresponding timbre matrix is finally generated. This timbre matrix can comprehensively and accurately describe the user's voice characteristics, providing a solid data foundation for subsequent speech recognition and personalized processing.
[0105] S702: Associate each user voiceprint feature in the voiceprint library with its corresponding timbre matrix.
[0106] Finally, associate the voiceprint feature of each user in the voiceprint library with the timbre matrix obtained through feature extraction and modeling, and establish a mapping relationship between the two. This association not only enriches the expression form of user voiceprint information but also facilitates subsequent multi-dimensional analysis and precise matching based on the timbre matrix, thereby improving the effects of speech recognition and user differentiation.
[0107] Steps S601 - S602, by extracting and modeling timbre features for each user voiceprint feature in the voiceprint library, generate an exclusive timbre matrix and associate it with the corresponding voiceprint feature. This can not only accurately capture and represent the unique voice characteristics of each user but also support speech synthesis or audio output based on the user's own timbre features. This personalized timbre expression method improves the naturalness of speech interaction and the user experience.
[0108] In a possible implementation, step S204 performs sequential splicing on the language tokens in all output items corresponding to the same speaking user based on the time markers corresponding to each output item to obtain the target data of this speaking user, including: For the speech tokens in all output items corresponding to the same speaking user, based on the timbre matrix associated with this speaking user, perform sequential splicing and acoustic parameter adjustment processing on the speech text tokens to generate the converted audio of this speaking user; for the text tokens in all output items corresponding to the same speaking user, perform sequential splicing on all text tokens to generate the converted text of this speaking user.
[0109] Specifically, for the speech tokens in all output items corresponding to the same speaking user, based on the timbre matrix associated with this user, first perform sequential splicing on the speech text tokens, and then perform adjustment processing in combination with acoustic parameters to generate a converted audio that conforms to the user's voice characteristics and is natural and fluent; at the same time, for the text tokens in all output items corresponding to the same speaking user, splice them in order to form a complete converted text. This not only ensures the temporal coherence of the audio and text but also fully reflects the user's personalized voice characteristics, realizing accurate and consistent multi-modal speech information output.
[0110] In a possible implementation, the server includes multiple multimodal conversion models, and each multimodal conversion model corresponds to and is dedicated to processing a single and specific output language.
[0111] Specifically, multiple multimodal conversion models are deployed in the server, and each model is specifically optimized and processed for a single and specific output language to ensure high-quality voice and text conversion in that language. Through this dedicated design, the conversion requirements in different language environments can be accurately met, thereby improving the conversion accuracy of the server and the user experience.
[0112] In a possible implementation, step S202 uses a pre-constructed multimodal conversion model to perform multimodal parsing and translation processing on the environmental audio, obtaining multiple output items, including: The environmental audio is input into a multimodal conversion model whose output language type is consistent with the expected receiving language through a dynamic translation matrix.
[0113] The server can combine the expected receiving language selected by the user and use a dynamic translation matrix to match the input environmental audio with the output language of the multimodal conversion model, thereby accessing a multimodal conversion model consistent with the expected receiving language for processing. This dynamic translation matrix contains the mapping relationship between the expected receiving language and the output languages of each multimodal conversion model, and can dynamically determine and switch to the corresponding output language conversion model according to the expected receiving language selected by different users, realizing flexible, efficient and accurate cross-language voice conversion.
[0114] Among them, the expected receiving language is the receiving language preset by the user through the smart glasses.
[0115] In a possible implementation, the construction process of the dynamic translation matrix includes: A1: Collect the input and output language sets: Input language set (Source Languages): Use a multi-task language recognition model to detect all languages currently in use in the current session in real time.
[0116] Output language set (Target Languages): Based on user preferences (such as the receiving language selected by the smart glasses), system default settings or business rules, determine the list of languages to be output.
[0117] A2: Initialize the matrix structure: Create a two-dimensional matrix, where the rows represent the input languages and the columns represent the output languages.
[0118] The initial value of the matrix elements is usually 0, indicating that the corresponding translation path has not been activated.
[0119] A3: Fill the matrix elements: For each input language i and output language j: If i == j, translation is usually not required, and the element can be set to 0 or represented by a special marker to indicate skipping. If translation from i to j needs to be supported, the element is set to 1 (indicating that this translation path is activated).
[0120] In a possible implementation, this application provides a method for updating a dynamic translation matrix. Refer to Figure 8 , Figure 8 which is a flowchart of a method for updating a dynamic translation matrix provided by an embodiment of this application, and can be specifically implemented through steps S801 - S802: S801: If a new expected receiving language is recognized, determine the correspondence between the new expected receiving language and the output language of the multimodal conversion model.
[0121] If a new expected receiving language is recognized, the server will determine the correspondence between the new expected receiving language and the output language of the multimodal conversion model. Specifically, the server will look for an output language that is the same as the new expected receiving language and determine the corresponding multimodal conversion model. That is to say, when the server detects a new target language (i.e., the new expected receiving language), it will check whether there is an output language in the existing multimodal conversion models that matches the new language. In this way, the server can ensure that the input data can be correctly converted into the new expected receiving language, thus supporting a wider range of language conversion requirements.
[0122] Exemplarily, if French is newly added as an expected receiving language, the server will check whether there is a multimodal conversion model in the existing models that supports French output and establish the corresponding mapping relationship, so as to ensure that French can be correctly processed and converted.
[0123] S802: Update the dynamic translation matrix based on the new expected receiving language, the correspondence between the new expected receiving language and the output language of the multimodal conversion model, and the preset language mapping rules.
[0124] If a new expected receiving language is recognized, the server will determine the correspondence between the new expected receiving language and the output language of the multimodal conversion model, and update the dynamic translation matrix based on this information and the preset language mapping rules. Specifically, the server will look for the multimodal conversion model corresponding to the output language that is the same as the new expected receiving language, and integrate this information into the dynamic translation matrix according to the preset language mapping rules. In this way, the dynamic translation matrix can timely reflect the new language and its mapping relationship with the existing models, ensuring that the server can correctly handle new language conversion requirements.
[0125] Through steps S701 - S702, it is possible to flexibly adapt to newly added expected receiving languages and dynamically update the translation matrix to ensure that the multi - modal conversion model can accurately handle conversion tasks between different languages.
[0126] In a possible implementation, the method further includes: if it is recognized that a user is set as an ignored user through the smart glasses, the output item of this user is excluded from the subsequent processing flow.
[0127] If it is recognized that a user has set an ignored user through the smart glasses, the output item of this user will be excluded from the subsequent processing flow. Specifically, when the server detects that a user has set a certain speaking user as an ignored user in the smart glasses, it will remove the output item of this user from all subsequent processing steps, ensuring that the audio and text content of this user will not be further processed or displayed. In this way, the server can filter out unwanted content according to the user's preferences and improve the user experience.
[0128] In a possible implementation, the sending the converted audio and converted text corresponding to each speaking user to the smart glasses includes: Sending the converted audio and converted text of all remaining speaking users excluding the ignored user to the smart glasses.
[0129] The server sends the converted audio and converted text of all remaining speaking users excluding the ignored user to the smart glasses. Specifically, once the server recognizes and excludes the output item of the ignored user set by the user, it will collect the converted audio and converted text of all remaining speaking users and send them to the smart glasses for processing and display. In this way, the smart glasses will only receive the audio and text content of non - ignored users, ensuring that users can focus on the speeches they are interested in.
[0130] By allowing users to actively ignore specified objects, personalized filtering and management of voice information output are achieved. At the same time, by flexibly processing and displaying the audio and text content of different speaking users, the user experience and the satisfaction of personalized needs can be improved.
[0131] Based on the voice information processing method for smart glasses provided in the above - mentioned method embodiment, the present application embodiment also provides a voice information processing device for smart glasses. The device will be described below with reference to the accompanying drawings.
[0132] See Figure 9 as shown Figure 9 is a schematic diagram of a voice information processing device for smart glasses provided by the present application embodiment. As Figure 9 shown, the voice information processing device for smart glasses includes: A receiving unit 901, configured to receive the ambient audio collected by the smart glasses and marked with a recording time period; An output item obtaining unit 902, configured to perform multimodal parsing and translation processing on the ambient audio by using a pre-constructed multimodal conversion model to obtain a plurality of output items; each output item includes a language token and a timbre vector corresponding to the language token; the language token is marked with a time stamp, and the time stamp corresponds to the recording time period; A first speaker confirmation unit 903, configured to retrieve in a timbre library based on the included timbre vector for each output item to confirm the speaker corresponding to the output item; A conversion content generating unit 904, configured to perform sequential splicing on the language tokens in all output items corresponding to the same speaker based on the time stamps corresponding to the respective output items to obtain the target data of the speaker; the target data is marked with a speaking time, and the speaking time corresponds to the time stamp of the first language token in the target data to indicate the start time of the target data; An output display unit 905, configured to output and display the target data corresponding to each speaker through the smart glasses.
[0133] In a possible implementation manner, there are a plurality of multimodal conversion models, and each multimodal conversion model corresponds to and is dedicated to processing a single and specific output language.
[0134] In a possible implementation manner, the output item obtaining unit 902 is specifically configured to: Access the ambient audio to a multimodal conversion model whose output language type is consistent with the expected receiving language through a dynamic translation matrix; the dynamic translation matrix includes a mapping relationship between the expected receiving language and the output language types of the multimodal conversion models; the expected receiving language is the receiving language preset by the user through the smart glasses.
[0135] In a possible implementation manner, the apparatus further includes: A correspondence determining unit, configured to determine the correspondence between the newly added expected receiving language and the output language of the multimodal conversion model if a newly added expected receiving language is recognized; A dynamic translation matrix updating unit, configured to update the dynamic translation matrix based on the newly added expected receiving language, the correspondence between the newly added expected receiving language and the output language of the multimodal conversion model, and a preset language mapping rule.
[0136] In a possible implementation manner, the timbre library includes a plurality of user voiceprint record entries; each user voiceprint record entry includes a user identification ID, a user voiceprint feature, and a user profile.
[0137] In a possible implementation, the first speaker confirmation unit 903 specifically includes: A similarity calculation unit, configured to calculate the similarity between each user voiceprint feature in the voiceprint library and the voiceprint vector for each output item to obtain a plurality of similarities; A target voiceprint feature determination unit, configured to identify the similarities greater than the similarity threshold among the plurality of similarities, and determine the user voiceprint feature corresponding to the similarity greater than the similarity threshold as the target voiceprint feature; A speaker determination unit, configured to determine the speaker corresponding to the output item based on the user profile associated with the target voiceprint feature.
[0138] In a possible implementation, the device further includes: A prompt unit, configured to prompt, through the smart glasses, a user who has not entered a user voiceprint record entry in the current scene to set a user profile, and prompt the user who has not entered a user voiceprint record entry to issue a registration voice that conforms to the text template; An identification and assignment unit, configured to identify the user voiceprint feature of the registration voice and assign a user ID to the user voiceprint feature; An association storage unit, configured to associate the newly identified user voiceprint feature, the corresponding newly set user profile, and the newly assigned user ID to obtain a new user voiceprint record entry and store it in the voiceprint library; An execution unit, configured to execute again the step of identifying the similarities greater than the similarity threshold among the plurality of similarities. If there are still no similarities greater than the similarity threshold, repeat the step of prompting the user to set a user profile through the smart glasses and issue a registration voice that conforms to the text template until a similarity greater than the similarity threshold is found.
[0139] In a possible implementation, the device further includes: A voiceprint matrix construction unit, configured to perform feature extraction and modeling on the voiceprint features of each user voiceprint feature in the voiceprint library to obtain a voiceprint matrix; An association unit, configured to associate each user voiceprint feature in the voiceprint library with its corresponding voiceprint matrix.
[0140] In a possible implementation, the conversion content generation unit 904 is specifically configured to: For the speech tokens in all output items corresponding to the same speaking user, based on the timbre matrix associated with the speaking user, perform sequential splicing and acoustic parameter adjustment processing on the speech text tokens to generate a converted audio of the speaking user in the expected receiving language; for the text tokens in all output items corresponding to the same speaking user, perform sequential splicing on all text tokens to generate a converted text of the speaking user in the expected receiving language.
[0141] In a possible implementation, the multimodal conversion model includes a multimodal encoding unit and a multimodal decoding unit; the multimodal decoding unit includes multiple decoding units with multimodal cross-language understanding functions; the multiple decoding units are obtained by jointly and gradually performing speech synthesis training, language translation training, and cross-language conversion training by multiple different basic decoding units; the multimodal encoding unit is obtained by gradually performing text-audio alignment training, timbre alignment training, and language alignment training by the basic encoding unit.
[0142] In a possible implementation, the output item acquisition unit 902 is specifically configured to: Obtain the output language information in the language decoding unit and the converted timbre information in the timbre decoding unit, and use the multimodal encoding unit to perform encoding processing on the environmental audio to obtain the vector representation of the environmental audio; Use the timbre decoding unit to perform timbre feature decoding processing on the vector representation to obtain multiple timbre vectors; Use the speech decoding unit to perform speech decoding processing on the vector representation based on the converted timbre information and the output language information to obtain the speech tokens in the language tokens; Use the text decoding unit to perform text decoding processing on the vector representation based on the converted timbre information and the output language information to obtain the text tokens in the language tokens; For each speech token, combine the speech token, the text token corresponding to the speech token, and the timbre vector corresponding to the speech token based on its time stamp to obtain an output item.
[0143] In addition, an embodiment of the present application further provides a smart glasses (as Figure 10 shown), the smart glasses include an information processing system for identifying user speech and user manual configuration information (such as personal profile information manually output by the user and the conversion language selected by the user), and the information processing system includes: a display module, a transmission module, sensors, and a control module; the display module includes smart display lenses; the sensors include a microphone and a speaker module; the control module includes the computing unit and the user interaction control unit of the smart glasses The transmission module is configured to send environmental audio, registration speech, and user profiles, and receive converted audio and converted text; The intelligent display lens is used to present the converted text and interactive information; The microphone is used to collect ambient audio and user voice input; The speaker module is used to play the converted audio and prompt sounds; The computing unit is used to process the data collected by the sensors; The user interaction control unit is used to receive user input instructions and user profiles.
[0144] This application uses a multi-modal conversion model to convert the collected ambient audio into speech tokens, text tokens, and timbre vectors with time stamps, so as to achieve accurate recognition and translation of the speech content of multiple speakers. By retrieving based on the timbre vectors in the timbre library, the specific speaking user corresponding to each output item can be accurately identified, effectively avoiding the problem of speaker identity confusion in traditional devices. At the same time, the time stamps are used to splice the speech and text content of the same user in time sequence, generating continuous and synchronous conversion results to ensure the coherence and time sequence accuracy of the translated content.
[0145] It should be noted that the various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components described as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0146] The above is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for processing voice information for smart glasses, characterized in that, The method includes: Receiving the ambient audio collected by the smart glasses and marked with a recording time period; the ambient audio includes the speech of at least one speaking user; Using a pre-constructed multimodal conversion model to perform multimodal parsing and translation processing on the ambient audio to obtain a plurality of output items; each output item includes a language token and a timbre vector corresponding to the language token; the language token is marked with a time stamp, and the time stamp corresponds to the recording time period; For each output item, retrieving in the timbre library based on the included timbre vector to confirm the speaking user corresponding to the output item; For the language tokens in all output items corresponding to the same speaking user, performing sequential splicing based on the time stamps corresponding to each output item to obtain the target data of the speaking user; the target data is marked with a speaking time, and the speaking time corresponds to the time stamp of the first language token in the target data to indicate the start time of the target data; Outputting and displaying the target data corresponding to each speaking user through the smart glasses.
2. The method according to claim 1, characterized in that There are multiple multimodal conversion models, and each multimodal conversion model corresponds to and is dedicated to processing a single and specific output language; The method further includes: Connecting the ambient audio to a multimodal conversion model with an output language type consistent with the expected receiving language through a dynamic translation matrix; the dynamic translation matrix includes the mapping relationship between the expected receiving language and the output language types of each multimodal conversion model; the expected receiving language is the receiving language preset by the user through the smart glasses.
3. The method according to claim 2, wherein The method further includes: If a new expected receiving language is recognized, determining the corresponding relationship between the new expected receiving language and the output language of the multimodal conversion model; Updating the dynamic translation matrix based on the new expected receiving language, the corresponding relationship between the new expected receiving language and the output language of the multimodal conversion model, and the preset language mapping rules.
4. The method according to claim 1, wherein The timbre library includes multiple user voiceprint record entries; each user voiceprint record entry includes a user identification ID, a user voiceprint feature, and a user profile.
5. The method according to claim 4, wherein The step of, for each output item, retrieving in the timbre library based on the included timbre vector to confirm the speaking user corresponding to the output item includes: For each output item, calculating the similarity between each user voiceprint feature in the timbre library and the timbre vector to obtain a plurality of similarities; Identifying the similarities greater than the similarity threshold among the plurality of similarities, and determining the user voiceprint feature corresponding to the similarity greater than the similarity threshold as the target voiceprint feature; Based on the user profile associated with the target voiceprint feature, determining the speaking user corresponding to the output item.
6. The method according to claim 5, wherein If there is no similarity greater than the similarity threshold among the plurality of similarities, the method further includes: Prompting, through the smart glasses, the user who has not entered a user voiceprint record entry in the current scene to set a user profile, and prompting the user who has not entered a user voiceprint record entry to send a registration voice conforming to the text template; Identifying the user voiceprint feature of the registration voice and assigning a user ID to the user voiceprint feature; Associate the newly recognized user voiceprint features, the newly set corresponding user profile, and the newly assigned user ID to obtain a new user voiceprint record entry and store it in the tone color library; Execute again the step of identifying the similarities greater than the similarity threshold among the multiple similarities. If there is still no similarity greater than the similarity threshold, repeat the step of prompting the user to set the user profile through the smart glasses and sending a registration voice that conforms to the text template until a similarity greater than the similarity threshold is found.
7. The method according to claim 4, wherein The method further includes: For each user voiceprint feature in the tone color library, extract and model the tone color features of the user voiceprint feature to obtain a tone color matrix; Associate each user voiceprint feature in the tone color library with its corresponding tone color matrix; For the speech tokens and text tokens in all output items corresponding to the same speaking user, perform temporal splicing based on their time stamps to generate the converted audio and converted text of the speaking user, including: For the speech tokens in all output items corresponding to the same speaking user, based on the tone color matrix associated with the speaking user, perform temporal splicing and acoustic parameter adjustment processing on the speech text tokens to generate the converted audio of the speaking user; for the text tokens in all output items corresponding to the same speaking user, perform temporal splicing on all text tokens to generate the converted text of the speaking user.
8. The method according to claim 1, characterized in that, The multi-modal conversion model includes a multi-modal encoding unit and a multi-modal decoding unit; the multi-modal decoding unit includes multiple decoding units with multi-modal cross-language understanding functions; the multiple decoding units are jointly obtained by gradually performing speech synthesis training, language translation training, and cross-language conversion training on multiple different basic decoding units; The multi-modal encoding unit is obtained by gradually performing text-audio alignment training, tone color alignment training, and language alignment training through a basic encoding unit; Using the pre-constructed multi-modal conversion model, perform multi-modal parsing and translation processing on the environmental audio to obtain multiple output items, including: Obtain the output language information in the language decoding unit and the converted tone color information in the tone color decoding unit, and use the multi-modal encoding unit to perform encoding processing on the environmental audio to obtain the vector representation of the environmental audio; Use the tone color decoding unit to perform tone color feature decoding processing on the vector representation to obtain multiple tone color vectors; Use the speech decoding unit to perform speech decoding processing on the vector representation based on the converted tone color information and the output language information to obtain the speech tokens in the language tokens; Use the text decoding unit to perform text decoding processing on the vector representation based on the converted tone color information and the output language information to obtain the text tokens in the language tokens; For each speech token, combine the speech token, the text token corresponding to the speech token, and the tone color vector corresponding to the speech token based on its time stamp to obtain an output item.
9. A voice information processing device for smart glasses, characterized in that, The device includes: A receiving unit for receiving the environmental audio marked with a recording time period collected by the smart glasses; the environmental audio includes the speech of at least one speaking user; An output item acquisition unit, configured to use a pre-constructed multimodal conversion model to perform multimodal parsing and translation processing on the environmental audio to obtain a plurality of output items; each output item includes a language token and a timbre vector corresponding to the language token; the language token is marked with a time stamp, and the time stamp corresponds to the recording time period; A second speaker confirmation unit, configured to, for each output item, retrieve in a timbre library based on the included timbre vector to confirm the speaker corresponding to the output item; A conversion content generation unit, configured to, for the language tokens in all output items corresponding to the same speaker, perform sequential splicing based on the time stamps corresponding to each output item to obtain the target data of the speaker; the target data is marked with a speaking time, and the speaking time corresponds to the time stamp of the first language token in the target data to indicate the start time of the target data; An output display unit, configured to output and display the target data corresponding to each speaker through the smart glasses.
10. An intelligent glasses, characterized in that, The smart glasses include an information processing system for identifying user speech and user manual configuration information, and the information processing system includes: a display module, a transmission module, sensors, and a control module; the display module includes smart display lenses; the sensors include a microphone and a speaker module; the control module includes a computing unit and a user interaction control unit of the smart glasses The transmission module is configured to send environmental audio, registered speech, and user profiles, and receive converted audio and converted text; The smart display lenses are configured to present converted text and interaction information; The microphone is configured to collect environmental audio and user speech input; The speaker module is configured to play converted audio and prompt sounds; The computing unit is configured to process data collected by the sensors; The user interaction control unit is configured to receive user input instructions and user profiles.
Citation Information
Patent Citations
Translation model training method and device
CN110175335A
Translation method and device based on intelligent glasses, and computer readable storage medium
CN110188364A
Translation display method and device, head-mounted display equipment and storage medium
CN111597828A
Voice generation method and device, equipment and computer readable medium
CN111916053A
Voiceprint segmentation method, apparatus and device, and readable storage medium
CN112201275A