Speech translation processing method, electronic device, and readable storage medium

CN116882420BActive Publication Date: 2026-09-11GEER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310769143.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-27
Publication Date
2026-09-11
Estimated Expiration
2043-06-27

AI Technical Summary

Technical Problem

[0004]本申请的主要目的在于提供一种语音翻译处理方法、装置、电子设备及可读存储介质,旨在解决现有技术中XR设备的翻译准确度低的技术问题

Benefits of technology

[0044] This application provides a voice translation processing method. First, the application acquires the voice information of a target user and the biometric information of the target user when making the voice message based on an XR device. Then, based on the voice information, it generates text information to be displayed. Next, based on the biometric information and the voice information, it matches the corresponding target emoji in a preset emoji library for the text information, so as to represent the emotional expression expressed by the target user when making the voice message. Finally, the application synthesizes and displays the text information and the target emoji through the XR device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116882420B_ABST
    Figure CN116882420B_ABST
Patent Text Reader

Abstract

The application discloses a speech translation processing method, an electronic device and a readable storage medium, relates to the technical field of speech processing, and is applied to an XR device. The speech translation processing method comprises the following steps: acquiring speech information of a target user and biological feature information of the target user when the target user makes the speech information based on the XR device; generating text information according to the speech information; matching corresponding target emoticons for the text information in a preset emoticon library according to the biological feature information and the speech information; and synthesizing and displaying the text information and the target emoticons in the XR device. The application solves the problem of low translation accuracy of the XR device in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a speech translation processing method, electronic device, and readable storage medium. Background Technology

[0002] XR (Extended Reality) refers to the use of computers to combine the real and virtual worlds, creating a virtual environment that allows for human-computer interaction, thereby providing users with a sense of immersion through a seamless transition between the virtual and real worlds.

[0003] Currently, XR glasses can perform real-time speech recognition and translate the recognized speech into text in the target language for display. However, simply translating speech into text and displaying it cannot fully express the information conveyed by the speaker. Therefore, the translation accuracy of current XR devices is low. Summary of the Invention

[0004] The main objective of this application is to provide a speech translation processing method, apparatus, electronic device, and readable storage medium, aiming to solve the technical problem of low translation accuracy in existing XR devices.

[0005] To achieve the above objectives, this application provides a speech translation processing method applied to an XR device, the speech translation processing method comprising:

[0006] Based on the XR device, the target user's voice information and the target user's biometric information when making the voice message are obtained;

[0007] Based on the voice information, generate text information;

[0008] Based on the biometric information and the voice information, the corresponding target emoji is matched for the text information in a preset emoji library;

[0009] The text information and the target emoji are synthesized and displayed in the XR device.

[0010] Optionally, the XR device includes an acoustic sensor and an image sensor, and the step of acquiring the target user's voice information and the biometric information of the target user when making the voice statement based on the XR device includes:

[0011] The target user's voice information is collected based on the acoustic sensor;

[0012] The image sensor is used to acquire image information when the target user makes the voice message;

[0013] By extracting facial and limb features from the image information, the biometric information of the target user can be obtained.

[0014] Optionally, the step of matching the text information with a corresponding target emoji from a preset emoji library based on the biometric information and the voice information includes:

[0015] Perform voice emotion recognition on the voice information to determine whether there is voice emotion information in the voice information;

[0016] If the voice information contains the voice emotion information, then search the emoji library for the first set of emojis corresponding to the voice emotion information;

[0017] In the first set of emojis, the emoji with the highest matching degree with the biometric information is selected as the target emoji corresponding to the text information.

[0018] Optionally, the step of performing voice emotion recognition on the voice information to determine whether there is voice emotion information in the voice information includes:

[0019] The speech information is subjected to speech emotion recognition to obtain speech emotion type and text emotion type, wherein the speech emotion type refers to the emotion type represented by the degree of fluctuation of speech, and the text emotion type refers to the emotion type represented by the text semantics of speech.

[0020] Based on the similarity between the voice emotion type and the text emotion type, it is determined whether there is voice emotion information in the voice information.

[0021] Optionally, the step of determining whether there is voice emotion information in the voice information based on the similarity between the voice emotion type and the text emotion type includes:

[0022] Verify whether the similarity is greater than a preset similarity threshold;

[0023] If so, then it is determined that the voice emotion information exists in the voice information;

[0024] If not, then it is determined that the voice emotion information is not present in the voice information.

[0025] Optionally, the step of matching the text information with a corresponding target emoji from a preset emoji library based on the biometric information and the voice information includes:

[0026] Non-voice emotion recognition is performed on the biometric information to determine whether non-voice emotion information exists in the biometric information;

[0027] If the non-voice emotion information exists in the biometric information, then search the emoji library for the second set of emojis corresponding to the non-voice emotion information;

[0028] In the second set of emojis, the emoji with the highest matching degree to the voice information is selected as the target emoji corresponding to the text information.

[0029] Optionally, the step of performing non-voice emotion recognition on the biometric information to determine whether non-voice emotion information exists in the biometric information includes:

[0030] Non-voice emotion recognition is performed on the biometric information to obtain facial expression types and the amplitude of body movement changes, wherein the facial expression types include emotional expression types and non-emotional expression types;

[0031] Based on the facial expression type and the amplitude of the body movement changes, determine whether there is non-vocal emotion information in the biometric information.

[0032] Optionally, the step of determining whether non-vocal emotion information exists in the biometric information based on the facial expression type and the amplitude of the body movement change includes:

[0033] Verify whether the facial expression type is the emotional expression type or whether the amplitude of the body movement change is greater than a preset amplitude threshold;

[0034] If the facial expression type is the emotional expression type or the change in body movement is greater than the preset amplitude threshold, then it is determined that the non-vocal emotional information exists in the biometric information;

[0035] If the facial expression type is the non-emotional expression type and the amplitude of the body movement change is less than or equal to the preset amplitude threshold, then it is determined that the non-vocal emotion information does not exist in the biometric information.

[0036] This application also provides a speech translation processing device for use in an XR device, the speech translation processing device comprising:

[0037] The acquisition module is used to acquire the voice information of the target user and the biometric information of the target user when making the voice message based on the XR device.

[0038] The generation module is used to generate text information based on the voice information;

[0039] The matching module is used to match the text information with the corresponding target emoji from a preset emoji library based on the biometric information and the voice information.

[0040] A compositing module is used to compose and display the text information and the target emoji in the XR device.

[0041] This application also provides an electronic device, which is a physical device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the speech translation processing method described above.

[0042] This application also provides a readable storage medium, which is a computer-readable storage medium, storing a program implementing a speech translation processing method. The program implementing the speech translation processing method is executed by a processor to implement the steps of the speech translation processing method described above.

[0043] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the speech translation processing method described above.

[0044] This application provides a voice translation processing method. First, the application acquires the voice information of a target user and the biometric information of the target user when making the voice message based on an XR device. Then, based on the voice information, it generates text information to be displayed. Next, based on the biometric information and the voice information, it matches the corresponding target emoji in a preset emoji library for the text information, so as to represent the emotional expression expressed by the target user when making the voice message. Finally, the application synthesizes and displays the text information and the target emoji through the XR device.

[0045] Therefore, this application enables the XR device to simultaneously include text information and target emoticons in the result of translating user speech, that is, to simultaneously include text information and non-text information, so as to express the speaker's speech content through text information and the speaker's speaking emotions through non-text information, thereby completely and accurately expressing the information conveyed by the speaker, and solving the technical problem of low translation accuracy of XR devices in the prior art. Attached Figure Description

[0046] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a flowchart illustrating an embodiment of the speech translation processing method of this application.

[0049] Figure 2 This is a schematic diagram of part of the emoji library provided in Embodiment 1 of the speech translation processing method of this application;

[0050] Figure 3 This is a flowchart illustrating Embodiment 2 of the speech translation processing method of this application;

[0051] Figure 4 A simplified flowchart of Embodiment 2 of the speech translation processing method of this application;

[0052] Figure 5 This is a schematic diagram of the module structure of the speech translation processing device according to an embodiment of this application;

[0053] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the speech translation processing method in the embodiments of this application.

[0054] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0055] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] Example 1

[0057] XR refers to the use of computers to combine the real and virtual worlds, creating a virtual environment that allows for human-computer interaction, thereby providing users with a sense of immersion through a seamless transition between the virtual and real worlds.

[0058] Currently, XR glasses can perform real-time speech recognition and translate the recognized speech into text in the target language for display. However, simply translating speech into text and displaying it cannot fully express the information conveyed by the speaker. Therefore, the translation accuracy of current XR devices is low.

[0059] Based on this, this application proposes a speech translation processing method according to the first embodiment, applied to an XR device. Please refer to [link / reference]. Figure 1 The speech translation processing method includes:

[0060] Step S10: Based on the XR device, acquire the voice information of the target user and the biometric information of the target user when making the voice message;

[0061] It should be noted that the biometric information refers to the inherent physiological characteristics of the human body, such as facial features, limb features, and behavioral characteristics. The biometric information can include facial feature information and limb feature information. The facial feature information is used to characterize the target user's facial micro-expressions and facial movements, and the limb feature information is used to characterize the target user's limb movements.

[0062] As an example, the step of acquiring the target user's voice information and the biometric information of the target user when making the voice information based on the XR device can be acquiring the target user's voice information within a preset time period and the biometric information of the target user when making the voice information based on the XR device; or it can be acquiring the target user's voice information within a preset time point and the biometric information of the target user when making the voice information based on the XR device. This example does not limit this step.

[0063] Step S20: Generate text information based on the voice information;

[0064] It should be noted that the text information is the text content corresponding to the voice information. When the language involved in the voice information is the same as the language involved in the text information, the recognized voice information can be directly converted into text information. However, when the language involved in the voice information is different from the language involved in the text information, the recognized voice information needs to be converted into text information to be translated first, and then the text information to be translated into the target language to obtain the text information.

[0065] As one example, when the language involved in the speech information is the same as the language involved in the text information, the step of generating text information based on the speech information includes: extracting text from the speech information to obtain the text information; as another example, when the language involved in the speech information is different from the language involved in the text information, the step of generating text information based on the speech information includes: extracting text from the speech information to obtain text information to be translated; translating the text information to be translated into the target language to obtain the text information.

[0066] Step S30: Based on the biometric information and the voice information, match the corresponding target emoji for the text information in a preset emoji library;

[0067] It should be noted that, please refer to Figure 2 , Figure 2 The example demonstrates a portion of the emoji library, which records the mapping relationship between emojis and their meanings. This emoji library can be further categorized into emotion-based emoji libraries, action-based emoji libraries, and shape-based emoji libraries. Different emojis can be set for different emotions, or different emojis can be set for the same facial expression based on different emotion levels. This embodiment does not impose such limitations. For example, assuming the facial expression is a smiley face, if the emotion is "very happy," then the emoji corresponding to the meaning of "laughing" in a smiley face is matched; if the emotion is "happy," then the emoji corresponding to the meaning of "smiling" in a smiley face is matched.

[0068] Additionally, it should be noted that the number of target emojis can be one or more, and this embodiment does not limit this.

[0069] It is understandable that, in order to improve the accuracy of obtaining the target emoji matched by the text information, when matching the target emoji for the text information in a preset emoji library based on the biometric information and the voice information, the biometric information can be used as the primary matching condition and the voice information as the secondary matching condition to match the target emoji for the text information, or the voice information can be used as the primary matching condition and the biometric information as the secondary matching condition to match the target emoji for the text information. This embodiment does not limit this.

[0070] Step S40: The text information and the target emoji are synthesized and displayed in the XR device.

[0071] It should be noted that when synthesizing the text information and the target emoji, the target emoji can be inserted at the end of the text information, at the beginning of the text information, after the emotion word, or after the verb. This embodiment does not limit this.

[0072] Understandably, the text information is used to express the target user's speech content, and the target emoji is used as non-text information to express the target user's speaking emotions. By synthesizing the text information and the target emoji, and displaying the synthesized content as the translation result of the XR device, the information conveyed by the target user can be fully expressed, achieving accurate translation and vivid display of the voice content.

[0073] This application provides a voice translation processing method. First, the method acquires the voice information of a target user and the biometric information of the target user when making the voice message using an XR device. Then, based on the voice information, text information to be displayed is generated. Next, based on the biometric information and the voice information, a corresponding target emoji is matched for the text information in a preset emoji library so as to represent the emotional expression expressed by the target user when making the voice message. Finally, the text information and the target emoji are synthesized and displayed using the XR device.

[0074] Therefore, the embodiments of this application enable the XR device to simultaneously include text information and target emoticons in the result of translating user speech, that is, to simultaneously include text information and non-text information, so as to express the speaker's speech content through text information and the speaker's speaking emotions through non-text information, thereby completely and accurately expressing the information conveyed by the speaker, and solving the technical problem of low translation accuracy of XR devices in the prior art.

[0075] In one possible implementation, the XR device includes an acoustic sensor and an image sensor, and the step of acquiring voice information of a target user based on the XR device, as well as biometric information of the target user when making the voice statement, includes:

[0076] Step S11: Collect the voice information of the target user based on the acoustic sensor;

[0077] It should be noted that the acoustic sensor may include a microphone.

[0078] Step S12: Acquire image information of the target user when they make the voice message based on the image sensor;

[0079] It should be noted that the image sensor may include a camera, and the image information may include facial feature information and limb feature information. The facial feature information is used to characterize the facial expressions and facial movements of the target user, and the limb feature information is used to characterize the limb movements of the target user.

[0080] It is understood that when collecting image information of the target user based on the image sensor, image information at a single point in time or over a period of time can be collected to obtain an image video. This embodiment does not limit this.

[0081] Step S13: By extracting facial and limb feature information from the image information, the biometric information of the target user is obtained.

[0082] In this embodiment, the acoustic sensor in the XR device is used to collect the voice information of the target user, and the image sensor in the XR device is used to simultaneously collect the image information of the target user when he / she makes the voice information. Then, by extracting the facial feature information and limb feature information from the image information, the biometric information of the target user can be obtained. Thus, when the XR device performs vivid translation of the voice, no external equipment is needed. The voice translation processing flow of this application can be realized using the existing equipment.

[0083] In one possible implementation, the step of matching a target emoji for the text information in a preset emoji library based on the biometric information and the voice information includes:

[0084] Step A10: Perform voice emotion recognition on the voice information to determine whether there is voice emotion information in the voice information;

[0085] It should be noted that this voice emotion information is used to characterize the speaking emotion expressed in the voice information. If voice emotion recognition can identify a specific emotion type in the voice information, it means that voice emotion information exists in the voice information. If voice emotion recognition cannot identify a specific emotion type in the voice information, it means that voice emotion information does not exist in the voice information.

[0086] Step A20: If the voice information contains the voice emotion information, then search the first set of emoticons corresponding to the voice emotion information in the emoticon library;

[0087] It should be noted that all the emojis in the first set of emojis correspond to the emotional information in the speech. This can be understood as the emotional type of the speech emotional information being the same as the emotional type of the emojis in the first set of emojis. For example, they can all belong to the emotional type of "happy".

[0088] Step A30: Select the emoji with the highest matching degree with the biometric information from the first emoji set as the target emoji corresponding to the text information.

[0089] It should be noted that the matching degree between the biometric information and the emojis in the first emoji set can be the similarity between the facial expression in the biometric information and the expression corresponding to the emoji, or the similarity between the emotional meaning corresponding to the body movement in the biometric information and the emotional meaning corresponding to the emoji. This embodiment does not limit this.

[0090] It is understandable that in this emoji library, the same emotional meaning may exist in different emojis. For example, the emoji library found that the emotional meaning of "happy" includes emoji 1 (mouth not open) and emoji 2 (mouth open). However, the target user's biometric information shows an open mouth facial expression. Therefore, the target emoji that is finally matched for this text information is emoji 2.

[0091] In this embodiment, voice emotion recognition is first performed on the voice information to determine whether there is voice emotion information in the voice information. If the voice emotion information exists in the voice information, the first set of emojis corresponding to the voice emotion information is searched in the emoji library using the emotional meaning corresponding to the voice emotion information as an index. Then, the emoji with the highest matching degree with the biometric information is selected from the first set of emojis as the target emoji corresponding to the text information. Thus, by performing multi-level hierarchical matching in the emotion database, the accuracy of obtaining the target emoji corresponding to the text information is improved.

[0092] In one possible implementation, the step of performing voice emotion recognition on the voice information to determine whether there is voice emotion information in the voice information includes:

[0093] Step A11: Perform voice emotion recognition on the voice information to obtain voice emotion type and text emotion type, wherein the voice emotion type refers to the emotion type represented by the degree of fluctuation of the voice, and the text emotion type refers to the emotion type represented by the text semantics of the voice.

[0094] It should be noted that the voice emotion type can include excited and low-pitched types, with the low-pitched type having a lower fluctuation level than the excited type. The text emotion type can include happy text type, angry text type, sad text type, neutral text type, etc., and this embodiment does not limit it.

[0095] Step A12: Determine whether there is voice emotion information in the voice information based on the similarity between the voice emotion type and the text emotion type.

[0096] Understandably, the similarity between the voice emotion type and the text emotion type is used to characterize whether the voice emotion type and the text emotion type belong to the same emotion type. For example, the excited type belongs to the same emotion type as the happy and angry text types, the depressed type belongs to the same emotion type as the sad text type, and neither the excited type nor the depressed type belongs to the same emotion type as the neutral text type.

[0097] For example, assuming the voice emotion type is excited, then when the text emotion type is happy, the voice information contains voice emotion information of the happy emotion type; when the text emotion type is neutral, the recognized emotion result is no emotion type, and the voice information does not contain voice emotion information; when the text emotion type is sad, the recognized emotion result is no emotion type, and the voice information does not contain voice emotion information.

[0098] Further, the step of determining whether there is voice emotion information in the voice information based on the similarity between the voice emotion type and the text emotion type includes:

[0099] Step A121: Verify whether the similarity is greater than a preset similarity threshold;

[0100] Step A122: If yes, then determine that the voice emotion information exists in the voice information;

[0101] Step A123: If not, then determine that the voice emotion information does not exist in the voice information.

[0102] In this embodiment, voice emotion recognition is first performed on the voice information to obtain a voice emotion type for the degree of fluctuation in the voice and a text emotion type for the text semantics representing the voice. Then, based on the similarity between the voice emotion type and the text emotion type, it is determined whether there is voice emotion information in the voice information. If the similarity between the voice emotion type and the text emotion type is greater than a preset similarity threshold, it means that the voice emotion type and the text emotion type belong to the same emotion type, and it is determined that there is voice emotion information in the voice information. If the similarity between the voice emotion type and the text emotion type is not greater than the preset similarity threshold, it means that the voice emotion type and the text emotion type belong to different emotion types, and it is determined that there is no voice emotion information in the voice information. Thus, by jointly determining whether there is voice emotion information in the voice information through the degree of fluctuation in the voice and the text semantics, the accuracy of voice emotion information determination is improved.

[0103] Example 2

[0104] Based on the first embodiment of this application, in another embodiment of this application, the same or similar content as in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 The step of matching the text information with a corresponding target emoji from a preset emoji library based on the biometric information and the voice information includes:

[0105] Step B10: Perform non-voice emotion recognition on the biometric information to determine whether there is non-voice emotion information in the biometric information;

[0106] It should be noted that this non-vocal emotion information is used to characterize the speaking emotion expressed by the biometric information. If non-vocal emotion recognition of the biometric information can yield a specific emotion type, it indicates that non-vocal emotion information exists in the biometric information. If non-vocal emotion recognition of the biometric information cannot yield a specific emotion type, it indicates that non-vocal emotion information does not exist in the biometric information.

[0107] Step B20: If the non-voice emotion information exists in the biometric information, then search the emoji library for the second set of emojis corresponding to the non-voice emotion information.

[0108] It should be noted that all the emojis in this second set of emojis correspond to the non-vocal emotional information. This can be understood as the facial expression or body movement of the non-vocal emotional information being the same as the expression or movement of the emojis in this second set of emojis. For example, they can all belong to the expression "open mouth".

[0109] Step B30: Select the emoji with the highest matching degree to the voice information from the second emoji set as the target emoji corresponding to the text information.

[0110] It should be noted that the matching degree between the voice information and the emojis in the second emoji set refers to the similarity of the emotional meaning between the emotional meaning corresponding to the voice information and the emotional meaning corresponding to the emoji.

[0111] It is understandable that similar emojis in this emoji library may represent different emotional meanings. For example, the emoji library may contain two types of emojis: emoji 1 (meaning surprise) and emoji 2 (meaning happiness). If the target user's voice information corresponds to the emotion of happiness, then the target emoji to be matched for the text information is emoji 2.

[0112] In this embodiment, non-voice emotion recognition is first performed on the biometric information to determine whether non-voice emotion information exists in the biometric information. If non-voice emotion information exists in the biometric information, the facial expression or body movement corresponding to the non-voice emotion information is used as an index to search for a second set of emojis corresponding to the non-voice emotion information in the emoji library. Then, the emoji with the highest matching degree with the voice information is selected from the second set of emojis as the target emoji corresponding to the text information. Thus, by performing multi-level hierarchical matching in the emotion database, the accuracy of obtaining the target emoji corresponding to the text information is improved.

[0113] In one possible implementation, the step of performing non-voice emotion recognition on the biometric information to determine whether non-voice emotion information exists in the biometric information includes:

[0114] Step B11: Perform non-voice emotion recognition on the biometric information to obtain facial expression types and body movement change amplitudes, wherein the facial expression types include emotional expression types and non-emotional expression types;

[0115] The face is an extremely complex carrier of expressions. Features such as the eyebrows, forehead, and mouth can produce a wide variety of movements, which, when combined, express different emotions. It is estimated that the human face can display over ten thousand different expressions, many of which, such as happiness, sadness, anger, and fear, are universal across cultures.

[0116] It should be noted that the greater the range of changes in body language, the greater the fluctuation in the speaker's emotions. For example, if the speaker's facial expression is smiling, the greater the range of changes in body language, the higher the speaker's level of happiness.

[0117] When performing non-voice emotion recognition on the biometric information, specific body movements can also be identified. For example, when body movements with emotional meaning, such as "raising a thumb" or "blinking," are identified in the biometric information, it is also determined that there is non-voice emotion information in the biometric information.

[0118] Step B12: Based on the facial expression type and the amplitude of the body movement changes, determine whether there is non-vocal emotion information in the biometric information.

[0119] Furthermore, the step of determining whether non-verbal emotional information exists in the biometric information based on the facial expression type and the amplitude of the body movement changes includes:

[0120] Step B121: Verify whether the facial expression type is the emotional expression type or whether the amplitude of the body movement change is greater than a preset amplitude threshold.

[0121] Step B122: If the facial expression type is the emotional expression type or the change in body movement is greater than the preset amplitude threshold, then it is determined that the non-vocal emotional information exists in the biometric information.

[0122] Step B123: If the facial expression type is the non-emotional expression type and the amplitude of the body movement change is less than or equal to the preset amplitude threshold, then it is determined that the non-vocal emotion information does not exist in the biometric information.

[0123] In this embodiment, non-voice emotion recognition is first performed on the biometric information to obtain the facial expression type and the amplitude of body movement changes. Then, based on the facial expression type and the amplitude of body movement changes, it is determined whether there is non-voice emotion information in the biometric information. If the facial expression type is an emotional expression type, or the amplitude of body movement changes is greater than a preset amplitude threshold, it is determined that there is non-voice emotion information in the biometric information. If the facial expression type is a non-emotional expression type, and the amplitude of body movement changes is not greater than the preset amplitude threshold, it is determined that there is no non-voice emotion information in the biometric information. Thus, by jointly determining whether there is non-voice emotion information in the biometric information through facial expression type and the amplitude of body movement changes, the accuracy of non-voice emotion information determination is improved.

[0124] For example, to aid in understanding the technical concept or principles of this application, please refer to Figure 4 , Figure 4 A simplified flowchart of speech translation processing is provided below:

[0125] 1. Acquire the speaker's (target user's) facial expressions, facial movements, and body movements (the aforementioned biometric information) through a camera (the aforementioned image sensor);

[0126] 2. Acquire the speaker's audio signal (the aforementioned voice information) through a MIC (Microphone);

[0127] 3. Convert the audio signal into text in the target language (the aforementioned text information);

[0128] 4. Determine the speaker's emotional type (the above-mentioned voice emotion type) and body movements (the above-mentioned non-voice emotion type) by facial expressions, facial movements and audio signals;

[0129] 5. Based on the recognition results, retrieve the emoji with the highest matching degree from the emoji library of the target type;

[0130] 6. Add the selected emojis (may be multiple) to the appropriate positions in the translation result (the above text information);

[0131] 7. Output the translation results with added emojis to the XR display screen.

[0132] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the credit risk identification method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0133] Example 3

[0134] This invention also provides a voice translation processing device for use in XR devices. Please refer to [link / reference]. Figure 5 The speech translation processing device includes:

[0135] The acquisition module 10 is used to acquire the voice information of the target user and the biometric information of the target user when making the voice message based on the XR device.

[0136] The generation module 20 is used to generate text information based on the voice information;

[0137] The matching module 30 is used to match the text information with the corresponding target emoji from a preset emoji library based on the biometric information and the voice information.

[0138] The compositing module 40 is used to compose and display the text information and the target emoji in the XR device.

[0139] Optionally, the XR device includes an acoustic sensor and an image sensor, and the acquisition module 10 is further configured to:

[0140] The target user's voice information is collected based on the acoustic sensor;

[0141] The image sensor is used to acquire image information when the target user makes the voice message;

[0142] By extracting facial and limb features from the image information, the biometric information of the target user can be obtained.

[0143] Optionally, the matching module 30 is further configured to:

[0144] Perform voice emotion recognition on the voice information to determine whether there is voice emotion information in the voice information;

[0145] If the voice information contains the voice emotion information, then search the emoji library for the first set of emojis corresponding to the voice emotion information;

[0146] In the first set of emojis, the emoji with the highest matching degree with the biometric information is selected as the target emoji corresponding to the text information.

[0147] Optionally, the matching module 30 is further configured to:

[0148] The speech information is subjected to speech emotion recognition to obtain speech emotion type and text emotion type, wherein the speech emotion type refers to the emotion type represented by the degree of fluctuation of speech, and the text emotion type refers to the emotion type represented by the text semantics of speech.

[0149] Based on the similarity between the voice emotion type and the text emotion type, it is determined whether there is voice emotion information in the voice information.

[0150] Optionally, the matching module 30 is further configured to:

[0151] Verify whether the similarity is greater than a preset similarity threshold;

[0152] If so, then it is determined that the voice emotion information exists in the voice information;

[0153] If not, then it is determined that the voice emotion information is not present in the voice information.

[0154] Optionally, the matching module 30 is further configured to:

[0155] Non-voice emotion recognition is performed on the biometric information to determine whether non-voice emotion information exists in the biometric information;

[0156] If the non-voice emotion information exists in the biometric information, then search the emoji library for the second set of emojis corresponding to the non-voice emotion information;

[0157] In the second set of emojis, the emoji with the highest matching degree to the voice information is selected as the target emoji corresponding to the text information.

[0158] Optionally, the matching module 30 is further configured to:

[0159] Non-voice emotion recognition is performed on the biometric information to obtain facial expression types and the amplitude of body movement changes, wherein the facial expression types include emotional expression types and non-emotional expression types;

[0160] Based on the facial expression type and the amplitude of the body movement changes, determine whether there is non-vocal emotion information in the biometric information.

[0161] Optionally, the matching module 30 is further configured to:

[0162] Verify whether the facial expression type is the emotional expression type or whether the amplitude of the body movement change is greater than a preset amplitude threshold;

[0163] If the facial expression type is the emotional expression type or the change in body movement is greater than the preset amplitude threshold, then it is determined that the non-vocal emotional information exists in the biometric information;

[0164] If the facial expression type is the non-emotional expression type and the amplitude of the body movement change is less than or equal to the preset amplitude threshold, then it is determined that the non-vocal emotion information does not exist in the biometric information.

[0165] The speech translation processing device provided by this invention, employing the speech translation processing method in Embodiment 1 or Embodiment 2 described above, can solve the technical problem of low translation accuracy in XR devices in the prior art. Compared with the prior art, the beneficial effects of the speech translation processing device provided by this invention are the same as those of the speech translation processing method provided in the above embodiments, and other technical features in the speech translation processing device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0166] Example 4

[0167] This invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the speech translation processing method described in Embodiment 1 above.

[0168] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0169] like Figure 6 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. While electronic devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0170] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.

[0171] The electronic device provided by this invention, employing the speech translation processing method in the above embodiments, can solve the technical problem of low translation accuracy in existing XR devices. Compared with the prior art, the beneficial effects of the electronic device provided by this invention are the same as those of the speech translation processing method provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0172] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.

[0173] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0174] Example 5

[0175] This invention provides a computer-readable storage medium having computer-readable program instructions stored thereon, which are used to execute the speech translation processing method in Embodiment 1 above.

[0176] The computer-readable storage medium provided in this embodiment of the invention may be, for example, a USB flash drive, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0177] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0178] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: acquire voice information of a target user and biometric information of the target user when the voice information is made, based on the XR device; generate text information based on the voice information; match a corresponding target emoji in a preset emoji library for the text information based on the biometric information and the voice information; and synthesize and display the text information and the target emoji in the XR device.

[0179] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0180] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0181] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0182] The readable storage medium provided by this invention is a computer-readable storage medium that stores computer-readable program instructions for executing the above-described speech translation processing method, thereby solving the technical problem of low translation accuracy in XR devices in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this embodiment are the same as those of the speech translation processing method provided in Embodiment 1 or Embodiment 2, and will not be repeated here.

[0183] Example 6

[0184] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the speech translation processing method described above.

[0185] The computer program product provided in this application can solve the technical problem of low translation accuracy of XR devices in the prior art. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of the present invention are the same as the beneficial effects of the speech translation processing method provided in Embodiment 1 or Embodiment 2 above, and will not be repeated here.

[0186] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.

Claims

1. A speech translation processing method characterized by, The speech translation processing method, applied to XR devices, includes: Based on the XR device, the target user's voice information and the target user's biometric information when making the voice message are obtained; Based on the voice information, generate text information; Based on the biometric information and the voice information, the corresponding target emoji is matched for the text information in a preset emoji library; The text information and the target emoji are synthesized and displayed in the XR device; The step of matching the text information with a corresponding target emoji from a preset emoji library based on the biometric information and the voice information includes: Perform voice emotion recognition on the voice information to determine whether there is voice emotion information in the voice information; If the voice information contains the voice emotion information, then search the first set of emojis corresponding to the voice emotion information in the emoji library; In the first set of emojis, the emoji with the highest matching degree to the biometric information is selected as the target emoji corresponding to the text information; The step of performing voice emotion recognition on the voice information to determine whether there is voice emotion information in the voice information includes: The speech information is subjected to speech emotion recognition to obtain speech emotion type and text emotion type, wherein the speech emotion type refers to the emotion type represented by the degree of fluctuation of speech, and the text emotion type refers to the emotion type represented by the text semantics of speech. Based on the similarity between the voice emotion type and the text emotion type, it is determined whether there is voice emotion information in the voice information.

2. The speech translation processing method as described in claim 1, characterized in that, The XR device includes an acoustic sensor and an image sensor. The step of acquiring the target user's voice information and the biometric information of the target user when making the voice statement based on the XR device includes: The target user's voice information is collected based on the acoustic sensor; The image sensor is used to acquire image information when the target user makes the voice message; By extracting facial and limb features from the image information, the biometric information of the target user can be obtained.

3. The speech translation processing method as described in claim 1, characterized in that, The step of determining whether there is voice emotion information in the voice information based on the similarity between the voice emotion type and the text emotion type includes: Verify whether the similarity is greater than a preset similarity threshold; If so, then it is determined that the voice emotion information exists in the voice information; If not, then it is determined that the voice emotion information is not present in the voice information.

4. The speech translation processing method as described in claim 1, characterized in that, The step of matching the text information with a corresponding target emoji from a preset emoji library based on the biometric information and the voice information includes: Non-voice emotion recognition is performed on the biometric information to determine whether non-voice emotion information exists in the biometric information; If the non-voice emotion information exists in the biometric information, then search the emoji library for the second set of emojis corresponding to the non-voice emotion information; In the second set of emojis, the emoji with the highest matching degree to the voice information is selected as the target emoji corresponding to the text information.

5. The speech translation processing method as described in claim 4, characterized in that, The step of performing non-voice emotion recognition on the biometric information to determine whether non-voice emotion information exists in the biometric information includes: Non-voice emotion recognition is performed on the biometric information to obtain facial expression types and the amplitude of body movement changes, wherein the facial expression types include emotional expression types and non-emotional expression types; Based on the facial expression type and the amplitude of the body movement changes, determine whether there is non-vocal emotion information in the biometric information.

6. The speech translation processing method as described in claim 5, characterized in that, The step of determining whether non-vocal emotion information exists in the biometric information based on the facial expression type and the amplitude of the body movement changes includes: Verify whether the facial expression type is the emotional expression type or whether the amplitude of the body movement change is greater than a preset amplitude threshold; If the facial expression type is the emotional expression type or the change in body movement is greater than the preset amplitude threshold, then it is determined that the non-vocal emotional information exists in the biometric information; If the facial expression type is the non-emotional expression type and the amplitude of the body movement change is less than or equal to the preset amplitude threshold, then it is determined that the non-vocal emotion information does not exist in the biometric information.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps of the speech translation processing method as described in any one of claims 1 to 6.

8. A readable storage medium, characterized in that, The readable storage medium is a computer-readable storage medium, on which a program implementing the speech translation processing method is stored. The program implementing the speech translation processing method is executed by a processor to implement the steps of the speech translation processing method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent voice conversion system based on internet technology

    CN109949794A

  • Speech synthesis method and device, electronic equipment and storage medium

    CN114678003A