Speech processing method and apparatus, electronic device, and storage medium
By displaying real-time audio captions and providing AI translation capabilities in multilingual conferences, the problem of low communication efficiency in multilingual conference scenarios has been solved, achieving efficient audio translation and simultaneous interpretation, thereby improving the communication efficiency and user experience of conference participants.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2025-11-17
- Publication Date
- 2026-05-28
AI Technical Summary
In multilingual conference scenarios, communication efficiency among participants is low, requiring translators or translation machines to translate sentence by sentence, resulting in inefficiency.
After receiving the voice messages from participants, the system displays voice subtitles in real time, including translated text or voice text. It uses artificial intelligence technology to recognize and translate voice messages in multiple languages, and displays the voice recognition function interface in a floating window in the meeting interface. It supports object identification and translation style adjustment, and provides simultaneous interpretation and voice broadcast options.
It improves communication efficiency among participants speaking different languages, allowing users to view translated text intuitively without frequently switching interfaces. It supports diverse translation styles and voice broadcasting methods, ensuring the accuracy and flexibility of simultaneous interpretation.
Smart Images

Figure CN2025135297_28052026_PF_FP_ABST
Abstract
Description
Voice processing methods, devices, electronic devices and storage media
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411688519.3, filed on November 22, 2024, entitled “Speech Processing Method, Apparatus, Electronic Device and Storage Medium”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of communication technology, and in particular to a voice processing method, apparatus, electronic device, and storage medium. Background Technology
[0004] With the popularization of mobile internet and smart devices, users' demand for functions such as voice interaction, language translation, and real-time information processing is increasing in their daily work and life.
[0005] Currently, in meeting scenarios, if there are participants speaking different languages, each participant needs to rely on translators or translation machines to translate sentence by sentence to communicate, resulting in low communication efficiency between participants. Summary of the Invention
[0006] The purpose of this application is to provide a voice processing method, apparatus, electronic device, and storage medium to improve communication efficiency among participants speaking different languages.
[0007] In a first aspect, embodiments of this application provide a voice processing method, the method comprising:
[0008] Upon receiving voice messages from participants, display audio subtitles based on those messages.
[0009] The audio subtitles include either the audio-translated text or both the audio text and the audio-translated text.
[0010] Secondly, embodiments of this application provide a voice processing apparatus, the apparatus comprising:
[0011] The display module is used to display audio subtitles based on the voice of the participants when the audio of the participants is received.
[0012] The audio subtitles include either the audio-translated text or both the audio text and the audio-translated text.
[0013] Thirdly, embodiments of this application also provide an electronic device, the device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect.
[0014] Fourthly, embodiments of this application also provide an electronic device configured to perform the wake-up method as described in the first aspect.
[0015] Fifthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0016] In a sixth aspect, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0017] In a seventh aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0018] In this embodiment, when meeting participants speak different languages, and upon receiving their voice messages, the system can intuitively display the translated text of their voice messages, or display both the translated text and the voice message itself. This allows users to easily view the translated text or the translated text and the voice message, and further enables them to communicate with the participants based on the translated text or the translated text and the voice message, thus improving communication efficiency between participants speaking different languages. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 is a flowchart illustrating a speech recognition method provided in some embodiments of this application;
[0021] Figure 2 is a schematic diagram of the voice recognition function interface provided in some embodiments of this application;
[0022] Figure 3 is a schematic diagram of the voice recognition function interface provided in some embodiments of this application;
[0023] Figure 4 is a schematic diagram of the voice recognition function interface provided in some embodiments of this application;
[0024] Figure 5 is a schematic diagram of the voice recognition function interface provided in some embodiments of this application;
[0025] Figure 6 is a schematic diagram of the voice recognition function interface provided in some embodiments of this application;
[0026] Figure 7 is a schematic diagram of the voice recognition function interface provided in some embodiments of this application;
[0027] Figure 8 is a schematic diagram of the voice recognition function interface provided in some embodiments of this application;
[0028] Figure 9 is a schematic diagram of the voice recognition function interface provided in some embodiments of this application;
[0029] Figure 10 is a schematic diagram of the voice recognition function interface provided in some embodiments of this application;
[0030] Figure 11 is a schematic diagram of the voice recognition function interface provided in some embodiments of this application;
[0031] Figure 12 is a schematic diagram of the structure of a voice processing device shown in some embodiments of this application;
[0032] Figure 13 is a schematic diagram of the structure of an electronic device shown in some embodiments of this application;
[0033] Figure 14 is a schematic diagram of the hardware structure of an electronic device shown in some embodiments of this application. Detailed Implementation
[0034] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0035] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0036] As described in the background section, in existing solutions, communication efficiency is low when there are participants speaking different languages in a meeting setting. To address this issue, this application provides a voice processing method, apparatus, electronic device, and storage medium. When participants are speaking different languages, and the system receives their voice messages, it can visually display the translated text of their voice, or display both the translated text and the voice message itself. This allows users to easily view the translated text and the voice message, and then communicate with the participants based on these translations, thereby improving communication efficiency between participants speaking different languages.
[0037] The technical solution of this application embodiment can be applied to scenarios where the voice information of each person is recognized, translated, and the voice subtitles are displayed in multi-person, multilingual communication scenarios. This application embodiment uses a multi-person, multilingual conference communication scenario as an example for illustration; however, the solution of this application embodiment is not limited to multi-person, multilingual conference communication scenarios, and can also be used for distance education and cross-border online classrooms, multilingual customer service, international conferences and business negotiations, cross-border project collaboration meetings, multilingual press releases and media interviews, etc.
[0038] The speech processing method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0039] Figure 1 is a flowchart illustrating a voice processing method provided in an embodiment of this application. The execution subject of the voice processing method can be an electronic device, specifically an electronic device held by a user. The electronic device can be, but is not limited to, a personal computer (PC), a smartphone, a tablet computer, or a personal digital assistant (PDA).
[0040] The voice processing method provided in this application embodiment may include: upon receiving the voice of a participant, displaying voice subtitles based on the participant's voice; wherein the voice subtitles include voice-translated text, or voice text and voice-translated text.
[0041] The participants can be those attending the meeting, specifically including users and other members.
[0042] Audio captions can be captions corresponding to the voices of participants. These audio captions can include translated audio text, or both audio text and translated audio text.
[0043] The aforementioned voice translation text can be text information resulting from the translation of a participant's speech. Specifically, it can be text information translating a participant's speech in one language into another language desired by the user. For example, if participant A speaks English, but the user's preferred language is Chinese, then participant A's English speech can be converted into Chinese voice translation text.
[0044] Voice-to-text can be text converted from the voices of meeting participants.
[0045] It should be noted that in step 110, there can be multiple participants, and the participants' voices can be in multiple languages.
[0046] In some embodiments of this application, the voice of the participants may include the first voice of a first object and the second voice of a second object. Here, the first object and the second object may be participants, the first voice may be the voice of the first object, the second voice may be the voice of the second object, and the language of the first voice is the first language, and the language of the second voice is the second language. Here, the first language and the second language are two different languages, such as the first language being English and the second language being Chinese.
[0047] During the process of receiving the voices of the participants, the voices of each participant can be recognized separately. Specifically, the voices of each participant can be converted into text to obtain the corresponding voice text. Alternatively, the voices of each participant can be translated into another language to obtain the corresponding voice translation text. The voice translation text and / or voice text corresponding to each participant's voice can then be displayed.
[0048] In one example, referring to FIG. 2, taking the language corresponding to the speech translation text as Chinese for example, the speech information of participant A in the meeting is "When does the meeting start?", and the speech information of participant B in the meeting is "How’s the project progress?". Then, the speech information of participant A, "When does the meeting start?", can be translated into Chinese "会议什么时候开始?", and the speech information of participant B, "How’s the project progress?", can be translated into Chinese "项目进展如何?", and displayed.
[0049] In some embodiments of the present application, the speech received from the meeting participants can be the speech of the meeting participants received in real time through a speech receiving device, or can be the speech of the meeting participants played from other audio devices received through a speech receiving device in a meeting scenario, such as the speech of the meeting participants received from a recording device or the like. The specific way of obtaining the speech is not limited herein, and any way of obtaining the speech of the meeting participants in a meeting scenario belongs to the protection scope of the embodiments of the present application.
[0050] It should be noted that the speech received from the meeting participants is in the case where the meeting participants enter the meeting scenario. The meeting participants entering the meeting scenario can be that the meeting participants enter the meeting scenario through a meeting link. After the user enters the meeting scenario, the electronic device automatically starts the speech recognition function and displays the speech recognition function interface. The speech recognition function interface can be superimposed on the meeting interface in the form of a floating window. Specifically, the speech recognition function interface can be an interface designed based on artificial intelligence (AI) technology, that is, the speech recognition function interface can be an AI interaction interface.
[0051] The above speech recognition function can be a function developed based on AI technology. The speech recognition function can simultaneously recognize the speech of multiple languages and convert it into the corresponding speech text and / or speech translation text. Thus, compared with the existing speech recognition function with a single function, the speech recognition function provided by the embodiments of the present application can simultaneously recognize the speech of multiple languages, meeting the diverse needs of recognizing the speech of different languages in a multilingual environment.
[0052] In one example, referring to FIG. 3, after the user enters the meeting, the electronic device automatically starts the speech recognition function and displays the AI interaction interface 31. The AI interaction interface 31 is superimposed on the meeting interface in the form of a floating window.
[0053] In some embodiments of this application, displaying audio subtitles based on the voices of meeting participants may specifically include:
[0054] The AI interactive interface displays the audio captions of the attendees.
[0055] In some embodiments of this application, after the AI interactive interface is displayed, the voice captions of the participants can be displayed in the AI interactive interface, as shown in Figure 2.
[0056] In the embodiments of this application, by overlaying the audio subtitles of the participants on the AI interactive interface in the meeting interface, the users can view the audio subtitles of the participants without having to exit the meeting interface and switch to the AI interactive interface. In other words, the users can intuitively view the audio subtitles of the participants without having to switch back and forth between interfaces, thus improving the viewing efficiency of the audio subtitles of the participants.
[0057] In some embodiments of this application, to further improve meeting communication efficiency, the step of displaying audio subtitles based on the voices of meeting participants may specifically include:
[0058] Identify the target identifier of the participant based on their voice.
[0059] Displays participant identifiers and audio captions.
[0060] The object identifier can be an identifier used to uniquely identify the participants.
[0061] In some embodiments of this application, the timbre features of the participants' voices can be analyzed to identify which participant the received voice belongs to, thus obtaining the participant's object identifier. Therefore, when displaying the participants' voice captions, the participants' object identifiers can be displayed along with them. Specifically, the voice captions can be located in the associated area of the participants' object identifiers.
[0062] The aforementioned associated area can be a pre-defined area within the object identifier of the participant, such as 2mm below the object identifier of the participant.
[0063] Continuing to refer to the above example and Figure 2, after translating the voice of attendee A "When does the meeting start?" into Chinese "会议什么时候开始?" and the voice of attendee B "How’s the project progress?" into Chinese "项目进展如何?", they can be displayed. In addition, the timbre characteristics of the voices of attendee A and attendee B can be analyzed. If it is determined that the voice "When does the meeting start?" and the voice "How’s the project progress?" come from different attendees respectively, the object identifiers of the attendees corresponding to the voice "When does the meeting start?" and the voice "How’s the project progress?" can be displayed respectively. For example, the object identifier of the voice "When does the meeting start?" is object A, and the object identifier of the voice "How’s the project progress?" is object B.
[0064] It should be noted that when identifying the attendee corresponding to the voice, only that the attendees corresponding to different voices are different can be recognized, and it is not possible to determine which specific person said which voice. Therefore, when displaying the object identifier in Figure 2, it is only to distinguish that different voices come from different attendees, but it is not possible to determine who the specific attendee is. That is, as shown in Figure 2, only that the voice "When does the meeting start?" and the voice "How’s the project progress?" come from different attendees can be recognized, and different object identifiers are assigned to the voice "When does the meeting start?" and the voice "How’s the project progress?" respectively, but it is not determined who said the voice "When does the meeting start?", such as Zhang San or Li Si, and who said the voice "How’s the project progress?", such as Zhang San or Li Si.
[0065] In another embodiment, when displaying the object identifier of the attendee and the voice caption, if the user also speaks a voice, but it is a voice in the habitual language, the voice spoken by the user can also be translated into the language voice adopted by most other objects, that is, the voice information can be displayed bilingually in the electronic device.
[0066] Referring to FIG. 2, taking English as the language used by most objects as an example, after attendee A asks "When does the meeting start?", if the user replies "The meeting starts at 10 AM" in Chinese, the Chinese reply "The meeting starts at 10 AM" and the corresponding English "The meeting starts at 10AM" can also be displayed on the electronic device, so that the user can intuitively check whether the text after the voice of the Chinese he said is correctly translated into English.
[0067] In the embodiments of the present application, by determining the object identifier of the attendee according to the voice of the attendee, and then when displaying the voice subtitle of the attendee, the object identifier of the attendee can also be correspondingly displayed, and the voice subtitle is displayed in the associated area of the object identifier. In this way, the user can intuitively view the attendee corresponding to each voice subtitle, which is convenient for the user to understand the voices of different attendees and improves the meeting communication efficiency.
[0068] In some embodiments of the present application, in order to facilitate the user to intuitively view the actual attendee corresponding to each voice, after displaying the object identifier and voice subtitle of the attendee, the above-mentioned method may further include:
[0069] Receiving a fifth input;
[0070] In response to the fifth input, displaying an object identifier editing interface;
[0071] Receiving a sixth input from the user for the first object identifier among the object identifiers of each attendee;
[0072] In response to the sixth input, updating the first object identifier.
[0073] Wherein, the fifth input is an input to the AI interaction interface. The above fifth input is used to display the object identifier editing interface, and the fifth input may be a fifth operation. Exemplarily, the above fifth input includes but is not limited to: the user's touch input to the AI interaction interface through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements, and the embodiments of the present invention do not make limitations. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input or a click input of any number of times, and may also be a long press input or a short press input. For example, the above fifth input may be: the user's touch input to the "role editing" control in the AI interaction interface.
[0074] The object identifier editing interface is used to edit the object identifiers of each participant. This interface can include the object identifiers of each participant.
[0075] The first object identifier can be any object identifier among the object identifiers of the various participants included in the object identifier editing interface.
[0076] The sixth input is an input to the first object identifier, used to update the first object identifier. The sixth input can be a sixth operation. Exemplarily, the sixth input includes, but is not limited to: touch input from the user to the first object identifier via a finger or stylus, a voice command input by the user, a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage needs and are not limited in this embodiment. The specific gesture in this embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-change gesture, a double-press gesture, or a double-click gesture; the click input in this embodiment can be a single-click input, a double-click input, or any number of clicks, and can also be a long-press input or a short-press input. For example, the sixth input can be: a click input from the user to the first object identifier.
[0077] In some embodiments of this application, after displaying the object identifiers of the participants, if the user clearly knows who a certain participant is, the object identifier of that participant can be edited. Specifically, in response to the fifth input, an object identifier editing interface can be displayed, and then in response to the user's sixth input on the first object identifier in the object identifier editing interface, the first object identifier can be updated.
[0078] Continuing with the example above and referring to Figure 4, when the user clicks the "Role Editing" control 41, the object identifier editing interface 42 will be displayed. This object identifier editing interface 42 contains the object identifiers of each participant, namely, the object identifier "Object A" for participant A and the object identifier "Object B" for participant B. If the user knows that participant A is Xiaoming, they can click the "Edit Name" control 421 after the object identifier "Object A" for participant A, which will display the interface 43 for editing the object identifier "Object A" for participant A. In the interface 43, the user can enter the name "Xiaoming" for the object identifier "Object A" for participant A, thus updating the object identifier "Object A" for the participant corresponding to the voice message "When does the meeting start?" to "Xiaoming".
[0079] It should be noted that, referring to Figure 2, before displaying the “Role Editing” control 41 as shown in Figure 4, the user can click the floating ball “Function” 22 in the AI interactive interface 20, and then the function controls of each function in the voice recognition function can be displayed, such as the “Role Editing” control 41, the “Smart Reminder” control 45 and the “Summary” control 46 in Figure 4.
[0080] In the embodiments of this application, in response to the fifth input, an object identifier editing interface can be displayed. Then, in response to the user's sixth input on the first object identifier in the object identifier editing interface, the first object identifier can be updated. In this way, the object identifiers of the participants can be updated according to the user's needs, so that the user can intuitively see the participants corresponding to each voice.
[0081] In some embodiments of this application, in order to improve the diversity of speech-translated text, after step 110, the method described above may further include:
[0082] Receive the first input for the speech-to-text translation;
[0083] In response to the first input, display at least one translation style option;
[0084] Receive a second input for a first translation style option among at least one translation style option;
[0085] In response to the second input, the speech-translated text is updated according to the translation style corresponding to the first translation style option.
[0086] The first input is the input of the voice-translated text. This first input is used to display at least one translation style option and can be a first operation. Exemplarily, the first input includes, but is not limited to: touch input of the voice-translated text by the user using a finger or stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage needs and are not limited in this embodiment of the invention. The specific gestures in this embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-change gesture, a double-press gesture, or a double-click gesture; the click input in this embodiment can be a single-click input, a double-click input, or any number of clicks, and can also be a long-press input or a short-press input. For example, the first input can be: the user's click input of the voice-translated text of each participant that requires a different translation style.
[0087] At least one translation style option can be an option for at least one translation style of the voice-to-text translation. For example, at least one translation style option here could include a "more comprehensible" translation style option and a "more formal" translation style option.
[0088] The first translation style option can be any one of at least one translation style option.
[0089] The second input is the input to the first translation style option among at least one translation style option. This second input is used to update the speech-translated text according to the translation style corresponding to the first translation style option. The second input can be a second operation. For example, the second input includes, but is not limited to: touch input from a user using a finger or stylus to the first translation style option among at least one translation style option; voice commands input by the user; specific gestures input by the user; or other feasible inputs, which can be determined according to actual usage needs. This embodiment of the invention does not limit this. Specific gestures in this application embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-change gesture, a double-press gesture, or a double-tap gesture. Click input in this application embodiment can be a single-click input, a double-tap input, or any number of clicks, and can also be a long-press input or a short-press input. For example, the second input can be: a user's click input to the first translation style option among at least one translation style option.
[0090] In some embodiments of this application, if the speech translation text does not meet the user's needs, the user can adjust the speech translation style. Specifically, in response to a first input to the speech translation text, at least one translation style option can be displayed, and then in response to a second input from the user to a first translation style option among the at least one translation style options, the speech translation text can be updated according to the translation style corresponding to the first translation style option.
[0091] It should be noted that the first input for receiving the speech translation text mentioned above is not performed on the speech translation text of all participants, but rather on the speech translation text of each participant whose speech translation text the user is not satisfied with. For example, in Figure 2 above, if the user is not satisfied with the speech translation text of participant A's speech "When does the meeting start?", the user can perform the first input on the speech translation text "When does the meeting start?".
[0092] In one example, referring to FIG. 5, during a meeting, a participating member said a voice message "Let’s not try to boll the ocean with this project". After translation, the voice translation text is "我们不要试图在这个项目上大海捞针". Regarding this voice translation text "我们不要试图在这个项目上大海捞针", the user cannot understand its meaning. Therefore, the translation style of this voice translation text "我们不要试图在这个项目上大海捞针" can be updated. Specifically, the user can click on this voice translation text "我们不要试图在这个项目上大海捞针", and a translation style adjustment interface 51 will be displayed. The translation style adjustment interface 51 includes at least one translation style option, such as a "better understood" translation style option 511 and a "more formal" translation style option 512. If the user clicks on the "better understood" translation style option 511, the translation style of the voice translation text "我们不要试图在这个项目上大海捞针" can be updated to a better understood voice translation text "我们不要在这个项目上搞的太复杂了".
[0093] In an embodiment of the present application, according to user needs, the translation style of the voice translation text corresponding to the voice can be adjusted and updated, so that a voice translation text with a translation style that better meets user needs can be obtained, improving the diversity of the voice translation text.
[0094] In some embodiments of the present application, referring to FIG. 1, the method described above may include:
[0095] Step 110, when receiving the voice of a participating member, display voice subtitles according to the voice of the participating member; wherein, the voice subtitles include a voice translation text, or, a voice text and a voice translation text;
[0096] Step 120, play at least one of the voice of the participating member and the translated voice of the participating member.
[0097] It can be understood that there is no limitation on the execution order between the display of the voice subtitles and the playback of the voice.
[0098] In some embodiments of the present application, during a meeting, simultaneous interpretation is supported, that is, during the process of receiving the voice of each participating member, the voice of each participating member, and / or, the translated voice of each participating member can also be broadcast.
[0099] The aforementioned simultaneous interpretation can mean translating another person's speech into the user's desired language and broadcasting it, or translating the user's speech into another person's desired language and broadcasting it. That is, in a meeting, four voices will be broadcast: the user's speech, the translated speech, the other party's speech, and the translated speech.
[0100] In the embodiments of this application, during the voice reception process of each participant, the voice of each participant and / or the translated voice of each participant can be broadcast, so that the user can hear the voice of each participant and / or the translated voice of each participant in a timely manner, thereby improving the efficiency of meeting communication.
[0101] In some embodiments of this application, to enhance the diversity of voice broadcasts, the method described above may further include, before playing at least one of the voices of the participants and the translated voices of the participants:
[0102] Receive third input;
[0103] In response to a third input, display a list of voice prompts;
[0104] Receive a fourth input from the user for a first voice broadcast method among at least one voice broadcast method;
[0105] Playing at least one of the voice recordings of the participants and the translated voice recordings of the participants may include:
[0106] In response to the fourth input, at least one of the participant's voice and the participant's translated voice is played according to the first voice broadcast method.
[0107] The third input refers to input made to the AI interactive interface. This third input is used for voice-reading lists and can be a third operation. For example, the third input includes, but is not limited to: touch input from the user via a finger or stylus to the AI interactive interface, voice commands input by the user, specific gestures input by the user, or other feasible inputs, which can be determined according to actual usage needs and are not limited in this embodiment. The specific gestures in this embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-change gesture, a double-press gesture, or a double-tap gesture; the click input in this embodiment can be a single-click input, a double-tap input, or any number of clicks, and can also be a long-press input or a short-press input. For example, the third input can be: a user's click input on the "Settings" control in the AI interactive interface.
[0108] The voice broadcast list can be a list that displays voice broadcast methods, and the voice broadcast list can include at least one voice broadcast method.
[0109] In the case where the participants' voices include the first voice of a first object and the second voice of a second object, and the language of the first voice is the first language and the language of the second voice is the second language, the language of the translated voice of the first object can be the second language.
[0110] In the above scenario, taking the user of the electronic device corresponding to the voice processing method provided in this application embodiment as the first object and other participants in the meeting as the second object, at least one voice broadcasting method can include: a silent second voice message, a translation voice message corresponding to the silent second voice message, and a silent first voice message. This allows the user to select the desired voice broadcasting method according to their needs, improving the flexibility of voice broadcasting.
[0111] Mute the second speech can be to mute the speech of the second object.
[0112] The translation of the second speech can be silenced by muting the translation of the second speech.
[0113] Mute the first voice can mute the voice of the first object.
[0114] The first voice broadcast method can be any one of at least one voice broadcast method.
[0115] The fourth input is input to the first voice broadcast method. This fourth input is used to play at least one of the participant's voice and the participant's translated voice according to the first voice broadcast method. The fourth input can be a fourth operation. Exemplarily, the fourth input includes, but is not limited to: touch input from the user via a finger or stylus to the first voice broadcast method, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage needs. This embodiment of the invention does not limit this. The specific gesture in this application embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-change gesture, a double-press gesture, or a double-click gesture; the click input in this application embodiment can be a single-click input, a double-click input, or any number of clicks, and can also be a long-press input or a short-press input. For example, the fourth input can be: the user's click input to the first voice broadcast method.
[0116] In some embodiments of this application, when playing at least one of the participant's voice and the participant's translated voice, the playback can be performed according to a user-defined playback method. Specifically, in response to a third input, a voice playback list containing at least one voice playback method can be displayed, and then in response to a fourth input from the user for a first voice playback method among the at least one voice playback methods, at least one of the participant's voice and the participant's translated voice can be played according to the first voice playback method.
[0117] Referring to Figure 2, the user can click the "Settings" control 21, which will display the voice broadcast list 62 as shown in Figure 6. The voice broadcast list 62 contains four voice broadcast modes: "Mute the other party's original voice" 621, "Mute my original voice" 622, "Mute my translation broadcast on this side" 623, and "Simultaneous translation and delayed playback" 624. The user can choose one of the voice broadcast modes. For example, if the user chooses "Mute the other party's original voice" 621, then when the user listens to the voices of each participant in the meeting on their own electronic device, if the language of other participants' voices is not the user's usual language, then the voices of other participants will be muted. That is, the user cannot hear the voices of other participants on their own electronic device, but can only hear the translated voices of other participants.
[0118] It should be noted that the aforementioned "mute other party's original voice" refers to muting the original voice of other participants when their speech is in a language other than the user's preferred language. This is essentially the "mute second voice" mentioned above. For example, if the user's preferred language is Chinese, and a participant speaks in a language other than Chinese, such as English, and the user has set "mute other party's original voice," the user will only hear a translated version of that participant's English speech in Chinese.
[0119] The aforementioned "Mute my original voice" refers to muting the user's own voice, which is equivalent to "Mute first voice" mentioned above. During a meeting, if a user has set "Mute my original voice," their own voice will not be included in the audio broadcast.
[0120] The aforementioned "My translation broadcast is muted on this device" means that the translated audio of the user's speech is not broadcast on the user's side, which is equivalent to "mute the translated audio of the second speaker." Because during a meeting, after speaking, users don't need to hear their own translated audio, but only want to hear the translated audio of other participants, users can set "My translation broadcast is muted on this device" so that they won't hear their own translated audio on their own side.
[0121] In the embodiments of this application, the voice broadcast method can be set according to user needs, so that at least one of the voice of the participants and the translated voice of the participants can be played according to the first voice broadcast method, thereby improving the diversity of voice broadcast.
[0122] In some embodiments of this application, to ensure the accuracy of simultaneous interpretation, when the first speech is received before the second speech, i.e., when the first person speaks before the second person, at least one speech broadcasting method may further include: synchronous translation and delayed playback. This "synchronous translation and delayed playback" is designed to present the broadcast content in the order of speech. For example, participant A speaks in English, which needs to be translated into Chinese for playback. However, translation takes time. During the translation process, participant B speaks in Chinese. On participant B's electronic device, the Chinese speech can be played directly. In reality, participant A speaks first, and participant B speaks later. Due to the translation delay, during the speech broadcasting process, participant B's Chinese speech may appear before participant A's translated speech. If "synchronous translation and delayed playback" is set, participant B's Chinese speech can be played with a delay to ensure that participant A's translated speech precedes participant B's Chinese speech.
[0123] Playing at least one of the voice recordings of the participants and the translated voice recordings of the participants may include:
[0124] When the first voice broadcast method is synchronous translation and delayed playback, at least one of the first voice and the translated voice of the first voice is played first, and then at least one of the second voice and the translated voice of the second voice is played.
[0125] In some embodiments of this application, when the first voice is received before the second voice, that is, when the first object speaks before the second object, when broadcasting at least one of the voice of the participant and the translated voice of the participant, if the first voice is broadcast in a simultaneous translation and delayed playback mode, then at least one of the first voice and the translated voice of the first voice is played first, and then at least one of the second voice and the translated voice of the second voice is broadcast. This ensures that at least one of the first voice and the translated voice of the first voice can be broadcast before at least one of the second voice and the translated voice of the second voice.
[0126] It should be noted that after setting the simultaneous translation and delayed playback broadcast mode, when displaying the audio translation text of the first speech and the audio text of the second speech, the audio translation text of the first speech can also be displayed first, and then the audio text of the second speech can be displayed. That is, it can be ensured that the content displayed on the audio subtitles can also be presented in the order of speech.
[0127] In the embodiments of this application, by setting a broadcast method of simultaneous translation and delayed playback, when broadcasting at least one of the participants' voices and the translated voices, the broadcast can be carried out in the order of the participants' speeches, ensuring the accuracy of simultaneous interpretation.
[0128] In existing technologies, if a meeting is long, it is difficult for users to maintain a high level of concentration throughout the meeting, which may cause them to miss key information.
[0129] To address the aforementioned problem, prior to step 110, the method may further include:
[0130] Receive the seventh input;
[0131] In response to the seventh input, determine the reminder keywords and members to be reminded to speak;
[0132] If a participant's voice text includes a reminder keyword and the participant is the one who will remind them to speak, a prompt message will be displayed.
[0133] The seventh input refers to input made to the AI interactive interface. This seventh input is used to determine reminder keywords and members to be reminded to speak. The seventh input can be a seventh operation. Exemplarily, the seventh input includes, but is not limited to: touch input from the user via a finger or stylus to the AI interactive interface, voice commands input by the user, specific gestures input by the user, or other feasible inputs, which can be determined according to actual usage needs. This embodiment of the invention does not limit this. Specific gestures in this application embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-change gesture, a double-press gesture, or a double-tap gesture. Click input in this application embodiment can be a single-click input, a double-tap input, or any number of clicks, and can also be a long-press input or a short-press input. For example, the seventh input can be: the user's click input on the "Smart Reminder" control in the AI interactive interface.
[0134] Reminder keywords can be words used in the participants' voices to draw the user's attention.
[0135] A speaker reminder member can be a speaker who is used to draw the user's attention, meaning that the user should pay attention when that speaker is speaking.
[0136] The prompt message can be generated when the voice text of a participant includes reminder keywords, and the participant is a speaker reminder member. This prompt message can be used to remind users that the voice of the speaker reminder member includes reminder keywords.
[0137] In some embodiments of this application, to prevent users from missing important information during a meeting, words and speakers that the user should pay attention to can be pre-set before receiving the voices of the participants. That is, keywords and speakers to remind the user to speak can be set. When each participant speaks, the voice characteristics of each participant can be matched with the voice characteristics of the set speakers to remind the user to speak. If it is determined that the voice characteristics of a certain participant match the voice characteristics of the set speakers to remind the user to speak, the content of the speaker's speech can be matched with the reminder keywords. If it is determined that the content of the speaker's speech matches the reminder keywords, for example, if the content of the speaker's speech mentions the reminder keywords, then a prompt message can be generated to remind the user.
[0138] In one example, referring to Figure 4, the user can click the "Smart Reminder" control 45, and then, as shown in Figure 7, the Smart Reminder interface 71 will be displayed. In this Smart Reminder interface 71, the user can set the members to be reminded and the reminder keywords. For example, if the user sets the reminder keywords "report, to-do" and the members to be reminded are "Xiaoming, Xiaohong", then in subsequent speeches, if "Xiaoming" and / or "Xiaohong" mention "report" or "to-do", a prompt message can be generated to remind the user. For example, in a meeting, if Xiaoming says "The next to-do item is to actively promote Project X", then the prompt message "Xiaoming mentioned the next to-do item" will be generated.
[0139] It should be noted that when matching the content of a speaker's message with the reminder keyword, a fuzzy matching method can be used. This means that the message doesn't necessarily have to mention the reminder keyword; it could also contain similar words, or the message's meaning is the same as the reminder keyword. For example, if a user sets the reminder keyword "to-do," and a speaker's message is "Next, we need to actively promote Project X," then this message expresses the meaning of upcoming tasks, which is the same as the reminder keyword "to-do." Therefore, a reminder message will be generated to alert the user.
[0140] In the embodiments of this application, by setting speaking reminder members and reminder keywords, when the voice text of a participant includes reminder keywords and the participant is a speaking reminder member, a prompt message can be displayed to avoid the user missing important information.
[0141] In some embodiments of this application, to facilitate users in tracing the statements made by each participant in a meeting, the methods described above may further include:
[0142] If the meeting has ended, receive the eighth input;
[0143] In response to the eighth input, the meeting minutes generation interface is displayed;
[0144] Receive the ninth input for the first meeting minutes type option in at least one meeting minutes type option;
[0145] In response to the ninth input, the meeting minutes are generated based on the voice of each participant, according to the meeting minutes type corresponding to the first meeting minutes type option.
[0146] The eighth input refers to input made to the AI interactive interface. This eighth input is used to display the meeting minutes generation interface and can be the eighth operation. For example, the eighth input includes, but is not limited to: touch input from the user via a finger or stylus to the AI interactive interface, voice commands input by the user, specific gestures input by the user, or other feasible inputs, which can be determined according to actual usage needs and are not limited in this embodiment. Specific gestures in this embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-change gesture, a double-press gesture, or a double-tap gesture; the click input in this embodiment can be a single-click input, a double-tap input, or any number of clicks, and can also be a long-press input or a short-press input. For example, the eighth input can be: user click input to the AI interactive interface.
[0147] The meeting minutes generation interface can be an interface used to generate meeting minutes. This meeting minutes generation interface can include at least one meeting minutes type option, where the meeting minutes type option corresponds to the type of meeting minutes to be generated.
[0148] The first meeting minutes type option can be any one of at least one meeting minutes type option.
[0149] The ninth input is an input to the first meeting minutes generation type option. This ninth input is used to generate meeting minutes based on the voice recordings of each participant according to the meeting minutes type corresponding to the first meeting minutes type option. The ninth input can be a ninth operation. Exemplarily, the ninth input includes, but is not limited to: touch input from the user via a finger or stylus to the first meeting minutes generation type option, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage needs. This embodiment of the invention does not limit this. The specific gesture in this application embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-change gesture, a double-press gesture, or a double-click gesture; the click input in this application embodiment can be a single-click input, a double-click input, or any number of clicks, and can also be a long-press input or a short-press input. For example, the ninth input can be: the user's click input to the first meeting minutes generation type option.
[0150] In some embodiments of this application, upon the conclusion of a meeting, in response to an eighth input from the user, a meeting minutes generation interface containing at least one meeting minutes type option can be displayed. Then, in response to a ninth input from the user to a first meeting minutes type option among the at least one meeting minutes type options, meeting minutes can be generated according to the meeting minutes type corresponding to the first meeting minutes type option, based on the voices of each participant.
[0151] Referring to Figure 4, when the user clicks the "Summary" control 46, the meeting minutes generation interface 81 will be displayed as shown in Figure 8. This meeting minutes generation interface 81 includes two meeting minutes generation type options: "Full Text Summary" option 811 and "Custom Summary" option 812. The user can select either "Full Text Summary" option 811 or "Custom Summary" option 812. For example, if the user selects "Full Text Summary" option 811, the meeting minutes will be generated according to the voices of each participant in the manner corresponding to the full text summary option.
[0152] It should be noted that when generating meeting minutes, a pre-trained large model can be used to translate the speech of each participant into text.
[0153] It should be noted that the generated meeting minutes can be structured, such as including the meeting's theme, summary, key points, and to-do list.
[0154] In the embodiments of this application, when the meeting ends, in response to the user's eighth input, a meeting minutes generation interface containing at least one meeting minutes type option is displayed. Then, in response to the user's ninth input for the first meeting minutes type option among the at least one meeting minutes type option, meeting minutes can be generated according to the meeting minutes type corresponding to the first meeting minutes type option, based on the voices of each participant. This facilitates the user's subsequent tracing of the speech of each participant in the meeting. Furthermore, when generating meeting minutes, the meeting minutes can be generated according to the user's required meeting minutes type, thus improving the flexibility of meeting minutes generation.
[0155] In some embodiments of this application, to further enhance the flexibility of meeting minutes generation, the step of generating meeting minutes based on the voice recordings of each participant according to the meeting minutes type corresponding to the first meeting minutes type option may include:
[0156] If the meeting minutes type corresponding to the first meeting minutes type option is the first meeting minutes type, then the meeting minutes are generated based on the voices of each participant according to the first meeting minutes type.
[0157] If the meeting minutes type corresponding to the first meeting minutes type option is the second meeting minutes type, the meeting minutes will be generated based on the voice of the selected target participants.
[0158] The first meeting minutes type option and the second meeting minutes type option can be two different meeting minutes type options from at least one meeting minutes type option. For example, the first meeting minutes type option can be the "Full Summary" option 811 in Figure 8 above, and the second meeting minutes type option can be the "Custom Summary" option 812 in Figure 8.
[0159] The first meeting minutes type can be the meeting minutes type corresponding to the first meeting minutes type option. When the first meeting minutes type option can be the "Full Summary" option 811 in Figure 8 above, the first meeting minutes type is to organize the voices of all participants in the meeting to form a meeting minutes containing the speech content of all participants.
[0160] The second meeting minutes type can be the meeting minutes type corresponding to the second meeting minutes type option. When the second meeting minutes type option can be the "Custom Summary" option 812 in Figure 8 above, the second meeting minutes type is to organize the voices of the selected participants in the meeting to form a meeting minutes that only contains the speech content of the selected participants.
[0161] The target participants can be those selected by the user to form meeting minutes.
[0162] In some embodiments of this application, when the meeting minutes type corresponding to the first meeting minutes type option is the first meeting minutes type, meeting minutes can be generated according to the voice of each participant based on the first meeting minutes type.
[0163] Referring to Figure 8, taking the first meeting minutes type option as "Full Summary" option 811 in Figure 8 as an example, after the user selects "Full Summary" option 811, as shown in Figure 9, the voices of all participants in the meeting can be organized to form a meeting minutes containing the speeches of all participants.
[0164] It should be noted that the meeting minutes shown in Figure 9 may include the main content of the meeting, such as the theme, summary, key points, and tasks to be completed.
[0165] If the meeting minutes type corresponding to the first meeting minutes type option is the second meeting minutes type, the meeting minutes will be generated based on the voice of the selected target participants.
[0166] Referring to Figure 8, taking the first meeting minutes type option as "Custom Summary" option 812 in Figure 8 as an example, after the user selects "Custom Summary" option 812, the user can select the target participant. For example, if the selected target participant is Xiaoming, then as shown in Figure 10, the content of Xiaoming's speech in the meeting can be summarized to form a meeting minutes containing only Xiaoming's speech.
[0167] It should be noted that the meeting minutes shown in Figure 10 can include the main content of Xiaoming's speech, such as the theme, summary, key points, and tasks to be done.
[0168] In the embodiments of this application, different meeting minutes type options can be selected according to user needs, thereby obtaining meeting minutes of different types, thus improving the flexibility of meeting minutes generation.
[0169] In some embodiments of this application, the selection of target participants can be implemented in the following manner: that is, before generating meeting minutes based on the voice of the selected target participants, the above-mentioned method may further include:
[0170] Receive the tenth input for the second object identifier;
[0171] In response to the tenth input, the participant selection interface is displayed;
[0172] In response to the eleventh input from the user regarding the third object identifier in the object identifiers of each participant, the target participant is determined.
[0173] The second object identifier can be any one of the object identifiers of multiple speaking objects displayed in the AI interactive interface. For example, the second object identifier can be the object identifier "Object A" of participant A in Figure 2.
[0174] The tenth input is the input to the second object identifier. This tenth input is used to display the participant selection interface and can be the tenth operation. Exemplarily, the tenth input includes, but is not limited to: touch input from the user using a finger or stylus to the second object identifier, a voice command input by the user, a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage needs. This embodiment of the invention does not limit this. The specific gesture in this application embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long-press gesture, an area change gesture, a double-press gesture, or a double-click gesture. The click input in this application embodiment can be a single-click input, a double-click input, or any number of clicks, and can also be a long-press input or a short-press input. For example, the tenth input can be: the user's click input on the second object identifier.
[0175] The participant selection interface can be used to select target participants, and the participant selection interface can include the object identifier of each participant.
[0176] The third object identifier can be any object identifier among the object identifiers of each participant included in the participant selection interface.
[0177] The eleventh input is the input of the third object identifier in the object identifier of each participant. The eleventh input is used to determine the target participant and can be an eleventh operation. For example, the eleventh input includes, but is not limited to: touch input by the user using a finger or stylus to the third object identifier in the object identifier of each participant, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage needs. This embodiment of the invention does not limit this. The specific gesture in this application embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long-press gesture, an area change gesture, a double-press gesture, or a double-click gesture; the click input in this application embodiment can be a single-click input, a double-click input, or any number of clicks, and can also be a long-press input or a short-press input. For example, the eleventh input can be: the user's click input of the third object identifier in the object identifier of each participant.
[0178] In some embodiments of this application, in response to a tenth input of a second object identifier, a participant selection interface including object identifiers of each participant can be displayed. Then, in response to a user's eleventh input of a third object identifier among the object identifiers of each participant, the participant corresponding to the third object identifier can be determined as the target participant.
[0179] Referring to Figure 8, after the user selects the "Custom Summary" option 812, the participant selection interface 111 will be displayed as shown in Figure 11. This participant selection interface 111 includes the object identifier of each participant in the meeting. For example, there is the object identifier "Xiaoming" 1111 for participant Xiaoming, the object identifier "Xiaoming" 1112 for participant Xiaohong, and the object identifier "Zhangsan" 1113 for participant Zhangsan. If the user selects the object identifier "Xiaoming" 1111 for participant Xiaoming, then the target participant can be identified as Xiaoming. After the user clicks the "OK" control 112, the meeting minutes shown in Figure 10 will be generated.
[0180] In the embodiments of this application, in response to the tenth input of the second object identifier, a participant selection interface including the object identifiers of each participant can be displayed. Then, in response to the eleventh input of the user of the third object identifier among the object identifiers of each participant, the participant corresponding to the third object identifier can be determined as the target participant. In this way, according to the user's needs, the target participants involved in the meeting minutes of the meeting minutes type corresponding to the second meeting minutes type option can be selected, which improves the flexibility of target participant selection.
[0181] It should be noted that after generating the meeting minutes, the meeting minutes can also be exported. They can be exported as PDF versions, for example, by clicking the "Export PDF" control in Figures 9 and 10, or as Word versions, for example, by clicking the "Export Word" control in Figures 9 and 10. The specific format of the meeting minutes exported can be selected according to the user's needs, and is not limited in this embodiment.
[0182] The speech processing method provided in this application can be executed by a speech processing device. This application uses a speech processing device executing the speech processing method as an example to illustrate the speech processing device provided in this application.
[0183] Figure 12 is a schematic diagram of a speech processing device according to an exemplary embodiment. As shown in Figure 12, the speech processing device 1300 may include:
[0184] Display module 1310 is used to display audio subtitles based on the voice of a participant when the audio of the participant is received.
[0185] The audio subtitles include either the audio-translated text or both the audio text and the audio-translated text.
[0186] In this embodiment, when meeting participants speak different languages, and upon receiving their voice messages, the system can intuitively display the translated text of their voice messages, or display both the translated text and the voice message itself. This allows users to easily view the translated text or the translated text and the voice message, and further enables them to communicate with the participants based on the translated text or the translated text and the voice message, thus improving communication efficiency between participants speaking different languages.
[0187] In some embodiments of this application, the display module 1310 is specifically used for:
[0188] The AI interactive interface displays the audio captions of the participants.
[0189] In some embodiments of this application, the apparatus described above may further include:
[0190] The first receiving module is used to receive the first input to the speech-translated text;
[0191] The display module 1310 is also configured to display at least one translation style option in response to the first input;
[0192] The second receiving module is used to receive a second input from the user for a first translation style option among the at least one translation style options;
[0193] The first update module is used to update the speech-translated text in response to the second input, according to the translation style corresponding to the first translation style option.
[0194] In some embodiments of this application, the apparatus described above may further include:
[0195] The audio playback module is used to play at least one of the audio of the participating member and the translated audio of the participating member.
[0196] In some embodiments of this application, the apparatus described above may further include:
[0197] The third receiving module is used to receive the third input;
[0198] The display module 1310 is further configured to display a voice broadcast list in response to the third input, wherein the voice broadcast list includes at least one voice broadcast method;
[0199] The fourth receiving module is used to receive the user's fourth input to the first voice broadcast method among the at least one voice broadcast method;
[0200] The voice playback module is specifically used for:
[0201] In response to the fourth input, at least one of the participant's voice and the participant's translated voice is played according to the first voice broadcast method.
[0202] In some embodiments of this application, the voice of the participants includes a first voice of a first object and a second voice of a second object, wherein the language of the first voice is a first language, the language of the second voice is a second language, and the language of the translation voice of the first object is the second language;
[0203] The at least one voice broadcasting method includes: muting the second voice, muting the translated voice of the second voice, and muting the first voice.
[0204] In some embodiments of this application, when the first object speaks before the second object, the at least one voice broadcasting method further includes: synchronous translation and delayed playback;
[0205] The voice playback module is specifically used for:
[0206] When the first voice broadcasting method is synchronous translation and delayed playback, at least one of the first voice and the translated voice of the first voice is played first, and then at least one of the second voice and the translated voice of the second voice is broadcast.
[0207] In some embodiments of this application, the display module 1310 is specifically used for:
[0208] The object identifier of the participant is determined based on the participant's voice.
[0209] The system displays the object identifier of the participating member and the audio captions; the audio captions are located in the associated area of the object identifier of the participating member.
[0210] In some embodiments of this application, the apparatus described above may further include:
[0211] The fifth receiving module is used to receive the fifth input;
[0212] The display module 1310 is also configured to respond to the fifth input by displaying an object identifier editing interface, wherein the object identifier editing interface includes object identifiers of each participant;
[0213] The sixth receiving module is used to receive the user's sixth input for the first object identifier in the object identifier of each participating member;
[0214] The second update module is used to update the first object identifier in response to the sixth input.
[0215] In some embodiments of this application, the apparatus described above may further include:
[0216] The seventh receiving module is used to receive the seventh input;
[0217] The first determining module is used to determine the reminder keyword and the member who is required to speak in response to the seventh input.
[0218] The display module 1310 is also used to display a prompt message when the voice text of the participant includes the reminder keyword and the participant is the speaker reminder member.
[0219] In some embodiments of this application, the apparatus described above may further include:
[0220] The eighth receiving module is used to receive the eighth input when the meeting ends;
[0221] Display module 1310 is also configured to display a meeting minutes generation interface in response to the eighth input, the meeting minutes generation interface including at least one meeting minutes type option;
[0222] The ninth receiving module is used to receive a ninth input for the first meeting minutes type option in at least one meeting minutes type option;
[0223] The meeting minutes generation module is used to respond to the ninth input and generate meeting minutes based on the voice of each participant according to the meeting minutes type corresponding to the first meeting minutes type option.
[0224] In some embodiments of this application, the meeting minutes generation module is specifically used for:
[0225] If the meeting minutes type corresponding to the first meeting minutes type option is the first meeting minutes type, then the meeting minutes are generated according to the voice of each participant based on the first meeting minutes type.
[0226] When the meeting minutes type corresponding to the first meeting minutes type option is the second meeting minutes type, the meeting minutes are generated based on the voice of the selected target participants.
[0227] The voice processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.
[0228] The voice processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0229] The speech processing device provided in this application embodiment can implement the various processes implemented in the method embodiment of FIG1. To avoid repetition, it will not be described again here.
[0230] Optionally, as shown in FIG13, this application embodiment also provides an electronic device 1400, including a processor 1401 and a memory 1402. The memory 1402 stores a program or instructions that can run on the processor 1401. When the program or instructions are executed by the processor 1401, they implement the various steps of the above-described speech processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0231] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0232] Figure 14 is a schematic diagram of the hardware structure of an electronic device that implements an embodiment of this application.
[0233] The electronic device 1500 includes, but is not limited to, components such as: radio frequency unit 1501, network module 1502, audio output unit 1503, input unit 1504, sensor 1505, display unit 1506, user input unit 1507, interface unit 1508, memory 1509, and processor 1510.
[0234] Those skilled in the art will understand that the electronic device 1500 may also include a power supply (such as a battery) for powering various components. The power supply may be logically connected to the processor 1510 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The electronic device structure shown in Figure 14 does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0235] The display unit 1506 is used to display audio subtitles based on the voice of a participant when the audio of the participant is received; wherein the audio subtitles include audio translation text, or audio text and audio translation text.
[0236] In this way, when participants speak different languages, the system can intuitively display the translated text of their speech, or both the translated text and the original speech. This allows users to easily view the translated text and the original speech, and then communicate with the participants based on these texts, thus improving communication efficiency among participants speaking different languages.
[0237] Optionally, the display unit 1506 is also used to display the voice captions of the participants on the artificial intelligence (AI) interactive interface.
[0238] In this way, by displaying the audio subtitles of the participants in the AI interactive interface overlaid on the meeting interface, users can view the audio subtitles of the participants without having to exit the meeting interface and switch to the AI interactive interface. This means that users can intuitively view the audio subtitles of the participants without having to switch back and forth between interfaces, thus improving the viewing efficiency of the audio subtitles of the participants.
[0239] Optionally, the user input unit 1507 is configured to receive a first input to the speech-translated text;
[0240] The display unit 1506 is also configured to display at least one translation style option in response to the first input;
[0241] User input unit 1507 is also configured to receive a second input from a user for a first translation style option among the at least one translation style options;
[0242] Processor 1510 is configured to update the speech-translated text in response to the second input, according to the translation style corresponding to the first translation style option.
[0243] In this way, the translation style of the voice translation text can be adjusted and updated according to user needs, thus obtaining voice translation text with a translation style that better meets user needs and improving the diversity of voice translation text.
[0244] Optionally, the audio output unit 1503 is used to play at least one of the speech of the participant and the translated speech of the participant.
[0245] In this way, during the voice reception process of each participant, the voice of each participant and / or the translated voice of each participant can be broadcast, so that users can hear the voice of each participant and / or the translated voice of each participant in a timely manner, thereby improving the efficiency of meeting communication.
[0246] Optionally, the user input unit 1507 is also used to receive a third input;
[0247] The display unit 1506 is further configured to display a voice broadcast list in response to the third input, wherein the voice broadcast list includes at least one voice broadcast method;
[0248] The user input unit 1507 is also used to receive a fourth input from the user regarding the first voice broadcast method in the at least one voice broadcast method;
[0249] The audio output unit 1503 is further configured to, in response to the fourth input, play at least one of the participant's voice and the participant's translated voice according to the first voice broadcasting method.
[0250] In this way, the voice broadcasting method can be set according to user needs. At least one of the participants' voices and the participants' translated voices can be played according to the first voice broadcasting method, which improves the diversity of voice broadcasting.
[0251] Optionally, the voices of the participants include a first voice of a first object and a second voice of a second object, wherein the language of the first voice is a first language, the language of the second voice is a second language, and the language of the translated voice of the first object is the second language; the at least one voice broadcasting method includes: muting the second voice, muting the translated voice of the second voice, and muting the first voice.
[0252] In this way, users can choose the voice broadcast method they need, which improves the flexibility of voice broadcast.
[0253] Optionally, when the first object speaks before the second object, the at least one voice broadcasting method further includes: synchronous translation and delayed playback; the audio output unit 1503 is further configured to, when the first voice broadcasting method is synchronous translation and delayed playback, first play at least one of the first voice and the translated voice of the first voice, and then broadcast the second voice and at least one of the translated voice of the second voice.
[0254] Thus, by setting up a simultaneous translation and delayed playback broadcast mode, when broadcasting at least one of the participants' voices and the translated voices, the broadcast can be done in the order in which the participants speak, ensuring the accuracy of simultaneous interpretation.
[0255] Optionally, the processor 1510 is further configured to determine the object identifier of the participant based on the participant's voice;
[0256] The display unit 1506 is also used to display the object identifier of the participant and the voice subtitle; the voice subtitle is located in the associated area of the object identifier of the participant.
[0257] In this way, by determining the object identifier of the participant based on the participant's voice, and then displaying the participant's voice subtitles, the object identifier of the participant can also be displayed accordingly, and the voice subtitles can be displayed in the associated area of the object identifier. This allows users to intuitively see the participant corresponding to each voice subtitle, making it easier for users to understand the voice of different participants and improving the efficiency of meeting communication.
[0258] Optionally, the user input unit 1507 is also used to receive a fifth input;
[0259] The display unit 1506 is also configured to respond to the fifth input by displaying an object identifier editing interface, wherein the object identifier editing interface includes object identifiers of each participant;
[0260] The user input unit 1507 is also used to receive a sixth input from the user for the first object identifier in the object identifiers of each participant;
[0261] The processor 1510 is also configured to update the first object identifier in response to the sixth input.
[0262] Thus, by responding to the fifth input, the object identifier editing interface can be displayed. Then, by responding to the user's sixth input on the first object identifier in the object identifier editing interface, the first object identifier can be updated. In this way, the object identifiers of the participants can be updated according to the user's needs, so that the user can intuitively see the participants corresponding to each voice.
[0263] Optionally, the user input unit 1507 is also used to receive a seventh input;
[0264] The processor 1510 is also configured to, in response to the seventh input, determine the reminder keyword and the speaker reminder member;
[0265] The display unit 1506 is also configured to display a prompt message when the voice text of the participant includes the reminder keyword and the participant is the speaker reminder member.
[0266] In this way, by setting reminder members and reminder keywords, a prompt message can be displayed when the voice text of a participant includes the reminder keyword and the participant is a reminder member, so as to avoid users missing important information.
[0267] Optionally, the user input unit 1507 is also configured to receive an eighth input if the meeting ends;
[0268] The display unit 1506 is also configured to display a meeting minutes generation interface in response to the eighth input, the meeting minutes generation interface including at least one meeting minutes type option;
[0269] User input unit 1507 is also configured to receive a ninth input for a first meeting minutes type option among at least one meeting minutes type option;
[0270] The processor 1510 is also configured to, in response to the ninth input, generate meeting minutes based on the voice of each participant according to the meeting minutes type corresponding to the first meeting minutes type option.
[0271] Thus, upon the conclusion of the meeting, in response to the user's eighth input, a meeting minutes generation interface containing at least one meeting minutes type option is displayed. Then, in response to the user's ninth input regarding the first meeting minutes type option, meeting minutes can be generated based on the voices of each participant according to the meeting minutes type corresponding to the first meeting minutes type option. This facilitates subsequent retrieval of each participant's speech during the meeting. Furthermore, when generating meeting minutes, the meeting minutes can be generated according to the user's desired meeting minutes type, thereby enhancing the flexibility of meeting minutes generation.
[0272] Optionally, the processor 1510 is further configured to generate meeting minutes based on the voice of each participant according to the first meeting minutes type when the meeting minutes type corresponding to the first meeting minutes type option is a first meeting minutes type; and to generate meeting minutes based on the voice of the selected target participant when the meeting minutes type corresponding to the first meeting minutes type option is a second meeting minutes type.
[0273] In this way, users can select different meeting minutes type options according to their needs, thereby obtaining meeting minutes of different types, thus improving the flexibility of meeting minutes generation.
[0274] It should be understood that, in this embodiment, the input unit 1504 may include a graphics processing unit (GPU) 15041 and a microphone 15042. The GPU 15041 processes image data of still images or videos obtained by an image capture device (such as a color camera) in video capture mode or image capture mode. The display unit 1506 may include a display panel 15061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1507 includes at least one of a touch panel 15071 and other input devices 15072. The touch panel 15071 is also called a touch screen. The touch panel 15071 may include a touch detection device and a touch controller. Other input devices 15072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0275] The memory 1509 can be used to store software programs and various data. The memory 1509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0276] Processor 1510 may include one or more processing units; optionally, processor 1510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1510.
[0277] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described speech processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0278] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0279] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described speech processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0280] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0281] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described speech processing method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0282] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0283] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0284] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A speech processing method, the method comprising: Upon receiving voice messages from participants, display audio subtitles based on those messages. The audio subtitles include either the audio-translated text or both the audio text and the audio-translated text.
2. The method according to claim 1, wherein, The method of displaying audio subtitles based on the voices of the participants includes: The AI interactive interface displays the voice captions of the participants.
3. The method according to claim 1, further comprising: Receive the first input to the speech-translated text; In response to the first input, at least one translation style option is displayed; Receive a second input from the user for a first translation style option among the at least one translation style options; In response to the second input, the speech-translated text is updated according to the translation style corresponding to the first translation style option.
4. The method according to claim 1, further comprising: Play at least one of the participant's voice recording and the participant's translated voice recording.
5. The method according to claim 4, further comprising: Receive third input; In response to the third input, a voice broadcast list is displayed, wherein the voice broadcast list includes at least one voice broadcast method; Receive a fourth input from the user for the first voice broadcast method among the at least one voice broadcast methods; Playing at least one of the participant's voice and the participant's translated voice includes: In response to the fourth input, at least one of the participant's voice and the participant's translated voice is played according to the first voice broadcast method.
6. The method according to claim 5, wherein, The voices of the participants include the first voice of the first object and the second voice of the second object, wherein the language of the first voice is the first language, the language of the second voice is the second language, and the language of the translation voice of the first object is the second language; The at least one voice broadcasting method includes: muting the second voice, muting the translated voice of the second voice, and muting the first voice.
7. The method according to claim 6, wherein, When the first object speaks before the second object, the at least one voice broadcasting method further includes: simultaneous translation and delayed playback; Playing at least one of the participant's voice and the participant's translated voice includes: When the first voice broadcasting method is synchronous translation and delayed playback, at least one of the first voice and the translated voice of the first voice is played first, and then at least one of the second voice and the translated voice of the second voice is broadcast.
8. The method according to claim 1, wherein, The method of displaying audio subtitles based on the voices of the participants includes: The object identifier of the participant is determined based on the participant's voice. The system displays the object identifier of the participating member and the audio captions; the audio captions are located in the associated area of the object identifier of the participating member.
9. The method according to claim 8, further comprising: Receive the fifth input; In response to the fifth input, an object identifier editing interface is displayed, wherein the object identifier editing interface includes the object identifiers of each participating member; Receive the sixth input from the user regarding the first object identifier in the object identifiers of each participating member; In response to the sixth input, the first object identifier is updated.
10. The method according to claim 1, further comprising: Receive the seventh input; In response to the seventh input, the reminder keyword and the member who should speak are determined; If the voice text of a participant includes the reminder keyword, and the participant is the one who issued the speaking reminder, a prompt message will be displayed.
11. The method according to claim 1, further comprising: If the meeting has ended, receive the eighth input; In response to the eighth input, a meeting minutes generation interface is displayed, the meeting minutes generation interface including at least one meeting minutes type option; Receive the ninth input for the first meeting minutes type option in at least one meeting minutes type option; In response to the ninth input, meeting minutes are generated based on the voice of each participant according to the meeting minutes type corresponding to the first meeting minutes type option.
12. The method according to claim 11, wherein, The step of generating meeting minutes based on the voice recordings of each participant according to the meeting minutes type corresponding to the first meeting minutes type option includes: If the meeting minutes type corresponding to the first meeting minutes type option is the first meeting minutes type, then the meeting minutes are generated according to the voice of each participant based on the first meeting minutes type. When the meeting minutes type corresponding to the first meeting minutes type option is the second meeting minutes type, the meeting minutes are generated based on the voice of the selected target participants.
13. A voice processing apparatus, the apparatus comprising: The display module is used to display audio subtitles based on the voice of the participants when the audio of the participants is received. The audio subtitles include either the audio-translated text or both the audio text and the audio-translated text.
14. An electronic device comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the speech processing method as claimed in any one of claims 1-12.
15. An electronic device configured to perform the speech processing method as claimed in any one of claims 1-12.
16. A readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the speech processing method as claimed in any one of claims 1-12.
17. A computer program product, said computer program product being executed by at least one processor to implement the speech processing method as claimed in any one of claims 1-12.
18. A chip comprising a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the speech processing method as described in any one of claims 1-12.
Citation Information
Patent Citations
Simultaneous interpretation method and device for teleconference, electronic device and storage medium
CN112153323A
Simultaneous interpretation method and system with machine and manual cooperation mode
CN112232092A
Text display method and device, electronic equipment and storage medium
CN114373464A
Translation method and device, wearable device, terminal device and readable storage medium
CN118468898A
Voice processing method and device, electronic equipment and storage medium
CN119517036A