Instant messaging method and device based on AI subtitles and storage medium
By introducing AI subtitles function in instant messaging applications, the text entered by users is converted into audio data, which solves the problem that users cannot exchange messages in a timely manner during voice calls or video calls, and achieves more efficient communication.
Patent Information
- Application Number
- CN202311641845.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-11-30
AI Technical Summary
When using instant messaging applications for voice or video calls, users may be inconvenient to speak or have language dysfunction, resulting in the inability to exchange messages in time.
Using an instant communication method based on AI subtitles, the text input by the user is converted into audio data through the AI subtitle application, and the audio data is sent to the second device through the instant messaging application. The user using the second device can hear a sound corresponding to the audio data.
It solves the problem that users are unable to communicate messages during meetings or have language dysfunction or are unable to speak, ensures that information can be exchanged in a timely manner, and improves the communication efficiency of users in different environments.
Smart Images

Figure CN120111020A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminal technology, and in particular to an instant messaging method, device and storage medium based on artificial intelligence (AI) subtitles. Background Art
[0002] With the rapid development of terminal technology, various instant messaging applications have emerged one after another. Most instant messaging applications not only support private communication, but also support remote office scenarios such as group communication and task management.
[0003] At present, when users use instant messaging applications (such as conference applications) to communicate online, they can not only send various types of data such as voice, video, pictures, text and emoticons, but also share files and mark files in real time, thus meeting different communication needs. However, when using instant messaging applications for voice calls or video calls, users may not be able to speak or have language barriers, resulting in the inability to exchange messages in a timely manner. Summary of the invention
[0004] The present application provides an instant messaging method, device and storage medium based on AI subtitles, which solves the technical problem that users cannot exchange messages in time when using instant messaging applications for voice calls or video calls.
[0005] In order to achieve the above objectives, this application adopts the following technical solutions:
[0006] In the first aspect, an embodiment of the present application provides an instant messaging method based on AI subtitles. The method can be applied to a first device. The first device may include a conference application, an AI subtitle application, an audio data entity (AudioFlinger), and a first audio interface. The first audio interface is an audio primary interface or an audio remote processing (remote_submix) interface of a hardware abstraction layer (HAL layer). The method may include: the first device displays the interface of the conference application and the window of the AI subtitle application; the first device receives the first text entered by the user in the window of the AI subtitle application; the first device writes the first audio data to the first audio interface through AudioFlinger, and the first audio data is obtained after the first text is converted from text to speech (text to speech, TTS); in response to the recording request of the conference application, the first device reads the first audio data from the first audio interface through AudioFlinger; the first device sends the first audio data to the second device through the conference application. Among them, the first device and the second device are terminal devices, such as mobile phones, PCs or Pads, and the window of the AI subtitle application can be a floating window.
[0007] In the above scheme, the first device can convert the text input by the user into audio data based on the AI subtitle application, and send the audio data to the second device through the instant messaging application. The user using the second device can hear the sound corresponding to the audio data, thereby solving the problem that the user is inconvenient to speak in the meeting or the user has language dysfunction and cannot speak, resulting in the inability to communicate messages.
[0008] In a possible implementation, the method may further include: the first device displays a send message in a window of the AI subtitle application, where the send message includes the first text.
[0009] In the above scheme, after the first device successfully sends the first audio data, the first device can display the first text in the window of the AI subtitle application, such as "OK, I understand", so that the user using the first device can know that the audio message corresponding to the first text "OK, I understand" has been successfully sent.
[0010] In one possible implementation, the first audio interface is an audio remote processing (remote_submix) interface, the first device also includes a microphone and a media recorder (MediaRecorder), and the microphone is in a recording disabled state. In response to a recording request of a conference application, the first device reads the first audio data from the first audio interface through AudioFlinger, which may include: in response to a recording request initiated by the conference application through the media recorder (MediaRecorder), the first device calls a recording thread through AudioFlinger to read the first audio data from the audio remote processing (remote_submix) interface.
[0011] In the above scheme, when the microphone is in a recording-prohibited state, there is no need to consider the problem that the audio remote processing (remote_submix) interface cannot provide audio writing and reading services for the microphone, and the audio remote processing (remote_submix) interface can be used as the first audio interface. In addition, the first device can convert the text input by the user into audio data based on the AI subtitle application, and mix the audio data with the sound collected by the microphone, and then send an audio message carrying the mixed audio to the second device. It not only solves the problem that the user is inconvenient to speak in the meeting and the user has a language barrier and cannot speak, resulting in the inability to interact with messages, but also enables other participants to hear the background sound of the mobile phone. It should be understood that the mixing of background sound makes the sound heard by other participants closer to the sound of the user's real environment.
[0012] In one possible implementation, the first audio interface is an audio primary interface, the first device also includes a microphone and a media recorder (MediaRecorder), and the microphone is in a recording-allowed state. The method may also include: the first device writes the second audio data to the audio primary interface, and the second audio data is the audio data collected by the microphone. Accordingly, in response to the recording request of the conference application, the first device reads the first audio data from the first audio interface through AudioFlinger, which may include: in response to the recording request initiated by the conference application through the media recorder (MediaRecorder), the first device calls the recording thread through AudioFlinger to read the mixed data of the first audio data and the second audio data from the audio primary interface. The first device sends the first audio data to the second device through the conference application, which may include: the first device sends the mixed data of the first audio data and the second audio data to the second device through the conference application.
[0013] In the above scheme, when the microphone is in the recording-allowed state, the performance attributes of the audio remote processing (remote_submix) interface determine that it cannot provide audio write and read services for the microphone, while the performance attributes of the audio primary (primary) interface determine that it can provide audio write and read services for the microphone. In this case, the first device can use the audio primary (primary) interface to provide audio write and read services for the first audio data and the second audio data.
[0014] In one possible implementation, the first device may further include a second audio interface, which is an audio primary interface or an audio remote processing (remote_submix) interface of a hardware abstraction layer. The method may further include: the first device receives the third audio data from the second device through the conference application; the first device writes the third audio data to the second audio interface through AudioFlinger; in response to the play request of the AI subtitle application, the first device reads the third audio data from the second audio interface through AudioFlinger; the first device displays the received message in the window of the AI subtitle application, and the received message may include a second text, which is obtained after the third audio data is automatically recognized by speech recognition.
[0015] In the above scheme, the first device can convert the audio data from the second device into text based on the AI subtitle application and display the text, so that the user using the mobile phone can see the text, solving the problem of the user's hearing loss causing the inability to hear the voice from the other user.
[0016] In a possible implementation, the second audio interface is an audio remote processing (remote_submix) interface, and the first device may further include a speaker, and the speaker is in a sound-disabled state.
[0017] In the above scheme, when the speaker is in a sound-disabled state, there is no need to consider the problem that the audio remote processing (remote_submix) interface cannot provide audio writing and reading services for the speaker, and the audio remote processing (remote_submix) interface can be used as the second audio interface.
[0018] In a possible implementation, the second audio interface is a primary audio interface, and the first device may further include a speaker, which is in a state where sound is allowed to be emitted. The method may further include: the first device reads third audio data from the primary audio interface; and the first device outputs sound corresponding to the third audio data through the speaker.
[0019] In the above scheme, when the speaker is in a state where sound is allowed, the performance attributes of the audio remote processing (remote_submix) interface determine that it cannot provide audio writing and reading services for the speaker, while the performance attributes of the audio primary interface determine that it can provide audio writing and reading services for the speaker, so the audio primary interface can be used as the second audio interface. In addition, the first device can also display text and output audio, so that the text can be viewed in a noisy environment and the sound can be listened to in a quiet environment, thereby increasing the diversity of the user's way of obtaining information.
[0020] In a possible implementation, the first device may further include an audio manager (AudioManager). The method may further include: based on the status of the microphone and speaker of the first device, the first device formulates an audio policy for AudioFlinger through the audio manager (AudioManager). Wherein, if the audio policy is the first audio policy, the first audio interface and the second audio interface are audio remote processing (remote_submix) interfaces; if the audio policy is the second audio policy, the first audio interface is an audio remote processing (remote_submix) interface, and the second audio interface is an audio primary interface; if the audio policy is the third audio policy, the first audio interface and the second audio interface are audio primary interfaces.
[0021] In the above solution, the AudioManager is responsible for the strategy selection of audio device switching. AudioFlinger is the executor of the audio system strategy, responsible for the management of audio stream devices and the processing and transmission of audio stream data. The AudioManager controls the audio paths of the microphone and the speaker by indicating different audio strategies to AudioFlinger, meeting the audio usage needs of users in different conference scenarios.
[0022] In a possible implementation, the first device may further include a media player (MediaPlayer); the first device writes the third audio data to the second audio interface through AudioFlinger, which may include: the first device calls AudioFlinger through the media player (MediaPlayer) to write the third audio data to the second audio interface.
[0023] In a possible implementation, the first device may further include a media player (MediaPlayer); the first device writes the first audio data to the first audio interface through AudioFlinger, which may include: the first device calls AudioFlinger through the media player (MediaPlayer) to write the first audio data to the first audio interface.
[0024] In the second aspect, an embodiment of the present application provides an instant messaging method based on AI subtitles. The method can be applied to a first device. The first device may include a conference application and an AI subtitle application, and the method may include: the first device displays the interface of the conference application; the first device responds to the user's first operation, and suspends and displays the window of the AI subtitle application on the interface of the conference application; the first device receives the first text entered by the user in the window of the AI subtitle application; the first device sends the first audio data to the second device through the conference application, and the first audio data is obtained after the first text is converted from text to speech; the first device displays a send message in the window of the AI subtitle application, and the send message may include the first text.
[0025] In the above scheme, the first device can convert the text input by the user into audio data based on the AI subtitle application, and send the audio data to the second device through the instant messaging application. The user using the second device can hear the sound corresponding to the audio data, thereby solving the problem that the user is inconvenient to speak in the meeting or the user has language dysfunction and cannot speak, resulting in the inability to communicate messages.
[0026] In a possible implementation, the first device may further include a microphone. The method may further include: the first device acquires second audio data, where the second audio data is audio data collected by the microphone. Accordingly, the first device sends the first audio data to the second device through the conference application, which may include: the first device sends mixed audio data of the first audio data and the second audio data to the second device through the conference application.
[0027] In the above scheme, the first device can convert the text input by the user into audio data based on the AI subtitle application, mix the audio data with the sound collected by the microphone, and then send an audio message carrying the mixed audio to the second device. It not only solves the problem that users are inconvenient to speak in the meeting and users have language barriers and cannot speak, resulting in the inability to interact with messages, but also enables other participants to hear the background sound of the mobile phone. It should be understood that the mixing of background sound makes the sound heard by other participants closer to the sound of the user's real environment.
[0028] In a possible implementation, the method may also include: the first device receives the third audio data from the second device through the conference application; the first device displays the received message in the window of the AI subtitle application, and the received message may include a second text, and the second text is obtained after the third audio data is automatically recognized.
[0029] In the above scheme, the first device can convert the audio data from the second device into text based on the AI subtitle application and display the text, so that the user using the mobile phone can see the text, solving the problem of the user's hearing loss causing the inability to hear the voice from the other user.
[0030] In a third aspect, the present application provides a device, which includes a unit for executing the method in the first aspect or the second aspect. The device may correspond to executing the instant messaging method based on AI subtitles described in the first aspect or the second aspect. For the relevant description of the unit in the device, please refer to the description of the first aspect or the second aspect, which will not be repeated here for the sake of brevity.
[0031] The method described in the first aspect or the second aspect may be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a processing module or unit, a display module or unit, etc.
[0032] In a fourth aspect, the present application provides a terminal device, which includes a memory and one or more processors. The memory is used to store computer program code, and the computer program code includes computer instructions. When the computer instructions are called by the processor, the terminal device executes the instant messaging method based on AI subtitles provided in any one of the first aspect or the second aspect.
[0033] In a fifth aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium includes computer instructions. When the computer instructions are executed on a terminal device, the terminal device executes an instant messaging method based on AI subtitles provided in any possible implementation of the first aspect or the second aspect.
[0034] In a sixth aspect, the present application provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute an instant messaging method based on AI subtitles as provided in any possible implementation of the first aspect or the second aspect.
[0035] In a seventh aspect, the present application provides a chip system. The chip system includes one or more interface circuits and one or more processors. The interface circuit and the processor are interconnected by lines. The chip system can be applied to a terminal device including a communication module and a memory. The interface circuit is used to receive a signal from the memory of the terminal device and send the received signal to the processor, the signal including a computer instruction stored in the memory. When the processor calls the computer instruction, the terminal device can execute an instant messaging method based on AI subtitles as provided in any possible implementation of the first aspect or the second aspect.
[0036] It can be understood that the beneficial effects that can be achieved by the above-mentioned device of the third aspect, the terminal device of the fourth aspect, the computer-readable storage medium of the fifth aspect, the computer program product of the sixth aspect and the chip system of the seventh aspect can be referred to as the beneficial effects in any possible implementation method of the first aspect or the second aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram of the interface of AI subtitles in a video scenario provided by an embodiment of the present application;
[0038] Figure 2 A schematic diagram of the interface of ordinary subtitles in a conference scenario provided by an embodiment of the present application;
[0039] Figure 3 A schematic diagram of an interface for enabling AI subtitles provided in an embodiment of the present application;
[0040] Figure 4 Schematic diagram of three other interfaces for opening AI subtitles provided in the embodiments of the present application;
[0041] Figure 5 One of the schematic diagrams of the scenario using AI subtitles provided in the embodiment of the present application;
[0042] Figure 6 The second schematic diagram of a scenario using AI subtitles provided in an embodiment of the present application;
[0043] Figure 7 The third schematic diagram of a scenario using AI subtitles provided in an embodiment of the present application;
[0044] Figure 8 The fourth schematic diagram of a scenario using AI subtitles provided in an embodiment of the present application;
[0045] Fig. 9 The fifth schematic diagram of a scenario using AI subtitles provided in an embodiment of the present application;
[0046] Fig.10 The sixth schematic diagram of a scenario using AI subtitles provided in an embodiment of the present application;
[0047] Fig.11 The seventh schematic diagram of a scenario using AI subtitles provided in an embodiment of the present application;
[0048] Fig.12 The eighth schematic diagram of a scenario using AI subtitles provided in an embodiment of the present application;
[0049] Fig.13 The ninth schematic diagram of a scenario using AI subtitles provided in an embodiment of the present application;
[0050] Fig.14 A schematic diagram of the system architecture of a terminal device provided in an embodiment of the present application;
[0051] Fig.15 A module interaction diagram based on AI subtitles provided in an embodiment of the present application;
[0052] Fig.16 For Fig.15 The corresponding specific method flow chart;
[0053] Fig.17 Another module interaction diagram based on AI subtitles provided in an embodiment of the present application;
[0054] Fig.18 For Fig.17 The corresponding specific method flow chart;
[0055] Fig.19 Another module interaction diagram based on AI subtitles provided in an embodiment of the present application;
[0056] Fig. 20For Fig.19 The corresponding specific method flow chart;
[0057] Fig.21 A module interaction diagram corresponding to the parameter distribution strategy for downlink data provided in an embodiment of the present application;
[0058] Fig. 22 For Fig.21 The corresponding specific method flow chart;
[0059] Fig.23 A module interaction diagram corresponding to the parameter distribution strategy for uplink data provided in an embodiment of the present application;
[0060] Fig.24 For Fig.23 The corresponding specific method flow chart;
[0061] Fig.25 A schematic diagram of the hardware structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.
[0063] In the description of this application, unless otherwise specified, " / " means or, for example, A / B can mean A or B. In the description of this application, "and / or" is only a way to describe the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0064] The terms "first" and "second" in the specification and claims of this application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects. For example, a first operation and a second operation are used to distinguish different operations, rather than to describe a specific order of operations. In the embodiments of this application, "plurality" refers to two or more.
[0065] At present, some mobile phones have launched AI subtitles. When playing online videos on mobile phones, the AI subtitles application can convert the audio stream corresponding to the online video into text based on AI technology and display it on the screen in the form of subtitles. Figure 1As shown in (a) of FIG. 1 , if the user of the mobile phone has hearing impairment, then when the mobile phone plays the video advertisement screen 01, the user cannot hear the sound, causing inconvenience to the user's life. In this case, the user can use a quick operation to trigger the mobile phone to pull up the video advertisement screen 01. Figure 1 The floating window 02 of AI subtitles shown in (b) of FIG. The floating window 02 of AI subtitles can be used to display the text obtained after the audio stream is converted. In this way, even a person with hearing impairment may obtain the specific information of the video advertisement through the floating window 02 of AI subtitles. However, the existing AI subtitle function is limited to video applications and only supports the conversion of audio streams into text.
[0066] In addition to video applications, there is also a demand for using AI subtitles in instant messaging applications. Among them, instant messaging applications are a terminal service that allows two or more people to instantly transmit text messages, picture files, voice exchanges, and video exchanges based on the network. Instant messaging applications can be divided into enterprise instant messaging applications and website instant messaging applications according to their use. For example, instant messaging applications can be conference applications or voice applications. When multiple devices communicate online based on instant messaging applications (such as conference applications), multiple devices can send each other various types of data such as online voice, online video, video files, picture files, text, and emoticons, and can also share files and markup files in real time, thereby meeting the different communication needs of users.
[0067] For example, Figure 2 As shown in (a) of FIG. 1 , the user can click on the icon 03 of the conference application on the desktop. The conference application can be a third-party application or a system application. In response to the user's click operation, the mobile phone displays the following Figure 2 The conference homepage shown in (b) in FIG. 1 includes multiple conference functions such as Join Conference 04, Quick Conference, and Screen Sharing. For example, a user clicks Join Conference 04, and the phone updates to display the following Figure 2The interface for joining a meeting is shown in (c) of FIG. The interface for joining a meeting includes a meeting ID 05, a participant name 06, a speaker control 07, a microphone control 08, a camera control 09, a joining meeting control 10, and the like. Among them, the meeting ID 05 is used to input the account of this meeting, for example, the meeting ID 05 is "11111111111"; the participant name 06 is used to set the name of the participant, for example, the participant name 06 is "Wang Yiyi"; the speaker control 07 is used to turn the speaker on or off, for example, the speaker is turned off when the speaker control 07 slides to the left, or the speaker is turned on when the speaker control 07 slides to the right; the microphone control 08 is used to turn the microphone on or off, for example, the microphone is turned off when the microphone control 08 slides to the left, or the microphone is turned on when the microphone control 08 slides to the right; the camera control 09 is used to turn the camera on or off, for example, the camera is turned on when the camera control 09 slides to the right, or the camera is turned off when the camera control 09 slides to the left; the joining meeting control 10 is used to confirm joining the meeting after the user completes various settings. In response to the user clicking the join conference control 10, the mobile phone displays the following Figure 2 The conference interface 11 shown in (d) in FIG. The conference interface 11 includes the avatars of the participants, unmute controls, video start controls, screen sharing controls, member management controls, conference end controls, and subtitle controls 12. It should be noted that the subtitle control 12 only provides a functional entry for users to input chat content, and does not have AI subtitle functions, such as the function of converting audio streams into text. Therefore, it is called ordinary subtitles, not AI subtitles. If the user clicks the subtitle control 12, the phone will display the following: Figure 2 The chat window 13 shown in (e) in FIG. 1 is shown. The user can input the text "No problem" 14 on the virtual keyboard in the chat window 13 and send a text message "No problem" to other participants. If the user closes the chat window 13, the mobile phone can display the message "No problem" 15 in the form of subtitles in the conference interface 11, so that each participant can see this subtitle message.
[0068] However, when multiple users use instant messaging applications to make voice calls or video calls, users may not be able to exchange messages in time in any of the following scenarios: Scenario 1, one user is in a meeting and it is inconvenient for the user to speak; Scenario 2, one user has a language disorder and the user cannot speak; Scenario 3, one user has hearing loss and the user cannot hear the voice from the other user; Scenario 4, the current environment of one user is too noisy and the user cannot hear the voice from the other user clearly. It can be seen that there is also a demand for using AI subtitles in instant messaging applications, but the prior art has not yet applied AI subtitles to instant messaging applications. In view of these problems, the present application provides an instant messaging method based on AI subtitles. The method can be applied to a terminal device, which can provide AI subtitle services when the user uses an instant messaging application. When the terminal device has established a connection with the peer device through an instant messaging application, the terminal device can convert the text input by the user into audio data based on the AI subtitle application, and send the audio data to the peer device, so that the user using the peer device can hear the sound corresponding to the audio data, thereby solving the problems in the above-mentioned scenarios 1 and 2. In addition, the terminal device can also convert the audio data from the other device into text based on the AI subtitle application and display the text, so that the user using the terminal device can see the text, thereby solving the problems existing in the above-mentioned scenarios 3 and 4.
[0069] In some embodiments, the terminal device is also referred to as a terminal or user equipment (UE). For example, the terminal device may be a personal computer (PC), a mobile phone, a smart screen, a smart TV, a tablet computer (Pad), a wearable device, a computer with wireless transceiver function, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, or a wireless terminal in smart home, etc., or may be other devices with display and audio playback functions.
[0070] The following takes the mobile phone as an example. Figures 3 to 13The application scenarios of the instant messaging method based on AI subtitles provided in this application are exemplified.
[0071] Scenario Example 1
[0072] In some embodiments, the user can enable the AI subtitle function through the settings application of the mobile phone.
[0073] For example, Figure 3 As shown in (a) in FIG. 1 , the desktop may include icons of applications such as Honor Video, Sports Health, Weather, Browser, Settings, and Smart Life. The user can click on any icon to trigger the mobile phone to display the operation interface of the corresponding application. For example, the user can click on Settings 16.
[0074] After receiving the user's click operation on setting 16, the mobile phone displays the following Figure 3 The setting interface shown in (b) in the figure. The setting interface includes options such as Bluetooth, wireless LAN, phone, flight mode, healthy use of mobile phone, smart assistant 17 and auxiliary functions, and each option corresponds to a system-level service. The user can trigger the mobile phone to display the service setting interface corresponding to this option by clicking on an option. For example, the user can click on smart assistant 17.
[0075] After receiving the user's click operation on the smart assistant 17, the mobile phone displays the following Figure 3 The smart assistant interface shown in (c) in the figure. The smart assistant interface includes YOYO suggestions, negative one screen, Honor search, and Honor vision. In addition, the smart assistant interface can also include other smart functions, such as smart voice 18, smart perception, smart screen recognition, etc. Users can trigger the phone to display the setting interface corresponding to a smart function by clicking on a smart function. For example, the user can click on smart voice 18.
[0076] After receiving the user's click operation on the smart voice 18, the mobile phone displays the following Figure 3 The intelligent voice interface shown in (d) in FIG. The intelligent voice interface provides voice services such as AI subtitles 19 and voice assistants. Among them, AI subtitles 19 is used to set AI subtitles, and voice assistants are used to set voice assistants. For example, a user can click on AI subtitles 19.
[0077] After receiving the user's click operation on AI subtitles 19, the mobile phone displays the following Figure 3The AI subtitle interface is shown in (e) in the figure. The AI subtitle interface includes an AI subtitle control 20. When the AI subtitle control 20 slides to the left, the phone turns off the AI subtitle function. When the AI subtitle control 20 slides to the right, the phone turns on the AI subtitle function. The AI subtitle interface also includes options such as writing mode, font size, desktop shortcut, and background opacity. Among them, the desktop shortcut includes an add control 21. When the user clicks the add control 21, the phone will create a subtitle on the desktop or the negative one screen. Figure 4 The shortcut icon 29 shown in (d) in the figure can be used to trigger the mobile phone to quickly pull up the floating window of AI subtitles by operating the shortcut icon 29.
[0078] like Figure 3 As shown in (f) in the figure, when the AI subtitle control 20 slides to the right, the mobile phone turns on the AI subtitle function and displays a floating window 22 of the AI subtitle. The bottom area of the floating window 22 includes an input box 23 for editing text, a preset word control 24 for quick reply, a mute control 25 for turning the microphone on or off, and a send control 26 for confirming the text in the input box 23. The top area of the floating window 22 includes a close control 27 for canceling the floating window 22 and more options. It should be understood that the floating window 22 is only an exemplary description and does not limit the present application. In actual implementation, the floating window 22 may include more or fewer controls, and the position, size and style of each control in the floating window 22 may also be adjusted.
[0079] In some embodiments, when the user triggers the mobile phone to exit the settings interface and switch to other interfaces, such as switching from the settings interface to the interface of an instant messaging application, a video application, or a call application, the floating window 22 can be displayed superimposed on the switched application interface, that is, the mobile phone always keeps displaying the floating window 22 until the user clicks to close the control 27.
[0080] In other embodiments, the user can also change the position of the floating window 22 on the screen, such as moving the floating window 22 to the top or bottom of the screen, by dragging the floating window 22. In addition, the user can also change the size of the floating window 22 by operating the frame of the floating window 22, such as increasing or decreasing the size of the floating window 22.
[0081] In other embodiments, the user can also change the transparency of the floating window 22 by operating the background opacity option in the AI subtitle interface. For example, when the background opacity is set to 100%, the background of the floating window 22 is completely visible, so that the application interface below the floating window 22 will not be blocked. For another example, when the background opacity is set to 0%, the background is completely invisible, and only the text is retained above the application interface, which basically does not block the application interface below the floating window 22.
[0082] The above embodiment is described by taking the example of a user triggering the mobile phone to turn on the AI subtitle function by operating the setting application, which does not limit the present application. In actual implementation, the user can also trigger the mobile phone to quickly pull up the floating window 22 of the AI subtitle in other ways. As an example, Figure 4 As shown in (a) of FIG. 1 , when the mobile phone displays the conference interface 11, the user can slide his finger downward from the top of the screen. In response to the sliding operation, the mobile phone displays the following Figure 4 The control center list shown in (b) in FIG. The control center list includes WLAN controls, Bluetooth controls, mobile data controls, AI subtitle controls 28, flight mode controls, and brightness adjustment controls. If the user clicks on the AI subtitle control 28, then Figure 4 As shown in (c) in FIG. 1 , the mobile phone displays a floating window 22 of AI subtitles above the conference interface 11. As another example, Figure 4 As shown in (d) in FIG. 1 , the user can also trigger the mobile phone to quickly pull up the floating window 22 of the AI subtitles by clicking the shortcut icon 29 on the desktop. As another example, Figure 4 As shown in (e) in the figure, the user can also use the voice assistant to say "turn on AI subtitles" to the mobile phone, triggering the mobile phone to quickly pull up the floating window 22 of AI subtitles.
[0083] In the above embodiment, the user can pull up the floating window of AI subtitles by setting applications, shortcut icons on the desktop, AI subtitle controls in the control center, or voice assistants. In addition to the above methods, the mobile phone may also pull up the floating window of AI subtitles in other ways. In this way, during the meeting, AI subtitles can provide users with AI subtitle functions, such as converting text messages to be sent into audio messages, or converting received audio messages into text messages.
[0084] Scenario Example 2
[0085] In some embodiments, after the mobile phone displays a floating window 22 of AI subtitles, the user can manually enter text in the floating window 22, and the mobile phone converts the text into audio, and then sends an audio message carrying the audio to the devices used by other participants.
[0086] For example, during a meeting, when it is inconvenient for a user to speak, Figure 5 As shown in (a) of FIG. 1 , the user can click the input box 23 of the floating window 22. In response to the click operation, the mobile phone displays the following in the area below the floating window 22: Figure 5 The virtual keyboard 30 shown in (b) in FIG. Figure 5As shown in (c) in the Chinese mode, the user can select and input the virtual keys of the virtual keyboard 30, so that the mobile phone displays the alphabetic pinyin corresponding to the virtual keys selected by the user in the input box 23, such as "haode". Figure 5 As shown in (d) in FIG. 1 , after editing the text “OK, I understand”, the user clicks the send control 26. In response to the click operation on the send control 26, the mobile phone converts the text “OK, I understand” into audio and sends an audio message carrying the audio to the devices used by other participants. Figure 5 As shown in (e), after the audio message is successfully sent, the mobile phone can display the text "OK, I understand" 32 in the floating window 22, so that the user using the mobile phone can know that the audio message corresponding to the text "OK, I understand" has been successfully sent. On the other hand, after the devices used by other participants successfully receive the audio message, these devices can play the audio through speakers, receivers or headphones, or convert the audio into the text "OK, I understand" based on AI subtitles and then display it.
[0087] In the above embodiment, the mobile phone can convert the text input by the user into audio data based on the AI subtitle application, and send the audio data to the peer device, so that the user using the peer device can listen to the audio corresponding to the audio data, thereby solving the problem that the user is inconvenient to speak in the meeting or the user has language dysfunction and cannot speak, resulting in the inability to communicate messages.
[0088] Scenario Example 3
[0089] The above scenario example 2 is explained by converting the text input by the user in the floating window into audio, and directly sending an audio message carrying the audio to the devices used by other participants. In some embodiments, the mobile phone can convert the text input by the user in the floating window into audio, mix it with the sound collected by the microphone, and then send an audio message carrying the mixed audio to the devices used by other participants.
[0090] For example, Figure 6 As shown in (a) in the figure, the mute control 25 in the floating window 22 is in mute mode, indicating that the microphone is in a recording disabled state, that is, the microphone will not collect sound. If the user wants other participants to hear the background sound of their environment, the user can click the mute control 25; or, when multiple users use the same mobile phone to participate in the meeting but one user is not convenient to speak, the user who is not convenient to speak can click the mute control 25. After receiving the user's click operation on the mute control 25, the mute control 25 is changed from Figure 6 The mute form shown in (a) is changed to Figure 6The non-mute mode shown in (b) of FIG. 1 is shown in FIG. 1 , and the microphone of the mobile phone is turned on. As an example, Figure 6 As shown in (b) in the figure, after turning on the microphone, the mobile phone may also display a prompt message "Microphone turned on" on the screen; as another example, the mobile phone may not display the prompt message.
[0091] like Figure 6 As shown in (e) in FIG. 1 , the microphone can record the local background sound of the mobile phone to obtain audio 2. At the same time, Figure 6 As shown in (d) in FIG. 2 , the user can enter the text “OK, I understand” through the virtual keyboard, and then click the send control 26. The mobile phone can first convert the text “OK, I understand” into audio 1, and then mix audio 1 and audio 2 to obtain audio 3. Then, the mobile phone sends an audio message carrying audio 3 to the devices used by other participants. On the one hand, Figure 6 As shown in (f) in the figure, after successfully sending the audio message carrying audio 3, the mobile phone can display the text "OK, I got it" 34 in the floating window 22, so that the user using the mobile phone can know that audio 3 has been successfully sent. On the other hand, after the devices used by other participants successfully receive the audio message carrying audio 3, these devices can play audio 3 through speakers, receivers or headphones, or extract "OK, I got it" from audio 3 based on AI subtitles and display it in text form.
[0092] In other embodiments, the conference interface 11 also includes a mute control 33, and the user can also trigger turning the microphone on or off by operating the mute control 33 of the conference interface 11, thereby changing the audio mode of the AI subtitle application.
[0093] For example, Figure 6 As shown in (c) of FIG. 1 , the mute control 33 in the conference interface 11 is in mute mode, indicating that the microphone is in a recording-disabled state. If the user wants other participants to hear the background sound of their environment, the user can click the mute control 33. After receiving the user's click operation on the mute control 33, the mute control 33 is changed from Figure 6 The mute form shown in (c) is changed to Figure 6 Referring to the description of the above embodiment, the mobile phone can convert the text input by the user through the virtual keyboard into audio, mix it with the sound collected by the microphone, and then send an audio message carrying the mixed audio to the devices used by other participants, which will not be repeated here.
[0094] In other embodiments, referring to the description of the above embodiments, the conference interface 11 and the floating window 22 are respectively provided with a mute control, and the mute controls include a mute form and a non-mute form. In order to avoid logical conflicts caused by user operations, before the user triggers the mobile phone to pull up the floating window 22 by setting applications, desktop shortcut icons, AI subtitle controls in the control center, or voice assistants, the mobile phone can detect the status of the microphone and set the display form of the mute control 25 in the floating window 22 according to the status of the microphone. Specifically, if the microphone is in the off state, the mute control 25 in the floating window 22 is set as follows: Figure 6 The mute form shown in (a) indicates that the microphone is in a recording disabled state; if the microphone is in an on state, the mute control 25 in the floating window 22 is set to Figure 6 The non-mute form shown in (b) indicates that the microphone is in a recording-enabled state. Figure 2 As shown in (c), the microphone has been set before entering the meeting, so after entering the meeting, the display form of the mute control 33 is consistent with the state of the microphone. Finally, the display forms of the mute control 25 and the mute control 33 are consistent with the state of the microphone.
[0095] In other embodiments, after pulling up the floating window 22, if the mobile phone receives an operation of the user on the mute control 33 of the conference interface 11 or receives an operation of the user on the mute control 25 of the floating window 22, then the mobile phone can notify the conference application and the AI subtitle application to change the display form of the mute control, such as setting both mute controls to a non-mute form, so that the display of the mute control 33 of the conference interface 11 or the mute control 25 of the floating window 22 remains consistent, so that the user can understand the actual status of the microphone at the current moment.
[0096] In other embodiments, after the floating window 22 is pulled up, the mobile phone can notify the conference application to set the mute control 33 of the conference interface 11 to an unavailable state (i.e., the user is not allowed to turn on or off the microphone by operating the mute control 33), and only the mute control 25 of the floating window 22 is in an available state (i.e., the user is allowed to turn on or off the microphone by operating the mute control 25). In this case, the user can only operate the mute control 25 of the floating window 22, but cannot operate the mute control 33 of the conference interface 11, thereby avoiding logical conflicts caused by user operations.
[0097] In the above embodiment, the mobile phone can convert the text input by the user into audio data based on the AI subtitle application, and mix the audio data with the sound collected by the microphone, and then send an audio message carrying the mixed audio to the devices used by other participants. It not only solves the problem that users are inconvenient to speak in the meeting, and users have language barriers and cannot speak, resulting in the inability to interact with messages, but also enables other participants to hear the background sound of the mobile phone. The mixing of background sound makes the sound heard by other participants closer to the sound of the user's real environment.
[0098] Scenario Example 4
[0099] In some embodiments, after the mobile phone displays the floating window 22 of AI subtitles, if the mobile phone receives an audio message sent by a device used by other participants, the mobile phone can convert the audio message into a text message through the AI subtitles application and display the text message in the floating window 22.
[0100] For example, Figure 7 As shown in (a), the upper and lower frames of mobile phone 1 are provided with speakers 1 and 2, and conference interface 11 also includes a playback control 35, which is in the form of a handset, indicating that speakers 1 and 2 are in a playback-prohibited state. If mobile phone 1 receives audio data 1 from mobile phone 2, mobile phone 1 converts audio data 1 into text "Departure at 9 o'clock tomorrow" through the AI subtitle application, and displays the text "Departure at 9 o'clock tomorrow" in the floating window 22. Since speakers 1 and 2 are in a playback-prohibited state at this moment, mobile phone 1 cannot play audio data 1 through speakers 1 and 2.
[0101] like Figure 7 As shown in (b) of FIG. 1 , the user can click the playback control 35 of the conference interface 11. In response to the click operation, the mobile phone 1 changes the playback control 35 to the following: Figure 7 The speaker format shown in (c) in FIG. 1 is displayed, and a prompt message “The sound will be played from the speaker” 37 is displayed on the screen. Through this prompt message, the user can know that mobile phone 1 has successfully turned on speaker 1 and speaker 2. After turning on speaker 1 and speaker 2, if mobile phone 1 receives audio data 2 from mobile phone 2, then mobile phone 1 converts audio data 2 into text “Meet at the company gate” through the AI subtitle application, and displays it as Figure 7 The floating window 22 shown in (d) in FIG. 3 displays the text "Meet at the company gate" 38. In addition, since the speaker 1 and the speaker 2 are in the state of allowing the sound to be played, Figure 7 As shown in (d) in FIG. 1 , mobile phone 1 can also play audio “Meet at the company gate” through speaker 1 and speaker 2 .
[0102] It should be noted that the above Figure 7 The example of turning on the speaker by user operation in a meeting is used for illustration, which does not limit the present application. Figure 2 In (c), if the user slides the microphone control 08 to the right in the conference joining interface, the microphone is turned on when the mobile phone 1 enters the conference interface 11, and the playback control 35 is as follows: Figure 7 In the speaker form shown in (c), in this case, the mobile phone 1 can not only display the text corresponding to the received audio data, but also play the audio data through the speaker.
[0103] In some other embodiments, the user can also click Figure 7 The playback control 35 shown in (c) in FIG. 1 is thus changed from the speaker form to the playback control 35 shown in FIG. Figure 7 In this case, the speaker is prohibited from outputting audio data from the mobile phone 2.
[0104] In the above embodiment, the mobile phone can convert the audio data from the peer device into text based on the AI subtitle application and display the text, so that the user using the mobile phone can see the text, solving the problem that the user's hearing loss causes the user to be unable to hear the voice from the other user. In addition, the mobile phone can also display text and output audio, so that the text can be viewed in a noisy environment and the sound can be heard in a quiet environment, thereby increasing the diversity of the user's way of obtaining information.
[0105] Scenario Example 5
[0106] The above scenario example 4 is described by taking the operation of the playback control 35 of the conference interface 11 to control the speaker to be turned on or off as an example. In some embodiments, the mobile phone can also control the speaker to be turned on or off by pressing the "volume +" button and the "volume -" button set on the side frame.
[0107] As an example, Figure 8 As shown in (a) of FIG. 1 , the user can press the “volume +” button 39. In response to the user pressing the “volume +” button 39, if the speaker 1 and the speaker 2 are not turned on, the mobile phone 1 turns on the speaker 1 and the speaker 2, and displays the following on the screen: Figure 8 The prompt message "The sound will be played from the speaker" 37 shown in (b) of the figure; if speaker 1 and speaker 2 are already turned on, mobile phone 1 increases the volume of speaker 1 and speaker 2. When mobile phone 1 has turned on speaker 1 and speaker 2, if mobile phone 1 receives audio data 2 from mobile phone 2, then mobile phone 1 converts audio data 2 into text "Meet at the company gate" through the AI subtitle application, and Figure 8As shown in (c) in the floating window 22, the text "Gather at the company gate" 38 is displayed. Since the speaker 1 and the speaker 2 are in the state of allowing the sound to be played at this moment, the mobile phone 1 can also play the audio "Gather at the company gate" through the speaker 1 and the speaker 2.
[0108] As another example, Figure 8 As shown in (d) of FIG. 4 , the user may press the “volume-” button 40. In response to the user pressing the “volume-” button 40, if the speaker 1 and the speaker 2 are turned on, the mobile phone 1 turns off the speaker 1 and the speaker 2, and displays the following: Figure 8 The prompt message "The sound will be turned off" 41 is shown in (e) of FIG. 1. When mobile phone 1 has turned off speaker 1 and speaker 2, if mobile phone 1 receives audio data 2 from mobile phone 2, mobile phone 1 converts audio data 2 into text "Meet at the company gate" through the AI subtitle application, and Figure 8 As shown in (f) in the figure, the text "Gather at the company gate" 38 is displayed in the floating window 22. Since the speaker 1 and the speaker 2 are in the prohibited playback state at this moment, the mobile phone 1 does not need to play any audio through the speaker 1 and the speaker 2.
[0109] In the above embodiment, when the user presses the "volume +" button 39, it means that the user may want to increase the volume. At this time, the speaker can be turned on, and audio from other devices can be input through the speaker, and text corresponding to the audio data can be displayed through the floating window 22; when the user presses the "volume -" button 40, it means that the user may want to reduce the volume. At this time, the speaker can be turned off, and audio from other devices can be prohibited from being input through the speaker, and only the text corresponding to the audio data can be displayed through the floating window 22.
[0110] Scenario Example 6
[0111] In some embodiments, the floating window 22 may also include a playback control, and the user may control the speaker to be turned on or off by operating the playback control.
[0112] As an example, Fig. 9 As shown in (a) of FIG. 1 , the playback control 42 is in the form of a prompt opening, indicating that the speaker is in a playback-disabled state, that is, the speaker will not output sound. The user can press the playback control 42. In response to the user's pressing operation on the playback control 42, the mobile phone 1 turns on the speaker 1 and the speaker 2, and the screen displays the following: Fig. 9The prompt information "The sound will be played from the speaker" 37 shown in (b) of the figure is displayed, and the playback control 42 is changed to the prompt off form, indicating that the speaker is in the state of allowing playback, that is, the speaker can output sound. When the mobile phone 1 has turned on 1 and the speaker 2, if the mobile phone 1 receives the audio data 2 from the mobile phone 2, then the mobile phone 1 converts the audio data 2 into the text "Meet at the company gate" through the AI subtitle application, and Fig. 9 As shown in (c) in the floating window 22, the text "Gather at the company gate" 38 is displayed. Since the speaker 1 and the speaker 2 are in the state of allowing the sound to be played at this moment, the mobile phone 1 can also play the audio "Gather at the company gate" through the speaker 1 and the speaker 2.
[0113] As another example, Fig. 9 As shown in (d) of FIG. 1 , the playback control 42 is in the form of a prompt off, indicating that the speaker is in a state of allowing playback, that is, the speaker can output sound. The user can press the playback control 42. In response to the user's pressing operation on the playback control 42, the mobile phone 1 turns off the speaker 1 and the speaker 2, and displays the following on the screen: Fig. 9 The prompt message "The sound will be turned off" 41 shown in (e) of the figure is displayed, and the playback control 42 is changed to the prompt on form, indicating that the speaker is in a playback prohibited state, that is, the speaker will not output sound. When the mobile phone 1 has turned off 1 and the speaker 2, if the mobile phone 1 receives the audio data 2 from the mobile phone 2, then the mobile phone 1 converts the audio data 2 into the text "Meet at the company gate" through the AI subtitle application, and Fig. 9 As shown in (f) in the figure, the text "Gather at the company gate" 38 is displayed in the floating window 22. Since the speaker 1 and the speaker 2 are in the prohibited playback state at this moment, the mobile phone 1 does not need to play any audio through the speaker 1 and the speaker 2.
[0114] In the above embodiment, the floating window 22 may also include a playback control 42. The user may trigger the mobile phone to turn on or off the speaker by operating the playback control 42. Specifically, when the speaker is turned on, the mobile phone can not only input audio from other devices through the speaker, but also display text corresponding to the audio data through the floating window 22; when the speaker is turned off, the mobile phone may prohibit the output of audio from other devices through the speaker, and only display text corresponding to the audio data through the floating window 22.
[0115] It should be noted that the above-mentioned scenario examples 4, 5 and 6 are only exemplary descriptions. In actual implementation, the speaker can be turned on or off in other ways. As another example, the mobile phone 1 can also automatically turn the speaker on or off according to the current volume. For example, when the mobile phone 1 detects that the current system volume is greater than or equal to the preset volume, the speaker is turned on. In this way, the mobile phone can not only input audio from other devices through the speaker, but also display text corresponding to the audio data through the floating window 22; when the mobile phone 1 detects that the current system volume is less than the preset volume, the speaker is turned off. In this way, the mobile phone is prohibited from outputting audio from other devices through the speaker, and only displays text corresponding to the audio data through the floating window 22.
[0116] Scenario Example 7
[0117] In some embodiments, the floating window 22 may further include a preset word control 24. The user may quickly reply to other devices with preset words by operating the preset word control 24.
[0118] For example, Fig.10 As shown in (a) of FIG. 1 , the floating window 22 may also include a preset word control 24. During the conference, the user may click on the preset word control 24. In response to the click operation on the preset word control 24, the mobile phone displays the following Fig.10 The preset word box shown in (b) in FIG. The preset word box includes multiple preset words, such as "I am busy now, reply later", "No problem", "Under urgent processing", etc. For example, the user selected "No problem" based on the meeting scene at that time 43. Fig.10 As shown in (c) in FIG. 1 , the mobile phone displays the text “No problem” in the input box 23. Fig.10 As shown in (d) in FIG. 2 , after the user confirms that the text in the input box 23 is correct, he clicks the send control 26. In response to the click operation of the send control 26, the mobile phone converts the text "no problem" into audio and sends an audio message carrying the audio to the devices used by other participants. Fig.10 As shown in (e) in FIG. 1 , after the audio message is successfully sent, the mobile phone can display the text "No problem" 44 in the floating window 22, so that the user can know that the audio message corresponding to the text "OK, I understand" has been successfully sent. On the other hand, after the devices used by other participants successfully receive the audio message, these devices can play the audio through speakers, receivers or headphones, or convert the audio into the text "No problem" based on the AI subtitle application and then display it.
[0119] As an example, a mobile phone can be pre-set with a pre-set word library, which contains a large number of pre-set words and pre-set sentences. These pre-set words and pre-set sentences can be pre-set by the user, pushed by the server, or built into the mobile phone system. In a meeting scenario, the AI subtitle application can perform semantic analysis on the received audio data and display multiple pre-set words corresponding to the semantic analysis results in the pre-set word box.
[0120] In the above embodiment, since the AI subtitle application provides the pre-set word control 24 in the floating window 22, the user can quickly reply to other devices with pre-set words by operating the pre-set word control 24, thus solving the problem of low efficiency in inputting text through the virtual keyboard.
[0121] Scene Example 8
[0122] In some embodiments, the AI subtitle application also provides a translation function.
[0123] Exemplarily, as Fig.11 shown in (a) of, the floating window 22 further includes a more options 45. When the user wants to use more functions provided by the AI subtitle application, the user can click on the more options 45. In response to the click operation on the more options 45, as Fig.11 shown in (b) of, the mobile phone displays controls such as font, quick reply, and language type 46. For example, the user can click on the language type 46, and the mobile phone displays multiple language options as Fig.11 shown in (c) of, such as Chinese-English 47, Chinese-German, Chinese-French, Chinese-Russian, Chinese-Korean, and Chinese-Japanese. If the user uses Chinese while other participants use English, then the user can select Chinese-English 47. As Fig.11 shown in (d) of, the user can enter the Chinese word "很好" in the input box 23 and then click on the send control 26. The AI subtitle application can translate the Chinese word "很好" into the English word "excellent", convert the English word "excellent" into audio, and then send a message carrying the audio "excellent" to the devices used by other participants. On the one hand, as Fig.11 shown in (e) of, after successfully sending the audio message, the mobile phone can display the text "很好" 48 in the floating window 22, so that the user can know that the audio message corresponding to the text "很好" has been successfully sent. On the other hand, after the devices used by other participants successfully receive the audio message, these devices can play the audio "excellent" through the speaker, earpiece, or headphones.
[0124] In other embodiments, if the devices used by other participants send audio in the first language to the mobile phone, the mobile phone can first convert the audio in the first language into text in the first language, then convert the text in the first language into text in the second language, and then display the text in the second language in the floating window 22.
[0125] As an example, suppose the language type selected by the user is "Chinese and English", "Chinese" represents the language used by the mobile phone user, and "English" represents the language used by other participants. Fig.12 As shown, mobile phone 2 sends an audio message "After work dining together" to mobile phone 1. Mobile phone 1 may first convert the audio message "After work dining together" into text "After work dining together", then convert the text "After work dining together" into text "After get off work dining together", and then display the text "After get off work dining together" 49 in floating window 22.
[0126] As another example, suppose the language type selected by the user is "English-Chinese", "English" represents the language used by the mobile phone user, and "Chinese" represents the language used by other participants. Fig.13 As shown, mobile phone 2 sends the audio "After get off work dining together" to mobile phone 1. Mobile phone 1 can first convert the audio "After get off work dining together" into text "After get off work dining together", then convert the text "After work dining together" into text "After get off work dining together", and then display the text "After work dining together" 50 in the floating window 22.
[0127] In the above embodiment, when participants use different languages, based on the AI subtitle application, not only language translation can be achieved, but also text and audio conversion can be achieved, thereby improving the communication efficiency of the participants.
[0128] It should be noted that the above-mentioned scenario examples are all illustrated by taking the AI subtitle application enabling the conference application as an example, which does not limit this application. In actual implementation, the AI subtitle application can still enable audio and video applications, such as converting the audio of the audio and video application into text, and displaying the text in subtitles. In addition, the AI subtitle application can also use call applications, such as converting the audio received by the call application into text and displaying the text; or converting the user input text into audio and sending the audio to other devices.
[0129] In order to more clearly understand the above scenario examples, the following first illustrates the system architecture of the terminal device, and then illustrates the instant messaging method based on AI subtitles in combination with the system architecture.
[0130] The software system of the terminal device can adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. The embodiment of the present invention takes the Android system of the layered architecture as an example to exemplify the software structure of the terminal device. Fig.14 A schematic diagram of the system architecture of a terminal device provided in an embodiment of the present application.
[0131] like Fig.14 As shown, the terminal device can adopt a layered architecture to divide the software into several layers, each with clear roles and division of labor. The layers communicate through software interfaces. In some embodiments, the software layers of the software structure are divided from top to bottom into: application (APP) layer, application framework (FW) layer, system library (FWKLIB), hardware abstraction layer (HAL) layer and kernel layer. The above software architecture runs on the hardware layer, which may include speakers, microphones, USB interfaces, displays, sensors, etc.
[0132] The application layer, also known as the application layer, may include a series of application packages. For example, the application layer may include AI subtitle applications, voice processing applications, conference applications, call applications, game applications, video applications, and settings applications, etc. These applications may be system applications or third-party applications. The AI subtitle application may provide an AI subtitle function. The voice processing application may provide a text to speech (TTS) function and an automatic speech recognition (ASR) function. As an example, the TTS function and the ASR function may be merged into the voice processing application. As another example, the TTS function and the ASR function are provided by two applications respectively, that is, the TTS application and the ASR application are two applications. As another example, the TTS function and the ASR function may also be merged into the AI subtitle application. When these application packages are run, the various service modules provided by the application framework layer may be accessed through the application programming interface (API) and the corresponding intelligent services may be performed.
[0133] The application framework layer provides API and programming framework for the application layer applications. The application framework layer includes some predefined functions. Fig.14As shown, the application framework layer can include a media player (such as MediaPlayer), a media recorder (such as MediaRecorder), an audio data entity (such as AudioFlinger), and an audio manager (such as AudioManager). Among them, the audio manager can include an audio service (AudioService) and an audio policy (AudioPolicy) module. AudioFlinger and AudioPolicy are the two basic services of the audio system: AudioPolicy is the maker of the audio system policy, responsible for the policy selection of audio device switching, volume adjustment policy, etc.; AudioFlinger is the executor of the audio system policy, responsible for the management of audio stream devices and the processing and transmission of audio stream data, so AudioFlinger is also called the engine of the audio system.
[0134] Exemplarily, after the conference application receives audio data from the peer device, the conference application can send the audio data to the media player. The media player can send the audio data to AudioFlinger for audio processing. AudioFlinger can obtain the audio policy corresponding to the audio data from the audio manager. Among them, the conference application can call AudioService and AudioPolicy to set the corresponding audio policy for the audio data at runtime. AudioFlinger can process the audio data from the conference application according to the audio policy formulated by AudioPolicy, such as mixing, resampling and sound effect setting of the audio data.
[0135] The system library can include multiple functional modules, such as surface manager, media libraries, 2D graphics engine, 3D graphics processing library, etc. In the system library, Android Runtime includes core library and virtual machine. Android Runtime is responsible for scheduling and management of Android system. Virtual machine is used to perform object life cycle management, stack management, thread management, security and exception management, garbage collection and other functions.
[0136] The hardware abstraction layer has a standard interface implemented by hardware vendors. For example, the hardware abstraction layer may include audio HAL, display HAL, etc. In addition, the hardware abstraction layer may also include playback thread (DuplicatingThread), mixing thread (MixerThread) and recording thread (RecordThread). AudioFlinger calls these threads to send the processed audio data to the audio HAL, and the audio HAL sends the audio data to the corresponding audio output device (such as speakers, wired headphones or Bluetooth headphones, etc.) for playback. Fig.14 As shown, depending on the audio output device, Audio HAL can be divided into audio primary interface, USB interface, audio remote processing (remote submix) interface, A2dp interface, etc. AudioFlinger can send audio data to different audio interfaces in Audio HAL according to the audio policy formulated by AudioPolicy. For example, AudioFlinger can call the audio primary interface to output audio data to the speaker of the mobile phone; call the USB interface to output audio data to the wired headset connected to the mobile phone; call the audio remote processing (remote submix) interface to output audio data to the remote device; call the A2dp interface to output audio data to the Bluetooth headset connected to the mobile phone. In addition, AudioFlinger can also call these threads to read audio data from the audio interface.
[0137] The kernel layer is the layer between hardware and software and belongs to the bottom layer of the Android system. The kernel layer can contain various driver interfaces, such as speaker driver, microphone driver, USB driver, and display driver.
[0138] It should be noted that although the embodiments of the present application are described using the Android system as an example, its basic principles are also applicable to terminal devices based on operating systems such as iOS or Windows.
[0139] It is understandable that in order to implement the instant messaging method based on AI subtitles in the embodiment of the present application, the terminal device includes hardware and / or software modules corresponding to the execution of each function. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments.
[0140] Taking conference applications as an example, three instant messaging methods based on AI subtitles are introduced in combination with three embodiments. It should be understood that the instant messaging method can also be applied to other instant messaging applications.
[0141] Embodiment 1
[0142] Embodiment 1 provides a first audio strategy: the local device does not support collecting local sound through a microphone and transmitting it to the peer device, and can only convert the text input in the floating window of the AI subtitle application into audio data and send it to the peer device through the TTS function; the local device also does not support playing audio data from the peer device through a speaker, and can only convert the audio data from the peer device into text and display it in the floating window of the AI subtitle application through the ASR function.
[0143] For example, Fig.15 shows a module interaction diagram based on AI subtitles, Fig.16 Shown with Fig.15 Schematic diagram of the corresponding specific method flow chart.
[0144] like Fig.15 As shown, during the instant messaging process between mobile phone 1 and mobile phone 2, four data streams are involved in mobile phone 1.
[0145] 1. Data flow as shown in line 1.
[0146] The user of mobile phone 1 enters text in floating window 22, and the AI subtitle application calls the voice processing application to convert the text into audio data. After that, the AI subtitle application sends the audio data to AudioFlinger through the media player (MediaPlayer). AudioFlinger calls the playback thread (DuplicatingThread), and the playback thread (DuplicatingThread) calls the mixing thread (MixerThread) to send the audio data to the audio remote processing (remote_submix) interface, such as the MonoPipe interface of the audio remote processing (remote_submix) interface.
[0147] It should be noted that the above-mentioned DuplicatingThread is the main thread responsible for audio playback, and the DuplicatingThread can call the MixerThread and the RecorderThread, etc. Among them, the MixerThread can be responsible for mixing the audio from other devices, and the RecorderThread can be responsible for mixing the locally recorded audio and the audio obtained through TTS.
[0148] 2. Data flow as shown in line 2.
[0149] Usually, when the microphone is turned on, the microphone records the local audio, and then the microphone driver sends the audio data recorded by the microphone to the audio primary interface. AudioFlinger calls the recording thread (RecordThread) to read the audio data from the audio primary interface and send the audio data to the conference application through the media recorder (MediaRecorder). Fig.15 As shown, when the data path / data stream shown in line 2 is closed, the microphone is in a closed state, so that the conference application will not obtain the local audio collected by the microphone.
[0150] 3. Data flow as shown in line 3.
[0151] The conference application can call AudioFlinger and the recording thread (RecordThread) through the media recorder (MediaRecorder) to read data from the audio remote processing (remote_submix) interface, such as the MonoPipe interface of the audio remote processing (remote_submix) interface to read the stored data.
[0152] 4. Data flow as shown in line 4.
[0153] The conference application can send the audio data from mobile phone 2 to AudioFlinger through the media player (MediaPlayer). AudioFlinger calls the playback thread (DuplicatingThread) and the mixing thread (MixerThread) to send the audio data to the audio remote processing (remote_submix) interface, such as the MonoPipe interface of the audio remote processing (remote_submix) interface. Then, the AI subtitle application can read the audio data from the MonoPipe interface through AudioFlinger. After that, the AI subtitle application calls the voice processing application to convert the audio data into text. After that, the AI subtitle application displays the converted text in the floating window 22. However, if Fig.15 As shown, since the path of the speaker driver is closed and the speaker is in a closed state, the speaker driver cannot read the audio data from the MonoPipe interface and will not output the sound corresponding to the audio data through the speaker.
[0154] like Fig.16 As shown, the method is applied to a mobile phone 1, and the method includes A1 to A12. It should be noted that due to the limited size of the attached drawings, Fig.16The playback thread (DuplicatingThread), the mixing thread (MixerThread), the recording thread (RecordThread), etc. are not shown.
[0155] A1, the conference application receives audio data a from mobile phone 2.
[0156] Among them, mobile phone 2 can be a terminal device that establishes a conference link with mobile phone 1.
[0157] In actual implementation, mobile phone 1 may establish conference links with multiple devices, that is, multiple people may join the conference. At a certain moment, the conference application of mobile phone 1 may receive multiple audio data at the same time, such as audio data from mobile phone 2 and audio data from mobile phone 3. The way mobile phone 1 processes these multiple audio data is similar to the way it processes audio data a. Please refer to the description of A2 to A6 below, which will not be repeated here.
[0158] A2: The conference application sends audio data a to AudioFlinger through the media player (MediaPlayer).
[0159] A3, AudioFlinger calls the playback thread (DuplicatingThread) and the mixing thread (MixerThread) to send the audio data a to the audio remote processing (remote_submix) interface.
[0160] For example, the audio data a is sent to the MonoPipe interface of the audio remote processing (remote_submix) interface.
[0161] It should be noted that the audio primary interface and the audio remote processing (remote_submix) interface have different properties and support different data types. Speakers and microphones usually use the audio primary interface, but cannot use the audio remote processing (remote_submix) interface. In embodiment 1, since the local device does not support the audio data from the opposite device to be played by the speaker, in this case, the audio primary interface may not be used, and the audio remote processing (remote_submix) interface may be used to process the remote data, such as audio data a.
[0162] A4, audio remote processing (remote_submix) interface writes audio data a.
[0163] A5, the AI subtitle application reads audio data a from the MonoPipe interface through AudioFlinger.
[0164] As an example, the AI subtitle application can read audio data from the MonoPipe interface according to a preset period. For example, the AI subtitle application can send a read instruction to AudioFlinger every N milliseconds. If the MonoPipe interface stores new audio data, such as audio data a, AudioFlinger returns audio data a to the AI subtitle application. If the MonoPipe interface does not store new audio data, AudioFlinger returns empty data to the AI subtitle application, or does not return data.
[0165] A6, the AI subtitle application sends a request to the voice processing application to convert the audio data a into text. In response to the request, the voice processing application converts the audio data a into text A and returns text A to the AI subtitle application.
[0166] The above-mentioned voice processing application provides ASR service and TTS service. In the above-mentioned A6, the voice processing application can convert the audio data a into text based on the ASR service.
[0167] A7, AI subtitle application displays text A in the floating window.
[0168] like Figure 7 As shown in (d), after the AI subtitle application calls the voice processing application to convert the audio data a into the text "Gather at the company gate", the text "Gather at the company gate" 38 can be displayed in the floating window 22.
[0169] When the conference application of mobile phone 1 receives multiple audio data at the same time, the AI subtitle application of mobile phone 1 can convert each of the multiple audio data into a text message and display each text message in the floating window 22. Each text message carries the device identifier, user identifier or account identifier that sent the audio data. In this way, the user can distinguish the source of each text message by these identifiers.
[0170] A8, the AI subtitle application receives text B input by the user in the floating window.
[0171] like Figure 5 (a) to Figure 5 As shown in (e) in FIG. 1 , the user of the mobile phone 1 can input the text “OK, I understand” in the input box 23 of the floating window 22 and click the send control 26 , so that the mobile phone 1 can execute the following A9 to A12 .
[0172] A9, the AI subtitle application sends a request to the voice processing application to convert text B into audio data. In response to the request, the voice processing application converts text B into audio data b and returns the audio data b to the AI subtitle application.
[0173] A10, the AI subtitle application sends audio data b to AudioFlinger through the media player (MediaPlayer). AudioFlinger calls the playback thread (DuplicatingThread) and the mixing thread (MixerThread) to send the audio data to the audio remote processing (remote_submix) interface.
[0174] For example, the audio data b is sent to the MonoPipe interface of the audio remote processing (remote_submix) interface.
[0175] A11, the conference application calls the recording thread (RecordThread) through the media recorder (MediaRecorder) and AudioFlinger, and reads audio data b from the MonoPipe interface of the audio remote processing (remote_submix) interface.
[0176] A12, the conference application sends the audio data b to the mobile phone 2 via the network.
[0177] For example, mobile phone 1 converts the text "OK, I understand" into audio, and sends a message carrying the audio "OK, I understand" to mobile phone 2, and mobile phone 2 can play the audio "OK, I understand" through a speaker, receiver or headset. Figure 5 As shown in (f) in FIG. 1 , the mobile phone 1 may display the text “OK, I understand” 32 in the floating window 22 , so that the user using the mobile phone may know that the audio message corresponding to the text “OK, I understand” has been successfully sent.
[0178] In the above embodiment, if the user of mobile phone 1 does not want to hear the sound from mobile phone 2, nor does he want to send local sound (such as background sound) to mobile phone 2, then the user of mobile phone 1 can turn off the microphone and speaker by operation. In this case, the microphone and speaker are both in the off state, and neither local recording nor sound from mobile phone 2 can be played. However, the AI subtitles of mobile phone 1 support the conversion of audio from mobile phone 2 into text for display, and also support the conversion of locally input text into audio and sending it to the other end. That is, the user of mobile phone 1 can communicate with the user of mobile phone 2 only through text. It should be noted that for the specific implementation method of the user turning off the microphone and speaker, please refer to the description of the above-mentioned scenario examples 3 to scenario examples 6, which will not be repeated here.
[0179] Embodiment 2
[0180] Embodiment 2 provides a second audio strategy: the local device does not support collecting local sound through a microphone and transmitting it to the peer device, and can only convert the text input in the floating window of the AI subtitle application into audio data through the TTS function and send it to the peer device; the local device supports both playing the audio data from the peer device through the speaker, and also supports converting the audio data from the peer device into text through the ASR function and displaying it in the floating window of the AI subtitle application.
[0181] For example, Fig.17 Another module interaction diagram based on AI subtitles is shown. Fig.18 Shown with Fig.17 Schematic diagram of the corresponding specific method flow chart.
[0182] like Fig.17 As shown, during the instant messaging process between mobile phone 1 and mobile phone 2, four data streams are involved in mobile phone 1.
[0183] 1. Data flow as shown in line 1.
[0184] 2. Data flow as shown in line 2.
[0185] 3. Data flow as shown in line 3.
[0186] For the implementation of the data flow shown in line type 1, the data flow shown in line type 2, and the data flow shown in line type 3, reference may be made to the specific description of the above embodiment 1, which will not be repeated here.
[0187] 4. Data flow as shown in line 4.
[0188] The conference application can send the audio data from mobile phone 2 to AudioFlinger through the media player (MediaPlayer). AudioFlinger calls the playback thread (DuplicatingThread) and the mixing thread (MixerThread) to send the audio data to the audio primary interface, and the audio primary interface writes the audio data. On the one hand, the AI subtitle application can read the audio data from the audio primary interface through AudioFlinger and the mixing thread (MixerThread), and then call the voice processing application to convert the audio data into text, and then display the converted text in the floating window 22. On the other hand, the speaker driver can also read the audio data from the audio primary interface and play the sound corresponding to the audio data through the speaker.
[0189] like Fig.18As shown, the method is applied to a mobile phone 1, and the method includes B1 to B13. It should be noted that due to the limited size of the attached drawings, Fig.18 The playback thread (DuplicatingThread), the mixing thread (MixerThread), the recording thread (RecordThread), etc. are not shown.
[0190] B1, the conference application receives audio data a from mobile phone 2.
[0191] B2, the conference application sends the audio data a to AudioFlinger through the media player (MediaPlayer).
[0192] B3, AudioFlinger calls the playback thread (DuplicatingThread) and the mixing thread (MixerThread) to send the audio data a to the audio primary interface.
[0193] Referring to the description of the above-mentioned embodiment 1, the properties of the audio primary interface and the audio remote processing (remote_submix) interface are different, and the supported data types are also different. The speaker and the microphone usually use the audio primary interface, but cannot use the audio remote processing (remote_submix) interface. In embodiment 2, since the local device supports the speaker to play the audio data from the opposite device, in this case, the audio primary interface can be used to process the remote data, such as audio data a.
[0194] B4, the audio primary interface writes audio data a.
[0195] B5, the AI subtitle application reads audio data a from the audio primary interface through AudioFlinger.
[0196] As an example, the AI subtitle application can read audio data from the audio primary interface according to a preset period. For example, the AI subtitle application can send a read instruction to AudioFlinger every N milliseconds. If the audio primary interface stores new audio data, such as audio data a, AudioFlinger returns audio data a to the AI subtitle application. If the audio primary interface does not store new audio data, AudioFlinger returns empty data to the AI subtitle application, or does not return data.
[0197] B6, the AI subtitle application sends a request to the voice processing application to convert the audio data a into text. In response to the request, the voice processing application converts the audio data a into text A and returns text A to the AI subtitle application.
[0198] B7, AI subtitle application displays text A in the floating window.
[0199] For the specific implementation of B6 and B7, please refer to the description of A6 and A7 in Example 1, which will not be repeated here.
[0200] B8, the audio primary interface sends audio data a to the speaker driver.
[0201] The speaker is driven to output the sound corresponding to the audio data a.
[0202] It should be noted that the present application does not specifically limit the execution order of B5-B7 and B8. For example, B5-B7 and B8 can be executed synchronously, so that when the speaker plays the sound corresponding to the audio data a, the screen displays the text A corresponding to the audio data a.
[0203] B9, the AI subtitle application receives the text B entered by the user in the floating window.
[0204] B10, the AI subtitle application sends a request to the voice processing application to convert text B into audio data. In response to the request, the voice processing application converts text B into audio data b and returns the audio data b to the AI subtitle application.
[0205] B11, the AI subtitle application sends the audio data b to AudioFlinger through the media player (MediaPlayer). AudioFlinger calls the playback thread (DuplicatingThread) and the mixing thread (MixerThread) to send the audio data to the audio remote processing (remote_submix) interface.
[0206] For example, the audio data b is sent to the MonoPipe interface of the audio remote processing (remote_submix) interface.
[0207] B12, the conference application calls the recording thread (RecordThread) through the media recorder (MediaRecorder) and AudioFlinger, and reads data from the audio remote processing (remote_submix) interface, such as reading audio data b from the MonoPipe interface of the audio remote processing (remote_submix) interface.
[0208] B13, the conference application sends the audio data b to the mobile phone 2 via the network.
[0209] For the specific implementation of B9-B13, please refer to the description of A8-A12 in Example 2, which will not be repeated here.
[0210] In the above embodiment, if the user of mobile phone 1 wants to hear the sound from mobile phone 2, but does not want to send local sound (such as background sound) to mobile phone 2, the user of mobile phone 1 can turn off the microphone and turn on the speaker. In this case, the microphone cannot record locally, and the speaker can play the sound from mobile phone 2. In addition, AI subtitles support converting the audio from mobile phone 2 into text for display, and also support converting the locally input text into audio and sending it to the other end.
[0211] Embodiment 3
[0212] Example 3 provides a third audio strategy: the local device supports both collecting local sound through a microphone and converting text input in the floating window of the AI subtitle application into audio data through the TTS function; the local device supports both playing audio data from the other device through a speaker and converting audio data from the other device into text through the ASR function and displaying it in the floating window of the AI subtitle application.
[0213] For example, Fig.19 Another module interaction diagram based on AI subtitles is shown. Fig. 20 Shown with Fig.19 Schematic diagram of the corresponding specific method flow chart.
[0214] like Fig.19 As shown, during the instant messaging process between mobile phone 1 and mobile phone 2, four data streams are involved in mobile phone 1.
[0215] 1. Data flow as shown in line 1.
[0216] The user of mobile phone 1 enters text in floating window 22, and the AI subtitle application calls the voice processing application to convert the text into audio data. After that, the AI subtitle application sends the audio data to AudioFlinger through the media player (MediaPlayer). AudioFlinger calls the mixer thread (MixerThread) to send the audio data to the audio primary interface.
[0217] 2. Data flow as shown in line 2.
[0218] The microphone is turned on. The microphone records local audio. The microphone driver sends the audio data recorded by the microphone to the audio primary interface.
[0219] In this way, the recording thread (RecordThread) can obtain two kinds of audio data (i.e., audio data obtained through the data stream shown in line 1 and audio data obtained through the data stream shown in line 2) from the audio primary interface, and mix the two kinds of audio data to obtain mixed data.
[0220] 3. Data flow as shown in line 3.
[0221] The conference application can call AudioFlinger through the media recorder (MediaRecorder) to read the mixing data from the recording thread (RecordThread).
[0222] 4. Data flow as shown in line 4.
[0223] The conference application can send the audio data from mobile phone 2 to AudioFlinger through the media player (MediaPlayer). AudioFlinger calls the mixer thread (MixerThread) to send the audio data to the audio primary interface, and the audio primary interface writes the audio data. On the one hand, the AI subtitle application can read the audio data from the audio primary interface through AudioFlinger, and call the voice processing application to convert the audio data into text, and then display the converted text in the floating window 22. On the other hand, the speaker driver can also read the audio data from the audio primary interface and play the sound corresponding to the audio data through the speaker. It should be noted that the audio primary interface of the mixing thread is a write interface, and the audio primary interface of the recording thread is a read interface. These two audio primary interfaces belong to different types of interfaces.
[0224] like Fig. 20 As shown, the method is applied to a mobile phone 1, and the method includes C1 to C14. It should be noted that due to the limited size of the attached drawings, Fig. 20 The playback thread (DuplicatingThread), the mixing thread (MixerThread), the recording thread (RecordThread), etc. are not shown.
[0225] C1, the conference application receives audio data a from mobile phone 2.
[0226] C2, the conference application sends the audio data a to AudioFlinger through the media player (MediaPlayer).
[0227] C3, AudioFlinger calls the mixer thread (MixerThread) to send the audio data a to the audio primary interface.
[0228] Referring to the description of the above-mentioned embodiment 1, the properties of the audio primary interface and the audio remote processing (remote_submix) interface are different, and the supported data types are also different. Speakers and microphones usually use the audio primary interface, but cannot use the audio remote processing (remote_submix) interface. In embodiment 3, since the local device supports both the playback of audio data from the opposite device by the speaker and the collection of local sound by the microphone, in this case, the audio primary interface can be used to store the audio data from the opposite device and the local audio data collected by the microphone.
[0229] C4, the audio primary interface writes audio data a.
[0230] C5, the AI subtitle application reads audio data a from the audio primary interface through AudioFlinger.
[0231] For the specific implementation of C5, please refer to the description of B5 in the second embodiment, which will not be repeated here.
[0232] C6, the AI subtitle application sends a request to the voice processing application to convert the audio data a into text. In response to the request, the voice processing application converts the audio data a into text A and returns text A to the AI subtitle application.
[0233] C7, AI subtitle application displays text A in the floating window.
[0234] C8, the audio primary interface sends audio data a to the speaker driver.
[0235] The speaker is driven to output the sound corresponding to the audio data a.
[0236] For the specific implementation of C1-C8, please refer to the description of B1-B8 in Example 2, which will not be repeated here.
[0237] C9, the AI subtitle application receives text B input by the user in the floating window.
[0238] C10, the AI subtitle application sends a request to the voice processing application to convert text B into audio data. In response to the request, the voice processing application converts text B into audio data b and returns the audio data b to the AI subtitle application.
[0239] C11, the AI subtitle application sends the audio data b to AudioFlinger through the media player (MediaPlayer). AudioFlinger calls the mixer thread (MixerThread) to send the audio data to the audio primary interface.
[0240] C12, the microphone driver sends the audio data c recorded by the microphone to the audio primary interface.
[0241] C13, the conference application calls the recording thread (RecordThread) through the media recorder (MediaRecorder) and AudioFlinger to read the mixed data of audio data b and audio data c.
[0242] For example, the conference application can send a data read instruction to the recording thread (RecordThread) through the media recorder (MediaRecorder) and AudioFlinger at preset intervals. For example, if the recording thread (RecordThread) reads audio data b and audio data c from the audio primary interface in a certain period, the recording thread (RecordThread) mixes audio data b and audio data c to obtain audio data d, and then returns audio data d to the conference application through AudioFlinger and the media recorder (MediaRecorder).
[0243] C14, the conference application sends the mixed audio data of the audio data b and the audio data c to the mobile phone 2 through the network.
[0244] In the above embodiment, if the user of mobile phone 1 wants to hear the sound from mobile phone 2 and wants to send local sound (such as background sound) to mobile phone 2, the user of mobile phone 1 can turn on the microphone and speaker through operation. In this case, the microphone can perform local recording and the speaker can play the sound from mobile phone 2. In addition, AI subtitles support converting the audio from mobile phone 2 into text for display, and also support converting the locally input text into audio and sending it to the other end.
[0245] The above three embodiments introduce three audio strategies. As an implementation method, the mobile phone can pre-set an audio strategy, and the mobile phone can directly use this audio strategy during the conference, such as the first audio strategy. As another implementation method, the mobile phone can pre-set three audio strategies and formulate the use methods of these three audio strategies so that only one audio strategy can be used at a certain time, thereby avoiding logical conflicts caused by using multiple audio strategies at the same time.
[0246] Combine the following Figure 21 to Figure 24 , which introduces the specific usage of the three audio strategies.
[0247] For example, Fig.21 The module interaction diagram corresponding to the parameter distribution strategy of downlink data (data from mobile phone 2) is shown. Fig. 22 Shown with Fig.21 Schematic diagram of the corresponding specific method flow chart.
[0248] like Fig.21 and Fig. 22 As shown, during the instant messaging process between mobile phone 1 and mobile phone 2, the audio manager of mobile phone 1 can formulate a parameter distribution strategy corresponding to the downlink data. Figure 7 The playback control 35 shown in (c) is in the form of a speaker, or as Figure 8 As shown in (a) of FIG. 1 , the user presses the “volume +” button 39, or as shown in FIG. Fig. 9 The playback control 42 shown in (b) is in the form of prompt closing. In these cases, the speaker is turned on, and the audio manager of the mobile phone 1 sends parameter 1 to AudioFlinger; Figure 7 The playback control 35 shown in (a) is in the form of a handset, or as shown in FIG. Figure 8 As shown in (c) of FIG. 1 , the user presses the “volume -” button 40, or as shown in FIG. Fig. 9 The playback control 42 shown in (e) is in the prompt-on form. In these cases, the speaker is in the off state, and the audio manager of the mobile phone 1 sends parameter 2 to AudioFlinger.
[0249] After the conference application of mobile phone 1 receives the audio data a from mobile phone 2, the conference application of mobile phone 1 can send the audio data a to AudioFlinger through the media player (MediaPlayer). AudioFlinger executes the distribution strategy corresponding to the downlink parameter at the current moment. If the downlink parameter at the current moment is parameter 1, A1 to A7 are executed, that is, the local device does not support the audio data from the opposite device to be played by the speaker, and can only convert the audio data from the opposite device into text through the ASR function and display it in the floating window of the AI subtitle application; if the downlink parameter at the current moment is parameter 2, B1 to B7 are executed, that is, the local device supports both the audio data from the opposite device to be played by the speaker, and the audio data from the opposite device to be converted into text through the ASR function and displayed in the floating window of the AI subtitle application. For the specific implementation methods of A1 to A7 and B1 to B7, please refer to the description of the above embodiments, which will not be repeated here.
[0250] For example, Fig.23The module interaction diagram corresponding to the parameter distribution strategy of uplink data (data sent to mobile phone 2) is shown. Fig.24 Shown with Fig.23 Schematic diagram of the corresponding specific method flow chart.
[0251] like Fig.23 and Fig.24 As shown, during the instant messaging process between mobile phone 1 and mobile phone 2, the audio manager of mobile phone 1 can formulate a parameter distribution strategy corresponding to the uplink data. Figure 6 The mute control 25 shown in (a) is in mute mode, or as shown in Figure 6 The mute control 33 shown in (c) is in mute mode. In these cases, the microphone is turned off, and the audio manager of the mobile phone 1 sends parameter 3 to AudioFlinger; Figure 6 The play control 35 and the mute control 33 shown in (b) are in non-mute form. In these cases, the speaker is turned on and the audio manager of the mobile phone 1 sends parameter 4 to AudioFlinger.
[0252] The user of mobile phone 1 enters text in the floating window 22. The AI subtitle application calls the voice processing application to convert the text into audio data. Afterwards, the AI subtitle application sends the audio data to AudioFlinger through the media player (MediaPlayer). AudioFlinger executes the distribution strategy corresponding to the uplink parameter at the current moment. If the uplink parameter at the current moment is parameter 3, A8 to A12 are executed, that is, the local device does not support the collection of local sound through the microphone and transmit it to the opposite device, and can only convert the text into audio data and send it to the opposite device through the TTS function; if the uplink parameter at the current moment is parameter 4, C9 to C14 are executed, that is, the local device supports both the collection of local sound through the microphone and the conversion of text into audio data through the TTS function. For the specific implementation methods of A8 to A12 and C9 to C14, please refer to the description of the above embodiments, which will not be repeated here.
[0253] According to the description of the above embodiment, the downlink parameters include parameters 1 and 2, and the uplink parameters include parameters 3 and 4. Usually, the audio manager sends a set of parameters to AudioFlinger, and this set of parameters includes a downlink parameter and an uplink parameter. If this set of parameters includes parameters 1 and 3, the audio policy at this moment is called the first audio policy in the above embodiment one; if this set of parameters includes parameters 2 and 3, the audio policy at this moment is called the second audio policy in the above embodiment two; if this set of parameters includes parameters 2 and 4, the audio policy at this moment is called the third audio policy in the above embodiment three. It should be understood that in addition to the first audio policy, the second audio policy and the third audio policy, there may also be a fourth audio policy, including parameters 1 and 4. The implementation method of the fourth audio policy can refer to the description of the above embodiment, which will not be repeated here.
[0254] In the above embodiment, the audio manager can formulate different audio strategies by sending different uplink and downlink parameters to AudioFlinger, thereby realizing the control of the audio path of the microphone and the audio path of the speaker to meet the audio usage needs of users in different conference scenarios.
[0255] Fig.25 A schematic diagram of the hardware structure of a terminal device provided in an embodiment of the present application is shown.
[0256] Take the mobile phone as an example. Fig.25 As shown, the mobile phone 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0257] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a controller, a memory, a baseband processor, and the like.
[0258] The wireless communication function of the mobile phone 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, so that the mobile phone 100 can exchange instant communication data with other devices, such as voice, video, picture, text, emoticon and other types of data.
[0259] The mobile phone 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The display screen 194 is used to display content, such as displaying a video call interface, a voice call interface, a subtitle box, a control center, a setting application interface, etc.
[0260] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, the subtitle text obtained based on the AI subtitle application is saved in the external memory card.
[0261] The mobile phone 100 can realize audio functions through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, and an application processor. Among them, the audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals; the speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals, and the user can listen to call request messages, instant voice messages, etc. from the opposite device through the speaker 170A; the receiver 170B, also called a "earpiece", is used to convert audio electrical signals from the opposite device into sound signals; the microphone 170C, also called a "microphone" or "microphone", is used to convert locally collected sound signals into electrical signals. During the instant messaging process, the user can make a sound by approaching the microphone 170C with his mouth to input the sound signal into the microphone 170C; the earphone interface 170D is used to connect a wired headset to convert the audio electrical signals from the opposite device into sound signals.
[0262] It is to be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the mobile phone 100. In other embodiments of the present application, the mobile phone 100 may include more or fewer components than shown in the figure, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0263] An embodiment of the present application also provides a terminal device, including a processor, the processor is coupled to a memory, and the processor is used to execute a computer program or instruction stored in the memory, so that the terminal device implements the methods in the above embodiments.
[0264] The embodiment of the present application also provides a computer-readable storage medium, in which computer instructions are stored; when the computer-readable storage medium is run on a computer, the computer is caused to execute the method shown above. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that can be integrated with one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk or a tape), an optical medium or a semiconductor medium (e.g., a solid state drive (SSD)), etc.
[0265] An embodiment of the present application further provides a computer program product, which includes a computer program code. When the computer program code runs on a computer, the computer executes the methods in the above embodiments.
[0266] The embodiment of the present application also provides a chip, which is coupled to a memory, and the chip is used to read and execute a computer program or instruction stored in the memory to perform the methods in the above embodiments. The chip can be a general-purpose processor or a dedicated processor. It should be noted that the chip can be implemented using the following circuits or devices: one or more field programmable gate arrays (field programmable gate array, FPGA), programmable logic devices (programmable logic device, PLD), controllers, state machines, gate logic, discrete hardware components, any other suitable circuits, or any combination of circuits capable of performing the various functions described throughout this application.
[0267] The terminal device, computer-readable storage medium, computer program product and chip provided in the above-mentioned embodiments of the present application are all used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects corresponding to the method provided above, and will not be repeated here.
[0268] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0269] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. An instant messaging method based on AI subtitles, It is characterized in that The method is applied to a first device, the first device includes a conference application, an AI subtitle application, an audio data entity AudioFlinger and a first audio interface, the first audio interface is an audio main interface or an audio remote processing interface of a hardware abstraction layer, and the method includes: The first device displays the interface of the conference application and the window of the AI subtitle application; The first device receives a first text input by a user in a window of the AI subtitle application; The first device writes first audio data to the first audio interface through the AudioFlinger, where the first audio data is obtained after the first text is converted from text to speech; In response to a recording request of the conference application, the first device reads the first audio data from the first audio interface through the AudioFlinger; The first device sends the first audio data to the second device through the conference application.
2. The method according to claim 1, It is characterized in that The method further comprises: The first device displays a sending message in the window of the AI subtitle application, and the sending message includes the first text.
3. The method according to claim 1 or 2, It is characterized in that The first audio interface is the audio remote processing interface, the first device further comprises a microphone and a media recorder, and the microphone is in a recording-disabled state; The step of responding to the recording request of the conference application, the first device reading the first audio data from the first audio interface through the AudioFlinger, comprises: In response to a recording request initiated by the conference application through the media recorder, the first device calls a recording thread through the AudioFlinger to read the first audio data from the audio remote processing interface.
4. The method according to claim 1 or 2, It is characterized in that The first audio interface is the audio main interface, the first device further includes a microphone and a media recorder, and the microphone is in a state of allowing recording; The method further includes: the first device writing second audio data into the audio main interface, the second audio data being audio data collected by the microphone; The first device reading the first audio data from the first audio interface through the AudioFlinger in response to the recording request of the conference application includes: in response to the recording request initiated by the conference application through the media recorder, the first device reading the mixed data of the first audio data and the second audio data from the audio main interface through the AudioFlinger calling the recording thread; The first device sending the first audio data to the second device through the conference application includes: the first device sending mixed data of the first audio data and the second audio data to the second device through the conference application.
5. The method according to any one of claims 1 to 4, It is characterized in that The first device further includes a second audio interface, wherein the second audio interface is an audio main interface or an audio remote processing interface of the hardware abstraction layer; The method further comprises: The first device receives third audio data from the second device through the conference application; The first device writes the third audio data to the second audio interface through the AudioFlinger; In response to a playback request of the AI subtitle application, the first device reads the third audio data from the second audio interface through the AudioFlinger; The first device displays a received message in the window of the AI subtitle application, where the received message includes a second text, and the second text is obtained after automatic speech recognition of the third audio data.
6. The method according to claim 5, It is characterized in that The second audio interface is the audio remote processing interface, and the first device also includes a speaker, and the speaker is in a state of prohibiting sound.
7. The method according to claim 5, It is characterized in that The second audio interface is the audio main interface, the first device further includes a speaker, and the speaker is in a state of allowing sound to be emitted; The method further comprises: The first device reads the third audio data from the audio main interface; The first device outputs sound corresponding to the third audio data through the speaker.
8. The method according to any one of claims 5 to 7, It is characterized in that The first device also includes an audio manager; the method further includes: Based on the status of the microphone and the speaker of the first device, the first device formulates an audio strategy for the AudioFlinger through the audio manager; Among them, if the audio strategy is the first audio strategy, the first audio interface and the second audio interface are the audio remote processing interface; if the audio strategy is the second audio strategy, the first audio interface is the audio remote processing interface, and the second audio interface is the audio main interface; if the audio strategy is the third audio strategy, the first audio interface and the second audio interface are the audio main interface.
9. The method according to any one of claims 5 to 7, It is characterized in that The first device further includes a media player; and the first device writes the third audio data to the second audio interface through the AudioFlinger, including: The first device calls the AudioFlinger through the media player to write the third audio data to the second audio interface.
10. The method according to any one of claims 1 to 8, It is characterized in that The first device further includes a media player; and the first device writes first audio data to the first audio interface through the AudioFlinger, including: The first device calls the AudioFlinger through the media player to write the first audio data to the first audio interface.
11. An instant messaging method based on AI subtitles, It is characterized in that The method is applied to a first device, the first device includes a conference application and an AI subtitle application, and the method includes: The first device displays the interface of the conference application; The first device, in response to a first operation of the user, displays a window of the AI subtitle application in a floating manner on the interface of the conference application; The first device receives a first text input by a user in a window of the AI subtitle application; The first device sends first audio data to the second device through the conference application, where the first audio data is obtained after the first text is converted from text to speech; The first device displays a sending message in the window of the AI subtitle application, and the sending message includes the first text.
12. The method according to claim 11, It is characterized in that The first device also includes a microphone; The method further comprises: The first device acquires second audio data, where the second audio data is audio data collected by the microphone; The first device sending the first audio data to the second device through the conference application includes: The first device sends mixed data of the first audio data and the second audio data to the second device through the conference application.
13. The method according to claim 11 or 12, It is characterized in that The method further comprises: The first device receives third audio data from the second device through the conference application; The first device displays a received message in the window of the AI subtitle application, where the received message includes a second text, and the second text is obtained after automatic speech recognition of the third audio data.
14. A terminal device, It is characterized in that comprising a processor, and a memory coupled to the processor; Wherein, instructions are stored in the memory, and the processor calls the instructions so that the terminal device executes the instant messaging method based on AI subtitles as described in any one of claims 1 to 13.
15. A chip, It is characterized in that The chip is coupled to a memory, and the chip is used to read and execute a computer program stored in the memory to implement the instant messaging method based on AI subtitles as described in any one of claims 1 to 13.
16. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program runs on an electronic device, the electronic device executes the instant messaging method based on AI subtitles as described in any one of claims 1 to 13.
Citation Information
Patent Citations
Communication method and system for voice and text conversion
CN101123630A
Instant messaging method, terminal, equipment and computer readable medium
CN109600307A
Subtitle generation method and terminal
CN110324723A
Method and device for superposing real-time voice subtitles on video images
CN112511847A
Interaction method and device for realizing conversation of hearing-impaired people, terminal and storage medium
CN112887194A