Method and mobile terminal for text call

Dedicated voice data transmission paths and virtual audio devices in text call systems resolve display confusion by correctly segregating and displaying user input text and received speech on mobile terminals, improving user experience.

WO2026071811A1PCT designated stage Publication Date: 2026-04-02SAMSUNG ELECTRONICS CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing text call systems suffer from display confusion on mobile terminals due to automatic speech recognition mistakenly recognizing and displaying input text data as the speech of the other party, leading to user interface confusion.

Method used

Implementing dedicated voice data transmission paths for text calls, using virtual audio devices to separate and manage voice and text data streams, ensuring correct display of user input text and received speech on a text call interface.

Benefits of technology

Prevents erroneous display of user input text in the speech area of the other party, enhancing user experience by maintaining clear separation and accurate representation of communication data on the text call interface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025015336_02042026_PF_FP_ABST
    Figure KR2025015336_02042026_PF_FP_ABST
Patent Text Reader

Abstract

According to an embodiment of the present disclosure, a method for a text call by a mobile terminal is provided. The method may comprise displaying a text call interface based on a text call function of the mobile terminal being enabled and a preset application of the mobile terminal being in a call state; based on first text data being input into the text call interface, converting the first text data into first voice data to be sent by the preset application to an other party of the text call, and displaying the first text data in a first area of the text call interface; and based on receiving second voice data from the other party of the text call, converting the second voice data into second text data, and displaying the second text data in a second area of the text call interface.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND MOBILE TERMINAL FOR TEXT CALL

[0001] The present disclosure relates to the field of electronic technology, and more specifically, to a method and a mobile terminal for a text call.

[0002] With the development and popularization of various smart terminals, audio calls (i.e. VoIP calls, such as WeChat voice calls, QQ voice calls, etc.) as a means of remote communication and collaboration are constantly expanding in modern society. The VoIP calls enable people to communicate and collaborate in different locations through real-time audio transmission.

[0003] Currently, in order to resolve a problem of being inconvenient to use voice to reply to calls in certain scenarios and resolve a call problem of hearing-impaired or / and speech-impaired people, some applications support a text call function. For example, when it is inconvenient for a user A to speak upon receiving a WeChat voice call from a user B in a library, the user A can enable the text call function and directly input text information on a display interface (hereinafter referred to as a text call interface) of a text call application. The text information will be converted into voice information and sent to the user B, thus achieving smooth communication between the user A and user B in specific scenarios.

[0004] FIG. 1 is a diagram illustrating an example of an existing text call function.

[0005] As shown in FIG. 1, a user A of a mobile terminal enables a text call function on a WeChat voice call interface, and the mobile terminal displays a text call interface. At this point, the voice call is displayed in a floating window form on the text call interface. The user A inputs Text 1 ("How's the weather today?") on the text call interface, then Text 1 is converted into voice data Voice_data1 and sent to the other party of the call (i.e., a user B) through WeChat as an input for the WeChat voice call. In this way, the user B can hear corresponding voice content.

[0006] In FIG. 1, a box on the right side displays speech of the user A, while a box on the left side displays speech of the user B. After the user A inputs and sends Text 1 ("How's the weather today?"), Text 2 ("How's the weather today?") should not be displayed repeatedly in the box on the left side (i.e., a speech area of the user B), but only Text 3 ("The weather is nice today") converted from voice content of the user B should be displayed. However, when Text 1 is converted into the voice data Voice_data1 and transmitted to WeChat, if the voice data Voice_data1 is captured by a text call application (i.e., an application that performs the text call function), Voice_data1 will be mistakenly recognized as the voice data transmitted by the user B to the user A through WeChat. In this case, the voice data Voice_data1 will be converted into Text 2 through automatic speech recognition (ASR) technology and displayed in the box on the left side of the text call interface, resulting in display confusion on the text call interface.

[0007] Embodiments of the present disclosure provide a text call method and a text call device for a mobile terminal, which can effectively solve a display confusion phenomenon on a text call interface.

[0008] According to an aspect of the present disclosure, a method for a text call by a mobile terminal is provided. The method may comprise displaying a text call interface based on a text call function of the mobile terminal being enabled and a preset application of the mobile terminal being in a call state. The method may comprise based on first text data being input into the text call interface, converting the first text data into first voice data to be sent by the preset application to an other party of the text call, and displaying the first text data in a first area of the text call interface. The method may comprise based on receiving second voice data from the other party of the text call, converting the second voice data into second text data, and displaying the second text data in a second area of the text call interface.

[0009] The method may comprise based on the text call function being enabled, monitoring audio status events of applications on the mobile terminal. The method may comprise based on detecting that an audio status of the preset application changes, determining that the preset application is in the call state.

[0010] The method may comprise transmitting the first voice data to only the preset application and not transmitting the first voice data to any other application on the mobile terminal.

[0011] The first voice data may be transmitted to the preset application through a first dedicated voice data transmission path.

[0012] The second voice data may be obtained from the preset application through a second dedicated voice data transmission path. The second voice data may be played through a speaker of the mobile terminal.

[0013] Based on a mute function for the second voice data being enabled, the second voice data may be obtained from the preset application through the second dedicated voice data transmission path, and the second voice data may not be played through the speaker of the mobile terminal

[0014] The method may comprise based on determining that the preset application is an application that supports the text call function, establishing at least one of a first dedicated voice data transmission path for transmitting the first voice data to the preset application, or a second dedicated voice data transmission path for obtaining the second voice data from the preset application.

[0015] The method may comprise determining that the preset application is the application that supports the text call function, based on an unique identification information of the preset application being included in a unique identification information list for applications that support the text call function. The unique identification information list for applications that support the text call function may be constructed based on the unique identification information of the applications that support the text call function set by a user.

[0016] Establishing the first dedicated voice data transmission path for transmitting the first voice data to the preset application may comprise creating a first virtual audio device, routing a data stream of the first voice data to an output side of the first virtual audio device, and routing a recording data stream of the preset application to an input side of the first virtual audio device.

[0017] Establishing the second dedicated voice data transmission path for obtaining the second voice data from the preset application may comprise creating a second virtual audio device, routing a data stream of the second voice data to an input side of the second virtual audio device, and routing a playback data stream of the preset application to an output side of the second virtual audio device.

[0018] The method may comprise creating an audio continuity module to store at least one of the following parameters: package name information of an application executing the text call; unique identification information for applications that support the text call function; information indicating whether exiting the text call function or not; and information indicating whether enabling a mute function for the second voice data while establishing a first dedicated voice data transmission path for the first voice data and a second dedicated voice data transmission path for the second voice data.

[0019] The first area may indicate an area of the text call interface used to display sent messages. The second area may indicate an area of the text call interface used to display received messages.

[0020] According to an aspect of the present disclosure, a mobile device is provided. The mobile device may comprise memory storing instructions; and at least one processor operably coupled to the memory. The instructions, when executed by the at least one processor, may cause the mobile terminal to perform operations. The operations may comprise displaying a text call interface based on a text call function of the mobile terminal being enabled and a preset application of the mobile terminal being in a call state. The operations may comprise based on first text data being input into the text call interface, converting the first text data into first voice data to be sent by the preset application to an other party of the text call, and displaying the first text data in a first area of the text call interface. The operations may comprise based on receiving second voice data from the other party of the text call, converting the second voice data into second text data, and displaying the second text data in a second area of the text call interface.

[0021] According to an aspect of the present disclosure non-transitory computer readable storage medium storing instructions is provided. The instructions, when executed by at least one processor of the mobile terminal, may cause the mobile terminal to perform operations. The operations may comprise displaying a text call interface based on a text call function of the mobile terminal being enabled and a preset application of the mobile terminal being in a call state. The operations may comprise based on first text data being input into the text call interface, converting the first text data into first voice data to be sent by the preset application to an other party of the text call, and displaying the first text data in a first area of the text call interface. The operations may comprise based on receiving second voice data from the other party of the text call, converting the second voice data into second text data, and displaying the second text data in a second area of the text call interface.

[0022] The text call method and the text call device for the mobile terminal according to the embodiment of the present disclosure can ensure that the text data input by the user is correctly displayed in the speech area of the user when the text call function is enabled, and will not be erroneously displayed in the speech area of the other party of the call, thereby avoiding the display confusion on the text call interface and improving the user experience.

[0023] Additional aspects and / or advantages of general concept of the present disclosure will be set forth in part in the following description, and another part will be apparent from the description, or may be learned through the practice of the general concept of the present disclosure.

[0024] The above and other objects and features of the exemplary embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings that exemplarily illustrates embodiments, in which:

[0025] FIG. 1 is a diagram illustrating an example of an existing text call function;

[0026] FIG. 2 is a diagram illustrating an example of an audio processing architecture for executing a text call method for a mobile terminal according to an embodiment of the present disclosure;

[0027] FIG. 3 is a flowchart illustrating a text call method for a mobile terminal according to an embodiment of the present disclosure;

[0028] FIG. 4 is a diagram illustrating an example of enabling a text call function;

[0029] FIG. 5 is a diagram illustrating an example of determining one or more applications that support a text call function;

[0030] FIG. 6 is a diagram illustrating an example of enabling a mute function for second voice data;

[0031] FIG. 7 is a diagram illustrating an example of a text call interface displayed when a VoIP call is connected;

[0032] FIG. 8 is a diagram illustrating another example of an audio processing architecture for executing a text call method for a mobile terminal according to an embodiment of the present disclosure;

[0033] FIG. 9 is a diagram illustrating an example of disabling a text call interface according to an embodiment of the present disclosure;

[0034] FIG. 10 is a block diagram illustrating a text call device for a mobile terminal according to an embodiment of the present disclosure; and

[0035] FIG. 11 is a block diagram illustrating an electronic device according to an embodiment of the present disclosure.

[0036] Hereinafter, various embodiments of the present disclosure are described with reference to the accompanying drawings, in which like reference numerals are used to depict the same or similar elements, features, and structures. However, the present disclosure is not intended to be limited by the various embodiments described herein to a specific embodiment and it is intended that the present disclosure covers all modifications, equivalents, and / or alternatives of the present disclosure, provided they come within the scope of the appended claims and their equivalents. The terms and words used in the following description and claims are not limited to their dictionary meanings, but, are merely used to enable a clear and consistent understanding of the present disclosure. Accordingly, it should be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustration purpose only and not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.

[0037] It is to be understood that the singular forms "a," "an," and "the" include plural forms as well, unless the context clearly dictates otherwise. The terms "include," "comprise," and "have", used herein, indicate disclosed functions, operations, or the existence of elements, but does not exclude other functions, operations, or elements.

[0038] For example, the expressions "A or B," or "at least one of A and / or B" may indicate A and B, A, or B. For instance, the expression "A or B" or "at least one of A and / or B" may indicate (1) A, (2) B, or (3) both A and B.

[0039] In various embodiments of the present disclosure, it is intended that when a component (for example, a first component) is referred to as being "coupled" or "connected" with / to another component (for example, a second component), the component may be directly connected to the other component or may be connected through another component (for example, a third component). In contrast, when a component (for example, a first component) is referred to as being "directly coupled" or "directly connected" with / to another component (for example, a second component), another component (for example, a third component) does not exist between the component and the other component.

[0040] The expression "configured to", used in describing various embodiments of the present disclosure, may be used interchangeably with expressions such as "suitable for," "having the capacity to," "designed to," "adapted to," "made to," and "capable of", for example, according to the situation. The term "configured to" may not necessarily indicate "specifically designed to" in terms of hardware. Instead, the expression "a device configured to" in some situations may indicate that the device and another device or part are "capable of" For example, the expression "a processor configured to perform A, B, and C" may indicate a dedicated processor (for example, an embedded processor) for performing a corresponding operation or a general purpose processor (for example, a central processing unit (CPU) or an application processor (AP)) for performing corresponding operations by executing at least one software program stored in a memory device.

[0041] The terms used herein are to describe certain embodiments of the present disclosure, but are not intended to limit the scope of other embodiments. Unless otherwise indicated herein, all terms used herein, including technical or scientific terms, may have the same meanings that are generally understood by a person skilled in the art. In general, terms defined in a dictionary should be considered to have the same meanings as the contextual meanings in the related art, and, unless clearly defined herein, should not be understood differently or as having an excessively formal meaning In any case, even terms defined in the present disclosure are not intended to be interpreted as excluding embodiments of the present disclosure.

[0042] A text call method and a text call device for a mobile terminal according to embodiments of the present disclosure will be described in detail with reference to FIGs. 2 to 11.

[0043] FIG. 2 is a diagram illustrating an example of an audio processing architecture for executing a text call method for a mobile terminal according to an embodiment of the present disclosure.

[0044] Referring to FIG. 2, the audio processing architecture may include an application layer (Application), an application framework layer (Framework), and a hardware abstraction layer (HAL).

[0045] As shown in FIG. 2, various applications related to audio processing run at the application layer, such as a VoIP APP and a Controller APP shown in the figure. In addition, various applications related to audio processing further include an audio player, a Call App, a sound setting, a recorder, etc.

[0046] According to the embodiment of the present disclosure, the text call method for the mobile terminal according to an embodiment of the present disclosure is executed by constructing the Controller APP as a system level application for the mobile terminal. Herein, the Controller APP can be the system level APP specifically developed for executing the text call method, or it can be an existing system level APP on the mobile terminal for executing the text call method.

[0047] The VoIP App may include an APP playback (APP Playback) module and an APP recording (APP Recording) module. The VoIP APP receives voice data from the other party in a call and transmits the received voice data as playback data (e.g. audio data played through a speaker of the mobile terminal) to the application framework layer through the APP Playback module. The APP Recording module is used to obtain recording data (e.g. audio data inputted by a user recorded through a microphone of the mobile terminal) from the application framework layer and transmit the recording data to the other party in the call through the VoIP APP.

[0048] The Controller APP includes a text input interface (Input words), a text output interface (Output words), a TTS (Text to Speech) module, an ASR (Automatic Speech Recognition) module, an AudioTrack module, and an AudioRecord module. The text input interface is used to receive text data input by a user of the mobile terminal in a text call interface, the TTS module is used to convert the text data into audio data, and the AudioTrack module serves as an API (application programming interface) to transmit the audio data as playback data to the application framework layer. The AudioRecord module serves as an API to obtain the recording data from the application framework layer, the ASR module is used to convert the recording data into the text data, and the text output interface is used to display the text data on the text call interface.

[0049] The application framework layer is a core component for processing the audio data in an operating system of the mobile terminal (such as but not limited to an Android system). It allows applications to capture, play, process, and render the audio data to support various audio applications, such as the audio player, the speech recognition, the audio recording, the text call, etc. The application framework layer may include an AudioFlinger module, which is responsible for managing and scheduling the transmission of the audio data. The AudioFlinger module is responsible for the audio data transmission between the VoIP APP and the Controller APP. The AudioFlinger module may further include a PlaybackThread submodule and a RecordThread submodule. The PlaybackThread submodule is responsible for transmitting the playback data from the application layer to the hardware abstraction layer, while the RecordThread submodule is responsible for transmitting the recording data from the hardware abstraction layer to the application layer.

[0050] The hardware abstraction layer is generally responsible for interacting with a hardware apparatus such as but not limited to the speaker, the microphones, etc. The hardware abstraction layer may include a Remote Submix HAL module, which is used to implement an audio data connection between the VoIP APP and the Controller APP. The Remote Submix HAL module may further include MonoPipe submodules, each of which may connect a playback data stream with a recording data stream.

[0051] The Controller APP includes an AudioManager module, which is an audio control interface. The Controller APP may control whether to provide audio data streams of the VoIP APP and the Controller APP to the Remote Submix HAL module, as well as whether to connect the audio data streams of the VoIP APP and the Controller APP, by setting corresponding parameters through the AudioManager module.

[0052] The AudioFlinger module in the application framework layer may include two playback thread submodules (PlaybackThread 1 and PlaybackThread 2) and two recording thread submodules (RecordThread 1 and RecordThread 2). PlaybackThread 1 is used to transmit the audio data converted from the text data output by the Controller APP to the hardware abstraction layer, while RecordThread 1 is used to transmit the audio data converted from text data obtained from the hardware abstraction layer to the VoIP APP. PlaybackThread 2 is used to transmit the audio data obtained from the other party of the call output by the VoIP APP to the hardware abstraction layer, while RecordThread 2 is used to transmit the audio data obtained from the other party of the call obtained from the hardware abstraction layer to the Controller APP.

[0053] According to the embodiment of the present disclosure, the application framework layer may further include an AudioPolicy module and an AudioContinuity module, and the AudioFlinger module may further include a SecAudioParamFlinger submodule. The SecAudioParamFlinger submodule is used to parse the parameters set through the AudioManager submodule and provide the parsed parameters to the AudioContinuity module. The AudioContinuity module is specifically designed to store the parameters set through the AudioManager submodule for use by the AudioPolicy module. The AudioPolicy module uses the parameters stored in the AudioContinuity module to control which device and corresponding HAL layer to be selected for the audio data in the mobile terminal. In other words, the AudioPolicy module may select a voice data transmission path for the VoIP APP and the Controller APP based on the parameters stored in the AudioContinuity module.

[0054] The Remote Submix HAL module in the hardware abstraction layer may use two audio pipeline submodules (MonoPipe 1 and MonoPipe 2). MonoPipe 1 is used to connect the playback data stream of the Controller APP (i.e., the audio data converted from the text data) with the recording data stream of the VoIP APP, in order to send the audio data converted from the text data to the other party in the call. MonoPipe 2 is used to connect the playback data stream of the VoIP APP (i.e., the voice data received from the other party in a call) with the recording data stream of the Controller APP, so that the Controller APP may output the text data converted from the voice data.

[0055] According to the embodiment of the present disclosure, the AudioTrack module, PlaybackThread 1, MonoPipe 1, RecordThread 1, and the APP Recording module may form a first dedicated voice data transmission path (shown by blue lines in FIG. 2), which is dedicated to transmitting only the audio data converted from the text data from the text call APP to the VoIP APP; the APP Playback module, PlaybackThread 2, MonoPipe 2, RecordThread 2, and the AudioRecord module may form a second dedicated voice data transmission path (shown by purple lines in FIG. 2), which is dedicated to transmitting the voice data received by the VoIP APP from the other party of the call from the VoIP APP to the Controller APP.

[0056] FIG. 3 is a flowchart illustrating a text call method for a mobile terminal according to an embodiment of the present disclosure. The text call method according to the embodiment of the present disclosure may be executed by a Controller APP as described above.

[0057] Referring to FIG. 3, in step S301, in response to a text call function being enabled and a preset application (i.e. a VoIP APP) of the mobile terminal being in a call state, a text call interface is displayed.

[0058] According to the embodiment of the present disclosure, a user of the mobile terminal (i.e., a user A) may enable the text call function through operations shown in FIG. 4. FIG. 4 is a diagram illustrating an example of enabling a text call function. As shown in FIG. 4, in response to the user's operation, the text call function (i.e., a Text call) may transition from a disabling state shown on the left side of FIG. 4 to the enabling state shown on the right side of FIG. 4. Herein, the user's operation may include but are not limited to a sliding operation, a clicking operation, a gesture operation, etc.

[0059] Furthermore, when the text call function is enabled, a text call settings interface may be entered to determine one or more applications that support the text call function. FIG. 5 is a diagram illustrating an example of determining one or more applications that support a text call function. When the text call function is enabled, at least one VoIP APP that supports the text call function may be selected. Herein, in response to the user's selection, WeChat can be identified as the application that supports the text call function. On the other hand, in response to the user's selection, the applications such as QQ, Wecom, DingDing, and Tencent Meeting may be identified as the applications that do not support the text call function. Alternatively, after determining one or more applications that support the text call function, a list of APPs that support text call function (i.e., mPackageList) may be automatically generated. Alternatively, when the text call function is enabled, all VoIP APPs may be automatically selected as the applications that support the text call function.

[0060] Alternatively, when the text call function is enabled, it may be determined whether to enable a mute function (Mute other person's voice) for second voice data obtained from the VoIP APP (i.e., the voice data received by the VoIP APP from the other party of a voice call) after entering the text call settings interface. FIG. 6 is a diagram illustrating an example of enabling a mute function. As shown in FIG. 6, in response to the user's operation, the mute function can transition from a disabling state shown on the left side of FIG. 6 to an enabling state shown on the right side of FIG. 6. When the text call function is enabled, the mute function may be enabled or disabled. When the mute function is disabled, voice of the other party (i.e., a user B) receiving the call during a VoIP call may still be played through a speaker, that is to say, the user A may hear the voice of the user B while watching the text converted from the voice. On the other hand, when the mute function is enabled, the user A will not hear the voice of the user B, but only see the text converted from the voice of the user B.

[0061] According to the embodiment of the present disclosure, when the text call function is enabled, the Controller APP may register with an operating system of the mobile terminal to monitor audio status events, so as to monitor audio status of the application; when it is monitored that the audio status of the preset application changes, it may be determined that the preset application is in the call state. Herein, the application refers to the VoIP APP installed on the mobile terminal, but the present disclosure is not limited to this. In this way, in the state that the text call function is enabled, when the preset application is in the call state, it may be determined whether the preset application supports the text call function based on mPackageList. If the preset application supports the text call function, the text call interface may be displayed. If the preset application does not support the text call function, a normal VoIP call may continue to be executed. Alternatively, in the stated that the text call function is enabled, the text call interface may be directly displayed when the preset application supports the text call function, or displayed according to the user's selection (such as touching a specific icon).

[0062] According to the embodiment of the present disclosure, when the preset application enters the call state, the text call function of the preset application may be enabled in response to a preset operation of the user, thereby displaying the text call interface.

[0063] FIG. 7 is a diagram illustrating an example of a text call interface displayed when a VoIP call is connected. As shown in FIG. 7, the user may input text in the text input box of the text call interface, and when a send button (Send) is selected, a speech conversion is performed on text data. Alternatively, the text call interface may also be displayed in a floating window form.

[0064] Below, operations performed in the audio processing architecture when displaying the text call interface will be described with reference to FIG. 2.

[0065] In S1, while displaying the text call interface, the Controller APP may set parameters through the AudioManager module. Specifically, the set parameters may include package name information of the Controller APP, UIDs of applications that support the text call function, an audio continuity mode, etc. Herein, the audio continuity mode may include: CONTINUITY_APP_NONE, which indicates stopping audio continuity and restoring a normal state; CONTINUITY_APP_ALL, which indicates that the first dedicated voice data transmission path and the second dedicated voice data transmission path will be established separately, and an audio data stream output from the VoIP APP to the speaker will be truncated (i.e., the mute function is enabled), and the audio data stream input from the microphone to the VoIP APP will be truncated; and CONTINUITY_APP_ALL_AND_PLAY, which indicates that the first dedicated voice data transmission path and the second dedicated voice data transmission path will be established separately, and the audio data stream output from the VoIP APP to the speaker will be retained (i.e., the mute function is disabled), while the audio data stream input from the microphone to the VoIP APP will be truncated. That is to say, when the mute function is not enabled, the Controller APP sets the audio continuity mode to CONTINUITY_APP_ALL_AND_PLAY; and when the mute function is enabled, the Controller APP sets the audio continuity mode to CONTINUITY_APP_ALL.

[0066] In S2, the SecAudioParamFlinger submodule of the AudioFlinger module parses the parameters set through the AudioManager module. As a non-limiting example, the parameters set by the AudioManager module may include A, B, and C. Herein, the parameter A indicates that the set parameters are used for the audio continuity mode; the parameter B indicates UID information of VoIP APPs that allow the use of the text call function. Only when the VoIP APP has a UID represented by B, the first dedicated voice data transmission path and the second dedicated voice data transmission path may be established to transmit audio data between the VoIP APP and the Controller APP. The parameter C indicates whether the audio continuity mode is CONTINUITY_APP_ALL or CONTINUITY_APP_ALL_AND_PLAY. When the parameter C represents CONTINUITY_APP_ALL, the mute function is enabled. When the parameter C represents CONTINUITY_APP_ALL_AND_PLAY, the mute function is disabled.

[0067] Furthermore, the parameter C is used to limit the audio data transmission between the Controller APP and the VoIP APP that is currently in a VoIP call mode and supports the text call function.

[0068] Next, in S3, the AudioContinuity module stores the parameters parsed by the SecAudioParamFlinger submodule. As described above, the AudioManager module provides an API interface for the VoIP APP and does not have storage capabilities. On the other hand, the AudioPolicy module and AudioFlinger module are core modules in the audio processing architecture and are not suitable for storing large amounts of data. Therefore, according to the embodiment of the present disclosure, the AudioContinuity module is constructed to store the parameters set through the AudioManager module. In this way, the AudioPolicy module and the AudioFlinger module can easily access the AudioContinuity module and obtain the parameters stored therein, thereby improving a stability of the audio processing architecture and cleanliness of codes.

[0069] In S4, the AudioPolicy module establishes the first dedicated voice data transmission path (indicated by blue lines in FIG. 2) and the second dedicated voice data transmission path (indicated by purple lines in FIG. 2) based on the parameters stored in the AudioContinuity module. Herein, as an example, S4 is described as immediately following S3. In fact, the operation of establishing the first dedicated voice data transmission path in S4 may be performed when the AudioTrack module is enabled in the Controller APP, and the operation of establishing the second dedicated voice data transmission path in S4 may be performed when the AudioRecord module is enabled in the Controller APP.

[0070] When establishing the first dedicated voice data transmission path, it may be first determined whether the running VoIP APP is an APP that supports the voice call function based on unique identification information of the running VoIP APP (such as the UID or the package name as described above) and a unique identification information list of the VoIP APPs that support the text call function which is constructed in advance as described above; then, in response to determining that the running VoIP APP is the application that supports the voice call function, the first dedicated voice data transmission path may be established for connecting the playback data stream of the audio data (i.e., first voice data) converted from the text data and the recording data stream of the running VoIP APP. Below, it is described by taking the unique identifier information as UID as an example.

[0071] More specifically, the AudioPolicy module may create a virtual audio device 1 in the Remote Submix HAL module based on the parameters obtained from the AudioContinuity module, and route a data stream of the first voice data to an output side of the virtual audio device 1. Then, the UID of the running VoIP APP is matched with the unique identification information list (such as but not limited to mPackageList as described above) of the VoIP APPs that support the text call function contained in the parameters. If the running VoIP APP exists in the unique identification information list of VoIP APPs that support the text call function, matching is successful; otherwise, the matching fails. Through UID matching, it may be ensured that only the VoIP APP that is in the VoIP call and supports the text call function may be successfully matched, and other VoIP APPs (including the VoIP APP that is in the VoIP call but does not support the text call function) will not be matched.

[0072] If the matching is successful, the AudioPolicy module and the AudioFlinger module may route the recording data stream of the VoIP APP to an input side of the virtual audio device 1 according to address of the virtual audio device 1 (the address assigned when creating the virtual audio device 1), at this point, an independent first dedicated voice data transmission path has been established for connecting the playback data stream of the Controller APP and the recording data stream of the VoIP APP. In this way, by establishing the independent first dedicated voice data transmission path, the Controller APP may only transmit its playback data to the corresponding VoIP APP, without transmitting its playback data to any other applications on the mobile terminal (including the Controller APP itself). If the matching fails, no processing will be performed.

[0073] Alternatively, a pair of virtual audio devices may be created for each VoIP APP that supports the text call function, thereby establishing independent first dedicated voice data transmission path and second dedicated voice data transmission path for each VoIP APP that supports the text call function, but the present disclosure is not limited to this.

[0074] On the other hand, when establishing the second dedicated voice data transmission path, it may be first determined whether the running VoIP APP is the APP that supports the voice call function based on the unique identification information of the running VoIP APP and the unique identification information list of the VoIP APPs that support the text call function which is constructed in advance as described above; then, in response to determining that the running VoIP APP is the APP that supports the voice call function, the second dedicated voice data transmission path may be established for connecting the playback data stream (i.e., a data stream of second voice data) of the running VoIP APP and the recording data stream of the Controller APP.

[0075] More specifically, the AudioPolicy module may create a virtual audio device 2 in the Remote Submix HAL module based on the parameters obtained from the AudioContinuity module, and route the recording data stream of the Controller APP to an input side of the virtual audio device 2. Then, as described above, the UID of the running VoIP APP may be matched with the unique identification information list of the VoIP APPs that support the text call function contained in the parameters (such as but not limited to mPackageList as described above). If the running VoIP APP exists in the unique identification information list of the VoIP APPs that support the text call function, the matching is successful; otherwise, the matching fails. Through the UID matching, it may be ensured that only the VoIP APP that is in the VoIP call and supports the text call function may be successfully matched, and other VoIP APPs (including the VoIP APP that is in the VoIP call but does not support the text call function) will not be matched.

[0076] If the matching is successful, the AudioPolicy module and the AudioFlinger module may route the playback data stream of the VoIP APP to an output side of the virtual audio device 2 according to address of the virtual audio device 2 (the address assigned when creating the virtual audio device 2), at this point, an independent second dedicated voice data transmission path has been established for connecting the recording data stream of the Controller APP and the playback data stream of the VoIP APP. In this way, by establishing the independent second dedicated voice data transmission path, the Controller APP may obtain the voice data received by the VoIP APP from the VoIP APP. If the matching fails, no processing will be performed.

[0077] Referring to FIG. 3 again, in step S302, in response to first text data being input into the text call interface, the first text data is converted into first voice data and transmitted to the preset application, and the first text data is displayed only in a first area of the text call interface. The first voice data will be sent by the preset application to the other party of the call (i.e., the user B). Herein, the first area may indicate an area of the text call interface used to display sent messages.

[0078] Below, referring to FIG. 2, the operations performed in the audio processing architecture when inputting the text data in the text call interface will be described.

[0079] In S5-1-a, the user A inputs the text data in the text call interface (e.g. "How's the weather today?"). In S5-1-b, the Controller APP converts the input text data into voice data or a voice file through the TTS module, and the converted voice data or voice file may be used as the playback data of the Controller APP. In S5-1-c, the Controller APP transmits the playback data to the AudioFlinger module through the AudioTrack module. In S5-1-d, the AudioFlinger module transmits the playback data to the Remote Submix HAL module through the PlaybackThread submodule. In S5-1-e, the Remote Submix HAL module stores the playback data in the MonoPipe submodule of the virtual audio device 1. In S5-1-f, the AudioFlinger module obtains the playback data from the MonoPipe submodule of the virtual audio device 1 through the RecordThread submodule, and transmits the playback data to the App Recording module of the VoIP APP, so that the VoIP APP may send the playback data to user B. In the above operations, only the voice data converted from the input text data may be transmitted to the preset application through the first dedicated voice data transmission path.

[0080] According to the embodiment of the present disclosure, the Controller APP may convert the text data into the voice data in the following two ways: (1) converting the text data into voice stream; and (2) converting the text data into a voice file. Accordingly, when the Controller APP converts the text data into the voice stream, Controller APP may enable audio playback to trigger the establishment of the first dedicated voice data transmission path. In a case where the Controller APP adopts the method of converting the text data into a voice file, when the Controller APP plays the voice file converted by the TTS module, it may trigger the establishment of the first dedicated voice data transmission path.

[0081] Referring to FIG. 3 again, in step S303, in response to obtaining the second voice data from the preset application, the second voice data is converted into second text data, and the second text data is only displayed in a second area of the text call interface. Here, the second voice data is received by the preset application from the other party of the call. The second area indicates an area of the text call interface used to display received messages.

[0082] Below, referring to FIG. 2, the operations performed in the audio processing architecture when the preset application receives the voice data from the other party in the call will be described.

[0083] In S5-2-a, the VoIP APP transmits the received voice data (e.g. "The weather is nice today.") to the Playback Thread submodule of the AudioFlinger module through the APP Playback module. In S5-2-b, when the mute function is disabled, the AudioFlinger module transmits the voice data (i.e. playback data) to the Remote Submix HAL module through the PlaybackThread submodule. Meanwhile, an original playback channel of the speaker is preserved, so that the VoIP APP may normally play the voice data of the user B through the speaker. In S5-2-c, the Remote Submix HAL module stores the playback data in the MonoPipe submodule of the virtual audio device 2. In S5-2-d, the RecordThread submodule of the AudioFlinger module obtains the playback data from the MonoPipe submodule of the virtual audio device 2 and transmits the playback data to the AudioRecord module of the Controller APP. In S5-2-e, the Controller APP converts the obtained playback data into text data through the ASR module. In S5-2-f, the Controller APP displays the converted text data on the text call interface. In the above operations, since the mute function is disabled, the voice data received by the VoIP APP may be obtained from the VoIP APP through the second dedicated voice data transmission path on the one hand, and the voice data received by the VoIP APP may be played through the speaker on the other hand, that is to say, the voice data received by the VoIP APP may be transmitted to the speaker for playback through a first general voice data transmission path (i.e., the original playback channel of the speaker).

[0084] In the audio processing architecture shown in FIG. 2, the original playback channel of the speaker is retained due to the mute function being disabled. On the other hand, FIG. 8 is a diagram illustrating another example of an audio processing architecture for executing a text call method for a mobile terminal according to an embodiment of the present disclosure. In the audio processing architecture shown in FIG. 8, the original playback channel of the speaker is disconnected due to the mute function being enabled. In this case, the voice data received by the VoIP APP may only be obtained from the VoIP APP through the second dedicated voice data transmission path, and the voice data received by the VoIP APP is not played through the speaker, that is to say, it is prohibited to transmit the voice data received by the VoIP APP to the speaker through the first general voice data transmission path (i.e., the original playback channel of the speaker).

[0085] Alternatively, the text call method according to the embodiment of the present disclosure may further include the following step: disabling the text call interface in response to the preset application exiting the call state. FIG. 9 is a diagram illustrating an example of disabling a text call interface according to an embodiment of the present disclosure. Referring to FIG. 9, the left side of FIG. 9 illustrates the enabled text call interface when the text call function is enabled and the preset application is in the call state, while the right side of FIG. 9 illustrates the disabled text call interface when the preset application exits the call state. When the text call interface is disabled, it is not allowed to input the text data on the text call interface. As a non-limiting example, disabling the text call interface may indicate that the text call interface is no longer displayed.

[0086] FIG. 10 is a block diagram illustrating a text call device for a mobile terminal according to an embodiment of the present disclosure.

[0087] Referring to FIG. 10, the text call device 1000 may include a display module 1010, a text data processing module 1020, and a voice data processing module 1030.

[0088] Specifically, in response to a text call function being enabled and a preset application of the mobile terminal being in a call state, the display module 1010 may display a text call interface. Alternatively, in response to the text call function being enabled, the text call device 1000 may register with an operating system of the mobile terminal to monitor audio status events, so as to monitor audio status of various applications on the mobile terminal. In response to monitoring that the audio status of the preset application changes, the text call device 1000 may determine that the preset application is in the call state. In addition, in response to the preset application exiting the call state, the display module 1010 may disable the text call interface.

[0089] In response to first text data being input into the text call interface, text data processing module 1020 may convert the first text data into first voice data and transmit the first voice data to the preset application, and display the first text data only in a first area of the text call interface. Herein, the first voice data will be sent by the preset application to the other party of the call, and the first area may indicate an area of the text call interface used to display sent messages. Alternatively, the text data processing module 1020 may only transmit the first voice data to the preset application and not transmit the first voice data to any other application on the mobile terminal. For example, the text data processing module 1020 may only transmit the first voice data to the preset application through a first dedicated voice data transmission path.

[0090] In response to obtaining second voice data from the preset application, the voice data processing module 1030 may convert the second voice data into second text data, and display the second text data only in a second area of the text call interface. Herein, the second voice data is received by the preset application from the other party of the call, and the second area may indicate an area of the text call interface used to display received messages. Alternatively, the voice data processing module 1030 may obtain the second voice data from the preset application through a second dedicated voice data transmission path, or the voice data processing module 1030 may obtain the second voice data from the preset application through the second dedicated voice data transmission path and play the second voice data through a speaker of the mobile terminal. In addition, in response to a mute function for the second voice data being enabled, the voice data processing module 1030 may only obtain the second voice data from the preset application through the second dedicated voice data transmission path, and not play the second voice data through the speaker transmitted to the mobile terminal.

[0091] The text call method for the mobile terminal according to the embodiment of the present disclosure may be written as a computer program and stored on a computer readable storage medium. When the computer program is executed by a processor, the text call method for the mobile terminal described above may be implemented. Examples of computer-readable storage media include: Read Only Memory (ROM), Random Access Programmable Read Only Memory (RROM), Electrically Erasable Programmable Read Only Memory (EEPROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, Hard Disk Drive (HDD), Solid State Drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or extremely fast digital (xD) cards), magnetic tapes, floppy disks, magneto-optical data storage devices, optical data storage devices, hard disks, solid state disks, or any other devices that are configured to store computer programs and any associated data, data files, and data structures in a non-transitory manner, and provide the computer programs and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer programs. In one example, the computer programs and any associated data, data files, and data structures are distributed on a networked computer system, so that the computer programs and any associated data, data files, and data structures are stored, accessed, and executed through one or more processors or computers in a distributed manner.

[0092] FIG. 11 is a block diagram illustrating an electronic device according to an embodiment of the present disclosure.

[0093] Referring to FIG. 11, the electronic device 1100 may include a memory 1101 and a processor 1102. The memory 1101 stores computer programs, which, when executed by the processor 1102, implement the text call method for the mobile terminal according to the embodiment of the present disclosure.

[0094] As an example, the electronic device 1100 may be a smartphone, a tablet device, a personal digital assistant, or other mobile electronic device capable of executing the aforementioned computer program. In the electronic device 1100, the processor 1102 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may further include analog processors, digital processors, microprocessors, multi-core processors, processor arrays, network processors, and the like. The processor 1102 can execute instructions or codes stored in the memory 1101, wherein the memory 1101 can also store data. Instructions and data may also be transmitted and received through a network via a network interface device, wherein the network interface device may use any known transmission protocol. The memory 1101 may be integrated with the processor 1102, for example, a RAM or a flash memory is arranged in an integrated circuit microprocessor or the like. In addition, the memory 1101 may include an independent device, such as an external disk drive, a storage array, or other storage device that can be used by any database system. The memory 1101 and the processor 1102 may be operatively coupled, or may communicate with each other, for example, through an I / O port, a network connection, or the like, so that the processor 1102 can read files stored in the memory.

[0095] In addition, the electronic device 1100 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, touch input device, etc.). All components of the electronic device 1100 may be connected to each other via a bus and / or network.

[0096] The text call method and the text call device for the mobile terminal according to the embodiment of the present disclosure can ensure that the text data input by the user is correctly displayed in the speech area of the user when the text call function is enabled, and will not be erroneously displayed in the speech area of the other party of the call, thereby avoiding the display confusion on the text call interface and improving the user experience.

[0097] Those skilled in the art will easily think of other embodiments of the disclosure after considering the description and carrying out the disclosure disclosed herein. The present disclosure aims to contain any variation, use or adaptive change of the present disclosure, this variation, use or adaptive change follows the general principles of the present disclosure, and includes well-known knowledge or conventional technical means in the technical field, which are not disclosed in the present disclosure. The description and embodiments are only considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0098] It should be understood that the present disclosure is not limited to the accurate structures described above and shown in the drawings, and various modifications and changes may be made without departing from scope thereof. The scope of the present disclosure is limited only by the appended claims and their equivalents.

Claims

1.A method for a text call by a mobile terminal (1100), the method comprising:displaying (S301) a text call interface based on a text call function of the mobile terminal being enabled and a preset application of the mobile terminal being in a call state;based on first text data being input into the text call interface, converting (S302) the first text data into first voice data to be sent by the preset application to an other party of the text call, and displaying the first text data in a first area of the text call interface; andbased on receiving second voice data from the other party of the text call, converting (S303) the second voice data into second text data, and displaying the second text data in a second area of the text call interface.2.The method of claim 1, further comprising:based on the text call function being enabled, monitoring audio status events of applications on the mobile terminal; andbased on detecting that an audio status of the preset application changes, determining that the preset application is in the call state.3.The method of claim 1, further comprises:transmitting the first voice data to only the preset application and not transmitting the first voice data to any other application on the mobile terminal.4.The method of claim 3, wherein the first voice data is transmitted to the preset application through a first dedicated voice data transmission path.5.The method of claim 1, wherein:the second voice data is obtained from the preset application through a second dedicated voice data transmission path, orthe second voice data is played through a speaker of the mobile terminal.6.The method of claim 5, wherein based on a mute function for the second voice data being enabled, the second voice data is obtained from the preset application through the second dedicated voice data transmission path, and the second voice data is not played through the speaker of the mobile terminal.7.The method of claim 1, further comprising:based on determining that the preset application is an application that supports the text call function, establishing at least one of a first dedicated voice data transmission path for transmitting the first voice data to the preset application, or a second dedicated voice data transmission path for obtaining the second voice data from the preset application.8.The method of claim 7, further comprising determining that the preset application is the application that supports the text call function, based on an unique identification information of the preset application being included in a unique identification information list for applications that support the text call function,wherein the unique identification information list for applications that support the text call function is constructed based on the unique identification information of the applications that support the text call function set by a user.9.The method of claim 7, wherein establishing the first dedicated voice data transmission path for transmitting the first voice data to the preset application comprises:creating a first virtual audio device, routing a data stream of the first voice data to an output side of the first virtual audio device, and routing a recording data stream of the preset application to an input side of the first virtual audio device.10.The method of claim 7, wherein establishing the second dedicated voice data transmission path for obtaining the second voice data from the preset application comprises:creating a second virtual audio device, routing a data stream of the second voice data to an input side of the second virtual audio device, and routing a playback data stream of the preset application to an output side of the second virtual audio device.11.The of claim 1, further comprising:creating an audio continuity module to store at least one of the following parameters: package name information of an application executing the text call; unique identification information for applications that support the text call function; information indicating whether exiting the text call function or not; and information indicating whether enabling a mute function for the second voice data while establishing a first dedicated voice data transmission path for the first voice data and a second dedicated voice data transmission path for the second voice data.12.The text call method of claims 1, wherein the first area indicates an area of the text call interface used to display sent messages, and the second area indicates an area of the text call interface used to display received messages.13.A mobile terminal (1100) comprising:memory (1101) storing instructions; andat least one processor (1102) operably coupled to the memory, wherein the instructions, when executed by the at least one processor, cause the mobile terminal to perform operations comprising:displaying (S301) a text call interface based on a text call function of the mobile terminal being enabled and a preset application of the mobile terminal being in a call state;based on first text data being input into the text call interface, converting (S302) the first text data into first voice data to be sent by the preset application to an other party of the text call, and displaying the first text data in a first area of the text call interface; andbased on receiving second voice data from the other party of the text call, converting (S303) the second voice data into second text data, and displaying the second text data in a second area of the text call interface.14.The mobile terminal of claim 13, wherein the operations further comprise at least one operation according to a method in one of claims 2 to 12.15.A non-transitory computer readable storage medium storing instructions which, when executed by at least one processor of a mobile terminal (1100), cause the mobile terminal to perform operations according to a method in one of claims 1 to 12.

Citation Information

Patent Citations

  • Text call method and text call device for mobile terminal

    CN119342139A

  • Heterogeneous data conversion system and method of the same

    KR1020150054561A

  • Mobile communication terminal and method for switching over to text chat during telephone conversation, and program stored in medium for executing the method

    KR1020170062693A

  • System for providing text-to-speech service

    KR102166264B1

  • Method and apparatus for live call text-to-speech

    US20160035343A1