An accessible call method, device, equipment, storage medium and product
By recognizing the language of the called terminal in the call system and translating audio and video data in real time, communication barriers caused by language differences are resolved, terminal device dependence and call signaling processes are simplified, and user experience and operational efficiency are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE GROUP DESIGN INST
- Filing Date
- 2026-01-30
- Publication Date
- 2026-06-05
AI Technical Summary
Existing call systems suffer from communication barriers due to language differences, and their reliance on video services increases the requirements for terminal devices, affecting user experience and increasing maintenance difficulty.
After the call is connected between the calling terminal and the called terminal, the system obtains the media stream information of the called terminal, identifies its language, and updates the translation module to the media stream path when necessary to perform real-time translation of audio and video data.
It enables simple and quick implementation of barrier-free calling, reduces dependence on terminal devices, simplifies call signaling processes, and improves user experience and operational efficiency.
Smart Images

Figure CN122160456A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of communications, specifically relating to a barrier-free calling method, apparatus, device, storage medium, and product. Background Technology
[0002] In existing calls, communication barriers arise due to the use of different languages or dialects by the two parties. To address this issue, 5G New Voice has introduced intelligent translation services. Currently, the intelligent translation process based on 5G New Voice is divided into video-based and data channel (DC)-based methods. However, regardless of whether it's based on video transmission or data channel (DC), it must be implemented based on terminal video. This makes the system require video services from the terminal, increasing dependence on terminal devices and thus affecting user experience. Furthermore, the call signaling process is relatively complex, involving multiple nodes, which increases the difficulty of maintenance during troubleshooting.
[0003] Therefore, a simple and quick method is needed to achieve barrier-free communication. Summary of the Invention
[0004] This application provides an accessible calling method that can quickly and easily achieve accessible calling.
[0005] In a first aspect, embodiments of this application provide an accessibility call method, the method comprising: after a call is connected between a calling terminal and a called terminal, acquiring first media stream information of the called terminal, wherein the first media stream information is call information sent by the called terminal to the calling terminal; determining a first call language of the called terminal based on the first media stream information, so as to determine whether accessibility call processing needs to be performed based on the first call language, wherein the first call language is the language used by the called user during the call; if so, updating a preset translation module to a media stream path, wherein the media stream path is used to transmit media stream information during the call between the calling terminal and the called terminal; translating the audio and video data in the first media stream information through the translation module in the media stream path, and sending the translation result to the calling terminal.
[0006] Secondly, embodiments of this application provide an accessibility calling device, comprising: a first acquisition module, configured to acquire first media stream information of the called terminal after a call is connected between a calling terminal and a called terminal, wherein the first media stream information is call information sent by the called terminal to the calling terminal; a first determination module, configured to determine a first call language of the called terminal based on the first media stream information, so as to determine whether accessibility calling processing needs to be performed based on the first call language, wherein the first call language is the language used by the called user during the call; a first update module, configured to update a preset translation module to the media stream path when accessibility calling processing needs to be performed, wherein the media stream path is used to transmit the media stream information during the call between the calling terminal and the called terminal; and a first translation module, configured to perform a translation operation on the audio and video data in the first media stream information through the translation module in the media stream path, and send the translation result to the calling terminal.
[0007] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0009] Fifthly, embodiments of this application provide a computer program product that, when executed by a processor, implements the steps of the method described in the first aspect.
[0010] In a sixth aspect, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0011] In this embodiment, after the caller terminal and the called terminal connect, the first media stream information of the called terminal is obtained. The first media stream information is the call information sent by the called terminal to the caller terminal. Based on the first media stream information, the first call language of the called terminal is determined to determine whether accessibility call processing needs to be performed. The first call language is the language used by the called user during the call. If so, a preset translation module is updated to the media stream path. The media stream path is used to transmit the media stream information during the call between the caller terminal and the called terminal. Through the translation module in the media stream path, the audio and video data in the first media stream information are translated, and the translation result is sent to the caller terminal, which can easily and quickly achieve accessibility call processing. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating an accessible calling method provided in an embodiment of this application; Figure 2 This is a flowchart illustrating the second accessible calling method provided in the embodiments of this application; Figure 3 This is a swimlane diagram of an accessible communication method provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an accessible communication system provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an accessible communication device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an accessible communication device provided in an embodiment of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0015] The barrier-free calling method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0016] Figure 1 This illustration shows an embodiment of an accessible calling method provided by the present invention. The method can be executed by an electronic device, which may include a server and / or a terminal device, wherein the terminal device may be, for example, an in-vehicle terminal or a mobile phone terminal. In other words, the method can be executed by software or hardware installed in an accessible calling device, and the method includes the following steps: Step 102: After the caller terminal and the called terminal connect, obtain the first media stream information of the called terminal.
[0017] The first media stream information is the call information sent by the called terminal to the calling terminal.
[0018] The entity executing the accessible calling method described in this application can be an accessible calling system, accessible calling software, or other entities. This application's embodiments will use an accessible calling system (hereinafter referred to as the calling system) as an example for illustration.
[0019] After the call is established between the calling terminal and the called terminal, the call system first obtains the first media stream information of the called terminal. The first media stream information is the call information sent by the called user of the called terminal to the calling user of the calling terminal. This call information is media stream information. The calling terminal is the terminal that actively makes the call, and the calling user is the user on the calling terminal side. The called terminal is the terminal that passively receives the call, and the called user is the user on the called terminal side.
[0020] After the calling terminal requests a call with the called terminal and the call is connected, the call system obtains the first media stream information sent by the called user at the called terminal. This first media stream information is determined based on the call information sent by the called user to the calling user, which can be "hello" or "hello", etc.
[0021] Furthermore, before initiating an accessibility call, the call system can first determine whether the calling terminal or user has subscribed to the accessibility call service, and only execute the accessibility call procedure after confirming that the calling terminal or user has subscribed. Therefore, the call system also includes a VoLTE AS network element, which can be used to determine whether the calling terminal or user has subscribed to the accessibility call service. The VoLTE AS network element can determine whether the calling terminal or user is an accessibility call terminal or user based on the user's subscription information, in order to proceed with subsequent services. When it is determined that an accessibility call is not needed, the user call procedure is consistent with the existing network; when it is determined that an accessibility call is needed, subsequent procedure steps are executed.
[0022] Furthermore, the call system also includes a called terminal detection module, which is installed on the called terminal. The call system obtains the first media stream information of the called terminal through this module. After receiving the media stream language detection notification from the VoLTE AS network element, the called terminal detection module performs media stream language detection and returns a detection response. The called terminal detection module reports a call detection event notification (Begin) to the VoLTE AS network element through the CSCF. The reported notification includes the calling number, the called number, and the SBC address. After receiving the reported notification, the VoLTE AS network element returns a response to the called terminal.
[0023] Specifically, the steps for the call system to obtain the first media stream information through the called terminal detection model can be divided into: setting up a packet capture tool to capture the RTP stream during the call; identifying the RTP stream to determine the type of each RTP message; and extracting RTP streams of a preset type, which is the RTP stream type related to the data sent by the user, i.e., the first media stream information. More specifically, the called terminal detection module uses a network packet capture tool, such as Wireshark, to capture the RTP stream. A filter is set in Wireshark, for example, using the following filter to only display RTP streams: udp.port == [RTP port]. After obtaining the packet capture results, the called terminal detection module finds the RTP packets in the capture results. RTP packets are usually transmitted in the UDP protocol and contain keywords such as sequence number, timestamp, and payload type. After identifying the RTP packets, the called terminal detection module uses Wireshark's "Stream" function to select a specific RTP stream, and then uses the "Export" function to export the data of that stream, identifying the exported data as the first media stream information.
[0024] Step 104: Determine the first call language of the called terminal based on the first media stream information, so as to determine whether accessibility call processing needs to be performed based on the first call language.
[0025] The first call language is the language used by the called user during the call.
[0026] After acquiring the first media stream information, the call system determines the first call language of the called user on the called terminal side based on the first media stream information. Then, it determines whether accessibility call processing is required based on the first call language. The first call language is the language used by the called user during the call, such as English or Chinese. Furthermore, the first call language determined by the call system can be not only different languages but also different regional languages of the same language, such as Henan dialect and Cantonese.
[0027] When determining the first call language based on the first media stream information, the call system can reconstruct the video or audio based on the first media stream information, and then perform speech recognition on the speech in the video or audio to obtain the speech recognition result, i.e., the first call language. Specifically, when the first media stream information is audio stream information, the called terminal detection module in the call system can use codecs such as G.711, G.729, etc. In Wireshark, the codec is determined according to the payload type in the RTP packet, and then an appropriate tool (such as FFmpeg or Audacity) is used to decode and play the audio; when the first media stream information is video stream information, it ensures that tools supporting the video encoding format can be used to process it according to the encoding type used (such as H.264), and finally plays the exported file through an audio or video player to obtain the reconstructed audio and video data. After obtaining the reconstructed audio and video, the called terminal detection module inputs the reconstructed audio and video into a pre-trained machine learning model, so that the model can determine the language type by analyzing the features in the audio and video stream and output the speech recognition result.
[0028] After determining the first call language, the called terminal detection module sends the first call language to the VoLTE AS network element, so that the VoLTE AS network element can determine whether accessibility call processing is required based on the first call language. Specifically, the VoLTE AS network element can determine whether the first call language is a preset language to be translated (e.g., English). If so, it determines that accessibility call processing is required.
[0029] In other words, the call system determines whether the first language of the call is a preset language that needs to be translated (the language to be translated). If so, it translates directly. The language to be translated is the language set in advance by the calling user. Therefore, when the system detects that the called user is using this language, it can determine whether it is a language that needs to be translated based on the calling user's pre-set information, and if it is, it determines that accessibility call processing needs to be performed.
[0030] Step 106: If so, update the preset translation module to the media stream path.
[0031] The media stream path is used to transmit media stream information when the calling terminal and the called terminal are having a call.
[0032] When it is determined that barrier-free call processing needs to be performed, the call system updates the preset translation module to the media stream path. The media stream path is the path used to transmit media stream information when the calling terminal and the called terminal are talking. The translation module is used to convert audio and video data in one language into audio and video data in another language, such as converting English audio and video data into Chinese audio and video data.
[0033] In other words, when the calling terminal and the called terminal are having a routine call, the call system transmits the media stream information during the call through the media stream path; when the calling terminal needs to make an accessible call, the call system updates the preset translation module into the media stream path so as to translate the audio and video data in the first media stream information during the transmission of the first media stream information.
[0034] Step 108: Translate the audio and video data in the first media stream information through the translation module in the media stream path, and send the translation result to the calling terminal.
[0035] After updating the translation module to the media stream path, the call system uses the translation module to translate the audio and video data in the first media stream information, obtains the translation result, and sends the translation result to the calling terminal. In other words, the call system first sends the first media stream information sent by the called terminal to the translation module via the media stream path, the translation module translates the audio and video data in the first media stream information to obtain the translation result, and then sends the translation result to the calling terminal via the media stream path to achieve barrier-free calling.
[0036] Specifically, when the translation module performs translation operations on the audio and video data in the first media stream information, the translation module first performs restoration processing on the first media stream information to obtain the original audio and video data sent by the called user during the call. Then, it performs translation operations on the original audio and video data to translate it into audio and video data in a preset language. The preset language is the language set by the calling terminal (for example, when the calling user sets the language to Chinese, the original audio and video data input by the called user will be converted into Chinese audio and video data), and the translated audio and video data is sent to the calling terminal.
[0037] More specifically, the call system also includes a translation control plane module. The VoLTE AS network element instructs the translation control plane module to continue translating signaling processing via call event control (Continue). Upon receiving the translation notification, the translation control plane module performs translation control processing through the pre-translation module according to the target language of the notification, and then instructs the translation media plane module to perform translation media processing according to the target call language via call event control (Continue).
[0038] Furthermore, when the call system acquires the first media stream information, it also acquires the attribute information of the first media stream. The attribute information includes the sequence number data of each RTP packet (used to detect packet loss), timestamp data (related to playback rate, helping to maintain time synchronization during audio or video playback), and payload type (indicating the type of encoding used). Then, the audio and video data characteristics are determined based on the attribute information. The audio and video data characteristics are used to characterize the feature information of the audio and video data at each time point. Finally, when the call system performs translation processing on the audio and video data, it can perform translation based on the audio and video data characteristics and the audio and video data to obtain the translation result. In this way, the translation result can ensure that the translated audio and video data is time-synchronized with the original audio and video data.
[0039] The barrier-free calling method provided in this embodiment of the invention first obtains the first media stream information of the called terminal after the call is connected between the calling terminal and the called terminal. The first media stream information is the call information sent by the called terminal to the calling terminal. Then, the first call language of the called terminal is determined based on the first media stream information to determine whether barrier-free calling processing needs to be performed. The first call language is the language used by the called user during the call. Next, when barrier-free calling processing needs to be performed, a preset translation module is updated to the media stream path. The media stream path is used to transmit the media stream information during the call between the calling terminal and the called terminal. Finally, the translation module in the media stream path translates the audio and video data in the first media stream information and sends the translation result to the calling terminal. Thus, when barrier-free calling processing is required, a translation module is added to the media stream path to translate the audio and video data of the called user and send it to the calling user, which can achieve barrier-free calling simply and quickly.
[0040] In one implementation, updating the preset translation module to the media stream path (step 106) can be performed using step A1: Step A1: Connect the translation module to the media stream path.
[0041] When updating the preset translation module to the media stream path, the call system can connect the translation module to the media stream path in series. In other words, there are multiple ways to update the translation module to the media stream path; the embodiment of this application updates the translation module by connecting it in series to the media stream path.
[0042] Specifically, the translation module in the call system is connected to the media stream path in series. The SBC (Single Controller) flows the media stream through the newly added translation module node in the path. The calling party's anchor node remains unchanged at the SBC end, and the called party's anchor node remains unchanged at the called party's SBC end. After the INVITE message from both the calling and called parties, a 183 message is sent to update the media stream path and add the translation media module node. The calling party sends a PRACK to confirm. The called party receives a PRACK response of 200 OK. After the calling party's call node is added, the calling party sends an UPDATE message to the called party, completing the addition of the called party's resource path.
[0043] The step of translating the audio and video data in the first media stream information through the translation module in the media stream path and sending the translation result to the calling terminal (step 108) can be executed via steps A2-A3: Step A2: The first media stream information sent by the called terminal is sent to the translation module through the media stream path, so that the translation module translates the audio and video data in the first media stream information and outputs the translation result.
[0044] After the translation module is connected to the media stream path, the call system sends the first media stream information sent by the called terminal to the translation module through the media stream path, so that the translation module can translate the audio and video data in the first media stream information, obtain the translated audio and video data, and output it as the translation result.
[0045] Step A3: Send the translation result output by the translation module to the calling terminal through the media stream path.
[0046] After the translation module outputs the translation result, the call system sends the translation result (translated audio and video data) to the calling terminal via the media stream path. Specifically, the call system performs media stream conversion processing on the translation result, i.e., the translated audio and video data, and sends the converted media stream information to the calling terminal so that the user at the calling terminal can receive the translated audio and video data.
[0047] Specifically, when performing language translation, the translation module translates the audio and video data sent by the called user into audio and video data in the language that the calling user expects to hear, and then sends the media stream of the translated audio and video data to the SBC network element of the calling terminal.
[0048] In one implementation, updating the preset translation module to the media stream path (step 106) can be performed by step B1: Step B1: Attach the translation module to the media stream path.
[0049] When updating the preset translation module to the media stream path, the call system can attach the translation module to the media stream path. That is, when updating the translation module to the media stream path, there are multiple ways to do so. In this embodiment, the translation module is updated to the media stream path by attaching it to the media stream path.
[0050] Specifically, the call system attaches the translation module to the media stream path. There is no call resource update information on either the calling terminal or the called terminal. The calling terminal remains anchored at the SBC node, and the SBC directly copies the media stream from the called terminal to the translation media module.
[0051] The step of translating the audio and video data in the first media stream information through the translation module in the media stream path and sending the translation result to the calling terminal (step 108) can be executed via steps B2-B3: Step B2: Copy the first media stream information and send it to the translation module so that the translation module can translate the audio and video data in the first media stream information and output the translation result.
[0052] After the translation module is attached to the media stream path, the SBC where the subscription is located copies the media stream to the translation module node, thus realizing the copying of the media stream. That is, the translation module is attached to the SBC side. The SBC copies the media stream in the call to the translation module so that the translation module can translate the audio and video data in the first media stream information, obtain the translated audio and video data, and output it as the translation result.
[0053] Step B3: Replace the first media stream information with the translation result, and send the translation result to the calling terminal.
[0054] After the translation module outputs the translation result, the call system replaces the first media stream information with the translation result and sends the translation result (translated audio and video data) to the calling terminal. Specifically, the call system performs media stream conversion processing on the translation result, i.e., the translated audio and video data, and sends the converted media stream information to the calling terminal so that the user at the calling terminal can receive the translated audio and video data.
[0055] Specifically, the first media stream information on the called terminal side is sent to the translation module after passing through the SBC. The translation module performs language translation, translating it into the language type required by the calling user. The translation module then sends the translated media stream information to the SBC network element on the calling terminal side, so that the SBC network element on the calling terminal side replaces the original first media stream with the translated media stream information, making it the media stream information output by the translation module.
[0056] More specifically, after the call is connected, the call system determines the terminal user's subscription information, issues a detection notification, and detects the call media stream information cleared by the Session Border Controller (SBC) through the calling terminal detection module or the called terminal detection module. It detects the language information of both parties and reports it to the call system in real time. If the call system determines that the languages of the two parties are inconsistent, it notifies the translation module to perform translation. Through the concatenation or duplication process of the media stream, this embodiment of the application realizes the language translation of the media stream. The translated media stream is sent to the subscribed terminal, and the terminal receives the translated media information, enabling barrier-free communication between the two parties. This embodiment meets the user's need for translation services without forcibly relying on video services, improves the user experience, and shortens the lengthy steps and number of nodes in the existing 5G Xintonghua translation service process, optimizing the efficiency of fault diagnosis.
[0057] Figure 2 This is a flowchart illustrating a second accessible calling method provided in one embodiment of this specification, as shown below. Figure 2 As shown, the schematic diagram includes: Step 202: After the caller terminal and the called terminal connect, obtain the first media stream information of the called terminal.
[0058] The first media stream information is the call information sent by the called terminal to the calling terminal.
[0059] Step 204: Determine the first call language of the called terminal based on the first media stream information, so as to determine whether accessibility call processing needs to be performed based on the first call language.
[0060] The first call language is the language used by the called user during the call.
[0061] Step 206: If yes, then connect the translation module in series to the media stream path, or attach the translation module to the media stream path.
[0062] The media stream path is used to transmit media stream information when the calling terminal and the called terminal are having a call.
[0063] Step 208: When the translation module is connected to the media stream path, the first media stream information sent by the called terminal is sent to the translation module through the media stream path, so that the translation module translates the audio and video data in the first media stream information and outputs the translation result.
[0064] Step 210: Send the translation result output by the translation module to the calling terminal through the media stream path.
[0065] Step 212: When the translation module is attached to the media stream path, the first media stream information is copied and sent to the translation module so that the translation module can translate the audio and video data in the first media stream information and output the translation result.
[0066] Step 214: Replace the first media stream information with the translation result, and send the translation result to the calling terminal.
[0067] In the embodiments described in the specification, the translation module is updated to the media stream path by means of serialization or side-by-side, which reduces the complexity of the call signaling process during the call and reduces the data transmission process, enabling barrier-free calls to be achieved simply and quickly.
[0068] In one implementation, determining the first call language of the called terminal based on the first media stream information (step 104) can be performed via steps C1-C2: Step C1: Perform a data restoration operation on the first media stream information to obtain the first audio and video data.
[0069] The first audio and video data is the audio and video data sent by the called user to the calling user.
[0070] When determining the first call language of the called terminal based on the first media stream information, the call system performs a data restoration operation on the first media stream information and determines the restoration result as the first audio and video data, which is the audio and video data sent by the called user to the calling user.
[0071] Step C2: Perform language recognition on the first audio and video data to determine the language of the first call.
[0072] After determining the first audio and video data, the call system can perform language recognition on the first audio and video data to determine the language used in the first audio and video data sent by the called user, i.e., the first call language.
[0073] Specifically, after determining the first audio and video data, the call system can also send the first audio and video data directly to the translation module when it is determined that barrier-free communication is required, so that the translation module can directly translate the first audio and video data.
[0074] In one implementation, the translation module in the media stream path translates the audio and video data in the first media stream information and sends the translation result to the calling terminal (step 108), which can be executed via steps D1-D5: Step D1: Obtain the second media stream information.
[0075] The second media stream information is the call information sent by the calling terminal to the called terminal.
[0076] When the translation module translates the audio and video data (i.e., the first audio and video data) in the first media stream information and sends the translation result to the calling terminal, the call system can also first obtain the second media stream information. This second media stream information is the call media stream information sent by the calling terminal (or calling user) to the called terminal (or called user). In other words, after the calling terminal and the called terminal establish a call, the call system not only obtains the first media stream information of the called terminal, but also obtains the second media stream information of the calling terminal.
[0077] Specifically, the call system installs a detection module not only on the called terminal side but also on the calling terminal side. In other words, the call system also includes a calling terminal detection module, which is installed on the calling terminal. The call system obtains the calling terminal's second media stream information through this module.
[0078] Step D2: Perform a data restoration operation on the second media stream information to obtain the second audio and video data.
[0079] The second audio and video data refers to the audio and video data sent by the calling user to the called user.
[0080] After acquiring the second media stream information, the call system performs a data restoration operation on the second media stream information and identifies the restoration result as the second audio and video data. This second audio and video data is the audio and video data sent by the calling user to the called user. Specifically, after identifying the second media stream information, the calling terminal detection module of the call system performs a data restoration operation on the second media stream information to obtain the second audio and video data.
[0081] Step D3: Perform language recognition on the second audio and video data to determine the second call language.
[0082] The second call language is the language used by the calling user during the call.
[0083] After determining the second audio and video data, the call system performs language recognition on the second audio and video data to determine the second call language, which is the language used by the calling user and the called terminal during the call, such as Chinese.
[0084] Specifically, the calling terminal detection module of the call system performs language recognition on the second audio and video data to obtain the second call language. After determining the second call language, the calling terminal detection module sends the determined second call language to the VoLTE AS network element, so that the VoLTE AS network element can perform further processing based on the second call language.
[0085] Step D4: Determine whether translation is required based on the first and second call languages.
[0086] After determining the first call language, the call system determines whether translation of the first audio / video data (or first media stream information) is needed based on the first and second call languages. Furthermore, the called terminal sends the first call language to the VoLTE AS network element, and the calling terminal sends the second call language to the VoLTE AS network element. Therefore, the VoLTE AS network element in the call system can determine whether translation is needed based on the first and second call languages.
[0087] Specifically, the call system can determine whether the first and second call languages are the same. If they are, no translation is needed; otherwise, translation is required. For example, if both the first and second call languages are Chinese, no translation is needed; if the first call language is Chinese and the second call language is English, translation is required. The call system can also determine if translation is required when the first call language is a preset fourth language and the second call language is a preset fifth language. The fourth and fifth languages are languages pre-set by the calling user. For example, if the calling user pre-sets the fourth language to English and the fifth language to Chinese, when the call system detects that the called user is using English and the calling user is using Chinese, it will automatically translate the audio / video data containing English input by the called user into Chinese audio / video data.
[0088] Step D5: If so, perform a language conversion operation on the first audio and video data to obtain the translation result.
[0089] After determining that a translation operation needs to be performed on the first audio and video data (or the first media stream information), the call system can perform a translation operation on the first audio and video data or the first media stream information and translate it into audio and video data in a preset language to obtain a translation result.
[0090] Specifically, when the data acquired by the call system is first audio and video data, the call system can directly perform language conversion on the first audio and video data through the translation module and obtain the translation result; when the data acquired by the call system is first media stream information, the call system restores the first media stream information through the translation module to obtain the first audio and video data, and then performs language conversion on the first audio and video data to obtain the translation result.
[0091] More specifically, the calling terminal detection module and the called terminal detection module complete the Mb interface media plane detection, parse out the corresponding user-level media stream information such as RTP / UDP, TRP / TCP, and MSRP / TCP, and classify the languages in the call to obtain the first and second call languages. Then, the calling terminal detection module and the called terminal detection module output the SBC node information and the language types of both parties to the VoLTE AS network element through the Mw interface. The VoLTE AS network element and the translation control plane module interact to instruct the translation module to complete the media stream anchoring, that is, the user media stream flows through the translation module. Finally, the VoLTE AS service network element determines the consistency of the languages of both parties in the call, and whether the language involved in the called party is consistent with the language list output by the accessibility service signed by the calling party. If the languages are inconsistent, the subsequent translation process is triggered.
[0092] Figure 3 This is a flowchart illustrating an embodiment of an accessible calling method provided in this specification, as shown below. Figure 3 As shown, the schematic diagram includes: Step 3.1: After the caller terminal and the called terminal connect, the called terminal detection module obtains the first media stream information of the called terminal.
[0093] The first media stream information is the call information sent by the called terminal to the calling terminal.
[0094] Step 3.2: The called terminal detection module determines the first call language based on the first media stream information.
[0095] The first call language is the language used by the called user during the call.
[0096] Step 3.3: The called terminal detection module sends the first call language to the VoLTE AS network element.
[0097] Step 3.4: After the caller terminal and the called terminal connect, the calling terminal detection module obtains the second media stream information of the calling terminal.
[0098] The second media stream information is the call information sent by the calling terminal to the called terminal. Step 3.5: The calling terminal detection module determines the second call language based on the second media stream information.
[0099] The second call language is the language used by the calling user when making a call.
[0100] Step 3.6: The calling terminal detection module sends the second call language to the VoLTE AS network element.
[0101] Step 3.7: The VoLTE AS network element determines whether accessibility call processing is required based on the first and second call languages; Step 3.8: When it is determined that accessibility call processing needs to be performed, the VoLTE AS network element updates the preset translation module to the media stream path.
[0102] The media stream path is used to transmit media stream information when the calling terminal and the called terminal are having a call.
[0103] Step 3.9: The calling terminal sends the first media stream information to the translation module through the media stream path.
[0104] Step 3.10: The translation module performs restoration processing on the first media stream information to obtain the first audio and video data.
[0105] The first audio and video data is the audio and video data sent by the called user to the calling user.
[0106] Step 3.11: The translation module processes the first audio and video data to obtain the translation result.
[0107] Step 3.12: The translation module sends the translation result to the calling terminal.
[0108] In the embodiments described in the specification, by determining whether accessibility call processing is required based on a first call language and a second call language, accessibility call processing is performed when required. This enables accessibility call processing based on the languages used by the calling and called users during the call, thereby improving the user experience.
[0109] In one implementation, steps E1-E2 can also be performed: Step E1: Receive the service subscription request sent by the calling terminal.
[0110] The service subscription request is used to request that the caller terminal activate the barrier-free calling service.
[0111] Before the calling terminal initiates an accessibility call, the calling system first receives a service subscription request sent by the calling terminal. This service subscription request is a request sent by the calling terminal to the calling system, which requests the calling system to activate the accessibility call service for the calling terminal.
[0112] In other words, before a calling terminal can make an accessible call, it must first activate the accessible calling service. Therefore, the calling terminal sends a service subscription request to the calling system to request the activation of the accessible calling service. When the calling system receives the service subscription request from the calling terminal, it activates the accessible calling service for the calling terminal so that the calling terminal can make an accessible call with the called terminal.
[0113] Step E2: Receive the language conversion request sent by the calling terminal.
[0114] The language conversion request is used to request the translation of audio and video data in a third call language into audio and video data in the second call language, wherein the third call language is the language that the calling user expects to be translated during the call.
[0115] After the caller terminal and the called terminal connect, the call system can receive a language conversion request sent by the caller terminal. The language conversion request is used to request the call system to translate the audio and video data of the third call language into the audio and video data of the second call language. The third call language is the language that the caller user expects to be translated during the call.
[0116] In other words, the calling user can pre-set the language they want to be translated during the call. For example, the calling user can pre-set the English in the call to be converted to Chinese. Therefore, the calling terminal can send a language conversion request to the call system to request that the language of the called terminal in the call be converted to the language they need to make, so as to make a barrier-free call.
[0117] Specifically, the call system can detect the first media stream information and the second media stream information in real time, and translate the first media stream information according to the determined first and second call languages to achieve barrier-free communication between the calling terminal and the called terminal. The call system can also receive language conversion requests sent by the calling terminal, and convert the audio and video data in the third language in the first media stream information (or the first audio and video data) into the audio and video data in the calling user's preset language according to the calling user's request.
[0118] Specifically, Figure 4This is a schematic diagram of an accessible calling system provided in an embodiment of this application. The accessible calling system includes a calling terminal detection module, a VoLTE AS network element, a called terminal detection module, a translation control plane module, and a translation module. The calling terminal detection module and the called terminal detection module communicate with service network elements using RESTFUL / HTTP messages to complete call language detection event notification; they also communicate with core network CSCF network elements using SIP messages to complete detection language signaling processing. The translation control plane module communicates with service network elements using RESTFUL / HTTP messages to complete language detection event notification, language detection event subscription, and capability invocation functions; it also communicates with core network CSCF network elements using SIP messages to complete signaling processing functions; and communicates with the translation module to complete media negotiation functions for language control. The translation module communicates with service network elements using RESTFUL / HTTP messages to complete media download functions; it communicates with the control plane module to complete media negotiation functions for language control; and it communicates with core network SBC network elements using RTP messages to complete media language overlay and synthesis functions.
[0119] Specifically, after the call is connected and begins, the VoLTE AS network element, based on the terminal subscription information, notifies the calling terminal detection module and the called terminal detection module to perform call language detection, returning the language type and "Media Capability C1," including audio and video media formats, SBC addresses, and ports. The calling and called terminal detection modules perform language detection on the call media stream and return the results to the VoLTE AS network element, carrying the language type and "Media Capability C1," including audio and video media formats, SBC addresses, and ports. The VoLTE AS network element compares the downlink language information received by the currently subscribed user with the subscribed language list. If the current downlink language does not match the subscribed list, it translates to the language corresponding to the subscribed list. The VoLTE AS network element notifies the translation control plane to perform translation, carrying the translated language type and "Media Capability C1," including audio and video media formats, SBC addresses, and ports. The translation control plane completes resource preparation and returns the address of the translated media plane, carrying "Media Capability U1," including the audio and video media formats, addresses, and ports of the translated media plane. The translation control plane notifies the translation module to perform translation, providing the audio / video media format, SBC address, and port information. The translated media stream is then continued into the call flow. The translation module returns a response. The translation control plane sends a disconnect message to the translation module. The translation module returns the result, releases resources, and ends the translation.
[0120] It should be noted that the accessible calling method provided in this application embodiment can be executed by an accessible calling device or a control module within that accessible calling device for executing the accessible calling method. This application embodiment uses an accessible calling device executing the accessible calling method as an example to illustrate the accessible calling device provided in this application embodiment.
[0121] Figure 5 This is a schematic diagram of the structure of an accessible communication device according to an embodiment of the present invention. Figure 5 As shown, the barrier-free communication device includes: a first acquisition module 502, a first determination module 504, a first update module 506, and a first translation module 508.
[0122] The first acquisition module 502 is used to acquire the first media stream information of the called terminal after the caller terminal and the called terminal have connected. The first media stream information is the call information sent by the called terminal to the caller terminal. The first determining module 504 is used to determine the first call language of the called terminal based on the first media stream information, so as to determine whether accessibility call processing needs to be performed based on the first call language, wherein the first call language is the language used by the called user when making a call. The first update module 506 is used to update the preset translation module to the media stream path when accessibility call processing is required. The media stream path is used to transmit media stream information when the calling terminal and the called terminal are making a call. The first translation module 508 is used to translate the audio and video data in the first media stream information through the translation module in the media stream path, and send the translation result to the calling terminal.
[0123] The barrier-free calling device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0124] The accessible communication device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0125] The barrier-free communication device provided in this application embodiment can achieve... Figures 1 to 4 The various processes implemented in the method embodiments are not described in detail here to avoid repetition.
[0126] Based on the same technical concept, embodiments of this application also provide an electronic device for performing the above-described barrier-free calling method. Figure 6 This is a schematic diagram of the structure of an electronic device to implement various embodiments of this application. The electronic device can vary significantly due to differences in configuration or performance, and may include a processor 602, a communications interface 604, a memory 606, and a communication bus 608. The processor 602, communications interface 604, and memory 606 communicate with each other via the communication bus 608. The processor 602 can call a computer program stored in the memory 606 and executable on the processor 602 to perform the following steps: After the call is connected between the calling terminal and the called terminal, the first media stream information of the called terminal is obtained. The first media stream information is the call information sent by the called terminal to the calling terminal. Based on the first media stream information, the first call language of the called terminal is determined, so as to determine whether accessibility call processing needs to be performed based on the first call language. The first call language is the language used by the called user when making a call. If so, the preset translation module will be updated to the media stream path, which is used to transmit media stream information when the calling terminal and the called terminal are talking. The translation module in the media stream path performs a translation operation on the audio and video data in the first media stream information and sends the translation result to the calling terminal.
[0127] In one implementation, updating the preset translation module to the media stream path includes: Connect the translation module to the media stream path; The step of translating the audio and video data in the first media stream information through the translation module in the media stream path and sending the translation result to the calling terminal includes: The first media stream information sent by the called terminal is sent to the translation module through the media stream path, so that the translation module translates the audio and video data in the first media stream information and outputs the translation result; The translation result output by the translation module is sent to the calling terminal via the media stream path.
[0128] In one implementation, updating the preset translation module to the media stream path includes: The translation module is attached to the media stream path; The step of translating the audio and video data in the first media stream information through the translation module in the media stream path and sending the translation result to the calling terminal includes: The first media stream information is copied and sent to the translation module, so that the translation module translates the audio and video data in the first media stream information and outputs the translation result; The first media stream information is replaced with the translation result, and the translation result is sent to the calling terminal.
[0129] In one implementation, determining the first call language of the called terminal based on the first media stream information includes: Perform a data restoration operation on the first media stream information to obtain the first audio and video data, which is the audio and video data sent by the called user to the calling user. Perform language recognition on the first audio and video data to determine the language of the first call.
[0130] In one implementation, the step of translating the audio and video data in the first media stream information through the translation module in the media stream path and sending the translation result to the calling terminal includes: Acquire second media stream information, which is the call information sent by the calling terminal to the called terminal; Perform a data restoration operation on the second media stream information to obtain second audio and video data, which is the audio and video data sent by the calling user to the called user; Perform language recognition on the second audio and video data to determine the second call language, which is the language used by the calling user during the call. Determine whether translation is required based on the first and second call languages; If so, then a language conversion operation is performed on the first audio and video data to obtain the translation result.
[0131] In one implementation, the method further includes: Receive a service subscription request sent by the calling terminal, the service subscription request being used to request the activation of the barrier-free calling service for the calling terminal; The system receives a language conversion request sent by the calling terminal. The language conversion request is used to request the translation of audio and video data in a third call language into audio and video data in a second call language. The third call language is the language that the calling user expects to be translated during the call.
[0132] The specific execution steps can be found in the various steps of the above-described barrier-free calling method embodiment, and can achieve the same technical effect. To avoid repetition, they will not be repeated here.
[0133] It should be noted that the electronic devices in the embodiments of this application include: servers, terminals, or other devices besides terminals.
[0134] The above electronic device structure does not constitute a limitation on the electronic device. An electronic device may include more or fewer components than illustrated, or combine certain components, or arrange them differently. For example, an input unit may include a Graphics Processing Unit (GPU) and a microphone, and a display unit may use a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar display panels. User input units include at least one of a touch panel and other input devices. A touch panel is also called a touchscreen. Other input devices may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be elaborated further here.
[0135] Memory can be used to store software programs and various data. Memory can primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area can store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, memory can include volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM).
[0136] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor, wherein the application processor mainly handles operations related to the operating system, user interface, and applications, while the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor.
[0137] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described barrier-free calling method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0138] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0139] This application also provides a computer program product. When the computer program product is executed by a processor, it implements the various processes of the above-described barrier-free calling method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0140] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described barrier-free calling method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0141] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0142] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0143] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0144] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A barrier-free calling method, characterized in that, include: After the call is connected between the calling terminal and the called terminal, the first media stream information of the called terminal is obtained. The first media stream information is the call information sent by the called terminal to the calling terminal. Based on the first media stream information, the first call language of the called terminal is determined, so as to determine whether accessibility call processing needs to be performed based on the first call language. The first call language is the language used by the called user when making a call. If so, the preset translation module will be updated to the media stream path, which is used to transmit media stream information when the calling terminal and the called terminal are talking. The translation module in the media stream path performs a translation operation on the audio and video data in the first media stream information and sends the translation result to the calling terminal.
2. The method according to claim 1, characterized in that, The step of updating the preset translation module to the media stream path includes: Connect the translation module to the media stream path; The step of translating the audio and video data in the first media stream information through the translation module in the media stream path and sending the translation result to the calling terminal includes: The first media stream information sent by the called terminal is sent to the translation module through the media stream path, so that the translation module translates the audio and video data in the first media stream information and outputs the translation result; The translation result output by the translation module is sent to the calling terminal via the media stream path.
3. The method according to claim 1, characterized in that, The step of updating the preset translation module to the media stream path includes: The translation module is attached to the media stream path; The step of translating the audio and video data in the first media stream information through the translation module in the media stream path and sending the translation result to the calling terminal includes: The first media stream information is copied and sent to the translation module, so that the translation module translates the audio and video data in the first media stream information and outputs the translation result; The first media stream information is replaced with the translation result, and the translation result is sent to the calling terminal.
4. The method according to claim 1, characterized in that, Determining the first call language of the called terminal based on the first media stream information includes: Perform a data restoration operation on the first media stream information to obtain the first audio and video data, which is the audio and video data sent by the called user to the calling user. Perform language recognition on the first audio and video data to determine the language of the first call.
5. The method according to claim 4, characterized in that, The step of translating the audio and video data in the first media stream information through the translation module in the media stream path and sending the translation result to the calling terminal includes: Acquire second media stream information, which is the call information sent by the calling terminal to the called terminal; Perform a data restoration operation on the second media stream information to obtain second audio and video data, which is the audio and video data sent by the calling user to the called user; Perform language recognition on the second audio and video data to determine the second call language, which is the language used by the calling user during the call. Determine whether translation is required based on the first and second call languages; If so, then a language conversion operation is performed on the first audio and video data to obtain the translation result.
6. The method according to claim 5, characterized in that, The method further includes: Receive a service subscription request sent by the calling terminal, the service subscription request being used to request the activation of the barrier-free calling service for the calling terminal; The system receives a language conversion request sent by the calling terminal. The language conversion request is used to request the translation of audio and video data in a third call language into audio and video data in a second call language. The third call language is the language that the calling user expects to be translated during the call.
7. An accessible communication device, characterized in that, include: The first acquisition module is used to acquire the first media stream information of the called terminal after the caller terminal and the called terminal have connected. The first media stream information is the call information sent by the called terminal to the caller terminal. The first determining module is used to determine the first call language of the called terminal based on the first media stream information, so as to determine whether accessibility call processing needs to be performed based on the first call language, wherein the first call language is the language used by the called user when making a call; The first update module is used to update the preset translation module to the media stream path when accessibility call processing is required. The media stream path is used to transmit media stream information when the calling terminal and the called terminal are making a call. The first translation module is used to translate the audio and video data in the first media stream information through the translation module in the media stream path, and send the translation result to the calling terminal.
8. An electronic device, characterized in that, The device includes: Processor; and A memory configured to store computer-executable instructions configured to be executed by the processor, the executable instructions including steps for performing the accessibility calling method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is used to store computer-executable instructions that cause the computer to perform the accessibility calling method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the accessibility call method according to any one of claims 1 to 6.