Voicemail Transcription Automatic Translation System

A system using NLP and machine learning translates voicemails into users' preferred languages, addressing misrepresentation issues in current translation tools, ensuring clear and timely message understanding.

KR1020260113031APending Publication Date: 2026-07-21T MOBILE US INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
T MOBILE US INC
Filing Date
2024-10-22
Publication Date
2026-07-21

Smart Images

  • Figure PCT00003_ABST
    Figure PCT00003_ABST
Patent Text Reader

Abstract

The disclosed technology includes a voicemail translation service for a telecommunications network. The voicemail translation service may receive a voicemail message being communicated from a sender to a recipient. The voicemail translation service calls an application programming interface (API) to upload the voicemail message to a transcription service, which completes the transcription of the voicemail message in a default language. In response to determining that the default language of the completed transcription does not match the recipient's target language, the voicemail translation service triggers another API to generate a translation of the transcription according to the target language. Subsequently, the voicemail translation service stores the translated transcription in a voicemail storage system accessible to the recipient.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Machine translation is the use of rule-based or probabilistic (e.g., statistical, and most recently, neural network-based) machine learning approaches to translate text or speech from one language to another, including the contextual, idiomatic, and pragmatic nuances of both languages.

[0002] Unedited machine translations are publicly available through tools on the internet such as Google Translate, Babylon, DeepL Translator, and StarDict. These tools generate approximate translations that "provide the gist" of the original text under favorable conditions. Through the internet, translation software can help non-native speakers understand web pages published in other languages. However, full-page translation tools have limited utility because they provide only a limited potential understanding of the original author's intent and context; furthermore, translated pages tend to be misrepresented and confusing rather than helpful for understanding. Brief explanation of the drawing

[0003] A detailed description of embodiments of the present invention will be described and explained with reference to the attached drawings. FIG. 1 is a block diagram illustrating a wireless communication system capable of implementing embodiments of the present technology. FIG. 2 is a block diagram illustrating a system for translating a voice message into a target language. Figure 3 is a flowchart illustrating a method performed by a voicemail translation service of a telecommunications network configured to translate voicemail messages. FIG. 4 is a block diagram illustrating an example of a computer system in which at least some of the operations described in this specification can be implemented. The technology described herein will become more apparent to those skilled in the art by reviewing the detailed description together with the drawings. Embodiments or embodiments describing aspects of the invention are illustrated by way of example, and the same reference numerals may denote similar elements. Although the drawings illustrate various embodiments for illustrative purposes, those skilled in the art will recognize that alternative embodiments may be used without departing from the principles of the technology. Accordingly, while specific embodiments are illustrated in the drawings, the technology is subject to various modifications. Specific details for implementing the invention

[0004] The disclosed technology includes a system that provides users with the ability to receive voicemail transcriptions in a user-specific language. The system may operate as a telecommunications service including communications between at least one pair of user devices (e.g., smartphones). The system includes a server that manages a process of delivering translated voicemail transcriptions to user devices. For example, the server may receive voicemail from a first smartphone and generate a transcription of the voicemail by calling an Application Programming Interface (API) to do so. The server may generate a translation of the transcription and provide the translated transcription to a second smartphone.

[0005] The system can provide users with automatically translated transcripts by transcribing voicemails using natural language processing (NLP) and machine learning. In one example, the system utilizes automatic speech recognition (ASR) technology to transcribe spoken words from voicemails into written text. After being trained on a large volume of speech data from multiple different languages, the system's model is capable of determining the language spoken in the voicemail by recognizing language-specific acoustic features. After transcribing the voicemail, NLP technology is applied to preprocess the text by correctly identifying words in the case of homonyms and correctly identifying the punctuation to be used in the voicemail transcription. Once transcription is complete, the server can compare the detected language with the user-selected language. In response to determining that the detected language does not match the user-selected language, the system can call an API to translate the transcription using translation models. In one example, the system utilizes neural networks trained for language translation, such as neural machine translation algorithms, to translate a transcript. The system can consider language-specific rules, such as grammatical rules and common expressions; thus, it increases the accuracy of the generated translations.

[0006] By calling an API to translate the transcript of a received voicemail, the system simplifies the transcription translation process. In addition, the API enables the system to interact with existing software and adapt easily to utilizing those software programs, thereby reducing development time. Overall, the API enables seamless communication and integration among multiple programs.

[0007] The disclosed technology addresses the challenges associated with bilingual or multilingual users receiving phone calls and voicemails in languages ​​other than their native tongues. Currently, there is no system that delivers voicemail transcripts in the user's preferred language. This can be problematic for busy users who are fluent in another language but have limited reading ability in that same language. This can lead to communication issues, potentially preventing users from understanding important messages. Additionally, by transcribing and translating voicemails, users who are busy with conference calls, meetings, or in public places where incoming voicemail cannot be played can easily read incoming voicemail transcripts in their chosen translated language while protecting their privacy. For example, a user might receive voicemails regarding medical test results. By receiving the transcribed message of the voicemail in a doctor's office, the user can determine whether they need to call back immediately. Thus, this technology enhances users' ability to communicate quickly and efficiently with others.

[0008] The description and associated drawings are exemplary examples and should not be construed as limiting. The present disclosure provides specific details to enable a full understanding and explanation of such examples. However, those skilled in the art will understand that the invention may be practiced without such details. Likewise, those skilled in the art will understand that, to avoid unnecessary obscurity in the description of the examples, the invention may include well-known structures or features that are not depicted or described in detail.

[0009] wireless communication system

[0010] FIG. 1 is a block diagram illustrating a wireless telecommunications network (100) (“Network (100)”) in which embodiments of the disclosed technology are integrated. The Network (100) includes base stations (102-1 to 102-4) (individually referred to as “Base Station (102)” or collectively as “Base Stations (102)”). A base station is a type of network access node (NAN) which may also be referred to as a cell site, a base station transceiver, or a radio base station. The Network (100) may include any combination of NANs including an access point, a radio transceiver, a gNodeB (gNB), a NodeB, an eNodeB (eNB), a home NodeB or a home eNodeB, etc. In addition to being a wireless wide area network (WWAN) base station, a NAN can be a wireless local area network (WLAN) access point, such as an IEEE (Institute of Electrical and Electronics Engineers) 802.11 access point.

[0011] The NANs of the network (100) formed by the network (100) also include wireless devices (104-1 to 104-7) (individually referred to as "wireless devices (104)" or collectively as "wireless devices (104)") and a core network (106). The wireless devices (104) may correspond to or include network (100) entities capable of communicating using various connectivity standards. For example, a 5G communication channel may use a millimeter wave (mmW) access frequency of 28 GHz or higher. In some embodiments, the wireless devices (104) may be operablely coupled to a base station (102) via a long-term evolution / long-term evolution-advanced (LTE / LTE-A) communication channel referred to as a 4G communication channel.

[0012] The core network (106) provides, manages, and controls security services, user authentication, access authorization, tracking, Internet Protocol (IP) connectivity, and other access, routing, or mobility functions. The base station (102) interfaces with the core network (106) through a first set of backhaul links (e.g., S1 interface) and can perform radio configuration and scheduling for communication with a wireless device (104) or operate under the control of a base station controller (not shown). In some examples, the base station (102) can communicate with each other directly or indirectly (e.g., through the core network (106)) through a second set of backhaul links (110-1 to 110-3) (e.g., X1 interface) which may be wired or wireless communication links.

[0013] A base station (102) can communicate wirelessly with a wireless device (104) through one or more base station antennas. Cell sites can provide communication coverage for geographical coverage areas (112-1 to 112-4) (also referred to individually as "coverage area (112)" or collectively as "coverage areas (112)"). The coverage area (112) for a base station (102) may be divided into sectors that constitute only a part of the coverage area (not shown). The network (100) may include different types of base stations (e.g., macro cell base stations and / or small cell base stations). In some embodiments, there may be overlapping coverage areas (112) for different service environments (e.g., Internet of Things (IoT), mobile broadband (MBB), vehicle-to-everything (V2X), machine-to-machine (M2M), machine-to-everything (M2X), ultra-reliable low-latency communication (URLC), machine-type communication (MTC)).

[0014] The network (100) may include a 5G network (100) and / or an LTE / LTE-A or other network. In an LTE / LTE-A network, the term "eNBs" is used to describe base stations (102), and in 5G new radio (NR) networks, the term "gNBs" is used to describe base stations (102) that may include mmW communications. Thus, the network (100) may form a heterogeneous network (100) in which different types of base stations provide coverage for various geographical areas. For example, each base station (102) may provide communication coverage for a macro cell, a small cell, and / or other types of cells. As used herein, the term "cell" may, depending on the context, relate to a base station, a carrier or component carrier associated with the base station, or a coverage area (e.g., a sector) of the carrier or base station.

[0015] Macro cells generally cover a relatively wide geographical area (e.g., a radius of several kilometers) and can allow access by wireless devices subscribed to the wireless network (100) service provider. As previously described, small cells are low-power base stations compared to macro cells and can operate in frequency bands that are the same or different (e.g., licensed, unlicensed) as those of macro cells. Examples of small cells include pico cells, femto cells, and micro cells. Generally, pico cells can cover a relatively smaller geographical area and can allow unrestricted access by wireless devices subscribed to the network (100) provider. Femto cells cover a relatively smaller geographical area (e.g., a home) and can provide restricted access by wireless devices associated with the femto unit (e.g., wireless devices within a closed subscriber group (CSG), wireless devices for users within the home). A base station can support one or multiple (e.g., two, three, four, etc.) cells (e.g., component carriers). All fixed transceivers mentioned in this specification that can provide access to the network (100) are NANs that include small cells.

[0016] A communication network accommodating the various examples disclosed may be a packet-based network operating according to a layered protocol stack. In the user plane, communication at the bearer or PDCP (Packet Data Convergence Protocol) layer may be IP-based. Subsequently, the Radio Link Control (RLC) layer performs packet segmentation and reassembly to communicate over a logical channel. The Medium Access Control (MAC) layer may perform priority handling and multiplexing of the logical channel into a transport channel. Additionally, the MAC layer may provide retransmission at the MAC layer using Hybrid ARQ (HARQ) to improve link efficiency. In the control plane, the Radio Resource Control (RRC) protocol layer provides the establishment, configuration, and maintenance of an RRC connection between a wireless device (104) and a base station (102) or core network (106) supporting a radio bearer for user plane data. In the physical (PHY) layer, the transport channel is mapped to a physical channel.

[0017] Wireless devices may be integrated with or embedded in other devices. As illustrated, wireless devices (104) are distributed across the entire network (100), and each wireless device (104) may be stationary or mobile. For example, wireless devices may include handheld mobile devices (104-1, 104-2) (e.g., smartphones, portable hotspots, tablets, etc.); laptops (104-3); wearables (104-4); drones (104-5); vehicles with wireless connectivity (104-6); head-mounted displays (104-7) with wireless augmented reality / virtual reality (AR / VR) connectivity; portable game consoles; wireless routers, gateways, modems, and other stationary-wireless access devices; wirelessly connected sensors that provide data to a remote server over the network; IoT devices such as wirelessly connected smart home appliances; etc.

[0018] Wireless devices (e.g., wireless devices (104)) may be referred to as user equipment (UE), customer premise equipment (CPE), mobile station, subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, handheld mobile device, remote device, mobile subscriber station, terminal equipment, access terminal, mobile terminal, wireless terminal, remote terminal, handset, mobile client, client, etc.

[0019] A wireless device can communicate with various types of base stations and network (100) equipment at the edge of the network (100), including macro eNB / gNB, small cell eNB / gNB, relay base stations, etc. The wireless device can also communicate with other wireless devices within or outside the same coverage area of ​​the base station through device-to-device (D2D) communications.

[0020] Communication links (114-1 to 114-9) illustrated in the network (100) (also referred to individually as "communication links (114)" or collectively as "communication links (114)") include uplink (UL) transmissions from a wireless device (104) to a base station (102) and / or downlink (DL) transmissions from the base station (102) to a wireless device (104). Downlink transmissions may also be forward link transmissions, and uplink transmissions may also be reverse link transmissions. Each communication link (114) includes one or more carriers, wherein each carrier may be a signal composed of a plurality of subcarriers (e.g., waveform signals of different frequencies) modulated according to various radio technologies. Each modulated signal may be transmitted over a different subcarrier and may carry control information (e.g., reference signals, control channels), overhead information, user data, etc. The communication link (114) can transmit bidirectional communications using frequency division duplex (FDD) (e.g., using paired spectrum resources) or time division duplex (TDD) operation (e.g., using unpaired spectrum resources). In some embodiments, the communication links (114) include LTE and / or mmW communication links.

[0021] In some embodiments of the network (100), base stations (102) and / or wireless devices (104) include a plurality of antennas for employing antenna diversity schemes to improve communication quality and reliability between base stations (102) and wireless devices (104). Additionally or alternatively, base stations (102) and / or wireless devices (104) may employ multiple-input, multiple-output (MIMO) techniques that can utilize multipath environments to transmit multiple spatial layers carrying the same or different coded data.

[0022] In some examples, the network (100) implements 6G technologies including increased density or diversification of network nodes. The network (100) can enable terrestrial and non-terrestrial transmissions. In this context, a non-terrestrial network (NTN) is enabled by one or more satellites, such as satellites (116-1, 116-2), to provide services anytime and anywhere and to provide coverage of areas that are not reachable by any conventional terrestrial network (TN). The 6G implementation of the network (100) can support terahertz (THz) communication. This can support wireless applications requiring ultra-high quality of service (QoS) requirements and multi-terabit data transmission per second in the 6G and beyond eras, such as terabit-per-second backhaul systems, ultra-high definition content streaming between mobile devices, AR / VR, and wireless high-bandwidth secure communications. In another example of 6G, the network (100) may implement a converged Radio Access Network (RAN) and core architecture to achieve Control and User Plane Separation (CUPS) and achieve extremely low user plane latency. In yet another example of 6G, the network (100) may implement a converged Wi-Fi and core architecture to increase and improve indoor coverage.

[0023] Automatic translation system

[0024] FIG. 2 illustrates an automatic translation system (200) comprising a voicemail application server (202) configured to receive and manage requests for transcribing and updating voicemails. The voicemail application server (202) includes processing components, software, and other components configured to connect with other devices and systems and exchange data via the Internet or other communication networks. As such, the voicemail application server (202) may be connected to a mobile device (204), a voicemail profile database (206), a voicemail message repository (208), a short message service center (214), a message transfer agent (MTA) (216), an IP Multimedia Subsystem (IMS) network (222), an object storage service, a transcription service (226), a translation service (228), and a mobile device (230). As such, users may receive voicemail transcriptions in their preferred language.

[0025] Voicemail is a digital record message that a first user (e.g., the calling party) leaves for a second user (e.g., the called party) when the second user is unable to answer a phone call. It serves as a method for leaving a verbal message when the recipient is absent or chooses not to answer the call. Voicemail messages are typically stored in the recipient's voicemail box on a voicemail system hosted by a telecommunications service provider. A voicemail application server (202) can receive voicemail messages transmitted via a telecommunications network from a first user device, such as a mobile device (204) associated with the first user (e.g., user (232)), to a second user device, such as a mobile device (230) associated with the second user (e.g., user (234)). For example, the mobile device (204) can initiate a phone call to the mobile device (230). Telephone calls are routed through an IMS network, such as an IMS network (222) associated with a telecommunications network. An IMS network refers to an architectural framework in telecommunications for delivering telephone calls. For example, when a user (232) initiates a call to a user (234), the IMS network (222) can manage the exchange of information between a mobile device (204) and a mobile device (230) based on the Session Initiation Protocol (SIP) and the Real-time Transport Protocol (RTP). This enables the telecommunications network to initiate and maintain real-time multimedia communication sessions through IP networks, such as Wi-Fi and cellular networks. The IMS network (222) enables the delivery of voicemails to a voicemail application server (202) through IP-based networks.The IMS network (222) can also prioritize voice traffic. This enables the voicemail application server (202) to receive voicemails with high-quality audio.

[0026] The IMS network (222) can determine when a call from the user (232) is not accepted by the user (234). For example, during a call, the IMS network (222) sends an alarm signal to the mobile device (230). The mobile device (230) may send a response in response to the alarm signal. In one embodiment, the mobile device (230) may send a rejection signal to the IMS network (222). In another embodiment, the mobile device (230) may not send any signal. If the IMS network (222) does not receive a response within a predetermined time frame, the IMS network (222) determines that the call is not being accepted. In response to the determination that the call is not being accepted, the IMS network (222) routes the call to the voicemail application server (202). The voicemail application server (202) may record a voicemail message from the user (232). The voicemail application server (202) can save a voicemail message as an audio file by calling a storage API for an object storage service (224) when the user (234) is subscribed to the voicemail transcription service.

[0027] After recording a voicemail message and uploading it to the object storage service (224), the voicemail application server (202) can access the voicemail message from the object storage service (224) and call a transcription API to process it for transcription using the transcription service (226). An API is a set of protocols that enable different software applications to communicate with each other. By doing so, APIs facilitate interoperability between different systems. For example, the voicemail application server (202) can record a voicemail message from a user (232) in English as "How are you". The voicemail application server (202) can upload the audio file to the object storage service (224) using the PutObject REST API. For example, the API can authenticate the credentials required to access the object storage service (224) to upload the voicemail audio file. The transcription service (226) is a cloud-based automatic speech recognition service. The transcription service (226) can process the audio file and convert it into transcribed text. Using ASR technology, the transcription service (226) can identify spoken words, correct punctuation, and other linguistic elements from the voicemail message. In response to the voicemail message being uploaded to the storage (224), the voicemail application server (202) can initiate transcription of the voicemail message by the transcription service (226) by calling an API that accesses the location of the audio file on the object storage service (224) using the MediaFileUri (Uniform Resource Identifier). For example, after determining that a successful upload has occurred, the voicemail application server (202) can create a transcription job by calling the StartTranscriptionJob REST API.For example, the API can authenticate credentials required to access the enterprise service (226) to start a transcription operation.

[0028] In some embodiments, by sending a command to the transcription service (226) to initiate the transcription of a voicemail message, the voicemail application server (202) may receive an acknowledgment signal from the transcription service (226) via an API. The acknowledgment signal indicates the successful reception and processing of the entire voicemail message. For example, the voicemail application server (202) may receive an API response containing details regarding the transcription task, such as the task name and its status, and associated timestamps. For example, the API response may include a status as in progress and a timestamp at which the transcription server (226) started the task. By doing so, the voicemail application server (202) can ensure that the transcription service (226) is functioning properly. Additionally, the voicemail application server (202) checks that the entire voicemail is being transcribed.

[0029] In some embodiments, the voicemail application server (202) may invoke another API to check the progress of the transcription of a voicemail message by the transcription service (226). The voicemail application server (202) may receive, via the API, an indication of an incomplete state regarding the transcription of a voicemail message by the transcription service (226). In response to receiving an indication of an incomplete state, the API is called to check the progress of the transcription of the voicemail message after a set time interval and until the transcription of the voicemail message is completed. For example, to know the status of a transcription job, the voicemail application server (202) calls the GetTranscriptionJob REST API. If the status of the transcription job is in progress, the voicemail application server (202) calls this API again after a configured time interval until it obtains a status as completed. By doing so, the voicemail application server (202) can use APIs to continuously check the transcription of voicemail messages.

[0030] The voicemail application server (202) can receive a completed transcription of a voicemail message with a default language value from the transcription service (226). For example, the default language value refers to the language detected from the audio file. For example, if the user (232) spoke in English, the transcription service (226) transcribes the voicemail in English. For example, the voicemail application server (202) can retrieve a voicemail transcription from the object storage service (224) using the GetObject REST API. The GetObject REST API can access the location of the transcription file on the object storage service (224) using the TranscriptFileUri (Uniform Resource Identifier) ​​specified in the GetTranscriptionJob API "Completed" status response. This API allows for efficient data retrieval for large files exchanged with various applications.

[0031] The voicemail application server (202) can retrieve a mark of the target language value for a completed transcription from the voicemail profile database (206). The voicemail profile database (206) can store multiple user profiles for subscribers of the voicemail translation service. The target language value corresponds to the user-selected language value within the user profile of the user (234). For example, the voicemail application server (202) can query the voicemail profile database (206) to determine that the recipient (e.g., user (234)) is a user of the voicemail translation service with Spanish as the selected translation language.

[0032] The voicemail application server (202) can compare the default language value of the completed transcript with the target language value. For example, the voicemail application server (202) can compare the English language value of the transcript with Spanish, which is the translation language selected by the user (234). In response to determining that the default language value of the completed transcript (e.g., English) does not match the target language value (e.g., Spanish), the voicemail application server (202) can trigger an API to generate a translation of the completed transcript according to the target language value. For example, the voicemail application server (202) can trigger an API request to process the English transcript and receive Spanish text as a return. The API can send the completed transcript to a translation service (228). The translation service (228) is a cloud-based machine translation service. The translation service (228) can use machine learning to generate the translated transcript of the voicemail. As used herein, the term “model” may refer to a configuration trained using training data to perform predictions or provide probabilities for new data items, whether or not new data items are included in the training data. For example, training data for supervised learning may include items having various parameters and assigned classifications. New data items may have parameters that the model can use to assign a certain classification to the new data items. As another example, the model may be a probability distribution derived from an analysis of the training data, such as the probability of an n-gram occurring in a given language based on the analysis of a large corpus of that language.Examples of models include neural networks, support vector machines, decision trees, Parzen windows, Bayes, clustering, reinforcement learning, probability distributions, decision trees, decision tree forests, etc. Models can be configured for various situations, data types, sources, and output formats.

[0033] In some embodiments, the machine learning model may be a neural network having multiple input nodes that receive data inputs, such as large amounts of multilingual text data. The input nodes may correspond to functions that receive inputs and generate results. These results may be provided to one or more levels of intermediate nodes, each generating additional results based on a combination of lower-level node results. Weighting factors may be applied to the output of each node before the result is passed to the next layer node. In the final layer ("output layer"), one or more nodes may generate values ​​that classify the input, which can be used to generate translations once the model is trained. In some embodiments, such neural networks, known as deep neural networks, may have multiple layers of intermediate nodes with different configurations, may be a combination of models receiving inputs from different parts of the deep neural network and / or different parts of the input, or may be convolutions that partially use outputs from previous iterations that apply the model as additional inputs to generate results for the current input.

[0034] The translation service (228) may utilize a neural machine translation (NMT) model to generate a translated transcript of a voicemail. For example, NMT uses artificial neural networks to translate one language into another. The translation service (228) may include an encoder-decoder architecture. For example, the encoder converts the transcript into a vector. The decoder takes the vector as input and outputs a translated transcript of the voicemail. The decoder generates one word at a time and considers the context of the previous words and sentences that were generated.

[0035] The voicemail application server (202) can store the translation of the completed transcription in the voicemail message storage (208). The voicemail message storage (208) is accessible by the user (234). The language value of the completed transcription translation is matched with the target language value. For example, the voicemail application server (202) receives "How are you" from the user (232) from the translation service (228) in Spanish as " A voicemail message transcription translated into " can be received. After that, the voicemail application server (202) can store the translated transcription in the voicemail message storage (208).

[0036] After saving the translated transcript, the voicemail application server (202) can generate a notification to display a new message to the user (234). After that, the voicemail application server (202) can determine how to present the voicemail message to the user (234). For example, the voicemail application server (202) can retrieve an indication of user preference for the delivery method for notifications of the translated voicemail messages from the voicemail profile database (206).

[0037] In one embodiment, the voicemail application server (202) may determine that the user preference includes a Short Message Service (SMS). In response, the voicemail application server (202) may send an SMS message to a second user device (e.g., a mobile device (230)) containing a notification that the translation of the completed transcription is stored in the voicemail message storage (208). For example, an SMS voicemail message waiting indicator (MMI) notification is sent by the voicemail application server (202) to a Short Message Service Center (SMSC) (214). The Short Message Service Center (214) is an important component of the SMS infrastructure in telecommunications networks. It is responsible for managing SMS messages between mobile devices. The SMSC (214) delivers the SMS voicemail MWI along with an SMS containing the voicemail transcription to the mobile device (230).

[0038] In another embodiment, the voicemail application server (202) may determine that user preferences include an e-mail box. In response to this determination, the voicemail application server (202) may send a push notification to a secure push proxy server (212). The secure push proxy server (212) manages push services in telecommunication networks. The secure push proxy server (212) delivers the push notification to the second user's e-mail box. The push notification indicates that a translation of a completed transcription is stored in the voicemail message store (208). For example, in response to the determination that the user (234) is registered for push notifications, the voicemail message store (208) sends a push notification to the secure push proxy server (216). The secure push proxy (216) delivers the push notification to a mobile device (230) via cloud messaging (220).

[0039] In some embodiments, the voicemail application server (202) may determine that user preferences display mail. In response to this determination, the voicemail application server (202) may send a notification to the message transfer agent (216) that a translation of the completed transcription is stored in the voicemail message store (208). The message transfer agent (216) delivers the notification to a second user device (e.g., a mobile device (230)) via the message delivery agent (218). The message transfer agent (216) and the message delivery agent (218) are two core components in email communication systems. The message transfer agent (216) manages the routing of email messages between servers. The message delivery agent (218) places email messages from the MTA (216) into a mailbox. For example, a user (234) may set their email address in a voicemail profile as their preferred method of receiving messages. The voicemail application server (202) can send an email notification to a message transmission agent (216). The MTA (216) delivers an email containing the translated transcription text and the attached voicemail to an email inbox on a mobile device (230) via a message delivery agent (218).

[0040] According to user preference, when a notification is received, the mobile device (230) performs authentication and sends an API request to retrieve a voicemail message containing a transcription from the voicemail message storage (208) through the web service gateway (210). After receiving the voicemail message, the mobile device (230) displays the voicemail message for the user (234).

[0041] In some embodiments, the voicemail application server (202) may remove uploaded messages and completed transcripts from the transcription service (226). The voicemail application server (202) uses an API to delete both uploaded audio files and voicemail transcripts from the object storage service (224) to protect user privacy.

[0042] FIG. 3 is a flowchart illustrating a method (300) performed by a voicemail translation service of a telecommunications network configured to translate voicemail messages. The method (300) may be performed by a system including, for example, a handheld mobile device (e.g., a smartphone) and / or a server coupled to the handheld mobile device via a communication network (e.g., a telecommunications network). The handheld mobile device and / or server include at least one hardware processor and at least one non-transient memory that stores instructions that cause the system to perform the method (300) when executed by the at least one hardware processor.

[0043] In 302, the system may receive voicemail messages communicated from a first user device to a second user device. In one example, the system may receive voicemail messages communicated via a telecommunications network from a first user device, such as a mobile device (204) associated with a first user (e.g., user (232)), to a second user device, such as a mobile device (230) associated with a second user (e.g., user (234)). For example, the mobile device (204) may initiate a phone call to the mobile device (230). The phone call is routed through an IMS network, such as an IMS network (222) associated with the telecommunications network.

[0044] In 304, the system can call an API to upload a voicemail message to an object storage service. In one example, the system can call an API to upload a voicemail message to a transcription service (226). An API is a set of protocols that enable different software applications to communicate with each other. By doing so, APIs facilitate interoperability between different systems. For example, a voicemail application server (202) can transcribe a voicemail message from a user (232) in English as "How are you". The voicemail application server (202) can upload the audio file to the object storage service (224) using the PutObject REST API. The API can transmit the audio file of the voicemail message from the object storage service (224) to the transcription service (226). The transcription service (226) is a cloud-based automatic speech recognition service. The transcription service (226) can take the audio file and convert it into transcribed text. Using ASR technology, the transcription service (226) can identify spoken words, correct punctuation, and other linguistic elements from a voicemail message.

[0045] In 306, the system may call an API to initiate the transcription of a voicemail message. The system may send a command to initiate the transcription of a voicemail message. In one example, the system may send a command to the transcription service (226) to initiate the transcription of a voicemail message. For example, after determining that a successful upload has occurred, the voicemail application server (202) may create a transcription job by calling the StartTranscriptionJob REST API. For example, the API may authenticate the credentials required to access the transcription service (226) to start the transcription job.

[0046] In 308, the system can receive a completed transcription of a voicemail message with a default language value from a transcription service. In one example, the system can receive a completed transcription of a voicemail message with a default language value from a transcription service (226). For example, the default language value refers to the language detected from the audio file. For example, if the user (232) spoke in English, the transcription service (226) transcribes the voicemail in English. For example, a voicemail application server (202) can retrieve a voicemail transcription from an object storage service (224) using the GetObject REST API. This API allows for efficient data retrieval for large files exchanged with various applications.

[0047] In 310, the system can retrieve a mark of the target language value for the completed transcription. In one example, the system can retrieve a mark of the target language value for the completed transcription from the voicemail profile database (206). The voicemail profile database (206) may store multiple user profiles for subscribers of the voicemail translation service. The target language value corresponds to the user-selected language value within the user profile of the user (234). For example, the voicemail application server (202) may query the voicemail profile database (206) to determine that the recipient (e.g., user (234)) is a user of the voicemail translation service with Spanish as the selected translation language.

[0048] In 312, the system can compare the default language value of the completed transcription with the target language value. In one example, the system can compare the default language value of the completed transcription with the target language value. For example, the voicemail application server (202) can compare the language value of the English transcription with Spanish, which is the translation language selected by the user (234).

[0049] In 314, the system may trigger another API to generate a translation of the completed transcript. In one example, the system may trigger an API to generate a translation of the completed transcript based on a target language value. For example, a voicemail application server (202) may trigger an API request to process an English transcript and receive Spanish text as a return thereof. The API may send the completed transcript to a translation service (228).

[0050] In 316, the system may store the translation of the completed transcript in a voicemail storage system. In one example, the system may store the translation of the completed transcript in a voicemail message storage (208). The voicemail message storage (208) is accessible by the user (234). The language value of the completed transcript translation is matched with the target language value. For example, the voicemail application server (202) translates "How are you" from the user (232) into Spanish " A voicemail message transcription translated into " can be received. After that, the voicemail application server (202) can store the translated transcription in the voicemail message storage (208).

[0051] computer system

[0052] FIG. 4 is a block diagram illustrating an example of a computer system (400) in which at least some of the operations described herein may be implemented. As illustrated, the computer system (400) may include one or more processors (402), main memory (406), non-volatile memory (410), network interface device (412), video display device (418), input / output device (420), control device (422) (e.g., keyboard and pointing device), a driving unit (424) including a machine-readable (storage) medium (426), and a signal generating device (430) communicably connected to a bus (416). The bus (416) represents one or more physical buses and / or point-to-point connections connected by appropriate bridges, adapters, or controllers. Various common components (e.g., cache memory) are omitted from FIG. 4 for brevity. Instead, the computer system (400) is intended to illustrate a hardware device in which the components illustrated or described in the examples of the drawings and any other components described herein may be implemented.

[0053] The computer system (400) may take any suitable physical form. For example, the computing system (400) may share an architecture similar to that of a server computer, a personal computer (PC), a tablet computer, a mobile phone, a game console, a music player, a wearable electronic device, a network-connected ("smart") device (e.g., a television or home assistant device), AR / VR systems (e.g., a head-mounted display), or any electronic device capable of executing a set of commands that specify the action(s) to be taken by the computing system (400). In some embodiments, the computer system (400) may be an embedded computer system, a system-on-chip (SOC), a single-board computing (SBC) system, or a distributed system such as a mesh of computer systems, or it may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems (400) may perform operations in real-time, near-real-time, or in batch mode.

[0054] A network interface device (412) enables the computing system (400) to mediate data in a network (414) with an entity outside the computing system (400) through any communication protocol supported by the computing system (400) and the external entity. Examples of network interface devices (412) include network adapter cards, wireless network interface cards, routers, access points, wireless routers, switches, multilayer switches, protocol converters, gateways, bridges, bridge routers, hubs, digital media receivers and / or repeaters, as well as all wireless elements mentioned herein.

[0055] Memory (e.g., main memory (406), non-volatile memory (410), machine-readable medium (426)) may be local, remote, or distributed. Although illustrated as a single medium, the machine-readable medium (426) may include multiple media (e.g., centralized / distributed databases and / or associated caches and servers) that store one or more sets of instructions (428). The machine-readable medium (426) may include any medium capable of storing, encoding, or returning sets of instructions for execution by the computing system (400). The machine-readable medium (426) may include non-transient or non-transient devices. In this context, a non-transient storage medium may include a tangible device, meaning that the device has a specific physical form, provided that the device may change its physical state. Thus, for example, non-transient means a device that remains in a tangible form despite such changes in state.

[0056] Although embodiments have been described in the context of fully functional computing devices, various examples may be distributed as program products of various forms. Examples of machine-readable storage media, machine-readable media, or computer-readable media include volatile and non-volatile memory (410), removable flash memory, hard disk drives, optical disks, and other writable media, and transmitting media such as digital and analog communication links.

[0057] Generally, routines executed to implement the examples in this specification may be implemented as part of an operating system or a specific application, component, program, object, module, or sequence of instructions (collectively referred to as “computer programs”). Computer programs typically include one or more instructions (e.g., instructions (404, 408, 428)) set at various times in various memory and storage devices within computing device(s). When read and executed by a processor (402), the instruction(s) cause the computing system (400) to perform operations for executing elements related to various aspects of this disclosure.

[0058] note

[0059] The terms “Example,” “Example,” and “Embodiment” are used interchangeably. For example, references to “one example” or “one example” in this disclosure may refer to the same embodiment, but are not necessarily so; such references mean at least one of the embodiments. The appearance of the phrase “in one example” does not necessarily refer to the same example, nor are separate or alternative examples mutually exclusive from other examples. Features, structures, or characteristics described in connection with one example may be included in other examples of this disclosure. Additionally, various features that may appear in some examples and not in others are described. Similarly, various requirements that may be requirements in some examples but not in others are described.

[0060] The terms used in this specification should be interpreted in the broadest and most reasonable manner, even when used with specific examples of the invention. The terms used in this disclosure generally have their ordinary meanings in the art, within the context of this disclosure, and in the specific context in which each term is used. The use of alternative language or synonyms does not exclude the use of other synonyms. No special significance should be attached to whether a term is described or discussed in detail in this specification. The use of emphasis does not affect the scope and meaning of the terms. It will also be understood that the same thing may be expressed in more than one way.

[0061] Unless clearly required otherwise by the context, throughout the specification and claims, words such as “comprising,” “comprising,” etc., shall be interpreted in an inclusive sense, that is, “comprising but not limited thereto,” rather than in an exclusive or complete sense. As used herein, the terms “connected,” “combined,” and any variations thereof mean any direct or indirect connection or combination between two or more elements; the combination or connection between elements may be physical, logical, or a combination thereof. Additionally, “in this specification,” “above,” “below,” and words of a similar meaning may refer to the entire application rather than any specific part of the application. Where the context permits, words using the singular or plural in the above detailed description may each include the plural or singular. With respect to a list of two or more items, the word “or” encompasses all of the following interpretations of the word: any of the items in the list, all of the items in the list, and any combination of the items in the list. The term "module" broadly refers to software components, firmware components, and / or hardware components.

[0062] Although specific examples of the technology have been described above for illustrative purposes, as will be recognized by those skilled in the art, various equivalent modifications are possible within the scope of the invention. For example, while processes or blocks are provided in a given order, alternative embodiments may employ systems having steps or blocks in a different order, and some processes or blocks may be deleted, moved, added, subdivided, combined, and / or modified to provide alternative or sub-combinations. Each of these processes or blocks may be implemented in various different ways. Also, while processes or blocks may be depicted as being performed sequentially, such processes or blocks may be performed or implemented in parallel or at different times. Furthermore, any specific numbers mentioned herein are merely illustrative, and different values ​​or ranges may be used in alternative embodiments.

[0063] The details of the disclosed embodiments may vary substantially in specific embodiments but may still be included in the disclosed teachings. As mentioned above, specific terms used when describing features or aspects of the invention should not be construed as implying that such terms are redefined in this specification to be limited to the specific characteristics, features, or aspects of the invention to which they are associated. Generally, unless such terms are explicitly defined in the above detailed description, the terms used in the following claims should not be interpreted as limiting the invention to the specific examples disclosed in this specification. Accordingly, the actual scope of the invention encompasses not only the disclosed examples but also all equivalent ways of carrying out or implementing the invention under the claims. Some alternative embodiments may include additional or fewer elements than the embodiments described above.

[0064] Any patents, applications, and other references mentioned above, and any materials that may be described in the attached application documents, are incorporated herein by reference in their entirety, except where subject matter disclaimers or disclaimers exist, and where the combined material does not correspond to the disclosures set forth in this specification, in which case the language of the disclosures shall prevail. Aspects of the invention may be modified to use the systems, functions, and concepts of the various references described above to provide further embodiments of the invention.

[0065] To reduce the number of claims, specific embodiments are presented below in specific claim forms, but the applicant is considering various aspects of the invention in other forms. For example, aspects of the claims may be described in other forms, such as in functional forms or embodied in computer-readable media. Claims intended to be interpreted as functional claims will use the word "means for." However, the use of the term "for" in any other context is not intended to evoke a similar interpretation. The applicant reserves the right to pursue such additional claim forms in this application or in subsequent applications.

Claims

Claim 1 A method performed by a voicemail translation service of a telecommunications network configured to translate voicemail messages, comprising: receiving a voicemail message transmitted through the telecommunications network from a first user device associated with a first user to a second user device associated with a second user; calling a first Application Programming Interface (API) to upload the voicemail message to a transcription service; transmitting a command to initiate transcription of the voicemail message by the transcription service in response to the uploading of the voicemail message to the transcription service; receiving a completed transcription of the voicemail message with a default language value from the transcription service; retrieving a mark of a target language value for the completed transcription from a database—wherein the database stores a plurality of user profiles for subscribers of the voicemail translation service, and the target language value corresponds to a user-selected language value within the user profile of the second user—; and comparing the default language value of the completed transcription with the target language value. A method comprising: a step of triggering a second API to generate a translation of the completed transcription according to the target language value in response to determining that the default language value of the completed transcription does not match the target language value; and a step of storing the translation of the completed transcription in a voicemail storage system accessible by the second user, wherein the language value of the translation of the completed transcription matches the target language value. Claim 2 A method according to claim 1, further comprising the step of receiving an acknowledgment signal from the transcription service via the first API before transmitting a command to initiate transcription of the voicemail message by the transcription service, wherein the acknowledgment signal indicates successful reception and processing of the entire voicemail message. Claim 3 A method according to claim 1, further comprising: invoke a third API to check the progress of the transcription of the voicemail message by the transcription service; receive, through the third API, an indication of an incomplete state regarding the transcription of the voicemail message by the transcription service; and in response to receiving the indication of an incomplete state, call the third API to check the progress of the transcription of the voicemail message after a set time interval and until the transcription of the voicemail message is completed. Claim 4 A method according to claim 1, further comprising the step of extracting an indication of user preference from the database, wherein the user preference corresponds to a user-selected delivery method for notifications of translated voicemail messages. Claim 5 A method according to claim 4, further comprising: determining that the user preference includes a short message service (SMS); and transmitting an SMS message to the second user device that includes a notification that the translation of the completed transcription is stored in the voicemail storage system. Claim 6 A method according to claim 4, further comprising: determining that the user preference includes an e-mail box; and transmitting a push notification to a secure push proxy server, wherein the secure push proxy server delivers the push notification to the second user's e-mail box, and the push notification indicates that the translation of the completed transcription is stored in the voicemail storage system. Claim 7 A method according to claim 4, further comprising: a step of determining that the user preference includes email; and a step of transmitting a notification to a message transmission agent regarding that the translation of the completed transcription is stored in the voicemail storage system, wherein the message transmission agent transmits the notification to the second user device through a message delivery agent. Claim 8 A method according to claim 1, further comprising: a step of generating a notification indicating that the translation of the completed transcription is stored in the voicemail storage system; and a step of communicating the notification to the second user device through the remote communication network. Claim 9 A method according to claim 1, further comprising the step of removing the uploaded message and the completed transcription from the transcription service. Claim 10 A system comprising at least one hardware processor; and at least one non-transient memory for storing instructions, wherein when the instructions are executed by the at least one hardware processor, the system, To receive voicemail messages for the user of the user device through a remote communication network; The transcription service transcribes the above voicemail message with the default language value; To extract the display of the target language value for the user of the above user device; In response to determining that the default language value of the above transcription does not match the target language value, trigger an application programming interface (API) configured to generate a translation of the above transcription according to the target language value; The translation of the above transcript is stored in a voicemail storage system accessible to the user, and A system in which the language value of the translation of the above transcription matches the target language value. Claim 11 A system according to claim 10, further comprising: invoking an additional API to inspect the progress of the transcription of the voicemail message by the transcription service; receiving, through the additional API, an indication of an incomplete transcription of the voicemail message by the transcription service; and in response to receiving the indication of an incomplete transcription, calling the additional API to inspect the progress of the transcription of the voicemail message after a set time interval and until the transcription of the voicemail message is completed. Claim 12 In paragraph 10, additionally, a system that extracts an indication of user preference from a database of user profiles, said user preference corresponding to a user-selected delivery method for notifications of translated voicemail messages. Claim 13 A system according to claim 12, further comprising determining that the user preference includes a short message service (SMS); and transmitting an SMS message to the user device that includes a notification regarding that the translation of the transcription is stored in the voicemail storage system. Claim 14 In paragraph 12, additionally, the system determines that the user preference includes an e-mail box; transmits a push notification to a secure push proxy server, the secure push proxy server delivers the push notification to the user's e-mail box, and the push notification indicates that the translation of the transcription is stored in the voicemail storage system. Claim 15 In paragraph 12, additionally, a system that determines that the user preference includes email; and transmits a notification to a message transmission agent indicating that the translation of the transcription is stored in the voicemail storage system. Claim 16 In paragraph 10, additionally, a system that generates a notification indicating that the translation of the above transcription is stored in the voicemail storage system; and communicates the notification to the user device through the remote communication network. Claim 17 A non-transient computer-readable storage medium in which instructions are recorded, wherein the instructions, when executed by at least one data processor of a system, cause the system to receive a voicemail message indication for a user of a user device; cause a transcribing service to transcribe the voicemail message in a default language; cause a indication of a target language for the user to be retrieved; cause the default language of the transcription to be compared with the target language; determine that the default language does not match the target language; cause an application programming interface (API) configured to generate a translation of the transcription according to the target language; and cause the translation of the transcription to be stored in a voicemail storage system accessible by the user. Claim 18 A non-transient computer-readable storage medium according to claim 17, wherein the system further enables the system to invoke an additional API to inspect the progress of the transcription of the voicemail message by the transcription service; to receive, through the additional API, an indication of an incomplete state regarding the transcription of the voicemail message by the transcription service; and, in response to receiving the indication of an incomplete state, to call the additional API to provide feedback regarding the progress of the transcription of the voicemail message. Claim 19 In paragraph 17, a non-transient computer-readable storage medium that further enables the system to retrieve a user profile of the user, an indication of user preferences, and the user preferences indicate a user selection notification method for translated voicemail messages. Claim 20 In claim 17, a non-transient computer-readable storage medium that causes the system to additionally communicate to the user device a notification indicating that the translation of the transcription is stored in the voicemail storage system.