Synchronous translation system, method, electronic device, medium and computer program product

By using media stream processing devices on the operator's network side to simultaneously translate and display the voice stream, the problem of users needing to change their terminal devices is solved, achieving simultaneous translation without changing devices, thus improving communication efficiency and user experience.

CN121125694APending Publication Date: 2025-12-12CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510252256.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

In existing technologies, users need to switch to a terminal device that supports simultaneous translation when they need to perform simultaneous translation, resulting in a poor user experience and low communication efficiency.

Method used

The voice stream of the first terminal is translated synchronously by the media stream processing device on the operator's network side, and the translation result is provided. The translation result is then transmitted and displayed through the existing network element components on the operator's network side, including video synthesis and audio mixing, to meet the service needs of different terminals.

Benefits of technology

Synchronous translation can be achieved without users changing their terminal devices, improving communication efficiency and user experience between users of different languages, solving language barriers, and promoting cross-cultural communication and understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125694A_ABST
    Figure CN121125694A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a synchronous translation system and method, electronic equipment, a medium and a computer program product, applied to an operator network side, the synchronous translation system comprises a media stream processing device, the media stream processing device is used for synchronously translating a first voice stream sent by a first terminal to obtain a translation result; sending the translation result to a second terminal; the first voice stream is an original voice stream sent by the first terminal to the second terminal.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of core networks, and particularly relates to a synchronous translation system and method, an electronic device, a medium and a computer program product. BACKGROUND

[0002] At present, synchronous translation based on user conversation is mainly realized by terminal devices such as mobile phones or earphones. For example, a mobile phone receiver is used as an audio collection device, the audio is transmitted to an end-side processing module for translation processing, and then the translation result is played out by the mobile phone receiver. When a user needs to perform synchronous translation on a conversation, a terminal device supporting synchronous translation needs to be replaced. SUMMARY

[0003] The application provides a synchronous translation system, method, electronic device, medium and computer program product.

[0004] The application provides a synchronous translation system applied to an operator network side, and the system comprises a media stream processing device, wherein the media stream processing device is configured to perform synchronous translation on a first voice stream sent by a first terminal to obtain a translation result, and send the translation result to a second terminal.

[0005] The media stream processing device is configured to perform synchronous translation on the first voice stream to obtain a translation result, and send the translation result to a second terminal.

[0006] In some embodiments, the media stream processing device is specifically configured to perform synchronous translation on the first voice stream to obtain a translation subtitle, perform video synthesis on the translation subtitle to obtain the translation result, and send the translation result to the second terminal.

[0007] It can be seen that, by the method provided in this embodiment, the first voice stream can be translated synchronously, and the translation result can be displayed in the form of a video.

[0008] In some embodiments, the media stream processing device is specifically configured to perform synchronous translation on the first voice stream to obtain a second voice stream, perform audio mixing on the first voice stream and the second voice stream to obtain a mixed audio result, and send the mixed audio result as the translation result to the second terminal.

[0009] It can be seen that, by the method provided in this embodiment, the first voice stream and the second voice stream obtained by translating the first voice stream can be mixed, and the mixed audio result can be sent to the second terminal, so that the form diversity of the translation result is realized.

[0010] In some embodiments, the system further comprises: a service network element; the service network element is configured to: determine whether the second terminal meets a synchronous translation condition after receiving the first off-hook response; determine a type of translation result in a case that the second terminal meets the synchronous translation condition, the type of translation result comprising a voice result or a subtitle result; and instruct the media stream processing device to perform synchronous translation on the first voice stream based on the type of translation result.

[0011] It can be seen that, by determining whether the second terminal meets the synchronous translation condition after obtaining the first off-hook response, the service of the second terminal can be checked, so that the synchronous translation result meets the service requirement of the second terminal. By determining the type of translation result, the translation result meeting the requirement of the second terminal can be obtained.

[0012] In some embodiments, the service network element is further configured to: copy the first voice stream to the media stream processing device after receiving the first voice stream; and send a synchronous translation request to the media stream processing device; and the media stream processing device is specifically configured to: perform synchronous translation on the first voice stream based on the type of synchronous translation result after receiving the synchronous translation request.

[0013] It can be seen that, by copying the first voice stream to the media stream processing device, the media stream processing device can process the copied first voice stream to obtain the synchronous translation result.

[0014] In some embodiments, the service network element comprises a first network element and a first server; the media stream processing device comprises a VoNR media surface and a media capability platform; the first network element is configured to: copy the first voice stream to the media capability platform; and send a first address of the media capability platform and a second address of the VoNR media surface to the first server; the first server is configured to: send a synchronous translation request to the media capability platform according to the first address; the synchronous translation request comprises the second address; the media capability platform is configured to: perform synchronous translation on the first voice stream according to the synchronous translation request to obtain a translated subtitle or a second voice stream; and send the translated subtitle or the second voice stream to the VoNR media surface based on the second address; and the VoNR media surface is configured to obtain a translation result according to the translated subtitle or the second voice stream.

[0015] It can be seen that, by combining the first address of the media capability platform and the second address of the VoNR media surface, the accurate transmission of the translation result can be achieved, and the efficiency of synchronous translation can be improved.

[0016] Embodiments of the present application further provide a synchronous translation method applied to an operator network side, the method comprising:

[0017] Acquire the first voice stream sent from the first terminal to the second terminal;

[0018] Simultaneous translation is performed on the first speech stream to obtain the translation result;

[0019] The translation result is sent to the second terminal.

[0020] In some embodiments, the step of synchronously translating the first audio stream to obtain a translation result and sending the translation result to the second terminal includes: synchronously translating the first audio stream to obtain translated subtitles; performing video synthesis on the translated subtitles to obtain a translation result; and sending the translation result to the second terminal.

[0021] As can be seen, the method given in this embodiment can be used to simultaneously translate the first audio stream and obtain the corresponding video translation result.

[0022] In some embodiments, the step of synchronously translating the first speech stream to obtain a translation result and sending the translation result to the second terminal includes: synchronously translating the first speech stream to obtain a second speech stream; mixing the first speech stream and the second speech stream to obtain a mixing result; and sending the mixing result as a translation result to the second terminal.

[0023] As can be seen, the method given in this embodiment can mix the first speech stream and the second speech stream obtained by translating the first speech stream, and send the mixing result to the second terminal, thus enriching the form of synchronous translation results.

[0024] In some embodiments, before acquiring the first voice stream sent by the first terminal to the second terminal, the method further includes: after receiving the first off-hook response sent by the first terminal, determining whether the second terminal meets the synchronous translation conditions; if the second terminal meets the synchronous translation conditions, determining the type of translation result, wherein the type of translation result includes voice result or subtitle result; and performing synchronous translation on the first voice stream based on the type of translation result.

[0025] It can be seen that by determining whether the second terminal meets the conditions for synchronous translation after receiving the first off-hook response, the service ordered by the second terminal can be verified, ensuring that the synchronous translation result meets the business requirements of the second terminal. Determining the type of translation result helps to obtain a translation result that meets the requirements of the second terminal.

[0026] This application provides an electronic device, which includes a processor and a memory for storing computer programs capable of running on the processor; wherein,

[0027] The processor is used to run the computer program to perform any of the above-described synchronous translation methods.

[0028] This application provides a computer storage medium storing a computer program that, when executed by a processor, implements any of the above-described synchronous translation methods.

[0029] This application provides a computer program product, including a computer program that, when executed by a processor, implements any of the above-described synchronous translation methods.

[0030] This application provides a synchronous translation system, method, electronic device, medium, and computer program product. Based on the synchronous translation system provided in this application, synchronous translation of user calls can be achieved through the operator network side, without requiring users to purchase or replace terminal equipment, thereby improving the efficiency of real-time communication between users and the user experience. Attached Figure Description

[0031] Figure 1 A schematic diagram of a synchronous translation system provided in an embodiment of this application;

[0032] Figure 2 A schematic diagram of synchronous translation network element interaction provided in an embodiment of this application;

[0033] Figure 3 A flowchart of a synchronous translation method provided in this application embodiment;

[0034] Figure 4 A flowchart of a network-side synchronous translation interaction is provided as an embodiment of this application;

[0035] Figure 5 This is a schematic diagram of the composition structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0036] Currently, simultaneous translation during user calls is primarily achieved by terminal devices, such as mobile phones or headsets. For example, the phone's earpiece acts as the audio capture device, transmitting the audio to the on-device processing module for translation processing, and then the translated result is played back through the phone's earpiece. When users require simultaneous translation during calls, they need to switch to a terminal device that supports simultaneous translation.

[0037] To address the aforementioned issues, the synchronous translation system provided in this application is applied to the operator's network side. It can utilize existing network elements on the operator's network side to achieve synchronous translation of terminal calls, eliminating the need for users to purchase separate terminal devices with real-time translation capabilities.

[0038] The embodiments of this application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the embodiments provided herein are merely illustrative of the embodiments of this application and are not intended to limit the embodiments of this application. Furthermore, the embodiments provided below are some embodiments for implementing this application, and not all embodiments for implementing this application. Unless otherwise specified, the technical solutions described in the embodiments of this application can be implemented in any combination.

[0039] It should be noted that, in the embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a method or apparatus that includes a list of elements includes not only the elements expressly described, but also other elements not expressly listed, or elements inherent to implementing the method or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of other related elements in the method or apparatus that includes that element (e.g., steps in the method or units in the apparatus; for example, a unit in the apparatus may be a portion of circuitry, a portion of a processor, a portion of a program or software, etc.).

[0040] The synchronous translation system provided in this application includes a series of modules and network elements. However, the system provided in this application is not limited to the modules and network elements explicitly described, but may also include modules or network elements that need to be set up to obtain relevant information or to process information. Similarly, the synchronous translation method provided in this application includes a series of steps, but the synchronous translation method provided in this application is not limited to the steps described.

[0041] This application provides a synchronous translation system applied to the operator's network side, such as... Figure 1 As shown, the synchronous translation system 10 provided in this application embodiment includes a media stream processing device 101. The media stream processing device 101 is used to synchronously translate a first voice stream sent by a first terminal to obtain a translation result; and send the translation result to a second terminal; the first voice stream is the original voice stream sent by the first terminal to the second terminal.

[0042] When the first terminal communicates with the second terminal, taking the first terminal as the called terminal and the second terminal as the calling terminal as an example, when the second terminal initiates a voice call to the first terminal, based on the current communication method, the call request initiated by the second terminal is converted into a Session Initiation Protocol (SIP) INVITE request and sent to the core network via the 5G network. The SIP INVITE request contains detailed information about the call, such as the called number, caller identity, media attributes, etc.

[0043] The SIP INVITE request is transmitted to the network location of the called terminal via the 5G wireless network, the 5G core network, and possibly the Internet Protocol Multimedia Subsystem (IMS). Before the call is initiated, the Voice over Long-Term Evolution Application Server (VoLTE AS) verifies the user identities of both the first and second terminals, ensuring they are registered with the VoLTE service and guaranteeing that only authorized users can initiate and receive VoLTE calls. The VoLTE AS handles the SIP INVITE request. When the VoLTE AS successfully locates the first terminal and sends a ringing signal, the called terminal begins ringing. If the first terminal accepts the call, it returns a 200 OK response to the VoLTE AS.

[0044] Once a call connection is established between the first terminal and the second terminal, the users of the first and second terminals can conduct voice calls using their respective terminal devices. When the second terminal user and the first terminal user use different languages, the first and second terminal users can use the system provided in this application embodiment to subscribe to a synchronous translation service provided by an operator to achieve synchronous translation of the call audio during the call.

[0045] Taking a second terminal user subscribing to a simultaneous translation service as an example, during a voice call between the second terminal user and the first terminal user, the first voice stream sent by the first terminal can be obtained through the media stream processing device 101 on the operator's network side, and the first voice stream can be simultaneously translated to obtain the translation result. Here, based on existing network elements or components, the media stream processing device 101 can specifically be one or more of the following in 5G (5th Generation Mobile Communication Technology): 5G Voice over New Radio (VoNR) media plane, VoNR+ media plane, media capability platform, and VoLTE media plane.

[0046] Taking a media capability platform as an example, this platform provides intelligent media applications within a 5G network. In this embodiment, the first speech stream can be processed based on the media capability platform. Preprocessing, such as noise reduction and enhancement, converts the preprocessed speech content into text. Necessary processing, such as word segmentation and part-of-speech tagging, is then performed on the identified text. The processed text is translated into the target language in real time to obtain the translation result. If the translation result is a subtitle, it is loaded into a second terminal for display. If the translation result is a speech, a speech stream is sent to the second terminal for synchronous translation.

[0047] This application provides a synchronous translation result that can synchronously translate the first voice stream sent by the first terminal using existing network elements or components on the operator's network side. This eliminates the need for the second terminal user to change their terminal equipment. Synchronous translation can be achieved based on existing VoLTE and VoNR terminal equipment, enabling users to perform synchronous translation during calls via the operator's network. This enhances communication between users speaking different languages, resolves language barriers, ensures the accuracy of information transmission, improves communication efficiency and user experience, and promotes cross-cultural exchange and understanding.

[0048] In some embodiments, the media stream processing apparatus described above is specifically used to synchronously translate a first audio stream to obtain translated subtitles; to perform video synthesis on the translated subtitles to obtain a translation result; and to send the translation result to a second terminal.

[0049] In the specific implementation process, the media stream processing device 101 can translate the first audio stream to obtain translated subtitles according to the specific service ordered by the second terminal, and synthesize the translated subtitles with preset videos in the preset video library to obtain a synthesized video containing translated subtitles, and send the synthesized video to the second terminal through a downlink video stream.

[0050] When the media stream processing device 101 includes a VoNR+ media surface and a media capability platform, it can translate the first audio stream through the media capability platform to obtain subtitle results, and send the subtitle results to the VoNR+ media surface. The VoNR+ media surface acquires a preset video through a data channel, and synthesizes the preset video with the subtitle results to obtain a translation result in video form. Based on the data channel established between the first terminal and the second terminal, the translation result in video form is sent to the second terminal. The user of the second terminal can see the translation result of the first audio stream in real time by reading the text in the translation result based on the translation result in video form.

[0051] In some embodiments, the media stream processing apparatus described above is specifically used to perform synchronous translation of the first audio stream to obtain a second audio stream; to perform audio mixing processing on the first audio stream and the second audio stream to obtain a mixing result; and to send the mixing result as a translation result to a second terminal.

[0052] In conjunction with the method described in the above embodiments, when a second terminal user requests voice translation, the media stream processing device 101 can also perform voice synthesis on the translation subtitles corresponding to the first voice stream to obtain a second voice stream, and send the second voice stream to the second terminal through the data channel, so that the second terminal user can obtain the translation result in voice form.

[0053] Furthermore, the media stream processing device 101 can also perform audio mixing on the first audio stream and the second audio stream. For example, it can create two mixing tracks, drag and drop the imported first audio stream and the second audio stream onto the created mixing tracks respectively, and adjust the order and position of the audio files as needed. This allows the audio in the mixing result to play a segment of the first audio stream first, and then play a segment of the second audio stream accordingly, so that the second terminal user can match the audio in the first audio stream with the audio in the synchronously translated second audio stream.

[0054] When the media stream processing device 101 includes a VoNR+ media plane and a media capability platform, the media capability platform can translate the first audio stream to obtain a second audio stream, which is then sent to the VoNR+ media plane. The VoNR+ media plane performs mixing processing based on the first and second audio streams to obtain a mixed result. Alternatively, the media capability platform can translate the first audio stream to obtain translated subtitles, which are then sent to the VoNR+ media plane. The VoNR+ media plane generates a second audio stream based on the translated subtitles and performs mixing processing.

[0055] The method provided in this embodiment can further provide users with different forms of synchronous translation results. Users can select different forms of translation results according to the method provided in this embodiment, thereby improving the user experience of synchronous translation.

[0056] In some embodiments, the system further includes: a service network element; the service network element is configured to, after receiving a first off-hook response sent by the first terminal, determine whether the second terminal meets the synchronous translation conditions; if the second terminal meets the synchronous translation conditions, determine the type of translation result, the type of translation result including voice result or subtitle result; and instruct the media stream processing device to perform synchronous translation on the first voice stream based on the type of translation result.

[0057] Based on the above discussion, after the second terminal initiates a call to the first terminal and the first terminal user accepts the call, the first terminal will return an off-hook response, that is, the first off-hook response to the service network element. Here, the first off-hook response can specifically be an INVITE 200OK response, and the specific service network element receiving the first off-hook response can be VoLTE AS.

[0058] Upon receiving the first off-hook response, the service network element further determines whether the second terminal meets the translation conditions. For example, it determines whether the second terminal user has purchased the synchronous translation service and whether the synchronous translation service is active. Simultaneously, it can also determine the translation result type as an audio result based on the second terminal's call status. For example, if the second terminal initiates a voice call, the translation result type can be determined as an audio result; if the second terminal initiates a video call, the translation result type can be determined as a subtitle result. The media stream processing device 101 is then instructed to send the corresponding translation result to the second terminal based on the determined translation result type and the method described in the above embodiments. When it is determined that the second terminal does not meet the synchronous translation conditions, a prompt can be issued to the second terminal, prompting the second terminal user to purchase the synchronous translation service or to enable the synchronous translation service.

[0059] In practical applications, the second terminal can also serve as the called terminal, and the first terminal as the calling terminal. Based on the synchronous translation method given in the embodiments of this application, synchronous translation is provided to the calling user.

[0060] In actual calls, the service network elements may specifically include VoLTE AS, VoNR+ capability network elements, etc. To control the synchronous translation service process, in this embodiment, the service network elements may further include a first server. Based on the synchronous translation system provided in this embodiment, after the VoLTE AS receives the first off-hook response, it sends a notification to the VoNR+ capability network element, informing it that the current call is a voice call or a video call. Then, the VoNR+ capability network element sends a notification to the first server, informing it that the current call is a voice call or a video call. Upon receiving the notification, the first server learns that a call connection has been established between the first terminal and the second terminal. Based on the services purchased by the second terminal user and the service status, the first server determines the type of translation result and instructs the service network element, such as the VoLTE AS, to anchor the voice stream between the first and second terminals. When the type of translation result is determined to be a subtitle result, the first server can also issue an instruction to the VoLTE AS to create a one-way video stream for the second terminal.

[0061] Once the service network element determines the type of translation result, it can send an acknowledgment character (ACK) to the first terminal in response to the first off-hook response, to confirm that the first terminal user has successfully off-hook and prepare for subsequent translation operations.

[0062] In this embodiment, since it is necessary to update the media streams of the first and second terminals—for example, to send a second audio stream or synthesized video to the second terminal—the service network element needs to perform media renegotiation with both the first and second terminals. Specifically, the service network element initiates a Re-INVITE request to the first terminal. This Re-INVITE request is used to update session parameters. Simultaneously, it can be configured so that the Re-INVITE request does not carry the Session Description Protocol (SDP), thus not changing the media description. The service network element initiates media renegotiation with the second terminal and sends a media update request to the second terminal, requesting an update to the audio or video stream to facilitate the subsequent sending of the synchronously translated second audio stream, mixing result, or synthesized video to the second terminal. During this process, when the service network element specifically includes VoLTE AS and VoNR+ capability network element, VoLTE AS anchors the bidirectional audio streams of the first terminal and the second terminal, and negotiates the unidirectional video of the second terminal. The UPDATE message for negotiating the unidirectional video can carry the Service-Interact-Info header field, indicating the service identifier (Identitydocument, ID) that currently triggers the unidirectional video, and indicating that the current request is a media renegotiation request after off-hook.

[0063] After the media renegotiation request is completed, a specific audio stream can be associated with a specific audio output device. In this embodiment, the second audio stream obtained after synchronous translation can be associated with a second terminal, so that the second audio stream can be output through the second terminal.

[0064] After the media renegotiation request is completed and the second audio stream is associated with the second terminal, a notification can be sent to the first server in the service network element to notify the first server that the media renegotiation has been completed, so that the first server can further issue a synchronous translation request.

[0065] When the translation result is determined to be a subtitle result, the service network element can instruct the media stream processing device 101 to obtain the preset video and related materials for video synthesis, and send the preset video to the second terminal for playback.

[0066] In order to detect the operations of the first and second terminal users on the terminal in real time during the call, the service network element can also send a dual-tone multifrequency (DTMF) detection request to detect whether the user sends relevant commands through the buttons of the first or second terminal.

[0067] In some embodiments, the service network element is further configured to copy the first voice stream to the media stream processing device after receiving the first voice stream; send a synchronous translation request to the media stream processing device; and the media stream processing device is specifically configured to perform synchronous translation processing on the first voice stream based on the type of the synchronous translation result after receiving the synchronous translation request.

[0068] Based on the method given in the above embodiments, after completing the media renegotiation request, during the call between the first terminal user and the second terminal user, the service network element copies the first voice stream received from the first terminal and the third voice stream received from the second terminal, and sends the copied voice stream to the media stream processing device 101. The media stream processing device 101 performs synchronous translation on the received voice stream to obtain a translation result of the same type as the synchronous translation result.

[0069] In some embodiments, the service network element includes a first network element and a first server; the media stream processing device includes a VoNR media plane and a media capability platform; the first network element is used to copy a first audio stream to the media capability platform, and send a first address of the media capability platform and a second address of the VoNR media plane to the first server; the first server is used to send a synchronous translation request to the media capability platform according to the first address; the synchronous translation request includes the second address; the media capability platform is used to perform synchronous translation on the first audio stream according to the synchronous translation request to obtain translated subtitles or a second audio stream; and send the translated subtitles or the second audio stream to the VoNR media plane based on the second address; the VoNR media plane is used to obtain the translation result according to the translated subtitles or the second audio stream.

[0070] In this embodiment, the first network element can specifically be composed of a service network element that handles call services. Specifically, the first network element may include a VoNR capability network element (or a VoNR+ capability network element), a VoLTE AS, etc. Based on the above embodiments, this embodiment further provides a synchronous translation system on the operator's network side.

[0071] After receiving the first audio stream, the first network element copies it to the media capability platform. To improve translation accuracy by incorporating contextual information, in practical applications, the first audio stream and the third audio stream from the second terminal can also be copied to the media capability platform together. Specifically, the first network element can instruct the copying of the first media stream to the VoNR media plane, and the VoNR media plane can send the copied result of the first media stream to the media capability platform, which then translates the received first audio stream. Simultaneously, the first network element sends the first address of the media capability platform and the second address of the VoNR media plane to the first server, enabling the first server to send a synchronous translation request to the media capability platform based on the first address, instructing the media capability platform to perform translation. Here, the first and second addresses can specifically be Uniform Resource Locators (URLs). When subtitle results are needed, the first server, based on the type of the translated result, initiates a subtitle result synthesis request to the first network element, requesting video synthesis of the translated subtitles. After receiving the subtitle result synthesis request, the first network element sends a subtitle result synthesis request to the VoNR media plane, requesting the VoNR media plane to perform video synthesis of the translated subtitles.

[0072] Specifically, the first server sends a synchronous translation request to the media capability platform based on the first address, instructing the media capability platform to perform synchronous translation. The media capability platform, based on the synchronous translation request, performs synchronous translation on the first audio stream, obtaining a second audio stream or translated subtitles. The media capability platform then sends the obtained second audio stream or translated subtitles to the VoNR media plane for further processing using the second address in the synchronous translation request.

[0073] When the VoNR media surface receives a second audio stream, it can perform speech enhancement processing on the second audio stream to obtain a processed second audio stream, and then send the processed second audio stream to the second terminal. The VoNR media surface can also perform audio mixing processing on the first and second audio streams according to instructions sent by the first server, and send the mixing result to the second terminal. When the VoNR media surface receives translated subtitles, it combines the subtitle result synthesis request with the acquired preset video to synthesize the preset video and translated subtitles, obtaining a synthesized video, which is then sent to the second terminal. In specific implementations, the VoNR media surface can be replaced with a VoNR+ media surface.

[0074] Based on the synchronous translation system described in the above embodiments, Figure 2A schematic diagram of synchronous translation network element interaction is shown. Based on existing network elements on the operator's network side, the functionality of the network-side media stream processing device 101 is upgraded by adding audio and video media speech recognition and speech synthesis functions to the media stream processing device 101 to achieve synchronous translation. Specifically, the upgraded media stream processing device 101 can acquire and analyze both the calling and called audio streams; it can obtain subtitles in the target language through video synthesis to produce a synthesized video; it can cut off the original audio in the call and play a second audio stream; and it can mix and play the first and second audio streams.

[0075] Combination Figure 2 As shown, the terminal can interact with the VoNR+ media surface. By sending the terminal's first audio stream to the VoNR+ media surface, the VoNR+ media surface can send the first audio stream to the media capability platform for synchronous translation and receive the second audio stream or translated subtitles returned by the media capability platform. The VoNR+ media surface processes the second audio stream or translated subtitles to obtain the translation result, and then sends the translation result back to the terminal.

[0076] Meanwhile, based on the synchronous translation system provided in the embodiments of this application, such as Figure 2 As shown, the existing VoNR+ capability network element can send a notification to the first server, enabling the first server to issue a synchronous translation request based on the notification; the VoNR+ capability network element can interact with the VoNR+ media plane and send a copied first voice stream to the VoNR+ media plane.

[0077] Based on the synchronous translation system proposed in the foregoing embodiments, this application also provides a synchronous translation method applied to the operator network side, such as... Figure 3 As shown, Figure 3 The synchronous translation methods shown include:

[0078] Step 301: Obtain the first voice stream sent from the first terminal to the second terminal.

[0079] Step 302: Simultaneously translate the first speech stream to obtain the translation result.

[0080] Step 303: Send the translation results to the second terminal.

[0081] Corresponding to the synchronous translation system described in the above embodiments, when the first terminal and the second terminal communicate, taking the first terminal as the called terminal and the second terminal as the calling terminal as an example, after establishing a call connection between the first terminal and the second terminal, the users of the first terminal and the second terminal can conduct voice calls based on their respective terminal devices. When the second terminal user and the first terminal user use different languages, the first terminal user and the second terminal user can, based on the method provided in the embodiments of this application, subscribe to the synchronous translation service provided by the operator to achieve synchronous translation of the call voice during the call.

[0082] Taking a second terminal user who has subscribed to a simultaneous translation service as an example, during a voice call between the second terminal user and the first terminal user, the first voice stream sent by the first terminal can be obtained through the media stream processing device on the operator's network side, and the first voice stream can be translated simultaneously to obtain the translation result.

[0083] Taking the media capability platform on the operator's network side as an example, the first audio stream can be processed based on the media capability platform. Preprocessing such as denoising and speech enhancement is performed on the first audio stream to convert the preprocessed audio content into text. The identified text is then processed with techniques such as word segmentation and part-of-speech tagging. The processed text content is translated into the target language in real time to obtain the translation result. If the translation result is a subtitle, the subtitle is loaded onto a second terminal for display; if the translation result is an audio result, the audio stream is sent to the second terminal to achieve synchronous translation.

[0084] This application provides a synchronous translation result that can synchronously translate the first voice stream sent by the first terminal using existing network elements or components on the operator's network side. This eliminates the need for the second terminal user to change their terminal equipment, enabling synchronous translation when users make calls through the operator's network. This enhances communication between users speaking different languages, improving communication efficiency and user experience.

[0085] In practical applications, steps 301 to 303 can be implemented based on a processor, which can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Synchronous Translator, Micro Synchronous Translator, and Microprocessor.

[0086] In some embodiments, the above-mentioned synchronous translation of the first audio stream to obtain a translation result and sending the translation result to the second terminal includes: synchronous translation of the first audio stream to obtain translated subtitles; video synthesis of the translated subtitles to obtain a translation result; and sending the translation result to the second terminal.

[0087] In the specific implementation process, the operator's network side can translate the first voice stream to obtain translated subtitles according to the specific service subscribed by the second terminal, and synthesize the translated subtitles with a preset video to obtain a composite video containing the translated subtitles, and send the composite video to the second terminal through a downlink video stream.

[0088] In some embodiments, the above-mentioned synchronous translation of the first speech stream to obtain a translation result and sending the translation result to the second terminal includes: synchronous translation of the first speech stream to obtain a second speech stream; mixing the first speech stream and the second speech stream to obtain a mixing result; and sending the mixing result as a translation result to the second terminal.

[0089] In conjunction with the method described in the above embodiments, when a second terminal user requests voice translation, the operator's network side can also perform voice synthesis based on the translation subtitles corresponding to the first voice stream to obtain a second voice stream, and send the second voice stream to the second terminal through the data channel, so that the second terminal user can obtain the translation result in voice form.

[0090] Furthermore, the operator's network side can also mix the first and second voice streams to obtain a mixed result. The audio in the mixed result can first play a segment of the first voice stream, and then play a corresponding segment of the second voice stream, so that the second terminal user can match the voice in the first voice stream with the voice in the synchronously translated second voice stream.

[0091] The method provided in this embodiment can enrich the translation formats of simultaneous translation. Users can select different forms of translation results according to the method provided in this embodiment, thereby improving the user experience of simultaneous translation.

[0092] In some embodiments, before obtaining the first voice stream sent from the first terminal to the second terminal, the method further includes: after receiving the first off-hook response sent by the first terminal, determining whether the second terminal meets the synchronous translation conditions; if the second terminal meets the synchronous translation conditions, determining the type of translation result, the type of translation result including voice result or subtitle result; and performing synchronous translation on the first voice stream based on the type of translation result.

[0093] Based on the above discussion, after the second terminal initiates a call to the first terminal and the first terminal user accepts the call, the first terminal will return an off-hook response, i.e., a first off-hook response, to the service network element. Upon receiving the first off-hook response, the operator network side further determines whether the second terminal meets the conditions for synchronous translation, such as whether the second terminal user has currently purchased a synchronous translation service and whether the synchronous translation service is active. Simultaneously, based on the call status of the second terminal—for example, if the second terminal initiates a voice call, the translation result type can be determined as a voice result; if the second terminal initiates a video call, the translation result type can be determined as a subtitle result—and based on the determined translation result type, the corresponding translation result is sent to the second terminal using the method provided in the above embodiments.

[0094] In practical applications, the second terminal can also act as the called terminal, and the first terminal as the calling terminal. Based on the synchronous translation method given in the embodiments of this application, synchronous translation is provided to the user of the calling terminal.

[0095] The synchronous translation method provided in this application eliminates the need for users to purchase or replace their terminals. Based on the large number of mobile terminals on the market that support VoLTE and VoNR, synchronous translation can be achieved during calls using the operator's network, improving communication efficiency and user experience, and promoting cross-cultural exchange and understanding.

[0096] When a second user using Chinese subscribes to the simultaneous translation service and speaks with a first user using English, the second user can hear the first user's English spoken in Chinese. Depending on the service settings, the first user can also hear the corresponding English audio from the second user, and the second user can see Chinese subtitles while the first user sees English subtitles. The second user can also use DTMF dialing to turn the simultaneous translation function on or off during the call.

[0097] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0098] Based on the synchronous translation method and synchronous translation system proposed in the foregoing embodiments, Figure 4 The flowchart of the synchronous translation interaction on the network side is shown. It can be seen that... Figure 4 Specifically, synchronous translation can be achieved by applying existing VoLTEAS, VoNR+ network elements, VoNR+ media plane, and media capability platforms. The specific synchronous translation interaction process includes:

[0099] Step 401: The second terminal initiates a call to the first terminal.

[0100] Step 402: The first terminal sends the first off-hook response to the VoLTE AS.

[0101] Step 403: The VoLTE AS sends a call event notification to the VoNR+ capability network element, and the VoNR+ capability network element sends a call event notification to the first server.

[0102] The VoNR+ capability network element sends a call event notification to the first server, informing the first server whether the current call is an audio call or a video call, and the first server responds.

[0103] Step 404: The first server makes a judgment.

[0104] The first server determines whether the current user, such as the second terminal user, has signed up for the synchronous translation service and whether the synchronous translation service has not been disabled. Based on the determination result, it decides whether to use a one-way video channel or an audio channel for synchronous translation.

[0105] Step 405: The first server sends call event control information to the VoNR+ capability network element, and the VoNR+ capability network element sends call event control information to the VoLTE AS.

[0106] The first server instructs the VoNR+ capability network element to perform bidirectional audio stream anchoring via call event control information, and the VoNR+ capability network element returns a response. The VoNR+ capability network element then sends call event control information to the VoLTE AS, instructing the VoLTE AS to perform bidirectional audio stream anchoring, and the VoLTE AS returns a response.

[0107] When the first server decides to use the one-way video channel method to achieve synchronous translation, the first server instructs the VoNR+ capability network element to create a one-way video stream, which in turn instructs the VoLTE AS to create a one-way video stream.

[0108] When the first server decides to use audio to achieve synchronous translation, the first server instructs the VoNR+ capability network element to create a voice stream, and the VoNR+ capability network element instructs the VoLTE AS to create a voice stream.

[0109] Step 406: The VoLTE AS sends an off-hook response (ACK) to the first terminal.

[0110] Step 407: VoLTE AS sends a media renegotiation to the first terminal.

[0111] Step 408: The first terminal responds to the VoLTE AS.

[0112] Step 409: VoLTE AS instructs VoNR+ media plane to create a media stream.

[0113] When the first server decides to use a one-way video channel to achieve synchronous translation, the VoLTE AS instructs the VoNR+ media plane to obtain a preset video and create a video for the current call based on the preset video.

[0114] Step 410: VoLTE AS makes a judgment.

[0115] When a VoLTE AS receives an SDP offer, it needs to initiate media renegotiation with the second terminal to negotiate a one-way video service request.

[0116] Step 411: The VoLTE AS sends a media update request to the second terminal.

[0117] The media update request will carry the Service-Interact-Info header field, which indicates the service ID that triggered the one-way video and indicates that the current request is an off-hook renegotiation request.

[0118] Step 412: The second terminal sends a media update response to the VoLTE AS.

[0119] Step 413: VoNR+ capability network elements and VoLTE AS alternately indicate the update of audio and video information and set association information.

[0120] The VoNR+ capability network element instructs the VoLTE AS to perform corresponding operations. The VoLTE AS updates audio and video information, sets association information, and ensures that audio and video information is correctly transmitted and processed during the call based on the association relationship.

[0121] Step 414: VoLTE AS sends a media renegotiation ACK to the first terminal.

[0122] Step 415: VoLTE AS sends an audio off-hook message INVITE 200OK to the second terminal.

[0123] Step 416: The second terminal sends an off-hook response ACK to the VoLTE AS.

[0124] Step 417: The VoLTE AS notifies the VoNR+ capability network element of the call control result, and the VoNR+ capability network element notifies the first server of the call control result.

[0125] The call control result informs the first server that the audio and video settings were successful.

[0126] Step 418: The first server requests background one-way video from the VoNR+ capability network element, and the VoNR+ capability network element requests background one-way video from the VoNR+ media plane.

[0127] Here, the background one-way video can be obtained based on a preset video. The background one-way video is used to display the preset video to the second terminal user before synchronous translation.

[0128] Step 419: The VoNR+ media sends a video playback status notification to the VoNR+ capability network element, and the VoNR+ capability network element sends a playback status notification to the first server.

[0129] The VoNR+ media plane acquires a preset video and sends a playback status notification of the preset video to the VoNR+ capability network element. The VoNR+ capability network element then sends the playback status notification of the preset video to the first server, enabling the first server to determine the information of the preset video provided to the second terminal.

[0130] Step 420: VoNR+ media sends a one-way video stream to the second terminal and transmits voice streams with the first and second terminals.

[0131] Before sending the translation results, VoNR+ Media sends a one-way video stream to the second terminal, that is, sends a preset video, and transmits the voice stream between the first terminal and the second terminal.

[0132] Step 421: The first server sends a DTMF detection request to the VoNR+ capability network element.

[0133] The first server sends a DTMF detection request, which includes information such as the DTMF code number to be detected and the URL1 address for reporting the DTMF detection results. URL1 is the primary address of the media capability platform. For example, if the server detects a DTMF code number sent by the user via a terminal key press, indicating that the synchronous translation function needs to be disabled, it will report the DTMF detection results to the media capability platform, causing the platform to terminate the translation of the first audio stream.

[0134] Step 422: VoNR+ capability network elements and VoLTE AS complete detection.

[0135] The DTMF detection command is issued through VoNR+ capability network elements and VoLTE AS.

[0136] Step 423: The VoNR+ capability network element sends a response to the first server.

[0137] Step 424: The first server determines whether the user has enabled the translation function.

[0138] The first server queries to determine whether the synchronous translation function of the current second terminal user is enabled by default.

[0139] Step 425: The first server instructs the first terminal to play sound.

[0140] If the first server determines that the synchronous translation function of the second terminal user is enabled by default and that a voice prompt switch is set, it considers that a voice prompt needs to be sent to the first terminal. The first server initiates a playback request to the VoNR+ capability network element, which carries the playback file ID. The VoNR+ capability network element controls the VoNR+ media plane to obtain the playback file based on the playback file ID and plays it to the first terminal, informing the first terminal user that the second terminal user is currently using the synchronous translation function.

[0141] Step 426: The first server sends a replication request to the VoNR+ capability network element.

[0142] The first server sends a copy request to the VoNR+ capability network element, requesting that the uplink voice stream of the first terminal, i.e. the first voice stream, be copied to the media capability platform.

[0143] Step 427: The VoNR+ capability network element initiates voice stream replication to the VoLTE AS.

[0144] Step 428: The VoNR+ capability network element sends a response to the first server.

[0145] Step 429: The VoNR+ media face copies the first audio stream.

[0146] VoNR media facets further replicate the first voice stream from the first terminal to the media capability platform.

[0147] Step 430: The first server performs business processing.

[0148] When the first server decides to use a one-way video channel to achieve synchronous translation, the first server determines, based on the previous call event notification, that the network side will synthesize the subtitle results into the downlink video stream of the second terminal. At the same time, it completes the association of the corresponding subtitle template according to the user settings.

[0149] Step 431: The first server sends a video compositing request to the VoNR+ capability network element.

[0150] Step 432: The VoNR+ capability network element requests video synthesis from the VoNR+ media plane.

[0151] Step 433: The VoNR+ capability network element responds to the first server.

[0152] Step 434: The first server performs media control on the media capability platform.

[0153] The first server initiates a synchronous translation request to URL1 and requests the media capability platform to send the subtitle results or the second audio stream to the second address URL2 of the VoNR+ media plane.

[0154] Step 435: The media capability platform responds to the first server.

[0155] Step 436: The first server begins recording.

[0156] Billing begins after the first server receives a response from the media capability platform.

[0157] Step 437: The media capability platform processes the first audio stream.

[0158] Step 438: The media capability platform returns the subtitle results or the second audio stream to the VoNR+ media surface.

[0159] Step 439: VoNR+ media responds to the media capability platform.

[0160] Step 440: VoNR+ media faces process the subtitle results or the second audio stream.

[0161] VoNR+ media streams process either the subtitles or the second audio stream to obtain the translation.

[0162] Step 441: VoNR+ media sends the translation results to the second terminal.

[0163] Based on the method given in the above embodiments, the media capability platform processes the first audio stream, sends the processed second audio stream or subtitle result to the VoNR+ media plane, and the VoNR+ media plane performs audio mixing or video synthesis to obtain the translation result. Finally, the translation result is sent to the second terminal to achieve synchronous translation.

[0164] It should be noted that the description of the above method embodiments is similar to the description of the above synchronous translation system embodiments, and has similar beneficial effects as the same system embodiments. For technical details not disclosed in the method embodiments of this application, please refer to the description of the system embodiments of this application for understanding.

[0165] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a terminal, server, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0166] This application also provides an electronic device. Figure 5 This is a schematic diagram of the composition structure of an electronic device provided in an embodiment of this application, as shown below. Figure 5 As shown, the electronic device 50 may include:

[0167] Memory 501 is used to store executable instructions.

[0168] The processor 502 is used to implement any of the above-described synchronous translation methods when executing executable instructions stored in the memory 501.

[0169] The processor 502 mentioned above can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, synchronous translator, micro synchronous translator, and microprocessor.

[0170] The aforementioned computer-readable storage medium or memory 501 may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; it may also be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0171] This application embodiment further provides a computer storage medium storing computer-executable instructions, which are used to implement any of the synchronous translation methods provided in the above embodiments.

[0172] Correspondingly, this application embodiment further provides a computer program product, the computer program product including computer executable instructions, which are used to implement any of the synchronous translation methods provided in the above embodiments.

[0173] In some embodiments, the synchronous translation system or method provided in this application may have functions or include modules that can be used to execute the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0174] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0175] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict.

[0176] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0177] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0178] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0179] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims. All of these forms are within the protection scope of this application.

Claims

1. A simultaneous translation system, characterized in that, Applied to the operator's network side, the system includes: a media stream processing device, wherein, The media stream processing device is used to synchronously translate a first audio stream sent by a first terminal to obtain a translation result; and send the translation result to a second terminal; the first audio stream is the original audio stream sent by the first terminal to the second terminal.

2. The system according to claim 1, characterized in that, The media stream processing device is specifically used to synchronously translate the first audio stream to obtain translated subtitles; to perform video synthesis on the translated subtitles to obtain the translation result; and to send the translation result to the second terminal.

3. The system according to claim 1, characterized in that, The media stream processing device is specifically used to synchronously translate the first audio stream to obtain a second audio stream; to perform audio mixing processing on the first audio stream and the second audio stream to obtain a mixing result; and to send the mixing result as the translation result to the second terminal.

4. The system according to claim 1, characterized in that, The system also includes: service network elements; The service network element is configured to, upon receiving a first off-hook response from the first terminal, determine whether the second terminal meets the conditions for synchronous translation; if the second terminal meets the conditions for synchronous translation, determine the type of translation result, wherein the type of translation result includes audio result or subtitle result; and instruct the media stream processing device to perform synchronous translation on the first audio stream based on the type of translation result.

5. The system according to claim 4, characterized in that, The service network element is also configured to, upon receiving the first voice stream, copy the first voice stream to the media stream processing device; and send a synchronous translation request to the media stream processing device. The media stream processing device is specifically used to perform synchronous translation processing on the first audio stream based on the type of the synchronous translation result after receiving the synchronous translation request.

6. The system according to claim 4 or 5, characterized in that, The service network element includes a first network element and a first server; the media stream processing device includes a VoNR media plane and a media capability platform; The first network element is used to copy the first voice stream to the media capability platform and send the first address of the media capability platform and the second address of the VoNR media plane to the first server. The first server is configured to send a synchronous translation request to the media capability platform based on the first address; the synchronous translation request includes the second address; The media capability platform is used to perform synchronous translation on the first audio stream according to the synchronous translation request, to obtain translated subtitles or a second audio stream; Based on the second address, the translated subtitles or the second audio stream are sent to the VoNR media surface; The VoNR media surface is used to obtain translation results based on the translated subtitles or the second audio stream.

7. A simultaneous translation method, characterized in that, Applied to the operator's network side, the method includes: Acquire the first voice stream sent from the first terminal to the second terminal; Simultaneous translation is performed on the first speech stream to obtain the translation result; The translation result is sent to the second terminal.

8. The method according to claim 7, characterized in that, The step of synchronously translating the first speech stream to obtain a translation result and sending the translation result to the second terminal includes: The first audio stream is translated synchronously to obtain translated subtitles; the translated subtitles are then combined with video to obtain a translation result; the translation result is then sent to the second terminal.

9. The method according to claim 7, characterized in that, The step of synchronously translating the first speech stream to obtain a translation result and sending the translation result to the second terminal includes: The first speech stream is translated synchronously to obtain a second speech stream; the first speech stream and the second speech stream are mixed to obtain a mixing result; the mixing result is sent to the second terminal as a translation result.

10. The method according to claim 7, characterized in that, Before acquiring the first voice stream sent from the first terminal to the second terminal, the method further includes: After receiving the first off-hook response from the first terminal, it is determined whether the second terminal meets the synchronous translation conditions; if the second terminal meets the synchronous translation conditions, the type of translation result is determined, the type of translation result includes voice result or subtitle result; based on the type of translation result, the first voice stream is synchronously translated.

11. An electronic device, characterized in that, The electronic device includes a processor and a memory for storing computer programs capable of running on the processor; wherein, The processor is used to run the computer program to perform the method according to any one of claims 7 to 10.

12. A computer storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 7 to 10.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 7 to 10.