Speech processing method and apparatus, related devices, storage medium, and computer program product
By anchoring the voice stream in the call signaling to a new call media node for encoding and decoding conversion via a satellite mobile communication terminal, the problem of conflict between voice quality and system capacity in satellite mobile communication is solved, and high-quality low-bit-rate voice encoding and decoding interoperability is achieved.
Patent Information
- Application Number
- CN202411514841.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2044-10-28
AI Technical Summary
In satellite mobile communications, existing technologies reduce communication bandwidth by modifying the terminals of both communicating parties to achieve a lower bit rate, but this still leads to a decrease in voice quality.
By using cell location information in the call signaling, the voice stream is anchored to a new call media node through the satellite mobile communication terminal, and the conversion between low bit rate voice codec and terrestrial general voice codec is performed to realize the voice stream codec conversion.
Without changing the existing terminals of both communicating parties, improve the quality of voice calls and achieve interoperability between low-bitrate voice codecs and the other end.
Smart Images

Figure CN119364295B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data services, and in particular to a voice processing method and device, related equipment, a storage medium and a computer program product. BACKGROUND
[0002] In a satellite mobile communication network, due to limited satellite link resources, if a traditional mobile communication voice codec is used, the problem of insufficient call system capacity will occur; if a narrowband voice codec is used, due to excessive compression of audio, the problem of degraded call voice quality will also occur.
[0003] To solve the above-mentioned conflict between voice quality and system capacity of satellite mobile communication, the related art mainly modifies the satellite mobile communication terminals of both parties (referred to as terminals of both parties) to reduce the communication bandwidth by using a low code rate while not reducing the voice quality as much as possible, thereby improving the system capacity of communication. However, using the related art solution requires the terminals of both parties to support a new low code rate voice codec, and the voice compression-based codec mode will always reduce the call voice quality. SUMMARY
[0004] To solve the technical problems in the related art, the embodiments of the present application provide a voice processing method and device, related equipment, a storage medium and a computer program product.
[0005] To achieve the above-mentioned purposes, the technical solutions of the embodiments of the present application are as follows:
[0006] In a first aspect, the embodiments of the present application provide a voice processing method applied to a first function, wherein the first function is used at least for processing voice service data, and the method comprises:
[0007] receiving a first request initiated by a second function through a third function; the first request is triggered by the second function based on first information, wherein the first information comprises location information of a first terminal; the first request is used for requesting to control a call event between the first terminal and a second terminal; the second function represents a service application server, and the third function is used at least for processing control plane data;
[0008] in response to the first request, controlling to anchor a voice stream to a fourth function based on the location information of the first terminal, so as to convert between a first codec and a second codec for the voice stream through the fourth function; the fourth function is used at least for processing user plane data.
[0009] In a second aspect, the embodiments of the present application further provide another voice processing method applied to a second function, wherein the second function represents a service application server, and the method comprises:
[0009] In a second aspect, the embodiments of the present application further provide another voice processing method applied to a second function, wherein the second function represents a service application server, and the method comprises:
[0010] initiating a first request from a third function to a first function; the first request is triggered based on first information, the first information comprising location information of a first terminal; the first request is used to request control over a call event between the first terminal and a second terminal; the first function is used to process at least voice service data, and the third function is used to process at least control plane data;
[0011] wherein the location information of the first terminal is used by the first function to control anchoring of a voice stream to a fourth function, so as to perform conversion between a first codec and a second codec for the voice stream by the fourth function; and the fourth function is used to process at least user plane data.
[0012] In a third aspect, another voice processing method is provided by embodiments of the present application, and is applied to a fourth function, the fourth function being used to process at least user plane data, the method comprising:
[0013] obtaining a voice stream; the voice stream is anchored by a first function based on location information of a first terminal in response to a first request, the first function being used to process at least voice service data;
[0014] performing conversion between a first codec and a second codec for the voice stream;
[0015] wherein the first request is sent by a second function to the first function through a third function, the first request is triggered by the second function based on first information, the first information comprising the location information of the first terminal; the first request is used to request control over a call event between the first terminal and a second terminal; the second function represents a service application server, and the third function is used to process at least control plane data.
[0016] In a fourth aspect, a voice processing apparatus is provided by embodiments of the present application, and is applied to a first function, the first function being used to process at least voice service data, the apparatus comprising:
[0017] a first receiving unit, configured to receive a first request initiated by a second function through a third function; the first request is triggered by the second function based on first information, the first information comprising location information of a first terminal; the first request is used to request control over a call event between the first terminal and a second terminal; the second function represents a service application server, and the third function is used to process at least control plane data;
[0018] The control unit is configured to control, in response to the first request, anchoring of a voice stream to a fourth function based on location information of the first terminal, so as to perform conversion between a first codec and a second codec for the voice stream by the fourth function; and the fourth function is configured to process at least user plane data.
[0019] In a fifth aspect, the embodiments of the present application further provide another voice processing apparatus applied to a second function, the second function representing a service application server, and the apparatus comprises:
[0020] The first sending unit is configured to initiate a first request to a first function through a third function; the first request is triggered based on first information, the first information comprising location information of a first terminal; the first request is configured to request control over a call event between the first terminal and a second terminal; the first function is configured to process at least voice service data, and the third function is configured to process at least control plane data.
[0021] The location information of the first terminal is used by the first function to control anchoring of a voice stream to a fourth function, so as to perform conversion between a first codec and a second codec for the voice stream by the fourth function; and the fourth function is configured to process at least user plane data.
[0022] In a sixth aspect, the embodiments of the present application further provide another voice processing apparatus applied to a fourth function, the fourth function being configured to process at least user plane data, and the apparatus comprises:
[0023] The first obtaining unit is configured to obtain a voice stream; the voice stream is anchored by a first function in response to a first request based on location information of a first terminal, and the first function is configured to process at least voice service data;
[0024] The conversion unit is configured to perform conversion between a first codec and a second codec for the voice stream.
[0025] The first request is sent by a second function to the first function through a third function, the first request is triggered by the second function based on first information, the first information comprising location information of the first terminal; the first request is configured to request control over a call event between the first terminal and a second terminal; the second function represents a service application server, and the third function is configured to process at least control plane data.
[0026] In a seventh aspect, the embodiments of the present application provide a first function, comprising: a first processor and a first memory configured to store a computer program capable of running on the first processor.
[0027] The first processor is configured to execute the steps of the first function side voice processing method when running the computer program.
[0028] In an eighth aspect, the embodiments of the present application further provide a second function, comprising: a second processor and a second memory configured to store a computer program capable of running on the second processor;
[0029] The second processor is configured to execute the steps of the second function side voice processing method when running the computer program.
[0030] In a ninth aspect, the embodiments of the present application further provide a fourth function, comprising: a third processor and a third memory configured to store a computer program capable of running on the third processor;
[0031] The third processor is configured to execute the steps of the fourth function side voice processing method when running the computer program.
[0032] In a tenth aspect, the embodiments of the present application further provide a storage medium having a computer program stored thereon, wherein the computer program is configured to implement the steps of the first function side voice processing method, or the steps of the second function side voice processing method, or the steps of the fourth function side voice processing method when executed by a processor.
[0033] In an eleventh aspect, the embodiments of the present application further provide a computer program product, comprising a computer program, wherein the computer program is configured to implement the steps of the first function side voice processing method, or the steps of the second function side voice processing method, or the steps of the fourth function side voice processing method when executed by a processor.
[0034] The voice processing method, device, related equipment, storage medium and computer program product provided by the embodiments of the present application are as follows: a first function receives a first request initiated by a second function through a third function; the first request is triggered by the second function based on first information, the first information comprising location information of a first terminal; the first request is used to request control over a call event between the first terminal and a second terminal; the second function represents a service application server, and the third function is used at least for processing control plane data; in response to the first request, the voice stream is anchored to a fourth function based on the location information of the first terminal to perform conversion between a first codec and a second codec for the voice stream through the fourth function; and the fourth function is used at least for processing user plane data. By using the voice processing method of the embodiments of the present application, the voice stream is anchored to a new call media node (i.e., the fourth function, which is used at least for processing user plane data) through the cell location information (i.e., the location information of the first terminal) of the satellite mobile communication terminal (i.e., the first terminal) in the call signaling, so that conversion between the first codec and the second codec for the voice stream is performed through the fourth function, i.e., conversion between the corresponding low-code-rate voice codec and the general ground voice codec is performed, thereby realizing that the first terminal can use the low-code-rate voice codec to interwork with the opposite terminal without changing the terminals of the existing communication parties, and improving the voice quality of the call. BRIEF DESCRIPTION OF DRAWINGS
[0035] Figure 1 Flowchart of the voice processing method of the embodiments of the present application Figure 1 ;
[0036] Figure 2 Flowchart of the voice processing method of the embodiments of the present application Figure 2 ;
[0037] Figure 3 Flowchart of the voice processing method of the embodiments of the present application Figure 3 ;
[0038] Figure 4 Composition structure diagram of the sending end encoding device of the embodiments of the present application
[0039] Figure 5 Composition structure diagram of the receiving end decoding device of the embodiments of the present application
[0040] Figure 6 Flowchart of the voice processing method of the embodiments of the present application Figure 4 ;
[0041] Figure 7 Composition structure diagram of the voice processing device of the embodiments of the present application Figure 1 ;
[0042] Figure 8 A schematic diagram of a component structure of a voice processing device according to an embodiment of the present application Figure 2 ;
[0043] Figure 9 A schematic diagram of a component structure of a voice processing device according to an embodiment of the present application Figure 3 ;
[0044] Figure 10 A schematic diagram of a hardware component structure of a first function according to an embodiment of the present application
[0045] Figure 11 A schematic diagram of a hardware component structure of a second function according to an embodiment of the present application
[0046] Figure 12 A schematic diagram of a hardware component structure of a fourth function according to an embodiment of the present application DETAILED DESCRIPTION
[0047] The present application will be further described by examples in conjunction with the accompanying drawings.
[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application.
[0049] At present, in order to solve the problem of conflict between voice quality and system capacity in satellite mobile communication, the scheme of the related technology mainly modifies the satellite mobile communication terminals of the two communication parties (referred to as the terminals of the two communication parties), specifically, respectively stores and processes the voice feature codes of the terminals of the two communication parties in a distributed manner, or uses Lyra AI audio codec (ultra-low bit rate audio codec based on voice compression) to realize the reduction of communication bandwidth by using low code rate while minimizing the reduction of voice quality, thereby improving the system capacity of communication.
[0050] However, using the above scheme of the related technology, the terminals of the two communication parties need to support new low code rate voice coding at the same time, and the coding mode based on voice compression will always reduce the voice quality of the call.
[0051] Based on this, the embodiment of the present application proposes a voice processing method. In various embodiments of the present application, a satellite mobile communication terminal (i.e., a first terminal) controls to anchor a voice stream to a new call media node (i.e., a fourth function, which is used at least for processing user plane data) through cell location information (i.e., location information of the first terminal) in call signaling, so as to convert between a first codec and a second codec for the voice stream through the fourth function, i.e., to convert between a corresponding low-code rate voice codec and a general ground voice codec, thereby realizing that the first terminal can use a low-code rate voice codec to interwork with a peer terminal without changing the terminals of the existing communication parties, and thereby improving the voice quality of a call.
[0052] The embodiment of the present application provides a voice processing method, which is applied to a first function, Figure 1 A flowchart of the voice processing method of the embodiment of the present application Figure 1 As shown in the figure, Figure 1 The method comprises the following steps.
[0053] Step 101: receiving a first request initiated by a second function through a third function; the first request is triggered by the second function based on first information, and the first information comprises location information of a first terminal.
[0054] In the embodiment of the present application, the first request is used to request control over a call event between the first terminal and a second terminal; and the second function represents an application server (AS, Application Server). The first function is used at least for processing voice service data, and an exemplary first function can be a voice over long term evolution (VoLTE, Voice over Long Term Evolution) AS. The third function is used at least for processing control plane data, and an exemplary third function can be a new call platform (NCP, New Calling Platform), which can also be referred to as a new call service control platform.
[0055] In an embodiment, the VoLTE AS receives a first request initiated by a service AS through the NCP, that is, the first request is initiated by the service AS, and the service AS first sends the first request to the NCP, and then forwards the first request to the VoLTE AS through the NCP, wherein the first request is used to request control over a call event between the first terminal and a second terminal.
[0056] Herein, the first terminal represents a satellite mobile communication terminal, and the second terminal represents a ground mobile communication terminal, wherein the voice call connection is established between the first terminal and the second terminal, that is, the first terminal triggers a call event through the first function, and the voice call connection between the first terminal and the second terminal is established based on the call event.
[0057] In actual application, the first request is not randomly initiated by the second function, but is conditionally triggered, specifically, the first request is triggered by the second function based on the first information, that is, the second function can only trigger the sending of the first request based on the first information after receiving the first information.
[0058] Based on this, in an embodiment, before the receiving of the first request initiated by the third function through the second function, the method further comprises: sending the first information to the third function, so that the third function sends the first information to the second function; wherein the first information represents a call event notification message.
[0059] Herein, in an embodiment, the sending of the first information to the third function comprises:
[0060] determining whether the second information is received; the second information represents a reply confirmation message sent by the second terminal, and the second information is generated by the second terminal based on the received second request initiated by the first terminal, and the second request represents a codec update request for the voice stream;
[0061] In the case of determining that the second information is received, the first information is sent to the third function.
[0062] Here, the first function triggers sending the first information to the third function, i.e., sending the call event notification message to the third function, only when it is determined that the second information is received. It should be noted that the second information is sent by the second terminal to the fifth function, and is forwarded to the first function by the fifth function. In the embodiments of the present application, the fifth function is at least used for processing call session data, and an exemplary fifth function can be an Interrogating / Serving-Call Session Control Function (I / S-CSCF). The second information represents a reply confirmation message sent by the second terminal, and an exemplary second information can be 200 OK. Specifically, the first terminal initiates a codec update request for a voice stream, i.e., an Update request. After receiving the Update request, the fifth function forwards the Update request to the second terminal. When the second terminal receives the Update request, it generates second information based on the Update request, i.e., generates feedback information for the Update request. In other words, the second terminal generates a reply confirmation message for the Update request after receiving the Update request, and then sends the reply confirmation message to the first function through the fifth function, to further trigger the first function to send a call event notification message to the third function.
[0063] In actual application, the first request is sent by the second function to the first function based on triggering of the first information.
[0064] Based on this, in an embodiment, the second function receiving the first request initiated by the third function includes: receiving the first request sent by the third function in a case where the second function receives the first information; and wherein the first request is sent by the second function to the third function based on triggering of the first information.
[0065] Step 102: in response to the first request, controlling anchoring of the voice stream to a fourth function based on location information of the first terminal, to perform conversion between the first codec and the second codec for the voice stream through the fourth function.
[0066] In the embodiments of the present application, the fourth function is at least used for processing user plane data, and an exemplary fourth function can be a Unified Media Function (UMF).
[0067] Here, since the first request is triggered by the second function based on the first information, the first information comprises the location information of the first terminal, therefore, after receiving the first request, the first function controls the call event between the first terminal and the second terminal based on the first request, specifically, the first function controls the re-anchoring of the voice stream, i.e. anchors the voice stream to a new call media node (i.e. the fourth function), through the cell location information of the satellite mobile communication terminal, i.e. the first terminal, in the call signaling, i.e. the location information of the first terminal, to convert the first codec and the second codec for the voice stream through the fourth function, i.e. to convert the corresponding low-code-rate voice codec and the general ground voice codec, so as to realize the intercommunication of the first terminal with the opposite terminal using the low-code-rate voice codec without changing the existing terminals of the communication parties, thereby improving the voice quality of the call.
[0068] In the embodiments of the present application, the first codec corresponds to the first terminal, the first terminal represents a satellite mobile communication terminal, the second codec corresponds to the second terminal, the second terminal represents a ground mobile communication terminal, wherein the first codec represents a low-code-rate voice codec, and the second codec represents a general ground voice codec.
[0069] It should be noted that, since the first terminal represents a satellite mobile communication terminal, the location information of the first terminal can be the location information of the satellite mobile communication terminal, which can be referred to as satellite location information.
[0070] In actual application, the location information of the first terminal can be carried in the first field.
[0071] Based on this, in an embodiment, the method further comprises: creating a first field in the first information; and adding the location information of the first terminal in the first field.
[0072] Here, the first function (i.e. VoLTE AS) sends the first information (i.e. call event notification message) to the third function (i.e. NCP) in the case of receiving the second information (i.e. reply confirmation message sent by the second terminal, for example, the second information can be 200OK), and the call event notification message carries the location information of the satellite mobile communication terminal.
[0073] Specifically, the first function creates a first field in the first information, for example, the first field can be an scgi field, specifically, the first field can be an scgi field set in the location type (Location Type), i.e. the location information of the satellite mobile communication terminal is carried in the scgi field in the Location Type.
[0074] In actual application, after the first function controls the voice stream to be anchored to the fourth function based on the location information of the first terminal, the first function controls the fourth function to create a media resource between the fourth function and the second terminal, and selects a corresponding voice codec.
[0075] Based on this, in an embodiment, after the first function controls the voice stream to be anchored to the fourth function based on the location information of the first terminal, the method further comprises:
[0076] initiating a third request to the fourth function; the third request is used to request the fourth function to create a first media resource; the first media resource represents a media resource between the fourth function and the second terminal, and the first media resource is used to determine that a first code stream is transmitted between the fourth function and the second terminal, and the first code stream is associated with the second codec;
[0077] receiving third information sent by the fourth function; the third information represents a confirmation message for creating the first media resource.
[0078] Here, by using the media resource (i.e., the first media resource) created between the fourth function and the second terminal, the voice stream transmitted between the fourth function and the second terminal can be determined to be the first code stream, and the first code stream is associated with the second codec, that is, the voice stream transmitted between the fourth function and the second terminal is voice encoded in the manner of a general terrestrial voice codec.
[0079] It should be noted that after the fourth function successfully creates the first media resource, a confirmation message will be generated for the event of successfully creating the first media resource, i.e., the third information is generated, and then the fourth function sends the third information to the first function to inform the first function that the first media resource is successfully created. For example, the third information can be 200OK.
[0080] In actual application, after the fourth function successfully creates the first media resource, the first media resource can also be updated to complete the conversion between the first codec and the second codec for the voice stream.
[0081] Based on this, in an embodiment, after the fourth function successfully creates the first media resource, the method further comprises:
[0082] initiating a third request to the fourth function; the third request is used to request the fourth function to create a first media resource; the first media resource represents a media resource between the fourth function and the second terminal, and the first media resource is used to determine that a first code stream is transmitted between the fourth function and the second terminal, and the first code stream is associated with the second codec;
[0083] receiving fourth information sent by the fourth function; the fourth information represents a confirmation message for updating the first media resource.
[0084] Here, after the first function controls the fourth function to create the first media resource between the fourth function and the second terminal, the first function can further control the fourth function to update the first media resource to obtain a second media resource, the second media resource representing a media resource between the fourth function and the first terminal, and the second media resource can be used to determine that the voice stream transmitted between the fourth function and the first terminal is a second code stream, and the second code stream is associated with the first codec, that is, the voice stream transmitted between the fourth function and the first terminal is encoded and decoded in the manner of low-code-rate voice codec, so that the first terminal can interwork with the opposite terminal using low-code-rate voice codec without changing the existing terminals of the communication parties, thereby improving the voice quality of the call.
[0085] It should be noted that after the fourth function completes the update of the first media resource, a confirmation message for the event of completing the update of the first media resource is generated, that is, the fourth information is generated, and then the fourth function sends the fourth information to the first function to inform the first function that the update of the first media resource is completed. Exemplarily, the fourth information can be 200OK.
[0086] The embodiments of the present application also provide another voice processing method, which is applied to a second function, Figure 2 The flow of the voice processing method of the embodiments of the present application is shown in Figure 2 ; as Figure 3 shown, the method comprises:
[0087] Step 201: initiating a first request to a first function through a third function; the first request is triggered based on first information, the first information comprising location information of a first terminal, the location information of the first terminal being used by the first function to control anchoring of a voice stream to a fourth function to convert the voice stream between a first codec and a second codec through the fourth function.
[0088] In the embodiments of the present application, the first request is used to request control over a call event between the first terminal and a second terminal; the first function is used to process at least voice service data, exemplarily, the first function can be a VoLTE AS; the second function represents a service AS; the third function is used to process at least control plane data, exemplarily, the third function can be an NCP; and the fourth function is used to process at least user plane data, exemplarily, the fourth function can be a UMF.
[0089] Here, the first terminal represents a satellite mobile communication terminal, and the second terminal represents a ground mobile communication terminal, wherein the voice call connection is established between the first terminal and the second terminal, that is, the first terminal triggers a call event through the first function, and the voice call connection between the first terminal and the second terminal is established based on the call event.
[0090] In actual application, the first request is not randomly initiated by the second function, but is conditionally triggered, specifically, the first request is triggered by the second function based on the first information, that is, the second function can trigger the sending of the first request based on the first information only after receiving the first information. Therefore, the second function can also receive the first information before sending the first request.
[0091] Based on this, in an embodiment, before the first request is initiated by the third function to the first function, the method further comprises: receiving the first information sent by the first function through the third function; wherein the first information represents a call event notification message.
[0092] In actual application, the first request is sent by the second function to the first function based on the trigger of the first information.
[0093] Based on this, in an embodiment, the first request is initiated by the third function to the first function, comprising: determining whether the first information is received; in the case where it is determined that the first information is received, triggering the sending of the first request to the third function based on the first information, so that the third function sends the first request to the first function.
[0094] Here, since the first request is triggered by the second function based on the first information, and the first information includes the location information of the first terminal, after receiving the first request, the first function controls the call event between the first terminal and the second terminal based on the first request, specifically, the first function controls the re-anchoring of the voice stream, that is, the voice stream is anchored to a new call media node (i.e. the fourth function), through the cell location information of the satellite mobile communication terminal (i.e. the first terminal) in the call signaling, that is, the location information of the first terminal, so as to convert the first codec and the second codec for the voice stream through the fourth function, that is, to convert the corresponding low-code-rate voice codec and the ground general voice codec, thereby realizing that the first terminal can use the low-code-rate voice codec to interwork with the opposite terminal without changing the existing communication terminals of both parties, and further improving the voice quality of the call.
[0095] In the embodiments of the present application, the first codec corresponds to the first terminal, the first terminal represents a satellite mobile communication terminal, the second codec corresponds to the second terminal, the second terminal represents a ground mobile communication terminal, and the first codec represents a low-code-rate speech codec, and the second codec represents a ground general speech codec.
[0096] It should be noted that, since the first terminal represents a satellite mobile communication terminal, the position information of the first terminal can be position information of the satellite mobile communication terminal, which is referred to as satellite position information for short.
[0097] The embodiments of the present application also provide another voice processing method, which is applied to a fourth function, Figure 3 The flow of the voice processing method of the embodiments of the present application is shown in Figure 4 FIG. 1. Figure 4 The method comprises the following steps.
[0098] Step 301: Obtain a voice stream.
[0099] In the embodiments of the present application, the voice stream is anchored by a first function based on position information of a first terminal in response to a first request. The first function is used for processing voice service data, for example, the first function can be a VoLTE AS. The fourth function is used for processing user plane data, for example, the fourth function can be a UMF.
[0100] Here, the first request is sent by a second function to the first function through a third function. The first request is triggered by the second function based on first information, and the first information comprises the position information of the first terminal. The first request is used to request control over a call event between the first terminal and a second terminal.
[0101] In the embodiments of the present application, the second function represents a service AS, and the third function is used for processing control plane data, for example, the third function can be a NCP. The first terminal represents a satellite mobile communication terminal, and the second terminal represents a ground mobile communication terminal. A voice call connection is established between the first terminal and the second terminal, that is, the first terminal triggers a call event through the first function, thereby establishing a voice call connection between the first terminal and the second terminal based on the call event.
[0102] It should be noted that, since the first terminal represents a satellite mobile communication terminal, the position information of the first terminal can be position information of the satellite mobile communication terminal, which is referred to as satellite position information for short.
[0103] In actual application, the position information of the first terminal can be carried in a first field, which is a field created by the first function in the first information. For example, the first field can be an scgi field, and specifically, the first field can be an scgi field set in the Location Type, that is, the position information of the satellite mobile communication terminal is carried in the scgi field in the Location Type.
[0104] Step 302: converting between the first codec and the second codec for the voice stream.
[0105] In the embodiments of the present application, the first codec corresponds to the first terminal, the first terminal represents a satellite mobile communication terminal, the second codec corresponds to the second terminal, the second terminal represents a ground mobile communication terminal, wherein the first codec represents a low-code-rate voice codec, and the second codec represents a ground general voice codec.
[0106] In actual application, after the first function controls the voice stream to be anchored to the fourth function based on the position information of the first terminal, the first function controls the fourth function to create a media resource (i.e., a first media resource) between the fourth function and the second terminal, and selects a corresponding voice codec based on the first media resource.
[0107] Based on this, in an embodiment, the conversion between the first codec and the second codec for the voice stream includes: receiving a third request initiated by the first function; the third request is used to request to create a first media resource, the first media resource representing a media resource between the fourth function and the second terminal; in response to the third request, creating the first media resource; based on the first media resource, determining that a first code stream is transmitted between the fourth function and the second terminal; the first code stream is associated with the second codec; and converting the second codec based on the first media resource.
[0108] Here, by using the media resource (i.e., the first media resource) created between the fourth function and the second terminal, it can be determined that the voice stream transmitted between the fourth function and the second terminal is a first code stream, and the first code stream is associated with the second codec, that is, the voice stream transmitted between the fourth function and the second terminal is voice coded in the manner of the ground general voice codec.
[0109] In an embodiment, after the first media resource is created in response to the third request, the method further includes: sending third information to the first function; the third information represents a confirmation message for creating the first media resource.
[0110] Here, after the fourth function successfully creates the first media resource, a confirmation message is generated for the event of successfully creating the first media resource, i.e., the third information is generated, and then the fourth function sends the third information to the first function to inform the first function that the first media resource is successfully created. Exemplarily, the third information can be 200 OK.
[0111] In actual application, after the fourth function successfully creates the first media resource, the first media resource can be updated to complete the conversion between the first codec and the second codec for the voice stream.
[0112] Based on this, in an embodiment, the converting the second codec based on the first media resource comprises: receiving a fourth request initiated by the first function; the fourth request is used to request to update the first media resource; updating the first media resource to obtain a second media resource in response to the fourth request; the second media resource represents the media resource between the fourth function and the first terminal; determining a second code stream transmitted between the fourth function and the first terminal based on the second media resource; the second code stream is associated with the first codec.
[0113] Here, after the first function controls the fourth function to create the first media resource between the fourth function and the second terminal, the first function can further control the fourth function to update the first media resource to obtain a second media resource, and the second media resource represents the media resource between the fourth function and the first terminal. Furthermore, the voice stream transmitted between the fourth function and the first terminal can be determined as a second code stream by using the second media resource, and the second code stream is associated with the first codec. That is, the voice stream transmitted between the fourth function and the first terminal is encoded by using the low-code-rate voice codec, so that the first terminal can interwork with the opposite terminal by using the low-code-rate voice codec without changing the existing terminals of the communication parties, thereby improving the voice quality of the call.
[0114] In an embodiment, after the first media resource is updated to obtain the second media resource in response to the fourth request, the method further comprises: sending fourth information to the first function; the fourth information represents a confirmation message for updating the first media resource.
[0115] Here, after the fourth function completes the update of the first media resource, a confirmation message is generated for the event of completing the update of the first media resource, i.e., the fourth information is generated, and then the fourth function sends the fourth information to the first function to inform the first function that the update of the first media resource is completed. Exemplarily, the fourth information can be 200 OK.
[0116] In actual application, after receiving the call control result notification message sent by the first function (VoLTE AS) through the third function (NCP), the second function (service AS) initiates a fifth request to the NCP and sends the fifth request to the fourth function (UMF) through the NCP, so that the UMF restores the voice feature of the first terminal. The call control result notification message includes an audio anchoring result, i.e., the voice stream is anchored to the fourth function, specifically, the voice stream is anchored to the UMF.
[0117] Based on this, in an embodiment, the method further includes: receiving a fifth request initiated by the third function; the fifth request is used to request to restore the voice feature of the first terminal; and in response to the fifth request, the voice feature of the first terminal is restored to obtain the original voice feature.
[0118] In an embodiment, the response to the fifth request to restore the voice feature of the first terminal to obtain the original voice feature includes: in response to the fifth request, the voiceprint feature of the first terminal is obtained by analyzing the fifth request; and based on the voiceprint feature of the first terminal, the voice feature of the first terminal is restored to obtain the original voice feature.
[0119] In the embodiments of the present application, the fifth request includes the voiceprint feature of the first terminal. For example, the voiceprint feature of the first terminal can be the voiceprint identity (ID) of the first terminal. The voice feature of the first terminal is a low-code-rate voice stream, which can also be referred to as a low-code-rate encoded stream.
[0120] Here, before the response to the fifth request to restore the voice feature of the first terminal to obtain the original voice feature, the method further includes: obtaining the voice feature of the first terminal by using a sending end encoding device, and sending the voice feature of the first terminal to a receiving end decoding device. In this way, the voice feature of the first terminal can be restored by using the receiving end decoding device of the fourth function to obtain the original voice feature.
[0121] Here, the sending end encoding device is the sending end encoding device of the fourth function, and the voice feature of the first terminal can be understood as the encoded voice feature of the first terminal. For obtaining the voice feature of the first terminal by using the sending end encoding device, the original audio stream collected by the first terminal can be received, the voice feature is extracted from the original audio stream, the voice feature is mapped to the corresponding phonemes or other speech units, and the voice feature is encoded by the phonemes or other speech units to obtain the encoded voice feature of the first terminal.
[0122] It should be noted that in the embodiments of the present application, the fourth function sending end encoding device can extract the sound features from the original audio stream by using the Mel-frequency cepstral coefficient (MFCC, Mel-Frequency Cepstral Coefficients), linear prediction cepstral coefficient (LPCC, Linear Prediction Cepstral Coefficients) or the like, and the embodiments of the present application do not limit this.
[0123] Here, the fourth function sending end encoding device sends the encoded sound features of the first terminal to the receiving end decoding device after obtaining the encoded sound features of the first terminal, so that the receiving end decoding device restores the sound features of the first terminal to obtain the original sound features.
[0124] Here, for the receiving end decoding device to restore the sound features of the first terminal to obtain the original sound features, the voiceprint features of the first terminal can be obtained first, the voiceprint features of the first terminal are fitted to obtain the fitted voiceprint features of the first terminal, and the sound features of the first terminal are restored based on the fitted voiceprint features of the first terminal to obtain the original sound features.
[0125] It should be noted that the voiceprint features of the first terminal can be system default voiceprint features, specifically, the first terminal user voiceprint features preloaded by phonemes or other speech units. After the receiving end decoding device obtains the original sound features, the method further comprises: sending the original sound features to a terminal playing device or a remote device.
[0126] By using the technical solution of the embodiments of the present application, the cell location information (i.e. the location information of the first terminal) in the call signaling of the satellite mobile communication terminal (i.e. the first terminal) is used to control the anchoring of the voice stream to the new call media node (i.e. the fourth function, which is used at least for processing user plane data), so that the fourth function is used to convert the voice stream between the first codec and the second codec, i.e. to convert between the corresponding low-code rate speech codec and the general ground speech codec, thereby realizing that the first terminal can use the low-code rate speech codec to interwork with the opposite terminal without changing the existing terminals of the communication parties, and further improving the call voice quality.
[0127] The present application will be described below in conjunction with application examples.
[0128] By using the related technical solution, the terminals of the communication parties need to support the new low-code rate speech codec at the same time, and the codec mode based on voice compression will always reduce the call voice quality. In order to solve this technical problem, the present application proposes the following corresponding solutions:
[0129] 1. Network dynamic transcoding, realizing intercommunication with existing mobile terminals (corresponding to the second terminal described above): using the new call network and service platform, anchoring the call audio stream for the satellite mobile communication terminal (corresponding to the first terminal described above), and performing corresponding conversion between the low-code-rate speech codec and the general speech codec on the ground, so as to realize intercommunication between the satellite mobile communication terminal and the existing mobile communication terminal using the low-code-rate speech codec without changing the terminals of the existing communication parties, thereby improving the call speech quality.
[0130] 2. Dynamic voice fitting, realizing high-quality speech under low-code-rate: using the new call network and service platform, when the satellite mobile communication terminal converts the low-code-rate speech codec into the general speech codec, dynamic fitting is performed using the voiceprint characteristics of the user, so that the satellite mobile communication terminal can receive and send high-quality speech under the premise of using the low-code-rate speech codec.
[0131] The present application proposes a method and system for low-code-rate speech call under the integration of satellite communication network and ground communication network (based on the integration of satellite communication network and ground communication network), which controls the anchoring of the speech stream to the new call media node (such as UMF) through the cell location information in the call signaling of the satellite mobile communication terminal on the existing new call network control node (such as NCP) and service control node (such as service AS), so as to subsequently convert and intercommunicate the low-code-rate speech codec of the satellite mobile communication terminal and the conventional general speech codec on the ground.
[0132] Among them, the system defines the mechanism of low-code-rate speech encoding and decoding for the satellite mobile communication terminal, as well as the network control nodes (such as Session Border Controller (SBC, Session Border Controller)-C (SBC control plane), VoLTE AS) and media nodes (such as SBC-U (SBC user plane) and UMF) in the network; at the same time, the network media node can use the voiceprint characteristics of the user under the control of the service control node to realize high-quality speech fitting for the low-code-rate speech encoding, and complete the conversion to the conventional general speech codec.
[0133] The low-code-rate speech encoding and decoding of the satellite mobile communication terminal and the network control node / media node are mainly completed by the sending end encoding device and the receiving end decoding device. Figure 5 The schematic diagram of the composition structure of the sending end encoding device of the embodiment of the present application is as follows: Figure 5As shown, the sending end coding device mainly includes four modules: (1) an audio stream receiving module, configured to receive original audio stream collected by a satellite mobile communication terminal or audio stream decoded from general voice coding; (2) a sound feature extraction module, configured to extract sound features from the audio stream by using MFCC, LPCC or the like; (3) a speech unit extraction module, configured to train in advance how to map sound features to corresponding phonemes or other speech units by using a large amount of speech data with labels by using a deep learning model such as a Hidden Markov Model (HMM) or a Recurrent Neural Network (RNN); and (4) a low-code-rate coding and sending module, configured to code sound features by using phonemes or other speech units and send the coded sound features (low-code-rate coded stream) to a remote end.
[0134] Figure 6 As shown in a constituent structure schematic diagram of a receiving end decoding device of an embodiment of the present application, Figure 4 the receiving end decoding device mainly includes four modules: (1) a low-code-rate coded stream receiving module, configured to receive a low-code-rate coded stream from a remote end; (2) a voiceprint feature fitting module, configured to restore original sound features from the received low-code-rate coded stream based on phonemes or other speech units and in combination with preloaded user voiceprint features. The terminal side can combine system default voiceprint features, and a network control node / media node can restore sound features according to user subscription voiceprint features indicated by a service node; (3) a sound feature restoring module, configured to restore original audio stream at the terminal side or re-encode into general voice coding and decoding at the network side based on the fitted voiceprint features; and (4) an audio stream sending module, configured to send the re-coded and decoded audio stream to a terminal playing device or a remote device.
[0135] The voice processing method of the present application is described in detail below by taking a first function as VoLTE AS, a second function as service AS, a third function as NCP, a fourth function as UMF, a first terminal as user equipment (UE) A (UE A is a satellite mobile communication terminal), and a second terminal as UE B (UE B is a ground mobile communication terminal) as examples. Figure 6 As shown in a flowchart of a voice processing method of an embodiment of the present application, Figure 7 the voice processing method is used to implement low-code-rate voice calls under a space-ground integrated networking, as shown in Figure 1 a service flow of the voice processing method is as follows:
[0136] Step 1-6: UE A sends an INVITE request to SBC-C, wherein the INVITE request is sent in the form of Session Description Protocol (SDP) 1, and the SDP 1 carries a low-code-rate speech codec SAT NB codec (corresponding to the first codec described above); the SBC-C acquires user location information, i.e., user location information of UE A (corresponding to the location information of the first terminal described above), saves the user location information in utran_cell_id_3gpp, and after supplementing a ground general-purpose speech codec (such as an Adaptive Multi-Rate (AMR) codec (corresponding to the second codec described above)) in the SDP 1, forwards the INVITE request to a mobile communication terminal UE B at the opposite end.
[0137] Steps 7-11: After UE B receives the INVITE request, the UE B selects the AMR codec as the preferred codec in the returned 183 message; the SBC-C forwards the 183 message to UE A, and selects the SAT NB codec as the preferred codec.
[0138] Steps 12-22: UE A initiates an Update request to the SBC-C, wherein the Update request is sent in the form of SDP 2, and the SDP 2 carries the SAT NB codec; the SBC-C forwards the Update request to UE B, and selects the AMR codec; after UE B receives the Update request, the UE B returns a 200 OK message to the I / S-CSCF and returns a 180 message to UE A. Wherein the SAT NB codec speech stream is transmitted between UE A and SBC-U, and the AMR codec speech stream is transmitted between SBC-U and UE B.
[0139] Steps 23-25: UE B replies with a 200 OK message to the I / S-CSCF, the 200 OK message is forwarded to the VoLTE AS through the I / S-CSCF, the VoLTE AS sends a call event notification (ANSWER, corresponding to the first information described above) to the NCP, and satellite location information of UE A is carried in the scgi field in the LocationType (the original ecgi is LTE location information, the ncgi is NR location information, and the scgi is newly added to represent satellite location information).
[0140] Steps 26-31: The NCP initiates call event control (ANCHOR, corresponding to the first request mentioned above) to the VoLTE AS; the VoLTE AS initiates a Re-Invite request to the UE B, the UE B updates the SDP and selects the preferred AMR codec, and the VoLTE AS controls the UMF to create media resources between the UMF and the UE B (corresponding to the first media resources mentioned above) and selects the AMR codec.
[0141] Steps 32-33: VoLTE AS initiates an Update request to UE A to update the SDP of UMF and select SAT_NB codec.
[0142] Steps 34-41: UE A replies with a 200 OK message to VoLTE AS, selects SAT_NB codec in SDP, and VoLTE AS controls UMF to update the media resources between UMF and UE A (corresponding to the aforementioned second media resources), and selects SAT_NB codec. After audio anchoring is completed, SAT_NB codec is used for the voice stream between UE A and UMF, and AMR codec is used for the voice stream between UMF and UE B.
[0143] Steps 42-46: After receiving the call control result notification (including audio anchoring result) sent by the VoLTE AS through the NCP, the service AS initiates a voiceprint restoration transcoding request to the NCP (corresponding to the aforementioned fifth request), and sends the voiceprint ID subscribed by UE A to the UMF through the NCP, so that the UMF can use the receiver decoding device to complete the re-encoding and decoding from SAT_NB codec to AMR codec based on the voiceprint model, and use the transmitter encoding device to complete the re-encoding and decoding from AMR codec to SAT_NB codec.
[0144] This application proposes a method for anchoring the voice stream to a new call media node (such as UMF) using cell location information in the call signaling of a satellite mobile communication terminal and by utilizing a new call network and service layer control. It also proposes a transmitting end encoding device and a receiving end decoding device for low bit rate voice calls, which transmits low bit rate voice through encoding and decoding in the satellite communication network, thereby saving communication bandwidth, while maintaining high-quality call voice through voiceprint feature fitting.
[0145] To implement the voice processing method on the first functional side of this application embodiment, this application embodiment also provides a voice processing apparatus. This apparatus is applied to the first function, which is at least used for processing voice service data. Figure 7 This is a schematic diagram of the composition structure of the voice processing device according to an embodiment of this application. Figure 8 ,like Figure 2 As shown, the device includes:
[0146] The first receiving unit 71 is configured to receive a first request initiated by a second function through a third function; the first request is triggered by the second function based on first information, the first information comprises location information of a first terminal; the first request is used to request control over a call event between the first terminal and a second terminal; the second function represents a service application server, and the third function is used to process at least control plane data;
[0147] The control unit 72 is configured to, in response to the first request, control anchoring of a voice stream to a fourth function based on the location information of the first terminal, so as to perform conversion between a first codec and a second codec for the voice stream through the fourth function; the fourth function is used to process at least user plane data.
[0148] In an embodiment, the apparatus further comprises a second sending unit; and wherein
[0149] The second sending unit is configured to, before the first receiving unit 71 receives the first request initiated by the second function through the third function, send the first information to the third function, so that the third function sends the first information to the second function.
[0150] The first information represents a call event notification message.
[0151] In an embodiment, the second sending unit is specifically configured to:
[0152] determine whether second information is received; the second information represents a reply confirmation message sent by the second terminal, and the second information is generated by the second terminal based on a received second request initiated by the first terminal, the second request representing a codec update request for the voice stream;
[0153] In a case where it is determined that the second information is received, the first information is sent to the third function.
[0154] In an embodiment, the first receiving unit 71 is specifically configured to:
[0155] In a case where the first information is received by the second function, the first request sent by the third function is received;
[0156] The first request is sent to the third function by the second function based on the first information.
[0157] In an embodiment, the apparatus further comprises a creating unit and an adding unit; and wherein
[0158] The creating unit is configured to create a first field in the first information.
[0159] The adding unit is configured to add the position information of the first terminal in the first field.
[0160] In an embodiment, the apparatus further comprises a third sending unit and a second receiving unit; wherein,
[0161] The third sending unit is configured to, after the control unit 72 controls the anchoring of the voice stream to the fourth function based on the position information of the first terminal in response to the first request, initiate a third request to the fourth function; the third request is configured to request the fourth function to create a first media resource; the first media resource represents a media resource between the fourth function and the second terminal, and the first media resource is configured to determine the transmission of a first code stream between the fourth function and the second terminal, the first code stream being associated with the second codec.
[0162] The second receiving unit is configured to receive third information sent by the fourth function; the third information represents a confirmation message for the creation of the first media resource.
[0163] In an embodiment, the apparatus further comprises a fourth sending unit and a third receiving unit; wherein,
[0164] The fourth sending unit is configured to, after the second receiving unit receives the third information sent by the fourth function, initiate a fourth request to the fourth function; the fourth request is configured to request an update of the first media resource to obtain a second media resource; the second media resource represents a media resource between the fourth function and the first terminal, and the second media resource is configured to determine the transmission of a second code stream between the fourth function and the first terminal, the second code stream being associated with the first codec.
[0165] The third receiving unit is configured to receive fourth information sent by the fourth function; the fourth information represents a confirmation message for the update of the first media resource.
[0166] In an embodiment, the first terminal represents a satellite mobile communication terminal, and the second terminal represents a ground mobile communication terminal; wherein the first codec corresponds to the first terminal, and the second codec corresponds to the second terminal.
[0167] In actual application, the first receiving unit 71 can be implemented by a communication interface in a voice processing apparatus; and the control unit 72 can be implemented by a processor in the voice processing apparatus.
[0168] In order to implement the voice processing method of the second function side according to the embodiments of the present application, another voice processing apparatus is further provided in the embodiments of the present application, which is applied to a second function, the second function representing a service application server, Figure 8Structure of a voice processing device according to an embodiment of the present application Figure 9 As shown in Figure 3 The device comprises:
[0169] A first sending unit 81 is configured to initiate a first request to a first function through a third function; the first request is triggered based on first information, the first information comprising location information of a first terminal; the first request is used to request control over a call event between the first terminal and a second terminal; the first function is used to process at least voice service data, and the third function is used to process at least control plane data.
[0170] The location information of the first terminal is used by the first function to control anchoring of a voice stream to a fourth function, so as to convert between a first codec and a second codec for the voice stream through the fourth function; the fourth function is used to process at least user plane data.
[0171] In an embodiment, the device further comprises a fourth receiving unit, wherein:
[0172] The fourth receiving unit is configured to receive the first information sent by the first function through the third function, before the first sending unit 81 initiates the first request to the first function through the third function.
[0173] The first information represents a call event notification message.
[0174] In an embodiment, the first sending unit 81 is specifically configured to:
[0175] determine whether the first information is received;
[0176] if it is determined that the first information is received, send the first request to the third function based on the first information, so that the third function sends the first request to the first function.
[0177] In an embodiment, the first terminal represents a satellite mobile communication terminal, and the second terminal represents a ground mobile communication terminal; the first codec corresponds to the first terminal, and the second codec corresponds to the second terminal.
[0178] In actual application, the first sending unit 81 can be implemented by a communication interface in the voice processing device.
[0179] In order to implement the voice processing method on the fourth function side according to an embodiment of the present application, another voice processing device is further provided in an embodiment of the present application, which is applied to a fourth function, the fourth function being used to process at least user plane data, Figure 9 Structure of a voice processing device according to an embodiment of the present applicationFigure 10 As shown in Figure 10 the apparatus comprises:
[0180] a first obtaining unit 91, configured to obtain a voice stream; the voice stream is controlled by a first function based on position information of a first terminal in response to a first request, the first function being used at least for processing voice service data;
[0181] a converting unit 92, configured to convert the voice stream between a first codec and a second codec;
[0182] wherein the first request is sent by a second function to the first function through a third function, the first request being triggered by the second function based on first information, the first information comprising the position information of the first terminal; the first request is used for requesting control over a call event between the first terminal and a second terminal; the second function represents a service application server, and the third function is used at least for processing control plane data.
[0183] In an embodiment, the converting unit 92 comprises a first receiving subunit, a creating subunit, a first determining subunit, and a converting subunit; wherein,
[0184] the first receiving subunit is configured to receive a third request initiated by the first function; the third request is used for requesting creation of a first media resource, the first media resource representing a media resource between the fourth function and the second terminal;
[0185] the creating subunit is configured to create the first media resource in response to the third request;
[0186] the first determining subunit is configured to determine, based on the first media resource, that a first code stream is to be transmitted between the fourth function and the second terminal; the first code stream is associated with the second codec;
[0187] the converting subunit is configured to convert the second codec based on the first media resource.
[0188] In an embodiment, the apparatus further comprises a first sending subunit; wherein,
[0189] the first sending subunit is configured to send third information to the first function after the creating subunit creates the first media resource in response to the third request; the third information represents a confirmation message for creating the first media resource.
[0190] In an embodiment, the converting subunit comprises a second receiving subunit, an updating subunit, and a second determining subunit; wherein,
[0191] The second receiving sub-unit is configured to receive a fourth request initiated by the first function, wherein the fourth request is used to request to update the first media resource.
[0192] The updating sub-unit is configured to update the first media resource to obtain a second media resource in response to the fourth request, wherein the second media resource represents a media resource between the fourth function and the first terminal.
[0193] The second determining sub-unit is configured to determine to transmit a second code stream between the fourth function and the first terminal based on the second media resource, wherein the second code stream is associated with the first codec.
[0194] In an embodiment, the apparatus further comprises a second sending sub-unit, and wherein
[0195] The second sending sub-unit is configured to send fourth information to the first function after the updating sub-unit updates the first media resource to obtain the second media resource in response to the fourth request, wherein the fourth information represents a confirmation message for updating the first media resource.
[0196] In an embodiment, the apparatus further comprises a fifth receiving unit and a restoration unit, and wherein
[0197] The fifth receiving unit is configured to receive a fifth request initiated by the third function, wherein the fifth request is used to request to restore a voice feature of the first terminal.
[0198] The restoration unit is configured to restore the voice feature of the first terminal to obtain an original voice feature in response to the fifth request.
[0199] In an embodiment, the restoration unit is specifically configured to:
[0200] restore a voiceprint feature of the first terminal in response to the fifth request;
[0201] restore the voice feature of the first terminal based on the voiceprint feature of the first terminal to obtain the original voice feature.
[0202] In an embodiment, the first terminal represents a satellite mobile communication terminal, and the second terminal represents a ground mobile communication terminal, wherein the first codec corresponds to the first terminal, and the second codec corresponds to the second terminal.
[0203] In actual application, the first acquiring unit 91 can be implemented by a communication interface in a voice processing apparatus, and the converting unit 92 can be implemented by a processor in the voice processing apparatus.
[0204] It should be noted that the voice processing apparatus provided in the above embodiments is only used as an example to illustrate the division of the above program modules, and in actual application, the above processing can be completed by different program modules according to needs, that is, the internal structure of the apparatus is divided into different program modules to complete all or part of the above-described processing. In addition, the voice processing apparatus and the voice processing method provided in the above embodiments belong to the same concept, and the specific implementation process is described in the voice processing method embodiments, which will not be repeated here.
[0205] Based on the hardware implementation of the above program modules, and in order to realize the first function side voice processing method of the embodiment of the present application, the embodiment of the present application further provides a first function, Figure 10 The schematic diagram of the hardware composition structure of the first function of the embodiment of the present application is shown in Figure 11 The first function 1000 includes:
[0206] The first communication interface 1001 can interact with other devices (such as the second function, the fourth function) to exchange information;
[0207] The first processor 1002 is connected with the first communication interface 1001 to realize information interaction with other devices (such as the second function, the fourth function), and is used to run the computer program to execute the above-mentioned first function side voice processing method, and the computer program is stored on the first memory 1003.
[0208] It should be noted that the specific processing process of the first communication interface 1001 and the first processor 1002 can be understood with reference to the above-mentioned first function side voice processing method.
[0209] Of course, in actual application, each component in the first function 1000 is coupled together through the first bus system 1004. It can be understood that the first bus system 1004 is used to realize the connection communication between the components. The first bus system 1004 includes not only the data bus, but also the power bus, the control bus and the state signal bus. However, in order to clearly illustrate, all kinds of buses are marked as the first bus system 1004 in Figure 11 .
[0210] The first memory 1003 in the embodiment of the present application is used to store various types of data to support the operation of the first function 1000. Examples of these data include: any computer program used to operate on the first function 1000.
[0211] The first function side voice processing method disclosed in the embodiments of the present application can be applied to the first processor 1002 or implemented by the first processor 1002. The first processor 1002 can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the first function side voice processing method disclosed above can be completed by integrated logic circuits of hardware in the first processor 1002 or instructions in the form of software. The first processor 1002 disclosed above can be a general processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The first processor 1002 can implement or execute the first function side voice processing method, steps and logic block diagram disclosed in the embodiments of the present application. The general processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the first function side voice processing method disclosed in the embodiments of the present application, the execution can be directly completed by a hardware decoding processor or a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the first memory 1003, and the first processor 1002 reads information in the first memory 1003 to complete the steps of the first function side voice processing method in combination with the hardware thereof.
[0212] In the exemplary embodiments, the first function 1000 can be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), general processors, controllers, micro controllers (MCUs), microprocessors (Microprocessors), or other electronic elements, for executing the first function side voice processing method disclosed above.
[0213] Based on the hardware implementation of the program module and in order to implement the second function side voice processing method of the embodiments of the present application, the embodiments of the present application further provide a second function, Figure 11 The hardware structure of the second function of the embodiments of the present application is shown in FIG. 11, which includes: Figure 12
[0214] The second communication interface 1101 is configured to interact with other devices (e.g., the first function) to exchange information.
[0215] The second processor 1102 is connected with the second communication interface 1101 to exchange information with other devices (e.g., the first function), and is configured to execute a computer program stored in the second memory 1103 to implement the second function side voice processing method.
[0216] It should be noted that the specific processing process of the second communication interface 1101 and the second processor 1102 can be understood with reference to the above-mentioned second function side voice processing method.
[0217] Of course, in actual application, various components in the second function 1100 are coupled together through the second bus system 1104. It can be understood that the second bus system 1104 is used to realize the connection and communication between the components. The second bus system 1104 includes not only a data bus, but also a power bus, a control bus and a status signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the second bus system 1104 in the Figure 12
[0218] The second memory 1103 in the embodiment of the present application is used to store various types of data to support the operation of the second function 1100. Examples of these data include: any computer program used for operation on the second function 1100.
[0219] The second function side voice processing method disclosed in the above-mentioned embodiments of the present application can be applied in or implemented by the second processor 1102. The second processor 1102 can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above-mentioned second function side voice processing method can be completed by integrated logic circuits or instructions in the form of software in the second processor 1102. The above-mentioned second processor 1102 can be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The second processor 1102 can implement or execute the second function side voice processing method, steps and logic block diagram disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the second function side voice processing method disclosed in the embodiments of the present application, the hardware decoding processor can be directly executed to complete, or the combination of hardware and software modules in the decoding processor can be executed to complete. The software module can be located in a storage medium, which is located in the second memory 1103, and the second processor 1102 reads the information in the second memory 1103 to complete the above-mentioned steps of the second function side voice processing method in combination with the hardware.
[0220] In an example embodiment, the second function 1100 can be implemented by one or more ASICs, DSPs, PLDs, CPLDs, FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic elements for performing the aforementioned second function side voice processing method.
[0221] Based on the hardware implementation of the above program module, and in order to implement the fourth function side voice processing method of the embodiments of the present application, the embodiments of the present application also provide a fourth function, Figure 12 The schematic diagram of the hardware composition structure of the fourth function of the embodiments of the present application is shown in The fourth function 1200 includes:
[0222] The third communication interface 1201 can interact with other devices (such as the first function) for information exchange;
[0223] The third processor 1202 is connected with the third communication interface 1201 to realize information exchange with other devices (such as the first function), and is used to run a computer program to execute the fourth function side voice processing method provided above, and the computer program is stored on the third memory 1203.
[0224] It should be noted that the specific processing process of the third communication interface 1201 and the third processor 1202 can be understood with reference to the above-mentioned fourth function side voice processing method.
[0225] Of course, in actual application, each component in the fourth function 1200 is coupled together through the third bus system 1204. It can be understood that the third bus system 1204 is used to realize the connection communication between the components. The third bus system 1204 includes not only a data bus, but also a power bus, a control bus and a state signal bus. However, in order to clearly illustrate, all kinds of buses are marked as the third bus system 1204 in .
[0226] The third memory 1203 in the embodiments of the present application is used to store various types of data to support the operation of the fourth function 1200. Examples of these data include: any computer program used to operate on the fourth function 1200.
[0227] The fourth function side voice processing method disclosed in the embodiments of the present application can be applied to the third processor 1202 or implemented by the third processor 1202. The third processor 1202 can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the fourth function side voice processing method can be completed by integrated logic circuit of hardware in the third processor 1202 or instructions in the form of software. The third processor 1202 can be a general processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The third processor 1202 can implement or execute each fourth function side voice processing method, step and logic block diagram disclosed in the embodiments of the present application. The general processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the fourth function side voice processing method disclosed in the embodiments of the present application, the hardware decoding processor can be directly implemented or executed by the hardware and software module combination in the decoding processor. The software module can be located in the storage medium, which is located in the third memory 1203. The third processor 1202 reads the information in the third memory 1203 and combines the hardware to complete the steps of the aforementioned fourth function side voice processing method.
[0228] In the exemplary embodiments, the fourth function 1200 can be implemented by one or more ASICs, DSPs, PLDs, CPLDs, FPGAs, general processors, controllers, MCUs, microprocessors, or other electronic elements, for executing the aforementioned fourth function side voice processing method.
[0229] It can be understood that the memory (including the first memory 1003, the second memory 1103 and the third memory 1203) of the embodiments of the present application can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM). The magnetic surface memory can be a magnetic disk memory or a magnetic tape memory. The volatile memory can be a random access memory (RAM) used as an external cache.By way of example and not limitation, many forms of RAM can be used, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), Direct Rambus Random Access Memory (DRRAM). The memory described in embodiments of the present application (including the first memory 1003, the second memory 1103, and the third memory 1203) is intended to include, but not be limited to, these and any other suitable types of memory.
[0230] In exemplary embodiments, the present application also provides a storage medium, i.e., a computer storage medium, specifically a computer readable storage medium, such as the first memory 1003 storing a computer program executable by the first processor 1002 in the first function 1000 to complete the steps of the voice processing method on the first function side described in the foregoing embodiments of the present application, or the second memory 1103 storing a computer program executable by the second processor 1102 in the second function 1100 to complete the steps of the voice processing method on the second function side described in the foregoing embodiments of the present application, or the third memory 1203 storing a computer program executable by the third processor 1202 in the fourth function 1200 to complete the steps of the voice processing method on the fourth function side described in the foregoing embodiments of the present application. The computer readable storage medium can be FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.
[0231] In an example embodiment, the embodiments of the present application also provide a computer program product comprising a computer program, which can be executed by the first processor 1002 in the first function 1000 to complete the steps of the voice processing method on the first function side as described in the foregoing embodiments of the present application, or can be executed by the second processor 1102 in the second function 1100 to complete the steps of the voice processing method on the second function side as described in the foregoing embodiments of the present application, or can be executed by the third processor 1202 in the fourth function 1200 to complete the steps of the voice processing method on the fourth function side as described in the foregoing embodiments of the present application.
[0232] It should be noted that "first", "second", "third", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0233] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0234] The above is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A speech processing method, characterized in that, Applied to a first function, which is at least used for processing voice service data, the method includes: The system receives a first request initiated by the second function through the third function; the first request is triggered by the second function based on first information, the first information including the location information of the first terminal; the first request is used to request control over a call event between the first terminal and the second terminal; the second function represents a business application server, and the third function is at least used to process control plane data. In response to the first request, the voice stream is anchored to a fourth function based on the location information of the first terminal, so as to perform a conversion between a first codec and a second codec on the voice stream through the fourth function; the fourth function is at least used to process user plane data.
2. The method according to claim 1, characterized in that, Before receiving the first request initiated by the second function through the third function, the method further includes: Send the first information to the third function, so that the third function sends the first information to the second function; The first information represents a call event notification message.
3. The method according to claim 2, characterized in that, Sending the first information to the third function includes: Determine whether the second information has been received; the second information represents a reply confirmation message sent by the second terminal, and the second information is generated by the second terminal based on a second request initiated by the first terminal, and the second request represents a codec update request for the voice stream; If the second information is received, the first information is sent to the third function.
4. The method according to claim 2, characterized in that, The receiving of the first request initiated by the second function through the third function includes: If the second function receives the first information, it receives the first request sent by the third function; The first request is sent to the third function by the second function based on the first information.
5. The method according to claim 1, characterized in that, The method further includes: Create a first field in the first information; Add the location information of the first terminal to the first field.
6. The method according to claim 1, characterized in that, After responding to the first request and anchoring the voice stream to the fourth function based on the location information of the first terminal, the method further includes: A third request is initiated to the fourth function; the third request is used to request the fourth function to create a first media resource; the first media resource represents the media resource between the fourth function and the second terminal, the first media resource is used to determine the transmission of a first bitstream between the fourth function and the second terminal, and the first bitstream is associated with the second codec. Receive the third information sent by the fourth function; the third information represents a confirmation message for creating the first media resource.
7. The method according to claim 6, characterized in that, After receiving the third information sent by the fourth function, the method further includes: A fourth request is initiated to the fourth function; the fourth request is used to request an update to the first media resource to obtain a second media resource; the second media resource represents the media resource between the fourth function and the first terminal, and the second media resource is used to determine the transmission of a second bitstream between the fourth function and the first terminal, and the second bitstream is associated with the first codec; Receive the fourth information sent by the fourth function; the fourth information represents a confirmation message for updating the first media resource.
8. The method according to any one of claims 1 to 7, characterized in that, The first terminal represents a satellite mobile communication terminal, and the second terminal represents a terrestrial mobile communication terminal; Wherein, the first codec corresponds to the first terminal, and the second codec corresponds to the second terminal.
9. A speech processing method, characterized in that, Applied to a second function, which represents a business application server, the method includes: A first request is initiated to the first function through the third function; the first request is triggered based on first information, the first information including the location information of the first terminal; the first request is used to request control over a call event between the first terminal and the second terminal; the first function is used to process at least voice service data, and the third function is used to process at least control plane data. The location information of the first terminal is used by the first function to anchor the voice stream to the fourth function, so that the fourth function can perform a conversion between the first codec and the second codec for the voice stream; the fourth function is at least used to process user plane data.
10. The method according to claim 9, characterized in that, Before initiating the first request to the first function via the third function, the method further includes: Receive the first information sent by the first function through the third function; The first information represents a call event notification message.
11. The method according to claim 9, characterized in that, The step of initiating a first request to the first function through the third function includes: Determine whether the first information has been received; Upon confirming receipt of the first information, the system triggers the sending of the first request to the third function based on the first information, so that the third function sends the first request to the first function.
12. The method according to any one of claims 9 to 11, characterized in that, The first terminal represents a satellite mobile communication terminal, and the second terminal represents a terrestrial mobile communication terminal; Wherein, the first codec corresponds to the first terminal, and the second codec corresponds to the second terminal.
13. A speech processing method, characterized in that, Applied to a fourth function, the fourth function being used to process user plane data, the method includes: Acquire a voice stream; the voice stream is controlled and anchored by a first function in response to a first request, based on the location information of a first terminal, and the first function is at least used to process voice service data; The speech stream is converted between a first codec and a second codec; Wherein, the first request is sent from the second function to the first function through the third function, the first request is triggered by the second function based on first information, the first information including the location information of the first terminal; the first request is used to request control over the call event between the first terminal and the second terminal; the second function represents the business application server, and the third function is at least used to process control plane data.
14. The method according to claim 13, characterized in that, The conversion between the first codec and the second codec for the speech stream includes: Receive a third request initiated by the first function; the third request is used to request the creation of a first media resource, the first media resource representing the media resource between the fourth function and the second terminal; In response to the third request, the first media resource is created; Based on the first media resource, it is determined that the fourth function transmits a first bitstream between itself and the second terminal; the first bitstream is associated with the second codec. Based on the first media resource, the second codec is converted.
15. The method according to claim 14, characterized in that, After creating the first media resource in response to the third request, the method further includes: Send a third message to the first function; the third message represents a confirmation message for the creation of the first media resource.
16. The method according to claim 14, characterized in that, The conversion of the second codec based on the first media resource includes: Receive a fourth request initiated by the first function; the fourth request is used to request an update to the first media resource; In response to the fourth request, the first media resource is updated to obtain the second media resource; the second media resource represents the media resource between the fourth function and the first terminal; Based on the second media resource, it is determined that the fourth function transmits a second bitstream between itself and the first terminal; the second bitstream is associated with the first codec.
17. The method according to claim 16, characterized in that, After updating the first media resource in response to the fourth request to obtain the second media resource, the method further includes: Send a fourth message to the first function; the fourth message represents a confirmation message for updating the first media resource.
18. The method according to claim 13, characterized in that, The method further includes: Receive a fifth request initiated by the third function; the fifth request is used to request the restoration of the sound features of the first terminal; In response to the fifth request, the sound features of the first terminal are restored to obtain the original sound features.
19. The method according to claim 18, characterized in that, In response to the fifth request, the sound features of the first terminal are restored to obtain the original sound features, including: In response to the fifth request, the fifth request is parsed to obtain the voiceprint features of the first terminal; Based on the voiceprint features of the first terminal, the sound features of the first terminal are restored to obtain the original sound features.
20. The method according to any one of claims 13 to 19, characterized in that, The first terminal represents a satellite mobile communication terminal, and the second terminal represents a terrestrial mobile communication terminal; Wherein, the first codec corresponds to the first terminal, and the second codec corresponds to the second terminal.
21. A voice processing device, characterized in that, The apparatus is applied to a first function, which is at least used for processing voice service data, and includes: The first receiving unit is configured to receive a first request initiated by the second function through the third function; the first request is triggered by the second function based on first information, the first information including the location information of the first terminal; the first request is used to request control over the call event between the first terminal and the second terminal; the second function represents a business application server, and the third function is at least used to process control plane data. A control unit is configured to, in response to the first request, control the anchoring of the voice stream to a fourth function based on the location information of the first terminal, so as to perform a conversion between a first codec and a second codec on the voice stream through the fourth function; the fourth function is at least used to process user plane data.
22. A voice processing device, characterized in that, The apparatus is applied to a second function, which represents a business application server, and includes: The first sending unit is configured to initiate a first request to the first function through a third function; the first request is triggered based on first information, the first information including the location information of the first terminal; the first request is used to request control over a call event between the first terminal and the second terminal; the first function is at least used to process voice service data, and the third function is at least used to process control plane data. The location information of the first terminal is used by the first function to anchor the voice stream to the fourth function, so that the fourth function can perform a conversion between the first codec and the second codec for the voice stream; the fourth function is at least used to process user plane data.
23. A voice processing device, characterized in that, The apparatus is applied to a fourth function, which is at least used for processing user plane data, and includes: The first acquisition unit is used to acquire a voice stream; the voice stream is controlled and anchored by a first function in response to a first request based on the location information of a first terminal, and the first function is at least used to process voice service data; A conversion unit is used to convert between a first codec and a second codec for the speech stream; Wherein, the first request is sent from the second function to the first function through the third function, the first request is triggered by the second function based on first information, the first information including the location information of the first terminal; the first request is used to request control over the call event between the first terminal and the second terminal; the second function represents the business application server, and the third function is at least used to process control plane data.
24. A first function, characterized in that, include: A first processor and a first memory for storing computer programs capable of running on the first processor; Wherein, when the first processor is used to run the computer program, it performs the steps of the method according to any one of claims 1 to 8.
25. A second function, characterized in that, include: A second processor and a second memory for storing computer programs capable of running on the second processor; Wherein, when the second processor is used to run the computer program, it performs the steps of the method according to any one of claims 9 to 12.
26. A fourth function, characterized in that, include: A third processor and a third memory for storing computer programs capable of running on the third processor; When the third processor runs the computer program, it performs the steps of the method according to any one of claims 13 to 20.
27. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8, or the steps of the method according to any one of claims 9 to 12, or the steps of the method according to any one of claims 13 to 20.
28. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8, or the steps of the method according to any one of claims 9 to 12, or the steps of the method according to any one of claims 13 to 20.
Citation Information
Patent Citations
Ground network in satellite communication system with nodes carrying speech at a high data rate
WO1997044918A1
Call processing method, apparatus and system
WO2023016172A1