Call processing method and apparatus, user equipment and network-side device

CN122845569APending Publication Date: 2026-09-29VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510385024.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]本申请实施例提供一种通话处理方法、装置、用户设备及网络侧设备,能够解决由于通话的网络环境变化,导致语音通话的效果较差的问题

Benefits of technology

[0033]本申请实施例中,通过在第一用户设备建立通话过程中,第一网络实体与所述第一用户设备进行会话描述协议SDP协商,所述SDP协商用于确定第一语音编码和第二语音编码。这样,由于在建立通话过程中协商了两种语音编码,从而可以实现语音编码的切换,以适应不同的网络环境的变化,提高语音通话的效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845569A_ABST
    Figure CN122845569A_ABST
Patent Text Reader

Abstract

The application discloses a call processing method and device, user equipment and network side equipment, and belongs to the technical field of communication. The call processing method comprises the following steps: in a call establishment process of first user equipment, a first network entity performs session description protocol (SDP) negotiation with the first user equipment, and the SDP negotiation is used for determining a first voice encoding and a second voice encoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technology, specifically relating to a call processing method, apparatus, user equipment, and network-side equipment. Background Technology

[0002] In communications, network-side devices typically establish voice bearers supporting a specific voice coding system for user equipment, and voice calls are made based on these bearers. However, changes in the network environment can lead to degraded call quality or even call interruption. Therefore, existing technologies suffer from the problem of poor voice call quality due to changes in the network environment. Summary of the Invention

[0003] This application provides a call processing method, apparatus, user equipment, and network-side equipment, which can solve the problem of poor voice call quality caused by changes in the network environment.

[0004] Firstly, a call processing method is provided, including:

[0005] During the establishment of a call by the first user equipment, the first network entity negotiates a Session Description Protocol (SDP) with the first user equipment. The SDP negotiation is used to determine the first voice codec and the second voice codec.

[0006] Secondly, a voice call processing method is provided, including:

[0007] The second network entity performs a first operation, which includes at least one of the following:

[0008] The first data received from the second user equipment is converted into second data, and the second data is sent to the first user equipment. The first data is the data corresponding to the first voice code, and the second data is the data corresponding to the second voice code.

[0009] The third data received from the first user equipment is converted into fourth data and sent to the second user equipment. The third data is the data corresponding to the second voice code, and the fourth data is the data corresponding to the first voice code.

[0010] Thirdly, a voice call processing method is provided, including:

[0011] During the establishment of a call by the first user equipment, the first user equipment negotiates a Session Description Protocol (SDP) with the first network entity. The SDP negotiation is used to determine the first voice code and the second voice code.

[0012] Fourthly, a call processing device is provided, comprising:

[0013] The first transmission module is used to negotiate Session Description Protocol (SDP) with the first user equipment during the establishment of a call. The SDP negotiation is used to determine the first voice code and the second voice code.

[0014] Fifthly, a voice call processing device is provided, comprising:

[0015] The second transmission module is configured to perform a first operation, the first operation including at least one of the following:

[0016] The first data received from the second user equipment is converted into second data, and the second data is sent to the first user equipment. The first data is the data corresponding to the first voice code, and the second data is the data corresponding to the second voice code.

[0017] The third data received from the first user equipment is converted into fourth data and sent to the second user equipment. The third data is the data corresponding to the second voice code, and the fourth data is the data corresponding to the first voice code.

[0018] Sixthly, a voice call processing device is provided, comprising:

[0019] The third transmission module is used to negotiate Session Description Protocol (SDP) with the first network entity during the call establishment process of the first user equipment. The SDP negotiation is used to determine the first voice codec and the second voice codec.

[0020] In a seventh aspect, a voice call processing apparatus is provided, the apparatus being configured to perform the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect, or to implement the steps of the method described in the third aspect.

[0021] Eighthly, a user equipment is provided, the user equipment including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the third aspect.

[0022] In a ninth aspect, a user equipment is provided, including a processor and a communication interface, wherein the communication interface is used to negotiate a Session Description Protocol (SDP) with a first network entity during the establishment of a call by the first user equipment, the SDP negotiation being used to determine a first voice codec and a second voice codec.

[0023] In a tenth aspect, a network-side device is provided, the network-side device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect, or implementing the steps of the method as described in the second aspect.

[0024] Eleventhly, a network-side device is provided, including a processor and a communication interface, wherein,

[0025] When the network-side device is the first network entity, the communication interface is used to negotiate the Session Description Protocol (SDP) with the first user equipment during the call establishment process. The SDP negotiation is used to determine the first voice code and the second voice code.

[0026] When the network-side device is a second network entity, the communication interface is used to perform a first operation, which includes at least one of the following:

[0027] The first data received from the second user equipment is converted into second data, and the second data is sent to the first user equipment. The first data is the data corresponding to the first voice code, and the second data is the data corresponding to the second voice code.

[0028] The third data received from the first user equipment is converted into fourth data and sent to the second user equipment. The third data is the data corresponding to the second voice code, and the fourth data is the data corresponding to the first voice code.

[0029] In a twelfth aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect, or the steps of the method described in the second aspect, or the steps of the method described in the third aspect.

[0030] In a thirteenth aspect, a wireless communication system is provided, comprising: a user equipment, a first network entity, and a second network entity, wherein the user equipment is configured to perform the steps of the method described in the third aspect, the first network entity is configured to perform the steps of the method described in the first aspect, and the second network entity is configured to perform the steps of the method described in the second aspect.

[0031] In a twelfth aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the method described in the first aspect, or the steps of the method described in the second aspect, or the steps of the method described in the third aspect.

[0032] In a thirteenth aspect, a computer program / program product is provided, the computer program / program product being stored in a storage medium, the computer program / program product being executed by at least one processor to implement the steps of the method as described in the first aspect, or the steps of the method as described in the second aspect, or the steps of the method as described in the third aspect.

[0033] In this embodiment, during the call establishment process of the first user equipment, the first network entity negotiates a Session Description Protocol (SDP) with the first user equipment. The SDP negotiation is used to determine a first voice codec and a second voice codec. Thus, because two voice codecs are negotiated during call establishment, switching between voice codecs can be achieved to adapt to changes in different network environments and improve the quality of voice calls. Attached Figure Description

[0034] Figure 1 This is a block diagram of a wireless communication system applicable to embodiments of this application;

[0035] Figure 2 This is a diagram of a traditional call setup process;

[0036] Figure 3 This is a schematic diagram of AGW transmitting data;

[0037] Figure 4 This is a flowchart illustrating a call processing method provided in an embodiment of this application;

[0038] Figure 5A This is one of the example diagrams of a call scenario in which a call processing method provided in this application embodiment is applied;

[0039] Figure 5B This is the second example diagram of a call scenario in which a call processing method provided in this application embodiment is applied;

[0040] Figures 6 to 14 This is a flowchart illustrating a call processing method provided in an embodiment of this application;

[0041] Figures 15 to 17 This is a schematic diagram of the structure of a call processing device provided in an embodiment of this application;

[0042] Figure 18 This is a schematic diagram of the structure of a network-side device provided in an embodiment of this application;

[0043] Figure 19 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application;

[0044] Figure 20 This is a schematic diagram of the structure of a communication device provided in an embodiment of this application. Detailed Implementation

[0045] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0046] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as the sender explicitly informing the receiver of specific information, the required operation, or the requested result in the instruction sent. An indirect instruction can be understood as the receiver determining the corresponding information based on the instruction sent by the sender, or making a judgment and determining the required operation or requested result based on the judgment result.

[0047] It is worth noting that the technologies described in this application are not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA), or other systems. The terms "system" and "network" in this application are often used interchangeably, and the described technologies can be used with the systems and radio technologies mentioned above, as well as with other systems and radio technologies. The following description describes New Radio (NR) systems for illustrative purposes, and the term NR is used in most of the following description; however, these technologies can also be applied to systems other than NR systems, such as 6th generation (6G) radio systems. th Generation 6G communication system.

[0048] Figure 1This diagram illustrates a block diagram of a wireless communication system applicable to embodiments of this application. The wireless communication system includes a terminal 11 and a network-side device 12. The terminal 11 can also be referred to as User Equipment (UE), and can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home devices (home appliances with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game consoles, personal computers (PCs), ATMs, or self-service machines, etc. Wearable devices include: smartwatches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. Furthermore, terminal 11 can be any of the terminals described above, or it can be a chip within a terminal, such as a modem chip, a system-on-chip (SoC), etc. It should be noted that the specific type of terminal 11 is not limited in this application embodiment. Network-side equipment 12 can include access network equipment or core network equipment, wherein access network equipment can also be referred to as Radio Access Network (RAN) equipment, radio access network function, or radio access network unit. Access network equipment can include base stations, Wireless Local Area Network (WLAN) access points (APs), or Wireless Fidelity (WiFi) nodes, etc.Among them, base stations can be referred to as Node B (NB), Evolved Node B (eNB), Next Generation Node B (gNB), New Radio Node B (NR Node B), Access Point, Relay Base Station (RBS), Serving Base Station (SBS), Base Transceiver Station (BTS), Radio Base Station, Radio Transceiver, Basic Service Set (BSS), Extended Service Set (ESS), Home Node B (HNB), Home Evolved Node B, Transmit / Receive Point (TRP), Non-Terrestrial Network (NTN) equipment (such as satellite or high altitude platform stations). The term "base station" can be any suitable term in the field, such as "station" or any other appropriate term in the relevant field, as long as the same technical effect is achieved. The term "base station" is not limited to specific technical terms. It should be noted that the embodiments of this application only use the base station in the NR system as an example for introduction, and do not limit the specific type of base station.

[0049] Core network equipment, also known as core network nodes, core network functions, or core network elements, includes, but is not limited to, at least one of the following: Mobility Management Entity (MME), Access and Mobility Management Function (AMF), Session Management Function (SMF), User Plane Function (UPF), Policy Control Function (PCF), Policy and Charging Rules Function (PCRF), Edge Application Server Discovery Function (EASDF), Unified Data Management (UDM), Unified Data Repository (UDR), Home Subscriber Server (HSS), Centralized network configuration (CNC), Network Repository Function (NRF), Network Exposure Function (NEF), Local NEF (L-NEF), and Binding Support. Functions include: BSF (Broadcast Function), Application Function (AF), Location Management Function (LMF), Gateway Mobile Location Centre (GMLC), Network Data Analytics Function (NWDAF), Serving Call Session Control Function (S-CSCF), Proxy Call Session Control Function (P-CSCF), Access Gateway (AGW), and Non-Terrestrial Network.NTN (Network Network Technology) equipment (such as satellites or high-altitude platform stations). It should be noted that this application embodiment only uses core network equipment in the NR system as an example and does not limit the specific type of core network equipment. If the name of the core network equipment mentioned in this application embodiment changes in subsequent protocol versions (e.g., 6G), it will still be within the scope of protection of this application.

[0050] Optionally, the core network equipment can be implemented by one or more functional modules in a single device, or by multiple devices working together; this application does not specifically limit this. It is understood that the aforementioned functional modules can be network elements in hardware devices, software functional modules running on dedicated hardware, or virtualized functional modules instantiated on a platform (e.g., a cloud platform).

[0051] For ease of understanding, the following describes some aspects of the embodiments of this application:

[0052] I. Call setup process.

[0053] The process of user equipment 1 initiating the establishment of an IP Multimedia Subsystem (IMS) voice call is as follows: Figure 2 As shown, user equipment 1 connects to the gNB via satellite and further connects to the IMS network element. The voice call establishment process includes:

[0054] Step 201: Establish a Protocol Data Unit (PDU) session, which is used to transmit IMS services.

[0055] Step 202: User equipment 1 initiates an IMS call establishment request by sending an Invite message 1 using Session Initiation Protocol (SIP) and carrying information about the voice media to be established in the requested Session Description Protocol (SDP).

[0056] Examples of SDP are as follows:

[0057] m = audio 0RTP / AVP 97;

[0058] a = rtpmap:97AMR / 8000 / 1.

[0059] The m-line can be called the media line, representing the type of media as voice service;

[0060] Row 'a' can be called the attribute row, representing the attribute information of this voice service.

[0061] In this context, "97" represents the Real-time Transport Protocol (RTP) payload type, "AMR" indicates that the encoding format is Adaptive Multi-Rate Narrowband (AMR-NB), "8000" represents the sampling rate, and "1" represents mono.

[0062] Step 203: The Session Border Controller (SBC) / P-CSCF sends the invite message 1 to the S-CSCF.

[0063] Optionally, Figure 2 The P-CSCF and S-CSCF are both IMS network elements that provide services to user equipment 1, while the IMS network elements that provide services to user equipment 2 are not shown.

[0064] Optionally, the SBC is an important network node in the IMS network, located at the IMS network boundary, playing a crucial role in connecting user equipment to the IMS core network. Its main functions include access permission control, network topology hiding, Network Address Translation (NAT) and NAT traversal, Quality of Service (QoS), and bandwidth policy adjustment. The SBC can include both a control plane and a user plane, or the control plane and user plane can be separated; for example, the SBC control plane (SBC-C) is used for control, and the SBC user plane (SBC-User plane) is used for transmitting media data.

[0065] Optionally, SBC / P-CSCF represents an IMS network element that includes both SBC and P-CSCF functions, such as an SBC that includes P-CSCF functions, or an IMS network element that combines SBC and P-CSCF.

[0066] Step 204: The S-CSCF sends invite message 1 to the application server (AS) that provides services to user equipment 1.

[0067] Step 205: AS sends the processed invite message 1 to S-CSCF, that is, AS processes invite message 1 and sends it back to S-CSCF.

[0068] Step 206: The S-CSCF sends the processed invite message 1 to the user equipment 2.

[0069] Step 207: User equipment 2 replies with a 183 message (i.e., a response message), which contains information about the voice services supported by SDP provided by user equipment 2.

[0070] Optionally, if user equipment 1 provides multiple voice encoding methods in step 2, such as supporting 4.75kbps, 12.2kbps, etc., user equipment 2 will select one of its supported voice encoding methods and carry it through the SDP answer of the 183 message.

[0071] Optionally, after receiving the 183 message, the P-CSCF providing the service for user equipment 2 triggers the establishment of a dedicated QoS flow for transmitting voice services for user equipment 2, which is the same as step 211.

[0072] Step 208: S-CSCF sends Message 183 to AS. AS will process Message 183 and obtain the processed Message 183.

[0073] Step 209: AS sends the processed 183 message to S-CSCF.

[0074] Step 210: The S-CSCF sends the processed 183 message to the SBC / P-CSCF.

[0075] Step 211: The P-CSCF sends a message to the 5G PCF based on the SDP answer to establish a dedicated QoS flow for transmitting voice services.

[0076] Step 212: After the QoS flow is established, the P-CSCF sends a 183 message to User Equipment 1.

[0077] Step 213: When user device 2 answers the call, user device 2 sends a 200 OK message. When the 200 OK message is sent to user device 1, user device 1 and user device 2 begin their call.

[0078] II. AGW transmission.

[0079] A schematic diagram of AGW data transmission is shown below. Figure 3As shown, T1 is the media transmission resource serving the user equipment side, used to receive media data from the user equipment and send media data to the user equipment. T2 is the media transmission resource serving the remote end, used to receive media data from the remote end and send media data to the remote end. The AGW receives the configuration from the network-side device (i.e., the configuration from the P-CSCF) to reserve and configure the resources of T1 and T2, and binds T1 and T2 to realize media interaction between the user equipment and the remote end. That is, the media data sent by the user equipment to T1 will be transferred by the AGW to T2 and then sent by T2 to the remote end, and vice versa.

[0080] Optionally, both TI and T2 correspond to a transport address, which can be understood as an IP address, or a combination of an IP address and a port number.

[0081] It should be noted that high-speed voice provides a better call experience, but it requires strong signal strength and is prone to dropped calls at cell edges; low-speed voice can maintain call continuity at cell edges, but the call experience is poor. Since only one voice codec can be used per call, the choice between high-speed and low-speed voice is a problem that needs to be solved. This application proposes a call processing method to address this issue.

[0082] Optionally, low bitrate speech refers to speech coding methods with a low coding rate. Low bitrate speech typically refers to speech codecs with a coding rate of less than or equal to 2.4kbps or 1.2kbps. Low bitrate speech can also be understood or replaced with low bitrate speech codec, ultra-low bitrate speech, ultra-low bitrate speech codec, narrowband codec speech, or narrowband speech.

[0083] The call processing method provided in this application will be described in detail below with reference to the accompanying drawings, through some embodiments and application scenarios.

[0084] Reference Figure 4 This application provides a call processing method, such as... Figure 4 As shown, the call processing method includes:

[0085] Step 401: During the call establishment process of the first user equipment, the first network entity negotiates the Session Description Protocol (SDP) with the first user equipment. The SDP negotiation is used to determine the first voice code and the second voice code.

[0086] In this embodiment, the first user equipment can be either the calling end or the called end, without further limitation. The first voice coding and the second voice coding are different; for example, their coding rates are different. Thus, since the first and second voice codings are determined during SDP negotiation, either the first or second voice coding can be initiated based on network quality. For example, in some embodiments, the coding rate of the first voice coding is higher than that of the second voice coding. That is, the first voice coding can be understood as conventional voice coding, and the second voice coding can be understood as narrowband voice coding.

[0087] Optionally, in some embodiments, the system can switch from a first voice codec to a second voice codec, or vice versa. The following explanation uses the example of a higher encoding rate for the first voice codec compared to the second voice codec. Switching from the first voice codec to the second voice codec can be understood as switching from a high-rate voice to a low-rate voice. For example, during communication, when network quality degrades, switching from a high-rate voice to a low-rate voice can ensure smooth voice calls and prevent interruptions. Switching from the second voice codec to the first voice codec can be understood as switching from a low-rate voice to a high-rate voice. For example, during communication, when network quality improves, switching from a low-rate voice to a high-rate voice can improve the audio quality of the voice call, thereby enhancing the overall voice call experience.

[0088] It should be noted that the speech encoding in this application can be understood or replaced as speech codec, or speech codec. The same applies thereafter, and will not be repeated hereafter.

[0089] In this embodiment, during the call establishment process of the first user equipment, the first network entity negotiates a Session Description Protocol (SDP) with the first user equipment. The SDP negotiation is used to determine a first voice codec and a second voice codec. Thus, because two voice codecs are negotiated during call establishment, switching between voice codecs can be achieved to adapt to changes in different network environments and improve the quality of voice calls.

[0090] Optionally, in some embodiments, the first network entity negotiates a Session Description Protocol (SDP) with the first user equipment, including at least one of the following:

[0091] The first network entity receives a first request message from the first user equipment, the first request message including a media line corresponding to the first speech code and a media line corresponding to the second speech code;

[0092] The first network entity sends a first message to the first user equipment, the first message including the media line corresponding to the first voice code and the media line corresponding to the second voice code.

[0093] In this embodiment, two voice coding schemes can be negotiated using two media lines. At this point, two RTP connections can be established based on the two media lines, as detailed below. Figure 5A As shown.

[0094] Optionally, the first request message can be an invite request message, meaning the first user equipment can include two media lines when initiating the invite request. For example, in some embodiments, the SDP offer in the invite request is as follows:

[0095] m=audio 49152RTP / AVP 97 98;

[0096] a = rtpmap:97EVS / 16000 / 1;

[0097] a=rtpmap:98AMR-WB / 16000 / 1;

[0098] m = audio 49153RTP / AVP 99;

[0099] a = rtpmap:99CODEC2 / 8000 / 1.

[0100] It should be noted that in the above examples, EVS and AMR can be understood as conventional speech coding, i.e., the first speech coding; CODEC2 can be understood as narrowband speech coding, i.e., the second speech coding. The two media lines can be understood or replaced as: two speech media lines, or media lines corresponding to two speech media, and the same applies thereafter, so it will not be elaborated further.

[0101] Optionally, the first message may be an invite request message sent by the first network entity to the first user equipment, a SIP update message sent by the first network entity to the user equipment, or a 183 message sent by the first network entity to the first user equipment.

[0102] Optionally, two voice coding schemes can be negotiated before establishing the voice bearer for voice transmission at the first user equipment. For example, when the first user equipment is the calling end, it can include two media lines when initiating an invite request. This means that the first network entity and the first user equipment negotiate the Session Description Protocol (SDP), including the first network entity receiving a first request message from the first user equipment. Furthermore, after establishing the voice bearer, the first network entity sends an 183 message to the first user equipment. This means that the first network entity and the first user equipment negotiate the SDP, including the first network entity sending a first message to the first user equipment.

[0103] For example, when the first user equipment is the called party, the first network entity, after receiving an invite request from the second user equipment, can include two voice-coded media lines in its invite request when sending the invite request to the first user equipment. In this case, the Session Description Protocol (SDP) negotiation between the first network entity and the first user equipment includes the first network entity sending a first message to the first user equipment.

[0104] Optionally, after establishing the voice bearer for voice transmission at the first user equipment, two voice coding schemes can be negotiated. For example, after establishing the voice bearer, the first network entity can send an 183 message to the first user equipment. Subsequently, the first network entity can initiate an update request to the user equipment, carrying media lines corresponding to the two voice coding schemes. At this time, the Session Description Protocol (SDP) negotiation between the first network entity and the first user equipment includes the first network entity sending a first message to the first user equipment.

[0105] It should be noted that the voice bearer (e.g., the first voice bearer or the second voice bearer) in the embodiments of this application can be an Evolved Packet System (EPS) bearer, or a QoS flow, or a DRB corresponding to the EPS bearer, or a DRB corresponding to the QoS flow.

[0106] Optionally, in some embodiments, the first request message further includes at least one of a first indication information and a second indication information;

[0107] Wherein, the first indication information is used to indicate at least one of the following:

[0108] The first user equipment supports multiple real-time transmission protocol for voice.

[0109] The first user equipment supports multiple real-time transmission protocol for audio (multiple RTP for audio);

[0110] The first user equipment supports multiple codecs for voice.

[0111] The first user equipment supports multiple codecs for audio.

[0112] The second indication information is used to indicate that the first user equipment supports the second voice coding.

[0113] It should be noted that "multi-channel" can be understood or replaced with: multiple or at least two. Specifically, "the first user equipment supports a multi-channel real-time transmission protocol for voice," which can be understood or replaced with: the first user equipment supports establishing multiple RTP connections for voice transmission. "The first user equipment supports a multi-channel real-time transmission protocol for audio," which can be understood or replaced with: the first user equipment supports establishing multiple RTP connections for audio transmission. "The first user equipment supports multiple codecs for voice," which can be understood or replaced with: the first user equipment supports multiple codecs for voice calls, or, the first user equipment supports negotiating multiple codecs for voice calls. "The first user equipment supports multiple codecs for audio," which can be understood or replaced with: the first user equipment supports multiple codecs for audio calls, or, the first user equipment supports negotiating multiple codecs for audio calls.

[0114] Optionally, support for a second voice codec can be indicated by carrying a new media feature tag.

[0115] Optionally, in some embodiments, the first and second indication information described above can be carried by the media feature tag of the contact header of the invite.

[0116] Optionally, in some embodiments, the method further includes:

[0117] The first network entity deletes the media line corresponding to the second speech code in the first request message to obtain the second request message;

[0118] The first network entity sends the second request message to the second user equipment.

[0119] In this embodiment, the aforementioned second request message can also be understood as an invite request message. Since the second user equipment may not support both voice encodings, or to reduce protocol changes, upon receiving the first request message, the media line corresponding to the second voice encoding in the first request message can be deleted, i.e., the media line corresponding to the second voice encoding in the SDP offer can be deleted.

[0120] Optionally, in some embodiments, the method further includes:

[0121] The first network entity receives a response message to the second request message from the second user equipment;

[0122] The first network entity adds a media line corresponding to the second voice code to the response message of the second request message, and sends the modified response message of the second request message to the first user equipment.

[0123] In this embodiment, the response message to the second request message may include an SDP answer. The response message to the first user equipment (UAE) sending the modified second request message can be understood as sending a 183 message to the UAE, which can be interpreted as a response message to the first request message. For example, the first network entity can add the media line corresponding to the second voice code to the SDP answer. After establishing a voice bearer for transmitting voice services, the first network entity can send a 183 message to the first UAE.

[0124] Optionally, in some embodiments, the method further includes:

[0125] The first network entity sends a third request message to the second network entity, the third request message being used to request the allocation of a first transmission address corresponding to the first voice code;

[0126] The media line corresponding to the second voice code in the response message of the modified second request message includes the first transmission address corresponding to the second voice code; the first transmission address is used for communication between the first user equipment and the second network entity.

[0127] In this embodiment of the application, the aforementioned third request message can also be used to request the allocation of a first transmission address corresponding to the second voice code, or the first network entity can request the second network entity to allocate a first transmission address corresponding to the second voice code through an additional request message.

[0128] Optionally, the second network entity mentioned above can be a P-CSCF.

[0129] Optionally, the aforementioned first transmission address can be understood as the T1 transmission address.

[0130] Optionally, in some embodiments, the first message is a first session initiation protocol (SIP) invitation message or a first SIP update message.

[0131] Optionally, in some embodiments, the method further includes:

[0132] The first network entity receives a response message to the first message from the first user equipment. The response message to the first message includes a media line corresponding to the first speech code and a media line corresponding to the second speech code.

[0133] In this embodiment, when the first user equipment is the calling end, the first message is an Update request message, and the response message can be a 200 OK message corresponding to the Update request message. After the second user equipment answers the call, the second user equipment replies with a 200 OK message, at which point the first user equipment and the second user equipment begin their conversation.

[0134] Optionally, in some embodiments, the first message further includes third indication information, which is used to indicate that the network supports a second speech code.

[0135] In this embodiment of the application, network support for the second speech code can be understood as AGW support for the second speech code.

[0136] Optionally, in some embodiments, the method further includes:

[0137] The first network entity sends a second instruction message to the second network entity, the second instruction message being used to instruct the second network entity to perform encoding conversion.

[0138] In this embodiment, the encoding conversion performed by the second network entity can be understood as the second network entity performing encoding conversion on the RTP data corresponding to the second voice code. For example, after receiving the RTP data corresponding to the second voice code, the second network entity can use the second voice code to decode the received data, and then use the voice data corresponding to the first voice code to encode the decoded data to obtain the encoded data. Finally, the encoded data is sent to the second user equipment.

[0139] Optionally, in some embodiments, the first network entity negotiates a Session Description Protocol (SDP) with the first user equipment, including at least one of the following:

[0140] The first network entity receives a fourth request message from the first user equipment. The fourth request message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0141] The first network entity sends a second message to the first user equipment. The second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0142] In this embodiment, two voice codes can be negotiated using one media line. At this point, one RTP connection can be established based on one media line, specifically as follows: Figure 5B As shown.

[0143] Optionally, the fourth request message mentioned above can be an invite request message, meaning the first user equipment can include one media line when initiating the invite request. For example, in some embodiments, the SDP offer in the invite request is as follows:

[0144] m=audio 49152RTP / AVP 97 98 99;

[0145] a = rtpmap:97EVS / 16000 / 1;

[0146] a=rtpmap:98AMR-WB / 16000 / 1;

[0147] a = rtpmap:99CODEC2 / 8000 / 1;

[0148] a=additional codec indication.

[0149] Optionally, the second message may be an invite request message sent by the first network entity to the first user equipment, a SIP update message sent by the first network entity to the user equipment, or a 183 message sent by the first network entity to the first user equipment.

[0150] Optionally, two voice codes can be negotiated before establishing the voice bearer for voice transmission at the first user equipment. For example, when the first user equipment is the calling end, it can carry attribute lines of two voice codes when initiating an invite request. That is, the first network entity and the first user equipment negotiate the Session Description Protocol (SDP), which includes the first network entity receiving a fourth request message from the first user equipment. Furthermore, after establishing the voice bearer for transmission, the first network entity sends an 183 message to the first user equipment. That is, the first network entity and the first user equipment negotiate the SDP, which also includes the first network entity sending a second message to the first user equipment.

[0151] For example, when the first user equipment is the called party, the first network entity, after receiving an invite request from the second user equipment, can include two voice-coded attribute lines in its invite request when sending the invite request to the first user equipment. In this case, the Session Description Protocol (SDP) negotiation between the first network entity and the first user equipment includes the first network entity sending a second message to the first user equipment.

[0152] Optionally, after establishing a voice bearer for voice transmission at the first user equipment, two voice coding schemes can be negotiated. For example, after establishing the voice bearer, the first network entity can send a 183 message to the first user equipment. Then, the first network entity can initiate an update request to the user equipment, carrying attribute lines for the two voice coding schemes. At this time, the Session Description Protocol (SDP) negotiation between the first network entity and the first user equipment includes the first network entity sending a second message to the first user equipment.

[0153] Optionally, in some embodiments, the fourth request message further includes at least one of the following:

[0154] The fifth instruction information is used to indicate the negotiation of two voice coding types;

[0155] The first indication information indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for voice; the first user equipment supports RTP for audio; the first user equipment supports multiplexing for voice; and the first user equipment supports multiplexing for audio.

[0156] The sixth indication information is used to indicate that the second speech encoding is an additional encoding;

[0157] Alternatively, the second message may include at least one of the following:

[0158] The third indication information is used to indicate that the network supports the second speech coding.

[0159] The seventh instruction information is used to indicate the negotiation of two voice coding types;

[0160] The eighth indication information is used to indicate that the second speech code is an additional code.

[0161] In this embodiment of the application, the second speech code can be indicated as an additional code or extra code by an additional codec indication, that is, the sixth and eighth indication information mentioned above are additional codec indications.

[0162] Optionally, the fifth indication information used to indicate the negotiation of two speech coding types can be understood as indicating that two speech coding types need to be negotiated. The two speech coding types may include high-rate speech coding (or conventional speech coding) and low-rate speech coding (or narrowband speech coding). For example, the high-rate speech coding may be the first speech coding mentioned above, and the low-rate speech coding may be the second speech coding mentioned above.

[0163] It should be noted that the additional encoding indication can be understood or replaced as: low bitrate speech indication, low bitrate speech codec indication, ultra-low bitrate speech indication, ultra-low bitrate speech codec indication, narrowband codec speech indication, or narrowband speech indication.

[0164] It should be noted that the additional encoding indicator is used to indicate that the speech encoding of the previous attribute row (i.e., row a) is additional encoding, extra encoding, low bitrate speech, low bitrate speech codec, ultra-low bitrate speech, ultra-low bitrate speech codec, narrowband codec speech, or narrowband speech.

[0165] Optionally, the aforementioned fifth, first, or sixth instruction information can be carried via the media feature tag in the contact header of the invite.

[0166] Optionally, in some embodiments, the method further includes:

[0167] The first network entity deletes the attribute row corresponding to the second voice code in the fourth request message to obtain the fifth request message;

[0168] The first network entity sends the fifth request message to the second user equipment.

[0169] In this embodiment of the application, the aforementioned fifth request message can also be understood as an invite request message. Since the second user equipment may not support both voice codecs, or to reduce protocol changes, after receiving the fourth request message, the attribute row corresponding to the second voice codec in the fourth request message can be deleted, that is, the attribute row corresponding to the second voice codec in the SDP offer can be deleted. For example:

[0170] m=audio 49152RTP / AVP 97 98 99

[0171] a = rtpmap:97EVS / 16000 / 1;

[0172] a=rtpmap:98AMR-WB / 16000 / 1;

[0173]

[0174] Optionally, in some embodiments, the method further includes:

[0175] The first network entity receives a response message to the fifth request message from the second user equipment;

[0176] The first network entity adds an attribute line corresponding to the second voice code to the response message of the fifth request message, and sends the modified response message of the fifth request message to the first user equipment.

[0177] In this embodiment, the response message to the fifth request message may include an SDP answer. The response message to the first user equipment (UAE) sending the modified fifth request message can be understood as sending a 183 message to the UAE, which can be understood as a response message to the fourth request message. For example, the first network entity can add the media line corresponding to the second voice code in the SDP answer. After establishing a voice bearer for transmitting voice services, the first network entity can send a 183 message to the first UAE.

[0178] Optionally, in some embodiments, the response message of the modified fifth request message further includes at least one of the following:

[0179] The seventh instruction information is used to indicate the negotiation of two voice coding types;

[0180] The eighth indication information is used to indicate that the second speech code is an additional code.

[0181] Optionally, in some embodiments, the second message is a second SIP invitation message or a second SIP update message.

[0182] Optionally, in some embodiments, the method further includes:

[0183] The first network entity receives a response message to the second message from the first user equipment. The response message to the second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0184] In this embodiment, when the first user equipment is the calling end, the second message is an Update request message, and the response message can be a 200 OK message corresponding to the Update request message. After the second user equipment answers the call, it replies with a 200 OK message, at which point the first and second user equipment begin their conversation.

[0185] Optionally, in some embodiments, the method further includes:

[0186] The first network entity receives first information from the first user equipment. The first information is information sent during the registration process. The first information includes first indication information, which indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for voice; the first user equipment supports RTP for audio; the first user equipment supports multiplexed codecs for voice; and the first user equipment supports multiplexed codecs for audio.

[0187] The first network entity sends a ninth indication message to the first user equipment, the ninth indication message indicating at least one of the following: the first network entity supports Real-Time Transport Protocol (RTP) for voice; the first network entity supports RTP for audio; the first network entity supports multiple codecs for voice; the first network entity supports multiple codecs for audio.

[0188] In this embodiment, the aforementioned first indication information can be carried through the media feature tag of the registered contact header. After receiving the first indication information and determining that multiple RTP or multiple codecs are supported, the first network entity can send the ninth indication information to the first user equipment. This allows the first user equipment to determine whether the second voice coding can be initiated subsequently, ensuring communication reliability.

[0189] Optionally, in some embodiments, the method further includes:

[0190] If the second voice coding is initiated, the first network entity sends a first notification message to the second network entity, which is used to notify the second network entity to initiate the second voice coding.

[0191] In this embodiment of the application, after receiving the first notification message, the second network entity can initiate the second voice coding, that is, use the second voice coding to communicate with the first user equipment.

[0192] It should be understood that after the second voice coding is started, if a notification is received again from the first network entity to switch voice coding, the user can switch from the second voice coding to the first voice coding.

[0193] Optionally, in some embodiments, the method further includes:

[0194] The first network entity receives a second notification message from the third network entity. The second notification message is used to notify the requirement information that meets the second voice code.

[0195] The first network entity determines to initiate the second voice coding based on the second notification message.

[0196] In this embodiment, the aforementioned third network entity can be a PCF. The aforementioned requirement information can be QoS parameters or a QoS configuration file.

[0197] Reference Figure 6 This application also provides a voice call processing method, such as... Figure 6 As shown, the voice call processing method includes:

[0198] Step 601, the second network entity performs a first operation, the first operation including at least one of the following:

[0199] The first data received from the second user equipment is converted into second data, and the second data is sent to the first user equipment. The first data is the data corresponding to the first voice code, and the second data is the data corresponding to the second voice code.

[0200] The third data received from the first user equipment is converted into fourth data and sent to the second user equipment. The third data is the data corresponding to the second voice code, and the fourth data is the data corresponding to the first voice code.

[0201] In this embodiment of the application, when the second network entity starts the first voice encoding, the received data can be forwarded directly. When the second network entity starts the second voice encoding, the data of the second voice encoding needs to be converted. This can ensure that the voice data of the second user equipment does not change.

[0202] In this embodiment of the application, the second network entity can perform voice encoding conversion to adapt to changes in different network environments, thereby improving the quality of voice calls.

[0203] Optionally, the method further includes:

[0204] The second network entity receives a first notification message from the first network entity, the first notification message being used to notify the second network entity to start the second voice coding.

[0205] Optionally, sending the second data to the first user equipment includes any of the following:

[0206] The second data is sent to the first user equipment based on the Real-Time Transport Protocol (RTP) corresponding to the second voice code.

[0207] The second data is sent to the first user equipment based on the RTP corresponding to the first voice code.

[0208] Reference Figure 7 This application also provides a voice call processing method, such as... Figure 7 As shown, the voice call processing method includes:

[0209] Step 701: During the call establishment process of the first user equipment, the first user equipment negotiates Session Description Protocol (SDP) with the first network entity. The SDP negotiation is used to determine the first voice codec and the second voice codec.

[0210] Optionally, the first user equipment negotiates a Session Description Protocol (SDP) with the first network entity, including at least one of the following:

[0211] The first user equipment sends a first request message to the first network entity, the first request message including the media line corresponding to the first speech code and the media line corresponding to the second speech code;

[0212] The first user equipment receives a first message from the first network entity, the first message including a media line corresponding to the first voice code and a media line corresponding to the second voice code.

[0213] Optionally, the first request message may further include at least one of a first indication information and a second indication information;

[0214] Wherein, the first indication information is used to indicate at least one of the following:

[0215] The first user equipment supports Real-Time Transport Protocol (RTP) for voice.

[0216] The first user equipment supports multiple RTP for audio;

[0217] The first user equipment supports multiplexing for voice;

[0218] The first user equipment supports multiplexing for audio;

[0219] The second indication information is used to indicate that the first user equipment supports the second voice coding.

[0220] Optionally, the first message is a first session initiation protocol (SIP) invitation message or a first SIP update message.

[0221] Optionally, the method further includes:

[0222] The first user equipment sends a response message to the first network entity for the first message, the response message for the first message including the media line corresponding to the first voice code and the media line corresponding to the second voice code.

[0223] Optionally, the first message may further include third indication information, which indicates that the network supports a second speech codec.

[0224] Optionally, the first user equipment negotiates a Session Description Protocol (SDP) with the first network entity, including at least one of the following:

[0225] The first user equipment sends a fourth request message to the first network entity. The fourth request message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0226] The first user equipment receives a second message from the first network entity. The second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0227] Optionally, the fourth request message further includes at least one of the following:

[0228] The fifth instruction information is used to indicate the negotiation of two voice coding types;

[0229] The first indication information indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for voice; the first user equipment supports RTP for audio; the first user equipment supports multiplexing for voice; and the first user equipment supports multiplexing for audio.

[0230] The sixth indication information is used to indicate that the second speech encoding is an additional encoding;

[0231] Alternatively, the second message may include at least one of the following:

[0232] The third indication information is used to indicate that the network supports the second speech coding.

[0233] The seventh instruction information is used to indicate the negotiation of two voice coding types;

[0234] The eighth indication information is used to indicate that the second speech code is an additional code.

[0235] Optionally, the second message is a second SIP invitation message or a second SIP update message.

[0236] Optionally, the method further includes:

[0237] The first user equipment sends a response message to the first network entity for the second message. The response message for the second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0238] Optionally, the method further includes:

[0239] The first user equipment sends first information to the first network entity. The first information is information sent during the registration process. The first information includes first indication information, which indicates at least one of the following: the first user equipment supports Real-time Transport Protocol for Voice (RTP); the first user equipment supports RTP for Audio; the first user equipment supports multiplexed codecs for Voice; and the first user equipment supports multiplexed codecs for Audio.

[0240] The first user equipment receives a ninth indication information from the first network entity, the ninth indication information being used to indicate at least one of the following: the first network entity supports Real-Time Transport Protocol (RTP) for voice; the first network entity supports RTP for audio; the first network entity supports multiple codecs for voice; the first network entity supports multiple codecs for audio.

[0241] Optionally, the method further includes:

[0242] When the first user equipment receives the second data, the first user equipment initiates the second voice encoding, wherein the second data corresponds to the second voice encoding.

[0243] Optionally, the method further includes at least one of the following:

[0244] If the first user equipment receives data based on the Real-time Transport Protocol (RTP) corresponding to the second voice code, it is determined that the second data has been received;

[0245] When the first user equipment receives data based on the RTP corresponding to the first voice code, it determines that the second data has been received based on the packet header of the received data.

[0246] Optionally, the method further includes:

[0247] The first user equipment sends third data based on the second voice encoding.

[0248] Optionally, the method further includes:

[0249] The first user equipment determines to initiate the second voice coding based on the obtained transmission bit rate.

[0250] Optionally, the method further includes:

[0251] The first user equipment receives a tenth indication information from a fourth network entity, the tenth indication information being used to indicate the transmission bit rate.

[0252] In this embodiment of the application, the fourth network entity can be a base station.

[0253] Optionally, in order to better understand this application, the following embodiments use 5G as an example for illustration. The solution of this application can also be extended to 4G, 6G and 7G, etc. For example, when extended to 4G, EPS bearer can be used instead of QoS flow, and Public Data Network (PDN) connection can be used instead of PDU session.

[0254] Example 1: Two voice codes are negotiated using two media lines. Assume the first user equipment is the calling end, such as... Figure 8 As shown, the communication process includes:

[0255] Step 801: The first user equipment completes IMS registration. During registration, it carries first indication information, which indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) multiplexing for voice; the first user equipment supports RTP multiplexing for audio; the first user equipment supports multiple codecs for voice; the first user equipment supports multiple codecs for audio.

[0256] Optionally, the network sends a network support indication, i.e., the ninth indication information, to the first user equipment via the media feature tag in the contact header of the registered device.

[0257] Step 802, the first user equipment sends a first request message (i.e., an invite request), the first request message including the media line corresponding to the first voice code and the media line corresponding to the second voice code; optionally, the first request message may also carry at least one of the first instruction information and the second instruction information.

[0258] Step 803: P-CSCF requests AGW to allocate a T2 transport address.

[0259] Step 804: P-CSCF deletes the media line corresponding to the second speech codec (i.e., codec-2) in the SDP offer of the first request message.

[0260] Step 805: P-CSCF sends a second request message (i.e., invite request) to the second user equipment, which contains only the media line corresponding to the first speech codec (i.e., codec-1) in SDPoffer.

[0261] Step 806: The second user equipment replies to the response message of the second request message, carrying the SDP answer.

[0262] Step 807: P-CSCF requests AGW to allocate a T1 transmission address corresponding to the first voice codec (e.g., a normal voice codec).

[0263] Step 808: P-CSCF requests AGW to allocate a T1 transmission address (i.e., the first transmission address) corresponding to the second speech codec (e.g., low-rate speech codec).

[0264] Step 809: Add the media line corresponding to the second speech codec (i.e. codec-2) to the SDP answer, and include the IP assigned in steps 807 and 808 in each media line.

[0265] Step 810: Establish a dedicated QoS flow for transmitting voice services.

[0266] In step 811, the P-CSCF replies to the first user equipment with message 183 (i.e., the response message to the first message or the modified second request message), carrying an SDP answer containing media lines corresponding to the two speech codes, and optionally, carrying an indication that supports the new media feature tag.

[0267] Step 812: The second user device answers the call and replies with a 200 OK message, and the first and second user devices begin a call.

[0268] Example 2: Two voice codes are negotiated using two media lines. Assume the first user equipment is the calling end, such as... Figure 9 As shown, the communication process includes:

[0269] Step 901: The first user equipment completes IMS registration. During registration, it carries first indication information, which indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for voice; the first user equipment supports RTP for audio; the first user equipment supports multiplexing / decoding for voice; the first user equipment supports multiplexing / decoding for audio.

[0270] Optionally, the network sends a network support indication, i.e., the ninth indication information, to the first user equipment via the media feature tag in the contact header of the registered device.

[0271] Step 902, the first user equipment sends an invite request message, the invite request carrying at least one of the first instruction information and the second instruction information.

[0272] Step 903: P-CSCF requests AGW to allocate a T2 transport address.

[0273] Step 904: P-CSCF sends an invite request message to the second user equipment.

[0274] Step 905: The second user equipment replies with a response message, carrying the SDP answer.

[0275] Step 906: P-CSCF requests AGW to allocate a T1 transmission address corresponding to the first voice codec (e.g., a normal voice codec).

[0276] Step 907: Establish a dedicated QoS flow for transmitting voice services.

[0277] In step 908, the P-CSCF sends an 183 message to the first user equipment, optionally carrying an indication that it supports the new media feature tag.

[0278] Step 909: P-CSCF requests AGW to allocate a T1 transmission address (i.e., the first transmission address) corresponding to the second speech codec (e.g., low-rate speech codec).

[0279] Optionally, P-CSCF instructs AGW to perform codec transcoding for RTP-2 data.

[0280] Step 910: P-CSCF sends an Update request (i.e., the first message) to the first user equipment. The Update request contains media lines corresponding to the two codecs.

[0281] Step 911: The first user equipment replies with a 200 OK message corresponding to the Update request, which includes media lines corresponding to the two codecs.

[0282] Step 912: The second user device answers the call and replies with a 200 OK message, and the first and second user devices begin their conversation.

[0283] The difference between this embodiment and Embodiment 1 is that Embodiment 1 modified the Invite message sent by the first user equipment to include media lines corresponding to two different codecs. Embodiment 2, following existing technology, includes only one media line corresponding to a single codec in the Invite message for the first user equipment. After determining that the first user equipment and the second user equipment have completed codec negotiation, the P-CSCF adds a media line corresponding to the second voice codec for the first user equipment. Compared to Embodiment 1, this embodiment reduces the modification of signaling to the first user equipment.

[0284] Example 3: Two voice codes are negotiated using two media lines. Assume the first user equipment is the called party, such as... Figure 10 As shown, the communication process includes:

[0285] Step 1001: The first user equipment completes IMS registration. During registration, it carries first indication information, which indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for voice; the first user equipment supports RTP for audio; the first user equipment supports multiplexing / decoding for voice; the first user equipment supports multiplexing / decoding for audio.

[0286] Optionally, the network sends a network support indication, i.e., the ninth indication information, to the first user equipment via the media feature tag in the contact header of the registered device.

[0287] Step 1002: The second user equipment sends an invite request message, which includes a media line corresponding to one audio (i.e., the media line corresponding to the first speech code).

[0288] Step 1003: P-CSCF requests AGW to allocate the T1 transmission address corresponding to the first voice code.

[0289] Step 1004: P-CSCF requests AGW to allocate the T1 transmission address (i.e., the first transmission address) corresponding to the second voice code.

[0290] Optionally, P-CSCF instructs AGW to perform codec transcoding on the RTP-2 data corresponding to the second speech code.

[0291] Step 1005: Add the media line corresponding to the second speech codec (low bitrate speech codec) to the SDP, and include the IP address allocated in steps 1003 and 1004 in each media line. For example:

[0292] m=audio 49152RTP / AVP 97 98;

[0293] a = rtpmap:97EVS / 16000 / 1;

[0294] a=rtpmap:98AMR-WB / 16000 / 1;

[0295] m = audio 49153RTP / AVP 99;

[0296] a = rtpmap:99CODEC2 / 8000 / 1.

[0297] Step 1006: P-CSCF sends a first message (i.e., invite request message) to the first user equipment.

[0298] Step 1007: The first user equipment replies to the P-CSCF with a 183 message (i.e., the response message of the first message), carrying an SDP answer containing two audio media lines (i.e., the media line of the first speech encoding and the media line of the second speech encoding), and optionally, carrying an indication that supports the new media feature tag.

[0299] Step 1008: Establish a dedicated QoS flow for transmitting voice services.

[0300] Step 1009: P-CSCF deletes the media line corresponding to the second speech code in the SDP answer.

[0301] Step 1010: P-CSCF requests AGW to allocate the T2 transmission address corresponding to the first voice code.

[0302] Step 1011: The P-CSCF replies with a 183 message to the second user equipment, carrying the SDPanswer corresponding to the first voice code.

[0303] Step 1012: The first user device answers the call and replies with a 200 OK message. The first user device and the second user device then begin a call.

[0304] Example 4: Two voice codes are negotiated using one media line. Assume the first user equipment is the calling end, such as... Figure 11 As shown, the communication process includes:

[0305] Step 1101: The first user equipment completes IMS registration. During registration, it carries first indication information, which indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for voice; the first user equipment supports RTP for audio; the first user equipment supports multiplexing / decoding for voice; the first user equipment supports multiplexing / decoding for audio.

[0306] Optionally, the network sends a network support indication, i.e., the ninth indication information, to the first user equipment via the media feature tag in the contact header of the registered device.

[0307] Step 1102: The first user equipment sends a fourth request message (i.e., an invite request). The media line in the fourth request message includes the attribute line corresponding to the first speech codec and the attribute line corresponding to the second speech codec; for example:

[0308] m=audio 49152RTP / AVP 97 98 99;

[0309] a = rtpmap:97EVS / 16000 / 1;

[0310] a=rtpmap:98AMR-WB / 16000 / 1;

[0311] a = rtpmap:99CODEC2 / 8000 / 1;

[0312] a=additional codec indication.

[0313] Optionally, the first request message may also carry at least one of a first instruction message and a second instruction message.

[0314] Step 1103: P-CSCF requests AGW to allocate a T2 transport address.

[0315] Step 1104: The P-CSCF deletes the attribute line corresponding to the second speech codec (i.e., codec-2) in the SDP offer of the fourth request message. For example:

[0316] m=audio 49152RTP / AVP 97 98 99;

[0317] a = rtpmap:97EVS / 16000 / 1;

[0318] a=rtpmap:98AMR-WB / 16000 / 1;

[0319]

[0320] Step 1105: P-CSCF sends a fifth request message (i.e., invite request) to the second user equipment, which contains only the attribute row corresponding to the first speech codec (i.e., codec-1) in SDPoffer.

[0321] Step 1106: The second user equipment replies to the response message of the second request message, carrying the SDP answer.

[0322] Step 1107: P-CSCF requests AGW to allocate a T1 transmission address corresponding to the first voice codec (e.g., a normal voice codec).

[0323] Step 1108: Add the attribute line corresponding to the second speech codec (i.e., codec-2) to the SDP answer.

[0324] Step 1109: Establish a dedicated QoS flow for transmitting voice services.

[0325] In step 1110, the P-CSCF replies to the first user equipment with message 183 (i.e., the response message to the second message or the modified fifth request message), carrying an SDP answer containing attribute lines corresponding to the two speech codes, and optionally, carrying an indication that supports the new media feature tag.

[0326] Step 1111: The second user device answers the call and replies with a 200 OK message. The first and second user devices then begin a call.

[0327] Example 5: Two voice codes are negotiated using one media line. Assuming the first user equipment is the calling end, the communication process is similar to Example 2, except that the media line in the SDPoffer carried by the Update request message in step 910 includes the attribute lines corresponding to the first and second voice codes. In step 911, the first user equipment replies with a 200 OK message corresponding to the Update request, containing the attribute lines corresponding to the two codecs.

[0328] Example 6: Two voice codes are negotiated using one media line. Assume the first user equipment is the called party, such as... Figure 12 As shown, the communication process includes:

[0329] Step 1201: The first user equipment completes IMS registration. During registration, it carries first indication information, which indicates at least one of the following: the first user equipment supports Real-time Transport Protocol (RTP) multiplexing for voice; the first user equipment supports RTP multiplexing for audio; the first user equipment supports multiple codecs for voice; the first user equipment supports multiple codecs for audio.

[0330] Optionally, the network sends a network support indication, i.e., the ninth indication information, to the first user equipment via the media feature tag in the contact header of the registered device.

[0331] Step 1202: The second user equipment sends an invite request message, which includes a media line corresponding to one audio (i.e., the media line corresponding to the first speech code).

[0332] Step 1203: P-CSCF requests AGW to allocate the T1 transmission address corresponding to the first voice code.

[0333] Step 1204: Add the attribute row corresponding to the second speech codec (low bit rate speech codec) in SDP.

[0334] Step 1205: The P-CSCF initiates a second message (i.e., an invite request message) to the first user equipment. The SDP offer in the second message includes attribute lines corresponding to the first and second voice codes. Optionally, the SDP may carry indication information of the codec from negotiation 2.

[0335] For example:

[0336] m=audio 49152RTP / AVP 97 98 99;

[0337] a = rtpmap:97EVS / 16000 / 1;

[0338] a=rtpmap:98AMR-WB / 16000 / 1;

[0339] a = rtpmap:99CODEC2 / 8000 / 1;

[0340] a=additional codec indication.

[0341] Step 1206: The first user equipment replies to the P-CSCF with message 183 (i.e., the response message of the second message), carrying an SDP answer containing attribute lines of two codecs (i.e., attribute lines of the first speech codec and attribute lines of the second speech codec), and optionally, carrying an indication that supports the new media feature tag.

[0342] Step 1207: Establish a dedicated QoS flow for transmitting voice services.

[0343] Step 1208: P-CSCF deletes the attribute row corresponding to the second speech code in the SDP answer.

[0344] Step 1209: P-CSCF requests AGW to allocate the T2 transmission address corresponding to the first voice code.

[0345] Step 1210: The P-CSCF replies with a 183 message to the second user equipment, carrying the SDPanswer corresponding to the first voice code.

[0346] Step 1211: The first user device answers the incoming call and replies with a 200 OK message. The first user device and the second user device then begin a call.

[0347] Optionally, in some embodiments, changes to the codec can be triggered by the network side. For example... Figure 13 As shown, the process may include the following:

[0348] Step 1301: The first user equipment communicates with the second user equipment using RTP-1 (the RTP corresponding to the second voice code), and the AGW performs transparent forwarding.

[0349] Step 1302, P-CSCF determines to start the second speech coding (codec-2).

[0350] For example, if the P-CSCF receives a notification from the PCF, it may be able to meet the QoS parameters corresponding to the second speech codec or may not be able to meet the QoS parameters corresponding to the first speech codec (e.g., a normal speech codec).

[0351] Step 1303, P-CSCF, notifies AGW to start codec conversion.

[0352] Step 1304: The second user equipment sends the first data to the AGW. The first data is the data corresponding to the first voice code.

[0353] Step 1305: AGW converts the first data into the second data, which is the data corresponding to the second voice code.

[0354] Step 1306: AGW sends the second data to the user equipment.

[0355] Optionally, for embodiments one to three above, the first data can be encapsulated into a quintuple corresponding to RTP-2 to obtain the second data, and the second data can be sent through RTP-2 (i.e., the RTP corresponding to the second voice code).

[0356] Optionally, for embodiments four to six above, the first data can be encapsulated into a 5-tuple corresponding to RTP-1 (i.e., the RTP corresponding to the first voice code) to obtain the second data, and the second data can be sent via RTP-1. The packet headers (5-tuples) of the RTP-1 and RTP-2 data packets are different.

[0357] Step 1307: When the first user equipment receives data via RTP-2 or RTP-1, it starts the encoder corresponding to the second voice codec (codec-2).

[0358] Optionally, the first user equipment determines the encoder corresponding to codec-2 to be started using the following method:

[0359] Method 1: The first user equipment determines the start codec-2 based on the data received via different RTPs;

[0360] Method 2: The first user equipment determines the startup codec-2 based on the header of the received RTP data packet.

[0361] Method 1 is only applicable to the case of establishing two RTP connections in Examples 1 to 3.

[0362] Step 1308: The first user equipment sends third data, which is the data corresponding to the second voice code.

[0363] Step 1309: AGW converts the third data into the fourth data;

[0364] Step 1310: AGW sends the fourth data to the second user equipment.

[0365] Optionally, in some embodiments, changes to the codec can be triggered by the user equipment. For example... Figure 14 As shown, the process may include the following:

[0366] Step 1401: The first user equipment communicates with the second user equipment using RTP-1 (the RTP corresponding to the second voice code), and the AGW performs transparent forwarding.

[0367] Step 1402: The base station determines that the current transmission bit rate cannot be met and determines a satisfactory transmission bit rate.

[0368] In step 1403, optionally, the base station sends a recommended transmission bit rate (i.e., a satisfactory bit rate) to the first user equipment via the Media Access Control (MAC) control element (CE).

[0369] Step 1404: The first user equipment determines the voice coding to switch based on the MAC CE or local statistical results, such as switching from the second voice coding to the first voice coding.

[0370] Step 1405: The first user equipment sends third data to the AGW, which is the data corresponding to the second voice code.

[0371] Optionally, for embodiments one to three above, the first user equipment can use the five-tuple encapsulation of voice packets corresponding to RTP-2 to obtain the third data.

[0372] Optionally, for embodiments four to six above, the first user equipment can use the five-tuple encapsulation of voice packets corresponding to RTP-1 to obtain the third data.

[0373] Step 1406: AGW converts the third data into the fourth data, which is the data corresponding to the first voice code.

[0374] Step 1407: AGW sends the fourth data to the second user equipment.

[0375] Step 1408: The second user equipment sends the first data to the AGW. The first data is the data corresponding to the first voice code.

[0376] Step 1309: AGW converts the first data into the second data, which is the data corresponding to the second voice code.

[0377] Step 1310: AGW sends the second data to the user equipment.

[0378] Optionally, for embodiments one to three above, the first data can be encapsulated into a quintuple corresponding to RTP-2 to obtain the second data, and the second data can be sent through RTP-2 (i.e., the RTP corresponding to the second voice code).

[0379] Optionally, for embodiments four to six above, the first data can be encapsulated into a 5-tuple corresponding to RTP-1 (i.e., the RTP corresponding to the first voice code) to obtain the second data, and the second data can be sent through RTP-1.

[0380] It should be noted that the first user equipment can switch from the second voice codec to the first voice codec, or from the first voice codec to the second voice codec. The switching method is similar and will not be described in detail here.

[0381] It should be understood that in some embodiments, IMS AS can be used instead of P-CSCF to perform back-to-back user agent (B2BUA) functions, and Multimedia Resource Function Controller (MRFC) can be used instead of AGW.

[0382] The call processing method provided in this application can be executed by a call processing device. This application uses the example of a call processing device executing the call processing method to illustrate the call processing device provided in this application.

[0383] This application provides a call processing apparatus. As an example, the call processing apparatus may be a communication device or a component within a communication device, such as a chip. The communication device may be a user equipment, a network-side device, or a server, etc. Exemplarily, the user equipment may include, but is not limited to, the type of user equipment 11 listed above, and the network-side device may include, but is not limited to, the type of network-side device 12 listed above. This application does not impose specific limitations.

[0384] The call processing device includes a receiving module, a transmitting module, and a processing module. These modules can be implemented in software or hardware. When implemented in hardware, the processing module can be implemented by a processor. For example, the processor can include a general-purpose processor, a special-purpose processor, such as a Central Processing Unit (CPU), a microprocessor, a Digital Signal Processor (DSP), an Artificial Intelligence (AI) processor, a Graphics Processing Unit (GPU), an Application Specific Integrated Circuit (ASIC), a Network Processor (NP), a Field Programmable Gate Array (FPGA), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc. The receiving and transmitting modules can be implemented by a communication interface, which can include one or more of the following: transceiver, pins, circuits, buses, radio frequency units, etc.

[0385] See Figure 15 When the call processing device is a network-side device or a component within a network-side device, the call processing device 1500 includes:

[0386] The first transmission module 1501 is used to negotiate Session Description Protocol (SDP) with the first user equipment during the establishment of a call, wherein the SDP negotiation is used to determine the first voice code and the second voice code.

[0387] Optionally, the first transmission module 1501 is specifically configured to perform at least one of the following:

[0388] Receive a first request message from the first user equipment, the first request message including a media line corresponding to the first speech code and a media line corresponding to the second speech code;

[0389] Send a first message to the first user equipment, the first message including the media line corresponding to the first voice code and the media line corresponding to the second voice code.

[0390] Optionally, the first request message may further include at least one of a first indication information and a second indication information;

[0391] Wherein, the first indication information is used to indicate at least one of the following:

[0392] The first user equipment supports Real-Time Transport Protocol (RTP) for voice.

[0393] The first user equipment supports multiple RTP for audio;

[0394] The first user equipment supports multiplexing for voice;

[0395] The first user equipment supports multiplexing for audio;

[0396] The second indication information is used to indicate that the first user equipment supports the second voice coding.

[0397] Optionally, the voice call processing device 1500 further includes:

[0398] The first processing module is used to delete the media line corresponding to the second speech code in the first request message to obtain the second request message;

[0399] The first transmission module 1501 is further configured to send the second request message to the second user equipment.

[0400] Optionally, the first transmission module 1501 is further configured to receive a response message of the second request message from the second user equipment;

[0401] The first processing module is further configured to add a media line corresponding to the second voice encoding to the response message of the second request message, and send the modified response message of the second request message to the first user equipment.

[0402] Optionally, the first transmission module 1501 is further configured to send a third request message to the second network entity, the third request message being used to request the allocation of a first transmission address corresponding to the first voice code;

[0403] The media line corresponding to the second voice code in the response message of the modified second request message includes the first transmission address corresponding to the second voice code; the first transmission address is used for communication between the first user equipment and the second network entity.

[0404] Optionally, the first message is a first session initiation protocol (SIP) invitation message or a first SIP update message.

[0405] Optionally, the first transmission module 1501 is further configured to receive a response message of the first message from the first user equipment, the response message of the first message including a media line corresponding to the first voice encoding and a media line corresponding to the second voice encoding.

[0406] Optionally, the first message may further include third indication information, which indicates that the network supports a second speech codec.

[0407] Optionally, the first transmission module 1501 is further configured to send a fourth indication information to the second network entity, the fourth indication information being used to instruct the second network entity to perform encoding conversion.

[0408] Optionally, the first transmission module 1501 is specifically configured to perform at least one of the following:

[0409] A fourth request message is received from the first user equipment, the fourth request message including a media line, and the media line including an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code;

[0410] A second message is sent to the first user equipment. The second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0411] Optionally, the fourth request message further includes at least one of the following:

[0412] The fifth instruction information is used to indicate the negotiation of two voice coding types;

[0413] The first indication information indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for voice; the first user equipment supports RTP for audio; the first user equipment supports multiplexing for voice; and the first user equipment supports multiplexing for audio.

[0414] The sixth indication information is used to indicate that the second speech encoding is an additional encoding;

[0415] Alternatively, the second message may include at least one of the following:

[0416] The third indication information is used to indicate that the network supports the second speech coding.

[0417] The seventh instruction information is used to indicate the negotiation of two voice coding types;

[0418] The eighth indication information is used to indicate that the second speech code is an additional code.

[0419] Optionally, the voice call processing device 1500 further includes:

[0420] The first processing module is used to delete the attribute row corresponding to the second voice code in the fourth request message to obtain the fifth request message;

[0421] The first transmission module 1501 is also used to send the fifth request message to the second user equipment.

[0422] Optionally, the first transmission module 1501 is further configured to receive a response message of the fifth request message from the second user equipment;

[0423] The first processing module is further configured to add an attribute line corresponding to the second voice code to the response message of the fifth request message, and send the modified response message of the fifth request message to the first user equipment.

[0424] Optionally, the response message to the modified fifth request message further includes at least one of the following:

[0425] The seventh instruction information is used to indicate the negotiation of two voice coding types;

[0426] The eighth indication information is used to indicate that the second speech code is an additional code.

[0427] Optionally, the second message is a second SIP invitation message or a second SIP update message.

[0428] Optionally, the first transmission module 1501 is further configured to receive a response message of the second message from the first user equipment, the response message of the second message including a media line, and the media line including an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0429] Optionally, the first transmission module 1501 is further configured to receive first information from the first user equipment, the first information being information sent during the registration process, the first information including first indication information, the first indication information being configured to indicate at least one of the following: the first user equipment supports Real-time Transport Protocol for Voice (RTP); the first user equipment supports RTP for Audio; the first user equipment supports multiplexing for Voice; the first user equipment supports multiplexing for Audio; and send ninth indication information to the first user equipment, the ninth indication information being configured to indicate at least one of the following: the first network entity supports RTP for Voice; the first network entity supports RTP for Audio; the first network entity supports multiplexing for Voice; and the first network entity supports multiplexing for Audio.

[0430] Optionally, the first transmission module 1501 is further configured to, upon determining that the second voice coding is to be initiated, send a first notification message to the second network entity, wherein the first notification message is used to notify the second network entity to initiate the second voice coding.

[0431] Optionally, the first transmission module 1501 is further configured to receive a second notification message from a third network entity, the second notification message being used to notify the requirement information that meets the second voice coding.

[0432] The first processing module is used to determine to start the second voice encoding based on the second notification message.

[0433] See Figure 16 When the call processing device is a network-side device or a component within a network-side device, the call processing device 1600 includes:

[0434] The second transmission module 1601 is configured to perform a first operation, the first operation including at least one of the following:

[0435] The first data received from the second user equipment is converted into second data, and the second data is sent to the first user equipment. The first data is the data corresponding to the first voice code, and the second data is the data corresponding to the second voice code.

[0436] The third data received from the first user equipment is converted into fourth data and sent to the second user equipment. The third data is the data corresponding to the second voice code, and the fourth data is the data corresponding to the first voice code.

[0437] Optionally, the second transmission module 1601 is further configured to receive a first notification message from the first network entity, the first notification message being used to notify the second network entity to initiate the second voice encoding.

[0438] Optionally, the second transmission module 1601 is further configured to perform any of the following:

[0439] The second data is sent to the first user equipment based on the Real-Time Transport Protocol (RTP) corresponding to the second voice code.

[0440] The second data is sent to the first user equipment based on the RTP corresponding to the first voice code.

[0441] For details, see Figure 17 When the voice call processing device is a user equipment or a component within a user equipment, the voice call processing device 1700 includes:

[0442] The third transmission module 1701 is used to negotiate Session Description Protocol (SDP) with the first network entity during the call establishment process of the first user equipment. The SDP negotiation is used to determine the first voice code and the second voice code.

[0443] Optionally, the third transmission module 1701 is specifically configured to perform at least one of the following:

[0444] Send a first request message to the first network entity. The first request message includes a media line corresponding to the first speech code and a media line corresponding to the second speech code.

[0445] The first message is received from the first network entity. The first message includes a media line corresponding to the first speech code and a media line corresponding to the second speech code.

[0446] Optionally, the first request message may further include at least one of a first indication information and a second indication information;

[0447] Wherein, the first indication information is used to indicate at least one of the following:

[0448] The first user equipment supports Real-Time Transport Protocol (RTP) for voice.

[0449] The first user equipment supports multiple RTP for audio;

[0450] The first user equipment supports multiplexing for voice;

[0451] The first user equipment supports multiplexing for audio;

[0452] The second indication information is used to indicate that the first user equipment supports the second voice coding.

[0453] Optionally, the first message is a first session initiation protocol (SIP) invitation message or a first SIP update message.

[0454] Optionally, the third transmission module 1701 is further configured to send a response message of the first message to the first network entity, the response message of the first message including the media line corresponding to the first speech code and the media line corresponding to the second speech code.

[0455] Optionally, the first message may further include third indication information, which indicates that the network supports a second speech codec.

[0456] Optionally, the third transmission module 1701 is specifically configured to perform at least one of the following:

[0457] A fourth request message is sent to the first network entity. The fourth request message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0458] The second message is received from the first network entity. The second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0459] Optionally, the fourth request message further includes at least one of the following:

[0460] The fifth instruction information is used to indicate the negotiation of two voice coding types;

[0461] The first indication information indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for voice; the first user equipment supports RTP for audio; the first user equipment supports multiplexing for voice; and the first user equipment supports multiplexing for audio.

[0462] The sixth indication information is used to indicate that the second speech encoding is an additional encoding;

[0463] Alternatively, the second message may include at least one of the following:

[0464] The third indication information is used to indicate that the network supports the second speech coding.

[0465] The seventh instruction information is used to indicate the negotiation of two voice coding types;

[0466] The eighth indication information is used to indicate that the second speech code is an additional code.

[0467] Optionally, the second message is a second SIP invitation message or a second SIP update message.

[0468] Optionally, the third transmission module 1701 is specifically used to send a response message of the second message to the first network entity. The response message of the second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

[0469] Optionally, the third transmission module 1701 is specifically configured to send first information to the first network entity, the first information being information sent during the registration process, the first information including first indication information, the first indication information being configured to indicate at least one of the following: the first user equipment supports Real-time Transport Protocol for Voice (RTP); the first user equipment supports RTP for Audio; the first user equipment supports multiplexing for Voice; the first user equipment supports multiplexing for Audio; and receive ninth indication information from the first network entity, the ninth indication information being configured to indicate at least one of the following: the first network entity supports RTP for Voice; the first network entity supports RTP for Audio; the first network entity supports multiplexing for Voice; and the first network entity supports multiplexing for Audio.

[0470] Optionally, the voice call processing device further includes:

[0471] The second processing module is used to initiate a second voice encoding when the first user equipment receives the second data, wherein the second data corresponds to the second voice encoding.

[0472] Optionally, the second processing module is further configured to perform at least one of the following:

[0473] If the first user equipment receives data based on the Real-time Transport Protocol (RTP) corresponding to the second voice code, it is determined that the second data has been received;

[0474] When the first user equipment receives data based on the RTP corresponding to the first voice code, it determines that the second data has been received based on the packet header of the received data.

[0475] Optionally, the third transmission module 1701 is further configured to send third data based on the second voice encoding.

[0476] Optionally, the second processing module is further configured to determine the initiation of the second voice encoding based on the obtained transmission bit rate.

[0477] Optionally, the third transmission module 1701 is further configured to receive tenth indication information from the fourth network entity, the tenth indication information being used to indicate the transmission bit rate.

[0478] The call processing device provided in this application embodiment can achieve... Figure 4 , Figure 6 and Figure 7 The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.

[0479] like Figure 18 As shown, this application embodiment also provides a communication device 1800, including a processor 1801 and a memory 1802. The memory 1802 stores a program or instructions that can run on the processor 1801. When the program or instructions are executed by the processor 1801, they implement the various steps of the above-described call processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0480] This application embodiment also provides a terminal, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement, for example... Figure 7 The steps in the method embodiment shown are illustrated. This terminal embodiment corresponds to the user equipment-side method embodiment described above. All implementation processes and methods of the above method embodiments can be applied to this terminal embodiment and achieve the same technical effect. The terminal can be... Figure 17 The call processing device shown. Specifically, Figure 19 A schematic diagram of the hardware structure of a terminal to implement an embodiment of this application.

[0481] The terminal 1900 includes, but is not limited to, at least some of the following components: radio frequency unit 1901, network module 1902, audio output unit 1903, input unit 1904, sensor 1905, display unit 1906, user input unit 1907, interface unit 1908, memory 1909, and processor 1910.

[0482] Those skilled in the art will understand that the terminal 1900 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1910 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 19 The terminal structure shown does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0483] It should be understood that, in this embodiment, the input unit 1904 may include a graphics processor 19041 and a microphone 19042. The graphics processor 19041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1906 may include a display panel 19061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1907 includes at least one of a touch panel 19071 and other input devices 19072. The touch panel 19071 is also called a touch screen. The touch panel 19071 may include a touch detection device and a touch controller. Other input devices 19072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0484] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 1901 can transmit it to the processor 1910 for processing; in addition, the radio frequency unit 1901 can send uplink data to the network-side device. Typically, the radio frequency unit 1901 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.

[0485] The memory 1909 can be used to store software programs or instructions, as well as various data. The memory 1909 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1909 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1909 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0486] Processor 1910 may include one or more processing units; optionally, processor 1910 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1910.

[0487] The radio frequency unit 1901 is used to negotiate Session Description Protocol (SDP) with a first network entity during the establishment of a call by the first user equipment. The SDP negotiation is used to determine the first voice code and the second voice code.

[0488] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the user equipment side method embodiment and achieve the same or corresponding technical effects. To avoid repetition, it will not be described again here.

[0489] This application embodiment also provides a network-side device, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement, for example... Figure 4 or Figure 6 The steps of the method embodiment shown are illustrated. This network-side device embodiment corresponds to the above-described network-side device method embodiment. All implementation processes and methods of the above-described method embodiments can be applied to this network-side device embodiment and can achieve the same technical effect.

[0490] Specifically, embodiments of this application also provide a network-side device. For example... Figure 20 As shown, the network-side device 2000 includes: a processor 2001, a network interface 2002, and a memory 2003. This network-side device can be... Figure 15 Or the voice call processing device shown in Figure 16. The network interface 2002 is, for example, a common public radio interface (CPRI).

[0491] When the network-side device is the first network entity, the network interface 2002 is used to negotiate the Session Description Protocol (SDP) with the first user equipment during the call establishment process. The SDP negotiation is used to determine the first voice code and the second voice code.

[0492] When the network-side device is a second network entity, the network interface 2002 is used to perform a first operation, the first operation including at least one of the following:

[0493] The first data received from the second user equipment is converted into second data, and the second data is sent to the first user equipment. The first data is the data corresponding to the first voice code, and the second data is the data corresponding to the second voice code.

[0494] The third data received from the first user equipment is converted into fourth data and sent to the second user equipment. The third data is the data corresponding to the second voice code, and the fourth data is the data corresponding to the first voice code.

[0495] Furthermore, the network-side device 2000 in this embodiment of the application also includes: a program or instructions stored in a memory 2003 and executable on a processor 2001, wherein the processor 2001 calls the program or instructions in the memory 2003 to execute. Figure 15 or Figure 16 The methods executed by each module shown achieve the same technical effect, and to avoid repetition, they will not be described in detail here.

[0496] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described voice call processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0497] The processor mentioned above is either the processor in the user equipment or the processor in the network-side device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.

[0498] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described voice call processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0499] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0500] This application also provides a computer program / program product, which includes computer instructions. The computer program / program product is executed by at least one processor to implement the various processes of the above-described voice call processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0501] This application also provides a wireless communication system, including: a user equipment, a first network entity, and a second network entity. The user equipment can be used to execute the steps of the voice call processing method on the first user equipment side as described above. The first network entity can be used to execute the steps of the voice call processing method on the first network entity side as described above. The second network entity can be used to execute the steps of the voice call processing method on the second network entity side as described above.

[0502] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0503] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.

[0504] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.

Claims

1. A call processing method, characterized in that, include: During the establishment of a call by the first user equipment, the first network entity negotiates a Session Description Protocol (SDP) with the first user equipment. The SDP negotiation is used to determine the first voice codec and the second voice codec.

2. The method according to claim 1, characterized in that, The first network entity negotiates Session Description Protocol (SDP) with the first user equipment, including at least one of the following: The first network entity receives a first request message from the first user equipment, the first request message including a media line corresponding to the first speech code and a media line corresponding to the second speech code; The first network entity sends a first message to the first user equipment, the first message including the media line corresponding to the first voice code and the media line corresponding to the second voice code.

3. The method according to claim 2, characterized in that, The first request message also includes at least one of a first instruction message and a second instruction message; Wherein, the first indication information is used to indicate at least one of the following: The first user equipment supports Real-Time Transport Protocol (RTP) for voice. The first user equipment supports multiple RTP for audio; The first user equipment supports multiplexing for voice; The first user equipment supports multiplexing for audio; The second indication information is used to indicate that the first user equipment supports the second voice coding.

4. The method according to claim 2, characterized in that, The method further includes: The first network entity deletes the media line corresponding to the second speech code in the first request message to obtain the second request message; The first network entity sends the second request message to the second user equipment.

5. The method according to claim 4, characterized in that, The method further includes: The first network entity receives a response message to the second request message from the second user equipment; The first network entity adds a media line corresponding to the second voice code to the response message of the second request message, and sends the modified response message of the second request message to the first user equipment.

6. The method according to claim 5, characterized in that, The method further includes: The first network entity sends a third request message to the second network entity, the third request message being used to request the allocation of a first transmission address corresponding to the first voice code; The media line corresponding to the second voice code in the response message of the modified second request message includes the first transmission address corresponding to the second voice code; the first transmission address is used for communication between the first user equipment and the second network entity.

7. The method according to any one of claims 2 to 6, characterized in that, The first message is either a first session initiation protocol SIP invitation message or a first SIP update message.

8. The method according to claim 2 or 7, characterized in that, The method further includes: The first network entity receives a response message to the first message from the first user equipment. The response message to the first message includes a media line corresponding to the first speech code and a media line corresponding to the second speech code.

9. The method according to claim 2, 7 or 8, characterized in that, The first message also includes a third indication, which indicates that the network supports a second speech code.

10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: The first network entity sends a fourth instruction message to the second network entity, the fourth instruction message being used to instruct the second network entity to perform encoding conversion.

11. The method according to claim 1, characterized in that, The first network entity negotiates Session Description Protocol (SDP) with the first user equipment, including at least one of the following: The first network entity receives a fourth request message from the first user equipment. The fourth request message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code. The first network entity sends a second message to the first user equipment. The second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

12. The method according to claim 11, characterized in that, The fourth request message also includes at least one of the following: The fifth instruction information is used to indicate the negotiation of two voice coding types; The first indication information indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for multiplexing voice; the first user equipment supports RTP for multiplexing audio; the first user equipment supports multiplexing codecs for voice. The first user equipment supports multiplexing for audio; The sixth indication information is used to indicate that the second speech encoding is an additional encoding; Alternatively, the second message may include at least one of the following: The third indication information is used to indicate that the network supports the second speech coding. The seventh instruction information is used to indicate the negotiation of two voice coding types; The eighth indication information is used to indicate that the second speech code is an additional code.

13. The method according to claim 11 or 12, characterized in that, The method further includes: The first network entity deletes the attribute row corresponding to the second voice code in the fourth request message to obtain the fifth request message; The first network entity sends the fifth request message to the second user equipment.

14. The method according to claim 13, characterized in that, The method further includes: The first network entity receives a response message to the fifth request message from the second user equipment; The first network entity adds an attribute line corresponding to the second voice code to the response message of the fifth request message, and sends the modified response message of the fifth request message to the first user equipment.

15. The method according to claim 14, characterized in that, The response message to the modified fifth request message also includes at least one of the following: The seventh instruction information is used to indicate the negotiation of two voice coding types; The eighth indication information is used to indicate that the second speech code is an additional code.

16. The method according to claim 11, characterized in that, The second message is either a second SIP invitation message or a second SIP update message.

17. The method according to claim 11 or 16, characterized in that, The method further includes: The first network entity receives a response message to the second message from the first user equipment. The response message to the second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

18. The method according to any one of claims 1 to 17, characterized in that, The method further includes: The first network entity receives first information from the first user equipment. The first information is information sent during the registration process. The first information includes first indication information, which indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for voice; the first user equipment supports RTP for audio; the first user equipment supports multiplexed codecs for voice; and the first user equipment supports multiplexed codecs for audio. The first network entity sends a ninth indication message to the first user equipment, the ninth indication message indicating at least one of the following: the first network entity supports Real-Time Transport Protocol (RTP) for voice; the first network entity supports RTP for audio; the first network entity supports multiple codecs for voice; the first network entity supports multiple codecs for audio.

19. The method according to any one of claims 1 to 18, characterized in that, The method further includes: If the second voice coding is initiated, the first network entity sends a first notification message to the second network entity, which is used to notify the second network entity to initiate the second voice coding.

20. The method according to claim 19, characterized in that, The method further includes: The first network entity receives a second notification message from the third network entity. The second notification message is used to notify the requirement information that meets the second voice code. The first network entity determines to initiate the second voice coding based on the second notification message.

21. A voice call processing method, characterized in that, include: The second network entity performs a first operation, which includes at least one of the following: The first data received from the second user equipment is converted into second data, and the second data is sent to the first user equipment. The first data is the data corresponding to the first voice code, and the second data is the data corresponding to the second voice code. The third data received from the first user equipment is converted into fourth data and sent to the second user equipment. The third data is the data corresponding to the second voice code, and the fourth data is the data corresponding to the first voice code.

22. The method according to claim 21, characterized in that, The method further includes: The second network entity receives a first notification message from the first network entity, the first notification message being used to notify the second network entity to start the second voice coding.

23. The method according to claim 21 or 22, characterized in that, Sending the second data to the first user equipment includes any one of the following: The second data is sent to the first user equipment based on the Real-Time Transport Protocol (RTP) corresponding to the second voice code. The second data is sent to the first user equipment based on the RTP corresponding to the first voice code.

24. A voice call processing method, characterized in that, include: During the establishment of a call by the first user equipment, the first user equipment negotiates a Session Description Protocol (SDP) with the first network entity. The SDP negotiation is used to determine the first voice code and the second voice code.

25. The method according to claim 24, characterized in that, The first user equipment negotiates Session Description Protocol (SDP) with the first network entity, including at least one of the following: The first user equipment sends a first request message to the first network entity, the first request message including the media line corresponding to the first speech code and the media line corresponding to the second speech code; The first user equipment receives a first message from the first network entity, the first message including a media line corresponding to the first voice code and a media line corresponding to the second voice code.

26. The method according to claim 25, characterized in that, The first request message also includes at least one of a first instruction message and a second instruction message; Wherein, the first indication information is used to indicate at least one of the following: The first user equipment supports Real-Time Transport Protocol (RTP) for voice. The first user equipment supports multiple RTP for audio; The first user equipment supports multiplexing for voice; The first user equipment supports multiplexing for audio; The second indication information is used to indicate that the first user equipment supports the second voice coding.

27. The method according to claim 25 or 26, characterized in that, The first message is either a first session initiation protocol SIP invitation message or a first SIP update message.

28. The method according to claim 25 or 27, characterized in that, The method further includes: The first user equipment sends a response message to the first network entity for the first message, the response message for the first message including the media line corresponding to the first voice code and the media line corresponding to the second voice code.

29. The method according to claim 25, 27 or 28, characterized in that, The first message also includes a third indication, which indicates that the network supports a second speech code.

30. The method according to claim 24, characterized in that, The first user equipment negotiates Session Description Protocol (SDP) with the first network entity, including at least one of the following: The first user equipment sends a fourth request message to the first network entity. The fourth request message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code. The first user equipment receives a second message from the first network entity. The second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

31. The method according to claim 30, characterized in that, The fourth request message also includes at least one of the following: The fifth instruction information is used to indicate the negotiation of two voice coding types; The first indication information indicates at least one of the following: the first user equipment supports Real-Time Transport Protocol (RTP) for multiplexing voice; the first user equipment supports RTP for multiplexing audio; the first user equipment supports multiplexing codecs for voice. The first user equipment supports multiplexing for audio; The sixth indication information is used to indicate that the second speech encoding is an additional encoding; Alternatively, the second message may include at least one of the following: The third indication information is used to indicate that the network supports the second speech coding. The seventh instruction information is used to indicate the negotiation of two voice coding types; The eighth indication information is used to indicate that the second speech code is an additional code.

32. The method according to claim 30, characterized in that, The second message is either a second SIP invitation message or a second SIP update message.

33. The method according to claim 30 or 32, characterized in that, The method further includes: The first user equipment sends a response message to the first network entity for the second message. The response message for the second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

34. The method according to any one of claims 24 to 33, characterized in that, The method further includes: The first user equipment sends first information to the first network entity. The first information is information sent during the registration process. The first information includes first indication information, which indicates at least one of the following: the first user equipment supports Real-time Transport Protocol for Voice (RTP); the first user equipment supports RTP for Audio; the first user equipment supports multiplexed codecs for Voice; and the first user equipment supports multiplexed codecs for Audio. The first user equipment receives a ninth indication information from the first network entity, the ninth indication information being used to indicate at least one of the following: the first network entity supports Real-Time Transport Protocol (RTP) for voice; the first network entity supports RTP for audio; the first network entity supports multiple codecs for voice; the first network entity supports multiple codecs for audio.

35. The method according to any one of claims 24 to 34, characterized in that, The method further includes: When the first user equipment receives the second data, the first user equipment initiates the second voice encoding, wherein the second data corresponds to the second voice encoding.

36. The method according to claim 35, characterized in that, The method further includes at least one of the following: If the first user equipment receives data based on the Real-time Transport Protocol (RTP) corresponding to the second voice code, it is determined that the second data has been received; When the first user equipment receives data based on the RTP corresponding to the first voice code, it determines that the second data has been received based on the packet header of the received data.

37. The method according to claim 35 or 36, characterized in that, The method further includes: The first user equipment sends third data based on the second voice encoding.

38. The method according to any one of claims 24 to 37, characterized in that, The method further includes: The first user equipment determines to initiate the second voice coding based on the obtained transmission bit rate.

39. The method according to claim 38, characterized in that, The method further includes: The first user equipment receives a tenth indication information from a fourth network entity, the tenth indication information being used to indicate the transmission bit rate.

40. A call processing device, characterized in that, include: The first transmission module is used to negotiate Session Description Protocol (SDP) with the first user equipment during the establishment of a call. The SDP negotiation is used to determine the first voice code and the second voice code.

41. The apparatus according to claim 40, characterized in that, The first transmission module is specifically used to perform at least one of the following: Receive a first request message from the first user equipment, the first request message including a media line corresponding to the first speech code and a media line corresponding to the second speech code; Send a first message to the first user equipment, the first message including the media line corresponding to the first voice code and the media line corresponding to the second voice code.

42. The apparatus according to claim 40, characterized in that, The first transmission module is specifically used to perform at least one of the following: A fourth request message is received from the first user equipment, the fourth request message including a media line, and the media line including an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code; A second message is sent to the first user equipment. The second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

43. A voice call processing device, characterized in that, include: The second transmission module is configured to perform a first operation, the first operation including at least one of the following: The first data received from the second user equipment is converted into second data, and the second data is sent to the first user equipment. The first data is the data corresponding to the first voice code, and the second data is the data corresponding to the second voice code. The third data received from the first user equipment is converted into fourth data and sent to the second user equipment. The third data is the data corresponding to the second voice code, and the fourth data is the data corresponding to the first voice code.

44. A voice call processing device, characterized in that, include: The third transmission module is used to negotiate Session Description Protocol (SDP) with the first network entity during the call establishment process of the first user equipment. The SDP negotiation is used to determine the first voice codec and the second voice codec.

45. The apparatus according to claim 44, characterized in that, The third transmission module is specifically used to perform at least one of the following: Send a first request message to the first network entity. The first request message includes a media line corresponding to the first speech code and a media line corresponding to the second speech code. The first message is received from the first network entity. The first message includes a media line corresponding to the first speech code and a media line corresponding to the second speech code.

46. ​​The apparatus according to claim 44, characterized in that, The third transmission module is specifically used to perform at least one of the following: A fourth request message is sent to the first network entity. The fourth request message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code. The second message is received from the first network entity. The second message includes a media line, and the media line includes an attribute line corresponding to the first speech code and an attribute line corresponding to the second speech code.

47. A user equipment, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the call processing method as described in any one of claims 24 to 39.

48. A network-side device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the call processing method as described in any one of claims 1 to 23.

49. A chip, characterized in that, The chip includes a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the call processing method as described in any one of claims 1 to 39.

50. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the call processing method as described in any one of claims 1 to 39.