A communication method, apparatus and system

Transmitting facial video streams through video call media transmission channels solves the problems of resource waste and high skill requirements in existing technologies, and achieves efficient and convenient user identity authentication.

CN116866504BActive Publication Date: 2025-10-28HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210317245.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-28
Publication Date
2025-10-28
Estimated Expiration
2042-03-28

AI Technical Summary

Technical Problem

In existing technologies, user terminals need to install a face recognition app and perform complex operations, which consumes additional bandwidth and port resources for face image transmission, resulting in resource waste and high skill requirements.

Method used

The face video stream is transmitted through the video call media transmission channel, avoiding the need to establish an additional transmission channel. Face recognition is performed using the media server, eliminating the need to install an app or perform complicated operations.

Benefits of technology

It saves bandwidth and port resources, simplifies operation processes, reduces skill requirements, and improves identity authentication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116866504B_ABST
    Figure CN116866504B_ABST
Patent Text Reader

Abstract

A communication method, apparatus, and system are provided, relating to the field of communication technology, which can save bandwidth resources and terminal port resources occupied by facial recognition during a call. The method includes: establishing a video call media transmission channel for transmitting call video streams between a call terminal and a peer call terminal in a video call service, the call video stream including video content captured by either the call terminal or the peer call terminal; receiving a SIP message including a facial recognition request identifier from a media server, the facial recognition request identifier being used to request facial recognition of the user corresponding to the call terminal; sending a response message of the SIP message to the media server, the response message indicating that the user corresponding to the call terminal agrees to facial recognition; then sending a facial video stream, including a facial image of the user corresponding to the call terminal, to the media server through the video call media transmission channel; and finally receiving the facial recognition result from the media server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a communication method, apparatus and system. Background Technology

[0002] Based on the operator's network, facial recognition is used to authenticate the user's identity online during voice / video calls between the user and the customer service center via the terminal, providing the user with safe and convenient services.

[0003] Currently, when performing facial recognition on users online, an application (APP) for facial recognition needs to be installed on the user's terminal (hereinafter referred to as the user terminal). This APP is designated by the customer service center for performing facial recognition on users. Then, the user terminal collects the user's facial image and uploads it to the APP, which then completes the facial recognition, or the APP sends the facial image to the recognition server, which then completes the facial recognition.

[0004] The aforementioned facial recognition methods require the installation of an app on the user terminal and complex operations, demanding a high level of skill from personnel. Furthermore, a dedicated transmission channel for transmitting facial images needs to be established between the user terminal and the app. If facial recognition is performed during a call, establishing this dedicated transmission channel requires additional bandwidth resources, and transmitting facial images through this channel also requires additional port resources on the user terminal. Summary of the Invention

[0005] This application provides a communication method, apparatus, and system that can save bandwidth resources and terminal port resources occupied by face recognition during a call.

[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, embodiments of this application provide a communication method executed by a calling terminal. The method includes: establishing a video call media transmission channel for transmitting a video stream between the calling terminal and a peer calling terminal in a video call service; the video stream containing video content captured by either the calling terminal or the peer calling terminal; receiving a SIP message from a media server, the SIP message including a face recognition request identifier for requesting face recognition of a user corresponding to the calling terminal; sending a response message to the media server, the response message indicating that the user corresponding to the calling terminal agrees to face recognition; then sending a face video stream to the media server through the video call media transmission channel, the face video stream including a face image of the user corresponding to the calling terminal; and finally receiving a face recognition result from the media server.

[0008] The communication method provided in this application embodiment allows the call terminal to transmit the face video stream through the video call media transmission channel originally used for transmitting call video streams when face recognition is required during a call. In this way, there is no need to spend additional bandwidth resources to establish a dedicated transmission channel for transmitting face video streams, and there is no need to occupy additional port resources of the call terminal.

[0009] Furthermore, compared with existing communication methods, the technical solution provided in this application embodiment does not require the installation of a face recognition APP on the terminal. Thus, it does not require operators to perform complex related operations, does not require operators to have high operating skills, and does not require terminal calls.

[0010] In one possible implementation, the face video stream is a video stream obtained by capturing images through the camera device of the call terminal; or, the face video stream is a video stream obtained from the storage device of the call terminal.

[0011] In one possible implementation, the communication method provided in this application embodiment further includes: receiving source indication information from the media server, wherein the source indication information instructs the calling terminal to acquire a face video stream through the camera device of the calling terminal or to acquire a face video stream from the storage device of the calling terminal.

[0012] In one possible implementation, the face video stream is a video stream obtained from the storage device of the calling terminal. Before sending the face video stream to the media server through the video call media transmission channel, the communication method provided in this application embodiment further includes: receiving transmission channel indication information from the media server. The transmission channel indication information instructs the calling terminal to transmit the face video stream through the video call media transmission channel. This transmission channel indication information can be carried in a SIP message. The media server can explicitly indicate (i.e., send the transmission channel indication information) that the face video stream is transmitted through the video call media transmission channel. In some cases, the media server can also implicitly indicate that the face video stream is transmitted through the video call media transmission channel, for example, by carrying the media server's SDP information in the SIP message and the calling terminal's SDP information in the response message of the SIP message, to negotiate (or instruct) the use of a video call media transmission channel established based on this pair of SDP information (the media server's SDP information and the calling terminal's SDP) to transmit the face video stream. It should be understood that this video call media transmission channel was originally used for transmitting call video streams.

[0013] In one possible implementation, before sending the face video stream to the media server through the video call media transmission channel, the method further includes: stopping the transmission of the call video stream through the video call media transmission channel, so that the face video stream obtained from the storage device of the call terminal can be transmitted through the video call media transmission channel.

[0014] It should be noted that when the face video stream transmitted by the call terminal is captured by the camera device of the call terminal, the aforementioned video call media transmission channel does not actually stop transmitting the call video stream. That is, the face video stream is the call video stream. The difference is that the content of the call video stream may change from other video content to video content including face images.

[0015] In one possible implementation, the method provided in this application embodiment further includes: receiving posture indication information from the media server, wherein the posture indication information instructs the user corresponding to the call terminal to adjust the user's posture so that the face image meets preset conditions.

[0016] In this embodiment of the application, when the face video stream is a video stream captured by the camera device of the call terminal, the call terminal captures the user's face image in real time and transmits the face image to the media server. It is understood that when the camera device of the call terminal captures the face image, the captured face image may not meet the preset conditions due to the influence of the user's head posture. The preset conditions may include, but are not limited to, the position and angle of the user's head within a certain range (for example, the face should be located within a preset detection frame and the user's face should be facing the camera device). This may cause face recognition to fail or result in inaccurate face recognition results. Based on this, after receiving the face video stream sent by the calling terminal, if the media server determines that the face image in the face video stream does not meet the preset conditions, the media server can send posture indication information to the calling terminal (correspondingly, the calling terminal receives the posture indication information from the media server). This posture indication information instructs the user corresponding to the calling terminal to adjust the user's posture so that the face image meets the preset conditions. For example, when the user's face is not entirely within the preset detection frame, the posture indication information prompts the user to place the entire face within the detection frame. Or, when the angle of the user's face is not appropriate, the posture indication information prompts the user to turn their face in a certain direction (for example, prompting the user to turn their head to the right).

[0017] In one possible implementation, the face recognition result may include information indicating successful face recognition or information indicating failure. Optionally, when the face image extracted from the face video stream matches the face image registered in the operator's system, the face recognition result may also include the identification information of the registered user corresponding to the call terminal, such as the user's name, identification document, and other identity information.

[0018] In one possible implementation, the face recognition request identifier can be carried in the header field of the SIP message, or, if the SIP message includes the SDP information of the media server, the face recognition request identifier can also be carried in the SDP information of the media server.

[0019] Secondly, embodiments of this application provide a communication method executed by a calling terminal. The method includes: establishing a video call media transmission channel for transmitting a video stream between the calling terminal and a peer calling terminal in a video call service; the video stream containing video content captured by either the calling terminal or the peer calling terminal; sending a face recognition request to a media server, the face recognition request including a face recognition request identifier used to request face recognition of a user corresponding to the peer calling terminal in a call with the calling terminal; receiving a face recognition result from the media server through the video call media transmission channel, the face recognition result being the result of face recognition of the user corresponding to the peer calling terminal based on the face video stream; and processing a service request from the peer calling terminal based on the face recognition result.

[0020] The communication method provided in this application embodiment allows for the transmission of face recognition results through the video call media transmission channel originally used for transmitting call video streams when face recognition is required during a call. In this way, no additional bandwidth resources are required to establish a dedicated transmission channel for transmitting face recognition results, and no additional port resources of the call terminal are occupied.

[0021] Furthermore, compared with existing communication methods, the technical solution provided in this application embodiment does not require the installation of a face recognition APP on the terminal. Thus, it does not require operators to perform complex related operations, does not require operators to have high operating skills, and does not require terminal calls.

[0022] In one possible implementation, the face video stream is a video stream obtained by capturing images with the camera device of the peer communication terminal; or, the face video stream is a video stream obtained from the storage device of the peer communication terminal.

[0023] In one possible implementation, before receiving the face recognition result from the media server via the video call media transmission channel, the communication method provided in this application embodiment further includes: receiving transmission channel indication information from the media server, wherein the transmission channel indication information instructs the calling terminal to receive the face recognition result via the video call media transmission channel. Similar to the first aspect described above, the transmission channel indication information can be carried in a SIP message. The media server can explicitly indicate (i.e., send the transmission channel indication information) that the face recognition result is transmitted via the video call media transmission channel. In some cases, the media server can also implicitly indicate that the face recognition result is transmitted via the video call media transmission channel.

[0024] In one possible implementation, before receiving the face recognition result from the media server through the video call media transmission channel, the communication method provided in this application embodiment further includes: stopping the transmission of the call video stream through the video call media transmission channel, so that the face recognition result can be transmitted through the video call media transmission channel.

[0025] In one possible implementation, the face recognition result includes information indicating successful face recognition or information indicating failed face recognition. The step of processing the service request of the peer call terminal based on the face recognition result includes: processing the service request of the peer call terminal when the face recognition result includes the information indicating successful face recognition, thus ensuring secure processing of user services.

[0026] Thirdly, embodiments of this application provide a communication method executed by a media server. The method includes: establishing a first video call media transmission channel and a second video call media transmission channel, wherein the first video call media transmission channel is a video call media transmission channel between a calling terminal and the media server, and the second video call media transmission channel is a video call media transmission channel between the media server and a peer calling terminal; the first and second video call media transmission channels are used for transmitting a video stream between the calling terminal and the peer calling terminal in a video call service, the video stream containing video content captured by the calling terminal or the peer calling terminal; and then transmitting the video stream from the peer calling terminal... The terminal receives a face recognition request, which includes a face recognition request identifier used to request face recognition of the user corresponding to the call terminal in a call with the peer terminal; and receives a face video stream from the call terminal through the first video call media transmission channel, the face video stream including the face image of the user corresponding to the call terminal; then obtains a face recognition result, which is the result of face recognition of the user corresponding to the call terminal based on the face video stream; and then sends the face recognition result to the peer call terminal through the second video call media transmission channel to trigger the peer call terminal to process the call terminal's service request based on the face recognition result.

[0027] The communication method provided in this application embodiment allows the media server to receive the face video stream and send the face recognition result through the video call media transmission channel originally used for transmitting the call video stream when face recognition is required during a call. In this way, there is no need to spend additional bandwidth resources to establish a transmission channel dedicated to transmitting the face video stream, and there is no need to occupy additional port resources of the call terminal.

[0028] Furthermore, compared with existing communication methods, the technical solution provided in this application embodiment does not require the installation of a face recognition APP on the terminal. Thus, it does not require operators to perform complex related operations, does not require operators to have high operating skills, and does not require terminal calls.

[0029] In one possible implementation, after receiving a face recognition request from the peer terminal, the communication method provided in this application embodiment further includes: sending a Session Initiation Protocol (SIP) message to the peer terminal, the SIP message including a face recognition request identifier, the face recognition request identifier being used to request face recognition for the user corresponding to the peer terminal; and receiving a response message from the peer terminal to the SIP message, the response message indicating that the user corresponding to the peer terminal agrees to face recognition.

[0030] In this embodiment of the application, the face recognition request identifier can be carried in the header field of the SIP message, or, if the SIP message includes the SDP information of the media server, the face recognition request identifier can also be carried in the SDP information of the media server.

[0031] In one possible implementation, the face video stream is a video stream obtained by capturing images through the camera device of the call terminal; or, the face video stream is a video stream obtained from the storage device of the call terminal.

[0032] In one possible implementation, the communication method provided in this application embodiment further includes: sending source indication information to the calling terminal, wherein the source indication information instructs the calling terminal to acquire a face video stream through the camera device of the calling terminal or to acquire a face video stream from the storage device of the calling terminal.

[0033] In one possible implementation, the face video stream is a video stream obtained from the storage device of the calling terminal. Before receiving the face video stream from the calling terminal through the first video call media transmission channel, the communication method provided in this application embodiment further includes: sending a first transmission channel indication information to the calling terminal, wherein the first transmission channel indication information instructs the calling terminal to transmit the face video stream through the first video call media transmission channel.

[0034] In one possible implementation, before receiving the face video stream from the call terminal through the first video call media transmission channel, the communication method provided in this application embodiment further includes: stopping the transmission of the call video stream through the first video call media transmission channel.

[0035] In one possible implementation, before sending the face recognition result to the peer call terminal through the second video call media transmission channel, the communication method provided in this application embodiment further includes: sending a second transmission channel indication information to the peer call terminal, wherein the second transmission channel indication information instructs the peer call terminal to receive the face recognition result through the second video call media transmission channel.

[0036] In one possible implementation, before sending the face recognition result to the peer call terminal through the second video call media transmission channel, the communication method provided in this application embodiment further includes: stopping the transmission of the call video stream through the second video call media transmission channel.

[0037] In one possible implementation, after obtaining the face recognition result, the communication method provided in this application embodiment further includes: sending the face recognition result to the calling terminal.

[0038] In this embodiment, the face recognition result is the result of face recognition of the user corresponding to the call terminal based on the face video stream. The face recognition result may include information indicating successful face recognition or information indicating failure of face recognition. Optionally, when the face image extracted from the face video stream is consistent with the face image registered in the operator's system, the face recognition result may also include the identification information of the registered user corresponding to the call terminal. For example, the identification information of the registered user includes, but is not limited to, the user's name, identification document, and other identity information.

[0039] In one possible implementation, the communication method provided in this application embodiment further includes: sending posture indication information to the call terminal, wherein the posture indication information instructs the user corresponding to the call terminal to adjust the user's posture so that the face image meets preset conditions.

[0040] In one possible implementation, after receiving a face video stream from the calling terminal via the first video call media transmission channel, the communication method provided in this application embodiment further includes: extracting a target face image from the face video stream; and sending the target face image to a face recognition server to trigger the face recognition server to perform face recognition on the user corresponding to the calling terminal based on the target face image. Based on this, obtaining the face recognition result includes: receiving the face recognition result from the face recognition server. In this application embodiment, if the media server does not have face recognition functionality, the face recognition process is performed by a dedicated face recognition server. It is understood that this face recognition server is a server of the operator's system, and the face recognition server maintains the face images of users registered in the operator's system and other related information.

[0041] Optionally, the media server may also have a face recognition function. The media server maintains the face images of users registered in the operator's system and other related information. In this case, the media server extracts the target face image from the face video stream and performs face recognition on the user corresponding to the call terminal based on the target face image to obtain the face recognition result.

[0042] The relevant content and technical effects of the third aspect can be referenced from the content and technical effects described in any one of the first and second aspects and their possible implementation methods.

[0043] Fourthly, embodiments of this application provide a calling terminal, including a processing module, a receiving module, and a sending module. The processing module is used to establish a video call media transmission channel, which is used for transmitting a video stream between the calling terminal and a peer calling terminal in a video call service. The video stream includes video content captured by either the calling terminal or the peer calling terminal. The receiving module is used to receive a SIP message from a media server, the SIP message including a face recognition request identifier, which requests face recognition for the user corresponding to the calling terminal. The sending module is used to send a response message to the SIP message to the media server, the response message indicating that the user corresponding to the calling terminal agrees to face recognition. The sending module is also used to send a face video stream to the media server through the video call media transmission channel, the face video stream including a face image of the user corresponding to the calling terminal. The receiving module is also used to receive a face recognition result from the media server.

[0044] In one possible implementation, the face video stream is a video stream obtained by capturing images through the camera device of the call terminal; or, the face video stream is a video stream obtained from the storage device of the call terminal.

[0045] In one possible implementation, the receiving module is further configured to receive source indication information from the media server, the source indication information instructing the call terminal to acquire a face video stream through the call terminal's camera device or to acquire a face video stream from the call terminal's storage device.

[0046] In one possible implementation, the face video stream is a video stream obtained from the storage device of the call terminal, and the receiving module is further configured to receive transmission channel indication information from the media server, the transmission channel indication information instructing the call terminal to transmit the face video stream through the video call media transmission channel.

[0047] In one possible implementation, the processing module is further configured to control the receiving module or the sending module to stop transmitting the call video stream through the video call media transmission channel.

[0048] In one possible implementation, the receiving module is further configured to receive posture indication information from the media server, the posture indication information instructing the user corresponding to the call terminal to adjust the user's posture so that the face image meets preset conditions.

[0049] Fifthly, embodiments of this application provide a calling terminal, including a processing module, a sending module, and a receiving module. The processing module is used to establish a video call media transmission channel, which is used for transmitting a video stream between the calling terminal and a peer calling terminal in a video call service. The video stream includes video content captured by either the calling terminal or the peer calling terminal. The sending module is used to send a face recognition request to a media server. The face recognition request includes a face recognition request identifier, which is used to request face recognition of the user corresponding to the peer calling terminal in a call with the calling terminal. The receiving module is used to receive a face recognition result from the media server through the video call media transmission channel. The face recognition result is the result of face recognition of the user corresponding to the peer calling terminal based on the face video stream. The processing module is also used to process service requests from the peer calling terminal based on the face recognition result.

[0050] In one possible implementation, the face video stream is a video stream obtained by capturing images with the camera device of the peer communication terminal; or, the face video stream is a video stream obtained from the storage device of the peer communication terminal.

[0051] In one possible implementation, the receiving module is further configured to receive transmission channel indication information from the media server, the transmission channel indication information instructing the call terminal to receive the face recognition result through the video call media transmission channel.

[0052] In one possible implementation, the processing module is further configured to control the sending module or receiving module to stop transmitting the call video stream through the video call media transmission channel.

[0053] In one possible implementation, the face recognition result includes information indicating successful face recognition or information indicating failed face recognition. The processing module is further configured to process the service request of the peer call terminal when the face recognition result includes the information indicating successful face recognition.

[0054] Sixthly, this application provides a media server, including: a processing module, a receiving module, an acquiring module, and a sending module. The processing module is used to establish a first video call media transmission channel and a second video call media transmission channel. The first video call media transmission channel is a video call media transmission channel between a calling terminal and the media server, and the second video call media transmission channel is a video call media transmission channel between the media server and a peer calling terminal. The first and second video call media transmission channels are used for transmitting a video stream between the calling terminal and the peer calling terminal in a video call service. The video stream includes video content captured by the calling terminal or the peer calling terminal. The receiving module is used to receive a face recognition application from the peer calling terminal. Please include a face recognition application identifier, which is used to apply for face recognition of the user corresponding to the call terminal that is talking to the peer call terminal; and receive a face video stream from the call terminal through the first video call media transmission channel, the face video stream including the face image of the user corresponding to the call terminal; an acquisition module is used to acquire a face recognition result, the face recognition result being the result of face recognition of the user corresponding to the call terminal based on the face video stream; a sending module is used to send the face recognition result to the peer call terminal through the second video call media transmission channel to trigger the peer call terminal to process the call terminal's service request based on the face recognition result.

[0055] In one possible implementation, the sending module is further configured to send a Session Initiation Protocol (SIP) message to the calling terminal, the SIP message including a face recognition request identifier, the face recognition request identifier being used to request face recognition of the user corresponding to the calling terminal; the receiving module is further configured to receive a response message of the SIP message from the calling terminal, the response message of the SIP message indicating that the user corresponding to the calling terminal agrees to face recognition.

[0056] In one possible implementation, the face video stream is a video stream obtained by capturing images through the camera device of the call terminal; or, the face video stream is a video stream obtained from the storage device of the call terminal.

[0057] In one possible implementation, the sending module is further configured to send source indication information to the calling terminal, the source indication information instructing the calling terminal to acquire a face video stream through the camera device of the calling terminal or to acquire a face video stream from the storage device of the calling terminal.

[0058] In one possible implementation, the face video stream is a video stream obtained from the storage device of the call terminal, and the sending module is further configured to send a first transmission channel indication information to the call terminal, the first transmission channel indication information instructing the call terminal to transmit the face video stream through the first video call media transmission channel.

[0059] In one possible implementation, the processing module is further configured to control the sending module or receiving module to stop transmitting the call video stream through the first video call media transmission channel.

[0060] In one possible implementation, the sending module is further configured to send a second transmission channel indication information to the peer call terminal, the second transmission channel indication information instructing the peer call terminal to receive the face recognition result through the second video call media transmission channel.

[0061] In one possible implementation, the processing module is further configured to control the sending module or receiving module to stop transmitting the call video stream through the second video call media transmission channel.

[0062] In one possible implementation, the sending module is further configured to send the face recognition result to the calling terminal.

[0063] In one possible implementation, the sending module is further configured to send posture indication information to the calling terminal, the posture indication information instructing the user corresponding to the calling terminal to adjust the user's posture so that the face image meets preset conditions.

[0064] In one possible implementation, the processing module is further configured to extract a target face image from the face video stream; the sending module is further configured to send the target face image to a face recognition server to trigger the face recognition server to perform face recognition on the user corresponding to the call terminal based on the target face image; the acquisition module is specifically configured to receive the face recognition result from the face recognition server.

[0065] In a seventh aspect, embodiments of this application provide a call terminal, including a memory and at least one processor connected to the memory. The memory is used to store computer program code, which includes computer instructions. When the computer instructions are executed by the at least one processor, the call terminal causes the call terminal to perform the method described in the first aspect and any of its possible implementations.

[0066] Eighthly, embodiments of this application provide a call terminal, including a memory and at least one processor connected to the memory. The memory is used to store computer program code, which includes computer instructions. When the computer instructions are executed by the at least one processor, the call terminal causes the call terminal to perform the method described in the second aspect and any of its possible implementations.

[0067] Ninthly, embodiments of this application provide a media server, including a memory and at least one processor connected to the memory, the memory being used to store computer program code, the computer program code including computer instructions, which, when executed by the at least one processor, cause the media server to perform the method described in the third aspect and any of its possible implementations.

[0068] In a tenth aspect, embodiments of this application provide a computer-readable storage medium including computer instructions that, when executed on a calling terminal, cause the calling terminal to perform the method described in the first aspect and any of its possible implementations.

[0069] Eleventhly, embodiments of this application provide a computer-readable storage medium including computer instructions that, when executed on a calling terminal, cause the calling terminal to perform the method described in the second aspect and any of its possible implementations.

[0070] In a twelfth aspect, embodiments of this application provide a computer-readable storage medium including computer instructions that, when executed on a media server, cause the media server to perform the method described in any one of the third aspect and its possible implementations.

[0071] In a thirteenth aspect, embodiments of this application provide a computer program product that, when run on a computer, executes the method described in the first aspect and any of its possible implementations.

[0072] In a fourteenth aspect, embodiments of this application provide a computer program product that, when run on a computer, executes the method described in the second aspect and any of its possible implementations.

[0073] In a fifteenth aspect, embodiments of this application provide a computer program product that, when run on a computer, executes the method described in the third aspect and any of its possible implementations.

[0074] In a sixteenth aspect, embodiments of this application provide a chip including a memory and a processor. The memory stores computer instructions. The processor retrieves and executes the computer instructions from the memory, causing a telephony terminal to perform the method described in the first aspect and any of its possible implementations.

[0075] In a seventeenth aspect, embodiments of this application provide a chip including a memory and a processor. The memory stores computer instructions. The processor retrieves and executes the computer instructions from the memory, causing a telephony terminal to perform the method described in the second aspect and any of its possible implementations.

[0076] In an eighteenth aspect, embodiments of this application provide a chip including a memory and a processor. The memory stores computer instructions. The processor retrieves and executes the computer instructions from the memory, causing a media server to perform the methods described in the third aspect and any of its possible implementations.

[0077] In a nineteenth aspect, embodiments of this application provide a communication system including a calling terminal, a peer calling terminal, and a media server. The calling terminal executes the method described in the first aspect and any of its possible implementations; the peer calling terminal executes the method described in the second aspect and any of its possible implementations; and the media server executes the method described in the third aspect and any of its possible implementations.

[0078] It should be understood that the beneficial effects achieved by the above-mentioned fourth to nineteenth aspects of the technical solutions and their corresponding possible implementations can be referred to the above-mentioned technical effects of the first to third aspects and their corresponding possible implementations, and will not be repeated here. Attached Figure Description

[0079] Figure 1 This application provides an embodiment of a communication system architecture diagram for a customer service scenario using human agents.

[0080] Figure 2 A schematic diagram of a voice call process provided in an embodiment of this application;

[0081] Figure 3 A schematic diagram of a video call process provided in an embodiment of this application;

[0082] Figure 4A A hardware schematic diagram of a mobile phone provided for an embodiment of this application;

[0083] Figure 4B A schematic diagram of a mobile phone system architecture provided in this application embodiment;

[0084] Figure 5 A hardware schematic diagram of a server provided for an embodiment of this application;

[0085] Figure 6 This is one of the schematic diagrams of a communication method provided in an embodiment of this application;

[0086] Figure 7 This is a second schematic diagram of a communication method provided in an embodiment of this application;

[0087] Figure 8 This is a schematic diagram of the structure of a call terminal provided in an embodiment of this application;

[0088] Figure 9 This is a schematic diagram of another communication terminal provided in an embodiment of this application;

[0089] Figure 10 This is a schematic diagram of the structure of a call terminal provided in an embodiment of this application;

[0090] Figure 11 This is a schematic diagram of another communication terminal provided in an embodiment of this application;

[0091] Figure 12 This is a schematic diagram of the structure of a media server provided in an embodiment of this application;

[0092] Figure 13 This is a schematic diagram of another media server structure provided in an embodiment of this application. Detailed Implementation

[0093] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0094] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, rather than to describe a specific order of objects. For example, "first video call media transmission channel" and "second video call media transmission channel," etc., are used to distinguish different video call media transmission channels, rather than to describe a specific order of video call media transmission channels.

[0095] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0096] Currently, based on mobile networks, users can make voice or video calls through terminals. Taking a voice call as an example, user 1 can dial the number (e.g., a phone number) of user 2's terminal 2 through user 1's terminal 1. After user 2 answers through terminal 2, a call connection can be established between user 1's terminal 1 and user 2's terminal 2, so user 1 and user 2 can have a voice call. For example, terminal 1 collects user 1's voice and sends the collected voice to terminal 2, and terminal 2 collects user 2's voice and sends the collected voice to terminal 1.

[0097] Understandably, in a scenario where terminals are communicating, one of the terminals can be called the calling terminal, and the other terminal communicating with that calling terminal can be called the peer calling terminal. For example, if terminal 1 is the calling terminal, then terminal 2 is the peer calling terminal.

[0098] The following section uses two calling terminals (i.e., one calling terminal and one peer calling terminal) as an example to briefly introduce the transmission principle of media streams during voice and video calls between the calling terminal and the peer calling terminal.

[0099] When a terminal makes a voice call with another terminal, the media stream transmitted between them is an audio stream. After establishing a voice call media transmission channel (also referred to as an audio stream transmission channel) between the two terminals, they can transmit audio streams based on this channel. For example, the terminal sends the audio stream of its microphone, capturing the voice of its corresponding user, to the other terminal via the audio stream transmission channel. Conversely, the other terminal sends the audio stream of its microphone, capturing the voice of its corresponding user, to the terminal via the audio stream transmission channel.

[0100] When a calling terminal conducts a video call with a peer calling terminal, the media streams transmitted between them include voice and video streams, and the voice and video streams are transmitted through different channels. Specifically, after establishing a voice call media transmission channel (i.e., the voice stream transmission channel) and a video call media transmission channel (which can be simply referred to as the video stream transmission channel), the calling terminal and the peer calling terminal can transmit voice streams based on the voice stream transmission channel, and the calling terminal and the peer calling terminal can transmit video streams based on the video stream transmission channel. For example, the calling terminal sends the voice stream of the corresponding user captured by its microphone to the peer calling terminal through the voice stream transmission channel, and the peer calling terminal sends the voice stream of the corresponding peer user captured by its microphone to the calling terminal through the voice stream transmission channel; similarly, the calling terminal sends the video stream of the corresponding user captured by its camera to the peer calling terminal through the video stream transmission channel, and the peer calling terminal sends the video stream of the corresponding peer user captured by its camera to the calling terminal through the video stream transmission channel.

[0101] With the development of communication technology, user authentication is essential to ensure the security of user services during transactions. Authentication methods include password authentication, fingerprint authentication, and facial recognition authentication. Among these, facial recognition is becoming increasingly popular as a key means of verifying user identity. For example, when a user conducts a transaction offline (e.g., banking), the service provider's facial recognition equipment captures the user's facial image and verifies their identity based on this image. After successful authentication, the user can continue with the transaction. Similarly, when a user conducts a transaction online through an application (app) installed on their electronic device (e.g., a financial management app), the app uses the device's camera to capture the user's facial image and verifies their identity based on this image. After successful authentication, the user can continue with the online transaction.

[0102] In scenarios where two users are making voice or video calls, if one user needs to verify the other user's identity through facial recognition (e.g., a user holding the aforementioned peer-to-peer call terminal performs facial recognition on another user holding the same terminal), one implementation method is as follows: both the peer-to-peer call terminal and the call terminal need to have a facial recognition application (APP) installed. A communication connection is established between the facial recognition APP on the call terminal and the facial recognition APP on the peer-to-peer call terminal. During the facial recognition process, the call terminal uploads the captured facial video stream (including one or more frames of facial images) to the facial recognition APP on that call terminal. In the facial recognition APP, face detection, facial feature extraction, and facial feature comparison are performed based on the facial video stream (or the facial recognition APP interacts with a backend recognition server to achieve face detection, facial feature extraction, and facial feature comparison) to obtain the facial recognition result (the recognition result may include information indicating successful facial recognition or information indicating failure). Then, the peer-to-peer call terminal obtains the facial recognition result from the facial recognition APP on the call terminal based on the communication connection established between the APP on the peer-to-peer call terminal and the APP on the call terminal.

[0103] When the two users make a voice or video call and perform facial recognition, a dedicated transmission channel for transmitting facial video streams needs to be established between the call terminal and the facial recognition APP. This transmission channel is different from the voice stream transmission channel and video stream transmission channel mentioned above.

[0104] It should be understood that if facial recognition is performed between the two terminals during a call, additional bandwidth resources are required to establish a transmission channel for transmitting the facial video stream. Furthermore, either the calling terminal or the receiving terminal needs to use an additional port to send the facial video stream or receive the facial recognition result. For example, during a voice call, facial recognition is performed on the user corresponding to the calling terminal. The calling terminal and the receiving terminal transmit the voice stream via a dedicated voice transmission channel, while the calling terminal and the facial recognition app transmit the facial video stream via a dedicated transmission channel. Similarly, during a video call, facial recognition is performed on the user corresponding to the calling terminal. The calling terminal and the receiving terminal transmit the video stream captured by the camera via a dedicated video transmission channel, while the calling terminal and the facial recognition app transmit the facial video stream via a dedicated transmission channel. In summary, establishing a dedicated transmission channel for transmitting facial video streams requires additional bandwidth resources, and this facial recognition method requires additional port resources on the terminal.

[0105] Furthermore, during the initial call between the calling terminal and the peer terminal, when performing facial recognition, the calling terminal may not have the facial recognition app installed. In this case, the calling terminal may need to disconnect from the peer terminal, install the facial recognition app, complete the facial recognition process using the app, and then re-call the peer terminal. The peer terminal then obtains the facial recognition result from its app and, if successful, resumes the call. In other words, the facial recognition process may require interrupting the call before completing the facial recognition, which is cumbersome and inconvenient for the user.

[0106] To address the issue that establishing a dedicated transmission channel for transmitting facial video streams in existing technologies requires additional bandwidth and terminal port resources, this application provides a communication method, apparatus, and system. This communication method can be applied to facial recognition during calls between terminals. Specifically, the communication system establishes a video call media transmission channel between the calling terminal and a media server. This channel is used for transmitting call video streams between the calling terminal and the peer terminal in a video call service. These video streams contain video content captured by either the calling terminal or the peer terminal. Subsequently, after the peer terminal applies for facial recognition and the user corresponding to the calling terminal undergoes facial recognition, the calling terminal sends a facial video stream containing the facial image of the user corresponding to the calling terminal to the media server through the video call media transmission channel. The calling terminal then receives the facial recognition result from the media server. Through the technical solution provided by this application, during a call between the calling terminal and the peer terminal, the calling terminal can send the facial video stream to the media server through the existing video call media transmission channel, without consuming additional bandwidth resources to establish a dedicated transmission channel for facial video streams, and without occupying additional port resources of the calling terminal.

[0107] Furthermore, compared with existing communication methods, the technical solution provided in this application embodiment does not require the installation of a face recognition APP on the terminal. Thus, it does not require operators to perform complex related operations, does not require operators to have high operating skills, and does not require interruption of the call.

[0108] Optionally, the communication method provided in this application embodiment can be applied to video conferencing scenarios, customer service scenarios, etc. The customer service scenario refers to a call between a user and a customer service center (or customer service system). This call resolves the user's service needs; for example, customer service scenarios are involved in banking, insurance, securities, and mobile communication customer service. Generally, a user can call the customer service center and establish a connection. Currently, in practical applications, most customer service processes are as follows: the user calls the customer service center (i.e., makes a call), the customer service center answers, and then pushes (i.e. plays) some pre-stored prompts to the user. The user can then select the desired service option based on the prompts. The customer service center provides targeted service based on the selected option (e.g., responds to the user's selected option and resolves the user's question). Alternatively, the user can select human assistance based on the pushed prompts, in which case they will enter a human customer service scenario. The human customer service scenario refers to the situation where a user speaks with a staff member of the customer service center during a call. Specifically, after the user selects human service based on the prompts from the customer service center, the customer service center continues to call a staff member (specifically, through the number on the staff member's terminal). In the following embodiments, the customer service center can be simply referred to as customer service or the customer service system, and the staff member of the customer service center can be simply referred to as customer service personnel. Taking the human customer service scenario as an example, the user calls the customer service system through their own calling terminal (which can be simply referred to as the user's calling terminal). When the call is transferred to human service, the customer service personnel answer, and then the calling terminal held by the customer service personnel (which can be simply referred to as the customer service calling terminal) speaks with the user's calling terminal.

[0109] In this embodiment of the application, after the customer service terminal and the user's terminal start a call, the customer service personnel can request to perform facial recognition on the user, thereby executing the communication method provided in this embodiment of the application. After the facial recognition is successful, the customer service personnel will provide services based on the service requests made by the user.

[0110] In customer service scenarios, facial recognition can ensure the security of certain important services. Moreover, facial recognition can be completed during online calls, eliminating the need for users to visit a service center or meet with service personnel face-to-face to conduct business offline, thus providing efficient and convenient services to users.

[0111] It should be noted that this application embodiment describes the communication method provided in this application embodiment using a customer service scenario as an example. It is understood that in a customer service scenario, after a user's calling terminal initiates a call, the customer service system responds, and the user selects human assistance, the user's calling terminal and the customer service representative's terminal interact to perform facial recognition on the user.

[0112] Optionally, in this embodiment of the application, the call initiated by the user's calling terminal can be a voice call or a video call.

[0113] When a user initiates a voice call, after the customer service system responds, the user selects the manual service option based on the audio prompts pushed by the media server. Subsequently, the customer service terminal initiates a facial recognition request. Then, when the user on the user's terminal agrees to facial recognition, a video stream transmission channel is established between the user's terminal and the customer service terminal through media resource negotiation to convert the voice call into a video call. Then, a facial video stream containing the user's facial image is transmitted based on the video streaming media transmission channel corresponding to the video call.

[0114] When a user initiates a video call, after the customer service system responds, the user selects manual service based on the video prompts pushed by the media server. Subsequently, the customer service terminal initiates a facial recognition request. When the user on the call terminal agrees to facial recognition, a facial video stream containing the user's facial image can be transmitted through the video stream transmission channel corresponding to the video call.

[0115] The communication system in a customer service scenario can be viewed as a conference control system. This system involves the access network, the IP Multimedia Subsystem (IMS, including the 4G / 5G core network and the IMS core network), the customer service platform (also known as the customer service system), and business systems. The architecture of the communication system in a customer service scenario is described below. Figure 1 As shown, the communication system specifically includes: a user call terminal 101, an access network device 102, an IP multimedia subsystem 103, a customer service platform 104, a business system 105, and a customer service call terminal 106. The IP multimedia subsystem 103 includes a core network (which can be a 4G core network and / or a 5G core network) and an IMS core network. It should be understood that the 4G core network includes gateway devices (e.g., S-GW, P-GW), the 5G core network includes user plane functions (UPF), mobility management functions (AMF), etc., and the IMS core network includes a session border controller (SBC), a proxy-call session control function (P-CSCF), a call session control function (I-CSCF), and a serving call session control function (S-CSCF). The customer service platform 104 includes a media server.

[0116] SBC: Used to provide secure access and media processing.

[0117] P-CSCF: This is the entry node for user terminals to access the IMS core network, and it is mainly responsible for signaling and message brokering.

[0118] I-CSCF: It is the unified initial entry node of the IMS core network, responsible for assigning and querying S-CSCFs for user registration.

[0119] S-CSCF: It is the central node of the IMS core network, mainly used for user registration, authentication control, session routing and service triggering control, and maintaining session state information.

[0120] Media Server: In this embodiment, in the communication system corresponding to a traditional customer service scenario, the customer service platform 104 includes a control server (which can also be called a signaling server) and a media server. The signaling server is mainly responsible for signal negotiation and processing, and controlling the joining or leaving of calls by user calling terminals and customer service calling terminals. The media server is mainly responsible for audio and video processing and playback, application and release of call venues, audio encoding and decoding, video encoding and decoding, and facial recognition processing. In some implementations, the functions of the media server and the control server can be integrated into one server. In this embodiment, the communication method provided by this embodiment is described with the example that the functions of the media server and the control server are both integrated into the media server.

[0121] Business system: Responsible for determining and triggering different business processes based on the caller's (e.g., user's calling terminal) and the called number. Different businesses may include, but are not limited to, video calls, video advertisements, and enterprise video shows.

[0122] Combination Figure 1 The architecture of the communication system shown describes the process of a voice call, taking it as an example, based on the user's calling terminal accessing the access network and establishing a session through the 4G or 5G core network and the IMS core network, to facilitate understanding of the voice call process in a customer service scenario. (Reference) Figure 2 The voice call process includes:

[0123] S201. The user's voice terminal sends an invitation message to the media server through the IMS network element.

[0124] Specifically, in combination Figure 1The diagram illustrates the architecture of the communication system. IMS includes 4G / 5G core network elements (including gateway devices / user plane function elements), IMS core network SBC / P-CSCF elements, and I-CSCF / S-CSCF elements. In this embodiment, these network elements in IMS can be collectively referred to as IMS network elements. The user call terminal sending an invitation message to the media server through the IMS network elements specifically includes: the user call terminal according to... Figure 1 The architecture diagram shown sequentially sends the invitation message to the media server via the 4G / 5G core network elements, SBC / P-CSCF network elements, and I-CSCF / S-CSCF network elements. It should be noted that the IMS network element is used to transparently transmit messages between the user's calling terminal and the media server, but does not process the messages.

[0125] It should be noted that in the following embodiments, the messages or information sent or received through the IMS network element are similar to the invitation message transmitted through the IMS network element in S201. The IMS network element is used to transparently transmit messages or information, and will not be described one by one in the following embodiments.

[0126] It should be understood that after a user dials the customer service access code (which can be understood as the customer service system's phone number) through their user calling terminal, the user calling terminal executes the aforementioned S201. For example, the customer service representative could be from a telecommunications operator or an internet operator (e.g., customer service for banking services, customer service for insurance services, etc.), etc. This application does not limit the type of customer service representative.

[0127] It should be noted that the customer service call terminal in this embodiment refers to the call terminal corresponding to the customer service personnel in the customer service system, and this customer service call terminal is part of the customer service system. It can be understood that when a user calls the customer service system through their call terminal, after the customer service system answers, the media server in the customer service platform plays audio prompts related to the user's business to prompt the user to select the appropriate service according to their actual needs. If the user selects human service, the media server in the customer service system continues to call the customer service call terminal, as explained in the relevant steps of the following embodiments.

[0128] In this embodiment, the user's calling terminal sends the invitation message via the Session Initiation Protocol (SIP), which can also be understood as the invitation message being sent via SIP. The invitation message carries the user's calling terminal's Session Description Protocol (SDP) information, including the user's calling terminal's address information, audio port information, and audio codec format. This SDP information is used to negotiate media resources with the media server to establish a voice call media transmission channel between the user's calling terminal and the media server for transmitting the call audio stream. In this embodiment, the device's address information can be the device's IP address.

[0129] S202, The media server sends a ringing message to the user's calling terminal.

[0130] This ringing message indicates that the customer service call dialed by the user is being connected. At this time, the user's terminal is in a ringing state, waiting for a response from the customer service system (i.e., answering the call). This ringing message can be an 18* series message, such as 181 (call being forwarded, indicating the call is being forwarded) or 183 (indicating the progress of establishing the conversation). The ringing message carries the media server's SDP information, including the media server's IP address, audio port information, and audio codec format. This SDP information is used to negotiate media resources with the user's terminal to establish a voice call media transmission channel between the user's terminal and the media server for transmitting the call audio stream.

[0131] S203. The media server sends a response message to the user's call terminal through the IMS network element.

[0132] It should be understood that after the media server sends a ringing message to the user's calling terminal, the user's calling terminal waits for the customer service system to respond (i.e. waits for the call to be connected). During this process, the user can hear a "beep...beep..." waiting tone or a ringback tone. When the customer service system responds, the call is connected, and the media server executes the above-mentioned S203.

[0133] Similarly, the IMS network element is used to transparently transmit the response message.

[0134] In this embodiment, after the customer service system answers a call from a user's calling terminal, the media server can play audio prompts related to the user's services. Specifically, the media server sends the audio prompts to the user's calling terminal based on the established voice call media transmission channel. These audio prompts can prompt the user to select different service content according to their needs. For example, if the voice call is a call made by a user to a telecommunications operator, the audio prompts may include:

[0135] For phone bill and data usage inquiries, press "1"; for broadband services, press "2"; for recharge services, press "3"; for service inquiries and processing, press "4"; for password services, press "5"; for group services, press "6"; for customer service, press "0", etc. Optionally, the audio prompts may also include advertisements, promotions, etc. The audio prompts are related to the specific application scenario, and this application does not limit the audio prompt content.

[0136] When a user performs an operation following the prompts in the audio prompt and selects human assistance, the media server detects the selection of human assistance and assigns a customer service representative to the user (i.e., selects a corresponding customer service terminal for the user's calling terminal). Then, the media server executes the following S204.

[0137] S204. The media server sends an invitation message to the customer service call terminal.

[0138] This invitation message is used to initiate a voice call between the customer service terminal and the user's terminal. The message includes the media server's SDP information, which includes the media server's IP address, audio port information, and audio codec format. This SDP information is used to negotiate media resources with the customer service terminal to establish a voice call media transmission channel between the customer service terminal and the media server for transmitting the voice stream.

[0139] S205, The customer service call terminal sends a response message to the media server.

[0140] After the customer service terminal sends the response message, it joins the call with the user's terminal. The response message includes the customer service terminal's SDP information, which includes its IP address, audio port information, and audio codec format. This SDP information is used to negotiate media resources with the media server to establish a voice call media transmission channel between the customer service terminal and the media server for transmitting the call audio stream.

[0141] It should be understood that since the customer service call terminal is a new device in the customer service system that communicates with the user call terminal, in subsequent processes, to enable communication between the user call terminal and the customer service call terminal, media resource renegotiation is required. Specifically, the media server renegotiations media resources with the user call terminal (refer to S206-S207) and with the customer service call terminal (refer to S208-S209). Through this media resource renegotiation, a voice stream transmission channel (i.e., a voice call media transmission channel) can be established. It should be noted that the voice call media transmission channel established through S206-S209 is a transmission channel that requires the media server as an intermediary; that is, an indirect voice call media transmission channel. This voice call media transmission channel includes the voice call media transmission channel between the user call terminal and the media server, as well as the voice call media transmission channel between the media server and the customer service call terminal.

[0142] S206. The media server sends a reinvite message to the user's calling terminal through the IMS network element.

[0143] This re-invitation message is used to renegotiate media resources with the user's calling terminal in order to establish a voice call media transmission channel between the user's calling terminal and the media server. The re-invitation message includes the media server's SDP information, which includes the media server's IP address, audio port information, and audio codec format.

[0144] S207. The user's call terminal sends a response message to the media server through the IMS network element.

[0145] The response message includes the user's SDP information, which includes the user's IP address, audio port information, and audio codec format.

[0146] Through the media resource negotiation process described in S206-S207, the user's calling terminal can obtain the SDP information of the media server, and the media server can also obtain the SDP information of the user's calling terminal. In this way, a voice call media transmission channel is established between the user's calling terminal and the media server.

[0147] S208, The media server sends a reinvite message to the customer service call terminal.

[0148] This re-invitation message is used to renegotiate media resources with the customer service call terminal in order to establish a voice call media transmission channel between the customer service call terminal and the media server. The re-invitation message includes the media server's SDP information, which includes the media server's IP address, audio port information, and audio codec format.

[0149] S209. The customer service call terminal sends a response message to the media server.

[0150] The response message includes the SDP information of the customer service call terminal, which includes the terminal's IP address, audio port information, and audio codec format.

[0151] Through the media resource negotiation process described in S208-S209, the customer service call terminal can obtain the SDP information of the media server, and the media server can also obtain the SDP information of the customer service call terminal. In this way, a voice call media transmission channel is established between the customer service call terminal and the media server.

[0152] It should be understood that the voice call media transmission channels established through S206-S209 above (including the voice call media transmission channel between the user call terminal and the media server, and the voice call media transmission channel between the customer service call terminal and the media server) are used to transmit the call voice stream between the customer service call terminal and the user call terminal. For example, based on the established voice call media transmission channel, when the user call terminal sends a call voice stream to the customer service call terminal, the user call terminal sends the call voice stream to the media server based on the voice call media transmission channel between the user call terminal and the media server. Then, the media server sends the received call voice stream to the customer service call terminal based on the voice call media transmission channel between the media server and the customer service call terminal.

[0153] Optionally, in some cases, a direct voice call media transmission channel for transmitting voice streams between the user's calling terminal and the customer service calling terminal can be established through media resource negotiation. It should be noted that this direct voice call media transmission channel between the user's calling terminal and the customer service calling terminal does not require a media server as a relay device and is not necessarily a direct connection between the user's calling terminal and the customer service calling terminal. In this case, S206-S209 above can be replaced with S206'-S210'.

[0154] S206' The media server sends a reinvite message to the user's calling terminal through the IMS network element.

[0155] This re-invitation message is used to renegotiate media resources with the user's calling terminal. The re-invitation message includes the media server's SDP information, which includes the media server's IP address, audio port information, and audio codec format.

[0156] S207' The user's voice terminal sends a response message to the media server through the IMS network element.

[0157] The response message includes the user's SDP information, which includes the user's IP address, audio port information, and audio codec format.

[0158] S208', The media server sends a reinvite message to the customer service call terminal.

[0159] This re-invitation message is used to renegotiate media resources with the customer service call terminal. The re-invitation message includes the user call terminal's SDP information, which includes the user call terminal's IP address, audio port information, and audio codec format.

[0160] S209', The customer service call terminal sends a response message to the media server.

[0161] The response message includes the SDP information of the customer service call terminal, which includes the IP address, audio port information, and audio codec format of the customer service call terminal.

[0162] S210' The media server sends a response message carrying the SDP information of the customer service call terminal to the user's call terminal.

[0163] Through the media resource negotiation process S206'-S210, the user calling terminal can obtain the SDP information of the customer service calling terminal, and the customer service calling terminal can obtain the SDP information of the user calling terminal, thus establishing a voice call media transmission channel between the user calling terminal and the customer service calling terminal. Based on the established voice call media transmission channel, the user calling terminal and the customer service calling terminal can communicate directly without the need for a media server to forward the call voice stream. For example, the user calling terminal can directly send the call voice stream to the customer service calling terminal based on the voice call media transmission channel between the user calling terminal and the media server; similarly, the customer service calling terminal can also directly send the call voice stream to the user calling terminal based on the same voice call media transmission channel.

[0164] Combination Figure 1 The architecture of the communication system shown is based on the user's calling terminal accessing the access network, the 4G or 5G core network, and the IMS core network. Taking video calling as an example, the process of video calling is described to facilitate understanding of the video call process in a customer service scenario. This video call process is similar to the voice call process described above, and relevant content in the video call process can be referenced from the description of the voice call process. Figure 3 The video call process includes:

[0165] S301. The user's call terminal sends an invitation message to the media server through the IMS network element.

[0166] The invitation message carries the user's SDP information for the calling terminal. This SDP information includes the user's IP address, audio port information, audio codec format, video port information, and video codec format. This SDP information is used to negotiate media resources with the media server to establish a voice call media transmission channel for transmitting the call's audio stream and a video call media transmission channel for transmitting the call's video stream. It should be understood that video calls involve the transmission of both voice and video streams; therefore, compared to audio calls, the SDP information for video calls must also include video port information and video codec format.

[0167] S302, The media server sends a ringing message to the user's calling terminal.

[0168] The ringing message carries the SDP information of the media server, including the media server's IP address, audio port information, audio codec format, video port information, and video codec format. This SDP information is used to negotiate media resources with the user's calling terminal to establish a voice call media transmission channel for transmitting voice streams and a video call media transmission channel for transmitting video streams between the user's calling terminal and the media server.

[0169] S303, the media server sends a response message to the user's call terminal through the IMS network element.

[0170] In this embodiment of the application, when the user initiates a video call, after the customer service system responds to the call from the user's calling terminal, the media server can play video prompts related to the user's business. Specifically, the media server sends the video prompts to the user's calling terminal based on the established voice call media transmission channel and video call media transmission channel. The video prompts can prompt the user to select different service content according to their needs.

[0171] When a user performs an operation following the prompts in the video and selects human assistance, the media server detects the selection of human assistance and assigns a customer service representative to the user (i.e., selects a corresponding customer service terminal for the user's calling terminal). Then, the media server executes the following S304.

[0172] S304. The media server sends an invitation message to the customer service call terminal.

[0173] This invitation message is used to initiate a video session between the customer service terminal and the user's terminal. The message includes the media server's SDP information, which includes the media server's IP address, audio port information, audio codec format, video port information, and video codec format. This SDP information is used to negotiate media resources with the customer service terminal to establish a voice call media transmission channel for transmitting voice streams and a video call media transmission channel for transmitting video streams between the customer service terminal and the media server.

[0174] S305, The customer service call terminal sends a response message to the media server.

[0175] After the customer service terminal sends the response message, it joins the video call with the user's terminal. The response message includes the customer service terminal's SDP information, which includes its IP address, audio port information, audio codec format, video port information, and video codec format. This SDP information is used to negotiate media resources with the media server to establish a voice call media transmission channel for transmitting the voice stream and a video call media transmission channel for transmitting the video stream between the customer service terminal and the media server.

[0176] It should be understood that since the customer service call terminal is a new device in the customer service system that communicates with the user call terminal, in subsequent processes, to enable communication between the user call terminal and the customer service call terminal, media resource renegotiation is required. Specifically, the media server renegotiations media resources with the user call terminal (refer to S306-S307) and with the customer service call terminal (refer to S308-S309). Media resource renegotiation establishes voice and video call media transmission channels. It should be noted that the voice and video call media transmission channels established through S306-S309 require the media server as an intermediary; that is, indirect voice and video call media transmission channels. The voice call media transmission channel includes the voice call media transmission channel between the user call terminal and the media server, and the voice call media transmission channel between the media server and the customer service call terminal. Similarly, the video call media transmission channel includes the video call media transmission channel between the user call terminal and the media server, and the video call media transmission channel between the media server and the customer service call terminal.

[0177] S306. The media server sends a reinvite message to the user's calling terminal through the IMS network element.

[0178] This re-invitation message is used to renegotiate media resources with the user's calling terminal in order to establish a voice call media transmission channel between the user's calling terminal and the media server, as well as a video call media transmission channel between the user's calling terminal and the media server. The re-invitation message includes the media server's SDP information, which includes the media server's IP address, audio port information, audio codec format, video port information, and video codec format.

[0179] S307. The user's call terminal sends a response message to the media server through the IMS network element.

[0180] The response message includes the user's SDP information, which includes the user's IP address, audio port information, audio codec format, video port information, and video codec format.

[0181] Through the media resource negotiation process described in S306-S307, the user's calling terminal can obtain the SDP information of the media server, and the media server can also obtain the SDP information of the user's calling terminal. In this way, a voice call media transmission channel is established between the user's calling terminal and the media server, as well as a video call media transmission channel is established between the user's calling terminal and the media server.

[0182] S308, the media server sends a reinvite message to the customer service call terminal.

[0183] This re-invitation message is used to renegotiate media resources with the customer service call terminal in order to establish a voice call media transmission channel between the customer service call terminal and the media server, as well as a video call media transmission channel between the customer service call terminal and the media server. The re-invitation message includes the media server's SDP information, which includes the media server's IP address, audio port information, audio codec format, video port information, and video codec format.

[0184] S309, The customer service call terminal sends a response message to the media server.

[0185] The response message includes the SDP information of the customer service call terminal, which includes the terminal's IP address, audio port information, audio codec format, video port information, and video codec format.

[0186] Through the media resource negotiation process described in S308-S309, the customer service call terminal can obtain the SDP information of the media server, and the media server can also obtain the SDP information of the customer service call terminal. In this way, a voice call media transmission channel and a video call media transmission channel are established between the customer service call terminal and the media server.

[0187] It should be understood that the voice call media transmission channels established through S306-S309 above (including the voice call media transmission channel between the user call terminal and the media server, and the voice call media transmission channel between the customer service call terminal and the media server) are used to transmit the voice stream of the call between the customer service call terminal and the user call terminal; the video call media transmission channels established through S306-S309 above (including the video call media transmission channel between the user call terminal and the media server, and the video call media transmission channel between the customer service call terminal and the media server) are used to transmit the video stream of the call between the customer service call terminal and the user call terminal.

[0188] Similar to the voice call process, optionally, in some cases, a voice call media transmission channel for transmitting voice streams and a video call media transmission channel for transmitting video streams can be established between the user's call terminal and the customer service call terminal through media resource negotiation. In this case, S306-S309 above can be replaced by S306'-S310'.

[0189] S306' The media server sends a reinvite message to the user's calling terminal through the IMS network element.

[0190] This re-invitation message is used to renegotiate media resources with the user's calling terminal. The re-invitation message includes the media server's SDP information, which includes the media server's IP address, audio port information, audio codec format, video port information, and video codec format.

[0191] S307' The user's voice terminal sends a response message to the media server through the IMS network element.

[0192] The response message includes the user's SDP information, which includes the user's IP address, audio port information, audio codec format, video port information, and video codec format.

[0193] S308', the media server sends a reinvite message to the customer service call terminal.

[0194] This re-invitation message is used to renegotiate media resources with the customer service call terminal. The re-invitation message includes the user call terminal's SDP information, which includes the user call terminal's IP address, audio port information, audio codec format, video port information, and video codec format.

[0195] S309', The customer service call terminal sends a response message to the media server.

[0196] The response message includes the SDP information of the customer service call terminal, which includes the IP address, audio port information, audio codec format, video port information, and video codec format of the customer service call terminal.

[0197] S310' The media server sends a response message carrying the customer service call terminal's SDP information to the user's call terminal.

[0198] In summary, unlike the voice call process, all SDP information in this media negotiation process includes the device's video port information and video codec format.

[0199] Through the media resource negotiation process S306'-S310, the user calling terminal can obtain the SDP information of the customer service calling terminal, and the customer service calling terminal can obtain the SDP information of the user calling terminal. This establishes a direct voice and video call media transmission channel between the user calling terminal and the customer service calling terminal. Based on this established direct voice and video call media transmission channel, communication between the user calling terminal and the customer service calling terminal no longer requires a media server to forward the voice and video streams.

[0200] Optionally, the user call terminal is a call terminal and the customer service call terminal is a peer call terminal, or the customer service call terminal is a call terminal and the user call terminal is a peer call terminal. The specific choice depends on the actual situation and is not limited in this application embodiment.

[0201] In this embodiment, the aforementioned calling terminals (the calling terminal and the peer calling terminal) can be electronic devices such as mobile phones, tablets, or ultra-mobile personal computers (UMPCs). Alternatively, they can be other electronic devices such as desktop devices, laptop devices, handheld devices, wearable devices, smart home devices, and in-vehicle devices, such as netbooks, smartwatches, smart cameras, and personal digital assistants (PDAs). This embodiment does not limit the specific type and structure of the calling terminals.

[0202] Taking a mobile phone as the calling terminal as an example, Figure 4A This is a schematic diagram of the hardware structure of a mobile phone 400 provided in an embodiment of this application. The mobile phone 400 includes a processor 410, an external memory interface 420, an internal memory 421, a universal serial bus (USB) interface 430, a charging management module 440, a power management module 441, a battery 442, an antenna 1, an antenna 2, a mobile communication module 450, a wireless communication module 460, an audio module 470, a speaker 470A, a receiver 470B, a microphone 470C, a headphone jack 470D, a sensor module 480, buttons 490, a motor 491, an indicator 492, a camera 493, a display screen 494, and a subscriber identification module (SIM) card interface 495, etc.

[0203] It is understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the mobile phone 400. In other embodiments of this application, the mobile phone 400 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0204] Processor 410 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. The different processing units may be independent devices or integrated into one or more processors.

[0205] The controller can serve as the nerve center and command center of the mobile phone 400. The controller can generate operation control signals based on the instruction operation code and timing signals to control the fetching and execution of instructions.

[0206] The processor 410 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 410 is a cache memory. This memory can store instructions or data that the processor 410 has just used or that are used repeatedly. If the processor 410 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 410, and thus improves the efficiency of the system.

[0207] The charging management module 440 receives charging input from the charger. While charging the battery 442, the charging management module 440 can also supply power to electronic devices through the power management module 441.

[0208] The power management module 441 connects the battery 442, the charging management module 440, and the processor 410. The power management module 441 receives input from the battery 442 and / or the charging management module 440, providing power to the processor 410, internal memory 421, external memory, display screen 494, camera 493, and wireless communication module 460. The power management module 441 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 441 may also be located within the processor 410. In other embodiments, the power management module 441 and the charging management module 440 may be housed in the same device.

[0209] The wireless communication function of mobile phone 400 can be realized through antenna 1, antenna 2, mobile communication module 450, wireless communication module 460, modem processor and baseband processor.

[0210] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals.

[0211] The mobile communication module 450 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on the mobile phone 400. The mobile communication module 450 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 450 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 450 can be housed in the processor 410. In some embodiments, at least some functional modules of the mobile communication module 450 and at least some modules of the processor 410 can be housed in the same device.

[0212] The wireless communication module 460 can provide solutions for wireless communication applications on the mobile phone 400, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 460 can be one or more devices integrating at least one communication processing module. The wireless communication module 460 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 410. The wireless communication module 460 can also receive signals to be transmitted from processor 410, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0213] In some embodiments, the antenna 1 of the mobile phone 400 is coupled to the mobile communication module 450, and the antenna 2 is coupled to the wireless communication module 460, enabling the mobile phone 400 to communicate with the network and other devices through wireless communication technology.

[0214] The mobile phone 400 implements display functions through a GPU, a display screen 494, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 494 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 410 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0215] Display screen 494 is used to display images, videos, etc. Display screen 494 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, mobile phone 400 may include one or N displays 494, where N is a positive integer greater than 1.

[0216] The mobile phone 400 can achieve shooting functions through ISP, camera 493, video codec, GPU, display 494 and application processor.

[0217] The ISP is used to process the data fed back by the camera 493, which is used to capture still images or videos.

[0218] Digital signal processors are used to process digital signals. In addition to processing digital image signals, they can also process other digital signals (such as audio signals).

[0219] Video codecs are used to compress or decompress digital video. The mobile phone 400 can support one or more video codecs. Thus, the mobile phone 400 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0220] The external storage interface 420 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the mobile phone 400. The external storage card communicates with the processor 410 through the external storage interface 420 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.

[0221] Internal memory 421 can be used to store computer executable program code, which includes instructions. Processor 410 executes various functional applications and data processing of mobile phone 400 by running the instructions stored in internal memory 421. Internal memory 421 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of mobile phone 400 (such as audio data, phonebook, etc.). Furthermore, internal memory 421 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0222] The mobile phone 400 can achieve audio functions such as music playback and recording through an audio module 470, a speaker 470A, a receiver 470B, a microphone 470C, a headphone jack 470D, and an application processor.

[0223] The audio module 470 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. The audio module 470 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 470 may be located in the processor 410, or some functional modules of the audio module 470 may be located in the processor 410.

[0224] The speaker 470A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. Mobile phones 400 can use the speaker 470A to listen to music or make hands-free calls.

[0225] The receiver 470B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the mobile phone 400 answers a call or voice message, the receiver 470B can be brought close to the user's ear to listen to the voice.

[0226] Microphone 470C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 470C, inputting the sound signal into microphone 470C. Mobile phone 400 may have at least one microphone 470C. In some embodiments, mobile phone 400 may have two microphones 470C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, mobile phone 400 may also have three, four, or more microphones 470C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0227] The headphone jack 470D is used to connect wired headphones.

[0228] Keypad 490 includes a power button, volume buttons, etc. Mobile phone 400 can receive keypad input and generate key signal inputs related to user settings and function control.

[0229] Motor 491 can generate vibration alerts. Motor 491 can be used for incoming call vibration alerts or for touch vibration feedback.

[0230] Indicator 492 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0231] The SIM card interface 495 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 495 to make contact with or separate from the mobile phone 400.

[0232] It is understood that in the embodiments of this application, the mobile phone 400 described above can execute some or all of the steps in the embodiments of this application. These steps or operations are merely examples, and the mobile phone 400 can also perform other operations or variations thereof. Furthermore, the various steps can be executed in different orders as presented in the embodiments of this application, and it is not necessary to execute all the operations in the embodiments of this application. The embodiments of this application can be implemented individually or in any combination, and this application does not limit this.

[0233] The communication method provided in this application embodiment can be applied to applications with, for example, Figure 4A The hardware structure shown is for a calling terminal or a calling terminal with a similar structure. Alternatively, it can be applied to calling terminals with other structures, but this application does not limit the scope of the embodiments.

[0234] After introducing the hardware structure of the calling terminal, this application takes a mobile phone 400 as an example to describe the system architecture of the calling terminal provided in this application. The system architecture of the mobile phone 400 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application's embodiment uses a layered architecture... Taking the system as an example, the software structure of mobile phone 400 is illustrated. Figure 4B This is a software structure block diagram of the call terminal according to an embodiment of this application.

[0235] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0236] The application layer can include a series of application packages, such as Figure 4B As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0237] In this embodiment, the calling application (i.e., the calling APP) in the application layer of the mobile phone 400 can be used to make voice or video calls with other calling terminals. This calling application is an application that is already present in the mobile phone 400 at the factory, and no installation or configuration is required by the user.

[0238] It should be understood that in the communication method provided in this application embodiment, the face recognition function implemented by the calling terminal and the peer calling terminal during voice or video calls is based on the calling application on the calling terminal and the peer calling terminal. Alternatively, the calling terminal and the peer calling terminal in this application embodiment can be considered specifically as the calling terminal or the calling application on the peer calling terminal.

[0239] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0240] like Figure 4B As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0241] The window manager manages window programs. It can obtain screen size, determine the presence of a status bar, lock the screen, and capture screenshots. The content provider stores and retrieves data, making it accessible to applications. This data can include videos, images, audio, made and received calls, browsing history and bookmarks, phone books, etc. The view system includes visual controls, such as controls for displaying text and controls for displaying images. The view system can be used to build applications. The display interface can consist of one or more views. For example, a display interface including a text notification icon can include views for displaying text and views for displaying images. The phone manager provides communication functions for the calling terminal, such as managing call status (including connection, hang-up, etc.). The resource manager provides various resources to the application, such as localized strings, icons, images, layout files, video files, etc. The notification manager allows applications to display notification information in the status bar. It can be used to convey informational messages and can disappear automatically after a short pause without user interaction. For example, the notification manager is used to notify of download completion, message alerts, etc. The notification manager can also display notifications as icons or scrolling text in the system's top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting alert sounds, causing electronic devices to vibrate, and flashing indicator lights.

[0242] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0243] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0244] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0245] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0246] The Surface Manager manages the display subsystem and provides fusion of 2D and 3D layers for multiple applications. The Media Library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG. The 3D Graphics Processing Library implements 3D graphics drawing, image rendering, compositing, and layer processing. The 2D Graphics Engine is the drawing engine for 2D graphics.

[0247] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0248] The following example illustrates the workflow of the mobile phone's 400 software and hardware in the context of capturing and photographing scenes.

[0249] When the phone's 400 touch sensor receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, touch operation timestamp, etc.). The raw input event is stored in the kernel layer. The application framework layer retrieves the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking a single touch operation as an example, where the corresponding control is the camera application icon, the camera application calls the application framework layer's interface to launch the camera application, and then calls the kernel layer to launch the camera driver, capturing still images or videos through the camera 493.

[0250] In this embodiment of the application, the media server in the above-described communication system can be a hardware server or a software server. Taking a hardware server as an example, such as... Figure 5 As shown, this application embodiment provides a media server 500, which includes at least one processor 501 and a memory 502.

[0251] The processor 501 includes one or more central processing units (CPUs). The CPU can be a single-core CPU or a multi-core CPU.

[0252] The memory 502 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or optical memory. The memory 302 stores the operating system code.

[0253] Optionally, the processor 501 implements the method in the above embodiments by reading instructions stored in the memory 502, or the processor 501 implements the method in the above embodiments by internally stored instructions. When the processor 501 implements the method in the above embodiments by reading instructions stored in the memory 502, the memory 502 stores instructions for implementing the communication method provided in the embodiments of this application.

[0254] After the program code stored in memory 502 is read by at least one processor 501, the media server 500 performs the following operations: triggers the establishment of a video call media transmission channel, which is used to transmit a video stream containing video content captured by the call terminal; and receives a face recognition application, receives a face video stream through the video call media transmission channel, and returns the face recognition result to the call terminal and the other end of the call terminal.

[0255] Optionally, Figure 5 The media server 500 shown also includes a network interface 503. Network interface 503 is a wired interface, such as a fiber distributed data interface (FDDI) or a gigabit Ethernet (GE) interface. Alternatively, network interface 503 is a wireless interface. Network interface 503 is used to receive messages (e.g., SIP messages). Alternatively, network interface 503 is used to receive video or audio streams from a call.

[0256] The memory 502 is used to store the audio or video streams received by the network interface 503. At least one processor 501 further executes the method described in the above method embodiments based on the information stored in the memory 502. For more details on how the processor 501 implements the above functions, please refer to the descriptions in the preceding method embodiments, which will not be repeated here.

[0257] Optionally, the media server 500 also includes a bus 504, through which the processor 501 and memory 502 are typically interconnected, or in other ways.

[0258] Optionally, the media server 500 also includes an input / output interface 505, which is used to connect to an input device and receive instructions input by the user through the input device. Input devices include, but are not limited to, keyboards, touchscreens, microphones, etc. The input / output interface 505 is also used to connect to an output device to output the processing results of the processor 501. Output devices include, but are not limited to, monitors, printers, etc.

[0259] Based on the description of the above embodiments, this application provides a communication method that can be applied to a communication system having the above-described... Figure 4A The hardware structure shown and the above Figure 4B The system architecture shown includes a calling terminal (including the calling terminal and the peer calling terminal) with the above-mentioned features. Figure 5 The communication method is implemented in the media server with the hardware structure shown, through the interaction of various devices.

[0260] In a customer service scenario, the device that initiates the call is the user's calling terminal, and the device being called is the customer service system's terminal. In this embodiment, facial recognition needs to be performed on the user corresponding to the user's calling terminal. That is, the user's calling terminal corresponds to the calling terminal in the following embodiment, and the customer service system's terminal (e.g., the customer service calling terminal) is the peer calling terminal.

[0261] As can be seen from the above embodiments, in the customer service scenario, the face recognition process occurs during the manual service stage, that is, when the customer service personnel participate in the call (i.e., the customer service call terminal held by the customer service personnel participates in the call).

[0262] The communication method provided in the embodiments of this application will be described in detail below, such as Figure 6 As shown, the communication method provided in this application embodiment includes:

[0263] S601. Establish a video call media transmission channel.

[0264] In this embodiment, when a customer service representative participates in a call, the aforementioned video call media transmission channel is established through interaction between the call terminal, the media server, and the peer call terminal. This video call media transmission channel is used for transmitting the call video stream between the call terminal and the peer call terminal in a video call service. The call video stream includes video content captured by either the call terminal or the peer call terminal. It should be understood that in this embodiment, the call terminal and the peer call terminal are relative concepts. Of the two terminals participating in the call, if either terminal can be the call terminal, then the other terminal is the peer call terminal.

[0265] It should be noted that, in this embodiment of the application, the video call media transmission channel established by the aforementioned call terminal, the peer call terminal, and the media server is an indirect video call media transmission channel. The establishment of the video call media transmission channel specifically includes: establishing a first video call media transmission channel and establishing a second video call media transmission channel. The first video call media transmission channel is the video call media transmission channel between the call terminal and the media server, and the second video call media transmission channel is the video call media transmission channel between the peer call terminal and the media server.

[0266] In this embodiment, the process of establishing the video call media transmission channel can be referred to the description of S206-S209 in the above embodiment, and will not be repeated here.

[0267] S602. The peer terminal sends a face recognition request to the media server. Correspondingly, the media server receives the face recognition request sent by the peer terminal.

[0268] The facial recognition application includes a facial recognition application identifier, which is used to apply for facial recognition of the user corresponding to the call terminal, that is, to perform facial recognition on the call terminal during the call between the call terminal and the other end of the call terminal.

[0269] Similarly, alternatively, if the caller (i.e., the calling terminal) initiates a video call, after the calling terminal selects human service according to the video prompts, the media server in the customer service system calls the other calling terminal, and after the other calling terminal responds, the calling terminal, the other calling terminal, and the media server interact to establish a video call media transmission channel.

[0270] Optionally, if the caller initiates a voice call, after the call terminal selects human service according to the audio prompts, the media server in the customer service system calls the other end of the call terminal. After the other end of the call terminal responds, the call terminal, the other end of the call terminal, and the media server interact to establish a voice call media transmission channel. After the media server receives the face recognition request, the call terminal, the other end of the call terminal, and the media server interact to establish a video call media transmission channel.

[0271] S603. The media server sends a SIP message to the calling terminal. This SIP message includes a face recognition request identifier, which is used to request face recognition for the user corresponding to the calling terminal. Correspondingly, the calling terminal receives the SIP message from the media server.

[0272] S604. The calling terminal sends a SIP message response message to the media server, which indicates that the user corresponding to the calling terminal agrees to perform face recognition.

[0273] Specifically, the SIP message response includes a face recognition response identifier, which indicates that the user corresponding to the calling terminal agrees to face recognition.

[0274] In this embodiment of the application, the face recognition request identifier in the SIP message can be carried in the header field of the SIP message.

[0275] Optionally, the face recognition request identifier can be carried in the header field of the SIP message in the following two ways.

[0276] The first method of carrying the face recognition request identifier (denoted as FR) is to carry it in the Contact extension field of the SIP message.

[0277] Taking INVITE sip:02033296999@gd.ctcims.cn SIP / 2.0 as an example,

[0278] The Contact header field is:

[0279] <sip:172.27.10.10:5060;transport=udp;zte-did=26-3-20481-3629-12-890-3302;zte-uid=200001+861892222222;Hpt=8e48_16;CxtId=4;TRC=ffffffff-ffffffff> ; audio; video; FR; +g.3gpp.mid-call; +g.3gpp.srvcc-alerting; +g.3gpp.ps2cs-srvcc-orig-pre-alerting; +g.3gpp.icsi-ref="urn%3Aurn-7%3A3gpp-service.ims.icsi.mmtel";

[0280] Max-Forwards: 64.

[0281] The second method is to carry the Face Recognition Request Identifier (FR) in the Supported extension field of the SIP message.

[0282] Supported:100rel,histinfo,precondition,timer,FR.

[0283] Optionally, when the call initiated by the caller is a voice call, the SIP message may also include the SDP information of the media server, which includes the address information (e.g., IP address), audio port information, audio codec format, video port information, and video codec format of the media server.

[0284] When a SIP message includes SDP information of a media server, the response message of that SIP message also includes SDP information of the calling terminal. The SDP information of the calling terminal includes the address information (e.g., IP address), audio port information, audio codec format, video port information, and video codec format of the calling terminal.

[0285] In this embodiment, when the call initiated by the caller is a voice call, after the media server receives the face recognition request sent by the peer call terminal, the media server and the call terminal can negotiate media resources based on the SDP information of the media server and the call terminal in the SIP message and the response message of the SIP message in S603-S604 to establish a video call media transmission channel (i.e., the first video call media transmission channel). In this case, during the process of the call terminal, the peer call terminal, and the media server interacting to establish the video call media transmission channel in S601, the step of media resource negotiation between the media server and the call terminal can be replaced by steps S603-S604.

[0286] Optionally, the media server interacts with the peer call terminal (for example, the media server sends a SIP message to the peer call terminal, the SIP message including the media server's SDP information, and the peer call terminal sends a response message to the media server, the response message including the peer call terminal's SDP information) to negotiate media resources and establish a video call media transmission channel (i.e., the second video call media transmission channel).

[0287] Optionally, if the SIP message includes the SDP information of the media server, the aforementioned face recognition request identifier can also be carried in the SDP information of the media server.

[0288] When the Face Recognition Request Identifier (FR) is included in the SDP information, the video port information for transmitting the face video stream can be further specified in the extended fields of the SDP information. Specifically, it can be specified to use the same video port used for transmitting call video streams. The fields of the SDP information are illustrated below.

[0289] a = sendrecv; indicates a two-way video call.

[0290] a = sendonly / sendrecv; indicates a one-way video call.

[0291] a = FR; Face recognition request identifier

[0292] m = video 12082RTP / AVP 114 113; indicates that the video port used for transmitting call video streams is used to transmit face video streams.

[0293] v = 0

[0294] o=HuaWeiUAP9600 12 12IN IP4 10.137.2.167

[0295] s = Sip Call

[0296] c = IN IP4 10.137.2.176 / / IP address

[0297] t=0 0

[0298] m = audio 12080RTP / AVP 104 103 102 101 8 0 18 96 97 / / Audio port number

[0299] b = AS:41

[0300] b = RS:600

[0301] b = RR:2000

[0302] a = rtpmap:104AMR-WB / 16000 / 1 / / Audio encoding / decoding

[0303] a=fmtp:104mode-change-capability=2; max-red=0

[0304] a = rtpmap:103AMR-WB / 16000 / 1

[0305] a=fmtp:103octet-align=1; mode-change-capability=2; max-red=0

[0306] a = rtpmap:102AMR / 8000 / 1

[0307] a=fmtp:102mode-change-capability=2; max-red=0

[0308] a = rtpmap:101AMR / 8000 / 1

[0309] a = fmtp:101 octet-align=1; mode-change-capability=2; max-red=0

[0310] a = rtpmap:96 telephone-event / 16000

[0311] a = fmtp:96 0-15

[0312] a = rtpmap:97 telephone-event / 8000

[0313] a = fmtp:97 0-15

[0314] a = curr:qos local none

[0315] a = curr:qos remote none

[0316] a = des:qos mandatory local sendrecv

[0317] a = des:qos optional remote sendrecv

[0318] a = sendrecv

[0319] a = maxptime:240

[0320] a = ptime:20

[0321] m = video 12082 RTP / AVP 114 113 / / Video port number

[0322] b = AS:2154

[0323] b = RS:8000

[0324] b = RR:6000

[0325] a = rtpmap:114 H264 / 90000 / / Video codec

[0326] a = fmtp:114

[0327] profile-level-id = 42C01F; sprop-parameter-sets = Z0LAH9oC0ChoBtChNQ==,aM4G4g==; packetization-mode = 1; sar-understood = 16; sar-supported = 1

[0328] a=imageattr:114send[x=720,y=1280]recv[x=720,y=1280]

[0329] a=rtpmap:113H264 / 90000

[0330] a=fmtp:113

[0331] profile-level-id=42C01F;sprop-parameter-sets=Z0LAH9oC0ChoBtChNQ==,aM4G4g==;packetizat ion-mode=0;sar-understood=16;sar-supported=1

[0332] a=imageattr:113send[x=720,y=1280]recv[x=720,y=1280]

[0333] a=curr:qos local none

[0334] a=curr:qos remote none

[0335] a=des:qos mandatory local sendrecv

[0336] a=des:qos optional remote sendrecv

[0337] a=rtcp-fb:*nack

[0338] a=rtcp-fb:*nack pli

[0339] a=rtcp-fb:*ccm fir

[0340] a=rtcp-fb:*ccm tmmbr

[0341] a=sendrecv

[0342] a=FR

[0343] a=tcap:1RTP / AVPF

[0344] a=pcfg:1t=1

[0345] a=extmap:2urn:3gpp:video-orientation.

[0346] Understandably, based on the description of SDP information, it can indicate whether a video call is one-way or two-way. Taking the calling terminal and the peer terminal as an example, a one-way video call may only transmit the video stream of the calling terminal, without transmitting the video stream of the peer terminal. For example, the calling terminal sends the video content it has captured to the peer terminal, and the peer terminal displays the video content captured by the calling terminal. However, if the peer terminal does not capture video content or does not send its captured video content to the calling terminal, then the video content captured by the peer terminal is not displayed on the calling terminal.

[0347] S605, the media server sends a first transmission channel indication message to the calling terminal. This first transmission channel indication message instructs the calling terminal to transmit a face video stream through a video call media transmission channel (specifically, the aforementioned first video call media transmission channel). Correspondingly, the calling terminal receives the first transmission channel indication message from the media server.

[0348] The face video stream contains the face image of the user corresponding to the call terminal. Optionally, the face video stream is a video stream obtained by the camera device of the call terminal; or, the face video stream is a video stream obtained from the storage device of the call terminal.

[0349] Understandably, since facial video streams can be obtained through different means, the communication method provided in this application embodiment further includes the following: the media server sends source indication information to the call terminal, which instructs the call terminal to obtain the facial video stream through the call terminal's camera device or from the call terminal's storage device. Correspondingly, the call terminal receives the source indication information from the media server and obtains the corresponding facial video stream according to the instructions of the source indication information.

[0350] In one implementation, the first transmission channel indication information can be carried in a SIP message. This SIP message can be the same as the SIP message in S603, or it can be a different SIP message. This application embodiment does not limit this.

[0351] The media server can explicitly instruct the transmission of a face video stream through the video call media transmission channel using the method described in S605 (i.e., sending transmission channel instruction information). In some cases, the media server can also implicitly instruct the transmission of a face video stream through the video call media transmission channel. For example, the SIP message in S603 carries the SDP information of the media server, and the response message of the SIP message carries the SDP information of the calling terminal, to negotiate (or instruct) the use of the video call media transmission channel established based on this pair of SDP information (the SDP information of the media server and the SDP information of the calling terminal) to transmit the face video stream. It should be understood that this video call media transmission channel was originally intended for transmitting call video streams.

[0352] S606. The calling terminal sends a facial video stream to the media server through the video call media transmission channel. Correspondingly, the media server also receives the facial video stream from the calling terminal through the same video call media transmission channel.

[0353] S607, The media server obtains the face recognition results.

[0354] In this embodiment, the face recognition result is the result of face recognition of the user corresponding to the call terminal based on the face video stream. The face recognition result may include information indicating successful face recognition or information indicating failure of face recognition. Optionally, when the face image extracted from the face video stream is consistent with the face image registered in the operator's system, the face recognition result may also include the identification information of the registered user corresponding to the call terminal. For example, the identification information of the registered user includes, but is not limited to, the user's name, identification document, and other identity information.

[0355] In one implementation, if the media server does not have facial recognition functionality, the facial recognition process is performed by a dedicated facial recognition server. After the media server receives the facial video stream from the calling terminal through the first video call media transmission channel, the communication method provided in this embodiment further includes: the media server extracting a target facial image from the facial video stream; and sending the target facial image to the facial recognition server to trigger the facial recognition server to perform facial recognition on the user corresponding to the calling terminal based on the target facial image. It is understood that this facial recognition server is a server of the operator's system, and the facial recognition server maintains facial images and other relevant information of users registered in the operator's system. Detailed information regarding the facial recognition server performing facial recognition based on the target facial image can be found in the relevant descriptions of the above embodiments or in the relevant content of the prior art, and will not be repeated here.

[0356] In this case, the media server obtaining the face recognition result specifically includes: the media server receiving the face recognition result from the face recognition server.

[0357] In another implementation, the media server can also have a face recognition function. The media server maintains the face images of users registered in the operator's system and other related information. In this case, the media server extracts the target face image from the face video stream and performs face recognition on the user corresponding to the call terminal based on the target face image to obtain the face recognition result.

[0358] S608, the media server sends the face recognition result to the calling terminal. Correspondingly, the calling terminal receives the face recognition result sent by the media server.

[0359] It is understood that the transmission channel between the media server and the call terminal may include the aforementioned video call media transmission channel and signaling channel. Optionally, the media server may send the face recognition result to the call terminal through the aforementioned video call media transmission channel (i.e., the first video call media transmission channel) or through the signaling channel. This application embodiment does not limit this.

[0360] S609. The media server sends a second transmission channel indication message to the peer call terminal. Correspondingly, the peer call terminal receives the second transmission channel indication message from the media server. The second transmission channel indication message instructs the peer call terminal to receive the face recognition result through the second video call media transmission channel.

[0361] S610, the media server sends the face recognition result to the peer terminal via the video call media transmission channel (specifically, the second video call media transmission channel mentioned above). Correspondingly, the peer terminal receives the face recognition result from the media server via the same video call media transmission channel.

[0362] Of course, the media server can also send the face recognition results to the peer terminal through the signaling channel between the media server and the peer terminal.

[0363] S611. The peer terminal processes the call terminal's service requests based on the face recognition results.

[0364] As described in the above embodiments, the calling terminal can be a user calling terminal, and the other end calling terminal can be a customer service calling terminal. The above S601-S611 process is the process of performing face recognition on the user corresponding to the user calling terminal during the call between the customer service calling terminal and the user calling terminal. If the face recognition result received by the customer service calling terminal includes information indicating that the face recognition was successful, the customer service calling terminal determines that the user holding the user calling terminal is a legitimate user and is consistent with the registered user of the user calling terminal. Then, the customer service calling terminal begins to process the business request of the calling terminal, thereby completing the subsequent service. In this way, the secure processing of the user's business can be guaranteed.

[0365] Optionally, if facial recognition is successful, the interaction between the calling terminal and the peer calling terminal is restored to a video call or voice call. The calling terminal, media server, and peer calling terminal can transmit the call video stream through the video call media transmission channel, or the three parties can renegotiate media resources to establish a voice call media transmission channel and transmit the call voice stream through the voice call media transmission channel.

[0366] Optionally, when the face video stream is a video stream obtained from the storage device of the call terminal, the communication method provided in this application embodiment further includes S612 before transmitting the face video stream to the media server through the video call media transmission channel (i.e., S606).

[0367] S612, The call terminal and media server stop transmitting call video streams through the video call media transmission channel.

[0368] In this embodiment of the application, the transmission of call video stream (content captured by the call terminal or the other end of the call terminal) between the call terminal and the media server through the video call media transmission channel is stopped. In this way, the face video stream obtained from the storage device of the call terminal can be transmitted through the video call media transmission channel.

[0369] It should be noted that when the face video stream transmitted by the call terminal is captured by the camera device of the call terminal, the aforementioned video call media transmission channel does not actually stop transmitting the call video stream. That is, the face video stream is the call video stream. The difference is that the content of the call video stream may change from other video content to video content including face images.

[0370] Optionally, before receiving the face recognition result from the media server via the video call media transmission channel, the communication method provided in this application embodiment further includes S613.

[0371] S613, The media server and the peer terminal stop transmitting the call video stream through the video call media transmission channel.

[0372] In this embodiment, the media server and the peer call terminal stop transmitting call video streams (content captured by the call terminal or the peer call terminal) through the video call media transmission channel. In this way, the face recognition results can be transmitted through the video call media transmission channel.

[0373] In one implementation, when the face video stream is captured by the camera device of the call terminal, the call terminal captures the user's face image in real time and transmits the face image to the media server. Understandably, when the camera device of the call terminal captures the face image, the captured face image may not meet the preset conditions due to the influence of the user's head posture. The preset conditions may include, but are not limited to, the position and angle of the user's head within a certain range (for example, the face should be located within a preset detection frame and the user's face should be facing the camera device). This may cause face recognition to fail or result in inaccurate face recognition results. Based on this, after receiving the face video stream sent by the calling terminal, if the media server determines that the face image in the face video stream does not meet the preset conditions, the media server can send posture indication information to the calling terminal (correspondingly, the calling terminal receives the posture indication information from the media server). This posture indication information instructs the user corresponding to the calling terminal to adjust the user's posture so that the face image meets the preset conditions. For example, when the user's face is not entirely within the preset detection frame, the posture indication information prompts the user to place the entire face within the detection frame. Or, when the angle of the user's face is not appropriate, the posture indication information prompts the user to turn their face in a certain direction (for example, prompting the user to turn their head to the right).

[0374] In summary, the communication method provided in this application allows the call terminal to transmit the face video stream through the video call media transmission channel originally used for transmitting the call video stream when face recognition is required during a call. This eliminates the need to spend additional bandwidth resources to establish a dedicated transmission channel for transmitting the face video stream and avoids occupying additional port resources of the call terminal.

[0375] Furthermore, compared with existing communication methods, the technical solution provided in this application embodiment does not require the installation of a face recognition APP on the terminal. Thus, it does not require operators to perform complex related operations, does not require operators to have high operating skills, and does not require terminal calls.

[0376] In this embodiment of the application, in a human customer service scenario, the aforementioned call terminal is a user call terminal in the customer service scenario, and the peer call terminal is a customer service call terminal in the customer service scenario. Based on the relevant descriptions of the above embodiments, it can be understood that the user call terminal initiating the call can initiate a video call or a voice call. The following detailed description of the communication method provided in this embodiment of the application will take a voice call initiated by the user call terminal as an example.

[0377] like Figure 7 As shown, the communication method provided in this application embodiment includes:

[0378] S701: The user's voice terminal sends an invitation message to the media server through the IMS network element.

[0379] S702, the media server sends ringing messages to the user's calling terminal through the IMS network element.

[0380] S703, the media server sends response messages to the user's call terminal through the IMS network element.

[0381] Understandably, after the customer service system answers a user's call, the media server plays audio prompts related to the user's business to prompt the user to select different service options (such as selecting human assistance) according to their needs. When the user selects human assistance based on the prompts, the media server detects the selection and assigns a customer service representative to the user (i.e., selects a corresponding customer service terminal for the user's call terminal).

[0382] S704, the media server sends an invitation message to the customer service call terminal.

[0383] S705, the customer service call terminal sends a response message to the media server.

[0384] S706, the media server sends a reinvite message to the user's voice terminal through the IMS network element.

[0385] The reinvite message includes the media server's SDP information, which includes the media server's address information (such as IP address), audio port information, and audio codec format.

[0386] S707: The user's voice terminal sends a response message to the media server through the IMS network element.

[0387] The response message includes the user's voice terminal SDP information, which includes the user's voice terminal's address information (such as IP address), audio port information, and audio codec format.

[0388] The above S706-S707 describes the process by which the call terminal interacts with the media server to establish a voice call media transmission channel between the user's call terminal and the media server through media resource negotiation.

[0389] S708, the media server sends a reinvite message to the customer service call terminal.

[0390] The reinvite message includes the media server's SDP information. For a description of the media server's SDP information, please refer to S906.

[0391] S709, the customer service call terminal sends a response message to the media server.

[0392] The response message includes the SDP information of the customer service call terminal, which includes the terminal's address information (such as IP address), audio port information, and audio codec format.

[0393] The above S709-S709 describes the process by which the customer service call terminal interacts with the media server to establish a voice call media transmission channel between the customer service call terminal and the media server through media resource negotiation.

[0394] It is understandable that S701-S709 is the process by which a user's calling terminal calls a customer service calling terminal and establishes a voice call media transmission channel. The content carried in the messages of each step of S701-S709 can be found in the detailed description of S201-S209 above, and will not be repeated here.

[0395] S710: The customer service terminal sends a face recognition request to the media server. Correspondingly, the media server receives the face recognition request sent by the customer service terminal.

[0396] The facial recognition application includes a facial recognition application identifier, which is used to apply for facial recognition of the user corresponding to the call terminal, that is, to perform facial recognition on the call terminal during the call between the call terminal and the other end of the call terminal.

[0397] S711. The media server sends a SIP message to the user's calling terminal through the IMS network element. This SIP message includes a face recognition request identifier, which is used to request face recognition for the user corresponding to the calling terminal. Correspondingly, the user's calling terminal receives the SIP message from the media server.

[0398] The SIP message also includes the media server's SDP information, which includes the media server's address information, audio port information, audio codec format, video port information, and video codec format.

[0399] Optionally, the Face Recognition Request Identifier (FR) in the SIP message can be carried in the header field of the SIP message, or the Face Recognition Request Identifier can also be carried in the SDP information in the SIP message. For details, please refer to the relevant descriptions of S603-S604 in the above embodiments, which will not be repeated here.

[0400] S712. The user terminal sends a SIP message response to the media server through the IMS network element. This SIP message response indicates that the user corresponding to the terminal agrees to face recognition. Correspondingly, the media server receives the SIP message response sent by the user terminal.

[0401] Optionally, the response message of the SIP message includes a face recognition response identifier, which indicates that the user corresponding to the calling terminal agrees to face recognition.

[0402] The response message to the SIP message includes the SDP information of the user's calling terminal, which includes the address information (e.g., IP address), audio port information, audio codec format, video port information, and video codec format of the user's calling terminal.

[0403] It should be noted that the user's calling terminal initiates a voice call, and the steps S706-S709 in the above embodiment establish a voice call media transmission channel. Since the face recognition process transmits a face video stream, while the voice call process can only transmit the voice stream and not the video stream, after the media server receives the face recognition request and the user agrees to face recognition, the media server will trigger the establishment of a video call media transmission channel. That is, the voice call needs to be converted into a video call to establish a transmission channel that can transmit a video stream, and the face video stream is transmitted using this video call media transmission channel.

[0404] It is understood that the video call media transmission channel includes a first video call media transmission channel between the user call terminal and the media server, and a second video call terminal between the customer service call terminal and the media server. The first video call media transmission channel and the second video call media transmission channel exist in pairs and are used as transmission channels for communication between the customer service call terminal and the user call terminal.

[0405] In this embodiment of the application, in S711-S712 above, media resource negotiation is performed using the SDP information of the media server in the SIP message and the SDP information of the user call terminal in the response message of the SIP message to establish a video call media transmission channel (i.e., the first video call media transmission channel) between the user call terminal and the media server.

[0406] The process of establishing the second video call media transmission channel is as follows: S713-S714.

[0407] S713, The media server sends a SIP message to the customer service terminal. Correspondingly, the customer service terminal receives the SIP message from the media server. This SIP message includes the media server's SDP information.

[0408] The SDP information of this media server includes the media server's address information, audio port information, audio codec format, video port information, and video codec format. It should be noted that, unlike the SDP information in the media resource negotiation message for voice calls (which includes the device's address information, audio port information, and audio codec format), the SDP information in the media resource negotiation message for video calls also includes the device's video port information and video codec format.

[0409] S714. The customer service terminal sends a response message to the media server via a SIP message. Correspondingly, the media server receives a response message from the customer service terminal via a SIP message, which includes the customer service terminal's SDP information.

[0410] The SDP information of the customer service call terminal includes the terminal's address information, audio port information, audio codec format, video port information, and video codec format.

[0411] The media resource negotiation process of S713-S714 can establish a video call media transmission channel (i.e., a second video call media transmission channel) between the customer service call terminal and the media server.

[0412] S715, the media server sends a first transmission channel indication message to the user's calling terminal. This first transmission channel indication message indicates that the face video stream is transmitted through the video call media transmission channel. Accordingly, the user's calling terminal receives the transmission channel indication message from the media server.

[0413] At this point, a transmission channel for transmitting facial video streams has been established. This transmission channel is a video call media transmission channel, and facial recognition can be completed based on this video call media transmission channel.

[0414] It should be noted that when the face video stream is a video stream captured by the camera device of the user's call terminal, the face video stream still belongs to the call video stream. In this case, the media server may not send the first transmission channel indication information to the user's call terminal.

[0415] S716. The user's calling terminal sends a face video stream to the media server through the video call media transmission channel. Correspondingly, the media server receives the face video stream from the user's calling terminal through the same video call media transmission channel. This face video stream contains the face image of the user corresponding to the user's calling terminal.

[0416] In this embodiment, the user's calling terminal transmits a face video stream through the aforementioned video call media transmission channel for transmitting call video streams, according to the first transmission channel indication information it receives. The face video stream includes the face image of the user corresponding to the calling terminal. Optionally, the face video stream is a video stream obtained by the camera device of the calling terminal; or, the face video stream is a video stream obtained from the storage device of the calling terminal.

[0417] S717, the media server obtains the face recognition results.

[0418] S718, the media server sends the face recognition result to the user's calling terminal. Correspondingly, the user's calling terminal receives the face recognition result sent by the media server.

[0419] S719. The media server sends a second transmission channel indication message to the customer service call terminal. Correspondingly, the customer service call terminal receives the second transmission channel indication message from the media server. The second transmission channel indication message instructs the customer service call terminal to receive the face recognition result through the second video call media transmission channel.

[0420] S720: The media server sends the facial recognition result to the customer service terminal via the video call media transmission channel. Correspondingly, the customer service terminal receives the facial recognition result from the media server via the video call media transmission channel.

[0421] S721: The customer service call terminal processes business requests based on facial recognition results.

[0422] Optionally, the media server sends source indication information to the calling terminal, which instructs the calling terminal to acquire a facial video stream through the calling terminal's camera device or from the calling terminal's storage device.

[0423] Optionally, if the user's calling terminal and the media server transmit a face video stream through the first video call media transmission channel, and the face video stream is obtained from the storage device of the user's calling terminal, the calling terminal and the media server may stop transmitting the call video stream through the first video call media transmission channel.

[0424] Similarly, optionally, when the face recognition result is transmitted between the customer service call terminal and the media server through the second video call media transmission channel, the customer service call terminal and the media server may stop transmitting the call video stream through the second video call media transmission channel.

[0425] Furthermore, the media server can also send posture instruction information to the user's calling terminal to instruct the user to adjust their posture so that the facial image meets preset conditions.

[0426] For further details regarding S701-S721, please refer to the relevant descriptions in the above embodiments, which will not be repeated here.

[0427] Accordingly, this application provides a calling terminal. Based on the above method example, the calling terminal can be divided into functional modules. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division; in actual implementation, other division methods may be used.

[0428] When dividing each function into modules according to its corresponding function. Figure 8 A schematic diagram of a possible structure of the call terminal involved in the above embodiments is shown. For example... Figure 8As shown, the call terminal includes a processing module 801, a receiving module 802, and a sending module 803. The processing module 801 establishes a video call media transmission channel for transmitting video streams between the call terminal and the peer call terminal in a video call service. The video stream includes video content captured by either the call terminal or the peer call terminal, for example, executing S601 in the above method embodiment. The receiving module 802 receives a SIP message from a media server. The SIP message includes a face recognition request identifier, which requests face recognition for the user corresponding to the call terminal, for example, executing S603 in the above method embodiment. The sending module 803 sends a response message to the media server, indicating that the user corresponding to the call terminal agrees to face recognition. It also sends a face video stream to the media server through the video call media transmission channel. The face video stream includes the face image of the user corresponding to the call terminal, for example, executing S604 and S606 in the above method embodiment. The receiving module 802 is used to receive face recognition results from the media server, for example, by performing S608 in the above method embodiment.

[0429] Optionally, the receiving module 802 is further configured to receive source indication information from the media server, which instructs the calling terminal to acquire a face video stream through the camera device of the calling terminal or to acquire a face video stream from the storage device of the calling terminal.

[0430] Optionally, the face video stream is a video stream obtained from the storage device of the call terminal. The receiving module 802 is further configured to receive transmission channel indication information from the media server, which instructs the call terminal to transmit the face video stream through the video call media transmission channel, for example, by executing S605 in the above method embodiment.

[0431] The processing module 801 is also used to control the receiving module 802 or the sending module 803 to stop transmitting the call video stream through the video call media transmission channel and execute S612 in the above method embodiment.

[0432] The receiving module 802 is further configured to receive posture indication information from the media server, which instructs the user corresponding to the call terminal to adjust the user's posture so that the face image meets preset conditions.

[0433] The modules of the aforementioned call terminal can also be used to perform other actions in the above method embodiments. All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0434] When using integrated units, Figure 9A schematic diagram of another possible structure of the calling terminal involved in the above embodiments is shown. For example... Figure 9 As shown, the calling terminal provided in this application embodiment may include a processing module 901 and a communication module 902. The processing module 901 can be used to control and manage the actions of the calling terminal. For example, the processing module 901 can be used to support the calling terminal in executing S601, S612 in the above method embodiments, and / or other processes used in the technology described herein. The communication module 902 can be used to support communication between the calling terminal and other network entities. The communication module 902 integrates the functions of the above-mentioned sending module 803 and receiving module 802. The communication module 902 can be used to support the calling terminal in executing S603, S604, S605, S606, and S608 in the above method embodiments. Optionally, as... Figure 9 As shown, the call terminal may also include a storage module 903 for storing the program code and data of the call terminal, such as received face video streams or face recognition results.

[0435] The processing module 901 can be a processor, for example, the processor can be... Figure 4A The processor 410 is included. The communication module 902 can be a transceiver, transceiver circuit, or communication interface, for example... Figure 4A The mobile communication module 450 and / or wireless communication module 460, and the storage module 903 can be a memory, for example Figure 4A The internal memory 421 in it.

[0436] For more details on how the modules included in the aforementioned call terminal implement the above functions, please refer to the descriptions in the preceding method embodiments, which will not be repeated here.

[0437] Accordingly, embodiments of this application provide a calling terminal, which is the aforementioned Figure 8 or Figure 9 The other end of the call terminal shown can be divided into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0438] When dividing each function into modules according to its corresponding function. Figure 10 A schematic diagram of a possible structure of the call terminal involved in the above embodiments is shown. For example... Figure 10As shown, the call terminal includes a processing module 1001, a sending module 1002, and a receiving module 1003. The processing module 1001 establishes a video call media transmission channel for transmitting call video streams between the call terminal and the peer call terminal in a video call service. The call video stream includes video content captured by either the call terminal or the peer call terminal, for example, executing S601 in the above method embodiment. The sending module 1002 sends a face recognition request to a media server. The face recognition request includes a face recognition request identifier, which is used to request face recognition of the user corresponding to the peer call terminal in the call, for example, executing S602 and S710 in the above method embodiment. The receiving module 1003 is used to receive the face recognition result from the media server through the video call media transmission channel. The face recognition result is the result of face recognition of the user corresponding to the peer call terminal based on the face video stream, for example, by executing S610 and S720 in the above method embodiment; the processing module 1001 is also used to process the service request of the peer call terminal based on the face recognition result, for example, by executing S611 and S721 in the above method embodiment.

[0439] Optionally, the receiving module 1003 is further configured to receive transmission channel indication information from the media server, which instructs the call terminal to receive the face recognition result through the video call media transmission channel, for example, by executing S609 and S719 in the above method embodiments.

[0440] Optionally, the processing module 1001 is also used to control the sending module 1002 or the receiving module 1003 to stop transmitting the call video stream through the video call media transmission channel, for example, by executing S613 in the above method embodiment.

[0441] Optionally, the face recognition result includes information indicating successful face recognition or information indicating failed face recognition. The processing module 1001 is also used to process the service request of the peer call terminal when the face recognition result includes information indicating successful face recognition.

[0442] The modules of the aforementioned call terminal can also be used to perform other actions in the above method embodiments. All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0443] When using integrated units, Figure 11 A schematic diagram of another possible structure of the calling terminal involved in the above embodiments is shown. For example... Figure 11As shown, the calling terminal provided in this application embodiment may include a processing module 1101 and a communication module 1102. The processing module 1101 can be used to control and manage the actions of the calling terminal. For example, the processing module 1101 can be used to support the calling terminal in executing S601, S611, S613, S721 in the above method embodiments, and / or other processes used in the technology described herein. The communication module 1102 can be used to support communication between the calling terminal and other network entities. The communication module 1102 integrates the functions of the above sending module 1002 and receiving module 1003. The communication module 1102 can be used to support the calling terminal in executing S602, S609, S610, S710, S719, S720 in the above method embodiments. Optionally, as... Figure 11 As shown, the call terminal may also include a storage module 1103 for storing the program code and data of the call terminal.

[0444] The processing module 1101 can be a processor, for example, the processor can be a... Figure 4A The processor 410 is included. The communication module 1102 can be a transceiver, transceiver circuit, or communication interface, for example... Figure 4A The mobile communication module 450 and / or wireless communication module 460, and the storage module 1103 may be a memory, for example Figure 4A The internal memory 421 in it.

[0445] For more details on how the modules included in the aforementioned call terminal implement the above functions, please refer to the descriptions in the preceding method embodiments, which will not be repeated here.

[0446] Accordingly, this application provides a media server. Based on the above method example, the media server can be divided into functional modules. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this embodiment is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used.

[0447] When dividing each function into modules according to its corresponding function. Figure 12 This diagram illustrates a possible structure of the media server involved in the above embodiments. For example... Figure 12As shown, the media server includes a processing module 1201, a receiving module 1202, an acquisition module 1203, and a sending module 1204. The processing module 1201 is used to establish a first video call media transmission channel and a second video call media transmission channel. The first video call media transmission channel is a video call media transmission channel between the calling terminal and the media server, and the second video call media transmission channel is a video call media transmission channel between the media server and the peer calling terminal. The first and second video call media transmission channels are used for transmitting call video streams between the calling terminal and the peer calling terminal in a video call service. The call video stream includes video content captured by the calling terminal or the peer calling terminal, for example, by executing S601 in the above method embodiment. The receiving module 1202 is used to receive a face recognition application from the peer call terminal. The face recognition application includes a face recognition application identifier, which is used to request face recognition of the user corresponding to the call terminal communicating with the peer call terminal, for example, executing S602 and S710 in the above method embodiments. It also receives a face video stream from the call terminal through a first video call media transmission channel. The face video stream includes the face image of the user corresponding to the call terminal, for example, executing S606 and S716 in the above method embodiments. The obtaining module 1203 is used to obtain a face recognition result, which is the result of face recognition of the user corresponding to the call terminal based on the face video stream, for example, executing S607 and S717 in the above method embodiments. The sending module 1204 is used to send the face recognition result to the peer call terminal through a second video call media transmission channel to trigger the peer call terminal to process the call terminal's service request based on the face recognition result, for example, executing S608 and S720 in the above method embodiments.

[0448] Optionally, the sending module 1204 is further configured to send a SIP message to the calling terminal, the SIP message including a face recognition request identifier, the face recognition request identifier being used to request face recognition for the user corresponding to the calling terminal, for example, executing S603 and S711 in the above method embodiment; the receiving module 1202 is further configured to receive a response message of the SIP message from the calling terminal, the response message of the SIP message indicating that the user corresponding to the calling terminal agrees to face recognition, for example, executing S604 and S712 in the above method embodiment.

[0449] Optionally, the sending module 1204 is further configured to send source indication information to the calling terminal, the source indication information instructing the calling terminal to acquire a face video stream through the camera device of the calling terminal or to acquire a face video stream from the storage device of the calling terminal.

[0450] Optionally, the face video stream is a video stream obtained from the storage device of the call terminal. The sending module 1204 is also used to send a first transmission channel indication information to the call terminal, which instructs the call terminal to transmit the face video stream through a first video call media transmission channel, for example, by executing S605 and S715 in the above method embodiment.

[0451] Optionally, the processing module 1201 is also used to control the sending module 1204 or the receiving module 1202 to stop transmitting the call video stream through the first video call media transmission channel, for example, by executing S612 in the above method embodiment.

[0452] Optionally, the sending module 1204 is further configured to send a second transmission channel indication information to the peer call terminal, the second transmission channel indication information indicating that the peer call terminal receives the face recognition result through the second video call media transmission channel, for example, by executing S609 and S719 in the above method embodiments.

[0453] Optionally, the processing module 1201 is also used to control the sending module 1204 or the receiving module 1202 to stop transmitting the call video stream through the second video call media transmission channel, for example, by executing S613 in the above method embodiment.

[0454] Optionally, the sending module 1204 is also used to send the face recognition result to the call terminal, for example, by executing S608 and S718 in the above method embodiments.

[0455] Optionally, the sending module 1204 is further configured to send posture indication information to the call terminal, which instructs the user corresponding to the call terminal to adjust the user's posture so that the face image meets preset conditions.

[0456] Optionally, the processing module 1201 is further configured to extract the target face image from the face video stream; the sending module 1204 is further configured to send the target face image to the face recognition server to trigger the face recognition server to perform face recognition on the user corresponding to the call terminal based on the target face image; the aforementioned acquisition module 1203 is specifically configured to receive the face recognition result from the face recognition server.

[0457] The various modules of the media server described above can also be used to perform other actions in the above method embodiments. All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0458] When using integrated units, Figure 13 A schematic diagram of another possible structure of the media server involved in the above embodiments is shown. For example... Figure 13As shown, the media server provided in this application embodiment may include a processing module 1301 and a communication module 1302. The processing module 1301 can be used to control and manage the actions of the media server. For example, the processing module 1301 can be used to support the media server in executing S601, S607, S612, S613, S717 in the above method embodiments, and / or other processes used in the technology described herein. The communication module 1302 can be used to support communication between the media server and other network entities. The communication module 1302 integrates the functions of the above receiving module 1202 and sending module 1204. The communication module 1302 can be used to support the media server in executing S602, S603, S604, S605, S606, S608, S609, S710, S711, S712, S715, S716, S719, S720 in the above method embodiments. Optionally, as... Figure 13 As shown, the media server may also include a storage module 1303 for storing the program code and data of the media server.

[0459] The processing module 1301 can be a processor, for example, the processor can be... Figure 5 The processor 501 is located in the middle. The communication module 1302 can be a transceiver, transceiver circuit, or network interface, for example... Figure 5 The network interface 503 and the storage module 1303 can be a memory, for example... Figure 5 The memory 502 in the memory.

[0460] For more details on how the modules included in the media server implement the above functions, please refer to the descriptions in the previous method embodiments, which will not be repeated here.

[0461] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0462] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center integrating one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).

[0463] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0464] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0465] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0466] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0467] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.

[0468] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A communication method, characterized in that, The method is executed by the calling terminal, and the method includes: A video call media transmission channel is established, which is used for transmitting the call video stream between the call terminal and the peer call terminal in the video call service. The call video stream includes video content captured by the call terminal or the peer call terminal. Receive a Session Initiation Protocol (SIP) message from the media server. The SIP message includes a face recognition request identifier, which is used to request face recognition of the user corresponding to the call terminal. A response message to the SIP message is sent to the media server, the response message indicating that the user corresponding to the call terminal agrees to perform facial recognition; A facial video stream is sent to the media server through the video call media transmission channel. The facial video stream includes the facial image of the user corresponding to the call terminal. Receive the face recognition result from the media server.

2. The method according to claim 1, characterized in that, The facial video stream is a video stream captured by the camera device of the call terminal; or, The facial video stream is a video stream obtained from the storage device of the call terminal.

3. The method according to claim 2, characterized in that, The method further includes: The call terminal receives source indication information from the media server, which instructs the call terminal to acquire a face video stream through its camera device or from its storage device.

4. The method according to claim 3, characterized in that, The face video stream is a video stream obtained from the storage device of the call terminal. Before sending the face video stream to the media server through the video call media transmission channel, the method further includes: The call terminal receives transmission channel indication information from the media server, which instructs the call terminal to transmit the face video stream through the video call media transmission channel.

5. The method according to claim 4, characterized in that, Before sending the face video stream to the media server via the video call media transmission channel, the method further includes: Stop transmitting the call video stream through the video call media transmission channel.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: The user receives posture indication information from the media server, which instructs the user corresponding to the call terminal to adjust the user's posture so that the face image meets preset conditions.

7. A communication method, characterized in that, The method is executed by the calling terminal, and the method includes: A video call media transmission channel is established, which is used for transmitting the call video stream between the call terminal and the peer call terminal in the video call service. The call video stream includes video content captured by the call terminal or the peer call terminal. Send a face recognition request to the media server. The face recognition request includes a face recognition request identifier, which is used to request face recognition of the user corresponding to the peer call terminal that is talking to the call terminal. The face recognition result is received from the media server through the video call media transmission channel. The face recognition result is the result of face recognition of the user corresponding to the peer call terminal based on the face video stream. Based on the facial recognition results, the service requests of the peer call terminal are processed.

8. The method according to claim 7, characterized in that, The facial video stream is a video stream captured by the camera device of the peer communication terminal; or, The facial video stream is a video stream obtained from the storage device of the peer's calling terminal.

9. The method according to claim 7 or 8, characterized in that, Before receiving the face recognition result from the media server via the video call media transmission channel, the method further includes: The media server receives transmission channel indication information, which instructs the call terminal to receive the face recognition result through the video call media transmission channel.

10. The method according to any one of claims 7 to 9, characterized in that, Before receiving the face recognition result from the media server via the video call media transmission channel, the method further includes: Stop transmitting the call video stream through the video call media transmission channel.

11. The method according to any one of claims 7 to 10, characterized in that, The facial recognition result includes information indicating successful facial recognition or information indicating failed facial recognition. The step of processing the service request from the peer call terminal based on the facial recognition result includes: If the face recognition result includes information indicating successful face recognition, the service request from the peer call terminal is processed.

12. A communication method, characterized in that, The method is executed by a media server, and the method includes: A first video call media transmission channel and a second video call media transmission channel are established. The first video call media transmission channel is a video call media transmission channel between the call terminal and the media server, and the second video call media transmission channel is a video call media transmission channel between the media server and the peer call terminal. The first video call media transmission channel and the second video call media transmission channel are used for the transmission of call video streams between the call terminal and the peer call terminal in the video call service. The call video stream includes video content captured by the call terminal or the peer call terminal. The system receives a face recognition request from the peer call terminal. The face recognition request includes a face recognition request identifier, which is used to request face recognition of the user corresponding to the call terminal that is communicating with the peer call terminal. A facial video stream is received from the call terminal through the first video call media transmission channel, the facial video stream including the facial image of the user corresponding to the call terminal; Obtain a face recognition result, wherein the face recognition result is the result of face recognition of the user corresponding to the call terminal based on the face video stream; The face recognition result is sent to the peer call terminal through the second video call media transmission channel to trigger the peer call terminal to process the call terminal's service request based on the face recognition result.

13. The method according to claim 12, characterized in that, After receiving a face recognition request from the peer terminal, the method further includes: A Session Initiation Protocol (SIP) message is sent to the calling terminal. The SIP message includes a face recognition request identifier, which is used to request face recognition of the user corresponding to the calling terminal. The call terminal receives a response message to the SIP message, which indicates that the user corresponding to the call terminal agrees to perform facial recognition.

14. The method according to claim 12 or 13, characterized in that, The facial video stream is a video stream captured by the camera device of the call terminal; or, The facial video stream is a video stream obtained from the storage device of the call terminal.

15. The method according to claim 14, characterized in that, The method further includes: Send source indication information to the calling terminal, the source indication information instructing the calling terminal to acquire a face video stream through the calling terminal's camera device or to acquire a face video stream from the calling terminal's storage device.

16. The method according to claim 15, characterized in that, The face video stream is a video stream obtained from the storage device of the call terminal. Before receiving the face video stream from the call terminal through the first video call media transmission channel, the method further includes: A first transmission channel indication message is sent to the calling terminal, the first transmission channel indication message instructing the calling terminal to transmit the face video stream through the first video call media transmission channel.

17. The method according to claim 16, characterized in that, Before receiving the face video stream from the calling terminal via the first video call media transmission channel, the method further includes: Stop transmitting the call video stream through the first video call media transmission channel.

18. The method according to any one of claims 12 to 17, characterized in that, Before sending the face recognition result to the peer call terminal via the second video call media transmission channel, the method further includes: Send a second transmission channel indication message to the peer call terminal, the second transmission channel indication message instructing the peer call terminal to receive the face recognition result through the second video call media transmission channel.

19. The method according to any one of claims 12 to 18, characterized in that, Before sending the face recognition result to the peer call terminal via the second video call media transmission channel, the method further includes: Stop transmitting the call video stream through the second video call media transmission channel.

20. The method according to any one of claims 12 to 19, characterized in that, After obtaining the face recognition results, the method further includes: The facial recognition result is sent to the call terminal.

21. The method according to any one of claims 12 to 20, characterized in that, The method further includes: A posture indication message is sent to the call terminal, which instructs the user corresponding to the call terminal to adjust the user's posture so that the face image meets preset conditions.

22. The method according to any one of claims 12 to 21, characterized in that, After receiving a face video stream from the calling terminal via the first video call media transmission channel, the method further includes: Extract the target face image from the face video stream; Send the target face image to the face recognition server to trigger the face recognition server to perform face recognition on the user corresponding to the call terminal based on the target face image; The process of obtaining the face recognition result includes: Receive the face recognition result from the face recognition server.

23. A communication terminal, characterized in that, The device includes a memory and at least one processor connected to the memory, the memory being used to store computer program code, the computer program code including computer instructions, which, when executed by the at least one processor, cause the call terminal to perform the method as described in any one of claims 1 to 6.

24. A communication terminal, characterized in that, The device includes a memory and at least one processor connected to the memory, the memory being used to store computer program code, the computer program code including computer instructions, which, when executed by the at least one processor, cause the call terminal to perform the method as described in any one of claims 7 to 11.

25. A media server, characterized in that, The system includes a memory and at least one processor connected to the memory, the memory being used to store computer program code, the computer program code including computer instructions, which, when executed by the at least one processor, cause the media server to perform the method as described in any one of claims 12 to 22.

26. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on a call terminal, cause the call terminal to perform the method as described in any one of claims 1 to 6.

27. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on a call terminal, cause the call terminal to perform the method as described in any one of claims 7 to 11.

28. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on a server, cause the server to perform the method as described in any one of claims 12 to 22.

29. A communication system, characterized in that, It includes a calling terminal, a peer calling terminal, and a media server; the calling terminal performs the method as described in any one of claims 1 to 6, the peer calling terminal performs the method as described in any one of claims 7 to 11, and the media server performs the method as described in any one of claims 12 to 22.

Citation Information

Patent Citations

  • Call method for terminal and related devices

    CN107396328A

  • Call prompt method and device, memory medium and mobile terminal

    CN108696641A