Audio and video communication system and method
By storing the media format and identification supported by the terminal in the call domain server and the called domain server, the first terminal conducts audio and video communication with the second terminal based on the identification of the target media format, solving the computing resource consumption problem caused by multiple calls and media negotiation in the prior art, and realizing more efficient audio and video communication.
Patent Information
- Application Number
- CN202510347061.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art requires multiple call operations to be performed in each service execution in audio and video communication, and requires media negotiation and media negotiation feedback processes, resulting in increased computing resource consumption.
By saving the media formats supported by the first and second terminals and their corresponding identifiers in the call domain server and the called domain server, the first terminal conducts audio and video communication with the second terminal through the call domain server and the called domain server based on the identification of the target media format, avoiding multiple calls and media negotiation processes.
It saves the process of media negotiation and media negotiation feedback, reduces the consumption of computing resources, and improves the efficiency of audio and video communication.
Smart Images

Figure CN120201151A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to an audio-video communication system and method. Background Art
[0002] Currently, existing audio-video communication solutions (such as video conferencing, video command, video surveillance, etc.) generally adopt a "call" mode similar to that of a program-controlled telephone. Each participating service terminal device or user exists in the system in the form of a CALL ID@IP address similar to a telephone number. When the video communication system organizes a service, it calls the required audio-video resources in the form of a CALL ID@IP address. After the call, it can only be sent point-to-point (P2P) or forwarded through a central service device and then reaches the receiving terminal side. The audio-video on the receiving terminal side is also sent after making a call in the above manner. Under the existing technical system, each service execution requires one or more call operations, and also requires a process of media negotiation and media negotiation feedback, which increases the consumption of computing resources. Summary of the Invention
[0003] The purpose of the embodiments of this application is to provide an audio-video communication system and method to reduce the consumption of computing resources.
[0004] In a first aspect, the present invention provides an audio-video communication system, which includes: a first terminal, a calling domain server, a called domain server, and a second terminal that are communicatively connected in sequence; the calling domain server and the called domain server both store a first media format supported by the first terminal and a second media format supported by the second terminal; each first media format supported by the first terminal has its own corresponding first identifier; each second media format supported by the second terminal has its own corresponding second identifier; each first identifier and each second identifier are different; the first terminal is used to perform audio-video communication with the second terminal based on the identifier corresponding to the target media format through the calling domain server and the called domain server.
[0005] In an alternative embodiment, the target media format is a media format commonly supported by the first terminal and the second terminal; the calling domain server is used to send a first pull request to the first terminal; wherein, the first pull request carries a first target identifier; the first target identifier is an identifier corresponding to the target media format supported by the first terminal; the first terminal is used to encode the first media data to be sent according to the target media format to obtain the encoded first media data, and send the encoded first media data to the calling domain server; the called domain server is used to send a second pull request to the second terminal; wherein, the second pull request carries a second target identifier; the second target identifier is an identifier corresponding to the target media format supported by the second terminal; the second terminal is used to encode the second media data to be sent according to the target media format to obtain the encoded second media data, and send the encoded second media data to the called domain server; the calling domain server is used to send the encoded first media data to the called domain server, so as to send the encoded first media data to the second terminal through the called domain server; the called domain server is used to send the encoded second media data to the calling domain server, so as to send the encoded second media data to the first terminal through the calling domain server.
[0006] In an alternative embodiment, the target media format is determined in the following manner: the target domain server is used to determine at least one common media format commonly supported by the first terminal and the second terminal according to each first media format and each second media format; wherein, the target domain server is the calling domain server and / or the called domain server; the target media format is determined from the at least one common media format.
[0007] In an alternative embodiment, corresponding usage permissions are configured for at least a part of the first identifiers corresponding to the first media formats; corresponding usage permissions are configured for at least a part of the second identifiers corresponding to the second media formats.
[0008] In an alternative embodiment, the target domain server is used to: determine at least one common media format commonly supported by the first terminal and the second terminal according to each first media format and the usage permission corresponding to each first identifier, each second media format, and the usage permission corresponding to each second identifier.
[0009] In an alternative embodiment, the target media format is one of the first media formats supported by the first terminal; the second terminal is in a state of being called by other terminals; the first terminal is configured to send a third pull request to the call domain server, and the third pull request is sent to the second terminal through the call domain server and the called domain server in sequence; wherein, the third pull request carries a third target identifier, and the third target identifier is an identifier of the target media format corresponding to the second terminal; the second terminal is configured to encode the third media data to be sent according to the target media format to obtain the encoded third media data, and send the encoded third media data to the first terminal through the called domain server and the call domain server.
[0010] In an alternative embodiment, there are multiple called domain servers and multiple second terminals, and each called domain server is communicatively connected to a corresponding second terminal; the target media format is specified by the user through the first terminal; the call domain server is configured to send a fourth pull request to the first terminal; wherein, the fourth pull request carries a fourth target identifier; the fourth target identifier is an identifier of the target media format corresponding to the first terminal; the first terminal is configured to encode the fourth media data to be sent according to the target media format to obtain the encoded fourth media data, and send the encoded fourth media data to the call domain server; each called domain server is configured to send a fifth pull request to the corresponding second terminal; wherein, the fifth pull request carries a fifth target identifier; the fifth target identifier is an identifier of the target media format corresponding to the second terminal; each second terminal is configured to encode the fifth media data to be sent according to the target media format to obtain the encoded fifth media data, and send the encoded fifth media data to its corresponding called domain server; the call domain server is configured to send the encoded fourth media data to each called domain server, so that each called domain server sends the encoded fourth media data to the corresponding second terminal; each called domain server is configured to send the encoded fifth media data to the call domain server, so that the call domain server sends the encoded fifth media data to the first terminal and / or other second terminals except the current second terminal corresponding to the current called domain server.
[0011] In an alternative embodiment, the user corresponding to the first terminal is an anonymous user or a specified user.
[0012] In an alternative embodiment, the user corresponding to the second terminal is an anonymous user or a specified user.
[0013] Second aspect, the present invention provides an audio - video communication method, which is applied to an audio - video communication system. The system includes: a first terminal, a calling domain server, a called domain server, and a second terminal that are communicatively connected in sequence; both the calling domain server and the called domain server store a first media format supported by the first terminal and a second media format supported by the second terminal; each first media format supported by the first terminal has its corresponding first identifier; each second media format supported by the second terminal has its corresponding second identifier; each first identifier and each second identifier are different; the method includes: the first terminal performs audio - video communication with the second terminal through the calling domain server and the called domain server based on the identifier corresponding to the target media format.
[0014] In the audio - video communication system and method provided by the present invention, both the calling domain server and the called domain server store a first media format supported by the first terminal and a second media format supported by the second terminal, and each first media format has its corresponding first identifier; each second media format has its corresponding second identifier, and the first identifier and the second identifier are both unique. The first terminal and the second terminal can perform audio - video communication through the calling domain server and the called domain server based on the identifier corresponding to the target media format, saving the process of media negotiation and media negotiation feedback, and reducing the consumption of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a schematic diagram of the communication process of a two - way call service in a traditional mode;
[0017] Figure 2 It is a schematic diagram of the communication process of a hybrid service in a traditional mode;
[0018] Figure 3 It is a schematic diagram of the communication process of a multi - point communication service in a traditional mode;
[0019] Figure 4 It is a schematic diagram of an audio - video communication system provided by an embodiment of the present invention;
[0020] Figure 5 It is a schematic diagram of the communication process of a two - way call service provided by an embodiment of the present invention;
[0021] Figure 6Schematic diagram of a hybrid service communication process provided by an embodiment of the present invention;
[0022] Figure 7 Schematic diagram of a multi-point communication service communication process provided by an embodiment of the present invention. Detailed implementation manners
[0023] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] Currently, the common organization methods of audio and video services are divided into the following three types:
[0025] 1) Mesh architecture: In the Mesh architecture, each terminal establishes a connection with other terminals, forming a mesh structure. This architecture is suitable for small-scale multi-person communication. However, as the number of participants increases, the required number of connections increases exponentially, resulting in a sharp rise in the demand for bandwidth and processing power. Generally, it is suitable for small-scale communication;
[0026] 2) MCU (MultiPoint Control Unit) mode: In the MCU mode, all audio and video streams are forwarded and processed through a central node. This mode can support large-scale multi-person communication, but has high performance requirements for the central node and there is a risk of single-point failure;
[0027] 3) SFU (Selective Forwarding Unit) mode: The SFU mode is a compromise solution. It only forwards selected audio and video streams, rather than processing all streams like the MCU. This method can reduce the load on the central node and support more participants at the same time. Due to its scalable forwarding mechanism, it consumes a relatively high amount of bandwidth in the communication network.
[0028] In the above audio and video service organization, when a new service is launched for a certain terminal or user, a call operation is performed through a call instruction, and then media negotiation is carried out through the Session Description Protocol (SDP). According to the negotiation result, audio and video payload data is sent to the central service node or the peer terminal or user for peer-to-peer communication. For the launch of different heterogeneous services, multiple calls need to be executed and multiple selective negotiations need to be carried out.
[0029] (1) For the two-way call service in the prior art mode, such as Figure 1A schematic diagram of the communication process of two-way call service in a traditional mode is shown below. The call process is as follows:
[0030] 1. The calling user / terminal carries the called user number and makes a call request to the call domain server;
[0031] 2. After querying, the call domain server forwards the call request to the called domain server;
[0032] 3. The called domain server forwards the call request to the called user / terminal;
[0033] 4. After answering the call request, the called user / terminal sends a call response to the called domain server;
[0034] 5. The called domain server forwards the call response to the call domain server;
[0035] 6. The call domain server forwards the call response to the calling user / terminal;
[0036] 7. The calling user / terminal initiates media negotiation to the call domain server according to the supported media format;
[0037] 8. The call domain server forwards the media negotiation to the called domain server;
[0038] 9. The called domain server forwards the media negotiation to the called user / terminal;
[0039] 10. The called user / terminal takes the intersection of the media support formats of the calling user / terminal and its own supported media format, and then sends the negotiation result to the called domain server;
[0040] 11. The called domain server forwards the media negotiation result to the call domain server;
[0041] 12. The call domain server forwards the media negotiation result to the calling user / terminal;
[0042] 13. The calling user / terminal sends media data to the call domain server according to the negotiation result; the called user / terminal sends media data to the called domain server according to the negotiation result;
[0043] 14. The call domain server and the called domain server exchange media data;
[0044] 15. The call domain server sends the media data of the called user / terminal obtained after the exchange to the calling user / terminal; the called domain server sends the media data of the calling user / terminal obtained after the exchange to the called user / terminal.
[0045] (2) As Figure 2 A schematic diagram of the communication process of hybrid services in a traditional mode is shown below. The specific process is as follows:
[0046] 1. The new user / terminal C initiates a video-on-demand request (which can also be called a monitoring service) to request video-on-demand of the user / terminal B that is already in a call.
[0047] 2. The user / terminal C initiates a media negotiation request (media heterogeneity) based on the local media format support.
[0048] 3. The call domain server forwards the video-on-demand request to the called domain server.
[0049] 4. The call domain server forwards the media negotiation request to the called domain server.
[0050] 5. The called domain server negotiates the media format with the user / terminal B.
[0051] 6. The user / terminal B feeds back the media occupancy situation to the called domain server.
[0052] 7. The called domain server feeds back the media negotiation result to the call domain server.
[0053] 8. The call domain server forwards the media negotiation result to the user / terminal C.
[0054] 9. The call domain server initiates a heterogeneous media data request based on the negotiation result.
[0055] 10. The user / terminal B sends media data to the called domain server.
[0056] 11. The called domain server forwards the media data to the call domain server.
[0057] 12. The call domain server forwards the media data to the user / terminal C.
[0058] (3) As Figure 3 shown, a schematic diagram of the multi-point communication service process in a traditional mode is as follows:
[0059] 1. The user / terminal A initiates a multi-person service request to the call domain server.
[0060] 2. The call domain server synchronizes the service information with multiple called domain servers.
[0061] 3. The called domain server synchronizes the service information with multiple users / terminals (such as the user / terminal B, user / terminal C, and user / terminal D in the figure).
[0062] 4. The multiple users / terminals corresponding to the called domain server feed back the synchronized service status results to their respective called domain servers.
[0063] 5. Each called domain server synchronizes the service status results to the call domain server.
[0064] 6. The call domain server synchronizes the service status result to the user / terminal A;
[0065] 7. The user / terminal A initiates media negotiation;
[0066] 8. The call domain server forwards the media negotiation to each called domain server;
[0067] 9. Each called domain server sends the media negotiation to its corresponding user / terminal;
[0068] 10. Each called user / terminal feeds back the media negotiation result to the corresponding called domain service;
[0069] 11. The called domain server forwards the media negotiation result to the call domain server;
[0070] 12. The call domain server feeds back the media negotiation result to the user / terminal A;
[0071] 13. The user / terminal A sends user media data to the call domain server, and each called user / terminal sends user media data to the corresponding called domain service;
[0072] 14. According to the service scheduling relationship, the call domain server and each called domain server continue to exchange media data;
[0073] 15. The local domain service forwards the foreign domain media data to the local user / terminal.
[0074] The above existing technologies have the following objective drawbacks:
[0075] 1) Under the existing technology system, each service execution requires one or more call operations. At the same time, it is impossible to accurately determine whether there is a risk of media format heterogeneity before the call. If there is a risk of heterogeneity, transcoding functions need to be added at the central node or the terminal, which increases the consumption of computing resources. At the same time, invalid call signaling occupies the communication bandwidth, which is extremely unfriendly to narrowband communication;
[0076] 2) Under the existing communication system, the user's permission division cannot be refined. When managing permissions based on terminals or users, the permissions can generally only be controlled as "yes" or "no", that is, whether there is permission to call the terminal (user). It is impossible to achieve permission division for different levels of capabilities and different service types. To achieve the above permission management capabilities, a relatively complex permission management model needs to be organized, increasing the system overhead;
[0077] 3) Under the existing communication system, it is impossible to effectively aggregate and manage multiple heterogeneous video and audio resources. In the call mode, video and audio resources exist in the system data stream in the form of communication loads. For the semi-heterogeneous situation (the same video format but different audio formats or the same audio format but different video formats), it is impossible to perform refined transcoding processing. Instead, they can only enter the transcoding queue uniformly, and the video and audio data are transcoded synchronously, which greatly reduces the execution efficiency and increases the system latency. Based on this, the embodiments of the present invention provide an audio-video communication system and method, and this technology can be applied to the application scenarios of audio-video communication.
[0078] To facilitate the understanding of this embodiment, first, an audio-video communication system disclosed in the embodiments of the present invention will be introduced as follows Figure 4 As shown, the system includes: a first terminal 40, a call domain server 41, a called domain server 42, and a second terminal 43 that are communicatively connected in sequence; the call domain server 41 and the called domain server 42 both store a first media format supported by the first terminal 40 and a second media format supported by the second terminal 43; each first media format supported by the first terminal 40 has its corresponding first identifier; each second media format supported by the second terminal 43 has its corresponding second identifier; each first identifier and each second identifier are different; the first terminal 40 is used to perform audio-video communication with the second terminal 43 based on the identifier corresponding to the target media format through the call domain server 41 and the called domain server 42.
[0079] The above-mentioned first terminal 40 and second terminal 43 can be a computer, a mobile phone, a tablet computer, etc.; the above-mentioned calling domain server 41 and called domain server 42 can be physical servers or virtual servers, and can be specifically set according to actual needs, which are not limited here; the above-mentioned first media format and second media format can be video coding formats, such as H.264 / AVC (Advanced Video Coding), H.265 / HEVC (High Efficiency Video Coding), etc., or can be audio coding formats, such as MP3, etc.; the above-mentioned first identifier and second identifier can be in the form of UUID (Universally Unique IDentifier). UUID is a 128-bit identifier used to uniquely identify information in a distributed system. Each first identifier and each second identifier are different. That is, each first identifier corresponding to each first media format supported by the first terminal 40 is different, and each second identifier corresponding to each second media format supported by the second terminal 43 is different. The identifiers corresponding to the same media format supported by different terminals are also different. It can also be understood that each first identifier and each second identifier are unique; the above-mentioned target media format can belong to the first media format or the second media format, etc.; in actual implementation, the first terminal 40 and the second terminal 43 perform audio and video communication through the calling domain server 41 and the called domain server 42 based on the identifier corresponding to the target media format.
[0080] In the above audio and video communication system, the first media format supported by the first terminal and the second media format supported by the second terminal are both stored in the calling domain server and the called domain server, and each first media format has its own corresponding first identifier; each second media format has its own corresponding second identifier, and the first identifier and the second identifier are both unique. The first terminal and the second terminal can perform audio and video communication through the calling domain server and the called domain server based on the identifier corresponding to the target media format, saving the process of media negotiation and media negotiation feedback, and reducing the consumption of computing resources.
[0081] Furthermore, the target media format is a media format commonly supported by the first terminal and the second terminal; for example, if the first terminal supports three media formats a, b, and c, and the second terminal supports three media formats b, c, and d, then the media formats commonly supported by the first terminal and the second terminal are b and c. In actual use, b can be selected as the target media format, or c can be selected as the target media format.
[0082] The calling domain server is used to send a first pull request to the first terminal; wherein, the first pull request carries a first target identifier; the first target identifier is the identifier corresponding to the target media format supported by the first terminal; the first terminal is used to encode the to-be-sent first media data according to the target media format to obtain the encoded first media data, and send the encoded first media data to the calling domain server;
[0083] After determining the target media format, the calling domain server can send a first pull request to the first terminal, which carries the first target identifier corresponding to the target media format supported by the first terminal. After receiving the first target identifier, the first terminal can encode the to-be-sent first media data according to the target media format, and return the encoded first media data to the calling domain server.
[0084] The called domain server is used to send a second pull request to the second terminal; wherein, the second pull request carries a second target identifier; the second target identifier is the identifier corresponding to the target media format supported by the second terminal; the second terminal is used to encode the to-be-sent second media data according to the target media format to obtain the encoded second media data, and send the encoded second media data to the called domain server;
[0085] After determining the target media format, the called domain server can send a second pull request to the second terminal, which carries the second target identifier corresponding to the target media format supported by the second terminal. After receiving the second target identifier, the second terminal can encode the to-be-sent second media data according to the target media format, and return the encoded second media data to the called domain server.
[0086] The calling domain server is used to send the encoded first media data to the called domain server, so that the called domain server can send the encoded first media data to the second terminal; the called domain server is used to send the encoded second media data to the calling domain server, so that the calling domain server can send the encoded second media data to the first terminal.
[0087] When the calling domain server receives the encoded first media data and the called domain server receives the encoded second media data, the calling domain server can exchange its media data with the called domain server. In this way, the calling domain server can receive the encoded second media data and send the encoded second media data to the first terminal; the called domain server can receive the encoded first media data and send the encoded first media data to the second terminal, thereby realizing audio and video communication between the first terminal and the second terminal.
[0088] Further, the target media format is determined as follows: the target domain server is used to determine at least one common media format supported by the first terminal and the second terminal according to each first media format and each second media format; wherein, the target domain server is the calling domain server and / or the called domain server; the target media format is determined from at least one common media format.
[0089] In actual implementation, since the first media formats supported by the first terminal and the second media formats supported by the second terminal are both stored in the calling domain server and the called domain server, when the first terminal initiates a call and the second terminal responds to the call, establishing a communication connection relationship between the two parties, the calling domain server and / or the called domain server can determine at least one common media format supported by both parties according to each first media format supported by the first terminal and each second media format supported by the second terminal. For example, if the first terminal supports three media formats a, b, and c, and the second terminal supports three media formats b, c, and d, then the common media formats supported by the first terminal and the second terminal are b and c. In specific applications, one of the media formats b and c can be selected as the target media format. For example, b can be selected as the target media format, or c can be selected as the target media format.
[0090] Further, corresponding usage permissions are configured for at least a part of the first identifiers corresponding to the first media formats; corresponding usage permissions are configured for at least a part of the second identifiers corresponding to the second media formats. In actual implementation, according to actual application requirements, corresponding usage permissions can be configured for at least a part of the first identifiers corresponding to the first media formats supported by the first terminal. For example, the users allowed to access the first identifier and the users prohibited from accessing the first identifier can be clearly defined, etc. Similarly, according to actual application requirements, corresponding usage permissions can be configured for at least a part of the second identifiers corresponding to the second media formats supported by the second terminal. For example, the users allowed to access the second identifier and the users prohibited from accessing the second identifier can be clearly defined, etc.
[0091] Further, the target domain server is used to: determine at least one common media format supported by the first terminal and the second terminal according to each first media format, the usage permissions corresponding to each first identifier, each second media format, and the usage permissions corresponding to each second identifier.
[0092] In actual implementation, since at least a part of the first identifiers corresponding to the first media formats are configured with corresponding usage permissions, and at least a part of the second identifiers corresponding to the second media formats are configured with corresponding usage permissions, when the calling domain server and / or the called domain server need to determine at least one common media format supported by the first terminal and the second terminal, it is also necessary to consider the usage permissions corresponding to each first identifier and the usage permissions corresponding to each second identifier. For example, the first terminal supports three media formats a, b, and c, and the second terminal supports three media formats b, c, and d. However, in the first terminal, access to media format b by the second terminal is prohibited, and access to media format c by the second terminal is allowed. Therefore, at least one common media format supported by the first terminal and the second terminal determined is media format c.
[0093] For ease of understanding, refer to Figure 5 the schematic diagram of a two-way call service communication process shown in the figure. The system accesses audio and video resources according to the uplink UUIDs of two users respectively, that is, changes the two-way audio and video communication into two one-way audio and video communications. The call process is as follows:
[0094] 1. The calling user / terminal (corresponding to the above-mentioned first terminal) carries the number of the called user / terminal (corresponding to the above-mentioned second terminal) to send a call request to the calling domain server. For example, it can be a video call, an audio call, etc.;
[0095] 2. After querying, the calling domain server forwards the call request to the called domain server;
[0096] 3. The called domain server forwards the call request to the called user / terminal;
[0097] 4. After the called user / terminal answers the call, it sends a call response to the called domain server;
[0098] 5. The called domain server forwards the call response to the calling domain server;
[0099] 6. The calling domain server forwards the call response to the calling user / terminal;
[0100] 7. The calling domain server and / or the called domain server determine the target media formats supported by both the calling user / terminal and the called user / terminal. The calling domain server pulls the UUID (corresponding to the above-mentioned first target identifier) corresponding to the target media format in the calling user / terminal according to the target media formats supported by both parties; the called domain server pulls the UUID (corresponding to the above-mentioned second target identifier) corresponding to the target media format in the called user / terminal according to the target media formats supported by both parties;
[0101] 8. The calling user / terminal responds to the pull request of the calling domain server, encodes the first media data to be sent according to the target media format to obtain the encoded first media data, and sends it to the calling domain server; the called user / terminal responds to the pull request of the called domain server, encodes the second media data to be sent according to the target media format to obtain the encoded second media data, and sends it to the called domain server;
[0102] 9. The calling domain server and the called domain server exchange their respective media data;
[0103] 10. The calling domain server sends the encoded second media data obtained after the exchange to the calling user / terminal; the called domain server sends the encoded first media data obtained after the exchange to the called user / terminal.
[0104] Further, the target media format is one of the first media formats supported by the first terminal; for example, if the first media formats supported by the first terminal are a, b, and c, then the target media format can be any one of a, b, and c; the second terminal is in a state of being called by other terminals, which can also be understood as the second terminal is in the process of performing audio and video communication with other terminals; the first terminal is used to send a third pull request to the calling domain server, and the third pull request is sent to the second terminal through the calling domain server and the called domain server in sequence; wherein, the third pull request carries a third target identifier, and the third target identifier is the identifier of the target media format corresponding to the second terminal; the second terminal is used to encode the third media data to be sent according to the target media format to obtain the encoded third media data, and send the encoded third media data to the first terminal through the called domain server and the calling domain server.
[0105] In actual implementation, when the first terminal requests to on-demand the second terminal that is already in a call, for example, when the user corresponding to the first terminal needs to monitor the second terminal, the first terminal can send a third pull request to the call domain server. The third pull request carries the third target identifier of the target media format corresponding to the second terminal. The call domain server forwards the third pull request to the called domain server, and the called domain server then forwards the third pull request to the second terminal. After receiving the third target identifier, the second terminal can encode the third media data to be sent according to the target media format, and return the encoded third media data to the first terminal through the called domain server and the call domain server in sequence, thus completing the on-demand of the first terminal for the second terminal. In this process, the first terminal does not need to perform two-way interaction with the second terminal, that is, the user of the second terminal is unaware of this process. It should be noted that the target media format corresponding to the second terminal above can be a media format that the second terminal originally supports, or a media format that can be supported after transcoding by the called domain server.
[0106] For easy understanding, refer to Figure 6 the schematic diagram of a hybrid service communication process shown below. The hybrid service communication process is as follows:
[0107] 1. The new user / terminal C (corresponding to the first terminal above) requests to on-demand the user / terminal B (corresponding to the second terminal above) that is already in a call;
[0108] 2. The user / terminal C sends a third pull request to the call domain server according to the local media support situation. The third pull request carries the UUID (corresponding to the third target identifier above);
[0109] 3. The call domain server forwards the third pull request carrying the UUID to the called domain server;
[0110] 4. The called domain server forwards the third pull request carrying the UUID to the user / terminal B;
[0111] 5. The called user / terminal B encodes the third media data to be sent according to the target media format, and sends the encoded third media data to the called domain server;
[0112] 6. The called domain server forwards the encoded third media data to the call domain server;
[0113] 7. The call domain server forwards the encoded third media data to the user / terminal C.
[0114] Furthermore, there are multiple called domain servers and multiple second terminals. Each called domain server is communicatively connected to the corresponding second terminal. The target media format is specified by the user through the first terminal. In actual implementation, in a business scenario involving multiple participants, there are usually multiple second terminals, and each second terminal has its own corresponding called domain server. Considering that it is difficult to determine the media format jointly supported by the first terminal and each second terminal in such a multi - participant business scenario, in this scenario, usually the user A who initiates the service request specifies it through the first terminal.
[0115] The calling domain server is used to send a fourth pull request to the first terminal; wherein, the fourth pull request carries a fourth target identifier; the fourth target identifier is the identifier of the target media format corresponding to the first terminal; the first terminal is used to encode the to - be - sent fourth media data according to the target media format to obtain the encoded fourth media data, and send the encoded fourth media data to the calling domain server;
[0116] After the first terminal specifies the target media format, the calling domain server can send a fourth pull request to the first terminal, which carries the fourth target identifier of the target media format corresponding to the first terminal. After receiving the fourth target identifier, the first terminal can encode the to - be - sent fourth media data according to the target media format and return the encoded fourth media data to the calling domain server.
[0117] Each called domain server is used to send a fifth pull request to the corresponding second terminal; wherein, the fifth pull request carries a fifth target identifier; the fifth target identifier is the identifier of the target media format corresponding to the second terminal; each second terminal is used to encode the to - be - sent fifth media data according to the target media format to obtain the encoded fifth media data, and send the encoded fifth media data to its corresponding called domain server;
[0118] After specifying the target media format, for each called domain server, the called domain server can send a fifth pull request to the corresponding second terminal, which carries the fifth target identifier of the target media format corresponding to the second terminal. After receiving the fifth target identifier, the second terminal can encode the to - be - sent fifth media data according to the target media format and return the encoded fifth media data to the called domain server.
[0119] The calling domain server is used to send the encoded fourth media data to each called domain server, so that each called domain server can send the encoded fourth media data to the corresponding second terminal; each called domain server is used to send the encoded fifth media data to the calling domain server, so that the calling domain server can send the encoded fifth media data to the first terminal and / or other second terminals except the current second terminal corresponding to the current called domain server.
[0120] After the calling domain server receives the encoded fourth media data and each called domain server receives the encoded fifth media data, the calling domain server can exchange its media data with the called domain servers. In this way, the calling domain server can receive the encoded fifth media data corresponding to each second terminal, and can send each encoded fifth media data to the first terminal, and / or send each encoded fifth media data to other second terminals except the current second terminal corresponding to the current called domain server; after each called domain server receives the encoded fourth media data, it can send the encoded fourth media data to each second terminal, thereby realizing audio and video communication between the first terminal and each second terminal.
[0121] For easy understanding, see Figure 7 the schematic diagram of a multi-point communication service communication process shown below, specifically as follows:
[0122] 1. User / terminal A (corresponding to the above-mentioned first terminal) initiates a multi-person service request to the calling domain server;
[0123] 2. The calling domain server synchronizes service information with multiple called domain servers. For example, taking a meeting as an example, the service information may include information such as meeting time and meeting location;
[0124] 3. The called domain server synchronizes service information with multiple users / terminals (corresponding to the above-mentioned multiple second terminals);
[0125] 4. Multiple users / terminals in the called domain feedback the synchronization service status result to their respective called domain servers. For example, the synchronization service status result may be whether the user / terminal in the called domain enters the meeting, etc.;
[0126] 5. Each called domain server synchronizes the service status result to the calling domain server;
[0127] 6. The calling domain server synchronizes the service status result to user / terminal A;
[0128] 7. According to the interconnection relationship between the calling domain server and each called domain server, request the UUID of the target media format corresponding to the domain user;
[0129] 8. Each domain server obtains the media data corresponding to the UUID.
[0130] 9. Each user / terminal responds to the request of the local domain service and sends the media data encoded in the target media format to the local domain server.
[0131] 10. Media data is exchanged between each domain server.
[0132] 11. The user / terminal receives the media data of each external domain user / terminal.
[0133] In this mode, for services involving multiple participants, the system distributes the uplink audio and video resources of each participant to other domain servers, and each domain server independently selects to obtain some or all of the audio and video resources to implement services such as multi-person audio and video conferencing.
[0134] Furthermore, the user corresponding to the first terminal is an anonymous user or a designated user. The user corresponding to the second terminal is an anonymous user or a designated user.
[0135] The above anonymous user can be understood as a user who does not need to disclose their real identity during audio and video communication. The characteristic of this anonymous access is that a user can participate in communication without exposing their personal information, which can improve the user experience and security. The above designated user can be understood as clearly specifying the access rights and behavior rules of a certain user. This mechanism allows the administrator to configure detailed permissions for specific users to achieve more refined management and control. In actual implementation, the users operating the first terminal and the second terminal can both be anonymous users or designated users, etc.
[0136] In the above audio and video communication system, all user resources (such as video resources, audio resources, etc.) and capabilities (media formats supported by the terminals operated by users) in the system are managed in the form of a unified resource encoding (UUID). When a user initiates a service, the system automatically generates a UUID for the user's uplink audio and video resources, and other users with permissions can access the resources from the system through the UUID.
[0137] The audio and video communication technology solution designed in this scheme uniformly encodes all user resources, device resources, and capability channels in the system with the basic audio and video payload as the basic capability unit and manages them uniformly. It enables various heterogeneous resources to be managed in a refined and fragmented manner at the center. Each uplink audio or video capability unit has an independent UUID identifier, forming a resource hotspot. Other users or operating terminals with permissions can access the resources from the central system in the form of UUID and can accurately request the supported and accessible resource UUIDs according to the required resource data.
[0138] Meanwhile, this solution breaks the authorization control relationship with users as nodes in the traditional mode. The system only needs to manage the permission relationship between UUID and users, and can flexibly control access permissions to achieve various forms such as anonymous access, registered user access, specified user access, access by times, access by time, and access authorized by the resource owner.
[0139] An embodiment of the present invention discloses an audio and video communication method, which is applied to an audio and video communication system. The system includes: a first terminal, a calling domain server, a called domain server, and a second terminal that are sequentially communicatively connected; the calling domain server and the called domain server both store a first media format supported by the first terminal and a second media format supported by the second terminal; each first media format supported by the first terminal has its own corresponding first identifier; each second media format supported by the second terminal has its own corresponding second identifier; each first identifier and each second identifier are different; the method includes:
[0140] Based on the identifier corresponding to the target media format, the first terminal performs audio and video communication with the second terminal through the calling domain server and the called domain server.
[0141] In the above audio and video communication method, the calling domain server and the called domain server both store the first media format supported by the first terminal and the second media format supported by the second terminal, and each first media format has its own corresponding first identifier; each second media format has its own corresponding second identifier, and the first identifier and the second identifier are both unique. The first terminal and the second terminal can perform audio and video communication based on the identifier corresponding to the target media format through the calling domain server and the called domain server, saving the process of media negotiation and media negotiation feedback, and reducing the consumption of computing resources.
[0142] In the embodiments provided in this application, it should be understood that the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0143] Furthermore, the functional modules in each embodiment of this application can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0144] It should be noted that if a function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0145] In this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0146] The above description is only for the embodiments of the present application and is not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An audio and video communication system, characterized in that: The system comprises: a first terminal, a calling domain server, a called domain server and a second terminal which are sequentially connected in communication; the calling domain server and the called domain server both store a first media format supported by the first terminal and a second media format supported by the second terminal; each first media format supported by the first terminal has a first identifier corresponding to it; each second media format supported by the second terminal has a second identifier corresponding to it; each first identifier and each second identifier are different; The first terminal is used to perform audio and video communication with the second terminal through the calling domain server and the called domain server based on the identifier corresponding to the target media format.
2. The system according to claim 1, characterized in that The target media format is a media format commonly supported by the first terminal and the second terminal; The calling domain server is used to send a first pull request to the first terminal; wherein the first pull request carries a first target identifier; the first target identifier is an identifier corresponding to a target media format supported by the first terminal; The first terminal is used to encode the first media data to be sent according to the target media format to obtain the encoded first media data, and send the encoded first media data to the calling domain server; The called domain server is used to send a second pull request to the second terminal; wherein the second pull request carries a second target identifier; the second target identifier is an identifier corresponding to a target media format supported by the second terminal; The second terminal is used to encode the second media data to be sent according to the target media format to obtain the encoded second media data, and send the encoded second media data to the called domain server; The calling domain server is used to send the encoded first media data to the called domain server, so that the encoded first media data is sent to the second terminal through the called domain server; the called domain server is used to send the encoded second media data to the calling domain server, so that the encoded second media data is sent to the first terminal through the calling domain server.
3. The system according to claim 1, characterized in that The target media format is determined by: The target domain server is used to determine at least one common media format commonly supported by the first terminal and the second terminal according to each of the first media formats and each of the second media formats; wherein the target domain server is a calling domain server and / or a called domain server; A target media format is determined from the at least one common media format.
4. The system according to claim 3, characterized in that At least a portion of the first identifiers corresponding to the first media format are configured with corresponding usage permissions; and at least a portion of the second identifiers corresponding to the second media format are configured with corresponding usage permissions.
5. The system according to claim 4, characterized in that The target domain server is used to: At least one common media format commonly supported by the first terminal and the second terminal is determined according to the usage rights corresponding to each of the first media formats and each of the first identifiers, and the usage rights corresponding to each of the second media formats and each of the second identifiers.
6. The system according to claim 1, characterized in that The target media format is one of the first media formats supported by the first terminal; the second terminal is in a state of being called by another terminal; The first terminal is used to send a third pull request to the calling domain server, and the third pull request is sent to the second terminal through the calling domain server and the called domain server in sequence; wherein the third pull request carries a third target identifier, and the third target identifier is an identifier of a target media format corresponding to the second terminal; The second terminal is used to encode the third media data to be sent according to the target media format to obtain the encoded third media data, and send the encoded third media data to the first terminal through the called domain server and the calling domain server.
7. The system according to claim 1, characterized in that There are multiple called domain servers, and there are multiple second terminals. Each called domain server is in communication connection with a corresponding second terminal. The target media format is specified by a user through the first terminal. The calling domain server is used to send a fourth pull request to the first terminal; wherein the fourth pull request carries a fourth target identifier; the fourth target identifier is an identifier of a target media format corresponding to the first terminal; The first terminal is used for encoding the fourth media data to be sent according to the target media format to obtain the encoded fourth media data, and sending the encoded fourth media data to the calling domain server; Each of the called domain servers is used to send a fifth pull request to the corresponding second terminal; wherein the fifth pull request carries a fifth target identifier; the fifth target identifier is an identifier of a target media format corresponding to the second terminal; Each of the second terminals is used to encode the fifth media data to be sent according to the target media format to obtain the encoded fifth media data, and send the encoded fifth media data to the called domain server corresponding to each terminal; The calling domain server is used to send the encoded fourth media data to each of the called domain servers, so as to send the encoded fourth media data to the corresponding second terminal through each of the called domain servers; each of the called domain servers is used to send the encoded fifth media data to the calling domain server, so as to send the encoded fifth media data to the first terminal and / or other second terminals except the current second terminal corresponding to the current called domain server through the calling domain server.
8. The system according to claim 1, characterized in that The user corresponding to the first terminal is an anonymous user or a designated user.
9. The system according to claim 1, characterized in that The user corresponding to the second terminal is an anonymous user or a designated user.
10. An audio and video communication method, characterized in that: The method is applied to an audio and video communication system, the system comprising: a first terminal, a calling domain server, a called domain server, and a second terminal which are sequentially connected in communication; the calling domain server and the called domain server both store a first media format supported by the first terminal and a second media format supported by the second terminal; each first media format supported by the first terminal has a first identifier corresponding to each; each second media format supported by the second terminal has a second identifier corresponding to each; each first identifier and each second identifier are different; the method comprises: The first terminal performs audio and video communication with the second terminal through the calling domain server and the called domain server based on the identifier corresponding to the target media format.