Audio and video communication method between remote terminals, medium and system
By establishing a long websocket link on the target service node, the real-time problem caused by duplicate messages in audio and video communication is solved, and more efficient audio and video communication is achieved.
Patent Information
- Application Number
- CN202510435127.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-04
AI Technical Summary
The existing audio and video communication schemes generate a large number of repeated and useless messages during the broadcasting process, resulting in an increase in the transmission link of the message and affecting the real-time nature of the communication.
By establishing a long websocket link on the target service node, avoiding forwarding of messages between service nodes, and creating audio and video links through the same audio and video media service coturn address, exchanging audio and video signaling data, ensuring that the long websocket connections between the initiator and the receiver are allocated to the same websocket service.
It significantly reduces the transmission delay of signaling, improves the real-time and stability of audio and video communication, optimizes system resource utilization, and improves user experience.
Smart Images

Figure CN120263775A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and more particularly, to a method for audio and video communication between remote terminals, a computer-readable storage medium, and a system for audio and video communication between remote terminals. Background Art
[0002] In the audio and video communication among multiple terminals based on WebRTC, it involves the communication transmission problems of audio and video signaling and audio and video media streams among terminals.
[0003] When performing audio and video signaling transmission, in order for one terminal to perceive the signaling messages of another terminal, it can avoid polling to obtain the messages of the other party from the backend service by timing. To avoid the performance and latency problems brought by polling, a long connection websocket between the client and the server is often established for real-time communication. The backend service is generally deployed in a distributed manner. Since the sessions of the long connection are not shared, it often involves the forwarding of messages between backend services.
[0004] When performing audio and video stream transmission, a turnserver can be set up to transmit the media stream, and a front-end Load Balancer Server is used for load balancing. Due to the randomness of the service nodes loaded by each end, it also involves the forwarding of audio and video media streams among service nodes. At the same time, the commonly used ip hash load policy cannot ensure the overall balance of service node resource usage.
[0005] That is, the existing solutions generate a large number of duplicate and useless messages during the broadcast process, which has a negative impact on the stability and performance of the system. Secondly, the transmission link of the messages will also increase, and the transmission time of the messages within the system will also be extended, affecting the real-time performance of the communication. Summary of the Invention
[0006] The main purpose of this application is to provide a method for audio and video communication between remote terminals, a computer-readable storage medium, and a system for audio and video communication between remote terminals, so as to at least solve the problem that the existing solutions generate a large number of duplicate and useless messages during the broadcast process, resulting in an increase in the transmission link of the messages, and further affecting the real-time performance of the communication.
[0007] To achieve the above object, according to one aspect of the present application, there is provided a method for audio and video communication between remote terminals, the method comprising: receiving an audio and video communication request sent by an initiating end, creating a room number for audio and video communication between the initiating end and a receiving end according to the audio and video communication request, and determining a target service node, the audio and video communication request including a call type, a call time, and a unique identifier of the initiating end; after sending the address of the target service node and the room number to the initiating end, receiving a long link construction request sent by the initiating end, and creating a websocket long link between the initiating end and the receiving end based on the long link request and the room number; creating an audio and video link through the address of the target service node, and exchanging at least signaling data of audio and video through the created websocket long link to establish audio and video communication between the initiating end and the receiving end.
[0008] Optionally, determining the target service node includes: determining whether a websocket service node has been allocated to the room number in redis to obtain a query result; in the case where the query result indicates that a websocket service node has been allocated to the room number, performing a first processing method, the first processing method indicating that the target service node is determined to be the allocated websocket service node; in the case where the query result indicates that a websocket service node has not been allocated to the room number, performing a second processing method, the second processing method indicating pulling a list of websocket service nodes from a configuration center, and determining the target service node at least according to the current allocated link number of the websocket service nodes in the websocket service node list.
[0009] Optionally, determining the target service node at least according to the current allocated link number of the websocket service nodes in the websocket service node list includes: determining the websocket service node corresponding to the minimum value of the current allocated link number as the target service node.
[0010] Optionally, determining the target service node at least according to the current allocated link number of the websocket service nodes in the websocket service node list includes: determining a websocket service node to be allocated, the websocket service node to be allocated being the websocket service node whose current allocated link number is less than a preset allocated link number; determining the target service node according to the priority of the websocket service node.
[0011] Optionally, determining the target service node at least according to the current allocated link numbers of the websocket service nodes in the websocket service node list includes: determining the websocket service nodes to be allocated, where the websocket service nodes to be allocated are the websocket service nodes with the current allocated link numbers less than the preset allocated link numbers; determining the target service node according to the number of processing cores of the websocket service nodes.
[0012] Optionally, creating a websocket long link between the initiating end and the receiving end based on the long link request and the room number includes: determining the identifier of the room number in the header of the long link request; determining an index in a service node list according to the hash value of the identifier of the room number, and using the websocket service node corresponding to the index in the service node list as the target node for processing the long link request;
[0013] Forwarding the long link request to the target node so that requests with the same identifier of the room number are forwarded to the same websocket service node for processing.
[0014] Optionally, before receiving an audio and video communication request sent by the initiating end, the method further includes: registering the coturn service to the configuration center in the manner of the open api of the configuration center and registering it to the coturn service when starting the coturn service; removing the coturn service from the configuration center when the coturn service stops; after creating the websocket long link between the initiating end and the receiving end, the method further includes: storing the relationship of the sessions corresponding to the receiving end and the initiating end in the local memory of the websocket service.
[0015] According to another aspect of the present application, a method for audio and video communication between remote terminals is provided. The method includes: sending an audio and video communication request to a receiving end, so that the receiving end creates a room number for audio and video communication between the initiating end and the receiving end according to the audio and video communication request, and determines a target service node, where the audio and video communication request includes a call type, a call time, and a unique identifier of the initiating end; after receiving the address of the target service node and the room number sent by the receiving end, sending a long link construction request to the receiving end, so that the receiving end creates a websocket long link between the initiating end and the receiving end based on the long link construction request and the room number, so that the receiving end creates an audio and video link through the address of the target service node, and exchanges at least signaling data of audio and video through the created websocket long link to establish audio and video communication between the initiating end and the receiving end.
[0016] According to still another aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the above methods.
[0017] According to yet another aspect of the present application, a communication system for audio and video between remote terminals is provided. The system includes: an initiating end and a receiving end, the receiving end executes any one of the above methods, and the initiating end executes the above method.
[0018] Applying the technical solution of the present application, through the target service node, the websocket long connections of each end that need to establish audio and video are allocated to the same websocket service, thereby avoiding the forwarding of messages between service nodes. An audio and video link is created through the same audio and video media service coturn address, and signaling data of audio and video is exchanged through the created websocket long link to establish audio and video between both parties, thus completing the construction of audio and video communication, and further solving the problem that a large number of duplicate and useless messages are generated in the existing solution during the broadcast process, resulting in an increase in the message transmission link, and further affecting the real-time performance of communication. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The specification drawings constituting a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0020] Figure 1 A flowchart showing a first method for audio and video communication between remote terminals provided according to an embodiment of the present application is shown;
[0021] Figure 2 The communication link structure diagram of audio and video signaling and media streams provided according to an embodiment of the present application is shown;
[0022] Figure 3 The schematic flow diagram of establishing a long connection with a websocket service node provided according to an embodiment of the present application is shown;
[0023] Figure 4 The schematic flow diagram of establishing a long connection with a websocket service node when viewing that the room number is not assigned to a websocket service node in redis provided according to an embodiment of the present application is shown;
[0024] Figure 5 The structural block diagram of a receiving end provided according to an embodiment of the present application is shown;
[0025] Figure 6 The schematic flow diagram of the second communication method of audio and video between remote terminals provided according to an embodiment of the present application is shown;
[0026] Figure 7 The structural block diagram of a sending end provided according to an embodiment of the present application is shown. Detailed implementation manners
[0027] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0028] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0030] For ease of description, some nouns or terms related to the embodiments of this application are described below:
[0031] Redis is an open-source in-memory data structure storage system, commonly used as a database, cache, and message broker. It supports various data structures such as strings, hashes, lists, sets, and sorted sets. Redis is known for its high performance, flexibility, and rich functionality, and is particularly suitable for application scenarios that require rapid access and processing of large amounts of data.
[0032] As introduced in the background art, when the prior art performs audio and video stream transmission, it can transmit media streams by setting up a turnserver and perform load balancing through a front-end Load Balancer Server. Due to the randomness of the service nodes loaded by each end, it also involves the forwarding of audio and video media streams between various service nodes. At the same time, the common ip hash load balancing strategy cannot ensure the overall balance of service node resource usage. To solve the problem that a large number of duplicate and useless messages are generated during the broadcast process in the existing solution, resulting in an increase in the message transmission link and thus affecting the real-time performance of communication, the embodiments of this application provide a method for audio and video communication between remote terminals, a computer-readable storage medium, and a system for audio and video communication between remote terminals.
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0034] In this embodiment, a method for audio and video communication between remote terminals is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0035] Figure 1 is a schematic flowchart of the first method for audio and video communication between remote terminals provided according to the embodiments of this application. As Figure 1 shown, the method includes the following steps:
[0036] Step S101, receive an audio and video communication request sent by the initiating end, create a room number for audio and video communication between the initiating end and the receiving end according to the above audio and video communication request, and determine the target service node. The above audio and video communication request includes a call type, a call time, and a unique identifier of the initiating end;
[0037] Among them, determining the target service node includes: determining whether the above room number has been assigned a websocket service node in redis to obtain a query result; in the case that the above query result indicates that the above room number has been assigned the above websocket service node, performing a first processing method, where the first processing method indicates determining the above target service node as the assigned websocket service node; in the case that the above query result indicates that the above room number has not been assigned the above websocket service node, performing a second processing method, where the second processing method indicates pulling a list of websocket service nodes from the configuration center and determining the above target service node at least according to the current assigned link numbers of the websocket service nodes in the above websocket service node list.
[0038] For the room numbers with assigned service nodes, directly locate to the corresponding service nodes, which avoids unnecessary resource search and allocation processes, saves system resources, and at the same time reduces the latency caused by resource allocation, improving the efficiency and response speed of audio and video communication. For the room numbers without assigned service nodes, pull the service node list from the configuration center and determine the target service node based on the current link numbers of each node. This method ensures the load balance between service nodes. The service nodes with fewer link numbers are preferentially selected, effectively preventing the situation where some nodes are overloaded while others are idle, maximizing the utilization of the overall system performance. By avoiding frequent migrations and data forwarding between service nodes, it reduces the uncertainty of cross-node communication and the probability of communication failures caused by network fluctuations or service instability, thus enhancing the stability and reliability of the audio and video communication system.
[0039] Among them, determining the above target service node at least according to the current assigned link numbers of the websocket service nodes in the above websocket service node list has three specific embodiments:
[0040] The first specific embodiment: determining the websocket service node corresponding to the minimum value of the above current assigned link numbers as the above target service node.
[0041] Selecting the service node with the fewest link numbers as the target node can effectively disperse link requests, avoid resource tension or overload caused by some nodes bearing too many links, thus achieving load balance between service nodes and ensuring the stable operation of the overall system. Service nodes with fewer link numbers usually mean a lower processing burden. Therefore, these nodes can respond to new link requests faster, reducing user waiting time and enhancing the real-time performance and user experience of audio and video communication.
[0042] The second specific embodiment: Determine the websocket service node to be allocated, where the websocket service node to be allocated is the websocket service node whose current allocated link number is less than the preset allocated link number; Determine the target service node according to the priority of the websocket service node.
[0043] By setting the threshold of the pre-allocated link number, the system can more accurately identify which service nodes still have processing capacity, thus avoiding the extensive management of evenly distributing links among all nodes, improving the efficiency and accuracy of load balancing. Selecting service nodes with a current allocated link number less than the preset allocated link number effectively utilizes system resources, ensures that service nodes can handle more connections before reaching their maximum processing capacity, avoids resource waste, and at the same time reduces the risk of overload. Introducing a priority scheduling mechanism can determine the priority of service nodes according to characteristics such as the performance, stability, or geographical location of the service nodes, so as to preferentially select service nodes with a higher priority in high-concurrency scenarios, providing users with higher-quality audio and video communication services. Preferentially selecting service nodes with unsaturated processing capacity can reduce the waiting time of connections, make the establishment of audio and video communication more rapid, and at the same time reduce the delay of signaling transmission and media stream forwarding, improving the real-time performance of communication.
[0044] The third specific embodiment: Determine the websocket service node to be allocated, where the websocket service node to be allocated is the websocket service node whose current allocated link number is less than the preset allocated link number; Determine the target service node according to the number of processing cores of the websocket service node.
[0045] Select the WebSocket service node with the fewest connections as the target service node, which can ensure that newly added long connections are allocated to nodes with lighter loads, avoid over - concentration of resources on service nodes, and achieve optimal allocation and use of resources. By dynamically monitoring the connection numbers and the number of processing cores of each service node, the system can adjust the allocation strategy in real - time to ensure that the load among service nodes is always balanced, prevent the emergence of hot - spot nodes, and guarantee the stable operation of the entire audio - video communication system. Service nodes with fewer connections usually have lower latency because they handle relatively fewer requests. This makes the transmission of audio - video signaling faster, significantly improving the real - time performance and response speed of audio - video communication. Incorporating the number of processing cores, which is an indicator of the processing capacity of service nodes, into the determining factors for target service nodes can ensure that the quality of audio - video communication is guaranteed even in high - concurrency scenarios, avoiding the risk of service quality degradation that may be brought about by a single indicator (such as the number of connections). When a service node fails or needs maintenance, healthy nodes with more processing cores can take on the additional load faster, ensuring the continuity and coherence of audio - video communication and reducing the impact of failures on the user experience.
[0046] In an embodiment of the present application, before receiving an audio - video communication request sent by the initiating end, the above - mentioned method further includes: registering the Coturn service to the above - mentioned configuration center through the open API of the configuration center and registering it to the above - mentioned Coturn service when the Coturn service is started; removing the above - mentioned Coturn service from the above - mentioned configuration center when the above - mentioned Coturn service stops; after creating the WebSocket long connection between the above - mentioned initiating end and the above - mentioned receiving end, the above - mentioned method further includes: storing the relationship of the sessions corresponding to the above - mentioned receiving end and the above - mentioned initiating end in the local memory of the WebSocket service.
[0047] The dynamic registration and cancellation of the coturn service are realized through the configuration center, which means that the system can perceive the status changes of the coturn service in real time. When the service starts or stops, the configuration center can update the service list in a timely manner to ensure that all relevant components (such as load balancers) can obtain the latest service information, improving the adaptability and flexibility of the system. The dynamic registration mechanism is combined with the load balancing strategy, enabling the system to forward audio and video streams to the best service node according to the real-time availability and load of the coturn service. Through the session relationships stored in local memory, it can be quickly determined which coturn service nodes are carrying the communication of a specific room, thus more effectively performing load balancing and avoiding blind allocation of links and resource waste. Storing the session relationships between the receiving end and the initiating end in local memory ensures that once the websocket long connection is established, subsequent signaling exchanges and media stream transmissions will follow the same path without having to re-find service nodes or re-establish connections. This not only reduces communication latency but also enhances the stability and coherence of communication, providing users with a better audio and video communication experience. This method based on dynamic registration and storing session relationships simplifies the operation and maintenance management of the audio and video communication system. Operation and maintenance personnel do not need to manually intervene in the addition or removal of service nodes, and the system can automatically adjust to cope with the increase or decrease of nodes, reducing the complexity and labor cost of operation and maintenance and improving the operation and maintenance efficiency.
[0048] Register the coturn service to the above configuration center through the open api of the configuration center, and register it to the above coturn service when starting the coturn service; when the above coturn service stops, remove the above coturn service from the above configuration center. A specific usage scenario: In a system that supports large-scale real-time audio and video communication conferences, end users need to establish a websocket long connection with the system to perform real-time audio and video data exchange. Due to the high bandwidth requirements and real-time requirements of audio and video data, the system must be able to efficiently and stably handle a large number of concurrent connections and data transmissions. The coturn service, as the media relay service of the system, is responsible for forwarding audio and video stream data, and its performance and availability directly affect the quality of the conference and the user experience. By dynamically registering the coturn service through the Open API of the configuration center, the service automatically registers with the configuration center when starting, and automatically removes from the configuration center when stopping. This mechanism enables the configuration center to update the service list in real time, ensuring that all requests can be correctly routed to available service nodes, achieving dynamic service discovery. At the same time, by real-time monitoring the status and load of service nodes, the system can intelligently allocate requests to service nodes with lighter loads, achieving load balancing, and improving the overall processing capacity and response speed of the system. During the conference, if a certain coturn service node fails, the configuration center can immediately remove the node from the service list, preventing the node from continuing to receive new media stream data. At the same time, other healthy service nodes will automatically assume more loads, ensuring the continuity and stability of the conference. This mechanism improves the high availability of the system and enhances the disaster tolerance ability. The dynamic registration mechanism allows the system to dynamically adjust the number and allocation of coturn service nodes according to actual needs, avoiding waste of resources during low load and insufficient resources during high load. The dynamic increase and decrease of service nodes enables the system to always be in the optimal resource utilization state, improving resource utilization rate and reducing operating costs. The characteristics of automatic registration and cancellation reduce the management burden of operation and maintenance personnel. Without manually maintaining the service node list, the system can automatically adapt to the increase and decrease of service nodes, simplifying the operation and maintenance process, reducing maintenance costs, and improving operation and maintenance efficiency. The dynamic load balancing and high availability mechanism ensure the stable transmission of audio and video data. Even in high-concurrency scenarios, it can maintain high-quality audio and video communication, reduce latency, avoid conference interruption, and provide users with a smooth and high-quality audio and video conference experience.
[0049] Step S102, after sending the address of the above target service node and the above room number to the above initiating end, receive the long connection request sent by the above initiating end, and create a websocket long connection between the above initiating end and the above receiving end based on the above long connection request and the above room number;
[0050] Among them, based on the above long - link request and the above room number, a WebSocket long - link between the above initiating end and the above receiving end is created, including: determining the identifier of the above room number in the header of the above long - link request; determining an index in a service node list according to the hash value of the identifier of the above room number, and using the WebSocket service node corresponding to the above index in the above service node list as the target node for processing the above long - link request; forwarding the above long - link request to the above target node, so that requests with the same identifier of the room number are forwarded to the same WebSocket service node for processing.
[0051] Through the association of the hash value and the index, it is ensured that requests with the same room - number identifier are always forwarded to the same WebSocket service node for processing, achieving the consistency and stability of the connection. This is crucial for audio - video communication because it avoids unnecessary cross - node communication, reduces transmission latency, and ensures the smoothness of the call. Since the requests of all participants in the same room are directed to the same service node, session management becomes simpler and more effective. The service node only needs to focus on the room numbers it manages and does not need to process session information from other nodes, reducing the complexity and potential errors of session management. Although requests with the same room number are directed to the same service node, by pre - calculating the hash value of the room - number identifier to determine the index in the service node list, a hash - based load balancing is actually achieved. This can not only ensure the centralized processing of requests in the same room but also reasonably disperse the total request volume through the hash distribution algorithm, avoiding excessive load on individual service nodes. By reducing the number of cross - node forwards, as well as optimizing session management and load balancing, this strategy can effectively improve the overall throughput of the system. Audio - video communication data can be transmitted and processed in a shorter time, enhancing the processing capacity and response speed of the system.
[0052] Step S103, create an audio - video link through the address of the above target service node, and exchange at least the signaling data of the audio - video through the created above WebSocket long - link to establish the audio - video communication between the above initiating end and the above receiving end.
[0053] In the above steps, through the target service node, the WebSocket long - connections of each end that needs to establish audio - video are assigned to the same WebSocket service, thus avoiding the forwarding of messages between service nodes. Create an audio - video link through the same audio - video media service coturn address, and exchange the signaling data of the audio - video through the created WebSocket long - link to establish the audio - video between both parties, thus completing the construction of the audio - video communication, and further solving the problem that a large number of duplicate and useless messages are generated in the existing solution during the broadcast process, resulting in an increase in the message transmission link, and further affecting the real - time performance of the communication.
[0054] This application avoids redundant forwarding of audio and video signaling between distributed backend services by ensuring that the websocket long links of the initiator and the receiver are established on the same service node, thereby significantly reducing the transmission delay of the signaling and improving the real-time performance of audio and video communication. By determining the target service node and using a custom load balancing strategy, system resources can be allocated more reasonably. By selecting the service node with the least number of current links, it is possible to avoid overloading of a single node, maintain balanced resource utilization of the service node, and ensure stable operation of the system. It avoids the use of redis clusters or message queues for message transfer, simplifying the system architecture. At the same time, by concentrating the audio and video streams in the same room to the same coturn service node for forwarding, the media stream transmission across nodes is reduced, and the efficiency of data processing and transmission is improved. Shorter delays, more stable links, and higher transmission efficiency directly improve the quality of audio and video communications, providing users with a smoother and higher-quality audio and video call experience. Through dynamic registration and load balancing strategies, the system can better adapt to the needs of high-concurrency scenarios and increase the scalability of audio and video communications. At the same time, the dynamic selection and allocation of service nodes also enhances the flexibility and fault tolerance of the entire system.
[0055] By customizing the routing algorithm of the forwarding service, it is possible to allocate the long websocket connections of each end that needs to establish audio and video to the same websocket service. This avoids the forwarding of messages between service nodes, improves the efficiency of message transmission and shortens the transmission time of messages. The audio and video service nodes are dynamically registered to the service registration center to ensure the high availability of media services. At the same time, the streaming media data of all terminals in the same room are forwarded to the same media service node, and a load balancing strategy is established based on the number of online rooms allocated to each node to ensure that the overall use of service node resources remains balanced. This load balancing mechanism improves the stability and efficiency of streaming media transmission.
[0056] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the audio and video communication method between remote terminals of the present application will be described in detail below in combination with specific embodiments.
[0057] This embodiment relates to a specific method for audio and video communication between remote terminals, such as Figure 2 As shown, including:
[0058] The terminal that needs to conduct audio and video communication initiates a call request to the room service, including: call type, call time, and calling device identification.
[0059] The room service matches according to the corresponding business rules and creates a room number for both parties to conduct audio and video communication. It obtains the media service list from the configuration center nacos and selects a streaming media service node through a custom load balancing strategy.
[0060] After the terminal obtains the assigned room number and the address of the audio and video service node, the terminal requests the signaling service to establish a websocket long connection, sets the room number in the header of the request, and through a custom load balancing strategy when passing through the forwarding service, forwards the requests with the same room number to the same websocket server node.
[0061] After successfully establishing the connection, save the relationship of the session corresponding to the communication between the terminal device and the signaling service in the local memory of the websocket server.
[0062] The terminal (terminal 1 or terminal 2) creates a peerconnection audio and video link through the same audio and video media service coturn address, and exchanges audio and video sdp signaling data and service data through the established websocket long connection to establish the audio and video between both parties.
[0063] The specific implementation method of routing the websocket links established between the devices in the same room and the signaling service to the same service node is as follows:
[0064] The forwarding service can implement the forwarding of custom policies for websocket through a service component similar to Spring Cloud Gateway. The component defaults to using a polling strategy for forwarding requests:
[0065] That is, by inheriting the ReactorServiceInstanceLoadBalancer load balancing parent class and overriding Mono<Response <serviceinstance>> The "choose(Request request)" method is used to customize the forwarding strategy to forward requests with the same room number to the same node. For example, Figure 3 as shown below, the specific strategy is as follows:
[0066] The forwarding service obtains the room number from the information sent by the header terminal of the request;
[0067] Check in redis whether a websocket service node has been assigned to the room number;
[0068] In the case of an assigned node, query the websocket service node corresponding to the room number and establish a long connection with this websocket service node;
[0069] For example, Figure 4 as shown, in the case of an unassigned node, pull the list of websocket service nodes from the nacos configuration center, find the node with the fewest assigned links from the list, record the relationship between the room number and the websocket service node in redis, and increase the number of long connections assigned to this service node, and establish a long connection with this websocket service node.
[0070] The implementation method for routing audio and video media streams to the same media service coturn node:
[0071] 1) When the coturn service starts, register the service to nacos through the open api of nacos.
[0072] 2) When a call request is received, pull the service list of coturn from nacos and select the service node with the fewest assigned audio and video links.
[0073] 3) Each terminal uses this coturn service address as the service node for media forwarding for terminals in the same room.
[0074] Coturn can perform load balancing through a front-end Load Balancer Server. The present invention customizes the load balancing strategy by pulling the service list from nacos and selecting the service node with the fewest assigned links as the service node for media forwarding for terminals in the same room.
[0075] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0076] The embodiments of the present application also provide a receiving end. It should be noted that the receiving end of the embodiments of the present application can be used to execute the method for audio and video communication between remote terminals provided by the embodiments of the present application. The device for implementing the above embodiments and preferred embodiments has been described and will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0077] The receiving end provided by the embodiments of the present application will be introduced below.
[0078] Figure 5 It is a structural block diagram of a receiving end provided according to an embodiment of the present application. As Figure 5 shown, the device includes:
[0079] A first processing unit 51, configured to receive an audio and video communication request sent by an initiating end, create a room number for audio and video communication between the initiating end and the receiving end according to the audio and video communication request, and determine a target service node, where the audio and video communication request includes a call type, a call time, and a unique identifier of the initiating end;
[0080] A second processing unit 52, configured to receive a long link construction request sent by the initiating end after sending the address of the target service node and the room number to the initiating end, and create a websocket long link between the initiating end and the receiving end based on the long link request and the room number;
[0081] A third processing unit 53, configured to create an audio and video link through the address of the target service node, and exchange at least signaling data of audio and video through the created websocket long link to establish audio and video communication between the initiating end and the receiving end.
[0082] The above receiving end distributes the websocket long connections of each end that needs to establish audio and video to the same websocket service through the target service node, thereby avoiding the forwarding of messages between service nodes, creating an audio and video link through the same audio and video media service coturn address, and exchanging signaling data of audio and video through the created websocket long link to establish the audio and video of both parties, thus completing the construction of audio and video communication, and further solving the problem that a large number of duplicate and useless messages are generated in the existing solution during the broadcast process, resulting in an increase in the transmission link of messages, and further affecting the real-time performance of communication.
[0083] In an embodiment of the present application, the first processing unit includes a first processing module, a second processing module, and a third processing module. The first processing module is configured to determine whether the above room number has been assigned to a websocket service node in redis to obtain a query result; the second processing module is configured to execute a first processing method when the query result indicates that the above room number has been assigned to the above websocket service node, and the first processing method indicates determining the above target service node as the assigned websocket service node; the third processing module is configured to execute a second processing method when the query result indicates that the above room number has not been assigned to the above websocket service node, and the second processing method indicates pulling a list of websocket service nodes from the configuration center and determining the above target service node at least according to the current assigned link number of the websocket service nodes in the above list of websocket service nodes.
[0084] In an embodiment of the present application, the third processing module includes a first determination sub-module, configured to determine the websocket service node corresponding to the minimum value of the above current assigned link number as the above target service node.
[0085] In an embodiment of the present application, the third processing module includes a second determination sub-module and a third determination sub-module. The second determination sub-module is configured to determine a to-be-assigned websocket service node, and the to-be-assigned websocket service node is the above websocket service node whose current assigned link number is less than a preset assigned link number; the third determination sub-module is configured to determine the above target service node according to the priority of the above websocket service node.
[0086] In an embodiment of the present application, the third processing module includes a fourth determination sub-module and a fifth determination sub-module. The fourth determination sub-module is configured to determine a to-be-assigned websocket service node, and the to-be-assigned websocket service node is the above websocket service node whose current assigned link number is less than a preset assigned link number; the fifth determination sub-module is configured to determine the above target service node according to the number of processing cores of the above websocket service node.
[0087] In an embodiment of the present application, the second processing unit includes a fourth processing module, a fifth processing module, and a sixth processing module. The fourth processing module is used to determine the identifier of the room number in the header of the above long link request; the fifth processing module is used to determine an index in a service node list according to the hash value of the identifier of the room number, and use the websocket service node corresponding to the index in the service node list as the target node for processing the long link request; the sixth processing module is used to forward the long link request to the target node, so that requests with the same identifier of the room number are forwarded to the same websocket service node for processing.
[0088] In an embodiment of the present application, the above receiving end further includes a fourth processing unit, a fifth processing unit, and a sixth processing unit. The fourth processing unit is used to register the coturn service to the above configuration center by means of the open api of the configuration center before receiving the audio and video communication request sent by the initiating end, and register it to the above coturn service when starting the coturn service; the fifth processing unit is used to remove the coturn service from the above configuration center when the above coturn service stops; the sixth processing unit is used to store the relationship of the sessions corresponding to the receiving end and the initiating end in the local memory of the websocket service after creating the websocket long link between the initiating end and the receiving end.
[0089] The above receiving end includes a processor and a memory. The first processing unit, the second processing unit, the third processing unit, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions. The above modules are all located in the same processor; or, the above each module is located in different processors in any combination form.
[0090] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem that a large number of duplicate and useless messages are generated in the existing solution during the broadcast process, resulting in an increase in the message transmission link and thus affecting the real-time performance of communication, can be solved.
[0091] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0092] An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the method for audio and video communication between remote terminals.
[0093] An embodiment of the present invention provides a processor. The processor is used to run a program. When the program runs, it executes the method for audio and video communication between remote terminals.
[0094] An embodiment of the present invention provides a device. The device includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements at least the following steps: receiving an audio and video communication request sent by an initiating end, creating a room number for audio and video communication between the initiating end and a receiving end according to the audio and video communication request, and determining a target service node. The audio and video communication request includes a call type, a call time, and a unique identifier of the initiating end; after sending the address of the target service node and the room number to the initiating end, receiving a long link construction request sent by the initiating end, and creating a websocket long link between the initiating end and the receiving end based on the long link request and the room number; creating an audio and video link through the address of the target service node, and exchanging at least signaling data of audio and video through the created websocket long link to establish audio and video communication between the initiating end and the receiving end. The device herein can be a server, a PC, a PAD, a mobile phone, etc.
[0095] The present application also provides a computer program product. When executed on a data processing device, it is adapted to execute a program initialized with at least the following method steps: receiving an audio and video communication request sent by an initiating end, creating a room number for audio and video communication between the initiating end and a receiving end according to the audio and video communication request, and determining a target service node. The audio and video communication request includes a call type, a call time, and a unique identifier of the initiating end; after sending the address of the target service node and the room number to the initiating end, receiving a long link construction request sent by the initiating end, and creating a websocket long link between the initiating end and the receiving end based on the long link request and the room number; creating an audio and video link through the address of the target service node, and exchanging at least signaling data of audio and video through the created websocket long link to establish audio and video communication between the initiating end and the receiving end.
[0096] The present application also provides a second method for audio and video communication between remote terminals, as Figure 6 shown. The method includes the following steps:
[0097] Step S601: Send an audio-video communication request to the receiving end, so that the receiving end creates a room number for audio-video communication between the initiating end and the receiving end according to the audio-video communication request, and determines a target service node. The audio-video communication request includes a call type, a call time, and a unique identifier of the initiating end.
[0098] Step S602: After receiving the address of the target service node and the room number sent by the receiving end, send a long connection construction request to the receiving end, so that the receiving end creates a websocket long connection between the initiating end and the receiving end based on the long connection construction request and the room number, so that the receiving end creates an audio-video link through the address of the target service node, and exchanges at least signaling data of the audio-video through the created websocket long connection to establish audio-video communication between the initiating end and the receiving end.
[0099] In the above steps, through the target service node, the websocket long connections of each end that needs to establish audio-video are allocated to the same websocket service, thus avoiding the forwarding of messages between service nodes. An audio-video link is created through the same audio-video media service coturn address, and signaling data of the audio-video is exchanged through the created websocket long connection to establish the audio-video of both parties, thereby completing the construction of audio-video communication, and further solving the problem that a large number of duplicate and useless messages are generated in the existing solution during the broadcast process, resulting in an increase in the message transmission link, and further affecting the real-time performance of communication.
[0100] Among them, the receiving end determines the target service node, including: determining whether the room number has been allocated a websocket service node in redis to obtain a query result; in the case where the query result indicates that the room number has been allocated the websocket service node, performing a first processing method, the first processing method indicates determining the target service node as the allocated websocket service node; in the case where the query result indicates that the room number has not been allocated the websocket service node, performing a second processing method, the second processing method indicates pulling a list of websocket service nodes from the configuration center, and determining the target service node at least according to the current allocated link quantity of the websocket service nodes in the list of websocket service nodes.
[0101] Among them, the receiving end determines the target service node at least according to the current allocated link number of the websocket service nodes in the above websocket service node list, including: determining the websocket service node corresponding to the minimum value of the current allocated link number as the target service node.
[0102] Among them, the receiving end determines the target service node at least according to the current allocated link number of the websocket service nodes in the above websocket service node list, including: determining the websocket service node to be allocated, where the websocket service node to be allocated is the websocket service node whose current allocated link number is less than the preset allocated link number; determining the target service node according to the priority of the above websocket service node.
[0103] Among them, the receiving end determines the target service node at least according to the current allocated link number of the websocket service nodes in the above websocket service node list, including: determining the websocket service node to be allocated, where the websocket service node to be allocated is the websocket service node whose current allocated link number is less than the preset allocated link number; determining the target service node according to the number of processing cores of the above websocket service node.
[0104] Among them, the receiving end creates a websocket long link between the initiating end and the receiving end based on the above long link request and the room number, including: determining the identifier of the room number in the header of the above long link request; determining an index in a service node list according to the hash value of the identifier of the room number, and using the websocket service node corresponding to the above index in the service node list as the target node for processing the above long link request; forwarding the above long link request to the above target node so that requests with the same identifier of the room number are forwarded to the same websocket service node for processing.
[0105] Among them, before the receiving end receives the audio and video communication request sent by the initiating end, the coturn service is registered to the above configuration center through the open api of the configuration center and registered to the above coturn service when the coturn service is started; when the above coturn service stops, the above coturn service is removed from the above configuration center; after creating the websocket long link between the initiating end and the receiving end, the method further includes: storing the relationship of the sessions corresponding to the receiving end and the initiating end in the local memory of the websocket service.
[0106] The present application also provides an initiating end, as Figure 7 shown. The initiating end includes:
[0107] A seventh processing unit 71, configured to send an audio-video communication request to a receiving end, so that the receiving end creates a room number for audio-video communication between the initiating end and the receiving end according to the audio-video communication request, and determines a target service node. The audio-video communication request includes a call type, a call time, and a unique identifier of the initiating end;
[0108] An eighth processing unit 72, configured to send a long connection building request to the receiving end after receiving the address of the target service node and the room number sent by the receiving end, so that the receiving end creates a websocket long connection between the initiating end and the receiving end based on the long connection building request and the room number, so that the receiving end creates an audio-video link through the address of the target service node, and exchanges at least signaling data of the audio-video through the created websocket long connection, so as to establish audio-video communication between the initiating end and the receiving end.
[0109] The above-mentioned initiating end distributes the websocket long connections of each end that needs to establish audio-video to the same websocket service through the target service node, thereby avoiding the forwarding of messages between service nodes. An audio-video link is created through the same audio-video media service coturn address, and signaling data of the audio-video is exchanged through the created websocket long connection to establish the audio-video of both parties, thus completing the construction of audio-video communication, and further solving the problem that a large number of duplicate and useless messages are generated in the existing solution during the broadcast process, resulting in an increase in the message transmission link, and further affecting the real-time performance of communication.
[0110] Wherein, the receiving end determines the target service node, including: determining whether the room number has been assigned a websocket service node in redis to obtain a query result; in the case that the query result indicates that the room number has been assigned the websocket service node, performing a first processing method, where the first processing method indicates determining the target service node as the assigned websocket service node; in the case that the query result indicates that the room number has not been assigned the websocket service node, performing a second processing method, where the second processing method indicates pulling a list of websocket service nodes from a configuration center, and determining the target service node at least according to the current assigned link quantity of the websocket service nodes in the list of websocket service nodes.
[0111] Among them, the receiving end determines the target service node at least according to the current allocated link number of the websocket service nodes in the above websocket service node list, including: determining the websocket service node corresponding to the minimum value of the above current allocated link number as the above target service node.
[0112] Among them, the receiving end determines the target service node at least according to the current allocated link number of the websocket service nodes in the above websocket service node list, including: determining the websocket service node to be allocated, where the websocket service node to be allocated is the websocket service node whose above current allocated link number is less than the preset allocated link number; determining the above target service node according to the priority of the above websocket service node.
[0113] Among them, the receiving end determines the target service node at least according to the current allocated link number of the websocket service nodes in the above websocket service node list, including: determining the websocket service node to be allocated, where the websocket service node to be allocated is the websocket service node whose above current allocated link number is less than the preset allocated link number; determining the above target service node according to the number of processing cores of the above websocket service node.
[0114] Among them, the receiving end creates a websocket long link between the initiating end and the receiving end based on the above long link request and the above room number, including: determining the identifier of the above room number in the header of the above long link request; determining an index in a service node list according to the hash value of the identifier of the above room number, and using the websocket service node corresponding to the above index in the above service node list as the target node for processing the above long link request; forwarding the above long link request to the above target node, so that requests with the same identifier of the room number are forwarded to the same websocket service node for processing.
[0115] Among them, before the receiving end receives the audio and video communication request sent by the initiating end, the coturn service is registered to the above configuration center through the open api of the configuration center, and is registered to the above coturn service when the coturn service is started; when the above coturn service stops, the above coturn service is removed from the above configuration center; after creating the websocket long link between the initiating end and the receiving end, the method further includes: storing the relationship of the sessions corresponding to the receiving end and the initiating end in the local memory of the websocket service.
[0116] The present application also provides a communication system for audio and video between remote terminals. The system includes: a sending end and a receiving end. The receiving end executes any of the above methods, and the sending end executes the above method.
[0117] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0118] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0119] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0120] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0121] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the process Figure 1 in one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.
[0122] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0123] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0124] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0125] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, commodity or device including the element.
[0126] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0127] 1) The method for audio and video communication between remote terminals of the present application distributes the websocket long connections of each end that needs to establish audio and video to the same websocket service through the target service node, thus avoiding the forwarding of messages between service nodes. An audio and video link is created through the same audio and video media service coturn address, and the signaling data of audio and video is exchanged through the created websocket long connection to establish the audio and video of both parties, thereby completing the construction of audio and video communication, and further solving the problem that a large number of duplicate and useless messages are generated in the existing solution during the broadcast process, resulting in an increase in the message transmission link, and further affecting the real-time performance of communication.
[0128] 2) The audio and video communication system between remote terminals of the present application distributes the websocket long connections of each end that needs to establish audio and video to the same websocket service through the target service node, thus avoiding the forwarding of messages between service nodes. An audio and video link is created through the same audio and video media service coturn address, and the signaling data of audio and video is exchanged through the created websocket long connection to establish the audio and video of both parties, thereby completing the construction of audio and video communication, and further solving the problem that a large number of duplicate and useless messages are generated in the existing solution during the broadcast process, resulting in an increase in the message transmission link, and further affecting the real-time performance of communication.
[0129] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.< / serviceinstance>
Claims
1. A method for audio and video communication between remote terminals, characterized in that, including: Receiving an audio - video communication request sent by an initiating end, creating a room number for audio - video communication between the initiating end and a receiving end according to the audio - video communication request, and determining a target service node, where the audio - video communication request includes a call type, a call time, and a unique identifier of the initiating end; After sending the address of the target service node and the room number to the initiating end, receiving a long - link construction request sent by the initiating end, and creating a websocket long - link between the initiating end and the receiving end based on the long - link request and the room number; Creating an audio - video link through the address of the target service node, and exchanging at least signaling data of the audio - video through the created websocket long - link to establish audio - video communication between the initiating end and the receiving end.
2. The method according to claim 1, wherein Determining a target service node includes: Determining whether a websocket service node has been assigned to the room number in redis to obtain a query result; When the query result indicates that a websocket service node has been assigned to the room number, performing a first processing method, where the first processing method indicates determining the target service node as the assigned websocket service node; When the query result indicates that a websocket service node has not been assigned to the room number, performing a second processing method, where the second processing method indicates pulling a list of websocket service nodes from a configuration center and determining the target service node at least according to the current allocated link number of the websocket service nodes in the websocket service node list.
3. The method according to claim 2, wherein Determining the target service node at least according to the current allocated link number of the websocket service nodes in the websocket service node list includes: Determining the websocket service node corresponding to the minimum value of the current allocated link number as the target service node.
4. The method according to claim 2, characterized in that, Determining the target service node at least according to the current allocated link number of the websocket service nodes in the websocket service node list includes: Determining a websocket service node to be allocated, where the websocket service node to be allocated is the websocket service node whose current allocated link number is less than a preset allocated link number; Determining the target service node according to the priority of the websocket service node.
5. The method according to claim 3, characterized in that Determining the target service node at least according to the current allocated link number of the websocket service nodes in the websocket service node list includes: Determining a websocket service node to be allocated, where the websocket service node to be allocated is the websocket service node whose current allocated link number is less than a preset allocated link number; Determining the target service node according to the number of processing cores of the websocket service node.
6. The method according to claim 1, wherein Create a WebSocket long connection between the initiating end and the receiving end based on the long connection request and the room number, including: Determine the identifier of the room number in the header of the long connection request; Determine an index in a service node list according to the hash value of the identifier of the room number, and use the WebSocket service node corresponding to the index in the service node list as the target node for processing the long connection request; Forward the long connection request to the target node, so that requests with the same identifier of the room number are forwarded to the same WebSocket service node for processing.
7. The method according to any one of claims 1 to 6, wherein Before receiving an audio and video communication request sent by the initiating end, the method further includes: registering the Coturn service to the configuration center in the manner of the open API of the configuration center, and registering it to the Coturn service when the Coturn service is started; when the Coturn service stops, removing the Coturn service from the configuration center; After creating the WebSocket long connection between the initiating end and the receiving end, the method further includes: storing the relationship of the sessions corresponding to the receiving end and the initiating end in the local memory of the WebSocket service.
8. A method for audio and video communication between remote terminals, characterized in that, including: Send an audio and video communication request to the receiving end, so that the receiving end creates a room number for audio and video communication between the initiating end and the receiving end according to the audio and video communication request, and determines the target service node, the audio and video communication request including the call type, the call time and the unique identifier of the initiating end; After receiving the address of the target service node and the room number sent by the receiving end, send a long connection request to the receiving end, so that the receiving end creates a WebSocket long connection between the initiating end and the receiving end based on the long connection request and the room number, so that the receiving end creates an audio and video link through the address of the target service node, and exchanges at least the signaling data of the audio and video through the created WebSocket long connection to establish the audio and video communication between the initiating end and the receiving end.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 7.
10. A communication system for audio and video between remote terminals, characterized in that, including: An initiating end and a receiving end, the receiving end executes the method according to any one of claims 1 to 7, and the initiating end executes the method according to claim 8.