Real-time communication method, computer-readable storage medium, and terminal device

By grouping the terminal devices in the network conference room and determining the main terminal, direct data transmission between terminal devices is realized, the problem of insufficient network bandwidth and computing power is solved, and network quality and the performance of terminal devices are improved.

CN115695705BActive Publication Date: 2025-08-26GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110846062.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-26
Publication Date
2025-08-26
Estimated Expiration
2041-07-26

AI Technical Summary

Technical Problem

As the number of access terminals increases, the demand for network bandwidth increases sharply, and the terminal's computing power is insufficient, resulting in communication lag and battery life decline.

Method used

Group the terminal devices connected to the network conference room, and determine the main terminals of each group, and directly transmit data between the main terminal and the main terminal to reduce the amount of data transmitted on the network and reduce network load.

Benefits of technology

Through terminal packets and main terminal data transmission, the amount of data transmitted on the network is significantly reduced, network quality is improved, and terminal computing power overload and battery life problems are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115695705B_ABST
    Figure CN115695705B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time communication method and terminal device, including: after grouping terminal devices accessing a network conference room, determining the main terminal device of each group; obtaining audio and video input information of the main terminal device in each group, and receiving first valid code stream information sent by other terminal devices in the corresponding group through the main terminal device in each group; performing effective judgment on the audio and video input information to obtain second valid code stream information; mixing the first valid code stream information and the second valid code stream information to obtain first mixed stream data, and sending the first mixed stream data to a cloud server. Thus, by grouping the terminal devices accessing the network conference room and determining the main terminal in the group, data transmission is directly carried out between the main terminals, thereby reducing the amount of data transmitted through the network, alleviating the network load, and improving the network quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a real-time communication method, a computer-readable storage medium, and a terminal device. Background Art

[0002] With the advancement of technology, more and more people are using remote conferencing to communicate, which requires excellent real-time audio and video technology. Currently, a common implementation is for each participating terminal device to connect to the same online conference room (registered on a cloud server). Information uploaded by any terminal is transmitted via the cloud server to all other terminals connected to the same conference room. The cloud server then receives and displays the information uploaded by all other terminals locally.

[0003] In real-time audio and video communication, if N terminals register to access a network conference room, there are a total of N uplink signals. Each uplink signal needs to pass through the cloud server before being transmitted to the other (N-1) terminals. Therefore, there are actually N*(N-1) signals transmitted on the network. Each terminal needs to process one uplink signal and N-1 downlink signals. As can be seen, as the number of connected terminals increases, the required network bandwidth increases dramatically.

[0004] When the number of participants is small, the existing network bandwidth and terminal computing power are fully capable of meeting the requirements of real-time audio and video communication. However, as the number of connected terminals continues to increase, the required network bandwidth and terminal computing power will increase, resulting in severe communication lag or even failure. Each terminal needs to receive the data streams from all other terminals and perform decoding and various post-processing. This requires a lot of computing power, which will seriously affect the battery life of handheld devices. Summary of the Invention

[0005] The present invention aims to at least partially address one of the technical problems in the related art. To this end, a first object of the present invention is to provide a real-time communication method that groups terminal devices accessing a network conference room, identifies a master terminal within each group, and directly transmits data between master terminals. This method reduces the amount of data transmitted over the network, alleviates network load, and improves network quality.

[0006] A second object of the present invention is to provide a computer-readable storage medium.

[0007] The third objective of the present invention is to provide a terminal device.

[0008] To achieve the above-mentioned purpose, the first aspect of the present invention proposes a real-time communication method, including: after grouping the terminal devices accessing the network conference room, determining the main terminal device of each group; obtaining the audio and video input information of the main terminal device in each group, and receiving the first valid code stream information sent by other terminal devices in the corresponding group through the main terminal device in each group; performing a valid judgment on the audio and video input information to obtain the second valid code stream information; mixing the first valid code stream information and the second valid code stream information to obtain first mixed stream data, and sending the first mixed stream data to the cloud server.

[0009] According to the real-time communication method of an embodiment of the present invention, after grouping the terminal devices accessing the network conference room, the main terminal device of each group is determined; the audio and video input information of the main terminal device in each group is obtained, and the first valid code stream information sent by other terminal devices in the corresponding group is received by the main terminal device in each group; the audio and video input information is effectively determined to obtain the second valid code stream information; the first valid code stream information and the second valid code stream information are mixed to obtain first mixed stream data, and the first mixed stream data is sent to the cloud server. Therefore, by grouping the terminal devices accessing the network conference room and determining the main terminal in the group, the method directly transmits data between the main terminals, thereby reducing the amount of data transmitted through the network, alleviating the network load, and improving the network quality.

[0010] In addition, the real-time communication method according to the above embodiment of the present invention may also have the following additional technical features:

[0011] According to one embodiment of the present invention, the method of grouping terminal devices accessing the network conference room includes one or more of the following: grouping the terminal devices accessing the network conference room according to the access gateway information of each terminal device; grouping the terminal devices accessing the network conference room according to the location information of each terminal device; grouping the terminal devices accessing the network conference room according to the audio input information of each terminal device; grouping the terminal devices accessing the network conference room according to the video input information of each terminal device.

[0012] According to one embodiment of the present invention, determining the main terminal device of each group includes: comprehensively determining the main terminal device of each group based on the signal strength between each terminal device in each group and the gateway, the performance and load conditions of each terminal device, and whether each terminal device is connected to a power source.

[0013] According to one embodiment of the present invention, all terminal devices in each group are in the same local area network.

[0014] According to one embodiment of the present invention, before obtaining the audio and video input information of the main terminal device, it is also detected whether the current user of the main terminal device is a speaker, and when the current user of the main terminal device is a listener, the acquisition of the audio and video input information of the main terminal device is stopped; or only the video input information of the main terminal device is obtained; or only the degraded video input information of the main terminal device is obtained.

[0015] According to one embodiment of the present invention, before the main terminal device in each group receives the first valid code stream information sent by other terminal devices in the corresponding group, it is also detected whether the current user of the other terminal device is a speaker, and when the current user of the other terminal device is a speaker, the audio and video input information of the other terminal device is effectively judged.

[0016] According to one embodiment of the present invention, performing effective judgment on the audio and video input information of the other terminal devices includes: when the voice of the current user exists in the audio input information of the other terminal devices, converting the voice of the current user into text information, and performing effective judgment on the text information.

[0017] According to one embodiment of the present invention, detecting whether the current user is a speaker includes: collecting video information and / or audio information of the current user, and analyzing whether the current user is a speaker based on the video information and / or audio information of the current user.

[0018] According to one embodiment of the present invention, sending the first mixed stream data to the cloud server includes: encapsulating the first mixed stream data in RTP (Real-time Transport Protocol) format and then sending the data to the cloud server in UDP (User Datagram Protocol) format.

[0019] According to one embodiment of the present invention, the above-mentioned real-time communication method further includes: receiving a second mixed stream data packet sent by the cloud server using a UDP receiving method, and performing RTP depacketization and then decoding and packet loss compensation processing on the second mixed stream data packet to obtain third valid code stream information, and playing audio and video according to the third valid code stream information.

[0020] According to an embodiment of the present invention, after obtaining the third valid code stream information, the method further includes: receiving a data request sent by the other terminal device, and sending the third valid code stream information to the other terminal device for audio and video playback according to the data request.

[0021] To achieve the above-mentioned object, a second embodiment of the present invention provides a computer-readable storage medium on which a real-time communication program is stored. When the real-time communication program is executed by a processor, the above-mentioned real-time communication method is implemented.

[0022] The computer-readable storage medium of the embodiment of the present invention can reduce the amount of data transmitted through the network, alleviate the network load, and improve the network quality by executing the above-mentioned real-time communication method.

[0023] To achieve the above-mentioned purpose, a terminal device is proposed in an embodiment of the third aspect of the present invention, comprising a memory, a processor, and a real-time communication program stored in the memory and runnable on the processor. When the processor executes the real-time communication program, the above-mentioned real-time communication method is implemented.

[0024] The terminal device according to the embodiment of the present invention can reduce the amount of data transmitted through the network, alleviate the network load, and improve the network quality by executing the above-mentioned real-time communication method.

[0025] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flow chart of a real-time communication method according to an embodiment of the present invention;

[0027] Figure 2 A schematic diagram of a workflow of a real-time communication method according to an embodiment of the present invention;

[0028] Figure 3 FIG. 4 is a block diagram of a terminal device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0030] The following describes a real-time communication method, a computer-readable storage medium, and a terminal device proposed in embodiments of the present invention with reference to the accompanying drawings.

[0031] Figure 1 FIG. 4 is a flow chart of a real-time communication method according to an embodiment of the present invention.

[0032] In one embodiment of the present invention, the following real-time communication method is executed by a management node of an ad hoc network, or by an upper-layer server, or for a terminal device.

[0033] like Figure 1 As shown, the real-time communication method according to the embodiment of the present invention may include the following steps:

[0034] S1, after the terminal devices accessing the network conference room are grouped, a master terminal device of each group is determined.

[0035] According to one embodiment of the present invention, the method of grouping terminal devices accessing the network conference room includes one or more of the following: grouping the terminal devices accessing the network conference room according to the access gateway information of each terminal device; grouping the terminal devices accessing the network conference room according to the location information of each terminal device; grouping the terminal devices accessing the network conference room according to the audio input information of each terminal device; grouping the terminal devices accessing the network conference room according to the video input information of each terminal device.

[0036] Specifically, when a large number of terminals access the same conference room, although each terminal is independent and unrelated to the network conference room, most terminals are related to each other, for example, they are on the same local area network or are physically close to each other.

[0037] For example, Company A and Company B are holding a conference call. Company A has 10 computers in the same conference room, and Company B has two computers in the same conference room. To the network cloud server, there are a total of 12 terminals registered and connected to the conference room. These 12 terminals have equal status and independently perform upload and download operations, requiring 12 x 11 data streams to be transmitted on the network. For Company A's 10 computers, their primary purpose is to receive information from Company B, and they expect the information they receive to be identical. For Company B's two computers, their primary purpose is to receive information from Company A, and they also expect the information they receive to be identical. This shows that although the network is transmitting so many streams in real time, much of the information in these streams is redundant. It's impossible for all computers in Company A to upload information at the same time, and it's also impossible for all computers in Company B to upload information at the same time. This means that many of the streams transmitted on the network are meaningless. If Company A's 10 streams were merged into one, and Company B's two streams were merged into one, only 2 x 1 data streams would need to be transmitted on the network, significantly reducing network bandwidth requirements.

[0038] Therefore, to address these characteristics of online conferencing, network bandwidth requirements can be reduced by grouping. Specifically, before normal communication begins, all terminal devices connected to the online conference room are automatically grouped. Changes in each terminal are detected at regular intervals and the grouping is modified. Furthermore, the grouping is also modified when a terminal exits or joins the online conference room. The principles of grouping are as follows:

[0039] First, grouping is done based on gateway information. The most common way for terminals to access the internet is through routers or mobile communication base stations (including but not limited to these two methods; with technological advancements, more advanced connection methods may emerge). However, for terminals connected to the same gateway, all uploaded information will converge at this gateway before being transmitted over the network. All downloaded information will also first reach this gateway before being distributed to each terminal. Therefore, the routing gateway closest to the communication terminal in the network can be used as a means of grouping.

[0040] Second, auxiliary technologies such as Bluetooth distance testing, GPS (Global Positioning System) positioning, and terminal signal strength are used. That is, after the initial division through the first method, it is necessary to ensure that the terminals within the group are in adjacent physical locations. This can avoid situations where the terminal devices within the group use the same router but are physically far apart, resulting in slow data transmission.

[0041] Third, the audio information input from each terminal device is analyzed to further group them. Numerous audio technologies can be used to detect whether terminals are adjacent to each other. For example, for closely spaced terminals, the microphone inputs have similar amplitudes and strong correlations. Alternatively, all terminals in a group can play high-frequency sound waves imperceptible to the human ear (either in a round-robin fashion or with varying frequencies) to determine whether other terminals detect them. Of course, further grouping of terminals within the same group can be achieved through other audio analysis techniques.

[0042] Fourth, the video information input by each terminal device is analyzed to further group them. For closely located terminals, the images captured by their cameras have strong similarities, and this similarity can be used to determine the physical distance between the terminals. Terminals within the same group can also be further grouped using other video analysis techniques.

[0043] By following these four steps, all terminals can be grouped. Terminals in the same group share the same gateway and are physically close to each other (for example, several terminals in the same conference room connected to the same Wi-Fi (wireless communication technology) will be grouped together). The number of terminals in each group must also be controlled within a user-defined range to prevent an excessive number of terminals from affecting data transmission speeds.

[0044] It should be noted that the above grouping process is not mandatory. If the required grouping results can be confirmed in advance, the subsequent grouping determination can be omitted. If some grouping techniques are unavailable, they can also be skipped. If the analysis results of several different steps conflict, a comprehensive correction can be made based on certain optimization criteria. As technology advances, more new technologies can be adopted to further improve the grouping results. Manual grouping is also supported. If the user clearly knows the grouping information, they can manually set which terminals belong to the same group.

[0045] In one embodiment of the present invention, all terminal devices in each group are located in the same local area network. That is, only after all terminal devices have been grouped and all terminal devices in each group are located in the same local area network can data transmission be achieved in this group, without affecting the data transmission speed. It should be noted that the same local area network can include different gateways.

[0046] Furthermore, according to one embodiment of the present invention, determining the main terminal device of each group includes: comprehensively determining the main terminal device of each group based on the signal strength between each terminal device in each group and the gateway, the performance and load conditions of each terminal device, and whether each terminal device is connected to a power source.

[0047] Specifically, after successful grouping, each group selects a terminal as the master terminal. Because the master terminal requires more computation, occupies more CPU (central processing unit), and consumes more power, the master terminal selection process considers information such as the signal strength between each terminal and the routing gateway, each terminal's performance, load, and whether it is connected to a power source. If the user understands the performance of each terminal, they can also manually select the master terminal.

[0048] Each terminal no longer needs to interact with all other terminals; it only needs to interact with the master terminal within its group. This interaction takes place over the local area network, making network bandwidth more readily available. The master terminal, on the other hand, only needs to interact with a limited number of terminals within its group, as well as a limited number of master terminals in other groups, significantly reducing network bandwidth requirements. Because the master terminal is selected based on comprehensive considerations, its computing performance, power supply, and network connectivity are relatively guaranteed.

[0049] It should be noted that the selection of groups and main terminals is not static. The original selection of groups and main terminals will be rechecked and adjusted at regular intervals or when a user enters or exits.

[0050] S2, obtaining audio and video input information of the main terminal device in each group, and receiving first valid code stream information sent by other terminal devices in the corresponding group through the main terminal device in each group.

[0051] According to one embodiment of the present invention, before the main terminal device in each group receives the first valid code stream information sent by other terminal devices in the corresponding group, it is also detected whether the current user of the other terminal devices is a speaker, and when the current user of the other terminal devices is the speaker, the audio and video input information of the other terminal devices is effectively judged.

[0052] Furthermore, according to an embodiment of the present invention, effective judgment is performed on the audio and video input information of other terminal devices, including: when the current user's voice exists in the audio input information of other terminal devices, the current user's voice is converted into text information, and the text information is effectively judged.

[0053] According to an embodiment of the present invention, detecting whether the current user is a speaker includes: collecting video information and / or audio information of the current user, and analyzing whether the current user is a speaker based on the video information and / or audio information of the current user.

[0054] Specifically, the mode set by the user is first determined. If the user manually sets it to speaker mode or manually plays pictures and sounds, information transmission is performed unconditionally (if the user sets it to speaker mode and does not speak for a long time, the user can be prompted to exit speaker mode). In other cases, the user's video and audio information is collected through the camera and microphone. After processing the input video information, the facial image of the current user is obtained to analyze whether the user is speaking. If the user closes his lips, the current user can be considered as an onlooker. When the video cannot determine whether the user is an onlooker, audio information is obtained, and audio voice detection is performed after signal pre-processing. If no voice is detected, the current user is considered to be an onlooker. If it is still impossible to determine whether the user is an onlooker, the voice is converted into text, and the converted text is semantically analyzed. The rationality of the analysis result is used to determine whether the user is speaking. If after all the above judgments, it is still impossible to confirm whether it is an onlooker, the current user is considered to be a speaker. Any of the above steps can be omitted or implemented in a different order depending on the actual implementation situation. If the current user is determined to be speaking, the entire audio, text, and video are transmitted to the master terminal. If the current user is determined to be an observer, the master terminal controls the transmission of audio, text, and video, or transmits only video, or at a reduced quality. The master terminal dynamically adjusts control instructions based on detected network conditions to ensure a balance between data volume and network status. If video transmission is required, video compression can be performed on the terminal or on the master terminal.

[0055] When the current user of other terminal devices is the speaker, it is necessary to first make a valid judgment on the audio and video input information of other terminal devices. For example, the current user's voice information is converted into text information to determine whether the user's voice information is related to the meeting content based on the text information. If it is relevant, the audio and video input information is considered valid and recorded as the first valid code stream information. At this time, the terminal device sends the first valid code stream information to the main terminal in the group.

[0056] S3, performing a validity determination on the audio and video input information to obtain second valid code stream information.

[0057] According to one embodiment of the present invention, before obtaining the audio and video input information of the main terminal device, it is also detected whether the current user of the main terminal device is a speaker, and when the current user of the main terminal device is a listener, the acquisition of the audio and video input information of the main terminal device is stopped; or only the video input information of the main terminal device is obtained; or only the degraded video input information of the main terminal device is obtained.

[0058] Specifically, in video conferencing, people tend to focus solely on the speaker's voice and video (including images shared by the speaker, such as PowerPoint presentations). Observers, on the other hand, are not interested in their voice and video. At any given moment, there are always only a few people speaking, while the vast majority are observers. Speakers require full audio, text, and video transmission to each endpoint, while observers do not need to. Alternatively, they can transmit only video, or at a reduced quality (resolution and frame rate, etc.) to each endpoint. This can be adjusted based on network conditions and user settings.

[0059] Not only must other terminal devices be judged as speakers, but the master terminal device must also be judged as the speaker. If the master terminal device is determined to be the speaker, it will transmit the entire audio, text, and video stream for storage to obtain the second valid bitstream information. If the current user of the master terminal device is a listener, and if the acquisition of the master terminal's audio and video input information is stopped, the second valid bitstream information will be empty. If only the master terminal's video input information is acquired, the second valid bitstream information will be the video input information. If only the master terminal's degraded video input information is acquired, the second valid bitstream information will be degraded video input information.

[0060] S4: Mix the first valid code stream information and the second valid code stream information to obtain first mixed stream data, and send the first mixed stream data to the cloud server.

[0061] According to an embodiment of the present invention, sending the first mixed stream data to the cloud server includes: encapsulating the first mixed stream data into RTP packets and then sending the packets to the cloud server using UDP.

[0062] Specifically, after the master terminal receives information from all the terminal devices in the group, it will merge the information and send it to the master terminals of other groups through the network. It will also receive information from the master terminals of other groups, process it, and then send it to the terminals in the group (when there are too many terminals in the group and the local area network bandwidth is overloaded, the master terminal can also degrade the received information and re-encode it before sending it). Figure 2 As shown, it is assumed that there are N intra-group terminals in total (the master terminal itself is also an intra-group terminal).

[0063] The master terminal receives audio, text, and video input streams from all N terminals within the group. The terminals corresponding to the observers have no audio or text input, and the video stream is either a complete video stream, a degraded video stream, or no video stream, depending on control information. The master terminal mixes the N input streams based on network conditions (in cases of a good network and a small number of terminals, mixing may not even be necessary). Audio is mixed into one or more streams, text is directly combined or mixed after semantic analysis, and video is discarded, transmitted at a reduced quality, or transmitted entirely. This mixing process and strategy are dynamically adjusted based on input and network conditions.

[0064] After the mixing, the first mixed stream data (including audio, text, and video) is obtained. After the audio and video are encoded, RTP packaging is performed and the data is sent to the cloud server using UDP.

[0065] Continue to refer to Figure 2 The above-mentioned real-time communication method may further include: receiving a second mixed stream data packet sent by the cloud server using a UDP receiving method, performing RTP decompression and subsequent decoding and packet loss compensation on the second mixed stream data packet to obtain third valid code stream information, and playing audio and video according to the third valid code stream information. The second mixed stream data packet may be the same as or different from the first mixed stream data packet, depending on the grouping situation. If the data packet is divided into two groups, the second mixed stream data packet is the same as the first mixed stream data packet. If the data packet is divided into multiple groups, the first mixed stream data packet may be the same as or different from the first mixed stream data packet.

[0066] Furthermore, after obtaining the third valid code stream information, the method further includes: receiving a data request sent by other terminal devices, and sending the third valid code stream information to the other terminal devices for audio and video playback according to the data request.

[0067] Specifically, the downlink receiving module uses UDP to receive the packet data (second mixed stream data packet) sent by the cloud server. After RTP depacketization and NETEQ decoding and post-processing, traditional real-time audio and video technologies are used to obtain audio, text, and video information. The main terminal sends the text and video to each terminal or discards them as appropriate. The terminals in each group request the audio, text, or video stream from the main terminal for display and playback as appropriate. At the same time, if there are too many members in the group and the effective bandwidth of the local area network is exceeded, the main terminal can also re-encode the original audio and video before distributing it to ensure call quality.

[0068] It should be noted that although this application only uses local terminal devices as the main scenario terminals for explanation, in actual situations, some terminals can also be expanded into terminal devices in a broad sense. For example, in the case of using hierarchical servers to interact with audio and video streams or multi-level LAN routing, some upper-level gateway services or grassroots servers close to the end-user terminals can also be regarded as a terminal, assuming the work of the main terminal or auxiliary terminal of this application. As long as it adopts a method similar to the technology described in this patent, it falls within the scope of protection of this patent.

[0069] In summary, the method of the present invention can significantly reduce the amount of data that needs to be transmitted over the network when there are many participants, alleviating network load and improving network quality. It also avoids excessive CPU load caused by a single terminal receiving and processing data from all other terminals, or all mixed and split traffic being concentrated on one or several devices, which can lead to severe lag or a rapid decrease in battery life.

[0070] In summary, according to the real-time communication method of an embodiment of the present invention, after grouping the terminal devices accessing the network conference room, the main terminal device of each group is determined, wherein all the terminal devices in each group meet the preset requirements; the audio and video input information of the main terminal device in each group is obtained, and the first valid code stream information sent by other terminal devices in the corresponding group is received through the main terminal device in each group; the audio and video input information is effectively judged to obtain the second valid code stream information; the first valid code stream information and the second valid code stream information are mixed to obtain the first mixed stream data, and the first mixed stream data is sent to the cloud server. Therefore, the method groups the terminal devices accessing the network conference room and determines the main terminal in the group, and transmits data directly between the main terminals, thereby reducing the amount of data transmitted through the network, alleviating the network load, and improving the network quality.

[0071] Corresponding to the above embodiment, the present invention further proposes a computer-readable storage medium on which a real-time communication program is stored. When the real-time communication program is executed by a processor, the above-mentioned real-time communication method is implemented.

[0072] The computer-readable storage medium of the embodiment of the present invention can reduce the amount of data transmitted through the network, alleviate the network load, and improve the network quality by executing the above-mentioned real-time communication method.

[0073] Corresponding to the above embodiment, the present invention further proposes a terminal device.

[0074] like Figure 3As shown, the terminal device 100 of an embodiment of the present invention includes: a memory 110, a processor 120, and a real-time communication program stored in the memory 110 and executable on the processor 120. When the processor 120 executes the real-time communication program, the above-mentioned real-time communication method is implemented.

[0075] The terminal device according to the embodiment of the present invention can reduce the amount of data transmitted through the network, alleviate the network load, and improve the network quality by executing the above-mentioned real-time communication method.

[0076] It should be noted that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0077] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0078] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0079] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0080] In the present invention, unless otherwise specified or limited, the terms "installed," "connected," "connect," "fixed," etc. should be understood in a broad sense. For example, they can refer to fixed connection, detachable connection, or integration; mechanical connection, electrical connection; direct connection, or indirect connection through an intermediate medium; internal communication between two components, or interaction between two components, unless otherwise specified. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0081] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A real-time communication method, characterized in that: include: After the terminal devices connected to the network conference room are grouped, the main terminal device of each group is determined; Obtaining audio and video input information of the main terminal device in each group, and receiving first valid code stream information sent by other terminal devices in the corresponding group through the main terminal device in each group, wherein the first valid code stream information is information related to the content of the conference, wherein, before obtaining the audio and video input information of the main terminal device, it also includes detecting whether the current user of the main terminal device is a speaker, and when the current user of the main terminal device is an observer, stopping obtaining the audio and video input information of the main terminal device; or only obtaining the video input information of the main terminal device; or only obtaining the degraded video input information of the main terminal device, wherein the process for the other terminal devices to obtain the first valid code stream information includes: converting the voice information in the currently obtained audio and video input information into text information, so that when it is determined that the user's voice information is related to the content of the conference according to the text information, the obtained audio and video input information is considered valid, and is determined to be the first valid code stream information; Performing a validity determination on the audio and video input information to obtain second valid code stream information, wherein, when the current user of the main terminal device is a listener, if obtaining the audio and video input information of the main terminal device is stopped, the second valid code stream information is empty information; if only the video input information of the main terminal device is obtained, the second valid code stream information is the video input information; if only the degraded video input information of the main terminal device is obtained, the second valid code stream information is the degraded video input information; The first valid code stream information and the second valid code stream information are mixed to obtain first mixed stream data, and the first mixed stream data is sent to a cloud server.

2. The method according to claim 1, characterized in that The method of grouping the terminal devices accessing the network conference room may also include one or more of the following: grouping the terminal devices connected to the network conference room according to audio input information of each terminal device; The terminal devices accessing the network conference room are grouped according to the video input information of each terminal device.

3. The method according to claim 1, characterized in that Determine the primary terminal device for each group, including: The master terminal device of each group is determined comprehensively based on the signal strength between each terminal device in each group and the gateway, the performance and load of each terminal device, and whether each terminal device is connected to a power source.

4. The method according to claim 1, wherein All terminal devices in each group are in the same local area network.

5. The method according to any one of claims 1 to 4, characterized in that Before the main terminal device in each group receives the first valid code stream information sent by other terminal devices in the corresponding group, it is also detected whether the current user of the other terminal device is a speaker, and when the current user of the other terminal device is a speaker, the audio and video input information of the other terminal device is effectively judged.

6. The method according to claim 5, characterized in that Effectively determining the audio and video input information of the other terminal devices includes: When the voice of the current user exists in the audio input information of the other terminal device, the voice of the current user is converted into text information, and the text information is effectively determined.

7. The method according to claim 5, characterized in that Detecting whether the current user is a speaker includes: The video information and / or audio information of the current user is collected, and whether the current user is a speaker is analyzed based on the video information and / or audio information of the current user.

8. The method according to claim 1, characterized in that Sending the first mixed stream data to a cloud server includes: After the first mixed stream data is packaged into RTP packets, it is sent to the cloud server using UDP.

9. The method according to claim 8, characterized in that Also includes: The second mixed stream data packet sent by the cloud server is received using UDP receiving mode, and the second mixed stream data packet is RTP depacketized and then decoded and packet loss compensation is performed to obtain third valid code stream information, and audio and video playback is performed according to the third valid code stream information.

10. The method according to claim 9, characterized in that After obtaining the third valid code stream information, the method further includes: Receive a data request sent by the other terminal device, and send the third valid code stream information to the other terminal device for audio and video playback according to the data request.

11. A computer-readable storage medium, characterized in that A real-time communication program is stored thereon, and when the real-time communication program is executed by a processor, the real-time communication method according to any one of claims 1 to 10 is implemented.

12. A terminal device, characterized in that: The method comprises a memory, a processor, and a real-time communication program stored in the memory and executable on the processor. When the processor executes the real-time communication program, the real-time communication method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Video and audio processing method, multi-point control unit and video conference system

    CN101370114A

  • Conference cascading method and system

    CN101867771A

  • Layout method and device for videos and audios in immersive conference

    CN104735390A

  • Processing method of video conference and computer readable storage medium

    CN109688365A