Real-time communication method, computer-readable storage medium, and electronic device

By filtering out invalid audio and video input and performing mixed stream processing at the master terminal, the problem of insufficient network bandwidth and computing power under multi-terminal access is solved, and efficient real-time audio and video communication is achieved.

CN115695704BActive Publication Date: 2025-09-30GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110846059.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-26
Publication Date
2025-09-30
Estimated Expiration
2041-07-26

AI Technical Summary

Technical Problem

When multiple terminals are connected to the existing real-time audio and video communication system, the network bandwidth demand increases dramatically, resulting in excessive network load and insufficient terminal computing power, affecting communication quality and battery life.

Method used

The main control terminal makes a valid judgment on the audio and video input, filters out invalid input, and transmits the valid input to the cloud server after mixed stream processing, reducing the amount of network transmission data and using the UDP protocol to transmit mixed stream data.

Benefits of technology

It reduces the amount of data transmitted on the network and the pressure of mixed flow, alleviates the network load, improves the network quality, and increases the computing power and battery life of the terminal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115695704B_ABST
    Figure CN115695704B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time communication method, a computer-readable storage medium, and an electronic device. The method comprises: obtaining audio and video input information and receiving first valid code stream information sent by a first electronic device; performing a valid determination on the audio and video input information to obtain second valid code stream information; performing a mixed stream processing on the first valid code stream information and the second valid code stream information to obtain first mixed stream data, and transmitting the first mixed stream data to a cloud server. The real-time communication method of the present invention can first filter out invalid audio and video inputs and then transmit all valid audio and video inputs to the cloud server after mixing them at a master control terminal. This can reduce the pressure of mixing, reduce the amount of data transmitted through the network, alleviate the network load, and improve network quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a real-time communication method, a computer-readable storage medium, and an electronic device. Background Art

[0002] With the advancement of technology, more and more people are using remote conferencing to communicate, which requires advanced real-time audio and video technology. A common implementation currently involves a main device in the conference room connected to a cloud server for communication. This can meet the needs of small meetings (those with only a few people), but it falls short of larger meetings (those with dozens of people in larger rooms). When some users need to communicate on their own terminals, they need to connect several auxiliary terminals to the same conference room. However, when too many auxiliary terminals are connected, the network load is significantly increased.

[0003] In real-time audio and video communication, if N terminals register to access a network conference room, there are N uplink signals. Each uplink signal must pass through the cloud server before being transmitted to (N-1) other terminals. Therefore, there are actually N*(N-1) signals transmitted on the network. Each terminal needs to process one uplink signal and N-1 downlink signals. Therefore, as the number of connected terminals increases, the required network bandwidth increases dramatically.

[0004] When the number of participants is small, the existing network bandwidth and terminal computing power are fully capable of meeting the requirements of real-time audio and video communication. However, as the number of connected terminals continues to increase, the required network bandwidth increases, eventually leading to severe communication lags or even disruptions. Each terminal needs to receive the data streams from all other terminals and perform decoding and various post-processing. This requires a lot of computing power, which, for handheld devices, can severely impact battery life. Summary of the Invention

[0005] The present invention aims to at least partially address one of the technical problems in the related art. To this end, the first object of the present invention is to provide a real-time communication method that can filter out invalid audio and video inputs and then mix all valid audio and video inputs at a master terminal before transmitting them to a cloud server. This method can reduce the mixing pressure, reduce the amount of data transmitted over the network, alleviate network load, and improve network quality.

[0006] The second object of the present invention is to provide a real-time communication method.

[0007] A third object of the present invention is to provide a computer-readable storage medium.

[0008] A fourth object of the present invention is to provide an electronic device.

[0009] To achieve the above-mentioned objectives, the first embodiment of the present invention proposes a real-time communication method, including: obtaining audio and video input information, and receiving first valid code stream information sent by a first electronic device; performing effective judgment on the audio and video input information to obtain second valid code stream information; mixing the first valid code stream information and the second valid code stream information to obtain first mixed stream data, and sending the first mixed stream data to a cloud server.

[0010] According to the real-time communication method of an embodiment of the present invention, first, audio and video input information is obtained, and first valid code stream information is received from a first electronic device. Then, the audio and video input information is validated and a second valid code stream information is obtained. Finally, the first valid code stream information and the second valid code stream information are mixed to obtain first mixed stream data, which is then sent to a cloud server. Thus, this method can first filter out invalid audio and video inputs and then transmit all valid audio and video inputs to the cloud server after mixing them at the master terminal. This can reduce the mixing pressure, reduce the amount of data transmitted across the network, alleviate network load, and improve network quality.

[0011] In addition, the real-time communication method according to the above embodiment of the present invention may also have the following additional technical features:

[0012] According to one embodiment of the present invention, performing a validity determination on the audio and video input information to obtain second valid code stream information includes: performing a voice presence determination on the audio and video input information; when valid voice is present in the audio and video input information, converting the valid voice into text information, and using the valid voice and the text information as the second valid code stream information.

[0013] According to one embodiment of the present invention, before receiving the first valid code stream information sent by the first electronic device, it also includes: sending an audio decision signal to the first electronic device, so that the first electronic device sends the first valid code stream information when it determines that the signal similarity between the audio decision signal and the audio input signal obtained by itself meets a preset condition.

[0014] According to one embodiment of the present invention, sending the first mixed stream data to the cloud server includes: encapsulating the first mixed stream data in RTP (Real-time Transport Protocol) format and then sending the data to the cloud server in UDP (User Datagram Protocol) format.

[0015] According to one embodiment of the present invention, the above-mentioned real-time communication method also includes: receiving a second mixed stream data packet sent by the cloud server using a UDP receiving method; performing RTP depacketization on the second mixed stream data packet and then decoding and packet loss compensation processing to obtain third valid code stream information; and playing audio and video according to the third valid code stream information.

[0016] According to an embodiment of the present invention, after obtaining the third valid code stream information, the method further includes: receiving a data request sent by the first electronic device, and sending the third valid code stream information to the first electronic device for audio and video playback according to the data request.

[0017] To achieve the above-mentioned purpose, the second aspect of the present invention proposes a real-time communication method, including: obtaining audio and video input information, and performing effective judgment on the audio and video input information to obtain first effective code stream information; sending the first effective code stream information to a second electronic device, so that the second electronic device mixes the first effective code stream information and the second effective code stream information and sends them to a cloud server, wherein the second effective code stream information is obtained by the second electronic device by performing effective judgment on the audio and video input information obtained by itself.

[0018] According to the real-time communication method of an embodiment of the present invention, audio and video input information is first obtained, and the audio and video input information is effectively judged to obtain first effective code stream information; then the first effective code stream information is sent to the second electronic device, so that the second electronic device mixes the first effective code stream information and the second effective code stream information and sends them to the cloud server, wherein the second effective code stream information is obtained by the second electronic device by effectively judging the audio and video input information obtained by itself. Therefore, the method can first filter out invalid audio and video inputs, and transmit all valid audio and video inputs to the cloud server after mixing them at the main control terminal, which can reduce the mixing pressure, reduce the amount of data transmitted through the network, reduce the network load, and improve the network quality.

[0019] In addition, the real-time communication method according to the above embodiment of the present invention may also have the following additional technical features:

[0020] According to one embodiment of the present invention, performing a validity determination on the audio and video input information to obtain first valid code stream information includes: performing a voice presence determination on the audio and video input information; when valid voice is present in the audio and video input information, converting the valid voice into text information, and using the valid voice and the text information as the first valid code stream information.

[0021] According to one embodiment of the present invention, before effectively judging the audio and video input information, the method further includes: receiving an audio decision signal sent by the second electronic device; and determining a signal similarity between the audio decision signal and the audio input signal in the audio and video input information.

[0022] According to an embodiment of the present invention, when the signal similarity meets a preset condition, the audio and video input information is effectively determined; when the signal similarity does not meet the preset condition, the audio and video input information is discarded.

[0023] According to one embodiment of the present invention, when the second electronic device receives the third valid code stream information from the cloud server, it also includes: sending a data request to the second electronic device so as to receive the third valid code stream information sent by the second electronic device according to the data request; and playing audio and video according to the third valid code stream information.

[0024] To achieve the above objectives, a third embodiment of the present invention proposes a computer-readable storage medium on which a real-time communication program of a terminal device is stored. When the real-time communication program of the terminal device is executed by a processor, the above-mentioned real-time communication method is implemented.

[0025] The computer-readable storage medium of the embodiment of the present invention can reduce the pressure of mixed flow, reduce the amount of data transmitted through the network, alleviate the network load, and improve the network quality by executing the above-mentioned real-time communication method.

[0026] To achieve the above-mentioned purpose, the fourth aspect of the present invention proposes an electronic device, comprising a memory, a processor, and a real-time communication program stored in the memory and runnable on the processor. When the processor executes the real-time communication program, the above-mentioned real-time communication method is implemented.

[0027] The electronic device of the present invention can reduce mixed flow pressure, reduce the amount of data transmitted through the network, alleviate network load, and improve network quality through the above-mentioned real-time communication method.

[0028] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flow chart of a real-time communication method according to a first embodiment of the present invention;

[0030] Figure 2 A schematic diagram of a workflow of a real-time communication method according to an embodiment of the present invention;

[0031] Figure 3FIG. 1 is a flow chart of a real-time communication method according to a second embodiment of the present invention. DETAILED DESCRIPTION

[0032] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0033] The real-time communication method, computer-readable storage medium, and electronic device proposed in the embodiments of the present invention are described below with reference to the accompanying drawings.

[0034] In the present invention, the terminal devices are divided into a main control terminal and multiple auxiliary terminals. For the main control terminal, since the main control terminal needs to perform more calculations, it occupies more CPU (central processing unit) and consumes more power. Therefore, when selecting the main control terminal, a suitable terminal can be selected as the main control terminal based on the signal strength between each terminal and the routing gateway, the performance of each terminal, the load condition, whether the power supply is connected, and other information. If the user understands the performance of each terminal, the main control terminal can also be manually selected by the user. As for the auxiliary terminals, they are a supplement to the communication of the main control terminal. After performing basic signal processing on the microphone input signal, each auxiliary terminal determines whether it is a valid sound input. If so, it outputs the processed sound and text signals. The main control terminal is used to mix the valid sound and text signals sent by all auxiliary terminals and transmit them to the cloud server. Since the input signal of the auxiliary terminal is processed first, when it is invalid information, it will no longer be sent to the main control terminal, so the pressure of mixing the main control terminal is relatively small. At the same time, the master terminal receives the voice and text signals sent from the remote end, decodes and displays them at the master terminal, and sends the text and audio signals to the corresponding auxiliary terminal based on the auxiliary terminal's request. The following describes the implementation process on the master terminal side and the auxiliary terminal side respectively.

[0035] Figure 1 FIG. 1 is a flow chart of a real-time communication method according to a first embodiment of the present invention.

[0036] like Figure 1 As shown, the real-time communication method according to the embodiment of the present invention may include the following steps:

[0037] S1, obtaining audio and video input information, and receiving first valid code stream information sent by a first electronic device.

[0038] S2: Perform validity determination on the audio and video input information to obtain second valid code stream information.

[0039] S3: Mix the first valid code stream information and the second valid code stream information to obtain first mixed stream data, and send the first mixed stream data to the cloud server.

[0040] According to one embodiment of the present invention, performing a validity determination on audio and video input information to obtain second valid code stream information includes: performing a voice presence determination on the audio and video input information; when valid voice exists in the audio and video input information, converting the valid voice into text information, and using the valid voice and text information as the second valid code stream information.

[0041] Specifically, the multiple auxiliary terminals first receive audio and video input information through microphones and judge whether the audio input information is valid. For example, they can determine whether the input audio information is valid (e.g., relevant to the meeting content) by judging the voice information. If the input audio information is determined to be valid (first valid code stream information), it is sent to the main control terminal. It is understood that if the input audio information is invalid, an empty packet (first valid code stream information) can also be sent to the main control terminal.

[0042] After the audio signal is input, the main control terminal first performs front-end signal processing, such as echo cancellation and noise reduction. It then makes a judgment on the processed signal. If it determines that there is no voice, the audio input signal is discarded. If it determines that there is voice, the ASR voice conversion module converts the audio into text, and outputs the audio and text as the second valid code stream information. It then receives signals from all auxiliary terminals with valid voice input, then mixes the first valid code stream information with the second valid code stream information to obtain mixed stream data (first mixed stream data), which is then sent to the cloud server. If the first valid code stream information is an empty packet, the obtained first mixed stream data is the second valid code stream information.

[0043] The above-mentioned first valid code stream information may also depend on the second valid code stream information. According to one embodiment of the present invention, before receiving the first valid code stream information sent by the first electronic device, it also includes: sending an audio decision signal to the first electronic device, so that the first electronic device sends the first valid code stream information when it determines that the signal similarity between the audio decision signal and the audio input signal obtained by itself meets the preset conditions.

[0044] Specifically, the main control terminal sends the original audio and video input information (the audio to be judged) to each auxiliary terminal for judging the similarity between the audio input signal obtained by the auxiliary terminal itself and the original audio and video input signal of the main control terminal. For example, the similarity judgment can be made by directly analyzing the actual collected voice or by analyzing the deliberately created high-frequency sound that is imperceptible to the human ear. Specifically, if the signal amplitudes of the two are similar and have a great similarity, it is considered that the physical distance between the auxiliary terminal and the main control terminal is adjacent, and the audio input signals between the two are repeated. At this time, all audio inputs are discarded, that is, the first valid code stream information is empty; if the signal amplitudes of the two are greatly different or the similarity is small, it is considered that the physical distance between the auxiliary terminal and the main control terminal is far apart. At this time, the auxiliary terminal processes the audio input signal to obtain the first valid code stream information and mixes it with the second valid code stream in the main control terminal.

[0045] It should be noted that in the above embodiment, if the signal similarity between the audio decision signal and the audio input signal obtained by the auxiliary terminal does not meet the preset conditions, the auxiliary terminal may stop sending the code stream information, or may continue to send the code stream information, but the code stream information will be an empty packet. In addition, if the auxiliary terminal is manually set to mute the microphone, indicating that there is no audio or text input, the signal similarity comparison in the above embodiment is not required.

[0046] According to an embodiment of the present invention, sending the first mixed stream data to the cloud server includes: encapsulating the first mixed stream data into RTP packets and then sending the packets to the cloud server using UDP.

[0047] According to one embodiment of the present invention, the above-mentioned real-time communication method further includes: receiving a second mixed stream data packet sent by the cloud server using a UDP receiving method; performing RTP depacketization and then decoding and packet loss compensation on the second mixed stream data packet to obtain third valid code stream information; and playing audio and video according to the third valid code stream information.

[0048] According to an embodiment of the present invention, after obtaining the third valid code stream information, the method further includes: receiving a data request sent by the first electronic device, and sending the third valid code stream information to the first electronic device for audio and video playback according to the data request.

[0049] Specifically, if Figure 2As shown, first exclude all invalid audio and text input streams, then all valid audio signals and text input streams (including audio signals and text input streams generated by the main control terminal) are mixed in the multi-channel mixing module, and audio encoding is performed in the audio encoding module. The text is annotated with the corresponding terminal number information and then merged. The first mixed stream data (including audio signals and text signals) obtained is RTP packaged and sent to the cloud server using UDP.

[0050] At the same time, the master terminal receives the voice and text signals (second mixed stream data packets) sent by the cloud server via UDP. After performing standard signal processing operations such as RTP depacketization, NETEQ decoding, and packet loss compensation, it obtains the audio and text information (third effective code stream information), plays the audio, and displays the text on demand. Simultaneously, based on the request of each auxiliary terminal (first electronic device), the master terminal sends the audio and text information (third effective code stream information) to the corresponding auxiliary terminal for playback or display.

[0051] It should be noted that the above-mentioned audio and text communication process can also be applied to video communication after being combined with video, and the video is transmitted accordingly according to the audio judgment result. In addition, although this application only uses local terminal devices as the main scenario terminals for explanation, in actual situations, some terminals can also be expanded into terminal devices in a broad sense. For example, in the case of using hierarchical servers to interact with audio and video streams or multi-level LAN routing, some upper-level gateway services or grassroots servers close to the end-user terminals can also be regarded as a terminal, assuming the work of the main control terminal or auxiliary terminal of this application. As long as they adopt methods similar to the technology described in this patent, they fall within the scope of protection of this patent.

[0052] In summary, according to the real-time communication method of an embodiment of the present invention, first, audio and video input information is obtained, and first valid code stream information sent by a first electronic device is received. Then, the audio and video input information is validated and a second valid code stream information is obtained. Finally, the first valid code stream information and the second valid code stream information are mixed to obtain first mixed stream data, and the first mixed stream data is sent to a cloud server. As a result, this method can first filter out invalid audio and video inputs and then transmit all valid audio and video inputs to the cloud server after mixing them at the master terminal. This can reduce the mixing pressure, reduce the amount of data transmitted through the network, alleviate the network load, and improve network quality.

[0053] The following embodiment is a real-time communication method on the auxiliary terminal side.

[0054] Figure 3 FIG. 1 is a flow chart of a real-time communication method according to a second embodiment of the present invention.

[0055] like Figure 3 As shown, the real-time communication method according to the embodiment of the present invention may include the following steps:

[0056] S11, obtaining audio and video input information, and performing a validity determination on the audio and video input information to obtain first valid code stream information.

[0057] S12, sending the first valid code stream information to the second electronic device, so that the second electronic device mixes the first valid code stream information and the second valid code stream information and sends them to the cloud server, wherein the second valid code stream information is obtained by the second electronic device through valid determination of the audio and video input information obtained by itself.

[0058] According to one embodiment of the present invention, performing a validity determination on audio and video input information to obtain first valid code stream information includes: performing a voice presence determination on the audio and video input information; when valid voice exists in the audio and video input information, converting the valid voice into text information, and using the valid voice and text information as the first valid code stream information.

[0059] According to one embodiment of the present invention, before effectively judging the audio and video input information, the method further includes: receiving an audio decision signal sent by a second electronic device; and determining a signal similarity between the audio decision signal and the audio input signal in the audio and video input information.

[0060] According to an embodiment of the present invention, when the signal similarity meets a preset condition, the audio and video input information is effectively judged; when the signal similarity does not meet the preset condition, the audio and video input information is discarded.

[0061] According to one embodiment of the present invention, when the second electronic device receives the third valid code stream information from the cloud server, it also includes: sending a data request to the second electronic device to receive the third valid code stream information sent by the second electronic device according to the data request; and playing audio and video according to the third valid code stream information.

[0062] Specifically, the main control terminal acts as the second electronic device. After receiving an audio signal, the auxiliary terminal first performs front-end signal processing, such as echo cancellation and noise reduction. It then makes a decision on the processed signal. If no speech is detected, the audio input signal is discarded. If speech is detected, the ASR voice conversion module converts the audio into text, and the audio and text are used as the first valid bitstream information. The main control terminal processes the incoming audio and video information in the same way as the auxiliary terminal processes the audio signal, generating the second valid bitstream information.

[0063] Before sending the first valid code stream information, the auxiliary terminal also judges the original audio and video input information (audio to be judged) sent by the main control terminal and the auxiliary terminal's own audio input signal to obtain the similarity between the two signals. For example, the similarity judgment can be made by directly analyzing the actual collected voice, or by analyzing the high-frequency sound that is deliberately created and cannot be perceived by the human ear. Specifically, if the amplitudes of the two signals are similar and have a large similarity, it is considered that the physical distance between the auxiliary terminal and the main control terminal is adjacent, and the audio input signals between the two are repeated. At this time, all audio inputs are discarded, that is, the first valid code stream information is empty, or the first valid code stream information is not sent to the second electronic device; if the amplitudes of the two signals are significantly different or the similarity is small, it is considered that the physical distance between the auxiliary terminal and the main control terminal is far apart. At this time, the auxiliary terminal processes the audio input signal to obtain the first valid code stream information and sends it to the main control terminal so that the main control terminal can mix the first valid code stream and the second valid code stream, and process it through RTP packets and send it to the cloud server using UDP.

[0064] At the same time, the main control terminal receives the voice and text signals (third effective code stream information) sent by the cloud server via UDP. After performing standard signal processing operations such as RTP depacketization, NETEQ decoding, and packet loss compensation, the main control terminal obtains the audio and text information (third effective code stream information), plays the audio, and displays the text on demand. If each auxiliary terminal sends a data request to the main control terminal (the second electronic device), the main control terminal sends the audio and text information (third effective code stream information) to the corresponding auxiliary terminal based on the request status. The auxiliary terminal then plays or displays the audio and text information based on the third effective code stream information.

[0065] It should be noted that for details not disclosed in the real-time communication method of the second embodiment of the present invention, please refer to the details disclosed in the real-time communication method of the first embodiment of the present invention, and no further details will be given here.

[0066] According to the real-time communication method of an embodiment of the present invention, audio and video input information is first obtained, and the audio and video input information is effectively judged to obtain first effective code stream information; then the first effective code stream information is sent to the second electronic device, so that the second electronic device mixes the first effective code stream information and the second effective code stream information and sends them to the cloud server, wherein the second effective code stream information is obtained by the second electronic device by effectively judging the audio and video input information obtained by itself. Therefore, the method can first filter out invalid audio and video inputs, and transmit all valid audio and video inputs to the cloud server after mixing them at the main control terminal, which can reduce the mixing pressure, reduce the amount of data transmitted through the network, reduce the network load, and improve the network quality.

[0067] Corresponding to the above embodiment, the present invention further proposes a computer-readable storage medium on which a real-time communication program of a terminal device is stored. When the real-time communication program of the terminal device is executed by a processor, the above-mentioned real-time communication method is implemented.

[0068] The computer-readable storage medium of the embodiment of the present invention can reduce the pressure of mixed flow, reduce the amount of data transmitted through the network, alleviate the network load, and improve the network quality by executing the above-mentioned real-time communication method.

[0069] Corresponding to the above embodiment, the present invention also proposes an electronic device, including a memory, a processor, and a real-time communication program stored in the memory and executable on the processor. When the processor executes the real-time communication program, the above-mentioned real-time communication method is implemented.

[0070] The electronic device of the present invention can reduce mixed flow pressure, reduce the amount of data transmitted through the network, alleviate network load, and improve network quality through the above-mentioned real-time communication method.

[0071] It should be noted that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0072] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0073] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0074] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0075] In the present invention, unless otherwise specified or limited, the terms "installed," "connected," "connect," "fixed," etc. should be understood in a broad sense. For example, they can refer to fixed connection, detachable connection, or integration; mechanical connection, electrical connection; direct connection, or indirect connection through an intermediate medium; internal communication between two components, or interaction between two components, unless otherwise specified. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0076] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A real-time communication method, characterized in that: Applied to a second electronic device, the method includes: Acquiring audio and video input information, and receiving first valid code stream information sent by a first electronic device, the first valid code stream information being valid information determined by the first electronic device from the received audio and video input information, wherein the first valid code stream information is sent when the first electronic device determines that signal similarity between an audio decision signal and an audio input signal in the audio and video input information acquired by the first electronic device satisfies a preset condition, and the audio decision signal is sent by the second electronic device to the first electronic device, wherein when the first electronic device determines that signal similarity between the audio decision signal and the audio input signal in the audio and video input information acquired by the first electronic device does not satisfy a preset condition, the first valid code stream information is not sent, and the audio decision signal represents the audio and video input information collected by the second electronic device, and when the preset condition is not satisfied, it indicates that the first electronic device and the second electronic device are adjacent; Performing voice presence determination on the audio and video input information; When there is valid voice in the audio and video input information, convert the valid voice into text information, and use the valid voice and the text information as second valid code stream information; The first valid code stream information and the second valid code stream information are mixed to obtain first mixed stream data, and the first mixed stream data is sent to a cloud server.

2. The method according to claim 1, characterized in that Sending the first mixed stream data to a cloud server includes: After the first mixed stream data is packaged into RTP packets, it is sent to the cloud server using UDP.

3. The method according to claim 2, characterized in that Also includes: Receive the second mixed flow data packet sent by the cloud server using UDP receiving mode; Performing RTP depacketization, decoding, and packet loss compensation on the second mixed stream data packet to obtain third valid code stream information; Audio and video playback is performed according to the third valid code stream information.

4. The method according to claim 3, characterized in that After obtaining the third valid code stream information, the method further includes: The device receives a data request sent by the first electronic device, and sends the third valid code stream information to the first electronic device for audio and video playback according to the data request.

5. A real-time communication method, characterized in that: Applied to a first electronic device, the method includes: receiving an audio decision signal sent by a second electronic device, and determining a signal similarity between the audio decision signal and an audio input signal in the audio and video input information acquired by the first electronic device, wherein the audio decision signal represents the audio and video input information acquired by the second electronic device; When it is determined that the signal similarity between the audio decision signal and the audio input signal obtained by the first electronic device meets a preset condition, performing a voice presence determination on the audio and video input information obtained by the first electronic device, wherein when the first electronic device determines that the signal similarity between the audio decision signal and the audio input signal in the audio and video input information obtained by the first electronic device does not meet the preset condition, the first valid bitstream information is not sent, and when the preset condition is not met, it indicates that the first electronic device and the second electronic device are adjacent; When there is valid voice in the audio and video input information, convert the valid voice into text information, and use the valid voice and the text information as first valid code stream information; The first valid bitstream information is sent to a second electronic device so that the second electronic device performs mixed stream processing on the first valid bitstream information and the second valid bitstream information and sends the mixed stream information to a cloud server, wherein the second valid bitstream information is obtained by the second electronic device through valid determination of the audio and video input information obtained by the second electronic device.

6. The method according to claim 5, characterized in that When the second electronic device receives the third valid code stream information from the cloud server, the method further includes: Sending a data request to the second electronic device, so as to receive third valid code stream information sent by the second electronic device according to the data request; Audio and video playback is performed according to the third valid code stream information.

7. A computer-readable storage medium, characterized in that A real-time communication program is stored thereon, and when the real-time communication program is executed by a processor, the real-time communication method according to any one of claims 1 to 4 or the real-time communication method according to any one of claims 5 to 6 is implemented.

8. An electronic device, characterized in that: The method comprises a memory, a processor, and a real-time communication program stored in the memory and executable on the processor. When the processor executes the real-time communication program, the method implements the real-time communication method according to any one of claims 1 to 4 or the real-time communication method according to any one of claims 5 to 6.

Citation Information

Patent Citations

  • Video and audio processing method, multi-point control unit and video conference system

    CN101370114A

  • Layout method and device for videos and audios in immersive conference

    CN104735390A

  • Terminal communication based echo suppression method and device

    CN106341563A

  • Processing method of video conference and computer readable storage medium

    CN109688365A