Multi-terminal call method, storage medium, and electronic device
By obtaining and mixing voice signals online, the problem of unclear voice collection in online meetings is solved, and low-cost and flexible voice playback and collection are achieved, improving the conference experience.
Patent Information
- Application Number
- CN202111621019.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-19
- Filing Date
- 2021-12-27
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-12-27
AI Technical Summary
In online meetings, the speaker is far away from the conference equipment, resulting in unclear voice collection, which affects the call effect. The existing hardware expansion terminal is costly, difficult to use and low flexibility.
By participating in an online call conference at the extended terminal set as the second terminal at the first terminal, the voice signal mixing process is acquired and performed, and the voice signal is sent to the terminal other than the first and second terminals to play.
It improves the call effect of online meetings, optimizes the participation experience, and solves the problems of high cost and low flexibility in hardware expansion terminals.
Smart Images

Figure CN114979545B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a multi-terminal communication method, a storage medium, and an electronic device. Background Art
[0002] In current video conferencing systems, a single conference phone is typically used for both playback and audio capture. If the conference room is large or the phone is improperly positioned, such that the speaker is far away from the phone, other online participants may not be able to clearly hear the live speech.
[0003] Traditional conference room expansion terminals typically include microphones and speakers. With the increasing popularity of consumer electronics devices and the increasingly sophisticated audio and video transmission capabilities of mobile internet, it's possible to leverage the audio playback and collection capabilities of hardware expansion terminals, such as wired microphones, to supplement the existing conference room equipment with voice playback and collection. However, hardware expansion terminals have several drawbacks. First, they are expensive, with a wide variety of equipment and consumables, resulting in high procurement and maintenance costs. Second, they are difficult to use, requiring specialized equipment to be configured by professionals. Furthermore, expansion terminals are often fixed in number and location, reducing flexibility in meeting communication and leading to lower meeting efficiency.
[0004] For example, in a multi-person online classroom, device A is the teacher's laptop and the conference initiator. Joining the conference through the computer on the podium, the audio data of teacher A's lecture is collected and sent to other online students. Audio data of students' questions can also be received in real time. Limited by the effective pickup range of the laptop for playback and collection, and the fact that participants may not be able to move around, students in the back rows of the conference room may not be able to communicate with other online students. For example, user E is a student in the back row of the classroom. When he wants to discuss a certain issue with other online students, his voice cannot be effectively collected by computer A because he is far away from the main device A in the conference room. As a result, his online classmates cannot hear his speech clearly.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0006] The embodiments of the present invention provide a multi-terminal call method, a storage medium, and an electronic device to at least solve the technical problem of poor call quality in online conferences existing in the related art.
[0007] According to one aspect of an embodiment of the present invention, a multi-terminal call method is provided, including: when a first terminal is set as an extended terminal of a second terminal to participate in a target call conference conducted online, obtaining a first voice signal collected by the first terminal and a second voice signal collected by the second terminal, wherein the terminals participating in the target call conference online include multiple terminals, and the multiple terminals include the first terminal and the second terminal; performing target mixing processing based on the first voice signal and the second voice signal to obtain a third voice signal; and sending the third voice signal to terminals other than the first terminal and the second terminal among the multiple terminals for playback.
[0008] According to another aspect of an embodiment of the present invention, a multi-terminal communication device is provided, including:
[0009] an acquisition module, configured to acquire, when a first terminal is configured as an extension terminal of a second terminal to participate in a target call conference being conducted online, a first voice signal collected by the first terminal and a second voice signal collected by the second terminal, wherein the terminals participating in the target call conference online include a plurality of terminals, and the plurality of terminals include the first terminal and the second terminal;
[0010] a mixing module, configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal;
[0011] A sending module is used to send the third voice signal to terminals other than the first terminal and the second terminal among the multiple terminals for playing.
[0012] Optionally,
[0013] The apparatus is configured to obtain a first voice signal collected by the first terminal and a second voice signal collected by the second terminal in the following manner: when the first terminal includes N terminals, obtaining N channels of the first voice signals collected by the N terminals and one channel of the second voice signal collected by the second terminal, wherein N is a natural number of 1 or greater, and each of the N terminals is configured to collect one channel of the first voice signal;
[0014] The device is used to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner: performing the target mixing processing on the N channels of the first voice signals and the one channel of the second voice signal to obtain one channel of the third voice signal.
[0015] Optionally, the apparatus is configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner:
[0016] When a fourth voice signal received and played on the second terminal is collected by the first terminal, performing echo cancellation processing on the first voice signal based on the fourth voice signal sent to the second terminal to obtain a fifth voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals;
[0017] The target mixing process is performed on the fifth voice signal and the second voice signal to obtain the third voice signal.
[0018] Optionally, the apparatus is configured to perform echo cancellation processing on the first voice signal according to the fourth voice signal sent to the second terminal in the following manner to obtain a fifth voice signal:
[0019] The fourth voice signal sent to the second terminal is eliminated from the first voice signal to obtain the fifth voice signal.
[0020] Optionally, the apparatus is configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner:
[0021] performing the target mixing process on the first voice signal and the second voice signal to obtain the sixth voice signal;
[0022] When the fourth voice signal received and played on the second terminal is collected by the first terminal, the sixth voice signal is echo-cancelled according to the fourth voice signal sent to the second terminal to obtain the third voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals.
[0023] Optionally, the apparatus is configured to perform the target mixing process on the first voice signal and the second voice signal to obtain the sixth voice signal in the following manner:
[0024] In a case where the first terminal includes N terminals, the N terminals collect N first voice signals, and the second terminal collects one second voice signal, the target mixing processing is performed on the N first voice signals and the one second voice signal to obtain one sixth voice signal, wherein N is a natural number of 1 or greater, and each of the N terminals is used to collect one first voice signal.
[0025] Optionally, the apparatus is configured to perform echo cancellation processing on the sixth voice signal according to the fourth voice signal sent to the second terminal in the following manner to obtain the third voice signal:
[0026] The fourth voice signal sent to the second terminal is eliminated from the sixth voice signal to obtain the third voice signal.
[0027] Optionally, the apparatus is configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner:
[0028] Determining a delay duration of the first voice signal relative to the second voice signal according to the first voice signal and the second voice signal;
[0029] adjusting the first voice signal to a seventh voice signal aligned with the second voice signal according to the delay duration;
[0030] The seventh voice signal and the second voice signal are mixed to obtain a third voice signal.
[0031] Optionally, the apparatus is configured to determine a delay duration of the first voice signal relative to the second voice signal based on the first voice signal and the second voice signal in the following manner:
[0032] Extracting a first set of audio fingerprints and a first set of timestamps corresponding to the first set of audio fingerprints from the first speech signal, and extracting a second set of audio fingerprints and a second set of timestamps corresponding to the second set of audio fingerprints from the second speech signal, wherein the first set of audio fingerprints includes one or more audio fingerprints, and the second set of audio fingerprints includes one or more audio fingerprints;
[0033] If a first audio fingerprint in the first set of audio fingerprints matches a second audio fingerprint in the second set of audio fingerprints, obtaining a first timestamp in the first set of timestamps corresponding to the first audio fingerprint and a second timestamp in the second set of timestamps corresponding to the second audio fingerprint;
[0034] The delay duration is determined as the time interval between the first timestamp and the second timestamp.
[0035] Optionally, the apparatus is configured to adjust the first voice signal to a seventh voice signal aligned with the second voice signal according to the delay duration in the following manner:
[0036] The first speech signal is shifted forward or backward in time by the delay duration so that the first group of audio fingerprints and the second group of audio fingerprints are aligned in terms of timestamps, thereby obtaining the seventh speech signal aligned with the second speech signal.
[0037] Optionally, the device is further used for:
[0038] Acquiring a target setting instruction on a target display interface of the first terminal;
[0039] In response to the target setting instruction, the first terminal is set as an extended terminal of the second terminal to participate in the target call conference being conducted online.
[0040] Optionally, the apparatus is configured to obtain a target setting instruction on a target display interface of the first terminal in the following manner:
[0041] Displaying the participant identifiers of the multiple terminals on the target display interface of the first terminal, wherein the participant identifiers of the multiple terminals include the participant identifier of the second terminal; obtaining the target setting instruction on the target display interface, wherein the target setting instruction is used to select the participant identifier of the second terminal and instruct the first terminal to be set as an extended terminal of the second terminal to participate in the target call conference being conducted online; or
[0042] In the case where the first terminal obtains the target information pushed by the second terminal, in response to the target information, a target prompt option is displayed on the target display interface of the first terminal, wherein the target prompt option is used to prompt whether to agree to set the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online; the target setting instruction is obtained on the target display interface, wherein the target setting instruction is used to select the agree option in the target prompt option and instruct to set the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online.
[0043] Optionally, the device is further used for:
[0044] In a case where the first terminal and the second terminal perform near field communication, the first terminal obtains the target information pushed by the second terminal, wherein the target information is used to trigger display of the target prompt option on the target display interface.
[0045] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the multi-terminal call method when running.
[0046] According to another aspect of an embodiment of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the multi-terminal call method through the computer program.
[0047] In an embodiment of the present invention, when a first terminal is set as an extended terminal of a second terminal to participate in an online target call conference, a first voice signal collected by the first terminal and a second voice signal collected by the second terminal are obtained, wherein the terminals participating in the online target call conference include multiple terminals, and the multiple terminals include a first terminal and a second terminal; according to the first voice signal and the second voice signal, target mixing processing is performed to obtain a third voice signal; and the third voice signal is sent to terminals other than the first terminal and the second terminal in the multiple terminals for playback. By setting the first terminal as an extended terminal of the second terminal to participate in the online target call conference, and performing mixing processing according to the voice signals collected by the first terminal and the second terminal, so as to send them to other terminals in the target call conference for playback, the purpose of using the first terminal to assist the second terminal in collecting voice signals is achieved, thereby achieving the technical effect of improving the call effect of the online meeting and optimizing the participation experience of the online meeting, thereby solving the technical problem of poor call effect of the online meeting existing in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0049] Figure 1 is a schematic diagram of an application environment of an optional multi-terminal conversation method according to an embodiment of the present invention;
[0050] Figure 2 This is a flow chart of an optional multi-terminal call method according to an embodiment of the present invention;
[0051] Figure 3 is a schematic diagram of an optional multi-terminal call method according to an embodiment of the present invention;
[0052] Figure 4 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention;
[0053] Figure 5 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention;
[0054] Figure 6is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention;
[0055] Figure 7 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention;
[0056] Figure 8 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention;
[0057] Figure 9 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention;
[0058] Figure 10 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention;
[0059] Figure 11 This is a schematic structural diagram of an optional multi-terminal communication device according to an embodiment of the present invention;
[0060] Figure 12 FIG. 4 is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0061] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0062] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0063] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:
[0064] Co-location multi-terminal scenario: In the same location (the same room, or physically close enough to be considered the same location), multiple additional terminals are added to the same session to improve the quality of voice playback and collection. Terminals specifically refer to devices that support voice playback and collection in the meeting, such as speakers and microphones.
[0065] Full-duplex call: also known as two-way simultaneous call, refers to a call interaction mode in which both parties involved in the conversation can speak and listen to voice at the same time.
[0066] Acquisition overlap: The same sound is captured repeatedly by different microphones, causing the listener to hear multiple repetitions. For example, if a speaker says "ABCD" and two microphones are recording simultaneously, and if the network delays of the two microphones are not aligned, the listener may hear "AABBCDCD."
[0067] The present invention will be described below in conjunction with embodiments:
[0068] According to one aspect of an embodiment of the present invention, a multi-terminal call method is provided. Optionally, in this embodiment, the multi-terminal call method can be applied to Figure 1 In the hardware environment composed of the server 101 and the user terminal 103 shown in FIG. Figure 1 As shown, the server 101 is connected to the terminal 103 via a network and can be used to provide services to the user terminal or the client installed on the user terminal. The client can be a video client, an instant messaging client, a browser client, an education client, a game client, a conference client, etc. A database 105 can be set on the server or independently of the server to provide data storage services for the server 101, for example, a conference data storage server. The above-mentioned network can include but is not limited to: a wired network and a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The user terminal 103 can be a terminal configured with an application capable of conducting a target call conference, and can include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, an MID (Mobile Internet Device), a PAD, a desktop computer, a smart TV and other computer devices. The above-mentioned server can be a single server, a server cluster consisting of multiple servers, or a cloud server.
[0069] Combine Figure 1 As shown, the multi-terminal call method can be implemented on the server 101 through the following steps:
[0070] S1, when a first terminal is configured as an extended terminal of a second terminal to participate in a target call conference being conducted online, obtaining, on server 101, a first voice signal collected by the first terminal and a second voice signal collected by the second terminal, wherein the terminals participating in the target call conference online include multiple terminals, and the multiple terminals include the first terminal and the second terminal;
[0071] S2, performing target mixing processing on the server 101 according to the first voice signal and the second voice signal to obtain a third voice signal;
[0072] S3: The server 101 sends the third voice signal to terminals other than the first terminal and the second terminal among the multiple terminals for playing.
[0073] Optionally, in this embodiment, the multi-terminal conversation method may also be used through a client including but not limited to a client configured on a server.
[0074] The above is only an example and is not specifically limited in this embodiment.
[0075] Alternatively, as an optional implementation, Figure 2 As shown, the multi-terminal call method includes:
[0076] S202, when a first terminal is configured as an extension terminal of a second terminal to participate in a target call conference being conducted online, obtaining a first voice signal collected by the first terminal and a second voice signal collected by the second terminal, wherein the terminals participating in the target call conference online include a plurality of terminals, and the plurality of terminals include the first terminal and the second terminal;
[0077] S204, performing target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal;
[0078] S206: Send the third voice signal to terminals other than the first terminal and the second terminal among the multiple terminals for playing.
[0079] Optionally, in this embodiment, the first terminal may include but is not limited to a terminal that can join the target call conference and has a voice signal collection function, such as a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, an MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV and other computer devices. The second terminal may include but is not limited to the same as the first terminal, but is a terminal that participates in the target call conference conducted online, and may include but is not limited to a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server may be connected directly or indirectly via wired or wireless communication, and this application does not impose any restrictions on this.
[0080] Optionally, in this embodiment, the subject of the above-mentioned method for executing multi-terminal calls may include but is not limited to a server. The above-mentioned server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0081] For example, Figure 3 FIG. 1 is a schematic diagram of an optional multi-terminal call method according to an embodiment of the present invention. Figure 3 As shown, the above multi-terminal call method can be applied to the following architecture:
[0082] Application 302, an application for participating in the online target conference call;
[0083] Intelligent cloud mixing (server) 304: used as an execution subject to execute the above-mentioned multi-terminal call method;
[0084] Extension device 306: the first device mentioned above;
[0085] Main device 308: is the second device mentioned above.
[0086] Specifically, it can be applied to cloud conference application scenarios including but not limited to:
[0087] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides these resources is called the "cloud." To users, these resources appear infinitely scalable and can be accessed at any time, used on demand, expanded at any time, and paid for on a per-use basis.
[0088] As a provider of cloud computing infrastructure, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) is established. Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0089] Based on logical functional divisions, the PaaS (Platform as a Service) layer can be deployed on top of the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. SaaS can also be deployed directly on top of IaaS. PaaS is a platform for software execution, such as databases and web containers. SaaS is a variety of business software, such as web portals and text messaging tools. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0090] Cloud conferencing is an efficient, convenient, and low-cost conferencing format based on cloud computing technology. Users can quickly and efficiently share voice, data, and video with teams and clients around the world through a simple, easy-to-use internet interface. The cloud conferencing service provider handles the complex technical aspects of data transmission and processing.
[0091] At present, domestic cloud conferencing mainly focuses on service content based on the SaaS (Software as a Service) model, including telephone, network, video and other service forms. Video conferencing based on cloud computing is called cloud conferencing.
[0092] In the era of cloud conferencing, data transmission, processing, and storage are all handled by the computer resources of video conferencing manufacturers. Users no longer need to purchase expensive hardware or install cumbersome software. They only need to open a browser and log in to the corresponding interface to conduct efficient remote meetings.
[0093] The above is only an example, and this embodiment does not impose any specific limitation.
[0094] Optionally, in this embodiment, the first terminal is configured as an extension terminal of the second terminal to participate in the online target call conference, which may include but is not limited to the following: Figure 4 Set up as shown, where Figure 4 FIG. 1 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention. Figure 4 As shown, a display interface 402 is displayed on the first terminal. By performing interactive operations on the interactive objects on the display interface 402, the first terminal is set as an extended terminal of the second terminal to participate in the online target call conference.
[0095] Optionally, in this embodiment, acquiring the first voice signal collected by the first terminal and the second voice signal collected by the second terminal may include but is not limited to acquiring the first voice signal and the second voice signal uplinked by the first terminal and the second terminal.
[0096] Optionally, in this embodiment, the terminals participating in the target call conference online include multiple terminals, which may include but are not limited to the first terminal, the second terminal, and other terminals participating in the target call conference online and of the same type as the second terminal.
[0097] Through this embodiment, when the first terminal is set as an extended terminal of the second terminal to participate in the target call conference being conducted online, the first voice signal collected by the first terminal and the second voice signal collected by the second terminal are obtained, wherein the terminals participating in the target call conference online include multiple terminals, and the multiple terminals include a first terminal and a second terminal; according to the first voice signal and the second voice signal, target mixing processing is performed to obtain a third voice signal; and the third voice signal is sent to terminals other than the first terminal and the second terminal in the multiple terminals for playback. By setting the first terminal as an extended terminal of the second terminal to participate in the target call conference being conducted online, and performing mixing processing according to the voice signals collected by the first terminal and the second terminal, so as to send them to other terminals in the target call conference for playback, the purpose of using the first terminal to assist the second terminal in collecting voice signals is achieved, thereby achieving the technical effect of improving the call effect of the online meeting and optimizing the participation experience of the online meeting, thereby solving the technical problem of poor call effect of the online meeting existing in the related technology.
[0098] As an optional solution,
[0099] The obtaining of the first voice signal collected by the first terminal and the second voice signal collected by the second terminal includes: when the first terminal includes N terminals, obtaining N channels of the first voice signals collected by the N terminals and one channel of the second voice signal collected by the second terminal, wherein N is a natural number of 1 or greater, and each of the N terminals is used to collect one channel of the first voice signal;
[0100] The performing target mixing processing according to the first voice signal and the second voice signal to obtain a third voice signal includes: performing the target mixing processing according to the N channels of the first voice signals and the one channel of the second voice signal to obtain one channel of the third voice signal.
[0101] Optionally, in this embodiment, when the first terminal includes N terminals, obtaining N first voice signals collected by the N terminals and one second voice signal collected by the second terminal may include but is not limited to configuring the number of the first terminals to one or more, wherein, when configured to N, each of the N terminals is used to collect one first voice signal, and the N terminals are all extended terminals of the second terminal.
[0102] Optionally, in this embodiment, target mixing processing is performed based on N first voice signals and one second voice signal to obtain a third voice signal, which may include but is not limited to performing an echo cancellation operation and a delay alignment operation on each of the N first voice signals and the above-mentioned second voice signal, and mixing the above-mentioned processed N voice signals with the above-mentioned second voice signal to obtain the above-mentioned third voice signal. In other words, the above-mentioned target mixing processing may include but is not limited to voice processing operations such as echo cancellation operation, delay alignment operation, and mixing operation.
[0103] Figure 5 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention, such as Figure 5 As shown, including but not limited to the following:
[0104] S1, the intelligent mixer obtains a second voice signal uploaded by the main microphone (corresponding to the aforementioned second terminal);
[0105] S2: The intelligent mixer obtains first voice signals uploaded by n expansion microphones (corresponding to the aforementioned N first terminals);
[0106] S3, the intelligent mixer performs target mixing processing according to the N first voice signals and the one second voice signal, performs mixed output, and obtains one third voice signal.
[0107] The above is only an example, and this embodiment does not impose any specific limitation.
[0108] Through this embodiment, when the first terminal includes N terminals, N first voice signals collected by the N terminals and one second voice signal collected by the second terminal are obtained, and target mixing processing is performed based on the N first voice signals and one second voice signal to obtain a third voice signal. By performing target mixing processing on the N first voice signals and one second voice signal, the purpose of using the first terminal to assist the second terminal in collecting voice signals is achieved, thereby achieving the technical effect of improving the call effect of online meetings and optimizing the participation experience of online meetings, and thus solving the technical problem of poor call effect of online meetings existing in related technologies.
[0109] As an optional solution, performing target mixing processing according to the first voice signal and the second voice signal to obtain a third voice signal includes:
[0110] When a fourth voice signal received and played on the second terminal is collected by the first terminal, performing echo cancellation processing on the first voice signal based on the fourth voice signal sent to the second terminal to obtain a fifth voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals;
[0111] Target mixing processing is performed on the fifth voice signal and the second voice signal to obtain a third voice signal.
[0112] Optionally, in this embodiment, the above-mentioned fourth voice signal may include but is not limited to a voice signal transmitted by other terminals among multiple terminals except the first terminal and the second terminal, and used for playing on the second terminal. The above-mentioned fourth voice signal sent to the second terminal may include but is not limited to a voice signal sent by a mixer to the above-mentioned second terminal for playback.
[0113] Optionally, in this embodiment, the above-mentioned echo cancellation processing of the first voice signal based on the fourth voice signal sent to the second terminal to obtain the fifth voice signal may include but is not limited to using an echo canceller, using the above-mentioned fourth voice signal as a near-end reference signal to cancel the echo to obtain the above-mentioned fifth voice signal.
[0114] Optionally, in this embodiment, target mixing processing is performed on the fifth voice signal and the second voice signal to obtain a third voice signal, which may include but is not limited to performing a delay alignment operation on the fifth voice signal and the second voice signal, and mixing the fifth voice signal after the above processing with the above second voice signal to obtain the above third voice signal. In other words, the above target mixing processing may include but is not limited to voice processing operations such as delay alignment operation and mixing operation.
[0115] For example, when the target mixing process includes a delay alignment operation, the process may include but is not limited to the following steps:
[0116] Determining a delay duration of the fifth voice signal relative to the second voice signal according to the fifth voice signal and the second voice signal;
[0117] adjusting the fifth voice signal to be aligned with the second voice signal according to the delay duration;
[0118] The fifth voice signal aligned with the second voice signal is mixed with the second voice signal to obtain a third voice signal.
[0119] Optionally, in this embodiment, Figure 6 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention, which may include but is not limited to the following: Figure 6 As shown, when the fourth voice signal received by the second terminal is played and collected by the first terminal ( Figure 6 The main device playback signal played on the second terminal collected by the extended microphone shown), according to the fourth voice signal ( Figure 6 The main device plays a signal) to the first voice signal ( Figure 6 The fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals, and the fifth voice signal and the second voice signal are subjected to target mixing processing to obtain a third voice signal, which is then sent to the terminals other than the first terminal and the second terminal among the multiple terminals for playback.
[0120] Through this embodiment, when the fourth voice signal received and played on the second terminal is collected by the first terminal, the first voice signal is echo-cancelled according to the fourth voice signal sent to the second terminal to obtain a fifth voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among multiple terminals; target mixing processing is performed on the fifth voice signal and the second voice signal to obtain a third voice signal. By performing echo cancellation processing on the first voice signal, the purpose of using the first terminal to assist the second terminal in collecting voice signals to improve the quality of the voice signals is achieved, thereby achieving the technical effect of improving the call effect of online meetings and optimizing the participation experience of online meetings, and thus solving the technical problem of poor call effect of online meetings existing in related technologies.
[0121] As an optional solution, performing echo cancellation processing on the first voice signal according to the fourth voice signal sent to the second terminal to obtain a fifth voice signal includes:
[0122] The fourth voice signal sent to the second terminal is eliminated from the first voice signal to obtain the fifth voice signal.
[0123] Optionally, in this embodiment, the above-mentioned elimination of the fourth voice signal sent to the second terminal in the first voice signal to obtain the fifth voice signal can be understood as removing the voice signal for playback sent to the second terminal included in all voice signals collected by the first terminal to obtain the above-mentioned fifth voice signal for subsequent mixing processing.
[0124] Through this embodiment, the fourth voice signal sent to the second terminal is eliminated from the first voice signal to obtain the fifth voice signal. By removing the voice signal for playback sent to the second terminal included in all voice signals collected by the first terminal, the above-mentioned multi-terminal call method can be applied to full-duplex application scenarios, achieving the purpose of using the first terminal to assist the second terminal in collecting voice signals, thereby achieving the technical effect of improving the call effect of online meetings and optimizing the participation experience of online meetings, and thus solving the technical problem of poor call effect of online meetings existing in related technologies.
[0125] As an optional solution, performing target mixing processing according to the first voice signal and the second voice signal to obtain a third voice signal includes:
[0126] performing the target mixing process on the first voice signal and the second voice signal to obtain the sixth voice signal;
[0127] When the fourth voice signal received and played on the second terminal is collected by the first terminal, the sixth voice signal is echo-cancelled according to the fourth voice signal sent to the second terminal to obtain the third voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals.
[0128] Optionally, in this embodiment, the target mixing processing is performed on the first voice signal and the second voice signal to obtain the sixth voice signal, which may include but is not limited to performing a delay alignment operation on the first voice signal and the second voice signal, and mixing the first voice signal after the above processing with the second voice signal to obtain the above sixth voice signal. In other words, the target mixing processing may include but is not limited to voice processing operations such as delay alignment operation and mixing operation.
[0129] Optionally, in this embodiment, when the fourth voice signal received and played on the second terminal is collected by the first terminal, the sixth voice signal is echo-cancelled according to the fourth voice signal sent to the second terminal to obtain the third voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals, and may include but is not limited to when the fourth voice signal received and played on the second terminal is collected by the first terminal, the sixth voice signal is echo-cancelled according to the fourth voice signal sent to the second terminal to obtain a third voice signal, and then the third voice signal is sent to the terminals other than the first terminal and the second terminal among the multiple terminals for playback.
[0130] As an optional solution, performing the target mixing process on the first voice signal and the second voice signal to obtain the sixth voice signal includes:
[0131] In a case where the first terminal includes N terminals, the N terminals collect N first voice signals, and the second terminal collects one second voice signal, the target mixing processing is performed on the N first voice signals and the one second voice signal to obtain one sixth voice signal, wherein N is a natural number of 1 or greater, and each of the N terminals is used to collect one first voice signal.
[0132] Optionally, in this embodiment, performing target mixing processing on the N first voice signals and the one second voice signal to obtain the one sixth voice signal may include, but is not limited to, performing delay alignment processing on the N first voice signals and the one second voice signal. Specifically, the processing may include, but is not limited to, the following steps:
[0133] Determine, based on the N channels of first voice signals and the one channel of second voice signals, a delay duration of the N channels of first voice signals relative to the one channel of second voice signals;
[0134] Adjusting the N-channel first voice signals to align with the one-channel second voice signal according to the delay duration;
[0135] The N first voice signals aligned with the one second voice signal are mixed with the one second voice signal to obtain a sixth voice signal.
[0136] Through this embodiment, in the case where the first terminal includes N terminals, the N terminals collect N first voice signals, and the second terminal collects one second voice signal, target mixing processing is performed on the N first voice signals and the second voice signal to obtain a sixth voice signal, wherein N is a natural number of 1 or greater than 1, and each of the N terminals is used to collect one first voice signal. By performing target mixing processing on the N first voice signals and the second voice signal, the purpose of using the first terminal to assist the second terminal in collecting voice signals is achieved, thereby achieving the technical effect of improving the call effect of online meetings and optimizing the participation experience of online meetings, and thus solving the technical problem of poor call effect of online meetings existing in related technologies.
[0137] As an optional solution, performing echo cancellation processing on the sixth voice signal according to the fourth voice signal sent to the second terminal to obtain the third voice signal includes:
[0138] The fourth voice signal sent to the second terminal is eliminated from the sixth voice signal to obtain the third voice signal.
[0139] Optionally, in this embodiment, the above-mentioned elimination of the fourth voice signal sent to the second terminal in the sixth voice signal to obtain the third voice signal can be understood as removing the voice signal for playback sent to the second terminal included in all voice signals collected by the first terminal to obtain the above-mentioned third voice signal for subsequent mixing processing.
[0140] Through this embodiment, the fourth voice signal sent to the second terminal is eliminated from the sixth voice signal to obtain the third voice signal. By removing the voice signal for playback sent to the second terminal included in all voice signals collected by the first terminal, the above-mentioned multi-terminal call method can be applied to full-duplex application scenarios, achieving the purpose of using the first terminal to assist the second terminal in collecting voice signals, thereby achieving the technical effect of improving the call effect of online meetings and optimizing the participation experience of online meetings, and thus solving the technical problem of poor call effect of online meetings existing in related technologies.
[0141] As an optional solution, performing target mixing processing according to the first voice signal and the second voice signal to obtain a third voice signal includes:
[0142] Determining a delay duration of the first voice signal relative to the second voice signal according to the first voice signal and the second voice signal;
[0143] adjusting the first voice signal to a seventh voice signal aligned with the second voice signal according to the delay duration;
[0144] The seventh voice signal and the second voice signal are mixed to obtain a third voice signal.
[0145] Optionally, in this embodiment, the delay duration of the first voice signal relative to the second voice signal may be estimated by a delay estimation method, but is not limited to, for example, Figure 6 As shown, the delay estimation module estimates the delay Tn of each extension microphone (corresponding to the aforementioned first device) relative to the main microphone (corresponding to the aforementioned second device).
[0146] Optionally, in this embodiment, according to the delay duration, adjusting the first voice signal to be aligned with the seventh voice signal of the second voice signal may include but is not limited to adjusting the first voice signal to be aligned with the seventh voice signal of the second voice signal by a delay alignment method, for example, Figure 6 As shown, the delay alignment module dynamically adjusts the delay of different extension microphone voices with reference to the estimated delay Tn, so that the extension microphone voice content of each channel is consistent with the main microphone voice signal at the same time, compensating for the delay misalignment between the main microphone and the extension microphone caused by inconsistent uplink network conditions. The above-mentioned delay alignment method may include but is not limited to audio fingerprint technology.
[0147] Optionally, in this embodiment, the mixing of the seventh voice signal and the second voice signal to obtain the third voice signal may include but is not limited to the following: Figure 7 As shown, Figure 7 FIG. 1 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention, which includes but is not limited to the following steps:
[0148] 1) Multi-channel microphone collection (including main microphone and extension microphone, Figure 7 The example is 4 channels), sent to the pre-mixer (PREMIX), and the total energy without gain adjustment is calculated;
[0149] 2) The Ratio Comparator module calculates the energy ratio of each microphone to the total energy given in step 1)
[0150] 3) Gain Control module, which adjusts the energy of each microphone based on the energy ratio calculated in step 2)
[0151] 4) The audio signals after energy adjustment in step 3) are mixed and superimposed to achieve the above-mentioned mixing of the seventh speech signal and the second speech signal to obtain a third speech signal.
[0152] The above is only an example, and this embodiment does not impose any specific limitation.
[0153] Through this embodiment, the delay duration of the first voice signal relative to the second voice signal is determined based on the first voice signal and the second voice signal, and according to the delay duration, the first voice signal is adjusted to a seventh voice signal aligned with the second voice signal, and the seventh voice signal and the second voice signal are mixed to obtain a third voice signal. By delaying and aligning the first voice signal and the second voice signal to avoid the above-mentioned collection accent phenomenon, the purpose of using the first terminal to assist the second terminal in collecting voice signals is achieved, thereby achieving the technical effect of improving the call quality of online meetings and optimizing the participation experience of online meetings, and thus solving the technical problem of poor call quality of online meetings existing in related technologies.
[0154] As an optional solution, determining, based on the first voice signal and the second voice signal, a delay duration of the first voice signal relative to the second voice signal includes:
[0155] Extracting a first set of audio fingerprints and a first set of timestamps corresponding to the first set of audio fingerprints from the first speech signal, and extracting a second set of audio fingerprints and a second set of timestamps corresponding to the second set of audio fingerprints from the second speech signal, wherein the first set of audio fingerprints includes one or more audio fingerprints, and the second set of audio fingerprints includes one or more audio fingerprints;
[0156] If a first audio fingerprint in the first set of audio fingerprints matches a second audio fingerprint in the second set of audio fingerprints, obtaining a first timestamp in the first set of timestamps corresponding to the first audio fingerprint and a second timestamp in the second set of timestamps corresponding to the second audio fingerprint;
[0157] The delay duration is determined as the time interval between the first timestamp and the second timestamp.
[0158] Optionally, in this embodiment, the determination of the delay duration of the first voice signal relative to the second voice signal based on the first voice signal and the second voice signal may include but is not limited to being based on audio fingerprint technology.
[0159] Specifically, Figure 8 FIG. 1 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention. Figure 8 As shown, after obtaining the main microphone and extended microphone data streams, by analyzing and comparing the audio fingerprints of the two streams (corresponding to the aforementioned first set of audio fingerprints and the second set of audio fingerprints), it is determined whether the sound collected by the main terminal device and the sound collected by the extended terminal are similar (corresponding to the first audio fingerprint in the aforementioned first set of audio fingerprints and the second audio fingerprint in the second set of audio fingerprints match). If similar, the relative delay value of the two sounds is determined.
[0160] Audio fingerprinting involves analyzing and extracting key features from sound to create a low-dimensional, unique digital signature. This method is commonly used for fast song searches and music copyright protection. It is primarily used to compare the fingerprints of human voices during calls. Considering real-time communication scenarios, the duration of the sound clips to be fingerprinted is shortened, multiple sub-fingerprints are extracted from a single frame of data, and the number of features is reduced.
[0161] When two fingerprint pools have sub-fingerprints with matching degrees that meet the requirements, the timestamps corresponding to the two matching sub-fingerprints are recorded (corresponding to the first timestamp corresponding to the first audio fingerprint in the first set of timestamps and the second timestamp corresponding to the second audio fingerprint in the second set of timestamps).
[0162] Finally, the delay judgment module combines the short-term fingerprint comparison similarity, similar fingerprint timestamps, and the statistical probability of long-term high-similarity fingerprints to give a delay estimate (corresponding to the aforementioned time interval).
[0163] The above is only an example, and this embodiment does not impose any specific limitation.
[0164] Through this embodiment, a first group of audio fingerprints and a first group of timestamps corresponding to the first group of audio fingerprints are extracted from a first voice signal, and a second group of audio fingerprints and a second group of timestamps corresponding to the second group of audio fingerprints are extracted from a second voice signal. The first group of audio fingerprints includes one or more audio fingerprints, and the second group of audio fingerprints includes one or more audio fingerprints. When a first audio fingerprint in the first group of audio fingerprints matches a second audio fingerprint in the second group of audio fingerprints, a first timestamp corresponding to the first audio fingerprint in the first group of timestamps and a second timestamp corresponding to the second audio fingerprint in the second group of timestamps are obtained. The delay duration is determined as the time interval between the first timestamp and the second timestamp. By aligning the first voice signal and the second voice signal for time delay, the aforementioned acquisition overlap is avoided. This achieves the purpose of using the first terminal to assist the second terminal in acquiring voice signals, thereby achieving the technical effect of improving the call quality of online meetings and optimizing the participant experience of online meetings, thereby solving the technical problem of poor call quality of online meetings existing in the related art.
[0165] As an optional solution, performing target mixing processing according to the first voice signal and the second voice signal to obtain a third voice signal includes:
[0166] The first voice signal and the second voice signal are mixed to obtain the third voice signal.
[0167] Optionally, in this embodiment, the mixing of the first voice signal and the second voice signal to obtain a third voice signal may include but is not limited to the following: Figure 7 As shown, Figure 7 FIG. 1 is a schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention, which includes but is not limited to the following steps:
[0168] 1) Multi-channel microphone collection (including main microphone and extension microphone, Figure 7 The example is 4 channels), sent to the pre-mixer (PREMIX), and the total energy without gain adjustment is calculated;
[0169] 2) The Ratio Comparator module calculates the energy ratio of each microphone to the total energy given in step 1)
[0170] 3) Gain Control module, which adjusts the energy of each microphone based on the energy ratio calculated in step 2)
[0171] 4) The audio signals after energy adjustment in step 3) are mixed and superimposed to achieve the mixing of the first speech signal and the second speech signal to obtain a third speech signal.
[0172] As an optional solution, adjusting the first voice signal to a seventh voice signal aligned with the second voice signal according to the delay duration includes:
[0173] The first speech signal is shifted forward or backward in time by the delay duration so that the first group of audio fingerprints and the second group of audio fingerprints are aligned in terms of timestamps, thereby obtaining the seventh speech signal aligned with the second speech signal.
[0174] Optionally, in this embodiment, the first voice signal may be moved forward or backward in time by a delay length, including but not limited to, by a delay alignment module, so that the first set of audio fingerprints and the second set of audio fingerprints are aligned in terms of timestamps, thereby obtaining a seventh voice signal aligned with the second voice signal. Specifically, based on the delay estimated by the delay estimation module, the relative time delay of the audio of the main device and the extended device is adjusted to align the audio text content and avoid the occurrence of audio overlap due to the close proximity of multiple microphones. The main scheme is as follows:
[0175] Based on the delay of the main device, the audio collected by the extended device is aligned with the main device. An audio data buffer is established between the main device and the extended device to buffer audio data of a certain length, generally 20-300ms, which can be adjusted according to the usage scenario. If the audio data of the extended device arrives at the mixing server before the main device, the earlier arriving data is buffered. The buffer length is the estimated relative delay between the extended device and the main device. If the audio data of the extended device arrives at the mixing server later than the main device, the earlier arriving data in the extended device buffer is cleared. The data clearing length is the delay between the main device and the extended device.
[0176] The above solution uses the main device as the benchmark and only adjusts the data flow of the expansion device, which can ensure the smoothness of the audio flow as a whole, without noise caused by packet loss, etc. In addition, any expansion device can be used as the benchmark, and the implementation plan is similar to the above.
[0177] The above is only an example, and this embodiment does not impose any specific limitation.
[0178] Through this embodiment, the first voice signal is moved forward or backward in time by a delay length, so that the first group of audio fingerprints and the second group of audio fingerprints are aligned in timestamp, and a seventh voice signal aligned with the second voice signal is obtained. By delaying the first voice signal and the second voice signal to align them, the above-mentioned collection accent phenomenon is avoided, and the purpose of using the first terminal to assist the second terminal in collecting voice signals is achieved, thereby achieving the technical effect of improving the call quality of online meetings and optimizing the participation experience of online meetings, and thus solving the technical problem of poor call quality of online meetings in related technologies.
[0179] As an optional solution, the method further includes:
[0180] Acquiring a target setting instruction on a target display interface of the first terminal;
[0181] In response to the target setting instruction, the first terminal is set as an extended terminal of the second terminal to participate in the target call conference being conducted online.
[0182] Optionally, in this embodiment, the target display interface of the first terminal may include but is not limited to: Figure 4 As shown, obtaining the target setting instruction on the above-mentioned target display interface may include but is not limited to performing a target interaction operation for the interactive object "yes", as the above-mentioned target setting instruction, in response to the target setting instruction, setting the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online.
[0183] It should be noted that Figure 9 Schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention, such as Figure 9 As shown, taking the first terminal as the extension device and the second terminal as the main device as an example, combined with Figure 4 When user E is far away from the second device, a target setting instruction is obtained on the target display interface of the first terminal. In response to the target setting instruction, the first terminal is set as an extended terminal of the second terminal to participate in the online target call conference, where user B, user C, and user D are users whose voice signals can be clearly collected by the above-mentioned second device.
[0184] Through this embodiment, a target setting instruction is obtained on the target display interface of the first terminal, and in response to the target setting instruction, the first terminal is set as an extended terminal of the second terminal to participate in the target call conference conducted online. The first terminal is used as an extended device of the above-mentioned second terminal at a position far away from the second device and the voice signal cannot be effectively collected by the second device. The target setting instruction is displayed on the target display interface of the second terminal to set the first terminal as an extended terminal of the second terminal, thereby achieving the purpose of using the first terminal to assist the second terminal in collecting voice signals, thereby achieving the technical effect of improving the call effect of the online meeting and optimizing the participation experience of the online meeting, thereby solving the technical problem of poor call effect of the online meeting existing in the related technology.
[0185] As an optional solution, obtaining a target setting instruction on a target display interface of the first terminal includes:
[0186] Displaying the participant identifiers of the multiple terminals on the target display interface of the first terminal, wherein the participant identifiers of the multiple terminals include the participant identifier of the second terminal; obtaining the target setting instruction on the target display interface, wherein the target setting instruction is used to select the participant identifier of the second terminal and instruct the first terminal to be set as an extended terminal of the second terminal to participate in the target call conference being conducted online; or
[0187] In the case where the first terminal obtains the target information pushed by the second terminal, in response to the target information, a target prompt option is displayed on the target display interface of the first terminal, wherein the target prompt option is used to prompt whether to agree to set the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online; the target setting instruction is obtained on the target display interface, wherein the target setting instruction is used to select the agree option in the target prompt option and instruct to set the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online.
[0188] Optionally, in this embodiment, the conference participant identifier may include but is not limited to an identifier of an account used to log in to the second terminal displayed on the first terminal, where the account is an account used to log in to the application where the target conference call is located.
[0189] For example, Figure 4 As shown, the avatar of user A is the above-mentioned participant identification. By selecting the avatar of user A, the target prompt option "Join the meeting as an extended microphone of user A?" is displayed on the target display interface, and then the target setting instruction is obtained, indicating that the first terminal is set as an extended terminal of the second terminal to participate in the online target call conference.
[0190] It should be noted that the above-mentioned steps of instructing the first terminal to be set as an extended terminal of the second terminal to participate in the target call conference being conducted online may include but are not limited to one or two steps. For example, the first step is to double-click the user head; the second step is to select "Yes" to indicate that the first terminal is set as an extended terminal of the second terminal to participate in the target call conference being conducted online.
[0191] As an optional solution, the method further includes:
[0192] In a case where the first terminal and the second terminal perform near field communication, the first terminal obtains the target information pushed by the second terminal, wherein the target information is used to trigger display of the target prompt option on the target display interface.
[0193] Optionally, in this embodiment, the near-field communication between the first terminal and the second terminal may include but is not limited to methods such as NFC proximity or ultrasonic watermarking. The second device broadcasts an identification code, and the first device automatically responds to the identification code, prompting the user whether to join as an extended device.
[0194] For example, Figure 10 Schematic diagram of another optional multi-terminal call method according to an embodiment of the present invention, such as Figure 10 As shown, when the near field communication method is configured as NFC identification, the main device broadcasts the identification code as the target information as the second device. When user E holds the expansion device as the first device close to the main device, the identification code pushed by the second terminal is obtained as the target information, and the target prompt options are displayed on the expansion device. The target prompt options may include but are not limited to the following: Figure 4 shown.
[0195] The above is only an example, and this embodiment does not impose any specific limitation.
[0196] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0197] According to another aspect of the embodiment of the present invention, a multi-terminal communication device for implementing the multi-terminal communication method is also provided. Figure 11 As shown, the device includes:
[0198] An acquisition module 1102 is configured to acquire, when a first terminal is configured as an extension terminal of a second terminal to participate in a target call conference being conducted online, a first voice signal collected by the first terminal and a second voice signal collected by the second terminal, wherein the terminals participating in the target call conference online include a plurality of terminals, and the plurality of terminals include the first terminal and the second terminal;
[0199] A mixing module 1104 is configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal;
[0200] The sending module 1106 is configured to send the third voice signal to terminals other than the first terminal and the second terminal among the multiple terminals for playing.
[0201] As an optional solution,
[0202] The apparatus is configured to obtain a first voice signal collected by the first terminal and a second voice signal collected by the second terminal in the following manner: when the first terminal includes N terminals, obtaining N channels of the first voice signals collected by the N terminals and one channel of the second voice signal collected by the second terminal, wherein N is a natural number of 1 or greater, and each of the N terminals is configured to collect one channel of the first voice signal;
[0203] The device is used to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner: performing the target mixing processing on the N channels of the first voice signals and the one channel of the second voice signal to obtain one channel of the third voice signal.
[0204] As an optional solution, the apparatus is configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner:
[0205] When a fourth voice signal received and played on the second terminal is collected by the first terminal, performing echo cancellation processing on the first voice signal based on the fourth voice signal sent to the second terminal to obtain a fifth voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals;
[0206] The target audio mixing process is performed on the fifth speech signal and the second speech signal to obtain the third speech signal.
[0207] As an optional solution, the apparatus is configured to perform echo cancellation processing on the first voice signal according to the fourth voice signal sent to the second terminal in the following manner to obtain a fifth voice signal:
[0208] The fourth voice signal sent to the second terminal is eliminated from the first voice signal to obtain the fifth voice signal.
[0209] As an optional solution, the apparatus is configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner:
[0210] performing the target mixing process on the first voice signal and the second voice signal to obtain the sixth voice signal;
[0211] When the fourth voice signal received and played on the second terminal is collected by the first terminal, the sixth voice signal is echo-cancelled according to the fourth voice signal sent to the second terminal to obtain the third voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals.
[0212] As an optional solution, the apparatus is configured to perform the target mixing process on the first voice signal and the second voice signal in the following manner to obtain the sixth voice signal:
[0213] In a case where the first terminal includes N terminals, the N terminals collect N first voice signals, and the second terminal collects one second voice signal, the target mixing processing is performed on the N first voice signals and the one second voice signal to obtain one sixth voice signal, wherein N is a natural number of 1 or greater, and each of the N terminals is used to collect one first voice signal.
[0214] As an optional solution, the apparatus is configured to perform echo cancellation processing on the sixth voice signal according to the fourth voice signal sent to the second terminal in the following manner to obtain the third voice signal:
[0215] The fourth voice signal sent to the second terminal is eliminated from the sixth voice signal to obtain the third voice signal.
[0216] As an optional solution, the apparatus is configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner:
[0217] Determining a delay duration of the first voice signal relative to the second voice signal according to the first voice signal and the second voice signal;
[0218] adjusting the first voice signal to a seventh voice signal aligned with the second voice signal according to the delay duration;
[0219] The seventh voice signal and the second voice signal are mixed to obtain a third voice signal.
[0220] As an optional solution, the apparatus is configured to determine a delay duration of the first voice signal relative to the second voice signal based on the first voice signal and the second voice signal in the following manner:
[0221] Extracting a first set of audio fingerprints and a first set of timestamps corresponding to the first set of audio fingerprints from the first speech signal, and extracting a second set of audio fingerprints and a second set of timestamps corresponding to the second set of audio fingerprints from the second speech signal, wherein the first set of audio fingerprints includes one or more audio fingerprints, and the second set of audio fingerprints includes one or more audio fingerprints;
[0222] If a first audio fingerprint in the first set of audio fingerprints matches a second audio fingerprint in the second set of audio fingerprints, obtaining a first timestamp in the first set of timestamps corresponding to the first audio fingerprint and a second timestamp in the second set of timestamps corresponding to the second audio fingerprint;
[0223] The delay duration is determined as the time interval between the first timestamp and the second timestamp.
[0224] As an optional solution, the apparatus is configured to adjust the first voice signal to a seventh voice signal aligned with the second voice signal according to the delay duration in the following manner:
[0225] The first speech signal is shifted forward or backward in time by the delay duration so that the first group of audio fingerprints and the second group of audio fingerprints are aligned in terms of timestamps, thereby obtaining the seventh speech signal aligned with the second speech signal.
[0226] As an optional solution, the device is further used for:
[0227] Acquiring a target setting instruction on a target display interface of the first terminal;
[0228] In response to the target setting instruction, the first terminal is set as an extended terminal of the second terminal to participate in the target call conference being conducted online.
[0229] As an optional solution, the apparatus is configured to obtain a target setting instruction on a target display interface of the first terminal in the following manner:
[0230] Displaying the participant identifiers of the multiple terminals on the target display interface of the first terminal, wherein the participant identifiers of the multiple terminals include the participant identifier of the second terminal; obtaining the target setting instruction on the target display interface, wherein the target setting instruction is used to select the participant identifier of the second terminal and instruct the first terminal to be set as an extended terminal of the second terminal to participate in the target call conference being conducted online; or
[0231] In the case where the first terminal obtains the target information pushed by the second terminal, in response to the target information, a target prompt option is displayed on the target display interface of the first terminal, wherein the target prompt option is used to prompt whether to agree to set the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online; the target setting instruction is obtained on the target display interface, wherein the target setting instruction is used to select the agree option in the target prompt option and instruct to set the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online.
[0232] As an optional solution, the device is also used for:
[0233] In a case where the first terminal and the second terminal perform near field communication, the first terminal obtains the target information pushed by the second terminal, wherein the target information is used to trigger display of the target prompt option on the target display interface.
[0234] According to another aspect of the embodiment of the present invention, there is also provided an electronic device for implementing the above-mentioned multi-terminal call method, which may be Figure 1 The terminal device or server shown in FIG. This embodiment is described by taking the electronic device as a server as an example. Figure 12 As shown, the electronic device includes a memory 1202 and a processor 1204. The memory 1202 stores a computer program, and the processor 1204 is configured to execute the steps in any of the above method embodiments through the computer program.
[0235] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0236] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0237] S1, when a first terminal is configured as an extension terminal of a second terminal to participate in a target call conference being conducted online, obtaining a first voice signal collected by the first terminal and a second voice signal collected by the second terminal, wherein the terminals participating in the target call conference online include multiple terminals, and the multiple terminals include the first terminal and the second terminal;
[0238] S2, performing target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal;
[0239] S3: Send the third voice signal to terminals other than the first terminal and the second terminal among the multiple terminals for playing.
[0240] Alternatively, those skilled in the art will appreciate that Figure 12 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 12 It does not limit the structure of the electronic device. For example, the electronic device may also include Figure 12 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 12 Different configurations shown.
[0241] Among them, the memory 1202 can be used to store software programs and modules, such as the program instructions / modules corresponding to the multi-terminal call method and device in the embodiment of the present invention. The processor 1204 executes various functional applications and data processing by running the software programs and modules stored in the memory 1202, that is, realizing the above-mentioned multi-terminal call method. The memory 1202 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1202 may further include a memory remotely located relative to the processor 1204, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1202 can be used for information such as voice signals, but is not limited to it. As an example, if Figure 12 As shown, the memory 1202 may include, but is not limited to, the acquisition module 1102, the mixing module 1104, and the sending module 1106 in the multi-terminal communication device. In addition, it may also include, but is not limited to, other module units in the multi-terminal communication device, which will not be repeated in this example.
[0242] Optionally, the transmission device 1206 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1206 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1206 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0243] In addition, the electronic device further includes: a display 1208 for displaying the target display interface; and a connection bus 1210 for connecting various module components in the electronic device.
[0244] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes through network communication. The nodes may form a peer-to-peer (P2P) network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.
[0245] According to one aspect of the present application, a computer program product or computer program is provided. The computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the aforementioned multi-terminal call. The computer program is configured to execute the steps of any of the aforementioned method embodiments when executed.
[0246] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0247] S1, when a first terminal is configured as an extension terminal of a second terminal to participate in a target call conference being conducted online, obtaining a first voice signal collected by the first terminal and a second voice signal collected by the second terminal, wherein the terminals participating in the target call conference online include multiple terminals, and the multiple terminals include the first terminal and the second terminal;
[0248] S2, performing target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal;
[0249] S3: Send the third voice signal to terminals other than the first terminal and the second terminal among the multiple terminals for playing.
[0250] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0251] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0252] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the methods described in various embodiments of the present invention.
[0253] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0254] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0255] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0256] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0257] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A multi-terminal conversation method, characterized in that: include: Acquiring a target setting instruction on a target display interface of the first terminal; In response to the target setting instruction, setting the first terminal as an extended terminal of the second terminal to participate in the online target call conference, so as to use the first terminal to assist the second terminal in collecting voice signals; Obtaining a first voice signal collected by the first terminal and a second voice signal collected by the second terminal, wherein the terminals participating in the target call conference online include multiple terminals, the multiple terminals include the first terminal and the second terminal, and the first terminal and the second terminal are in a full-duplex scenario; performing target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal; when a fourth voice signal received and played on the second terminal is collected by the first terminal, performing echo cancellation processing on the first voice signal based on the fourth voice signal sent to the second terminal to obtain a fifth voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals; performing the target mixing processing on the fifth voice signal and the second voice signal to obtain the third voice signal; and sending the third voice signal to the terminals other than the first terminal and the second terminal among the multiple terminals for playback; The method also includes: adjusting the data flow of the first terminal based on the second terminal, establishing a buffer for the audio data of the second terminal and the audio data of the first terminal, and when the audio data of the first terminal arrives at the mixing server before the audio data of the second terminal, buffering the audio data of the first terminal, and the buffering length is the delay length; when the audio data of the first terminal arrives at the mixing server later than the audio data of the second terminal, clearing the audio data that arrived earlier in the buffer, and the data clearing length is the delay length, and the delay length is used to adjust the first voice signal to be aligned with the second voice signal.
2. The method according to claim 1, characterized in that The obtaining of the first voice signal collected by the first terminal and the second voice signal collected by the second terminal includes: when the first terminal includes N terminals, obtaining N channels of the first voice signals collected by the N terminals and one channel of the second voice signal collected by the second terminal, wherein N is a natural number of 1 or greater, and each of the N terminals is used to collect one channel of the first voice signal; The performing target mixing processing according to the first voice signal and the second voice signal to obtain a third voice signal includes: performing the target mixing processing according to the N channels of the first voice signals and the one channel of the second voice signal to obtain one channel of the third voice signal.
3. The method according to claim 1, characterized in that The performing echo cancellation processing on the first voice signal according to the fourth voice signal sent to the second terminal to obtain a fifth voice signal includes: The fourth voice signal sent to the second terminal is eliminated from the first voice signal to obtain the fifth voice signal.
4. The method according to claim 1, wherein The performing target mixing processing according to the first voice signal and the second voice signal to obtain a third voice signal includes: performing the target mixing process on the first voice signal and the second voice signal to obtain a sixth voice signal; When the fourth voice signal received and played on the second terminal is collected by the first terminal, the sixth voice signal is echo-cancelled according to the fourth voice signal sent to the second terminal to obtain the third voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals.
5. The method according to claim 4, characterized in that The performing the target mixing process on the first voice signal and the second voice signal to obtain the sixth voice signal includes: In a case where the first terminal includes N terminals, the N terminals collect N first voice signals, and the second terminal collects one second voice signal, the target mixing processing is performed on the N first voice signals and the one second voice signal to obtain one sixth voice signal, wherein N is a natural number of 1 or greater, and each of the N terminals is used to collect one first voice signal.
6. The method according to claim 4, characterized in that The performing echo cancellation processing on the sixth voice signal according to the fourth voice signal sent to the second terminal to obtain the third voice signal includes: The fourth voice signal sent to the second terminal is eliminated from the sixth voice signal to obtain the third voice signal.
7. The method according to claim 1, characterized in that The performing target mixing processing according to the first voice signal and the second voice signal to obtain a third voice signal includes: Determining a delay duration of the first voice signal relative to the second voice signal according to the first voice signal and the second voice signal; adjusting the first voice signal to a seventh voice signal aligned with the second voice signal according to the delay duration; The seventh voice signal and the second voice signal are mixed to obtain a third voice signal.
8. The method according to claim 7, characterized in that The determining, based on the first voice signal and the second voice signal, a delay duration of the first voice signal relative to the second voice signal includes: Extracting a first set of audio fingerprints and a first set of timestamps corresponding to the first set of audio fingerprints from the first speech signal, and extracting a second set of audio fingerprints and a second set of timestamps corresponding to the second set of audio fingerprints from the second speech signal, wherein the first set of audio fingerprints includes one or more audio fingerprints, and the second set of audio fingerprints includes one or more audio fingerprints; If a first audio fingerprint in the first set of audio fingerprints matches a second audio fingerprint in the second set of audio fingerprints, obtaining a first timestamp in the first set of timestamps corresponding to the first audio fingerprint and a second timestamp in the second set of timestamps corresponding to the second audio fingerprint; The delay duration is determined as the time interval between the first timestamp and the second timestamp.
9. The method according to claim 8, characterized in that The adjusting, according to the delay duration, the first voice signal to be a seventh voice signal aligned with the second voice signal includes: The first speech signal is shifted forward or backward in time by the delay duration so that the first group of audio fingerprints and the second group of audio fingerprints are aligned in terms of timestamps, thereby obtaining the seventh speech signal aligned with the second speech signal.
10. The method according to claim 1, characterized in that The method further comprises: Displaying the participant identifiers of the multiple terminals on the target display interface of the first terminal, wherein the participant identifiers of the multiple terminals include the participant identifier of the second terminal; obtaining the target setting instruction on the target display interface, wherein the target setting instruction is used to select the participant identifier of the second terminal and instruct the first terminal to be set as an extended terminal of the second terminal to participate in the target call conference being conducted online; or In the case where the first terminal obtains the target information pushed by the second terminal, in response to the target information, a target prompt option is displayed on the target display interface of the first terminal, wherein the target prompt option is used to prompt whether to agree to set the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online; the target setting instruction is obtained on the target display interface, wherein the target setting instruction is used to select the agree option in the target prompt option and instruct to set the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online.
11. The method according to claim 10, characterized in that The method further comprises: In a case where the first terminal and the second terminal perform near field communication, the first terminal obtains the target information pushed by the second terminal, wherein the target information is used to trigger display of the target prompt option on the target display interface.
12. A multi-terminal communication device, characterized in that: include: An acquisition module, configured to acquire a target setting instruction on a target display interface of the first terminal; In response to the target setting instruction, setting the first terminal as an extended terminal of the second terminal to participate in the online target call conference, so as to use the first terminal to assist the second terminal in collecting voice signals; Obtaining a first voice signal collected by the first terminal and a second voice signal collected by the second terminal, wherein the terminals participating in the target call conference online include multiple terminals, the multiple terminals include the first terminal and the second terminal, and the first terminal and the second terminal are in a full-duplex scenario; A mixing module is configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal; and when a fourth voice signal received and played on the second terminal is collected by the first terminal, perform echo cancellation processing on the first voice signal according to the fourth voice signal sent to the second terminal to obtain a fifth voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals; perform the target mixing processing on the fifth voice signal and the second voice signal to obtain the third voice signal; and a sending module is configured to send the third voice signal to a terminal other than the first terminal and the second terminal among the multiple terminals for playback; The device is also used to: adjust the data flow of the first terminal based on the second terminal, establish a buffer zone for the audio data of the second terminal and the audio data of the first terminal, and when the audio data of the first terminal arrives at the mixing server before the audio data of the second terminal, buffer the audio data of the first terminal, and the buffer length is the delay duration; when the audio data of the first terminal arrives at the mixing server later than the audio data of the second terminal, clear the audio data that arrived earlier in the buffer zone, and the data clearing length is the delay duration, and the delay duration is used to adjust the first voice signal to be aligned with the second voice signal.
13. The device according to claim 12, characterized in that The apparatus is configured to obtain a first voice signal collected by the first terminal and a second voice signal collected by the second terminal in the following manner: when the first terminal includes N terminals, obtaining N channels of the first voice signals collected by the N terminals and one channel of the second voice signal collected by the second terminal, wherein N is a natural number of 1 or greater, and each of the N terminals is configured to collect one channel of the first voice signal; The device is used to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner: performing the target mixing processing on the N channels of the first voice signals and the one channel of the second voice signal to obtain one channel of the third voice signal.
14. The device according to claim 12, characterized in that The apparatus is configured to perform echo cancellation processing on the first voice signal according to the fourth voice signal sent to the second terminal in the following manner to obtain a fifth voice signal: The fourth voice signal sent to the second terminal is eliminated from the first voice signal to obtain the fifth voice signal.
15. The device according to claim 12, characterized in that The apparatus is configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner: performing the target mixing process on the first voice signal and the second voice signal to obtain a sixth voice signal; When the fourth voice signal received and played on the second terminal is collected by the first terminal, the sixth voice signal is echo-cancelled according to the fourth voice signal sent to the second terminal to obtain the third voice signal, wherein the fourth voice signal is a voice signal collected by a terminal other than the first terminal and the second terminal among the multiple terminals.
16. The device according to claim 15, characterized in that The apparatus is configured to perform the target mixing process on the first voice signal and the second voice signal in the following manner to obtain the sixth voice signal: In a case where the first terminal includes N terminals, the N terminals collect N first voice signals, and the second terminal collects one second voice signal, the target mixing processing is performed on the N first voice signals and the one second voice signal to obtain one sixth voice signal, wherein N is a natural number of 1 or greater, and each of the N terminals is used to collect one first voice signal.
17. The device according to claim 15, characterized in that The apparatus is configured to perform echo cancellation processing on the sixth voice signal according to the fourth voice signal sent to the second terminal in the following manner to obtain the third voice signal: The fourth voice signal sent to the second terminal is eliminated from the sixth voice signal to obtain the third voice signal.
18. The device according to claim 12, characterized in that The apparatus is configured to perform target mixing processing on the first voice signal and the second voice signal to obtain a third voice signal in the following manner: Determining a delay duration of the first voice signal relative to the second voice signal according to the first voice signal and the second voice signal; adjusting the first voice signal to a seventh voice signal aligned with the second voice signal according to the delay duration; The seventh voice signal and the second voice signal are mixed to obtain a third voice signal.
19. The device according to claim 18, characterized in that The apparatus is configured to determine a delay duration of the first voice signal relative to the second voice signal based on the first voice signal and the second voice signal in the following manner: Extracting a first set of audio fingerprints and a first set of timestamps corresponding to the first set of audio fingerprints from the first speech signal, and extracting a second set of audio fingerprints and a second set of timestamps corresponding to the second set of audio fingerprints from the second speech signal, wherein the first set of audio fingerprints includes one or more audio fingerprints, and the second set of audio fingerprints includes one or more audio fingerprints; If a first audio fingerprint in the first set of audio fingerprints matches a second audio fingerprint in the second set of audio fingerprints, obtaining a first timestamp in the first set of timestamps corresponding to the first audio fingerprint and a second timestamp in the second set of timestamps corresponding to the second audio fingerprint; The delay duration is determined as the time interval between the first timestamp and the second timestamp.
20. The device according to claim 19, characterized in that The apparatus is configured to adjust the first voice signal to a seventh voice signal aligned with the second voice signal according to the delay duration in the following manner: The first speech signal is shifted forward or backward in time by the delay duration so that the first group of audio fingerprints and the second group of audio fingerprints are aligned in terms of timestamps, thereby obtaining the seventh speech signal aligned with the second speech signal.
21. The device according to claim 12, characterized in that The device is also used for: Displaying the participant identifiers of the multiple terminals on the target display interface of the first terminal, wherein the participant identifiers of the multiple terminals include the participant identifier of the second terminal; obtaining the target setting instruction on the target display interface, wherein the target setting instruction is used to select the participant identifier of the second terminal and instruct the first terminal to be set as an extended terminal of the second terminal to participate in the target call conference being conducted online; or In the case where the first terminal obtains the target information pushed by the second terminal, in response to the target information, a target prompt option is displayed on the target display interface of the first terminal, wherein the target prompt option is used to prompt whether to agree to set the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online; the target setting instruction is obtained on the target display interface, wherein the target setting instruction is used to select the agree option in the target prompt option and instruct to set the first terminal as an extended terminal of the second terminal to participate in the target call conference conducted online.
22. The device according to claim 21, characterized in that The device is also used for: In a case where the first terminal and the second terminal perform near field communication, the first terminal obtains the target information pushed by the second terminal, wherein the target information is used to trigger display of the target prompt option on the target display interface.
23. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein the program can be executed by a terminal device or a computer to execute the method described in any one of claims 1 to 11.
24. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 11 through the computer program.
Citation Information
Patent Citations
Voice communication method and device, electronic equipment and computer readable storage medium
CN110602327A
Sound mixing method, device, equipment and system and readable storage medium
CN110995946A
Audio synthesis method and device and computer readable storage medium
CN111640411A
Audio and video conference system and equipment
CN210112145U