Methods and devices for audio playback and device management based on group conversations
By adding audio watermarks to the audio data, the problems of echo and feedback in group sessions were solved, device management was enabled, and session quality was improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-18
- Publication Date
- 2026-04-03
AI Technical Summary
In group conversation scenarios, when multiple users are in the same room, the microphone repeatedly picks up the content played by the speakers of other users' terminals, resulting in echoes and feedback, which affects the quality of the conversation.
By adding audio watermarks to audio data and associating the watermarks with the device identifiers of the terminals, terminals in the same space can be identified and managed, such as muting or using headphones, to avoid echoes and feedback.
It effectively avoids echoes and feedback, improving the session quality of group conversations.
Smart Images

Figure CN113516991B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio data processing, and in particular to a method, apparatus, and computer device for audio playback and device management based on group conversations. Background Technology
[0002] With the development of internet and cloud computing technologies, group conversations relying on the internet and cloud servers are becoming increasingly common. In a group conversation scenario, when a user is speaking, the terminal used by that user sends the collected audio data to a cloud server, which then distributes the audio data to other users.
[0003] In the aforementioned group conversation scenario, when multiple users are in the same room and all users' microphones are on, the microphones will repeatedly pick up the content played by the speakers of other users' terminals. This will generate echoes and feedback, severely impacting conversation quality. Therefore, in group conversation scenarios, accurately identifying which terminals are in the same space to prevent the audio played by the speaker of one terminal in the same space from being repeatedly picked up by the microphones of other terminals, thus avoiding echoes and feedback during the conversation and improving conversation quality, is an important research direction. Summary of the Invention
[0004] This application provides an audio playback and device management method, apparatus, and computer device based on group sessions, which can avoid echo and feedback during sessions and improve session quality. The technical solution is as follows:
[0005] On the one hand, a method for audio playback based on group sessions is provided, which includes:
[0006] It is determined that the speaker of the first terminal is turned on, and the first terminal is the terminal participating in the target session;
[0007] An audio watermark is added to the first audio data to be played to obtain the second audio data. The audio watermark is determined based on the session identifier of the target session and the device identifier of the first terminal.
[0008] The second audio data is played through the speaker.
[0009] On the one hand, a device management method based on group sessions is provided, which includes:
[0010] The second terminal collects audio data and is the terminal participating in the target session;
[0011] In response to the collected audio data, watermark detection is performed on the audio data;
[0012] In response to the detection of an audio watermark in the audio data, it is determined that the second terminal and the terminal corresponding to the audio watermark are in the same space;
[0013] The system displays a first prompt message, which instructs the user to disable the voice function of the second terminal.
[0014] In one possible implementation, determining at least one watermark loading location in the audio data includes any of the following:
[0015] Obtain the cepstrum of the audio data, and determine the positions in the cepstrum where the peak value is greater than a first threshold as the watermark loading positions;
[0016] The audio data is subjected to discrete cosine transform to obtain the energy intensity corresponding to each position of the audio data, and the positions with energy intensity greater than the second threshold are determined as the watermark loading positions.
[0017] On the one hand, a method for audio playback based on group sessions is provided, which includes:
[0018] Receive watermark detection results and audio data sent by a second terminal, which is a terminal participating in the target session;
[0019] Based on the watermark detection result, it is determined that a target terminal exists among the participating terminals of the target session, and the target terminal and the second terminal are in the same space;
[0020] The audio data is forwarded to other participating terminals in the target session for playback. These other participating terminals are terminals other than the second terminal and the target terminal.
[0021] In one possible implementation, the method further includes:
[0022] Based on the watermark detection result, the mixing paths of the target terminal and the second terminal in the mixing topology are removed. The mixing topology includes the mixing paths between each terminal in the target session.
[0023] In one possible implementation, the method further includes:
[0024] Receive third audio data sent by a fourth terminal, which is a terminal in the target session that is in a different space from the target terminal and the second terminal;
[0025] In response to the fact that the speakers of both the target terminal and the second terminal are turned on, a data receiving terminal is determined from the target terminal and the second terminal based on the device types of the target terminal and the second terminal.
[0026] The third audio data is forwarded to the data receiving terminal.
[0027] On the one hand, an audio playback device based on group conversation is provided, the device comprising:
[0028] The determination module is used to determine that the speaker of the first terminal is in the on state, and that the first terminal is a terminal participating in the target session;
[0029] The watermarking module is used to add an audio watermark to the first audio data to be played, thereby obtaining the second audio data. The audio watermark is determined based on the session identifier of the target session and the device identifier of the first terminal.
[0030] A playback module for playing the second audio data through the speaker.
[0031] In one possible implementation, the watermark adding module includes:
[0032] The acquisition unit is used to obtain the watermark text based on the session identifier of the target session and the device identifier of the first terminal;
[0033] The encoding unit is used to perform source coding and channel coding on the watermark text to obtain the watermark sequence;
[0034] The loading unit is used to load the watermark sequence into the first audio data to obtain the second audio data.
[0035] In one possible implementation, the loading unit includes:
[0036] The location determination subunit is used to determine at least one watermark loading location in the first audio data based on the energy spectrum envelope of the first audio data.
[0037] A loading subunit is used to load the watermark sequence at at least one watermark loading position to obtain the second audio data.
[0038] In one possible implementation, this location-determining sub-unit is used for:
[0039] The energy spectrum envelope of the first audio data is compared with a reference threshold.
[0040] The position corresponding to the energy spectrum envelope of the first audio data that is greater than the reference threshold is determined as the at least one watermark loading position.
[0041] On the one hand, a device management device based on group sessions is provided, the device comprising:
[0042] The acquisition module is used to acquire audio data, and the second terminal is the terminal participating in the target session;
[0043] The detection module is used to detect watermarks on the collected audio data in response to the data being collected.
[0044] The determination module is used to determine, in response to the detection of an audio watermark in the audio data, that the second terminal and the terminal corresponding to the audio watermark are in the same space;
[0045] The display module is used to display a first prompt message, which indicates that the voice function of the second terminal should be turned off.
[0046] In one possible implementation, the detection module includes:
[0047] The demodulation unit is used to perform watermark demodulation on the audio data to obtain a watermark sequence.
[0048] The decoding unit is used to perform channel decoding and source decoding on the watermark sequence to obtain the watermark text, which includes the device identifier of the terminal participating in the target session.
[0049] In one possible implementation, the demodulation unit includes:
[0050] A location determination subunit is used to determine at least one watermark loading location in the audio data;
[0051] A demodulation subunit is used to perform watermark demodulation on the audio data based on the at least one watermark loading position to obtain a watermark sequence.
[0052] In one possible implementation, the location-determined subunit is used to perform any of the following:
[0053] Obtain the cepstral spectrum of the audio data, and determine the positions where the peak value in the cepstral spectrum is greater than the first threshold as the watermark loading positions;
[0054] Perform a discrete cosine transform on the audio data to obtain the energy intensity corresponding to each position of the audio data, and determine the positions where the energy intensity is greater than the second threshold as the watermark loading positions.
[0055] In one possible implementation, the device further includes:
[0056] The data processing module is used to process the audio data based on the watermark detection results;
[0057] The sending module is used to send the watermark detection result and the processed audio data to the server, which is used to forward the processed audio data based on the watermark detection result.
[0058] In one possible implementation, the data processing module is used to perform any of the following:
[0059] The audio energy of the audio data is attenuated.
[0060] Echo cancellation is performed on the audio data based on the watermark detection results;
[0061] The audio data is then denoised based on the watermark detection result.
[0062] Mute the audio data.
[0063] On the one hand, an audio playback device based on group conversation is provided, the device comprising:
[0064] The receiving module is used to receive the watermark detection result and audio data sent by the second terminal, which is the terminal participating in the target session;
[0065] The determination module is used to determine, based on the watermark detection result, that a target terminal exists among the participating terminals of the target session, and that the target terminal and the second terminal are in the same space;
[0066] The forwarding module is used to forward the audio data to other participating terminals of the target session for playback. These other participating terminals are terminals other than the second terminal and the target terminal.
[0067] In one possible implementation, the device further includes a sending module, configured to: send a second prompt message to the target terminal, the second prompt message indicating that the target terminal and the second terminal are in the same space; and send a third prompt message to a third terminal, the third terminal being the management terminal of the target session, the third prompt message indicating that the target terminal and the second terminal are in the same space and that the voice function of the target terminal or the second terminal needs to be turned off.
[0068] In one possible implementation, the device further includes a removal module for: removing the mixing paths of the target terminal and the second terminal in the mixing topology based on the watermark detection result, the mixing topology including the mixing paths between the terminals in the target session.
[0069] In one possible implementation, the receiving module is used to receive third audio data sent by a fourth terminal, which is a terminal in the target session that is in a different space from the target terminal and the second terminal.
[0070] The determining module is used to determine the data receiving terminal from the target terminal and the second terminal based on the device type of the target terminal and the second terminal in response to the fact that the speakers of the target terminal and the second terminal are both turned on.
[0071] This forwarding module is used to forward the third audio data to the data receiving terminal.
[0072] On one hand, a computer device is provided, the computer device including one or more processors and one or more memories, the one or more memories storing at least one piece of program code, the at least one piece of program code being loaded and executed by the one or more processors to implement the operations performed by the group session-based audio playback or device management method.
[0073] On one hand, a computer-readable storage medium is provided that stores at least one piece of program code, which is loaded and executed by a processor to implement the operations performed by the group session-based audio playback or device management method.
[0074] On one hand, a computer program product is provided, comprising at least one piece of program code stored in a computer-readable storage medium. A processor of a computer device reads the at least one piece of program code from the computer-readable storage medium and executes the at least one piece of program code, causing the computer device to perform the operations performed by the group session-based audio playback or device management method.
[0075] The technical solution provided in this application adds an audio watermark to the audio data to be played during a cloud-based group session. Since the audio watermark is associated with the device identifier of the terminal, it can indicate which terminal played the audio data. When other terminals collect the audio data, they can determine which terminals are in the same space based on the audio watermark. This facilitates subsequent device management by the user, such as muting the device or connecting headphones. This avoids the audio played by the speaker of one terminal in the same space being repeatedly collected by the microphone of other terminals, thus avoiding echoes and howling during the session and improving the session quality of the group session. Attached Figure Description
[0076] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0077] Figure 1 This is a schematic diagram of an implementation environment for a group session provided in an embodiment of this application;
[0078] Figure 2 This is a schematic diagram illustrating an audio watermark loading and recognition process provided in an embodiment of this application;
[0079] Figure 3This is a flowchart of an audio playback method based on a group session provided in an embodiment of this application;
[0080] Figure 4 This is a schematic diagram of a watermark loading unit provided in an embodiment of this application;
[0081] Figure 5 This is a schematic diagram of the structure of a source data frame provided in an embodiment of this application;
[0082] Figure 6 This is a schematic diagram of the structure of a channel-coded frame provided in an embodiment of this application;
[0083] Figure 7 This is a schematic diagram of a watermark loading method provided in an embodiment of this application;
[0084] Figure 8 This is a schematic diagram of a watermark loading method provided in an embodiment of this application;
[0085] Figure 9 This is a flowchart of a device management method based on group sessions provided in an embodiment of this application;
[0086] Figure 10 This is a schematic diagram of a watermark parsing unit provided in an embodiment of this application;
[0087] Figure 11 This is a schematic diagram of a session interface provided in an embodiment of this application;
[0088] Figure 12 This is a flowchart of an audio data forwarding and playback method provided in an embodiment of this application;
[0089] Figure 13 This is a schematic diagram of another session interface provided in an embodiment of this application;
[0090] Figure 14 This is a schematic diagram of yet another session interface provided in an embodiment of this application;
[0091] Figure 15 This is a schematic diagram of the structure of an audio playback device based on a group session, provided in an embodiment of this application;
[0092] Figure 16 This is a schematic diagram of the structure of a device management device based on group sessions provided in an embodiment of this application;
[0093] Figure 17 This is a schematic diagram of the structure of an audio playback device based on a group session, provided in an embodiment of this application;
[0094] Figure 18 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application;
[0095] Figure 19 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0096] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0097] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.
[0098] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form a resource pool, which can be used on demand, making it flexible and convenient.
[0099] The technical solutions provided in this application can be applied to cloud conferencing scenarios. Cloud conferencing is an efficient, convenient, and low-cost conferencing format based on cloud computing technology. Users only need to perform simple and easy-to-use operations through an internet interface to quickly and efficiently share voice, data files, and videos synchronously with teams and clients around the world. The complex technologies such as data transmission and processing during the meeting are handled by the cloud conferencing service provider. Currently, domestic cloud conferencing mainly focuses on services based on the SaaS (Software as a Service) model, including telephone, network, and video services. Video conferencing based on cloud computing is called cloud conferencing. In the era of cloud conferencing, data transmission, processing, and storage are all handled by the computer resources of the video conferencing vendor. Users no longer need to purchase expensive hardware or install cumbersome software; they only need to open a browser, log in to the corresponding interface, and conduct efficient remote meetings. The cloud conferencing system supports dynamic cluster deployment of multiple servers and provides multiple high-performance servers, greatly improving the stability, security, and availability of the meeting.
[0100] Figure 1 This is a schematic diagram of an implementation environment for a group session provided in an embodiment of this application. See also... Figure 1 The implementation environment includes at least two terminals 101 and a server 102.
[0101] In this embodiment, both terminals 101 are user-side devices, and each terminal 101 has a target application that supports group sessions installed and running, such as a social networking application or an instant messaging application. In this embodiment, the at least two terminals 101 are terminals participating in the same session. The terminals 101 can be smartphones, tablets, laptops, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc., and this embodiment does not limit the specific devices used.
[0102] Server 102 is used to provide background services for the target application running on terminal 101, such as providing support for group sessions. Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0103] The terminal 101 and the server 102 can be directly or indirectly connected via wired or wireless communication, and this application embodiment does not limit this.
[0104] This application provides a method for audio playback and device management based on group sessions. It uses audio watermarking to accurately locate multiple terminals in the same space within a group session, enabling device management for each terminal. This avoids echoes and feedback caused by close proximity of multiple terminals in group session scenarios, thus improving session quality. The technical solution provided in this application can be combined with various scenarios, such as cloud conferencing, online teaching, and telemedicine. Figure 2 This is a schematic diagram of an audio watermark loading and recognition process provided in an embodiment of this application. The following is in conjunction with... Figure 2This application provides a brief overview of its embodiments. In this application, the first terminal 201 participating in the target session inputs first audio data obtained from the server 202 into the downlink audio packet processing unit 203. The downlink audio packet processing unit 203 performs audio decoding, network jitter processing, mixing, and sound enhancement on the first audio data. The first terminal 201 inputs a data packet into the downlink data packet processing unit 204. This data packet includes the session identifier of the target session and the device identifier of the first terminal. The downlink data packet processing unit 204 outputs watermark text based on this data packet. The watermark loading unit 205 adds the watermark text to the audio data output by the downlink audio packet processing unit 203, resulting in second audio data with an added audio watermark. This second audio data is then played by the speaker of the first terminal 201. Simultaneously, the second terminal 206 participating in the target session collects audio data and inputs the collected audio data into the watermark parsing unit 207 and the uplink audio packet processing unit 208. The second terminal extracts the watermark text from the audio data through the watermark parsing unit 207, and inputs the parsed watermark text into the uplink data packet processing unit 209. The uplink data packet processing unit 209 analyzes the watermark text to obtain the watermark parsing result, which determines whether there is a terminal in the same space as the second terminal among the terminals participating in the target session. In this embodiment, if there is a terminal in the same space as the second terminal, the second terminal can display a prompt message to remind the user to mute or use headphones. In this embodiment, the uplink audio packet processing unit 208 can optimize the collected audio data based on the watermark detection result output by the uplink data packet processing unit 209. The second terminal 206 sends the optimized audio data and the watermark detection result to the server 202, which forwards the data. In this embodiment, the server 202 can also send a prompt message to the administrator terminal based on the watermark detection result to remind the administrator to manage multiple terminals in the same space. By applying the technical solution provided in the embodiments of this application, when multiple terminals are detected to be in the same space, a prompt message is displayed on the terminal to prompt the user to mute or use headphones, thereby avoiding the sound played by one terminal in the same space being repeatedly collected by other terminals, eliminating echo and howling in the conversation, and thus improving the conversation quality of group conversations.
[0105] Figure 3 This is a flowchart illustrating an audio playback method based on a group session, as provided in an embodiment of this application. This method can be applied to the aforementioned implementation environment. In this embodiment, the terminal is used as the execution subject to describe the process of adding a watermark to the audio data. See [link to relevant documentation]. Figure 3 This embodiment may specifically include the following steps:
[0106] 301. The first terminal determines that the speaker is in the on state.
[0107] In this embodiment, the first terminal is any terminal participating in the target session, which is a group session. During the session, the first user inputs voice data through a microphone or other voice input device on the first terminal. The first terminal sends the collected audio data to the server, which then forwards the audio data, allowing other terminals participating in the target session to access the audio data collected by the first terminal. The first terminal can also retrieve and play audio data collected by other terminals from the server.
[0108] In one possible implementation, after joining the target session, the first terminal can detect the audio playback device. In response to detecting that the speaker is on, i.e. the terminal is in audio playback mode, the first terminal needs to add a watermark to the audio data before playing the audio data, i.e., perform the following step 302; in response to detecting that the speaker is off, or detecting that the terminal is connected to headphones, the first terminal can directly play the audio data through the headphones, i.e., it is not necessary to perform the following audio watermark addition step 302.
[0109] 302. The first terminal adds an audio watermark to the first audio data to be played, thus obtaining the second audio data.
[0110] The first audio data is audio data obtained by the first terminal from the server. The audio watermark is determined based on the session identifier of the target session and the device identifier of the first terminal. When any terminal performs watermark detection on the audio data, it can determine which terminal played the audio data based on the audio watermark. The session identifier is used to uniquely identify a session, and the device identifier is used to uniquely identify a terminal participating in the session. In one possible implementation, when the target session is created, the server can assign a session identifier to the target session and assign device identifiers to each terminal participating in the target session. Alternatively, the user account identifier logged into the terminal can be used to identify each terminal; this application embodiment does not limit this approach. In this application embodiment, assigning device identifiers to each terminal is used as an example for explanation.
[0111] In one possible implementation, step 302 described above can be implemented by the watermark loading unit in the first terminal. Figure 4 This is a schematic diagram of a watermark loading unit provided in an embodiment of this application. See also... Figure 4 The watermark loading unit includes a source coding unit 401, a channel coding unit 402, an audio preprocessing unit 403, and a watermark generation unit 404. The following is a combination of... Figure 4 The process of adding an audio watermark to the first audio data is explained below:
[0112] Step 1: The first terminal receives the session identifier based on the target session and the device identifier of the first terminal to obtain the watermark text.
[0113] For example, the first terminal can concatenate the session identifier of the target session and the device identifier of the first terminal to obtain the watermark text. Of course, the watermark text may also contain other information, which is not limited in this embodiment of the application.
[0114] Step 2: The first terminal will perform source coding and channel coding on the watermarked text to obtain the watermark sequence.
[0115] The watermark sequence can be represented as a binary bit sequence.
[0116] In one possible implementation, after obtaining the watermark text, the first terminal first performs source encoding on the watermark text. For example, firstly, the first terminal determines the byte length of the watermark text; then, it divides the watermark text into content byte packets of byte length; finally, it adds the total byte length of the watermark text and the byte sequence number of the current content byte packet to the header of each content byte packet, and adds a checksum to the tail of each content byte packet to obtain the source data frame. The checksum can be a 32-bit CRC (Cyclic Redundancy Check) code, a parity check code, or a block check code, etc., and this embodiment does not limit the specific implementation. Figure 5 This is a schematic diagram of the structure of a source data frame provided in an embodiment of this application. See also... Figure 5 A source data frame includes a total byte length of 501 for the watermarked text, a byte sequence number of 502, a content byte of 503, and a checksum of 504.
[0117] In one possible implementation, the first terminal performs channel coding on each source data frame to improve the recognition rate and robustness of subsequent watermark parsing. For example, the first terminal adds a synchronization code to the header and an error correction code to the tail of each source data frame to obtain a channel-coded frame, which is the watermark sequence. The synchronization code is a preset reference code sequence used for frame synchronization during data transmission. The length and specific content of this reference code sequence are set by the developers, and this embodiment does not limit this. For example, the synchronization code can be a 13-bit Barker code. The error correction code is used to reduce the bit error rate at the receiver when the channel signal-to-noise ratio is poor. The length and specific content of this error correction code can also be set by the developers, and this embodiment does not limit this. For example, the error correction code can be a 63-bit BCH (Bose, Ray-Chaudhuri, Hocquenghem) code. Figure 6 This is a schematic diagram of the structure of a channel-coded frame provided in an embodiment of this application. See also... Figure 6 Each channel-coded frame includes a synchronization code 601, a data packet 602 corresponding to the source data frame, and an error correction code 603.
[0118] It should be noted that the above description of the source coding and channel coding methods is merely an illustrative example, and the embodiments of this application do not limit which specific method is used for source coding and channel coding. In the embodiments of this application, communication quality improvement methods such as synchronization, error detection, and error correction are applied in the source coding and channel coding stages to reduce the bit error rate of subsequent data transmission and improve the efficiency and accuracy of subsequent watermark detection.
[0119] In this embodiment, the channel coding unit needs to send each channel coding frame to the watermark generation unit. The watermark generation unit determines the watermark sequence based on the data in each channel coding frame. In one possible implementation, since packet loss and bit errors may occur during data transmission, the channel coding unit can repeatedly send channel coding frames to the watermark generation unit. The watermark generation unit performs data deduplication and data splicing based on the information in the packet header and packet tail of each channel coding frame to obtain a complete and accurate watermark sequence.
[0120] Step 3: The first terminal loads the watermark sequence into the first audio data to obtain the second audio data.
[0121] In one possible implementation, the first terminal obtains the energy spectrum envelope of the first audio data through an audio preprocessing unit. This energy spectrum envelope can be used to indicate the energy intensity of each audio frame. Based on the energy spectrum envelope of the first audio data, the first terminal determines at least one watermark loading position in the first audio data. For example, the first terminal compares the energy spectrum envelope of the first audio data with a reference threshold, and determines the position corresponding to the energy spectrum envelope of the first audio data that is greater than the reference threshold as the at least one watermark loading position. The reference threshold can be set by the developer, and this application embodiment does not limit this. In this application embodiment, determining the position with higher energy intensity in the audio data as the watermark loading position and performing watermark loading can effectively avoid the audio watermark interfering with the lower energy audio, avoid the loss of effective information in the audio frame, and thus ensure the accuracy of the subsequent decoding process.
[0122] In the embodiments of this application, the first terminal loads the watermark sequence at at least one watermark loading location to obtain the second audio data. In one possible implementation, the first terminal can load the audio watermark in the time domain based on the temporal masking characteristics of the human ear, converting the watermark sequence into early reflections with different delays, thereby hiding the watermark sequence in the audio data, i.e., applying a temporal watermark generation technique based on echo hiding. Figure 7This is a schematic diagram of a watermark loading method provided in an embodiment of this application. See also... Figure 7 The following example illustrates loading an audio watermark at a single watermark loading position. For instance, the first terminal can first encrypt the watermark sequence, converting each element in the watermark sequence into a PN (Pseudo-Noise Code) sequence 701. For each element in the watermark sequence, based on the watermark loading position 702 and the delay parameter 703 corresponding to that element, the PN sequence of that element is inserted into the audio data. Different elements can correspond to different delay parameters, and these delay parameters and the correspondence between them and the elements are set by the developers; this embodiment does not limit this.
[0123] In one possible implementation, the first terminal can load audio watermarks in the transform domain based on the frequency domain masking characteristics of the human ear, converting the watermark sequence into energy fluctuations on different frequency sub-bands, thereby hiding the watermark sequence in the audio data. This is an application of DCT (Discrete Cosine Transform) domain watermark generation technology based on the spread spectrum principle. Figure 8 This is a schematic diagram of a watermark loading method provided in an embodiment of this application. See also... Figure 8 For example, the first terminal performs a DCT domain transform on the audio data 801 to obtain the energy intensity sequence corresponding to the audio data 801. The first terminal encrypts the watermark sequence, converting each element in the watermark sequence into a PN (Pseudo-Noise Code) sequence 802. Then, based on the determined watermark loading position, it obtains an element 803 corresponding to a watermark loading position from the energy intensity sequence, multiplies this element 803 with an element 804 in the watermark sequence, and loads the multiplication result into the audio data to obtain the audio data 805 with the audio watermark added.
[0124] It should be noted that the above description of the method for adding an audio watermark to the first audio data is merely an illustrative example, and this application embodiment does not limit which specific method is used to add the audio watermark. Of course, before adding an audio watermark to the first audio data, the first terminal may also perform post-processing enhancements such as network impairment repair and sound enhancement on the first audio data, and this application embodiment does not limit this.
[0125] 303. The first terminal plays the second audio data through the speaker.
[0126] In this embodiment of the application, after the first terminal obtains the second audio data with the added audio watermark, it can play the second audio data through the speaker.
[0127] The technical solution provided in this application adds an audio watermark, imperceptible to the human ear, to the audio data to be played during a session. Since the audio watermark is associated with the device identifier of the terminal, it can indicate which terminal played the audio data. When other terminals collect the audio data, they can determine which terminals are in the same space based on the audio watermark, which facilitates subsequent device management by the user, such as muting the device or connecting headphones, thereby avoiding echoes and howling during the session.
[0128] The above embodiments mainly describe the process of adding audio watermarks to audio data. In this embodiment, since the audio watermark is associated with the terminal's device identifier, the terminal can perform watermark detection on the collected audio data during the session to determine whether the collected audio data includes audio data that has been played by other terminals, and which terminal specifically played this audio data. This prompts the user to manage the device, for example, prompting the user to mute the terminal or use headphones to avoid echoes and feedback in group sessions. Figure 9 This is a flowchart illustrating a device management method based on group sessions provided in an embodiment of this application. This method can be applied to... Figure 1 In the implementation environment shown, this application embodiment uses a terminal as the execution subject to describe the method. See [link to relevant documentation]. Figure 9 The method may include the following steps:
[0129] 901. The second terminal collects audio data.
[0130] The second terminal is any terminal participating in the target session, which is a group session. During the session, the second terminal collects audio data in real time via a microphone. This audio data may include the user's voice data or audio data played by speakers from other terminals.
[0131] 902. The second terminal responds to the collected audio data by performing watermark detection on the audio data.
[0132] In one possible implementation, step 902 above can be implemented by the watermark parsing unit in the second terminal. Figure 10 This is a schematic diagram of a watermark parsing unit provided in an embodiment of this application. See also... Figure 10 The watermark parsing unit includes a watermark demodulation unit 1001, a channel decoding unit 1002, and a source decoding unit 1003. The following is a combination of... Figure 10 The watermark detection process is explained below:
[0133] Step 1: The second terminal performs watermark demodulation on the audio data to obtain the watermark sequence.
[0134] In this embodiment, the second terminal first determines at least one watermark loading position in the audio data. For watermark sequences loaded in the time domain, cepstral analysis can be used to analyze the acquired audio data to determine the watermark loading position. For example, the second terminal obtains the cepstral spectrum of the audio data and determines the position where the peak value in the cepstral spectrum is greater than a first threshold as the watermark loading position. For audio watermarks loaded in the transform domain, the second terminal performs a discrete cosine transform (DCT transform) on the audio data to obtain the energy intensity corresponding to each position of the audio data, and determines the position where the energy intensity is greater than a second threshold as the watermark loading position. The first and second positions can be set by the developer, and this embodiment does not limit this. It should be noted that the above description of the method for determining the watermark loading position is only an exemplary description, and this embodiment does not limit which specific method is used to determine the watermark loading position.
[0135] In one possible implementation, the second terminal performs watermark demodulation on the audio data based on the at least one watermark loading location to obtain a watermark sequence, that is, extracts the hidden watermark sequence from the audio data. It should be noted that this application embodiment does not limit the specific method used by the second terminal for watermark demodulation.
[0136] Step 2: The second terminal performs channel decoding and source decoding on the watermark sequence to obtain the watermark text.
[0137] In one possible implementation, the second terminal performs channel decoding on the watermark sequence demodulated from the audio data, i.e., each channel-coded frame. For example, the second terminal first performs cross-device bit alignment based on the synchronization code in the header of the channel-coded frame, and then corrects the bit errors generated during channel transmission based on the error correction code in the tail of the channel-coded frame. If the error correction is successful, the decoded data is output to the source decoding unit. If the number of bit errors exceeds the error correction capability of the error correction code after correction, i.e., the error correction fails, the data table is discarded, and the terminal waits to decode the next channel-coded frame.
[0138] In one possible implementation, the second terminal performs source decoding on the bitstream output by the channel decoding unit to obtain the watermark text. This watermark text includes the device identifier of the terminal participating in the target session, and of course, also includes the session identifier of the target session and other information, which are not limited in this embodiment. For example, the second terminal performs source-side error checking based on the checksum in the bitstream. If the check passes, it performs data packet content parsing, that is, parsing the content of the source data frame to obtain the total byte length of the watermark text, the byte sequence number of the source data frame, and the byte content. If the check fails, the data packet is discarded, and the next data packet is waited for decoding.
[0139] 903. In response to detecting the presence of an audio watermark in the audio data, the second terminal determines that it is in the same space as the terminal corresponding to the audio watermark and displays the first prompt message on the session interface.
[0140] In one possible implementation, after the second terminal extracts the watermark text from the audio data, it compares the session identifier in the watermark text with the session identifier sent by the server. If the two session identifiers are the same, it is determined that among the terminals participating in the target session in the collected audio data, there is a terminal in the same space as the second terminal. Furthermore, the second terminal determines which specific terminal is in the same space as the second terminal based on the device identifier in the watermark text.
[0141] In one possible implementation, the second terminal may display a first prompt message on the session interface of the target session based on the device identifier in the watermark text. The first prompt message is used to instruct the user to turn off the voice function of the second terminal, for example, to prompt the user to mute or use a headset for the call. Figure 11 This is a schematic diagram of a session interface provided in an embodiment of this application. See also... Figure 11 The session interface displays a first prompt message 1101 to prompt the user to adjust the terminal's voice function settings. In this embodiment, when a terminal is detected to be in the same space as the second terminal, i.e., when the second terminal is detected to be in a multi-terminal state in the same location, the client interface UI prompt can be triggered to inform the user which terminals are currently closest to the user and prompt the user to check the microphone and speaker.
[0142] 904. In response to detecting the presence of an audio watermark in the audio data, the second terminal processes the audio data based on the watermark detection result and sends the watermark detection result and the processed audio data to the server for data forwarding.
[0143] In this embodiment, the second terminal can further process the collected audio data based on the watermark detection result, that is, optimize the audio data to eliminate echo and howling in the audio data, and then send the optimized audio data and watermark detection result to the server corresponding to the target session, so that the server can perform subsequent data forwarding steps.
[0144] In one possible implementation, the second terminal optimizes the audio data using any of the following methods.
[0145] Implementation Method 1: The second terminal attenuates the audio energy of the audio data. For example, the second terminal can attenuate the audio data energy using an attenuator. This application embodiment does not limit the specific method of attenuation processing. In this application embodiment, by attenuating the audio energy, the energy of the sound fed back by other terminals in the same space can be reduced, thereby preventing echo leakage and reducing the probability of howling.
[0146] Implementation Method Two: The second terminal performs echo cancellation on the audio data based on the watermark detection result. For example, the second terminal is equipped with an echo cancellation unit. Based on the watermark detection result, the second terminal adjusts various parameters of the echo cancellation unit to enhance the intensity of the post-processing filter, thereby filtering out more echoes in the audio data. It should be noted that the specific method of echo cancellation performed by the second terminal in this embodiment is not limited.
[0147] Implementation Method 3: The second terminal performs noise reduction on the audio data based on the watermark detection result. For example, the second terminal is equipped with a noise reduction unit. After determining that an audio watermark exists in the audio data, the second terminal can enhance the noise reduction level of the noise reduction unit to remove more noise from the audio data.
[0148] Implementation Method Four: The second terminal mutes the audio data. For example, the second terminal can adjust the audio detection threshold during the audio acquisition stage. This audio detection threshold can be used to limit the loudness, energy, etc. of the audio data. This embodiment does not limit this; the specific content of the audio detection threshold is set by the developer. In one possible implementation, the second terminal can adjust the audio detection threshold to a larger value, determining audio data with audio energy, loudness, etc., below the audio detection threshold as muted. This increases the probability that audio data played by other terminals in the same space will be muted. For audio data determined to be muted, it is not necessary to send it to the server.
[0149] It should be noted that the above description of the audio data processing method is merely an exemplary illustration of several possible implementations, and the embodiments of this application do not limit which specific audio data processing method is used. In the embodiments of this application, the above-mentioned multiple implementations can be combined arbitrarily. For example, the second terminal can first perform echo cancellation on the acquired audio data and then perform attenuation processing, or it can first perform noise reduction on the audio data and then perform attenuation processing. The embodiments of this application do not limit which specific combination of methods is used to process the audio data.
[0150] It should be noted that in the embodiments of this application, the execution order is described as follows: first, the step 903 of displaying the prompt information is executed, and then the step 904 of audio data processing is executed. In some embodiments, the step of audio data processing may be executed first, and then the step of displaying the prompt information may be executed. Alternatively, the two steps may be executed simultaneously. The embodiments of this application do not limit this.
[0151] The technical solution provided in this application embodiment allows the second terminal to identify the audio watermark in the collected audio data and determine that there is another terminal in the same space as the target terminal participating in the target session. This prompts the user to turn off the current voice function to prevent the audio played by the target terminal's speaker from being repeatedly collected by the second terminal's microphone, thereby avoiding echoes and howling during the session and improving session quality.
[0152] The above embodiments mainly introduce the process of adding and parsing audio watermarks. In this embodiment, after the second terminal sends the watermark detection result and the optimized audio data to the server, the server can forward the audio data based on the watermark detection result, and the terminal can play the forwarded audio data. Figure 12 This is a flowchart illustrating audio data forwarding and playback provided in an embodiment of this application. See also... Figure 12 The method may include the following steps:
[0153] 1201. The server receives the watermark detection result and audio data sent by the second terminal.
[0154] The second terminal can be any terminal participating in the target session.
[0155] 1202. Based on the watermark detection result, the server determines that the target terminal exists among the participating terminals of the target session, and the target terminal is the second terminal in the same space.
[0156] In one possible implementation, the server obtains the session identifier in the watermark detection result. In response to the fact that the session identifier is the same as the session identifier of the current target session, it is determined that there is a terminal in the same space as the second terminal among the participating terminals of the target session. Based on the device identifier in the watermark detection result, it is determined which specific terminal it is. In this embodiment of the application, the example of the target terminal and the second terminal being in the same space is used for illustration.
[0157] 1203. The server forwards the audio data to other participating terminals in the target session for playback. These other participating terminals are terminals other than the second terminal and the target terminal.
[0158] In this embodiment, audio data collected by multiple terminals in the same space is not forwarded between them. That is, audio data collected by the second terminal is not forwarded to the target terminal, and audio data collected by the target terminal is not forwarded to the second terminal. This data forwarding mechanism avoids the terminal from repeatedly playing the user's input speech in the current space, thus preventing echoes and feedback.
[0159] 1204. Based on the watermark detection result, the server sends a prompt message to the target terminal and the administrator terminal.
[0160] In one possible implementation, the server sends a second prompt message to the target terminal. This second prompt message indicates that the target terminal and the second terminal are in the same space, and may prompt the user of the target terminal to connect headphones for a conversation.
[0161] In one possible implementation, the server sends a third prompt message to a third terminal. For example, after receiving watermark detection results from multiple terminals participating in the target session, the server summarizes the results, generates a third prompt message, and sends this message to the administrator user of the target session. This third terminal is the management terminal for the target session, and the third prompt message indicates that the target terminal and the second terminal are in the same space and that the voice function of either the target terminal or the second terminal needs to be disabled. Figure 13 This is a schematic diagram of another session interface provided in an embodiment of this application. See also... Figure 13 This is the administrator's session interface, which displays a third-party prompt message 1301. Figure 14 This is a schematic diagram of another session interface provided in an embodiment of this application. See also... Figure 14 This is an administrator's session interface, which displays a third prompt message 1401. It should be noted that this embodiment does not limit the specific display method of the prompt message.
[0162] 1205. Based on the watermark detection result, the server removes the mixing paths of the target terminal and the second terminal from the mixing topology, and performs subsequent audio data forwarding steps based on the updated mixing topology.
[0163] In one possible implementation, the server stores a mixing topology, which includes mixing paths between various terminals in the target session. After receiving audio data from any terminal, the server can mix the audio data based on this mixing topology before forwarding it. In this embodiment, audio data collected by multiple terminals in the same space does not require mixing. If a terminal simultaneously receives audio data collected by multiple terminals in the same space, the server can select the highest quality audio data path for forwarding; that is, audio data collected by the target terminal and the second terminal will not be forwarded to other terminals simultaneously. The audio quality can be determined based on factors such as the type of audio acquisition device, audio energy intensity, and signal-to-noise ratio.
[0164] In one possible implementation, when the server receives third audio data sent by a fourth terminal, where the fourth terminal is a terminal in the target session that is located in a different space from the target terminal and the second terminal, the server only needs to select one terminal from the target terminal and the second terminal for forwarding. For example, in response to the fact that the speakers of the target terminal and the second terminal are both on, the server determines the data receiving terminal from the target terminal and the second terminal based on their device types; and forwards the third audio data to the data receiving terminal. For example, when the speakers of all terminals are on, the server can determine the data receiving terminal according to the priority of professional telephone > laptop > mobile phone speaker > headphones. If the priorities of all terminals are the same, the user can be prompted to specify a data receiving terminal, or the user can set the data receiving priority of each terminal. This application embodiment does not limit this.
[0165] It should be noted that the specific execution order of steps 1203, 1204, and 1205 described above is not limited in the embodiments of this application.
[0166] The technical solution provided in this application sends the watermark detection result to the server, enabling the server to obtain the location distribution of each terminal participating in the target session during audio forwarding. Based on the location distribution of each terminal, the server selectively forwards audio data, thereby eliminating echo and howling in the session from the data forwarding stage and improving session quality.
[0167] In this embodiment, when multiple terminals in the same location are detected in a group conversation scenario, on the one hand, prompts can be displayed on the user's conversation interface to remind the user to check their devices and prevent problems such as echo and howling that could damage the audio. On the other hand, when the collected audio data includes sounds played from other terminals, the audio data is optimized to eliminate sounds from other devices and prevent echo leakage. Furthermore, the watermark detection results are sent to the server. Based on the terminal distribution indicated by the watermark detection results, the server changes the mixing topology, routes the audio data uploaded by multiple terminals in the same space, selects the highest quality route to forward to other terminals, and removes the mixing paths of multiple terminals in the same location to avoid mixing and forwarding duplicate data, prevent repeated playback of audio data, and improve the conversation quality of the group conversation.
[0168] The technical solutions provided in this application can also be used to manage the cameras, projection devices, and screen sharing devices of various terminals participating in a session. For example, when sharing a screen, this solution identifies multiple terminals in the same space and selects one device to send the shared video stream based on the device type of each terminal. For instance, if the multiple terminals in the same space are a large-screen TV and a laptop, the user can be advised to share the video stream on the large-screen TV to improve the video viewing experience and thus enhance the session experience.
[0169] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0170] Figure 15 This is a schematic diagram of the structure of an audio playback device based on a group session, as provided in an embodiment of this application. See also... Figure 15 The device includes:
[0171] The determination module 1501 is used to determine that the speaker of the first terminal is in the on state, and the first terminal is a terminal participating in the target session;
[0172] The watermarking module 1502 is used to add an audio watermark to the first audio data to be played to obtain the second audio data. The audio watermark is determined based on the session identifier of the target session and the device identifier of the first terminal.
[0173] Playback module 1503 is used to play the second audio data through the speaker.
[0174] In one possible implementation, the watermark adding module 1502 includes:
[0175] The acquisition unit is used to obtain the watermark text based on the session identifier of the target session and the device identifier of the first terminal;
[0176] The encoding unit is used to perform source coding and channel coding on the watermark text to obtain the watermark sequence;
[0177] The loading unit is used to load the watermark sequence into the first audio data to obtain the second audio data.
[0178] In one possible implementation, the loading unit includes:
[0179] The location determination subunit is used to determine at least one watermark loading location in the first audio data based on the energy spectrum envelope of the first audio data.
[0180] A loading subunit is used to load the watermark sequence at at least one watermark loading position to obtain the second audio data.
[0181] In one possible implementation, this location-determining sub-unit is used for:
[0182] The energy spectrum envelope of the first audio data is compared with a reference threshold.
[0183] The position corresponding to the energy spectrum envelope of the first audio data that is greater than the reference threshold is determined as the at least one watermark loading position.
[0184] The device provided in this application adds an audio watermark to the audio data to be played during a session. Since the audio watermark is associated with the device identifier of the terminal, it can indicate which terminal played the audio data. When other terminals collect the audio data, they can determine which terminals are in the same space based on the audio watermark, which facilitates subsequent device management by the user, such as muting the device or connecting headphones, thereby avoiding echo and howling during the session.
[0185] It should be noted that the audio playback device based on group sessions provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio playback device based on group sessions provided in the above embodiments and the audio playback method embodiments based on group sessions belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0186] Figure 16 This is a schematic diagram of a device management device based on group sessions provided in an embodiment of this application. See also... Figure 16 The device includes:
[0187] The acquisition module 1601 is used to acquire audio data, and the second terminal is the terminal participating in the target session;
[0188] The detection module 1602 is used to perform watermark detection on the acquired audio data in response to the acquisition of audio data;
[0189] The determination module 1603 is used to determine, in response to detecting that there is an audio watermark in the audio data, that the second terminal and the terminal corresponding to the audio watermark are in the same space.
[0190] Display module 1604 is used to display a first prompt message, which is used to indicate that the voice function of the second terminal is turned off.
[0191] In one possible implementation, the detection module 1602 includes:
[0192] The demodulation unit is used to perform watermark demodulation on the audio data to obtain a watermark sequence.
[0193] The decoding unit is used to perform channel decoding and source decoding on the watermark sequence to obtain the watermark text, which includes the device identifier of the terminal participating in the target session.
[0194] In one possible implementation, the demodulation unit includes:
[0195] A location determination subunit is used to determine at least one watermark loading location in the audio data;
[0196] A demodulation subunit is used to perform watermark demodulation on the audio data based on the at least one watermark loading position to obtain a watermark sequence.
[0197] In one possible implementation, the location-determined subunit is used to perform any of the following:
[0198] Obtain the cepstral spectrum of the audio data, and determine the positions where the peak value in the cepstral spectrum is greater than the first threshold as the watermark loading positions;
[0199] Perform a discrete cosine transform on the audio data to obtain the energy intensity corresponding to each position of the audio data, and determine the positions where the energy intensity is greater than the second threshold as the watermark loading positions.
[0200] In one possible implementation, the device further includes:
[0201] The data processing module is used to process the audio data based on the watermark detection results;
[0202] The sending module is used to send the watermark detection result and the processed audio data to the server, which is used to forward the processed audio data based on the watermark detection result.
[0203] In one possible implementation, the data processing module is used to perform any of the following:
[0204] The audio energy of the audio data is attenuated.
[0205] Echo cancellation is performed on the audio data based on the watermark detection results;
[0206] The audio data is then denoised based on the watermark detection result.
[0207] Mute the audio data.
[0208] The device provided in this application embodiment allows the second terminal to identify audio watermarks in the collected audio data and determine that among the terminals participating in the target session, there is also a target terminal in the same space as the second terminal. This prompts the user to turn off the current voice function to prevent the audio played by the target terminal's speaker from being repeatedly collected by the second terminal's microphone, thereby avoiding echoes and howling during the session and improving session quality.
[0209] It should be noted that the device management device based on group sessions provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device management device based on group sessions provided in the above embodiments and the device management method embodiments based on group sessions belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0210] Figure 17 This is a schematic diagram of the structure of an audio playback device based on a group session, as provided in an embodiment of this application. See also... Figure 17 The device includes:
[0211] The receiving module 1701 is used to receive the watermark detection result and audio data sent by the second terminal, which is a terminal participating in the target session;
[0212] The determination module 1702 is used to determine, based on the watermark detection result, that there is a target terminal among the participating terminals of the target session, and that the target terminal and the second terminal are in the same space;
[0213] The forwarding module 1703 is used to forward the audio data to other participating terminals of the target session for playback of the audio data. The other participating terminals are terminals other than the second terminal and the target terminal.
[0214] In one possible implementation, the device further includes a sending module, configured to: send a second prompt message to the target terminal, the second prompt message indicating that the target terminal and the second terminal are in the same space; and send a third prompt message to a third terminal, the third terminal being the management terminal of the target session, the third prompt message indicating that the target terminal and the second terminal are in the same space and that the voice function of the target terminal or the second terminal needs to be turned off.
[0215] In one possible implementation, the device further includes a removal module for: removing the mixing paths of the target terminal and the second terminal in the mixing topology based on the watermark detection result, the mixing topology including the mixing paths between the terminals in the target session.
[0216] In one possible implementation, the receiving module 1701 is used to receive third audio data sent by a fourth terminal, which is a terminal in the target session that is in a different space from the target terminal and the second terminal.
[0217] The determining module 1702 is used to determine a data receiving terminal from the target terminal and the second terminal based on the device type of the target terminal and the second terminal in response to the fact that the speakers of the target terminal and the second terminal are both turned on.
[0218] The forwarding module 1703 is used to forward the third audio data to the data receiving terminal.
[0219] The apparatus provided in this application sends the watermark detection result to a server, enabling the server to obtain the location distribution of each terminal participating in the target session during audio forwarding. Based on the location distribution of each terminal, the server selectively forwards audio data, thereby eliminating echo and howling in the session from the data forwarding stage and improving session quality.
[0220] It should be noted that the audio playback device based on group sessions provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio playback device based on group sessions provided in the above embodiments and the audio playback method embodiments based on group sessions belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0221] The computer equipment provided by the above technical solution can be implemented as a terminal or a server, for example, Figure 18This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal 1800 can be: a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 1800 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.
[0222] Typically, terminal 1800 includes one or more processors 1801 and one or more memories 1802.
[0223] Processor 1801 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1801 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1801 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1801 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1801 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0224] The memory 1802 may include one or more computer-readable storage media, which may be non-transitory. The memory 1802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1802 is used to store at least one piece of program code, which is executed by the processor 1801 to implement the group session-based audio playback method or the group session-based device management method provided in the method embodiments of this application.
[0225] In some embodiments, the terminal 1800 may also optionally include a peripheral device interface 1803 and at least one peripheral device. The processor 1801, memory 1802, and peripheral device interface 1803 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1803 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: radio frequency circuitry 1804, display screen 1805, camera assembly 1806, audio circuitry 1807, and power supply 1809.
[0226] Peripheral device interface 1803 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1801 and memory 1802. In some embodiments, processor 1801, memory 1802 and peripheral device interface 1803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1801, memory 1802 and peripheral device interface 1803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0227] The radio frequency (RF) circuit 1804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1804 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1804 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0228] Display screen 1805 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1805 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1801 for processing. In this case, display screen 1805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1805, which serves as the front panel of terminal 1800; in other embodiments, there may be at least two display screens, respectively disposed on different surfaces of terminal 1800 or in a folded design; in some embodiments, display screen 1805 may be a flexible display screen, disposed on a curved or folded surface of terminal 1800. Furthermore, display screen 1805 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 1805 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0229] The camera assembly 1806 is used to acquire images or videos. Optionally, the camera assembly 1806 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1806 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0230] The audio circuit 1807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 1801 for processing, or to the radio frequency circuit 1804 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 1800. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1801 or the radio frequency circuit 1804 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1807 may also include a headphone jack.
[0231] The power supply 1809 is used to power the various components in the terminal 1800. The power supply 1809 can be AC power, DC power, a disposable battery, or a rechargeable battery. When the power supply 1809 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0232] In some embodiments, the terminal 1800 further includes one or more sensors 1810. The one or more sensors 1810 include, but are not limited to: an acceleration sensor 1811, a gyroscope sensor 1812, a pressure sensor 1813, an optical sensor 1815, and a proximity sensor 1816.
[0233] Accelerometer 1811 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 1800. For example, accelerometer 1811 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1801 can control display screen 1805 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1811. Accelerometer 1811 can also be used for games or for acquiring user motion data.
[0234] The gyroscope sensor 1812 can detect the orientation and rotation angle of the terminal 1800. The gyroscope sensor 1812, in conjunction with the accelerometer sensor 1811, can collect 3D motion data from the user on the terminal 1800. Based on the data collected by the gyroscope sensor 1812, the processor 1801 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0235] The pressure sensor 1813 can be disposed on the side bezel of the terminal 1800 and / or on the lower layer of the display screen 1805. When the pressure sensor 1813 is disposed on the side bezel of the terminal 1800, it can detect the user's grip signal on the terminal 1800, and the processor 1801 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1813. When the pressure sensor 1813 is disposed on the lower layer of the display screen 1805, the processor 1801 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1805. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0236] An optical sensor 1815 is used to collect ambient light intensity. In one embodiment, the processor 1801 can control the display brightness of the display screen 1805 based on the ambient light intensity collected by the optical sensor 1815. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1805 is increased; when the ambient light intensity is low, the display brightness of the display screen 1805 is decreased. In another embodiment, the processor 1801 can also dynamically adjust the shooting parameters of the camera assembly 1806 based on the ambient light intensity collected by the optical sensor 1815.
[0237] The proximity sensor 1816, also known as a distance sensor, is typically located on the front panel of the terminal 1800. The proximity sensor 1816 is used to detect the distance between the user and the front of the terminal 1800. In one embodiment, when the proximity sensor 1816 detects that the distance between the user and the front of the terminal 1800 is gradually decreasing, the processor 1801 controls the display screen 1805 to switch from a screen-on state to a screen-off state; when the proximity sensor 1816 detects that the distance between the user and the front of the terminal 1800 is gradually increasing, the processor 1801 controls the display screen 1805 to switch from a screen-off state to a screen-on state.
[0238] Those skilled in the art will understand that Figure 18 The structure shown does not constitute a limitation on terminal 1800 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0239] Figure 19 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1900 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1901 and one or more memories 1902. Each memory 1902 stores at least one line of program code, which is loaded and executed by the one or more processors 1901 to implement the methods provided in the various method embodiments described above. Of course, the server 1900 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1900 may also include other components for implementing device functions, which will not be elaborated upon here.
[0240] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including at least one line of program code, which can be executed by a processor to perform the group session-based audio playback method or the group session-based device management method in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0241] In an exemplary embodiment, a computer program product is also provided, comprising at least one piece of program code stored in a computer-readable storage medium. A processor of a computer device reads the at least one piece of program code from the computer-readable storage medium and executes the at least one piece of program code, causing the computer device to perform the operations performed by the group session-based audio playback or device management method.
[0242] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program with at least one piece of program code associated with the hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0243] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An audio playback method based on group conversation, characterized in that, Applied to a first terminal, the method includes: It is determined that the speaker of the first terminal participating in the target session is turned on; The first audio data and data packet to be played are obtained from the server. The session identifier of the target session and the device identifier of the first terminal in the data packet are concatenated to obtain the watermark text. Based on the watermarked text, an audio watermark is added to the first audio data to obtain the second audio data; the second audio data is then played through a speaker. If the second audio data is acquired by a second terminal located in the same space, the second terminal is configured to: determine at least one watermark loading position in the second audio data; based on the at least one watermark loading position, perform watermark demodulation on the second audio data to obtain a watermark sequence; perform channel decoding and source decoding on the watermark sequence to obtain watermark text; compare the session identifier in the watermark text with the session identifier sent by the server; if the two session identifiers are the same, it is determined that among the terminals participating in the target session, there is a terminal located in the same space as the second terminal; based on the device identifier in the watermark text, display a first prompt message on the session interface of the target session, the first prompt message being used to indicate which user's terminal is closer to the second terminal and therefore it is recommended to disable the voice function of the second terminal; The server is configured to: receive watermark detection results and audio data sent by the second terminal; obtain a session identifier from the watermark detection results; in response to the session identifier in the watermark detection results being the same as the session identifier of the target session, determine that there is a target terminal in the target session that is in the same space as the second terminal; based on the device identifier in the watermark detection results, determine the target terminal in the same space as the second terminal, and forward the audio data to other participating terminals in the target session other than the second terminal and the target terminal for playback; send a second prompt message to the target terminal and a third prompt message to the administrator terminal, wherein the second prompt message is used to indicate that the target terminal and the second terminal are in the same space. When both the target terminal and the second terminal are in the same space, it is recommended that the user of the target terminal connect headphones for the conversation. The third prompt message indicates that since the target terminal and the second terminal are in the same space, it is recommended to turn off the voice function of the target terminal or the second terminal. When the speakers of both the target terminal and the second terminal are on, the server is further configured to: select the terminal with the highest priority from the target terminal and the second terminal as the data receiving terminal based on the device type; when the priorities of all terminals are the same, prompt the user to specify a data receiving terminal, or allow the user to set the data receiving priority of each terminal; and send third audio data from a terminal in the target session that is in a different space from the target terminal and the second terminal to the data receiving terminal.
2. The method according to claim 1, characterized in that, The step of adding an audio watermark to the processed first audio data based on the watermarked text to obtain second audio data includes: The watermarked text is subjected to source coding and channel coding to obtain a watermark sequence; The watermark sequence is loaded into the first audio data to obtain the second audio data.
3. The method according to claim 2, characterized in that, The step of loading the watermark sequence into the first audio data to obtain the second audio data includes: Based on the energy spectrum envelope of the first audio data, at least one watermark loading location in the first audio data is determined; The watermark sequence is loaded at at least one watermark loading position to obtain the second audio data.
4. The method according to claim 3, characterized in that, Determining at least one watermark loading location in the first audio data based on the energy spectrum envelope of the first audio data includes: The energy spectrum envelope of the first audio data is compared with a reference threshold; The position corresponding to the energy spectrum envelope in the first audio data that is greater than the reference threshold is determined as the at least one watermark loading position.
5. A device management method based on group sessions, characterized in that, Applied to a second terminal, the method includes: Audio data is collected, and the second terminal is the terminal participating in the target session; In response to the acquisition of second audio data, at least one watermark loading position is determined in the second audio data. Based on the at least one watermark loading position, the second audio data is watermark demodulated to obtain a watermark sequence. The watermark sequence is then subjected to channel decoding and source decoding to obtain watermark text. The session identifier in the watermark text is compared with the session identifier sent by the server. If the two session identifiers are the same, it is determined that among the terminals participating in the target session, there is a terminal in the same space as the second terminal. Based on the device identifier in the watermark text, a first prompt message is displayed on the session interface of the target session. The first prompt message is used to indicate that the second terminal's voice function is recommended to be turned off because the second terminal is close to which user's terminal. The watermark text is obtained by splicing the session identifier of the target session and the device identifier of the first terminal in the data packet sent by the server. The second audio data is obtained by the first terminal by adding an audio watermark to the first audio data based on the watermark text. The first audio data is obtained from the server. The server is configured to: receive watermark detection results and audio data sent by the second terminal; obtain a session identifier from the watermark detection results; in response to the session identifier in the watermark detection results being the same as the session identifier of the target session, determine that there is a target terminal in the target session that is in the same space as the second terminal; based on the device identifier in the watermark detection results, determine the target terminal in the same space as the second terminal, and forward the audio data to other participating terminals in the target session other than the second terminal and the target terminal for playback; send a second prompt message to the target terminal and a third prompt message to the administrator terminal, wherein the second prompt message is used to indicate that the target terminal and the second terminal are in the same space. When both the target terminal and the second terminal are in the same space, it is recommended that the user of the target terminal connect headphones for the conversation. The third prompt message indicates that since the target terminal and the second terminal are in the same space, it is recommended to turn off the voice function of the target terminal or the second terminal. When the speakers of both the target terminal and the second terminal are on, the server is further configured to: select the terminal with the highest priority from the target terminal and the second terminal as the data receiving terminal based on the device type; when the priorities of all terminals are the same, prompt the user to specify a data receiving terminal, or allow the user to set the data receiving priority of each terminal; and send third audio data from a terminal in the target session that is in a different space from the target terminal and the second terminal to the data receiving terminal.
6. The method according to claim 5, characterized in that, After determining that among the terminals participating in the target session there is a terminal in the same space as the second terminal, the method further includes: The second audio data is processed based on the watermark detection results; The watermark detection result and the processed second audio data are sent to the server, which is used to forward the processed second audio data based on the watermark detection result.
7. The method according to claim 6, characterized in that, The data processing of the second audio data based on the watermark detection result includes any one of the following: The audio energy of the second audio data is attenuated. Echo cancellation is performed on the second audio data based on the watermark detection results; The second audio data is denoised based on the watermark detection result; The second audio data is muted.
8. The method according to claim 5, characterized in that, Determining at least one watermark loading location in the second audio data includes any of the following: Obtain the cepstral spectrum of the second audio data, and determine the positions in the cepstral spectrum where the peak value is greater than the first threshold as the watermark loading positions; Perform a discrete cosine transform on the second audio data to obtain the energy intensity corresponding to each position of the second audio data, and determine the positions where the energy intensity is greater than the second threshold as the watermark loading positions.
9. An audio playback method based on group conversation, characterized in that, Applied to a server, the method includes: The system receives watermark detection results and audio data sent by a second terminal, which is a terminal participating in the target session. The audio data is obtained based on the second audio data collected by the second terminal. The second audio data is obtained by the first terminal adding an audio watermark to the first audio data based on the watermark text. The watermark text is obtained by the first terminal splicing the session identifier of the target session and the device identifier of the first terminal in the data packet sent by the server. The process involves obtaining the session identifier from the watermark detection result. If the session identifier in the watermark detection result matches the session identifier of the target session, it is determined that a target terminal exists in the same space as the second terminal among the participating terminals in the target session. Based on the device identifier in the watermark detection result, the target terminal in the same space as the second terminal is identified. The audio data is forwarded to other participating terminals in the target session besides the second terminal and the target terminal for playback. A second prompt message is sent to the target terminal, and a third prompt message is sent to the administrator terminal. The second prompt message indicates that since the target terminal and the second terminal are in the same space, it is recommended that the user of the target terminal connect headphones for the session. The third prompt message indicates that since the target terminal and the second terminal are in the same space, it is recommended that the voice function of either the target terminal or the second terminal be turned off. When both the speaker of the target terminal and the speaker of the second terminal are turned on, the terminal with the highest priority is selected as the data receiving terminal based on the device type. When the priorities of all terminals are the same, the user is prompted to specify a data receiving terminal, or the user can set the data receiving priority of each terminal. Third audio data from a terminal in the target session that is in a different space from the target terminal and the second terminal is sent to the data receiving terminal. The second terminal is configured to: determine at least one watermark loading position in the second audio data; based on the at least one watermark loading position, perform watermark demodulation on the second audio data to obtain a watermark sequence; perform channel decoding and source decoding on the watermark sequence to obtain watermark text; compare the session identifier in the watermark text with the session identifier sent by the server; if the two session identifiers are the same, it is determined that among the terminals participating in the target session, there is a terminal in the same space as the second terminal; based on the device identifier in the watermark text, display a first prompt message on the session interface of the target session, the first prompt message being used to indicate that the second terminal's voice function is recommended to be turned off because the second terminal is close to which user's terminal.
10. The method according to claim 9, characterized in that, The method further includes: Based on the watermark detection result, the mixing paths of the target terminal and the second terminal in the mixing topology are removed. The mixing topology includes the mixing paths between each terminal in the target session.
11. An audio playback device based on group conversation, characterized in that, The device includes: The determination module is used to determine whether the speaker of the first terminal participating in the target session is turned on; The watermarking module is used to obtain the first audio data and data packet to be played from the server, and to concatenate the session identifier of the target session and the device identifier of the first terminal in the data packet to obtain the watermark text. The watermarking module is used to add an audio watermark to the first audio data based on the watermark text to obtain the second audio data. A playback module for playing the second audio data through a speaker; If the second audio data is acquired by a second terminal located in the same space, the second terminal is configured to: determine at least one watermark loading position in the second audio data; based on the at least one watermark loading position, perform watermark demodulation on the second audio data to obtain a watermark sequence; perform channel decoding and source decoding on the watermark sequence to obtain watermark text; compare the session identifier in the watermark text with the session identifier sent by the server; if the two session identifiers are the same, it is determined that among the terminals participating in the target session, there is a terminal located in the same space as the second terminal; based on the device identifier in the watermark text, display a first prompt message on the session interface of the target session, the first prompt message being used to indicate which user's terminal is closer to the second terminal and therefore it is recommended to disable the voice function of the second terminal; The server is configured to: receive watermark detection results and audio data sent by the second terminal; obtain a session identifier from the watermark detection results; in response to the session identifier in the watermark detection results being the same as the session identifier of the target session, determine that there is a target terminal in the target session that is in the same space as the second terminal; based on the device identifier in the watermark detection results, determine the target terminal in the same space as the second terminal, and forward the audio data to other participating terminals in the target session other than the second terminal and the target terminal for playback; send a second prompt message to the target terminal and a third prompt message to the administrator terminal, wherein the second prompt message is used to indicate that the target terminal and the second terminal are in the same space. When both the target terminal and the second terminal are in the same space, it is recommended that the user of the target terminal connect headphones for the conversation. The third prompt message indicates that since the target terminal and the second terminal are in the same space, it is recommended to turn off the voice function of the target terminal or the second terminal. When the speakers of both the target terminal and the second terminal are on, the server is further configured to: select the terminal with the highest priority from the target terminal and the second terminal as the data receiving terminal based on the device type; when the priorities of all terminals are the same, prompt the user to specify a data receiving terminal, or allow the user to set the data receiving priority of each terminal; and send third audio data from a terminal in the target session that is in a different space from the target terminal and the second terminal to the data receiving terminal.
12. The apparatus according to claim 11, characterized in that, The watermark adding module includes: The encoding unit is used to perform source coding and channel coding on the watermark text to obtain a watermark sequence; The loading unit is used to load the watermark sequence into the first audio data to obtain the second audio data.
13. The apparatus according to claim 12, characterized in that, The loading unit includes: The location determination subunit is used to determine at least one watermark loading location in the first audio data based on the energy spectrum envelope of the first audio data. A loading subunit is used to load the watermark sequence at the at least one watermark loading position to obtain the second audio data.
14. The apparatus according to claim 13, characterized in that, The position determination subunit is used for: The energy spectrum envelope of the first audio data is compared with a reference threshold; The position corresponding to the energy spectrum envelope in the first audio data that is greater than the reference threshold is determined as the at least one watermark loading position.
15. A device management device based on group sessions, characterized in that, The device includes: The acquisition module is used to acquire audio data, and the second terminal is the terminal that participates in the target session; The detection module is used to respond to the acquisition of second audio data, determine at least one watermark loading position in the second audio data, perform watermark demodulation on the second audio data based on the at least one watermark loading position to obtain a watermark sequence, and perform channel decoding and source decoding on the watermark sequence to obtain watermark text. The determination module is used to compare the session identifier in the watermark text with the session identifier sent by the server. If the two session identifiers are the same, it is determined that among the terminals participating in the target session, there is a terminal in the same space as the second terminal. The display module is used to display a first prompt message on the session interface of the target session based on the device identifier in the watermark text. The first prompt message indicates that the second terminal's voice function should be disabled because it is too close to another user's terminal. The watermark text is obtained by splicing the session identifier of the target session and the device identifier of the first terminal in the data packet sent by the server. The second audio data is obtained by the first terminal by adding an audio watermark to the first audio data based on the watermark text. The first audio data is obtained from the server. The server is configured to: receive watermark detection results and audio data sent by the second terminal; obtain a session identifier from the watermark detection results; in response to the session identifier in the watermark detection results being the same as the session identifier of the target session, determine that there is a target terminal in the target session that is in the same space as the second terminal; based on the device identifier in the watermark detection results, determine the target terminal in the same space as the second terminal, and forward the audio data to other participating terminals in the target session other than the second terminal and the target terminal for playback; send a second prompt message to the target terminal and a third prompt message to the administrator terminal, wherein the second prompt message is used to indicate that the target terminal and the second terminal are in the same space. When both the target terminal and the second terminal are in the same space, it is recommended that the user of the target terminal connect headphones for the conversation. The third prompt message indicates that since the target terminal and the second terminal are in the same space, it is recommended to turn off the voice function of the target terminal or the second terminal. When the speakers of both the target terminal and the second terminal are on, the server is further configured to: select the terminal with the highest priority from the target terminal and the second terminal as the data receiving terminal based on the device type; when the priorities of all terminals are the same, prompt the user to specify a data receiving terminal, or allow the user to set the data receiving priority of each terminal; and send third audio data from a terminal in the target session that is in a different space from the target terminal and the second terminal to the data receiving terminal.
16. The apparatus according to claim 15, characterized in that, The device further includes: The data processing module is used to process the second audio data based on the watermark detection result; The sending module is used to send the watermark detection result and the processed second audio data to the server, and the server is used to forward the processed second audio data based on the watermark detection result.
17. The apparatus according to claim 16, characterized in that, The data processing module is used to perform any of the following: The audio energy of the second audio data is attenuated. Echo cancellation is performed on the second audio data based on the watermark detection results; The second audio data is denoised based on the watermark detection result; The second audio data is muted.
18. The apparatus according to claim 15, characterized in that, The detection module is used to perform any of the following: Obtain the cepstral spectrum of the second audio data, and determine the positions in the cepstral spectrum where the peak value is greater than the first threshold as the watermark loading positions; Perform a discrete cosine transform on the second audio data to obtain the energy intensity corresponding to each position of the second audio data, and determine the positions where the energy intensity is greater than the second threshold as the watermark loading positions.
19. An audio playback device based on group conversation, characterized in that, The device includes: The receiving module is used to receive watermark detection results and audio data sent by a second terminal, which is a terminal participating in the target session; the audio data is obtained based on the second audio data collected by the second terminal; the second audio data is obtained by the first terminal adding an audio watermark to the first audio data based on the watermark text, and the watermark text is obtained by the first terminal splicing the session identifier of the target session and the device identifier of the first terminal in the data packet sent by the server; The determination module is used to obtain the session identifier in the watermark detection result. In response to the fact that the session identifier in the watermark detection result is the same as the session identifier of the target session, it is determined that there is a target terminal in the same space as the second terminal among the participating terminals of the target session; based on the device identifier in the watermark detection result, the target terminal in the same space as the second terminal is determined. The forwarding module is used to forward the audio data to other participating terminals in the target session, excluding the second terminal and the target terminal, for playback of the audio data; The sending module is configured to send a second prompt message to the target terminal and a third prompt message to the administrator terminal. The second prompt message indicates that, since the target terminal and the second terminal are in the same space, it is recommended that the user of the target terminal connect a headset for conversation. The third prompt message indicates that, since the target terminal and the second terminal are in the same space, it is recommended that the voice function of either the target terminal or the second terminal be turned off. The module for performing the following steps: when both the speaker of the target terminal and the speaker of the second terminal are turned on, selects the terminal with the highest priority from the target terminal and the second terminal as the data receiving terminal based on the device type; when the priorities of each terminal are the same, prompts the user to specify a data receiving terminal, or allows the user to set the data receiving priority of each terminal; and sends third audio data from a terminal in the target session that is in a different space from the target terminal and the second terminal to the data receiving terminal. The second terminal is configured to: determine at least one watermark loading position in the second audio data; based on the at least one watermark loading position, perform watermark demodulation on the second audio data to obtain a watermark sequence; perform channel decoding and source decoding on the watermark sequence to obtain watermark text; compare the session identifier in the watermark text with the session identifier sent by the server; if the two session identifiers are the same, it is determined that among the terminals participating in the target session, there is a terminal in the same space as the second terminal; based on the device identifier in the watermark text, display a first prompt message on the session interface of the target session, the first prompt message being used to indicate that the second terminal's voice function is recommended to be turned off because the second terminal is close to which user's terminal.
20. The apparatus according to claim 19, characterized in that, The device further includes a removal module for: Based on the watermark detection result, the mixing paths of the target terminal and the second terminal in the mixing topology are removed. The mixing topology includes the mixing paths between each terminal in the target session.
21. A computer device, characterized in that, The computer device includes one or more processors and one or more memories, wherein at least one piece of program code is stored in the one or more memories, and the at least one piece of program code is loaded and executed by the one or more processors to perform the operations performed by the group session-based audio playback method as described in any one of claims 1-4 or 9-10; or the operations performed by the group session-based device management method as described in any one of claims 5-8.
22. A computer-readable storage medium storing at least one piece of program code, the at least one piece of program code being loaded and executed by a processor to perform the operations performed by the group session-based audio playback method or device management method as described in any one of claims 1-10.
23. A computer program product comprising at least one piece of program code stored in a computer-readable storage medium, wherein a processor of a computer device reads the at least one piece of program code from the computer-readable storage medium, and the processor executes the at least one piece of program code to cause the computer device to perform the operations performed by the group session-based audio playback method or device management method as described in any one of claims 1-10.
Citation Information
Patent Citations
Mobile terminal and television interaction method and system based on interactive audio watermarks
CN108712666A
digital teleconferencing system, subscriber access and switching device, working method and computer program product
DE102014116610A1
Audio apparatus, audio distribution system and method of operation therefor
EP3594802A1
Systems and methods for mitigating and / or avoiding feedback loops during communication sessions
US20170346950A1
Potential echo detection and warning for online meeting
US20180077205A1