An echo delay determination method, apparatus, device and storage medium

By embedding watermark information into the audio signal to analyze the echo delay, the problems of low accuracy and high computational resource consumption in the existing technology for determining echo delay are solved, and efficient and accurate echo delay determination is achieved under the condition of multiple audio signals mixed together.

CN113707160BActive Publication Date: 2026-03-31TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-05
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in determining echo delay, especially when multiple audio signals are mixed, and consume a lot of computational resources, with poor compatibility and versatility.

Method used

Audio watermarking technology is used to embed watermark information that is inaudible to the human ear into the reference audio signal to be played. The watermark information is analyzed by playing and collecting near-end audio signals to determine the echo delay.

Benefits of technology

It can accurately determine echo delay in various scenarios without requiring a large amount of computing resources, and has good compatibility and versatility, making it suitable for different hardware devices and software applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113707160B_ABST
    Figure CN113707160B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an echo delay determination method and device, equipment and a storage medium, wherein the method comprises: embedding watermark information in a reference audio signal to be played to obtain a target audio signal; playing the target audio signal; collecting a near-end audio signal; performing watermark information analysis processing on the near-end audio signal; in a case where the watermark information is analyzed from the near-end audio signal through the watermark information analysis processing, determining an echo delay according to a position of the watermark information in the target audio signal and a position of the watermark information in the near-end audio signal. The method can accurately determine the echo delay, thereby helping to improve the echo cancellation effect, and by designing the embedding structure of the watermark information, the transmission quality of the watermark information embedded in the audio signal under strong attacks can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, device and storage medium for determining echo delay. Background Technology

[0002] Echoes are currently more commonly found in real-time communication (RTC) scenarios. Figure 1 This is a schematic diagram illustrating the principle of echo generation in an RTC scenario; for example... Figure 1 As shown, when user A and user B are having a real-time voice call, user A's voice 'a' is captured by terminal device 110 and transmitted to terminal device 120 via the network. Upon receiving voice 'a', terminal device 120 plays it back. Simultaneously, it captures user B's voice 'b' and its own played voice 'a', and transmits both voice 'b' and voice 'a' back to terminal device 110 for playback. At this point, user A will hear their own previously emitted voice 'a' – this is the echo. The presence of the echo interferes with the quality of voice communication and reduces the intelligibility of the voice.

[0003] To avoid echo interference in voice call quality, Acoustic Echo Cancellation (AEC) technology was developed. When a terminal device eliminates echo using AEC technology, it treats the audio signal sent by the other device that needs to be played as the far-end audio signal, and the audio signal it collects that needs to be sent to the other device as the near-end audio signal. The echo in the near-end audio signal is then filtered out based on the far-end audio signal.

[0004] Since echo formation typically involves three stages—audio signal playback, air propagation, and audio signal acquisition—the echo in the near-end audio signal lags behind the far-end audio signal. This lag is called echo delay. When terminal devices use echo cancellation technology to eliminate echoes in near-end audio signals, they usually first use the echo delay to align the near-end and far-end audio signals, and then eliminate the echo in the near-end audio signal based on the far-end audio signal. Therefore, determining echo delay as a pre-processing technique for echo cancellation, and the accuracy of the determined echo delay, significantly impacts the effectiveness of echo cancellation. Summary of the Invention

[0005] This application provides an echo delay determination method, apparatus, device, and storage medium that can accurately determine echo delay, thereby helping to improve echo cancellation effect.

[0006] In view of this, the first aspect of this application provides a method for determining echo delay, the method comprising:

[0007] Watermark information is embedded in the reference audio signal to be played to obtain the target audio signal;

[0008] Play the target audio signal; and acquire the near-end audio signal;

[0009] The near-end audio signal is processed by watermark information parsing;

[0010] When the watermark information is parsed from the near-end audio signal through the watermark information parsing process, the echo delay is determined based on the position of the watermark information in the target audio signal and the position of the watermark information in the near-end audio signal.

[0011] A second aspect of this application provides an echo delay determination apparatus, the apparatus comprising:

[0012] The watermark embedding module is used to embed watermark information into the reference audio signal to be played, so as to obtain the target audio signal.

[0013] An audio playback module is used to play the target audio signal;

[0014] The audio acquisition module is used to acquire near-end audio signals;

[0015] The watermark parsing module is used to parse and process the watermark information of the near-end audio signal;

[0016] The echo delay determination module is used to determine the echo delay based on the position of the watermark information in the target audio signal and the position of the watermark information in the near-end audio signal after parsing the watermark information from the near-end audio signal through the watermark information parsing process.

[0017] A third aspect of this application provides an electronic device, the electronic device including a processor and a memory:

[0018] The memory is used to store computer programs;

[0019] The processor is configured to perform the steps of the echo delay determination method as described in the first aspect above, according to the computer program.

[0020] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program for performing the steps of the echo delay determination method described in the first aspect.

[0021] A fifth aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the echo delay determination method described in the first aspect.

[0022] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0023] This application provides a method for determining echo delay, which innovatively applies audio watermarking technology to determine echo delay. Specifically, an audible watermark is first embedded into a reference audio signal to be played, resulting in a target audio signal. Then, the target audio signal is played, and a near-end audio signal is acquired. Next, the acquired near-end audio signal is processed to parse the watermark information. After retrieving the previously embedded watermark information from the near-end audio signal through the watermark information parsing process, the time lag of the echo in the near-end audio signal relative to the target audio signal is determined based on the position of the watermark information in the target audio signal and the position of the watermark information in the near-end audio signal; that is, the echo delay is determined. On one hand, determining echo delay based on watermark information in the near-end audio signal does not place particular requirements on the signal-to-noise ratio of the near-end audio signal. Therefore, even when the near-end audio signal contains multiple audio signals, the method provided in this application can accurately determine the echo delay. On the other hand, for the device executing the method provided in this application, it can accurately determine the echo delay without consuming a large amount of computing resources. Furthermore, the method provided in this application has good compatibility and versatility with different hardware devices and software applications. That is, the echo delay can be accurately determined by the method provided in this application for different hardware devices and software applications. Attached Figure Description

[0024] Figure 1 This is a schematic diagram illustrating the principle of echo generation in an RTC scenario.

[0025] Figure 2 This is a schematic diagram illustrating the working principle of an echo cancellation module in communication software.

[0026] Figure 3 A schematic diagram illustrating the principle of aligning far-end and near-end audio signals;

[0027] Figure 4 A schematic diagram illustrating the implementation principle of the echo cancellation module for eliminating echoes;

[0028] Figure 5A schematic diagram illustrating an application scenario of the echo delay determination method provided in this application embodiment;

[0029] Figure 6 A flowchart illustrating the echo delay determination method provided in this application embodiment;

[0030] Figure 7 A schematic diagram illustrating the process of generating a target audio signal provided in an embodiment of this application;

[0031] Figure 8 A schematic diagram of the frame structure of the watermark source coding frame provided in the embodiments of this application;

[0032] Figure 9 This is a schematic diagram of the frame structure of a channel-coded frame provided in an embodiment of this application;

[0033] Figure 10 This is a schematic diagram of the watermark information parsing and processing provided in the embodiments of this application;

[0034] Figure 11 A schematic diagram illustrating the implementation principle of watermark information injection at the playback end provided in this application embodiment;

[0035] Figure 12 A schematic diagram illustrating the implementation principle of watermark information parsing at the recording end provided in this application embodiment;

[0036] Figure 13 A schematic diagram of the structure of the first echo delay determination device provided in the embodiments of this application;

[0037] Figure 14 This is a schematic diagram of the structure of the second echo delay determination device provided in the embodiments of this application;

[0038] Figure 15 A schematic diagram of the structure of the third echo delay determination device provided in the embodiments of this application;

[0039] Figure 16 A schematic diagram of the structure of the fourth echo delay determination device provided in the embodiments of this application;

[0040] Figure 17 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0041] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0042] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0043] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0044] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0045] Key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Voiceprint Recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech emerging as one of the most promising methods.

[0046] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0047] The solution provided in this application relates to artificial intelligence voice technology, which is specifically illustrated through the following embodiments:

[0048] To eliminate echoes in the acquired near-end audio signals, hardware devices and software applications with real-time communication capabilities typically include echo cancellation modules. Figure 2 This is a schematic diagram illustrating the working principle of an echo cancellation module in a certain communication software; for example... Figure 2 As shown, when a user makes a real-time voice call through the communication software, the far-end audio signal received by the communication software is played out through the broadcast module framework. After the far-end audio signal is played out, the communication software can collect the played far-end audio signal when collecting the near-end audio signal through the recording module framework. The collected far-end audio signal is the echo. Then, the echo cancellation module in the communication software will use the echo cancellation kernel to cancel the echo in the near-end audio signal based on the far-end audio signal previously received by the communication software, and then send out the echo-free near-end audio signal.

[0049] from Figure 2 As shown in the schematic diagram of the working principle of the echo cancellation module, the echo formation of the far-end audio signal needs to go through three stages: the broadcast module framework (including software channels and hardware channels), sound propagation through the air, and the recording module framework (including software channels and hardware channels). Based on this, the echo in the near-end audio signal will lag behind the far-end audio signal. This lag is called the echo delay, and the expression for the echo delay is shown in Equation (1):

[0050] echo delay=playback delay+broadcast delay+record delay (1)

[0051] Playback delay is the time from when the remote audio signal is recovered to when it is played by an audio playback device (such as a speaker). Playback delay varies greatly depending on the operating system. Typically, the playback delay for Android is 100-300ms, while for iOS it is 50-80ms.

[0052] Broadcast delay is the time it takes for a distant audio signal to travel through the air from an audio playback device (such as a speaker) to an audio acquisition device (such as a microphone). This delay is related to the length of the physical path the audio signal travels. Since the distance between the audio playback device and the audio acquisition device of the terminal device is generally short, broadcast delay can usually be ignored.

[0053] The record delay is the time from when the audio acquisition device acquires the far-end audio signal to when the echo cancellation module obtains the near-end audio signal including the far-end audio signal, which is usually about 10ms.

[0054] Echo delay estimation (also known as echo delay estimator) is used as a preprocessing technique for echo cancellation to determine the time difference between the echo in the near-end audio signal and the far-end audio signal. In practice, the echo cancellation module needs to align the far-end and near-end audio signals based on the determined echo delay. Figure 3 As shown; furthermore, the near-end audio signal is subjected to adaptive filtering and nonlinear processing sequentially based on the far-end audio signal to filter out echoes in the near-end audio signal, such as... Figure 4 As shown. In many cases, the ability to accurately determine the echo delay can significantly impact the performance of the echo cancellation module, that is, the effectiveness of echo cancellation.

[0055] In related technologies, echo delay is currently determined mainly through the following three methods:

[0056] The first implementation method first transforms the near-end and far-end audio signals to the frequency domain, obtaining the near-end spectrum and the far-end spectrum respectively. Then, it performs binary processing on the near-end and far-end spectra respectively, obtaining the near-end binary spectrum and the far-end binary spectrum. Finally, it estimates the echo delay by comparing the near-end binary spectrum and the far-end binary spectrum. This implementation method requires high signal-to-noise ratio of the near-end audio signal because it uses spectral energy binarization. Furthermore, it performs poorly in "two-way" scenarios (i.e., in a voice acquisition environment where multiple audio signals are mixed, resulting in multiple audio signals being mixed in the acquired near-end audio signal), and the accuracy of the estimated echo delay is low.

[0057] The second implementation method determines the echo delay based on the generalized cross-correlation function. Its basic principle is to obtain the cross-power spectrum between the near-end and far-end audio signals, then perform weighted operations in the frequency domain with different weights, and finally inversely transform to the time domain to obtain the cross-correlation function between the near-end and far-end audio signals. The time corresponding to the extreme value of this cross-correlation function is the echo delay. This implementation method outperforms the first method; however, determining the echo delay using this method requires a large number of domain transformation and cross-correlation operations, resulting in a significant computational burden and high computational requirements for the terminal device.

[0058] The third approach involves determining echo delay through machine learning. This involves training a specific model for each terminal device to obtain a neural network model that determines the echo delay for that device. While this approach can determine echo delay relatively accurately, the same model often cannot accurately determine the echo delay for multiple devices due to differences in hardware performance and other aspects. Therefore, it requires training a dedicated model for each terminal device, resulting in poor model versatility and compatibility.

[0059] To address the problems existing in the aforementioned related technologies, this application provides an echo delay determination method. This method can accurately determine the echo delay in various scenarios without consuming a large amount of computing resources from the terminal device, and has good versatility and compatibility.

[0060] Specifically, in the echo delay determination method provided in this application embodiment, firstly, watermark information that is inaudible to the human ear is embedded in the reference audio signal to be played to obtain the target audio signal; then, the target audio signal is played, and a near-end audio signal is acquired; furthermore, the acquired near-end audio signal is subjected to watermark information parsing processing; if the previously embedded watermark information is parsed from the near-end audio signal through the above watermark information parsing processing, the time lag of the echo in the near-end audio signal relative to the target audio signal is determined according to the position of the watermark information in the target audio signal and the position of the watermark information in the near-end audio signal, that is, the echo delay is determined.

[0061] The aforementioned echo delay determination method applies audio watermarking technology to determine echo delay. Based on the auditory masking mechanism of the human ear, the watermark information is embedded in the reference audio signal to obtain the target audio signal without affecting the audio playback quality and without being perceived by the human ear. Since the process from audio playback to echo acquisition is a closed loop, the near-end audio signal acquired when playing the target audio signal should also include the watermark information. Therefore, the echo delay can be determined based on the position of the watermark information in the target audio signal and the near-end audio signal. On the one hand, determining the echo delay based on the watermark information in the near-end audio signal does not have special requirements for the signal-to-noise ratio of the near-end audio signal. Therefore, even when the near-end audio signal is mixed with multiple audio signals, the method provided in this application embodiment can accurately determine the echo delay. On the other hand, for the device used to execute the method provided in this application embodiment, it can accurately determine the echo delay without consuming a large amount of computing resources. Furthermore, the method provided in this application has good compatibility and versatility with different hardware devices and software applications. That is, the echo delay can be accurately determined by the method provided in this application for different hardware devices and software applications.

[0062] It should be understood that, in practical applications, the echo delay determination method provided in this application embodiment can be applied to terminal devices, such as smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, etc., but is not limited to these.

[0063] To facilitate understanding of the echo delay determination method provided in the embodiments of this application, the application scenarios of the echo delay determination method provided in the embodiments of this application will be introduced by way of example below.

[0064] See Figure 5 , Figure 5 This is a schematic diagram illustrating an application scenario for the echo delay determination method provided in this application embodiment. For example... Figure 5 As shown, this application scenario includes terminal device 510 and terminal device 520. Both terminal devices 510 and 520 run target communication software and can communicate via a network. Both terminal devices 510 and 520 can be used to execute the echo delay determination method provided in this application embodiment. The following description uses terminal device 510 executing the echo delay determination method as an example.

[0065] In practical applications, user A using terminal device 510 and user B using terminal device 520 can conduct real-time voice calls through target communication software. During the real-time voice call, terminal device 510 receives an audio signal sent by terminal device 520 via the network and uses this audio signal as a reference audio signal to be played. Then, it embeds an inaudible watermark information into the reference audio signal to obtain the target audio signal. This watermark information can be pre-set. Subsequently, terminal device 510 can play the target audio signal.

[0066] During a real-time voice call, the terminal device 510 continuously collects audio signals from its surrounding environment; these collected audio signals are known as near-end audio signals. Furthermore, the terminal device 510 performs watermark information parsing on the collected near-end audio signals. If the watermark information embedded in the target audio signal is extracted from the near-end audio signal, the terminal device can determine the echo delay based on the position of the watermark information in both the target and near-end audio signals.

[0067] Furthermore, terminal device 510 can filter out the echo in the near-end audio signal based on the determined echo delay, and send the echo-filtered audio signal to terminal device 520 via the network. Similarly, terminal device 520 will also perform the above operations during a real-time voice call.

[0068] It should be noted that, Figure 5 The application scenarios shown are merely examples. In practical applications, the echo delay determination method provided in this application can be applied not only to two-person real-time voice calls but also to multi-person real-time voice calls. Furthermore, the echo delay determination method provided in this application can be applied not only to real-time voice calls via communication software but also to real-time voice calls based on hardware devices, such as making phone calls. Additionally, the echo delay determination method provided in this application can also be applied to real-time video scenarios. No limitations are made here regarding the application scenarios of the echo delay determination method provided in this application.

[0069] It should be noted that echo is also a problem that must be addressed in scenarios where users interact with smart devices (such as smart speakers, smart voice assistants in terminal devices, and in-vehicle voice recognition devices). Specifically, in many cases, the audio signal collected by a smart device may simultaneously include the user's voice signal and the voice signal previously emitted by the smart device itself (i.e., echo). If the smart device directly processes the audio signal it has collected, it is highly susceptible to a series of problems such as misrecognition, misunderstanding, and misresponse, resulting in a poor user experience. To avoid this situation, such smart devices often need to perform echo cancellation processing on the audio signal they collect. Accordingly, before performing echo cancellation processing, such smart devices can determine the echo delay using the echo delay determination method provided in the embodiments of this application.

[0070] The echo delay determination method provided in this application will be described in detail below through method embodiments.

[0071] See Figure 6 , Figure 6 This is a flowchart illustrating the echo delay determination method provided in an embodiment of this application. The following embodiments describe the method using a terminal device as the executing entity. Figure 6 As shown, the echo delay determination method includes the following steps:

[0072] Step 601: Embed watermark information into the reference audio signal to be played to obtain the target audio signal.

[0073] In application scenarios where echo cancellation is required, before the terminal device plays the reference audio signal, it needs to use audio watermarking technology to embed watermark information that is inaudible to the human ear into the reference audio signal to be played, thereby obtaining the corresponding target audio signal.

[0074] It should be noted that audio watermarking technology utilizes an auditory masking mechanism to embed watermark information, which is inaudible to the human ear, into the audio stream to be played. This embedded watermark information can be identified and authenticated at the decoding end. Currently, the main uses of audio watermarking technology include: protecting the copyright of audio works, preventing the unauthorized recording of live streams, and tracing the source of leaked information from recorded online meetings.

[0075] It should be noted that the aforementioned reference audio signal can be generated in different ways in different application scenarios. For example, in real-time communication applications, the reference audio signal should be the audio signal sent by the other party's terminal device. For instance, in a real-time voice call between user A and user B, the audio signal sent by user B's terminal device is the reference audio signal for user A's terminal device, and vice versa. It should be understood that real-time communication applications include, but are not limited to, two-person or multi-person real-time voice calls and two-person or multi-person real-time video calls. For example, in a dialogue between a user and a terminal device, the reference audio signal should be the audio signal generated by the terminal device in response to the user's dialogue. For instance, in a dialogue between a user and a smart speaker, after receiving the user's voice signal, the smart speaker can generate an audio signal to respond to that voice signal; this audio signal is the reference audio signal. Of course, in other application scenarios, the reference audio signal can also be an audio signal generated in other ways. No limitation is made here on the application scenarios of the embodiments of this application, nor on the way the reference audio signal is generated.

[0076] It should be noted that the watermark information embedded in the reference audio signal can be pre-set. For example, the watermark information can be pre-set text information or binary code. This application does not impose any limitations on the watermark information embedded in the reference audio signal.

[0077] In one possible implementation, when the terminal device embeds watermark information in the reference audio signal, it can... Figure 7 The process shown is implemented. Figure 7 This is a schematic diagram illustrating the process of generating a target audio signal according to an embodiment of this application. Figure 7 As shown, the process of generating the target audio signal includes the following steps:

[0078] Step 701: Perform source encoding on the watermark information to obtain a watermark source encoded frame.

[0079] Before embedding watermark information into the reference audio signal, the terminal device needs to perform source encoding on the pre-set watermark information (such as pre-set text information or binary code) to obtain the corresponding watermark source encoded frame, so as to realize the embedding of watermark information into the reference audio signal.

[0080] In practice, the terminal device can segment the watermark information into multiple sub-watermark information by a preset byte length. Then, for each sub-watermark information, source encoding is performed to obtain the watermark source encoding frame corresponding to the sub-watermark information. The watermark source encoding frame includes the byte length of the watermark information, the sequence number of the sub-watermark information in the watermark information, the sub-watermark information itself, and the check code.

[0081] For example, the terminal device can segment the watermark information into several sub-watermark information units, each byte in size. Then, according to a pre-defined watermark source coding frame structure, source coding is performed on each sub-watermark information to obtain the corresponding watermark source coding frame; using the watermark source coding frame structure as... Figure 8 Taking the frame structure shown as an example, when a terminal device constructs a corresponding watermark source encoding frame for a certain sub-watermark information, it can add the byte length of the watermark information (i.e., watermark length) and the sequence number of the sub-watermark information in the watermark information (i.e., byte number) to the frame header of the watermark source encoding frame, add the content of the sub-watermark information (i.e., byte content) to the frame body of the watermark source encoding frame, and add the checksum to the frame tail of the watermark source encoding frame. Assuming the default length of the watermark source encoding frame is 32 bits, the byte length of the watermark information can occupy 4 bits, the sequence number of the sub-watermark information in the watermark information can occupy 4 bits, the content of the sub-watermark information can occupy 8 bits, and the checksum can occupy 16 bits. Of course, in practical applications, the watermark source encoding frame structure can also be other forms of structure, and this application does not impose any limitations on the watermark source encoding frame structure.

[0082] It should be noted that the checksum in the aforementioned watermark source-encoded frame can be a Cyclic Redundancy Check (CRC) code, which is a checksum with error detection and correction capabilities. Alternatively, the checksum in the aforementioned watermark source-encoded frame can be a checksum generated through block check. This application does not impose any limitations on the checksum in the watermark source-encoded frame.

[0083] Step 702: Detect the target position in the reference audio signal for embedding watermark information.

[0084] In addition, before embedding the watermark information into the reference audio signal, the terminal device also needs to detect the target position in the reference audio signal that can be used to embed the watermark information, that is, to detect the target position that can be used to embed the watermark source coding frame.

[0085] In practice, the terminal device can detect the energy spectrum envelope of the reference audio signal, and then determine the position in the reference audio signal where the energy spectrum envelope exceeds a preset energy threshold as the target position that can be used to embed watermark information, and mark the watermark loading enable flag for the target position.

[0086] For example, after acquiring a reference audio signal, the terminal device can detect its energy spectrum envelope, which characterizes the energy level at various locations within the reference audio signal. Furthermore, based on the energy spectrum envelope, it detects locations where the energy spectrum envelope exceeds a preset energy threshold, i.e., locations with higher energy in the reference audio signal. These higher-energy locations are then identified as target locations for embedding watermark information, and a watermark loading enable flag is marked at these target locations. This prevents watermark information from being embedded into silent or low-energy audio segments of the reference audio signal, thus avoiding the loss of valid information at the audio signal decoding end.

[0087] It should be understood that in practical applications, the terminal device may execute step 701 first and then step 702, or step 702 first and then step 701, or both steps 701 and 702 may be executed simultaneously. This application does not impose any restrictions on the execution order of steps 701 and 702.

[0088] Step 703: Based on the target position in the reference audio signal, perform channel coding on the reference audio signal and the watermark source coding frame to obtain the target audio signal.

[0089] After the terminal device completes the source coding of the watermark information to obtain the watermark source coding frame, and detects the target position in the reference audio signal that can be used to embed the watermark information, it can perform channel coding on the reference audio signal and the watermark source coding frame based on the target position in the reference audio signal, so as to embed the watermark information into the reference audio signal and thus obtain the target audio signal.

[0090] In specific implementation, during the channel coding of the reference audio signal, the terminal device can determine whether the current coding position of the reference audio signal is a target position that can be used to embed watermark information. If the current coding position is a target position, the terminal device can embed a watermark source coding frame into the audio signal at the current coding position in the reference audio signal using a watermark modulation algorithm to obtain a first signal to be encoded; then, channel coding is performed on the first signal to be encoded to obtain the channel-coded frame corresponding to the current coding position. If the current coding position is not a target position, the terminal device can directly use the audio signal at the current coding position in the reference audio signal as a second signal to be encoded, and perform channel coding on the second signal to be encoded to obtain the channel-coded frame corresponding to the current coding position. Finally, the channel-coded frames corresponding to each coding position in the reference audio signal are combined to obtain the target audio signal.

[0091] For example, when a terminal device performs channel coding on a reference audio signal, it can detect whether the current coding position of the reference audio signal is marked with a watermark loading enable flag. If the current coding position is marked with a watermark loading enable flag, it can be determined that the current coding position is a target position that can be used to embed watermark information. Then, using a watermark modulation algorithm, the watermark source-coded frame previously obtained through source coding is added to the audio signal at the current coding position to obtain the first signal to be encoded. Then, according to a pre-set channel coding frame structure, the first signal to be encoded is channel-coded to obtain the channel-coded frame corresponding to the current coding position. Conversely, if the current coding position is not marked with a watermark loading enable flag, it can be determined that the current coding position is not a target position that can be used to embed watermark information. Then, the audio signal at the current coding position is directly used as the second information to be encoded, and according to a pre-set channel coding frame structure, the second signal to be encoded is channel-coded to obtain the channel-coded frame corresponding to the current coding position. Thus, by performing the above channel coding process on each coding position in the reference audio signal, the channel coded frames corresponding to each coding position are obtained. According to the arrangement order of each coding position in the reference audio signal, the channel coded frames corresponding to each coding position are combined accordingly to obtain the target audio signal.

[0092] It should be noted that the selection of the watermark modulation algorithm can be based on the actual needs of the scenario. As an example, this application embodiment can employ a robust, low-sound-quality-loss, and low-complexity time-domain bidirectional multi-core echo-hidden watermark modulation algorithm. The main principle of this watermark modulation algorithm is to utilize the time-series masking mechanism of human hearing to modulate the watermark information into early reflections that are indistinguishable to the human ear. Furthermore, the use of bidirectional echoes can resist interference caused by spatial multipath reflections, and the use of multi-cores can enhance the data transmission rate. Of course, in practical applications, terminal devices can also use other watermark modulation algorithms to add the watermark source encoded frame to the reference audio signal. This application does not impose any limitations on the watermark modulation algorithm used.

[0093] It should be noted that since the echo delay typically does not change significantly over a short period, in practical applications, the terminal device does not need to embed a watermark source-coded frame in the audio signal at every target location in the reference audio signal. For example, the terminal device can embed a watermark source-coded frame in the audio signal at a specific target location within the current reference audio signal at regular intervals, such as 1 minute or 30 seconds. Of course, to ensure the accuracy of the determined echo delay, the terminal device can also embed a watermark source-coded frame in the audio signal at every target location in the reference audio signal. This application does not impose any limitations on the specific method of embedding the watermark source-coded frame.

[0094] Furthermore, to improve the recognition rate of the watermark information parsing end and its robustness to different transmission channels, this application also proposes a channel-coded frame structure. Figure 9 The diagram shown illustrates the structure of the channel-coded frame. Figure 9 As shown, the header of the channel-coded frame carries the synchronization code; the body of the channel-coded frame carries data packets. When the channel-coded frame corresponds to a target location in the reference audio signal that can be used to embed watermark information, the data packet should include an audio signal with an embedded watermark source-coded frame, i.e., the first signal to be encoded. When the channel-coded frame corresponds to a location in the reference audio signal that cannot be used to embed watermark information, the data packet should include an audio signal without an embedded watermark source-coded frame, i.e., the second signal to be encoded. The body of the channel-coded frame carries error correction codes, which are generated based on the content carried by the header and body of the channel-coded frame.

[0095] For example, when a terminal device performs channel coding, it can add a synchronization code to the header of the channel-coded frame. This synchronization code can be a fixed string of codewords used for frame synchronization, and its specific length and content can be adjusted according to the actual channel conditions. Data packets can be added to the body of the channel-coded frame. If the channel-coded frame corresponds to a target location, the data packet can include the audio signal at that coding location in the reference audio signal and the watermark source coding frame; if the channel-coded frame does not correspond to a target location, the data packet can include the audio signal at that coding location in the reference audio signal. Error correction codes can be added to the end of the channel-coded frame. Setting these error correction codes can reduce the bit error rate at the decoding end and ensure signal transmission quality when the channel signal-to-noise ratio is poor. For example, the error correction code in the channel-coded frame can be a BCH error correction code. When the terminal device generates the BCH error correction code in the channel-coded frame, it can divide the information carried in the frame header and frame body of the channel-coded frame into several message groups according to a preset number of bits, and then transform each message group into a binary number group of a specific length, i.e., a codeword. The codewords corresponding to each message group can then form the BCH error correction code. Of course, in practical applications, the terminal device can also set other types of error correction codes in the channel-coded frame. This application does not impose any limitations on the type of error correction code in the channel-coded frame.

[0096] It should be understood that Figure 7 The implementation method for generating the target audio signal shown is only an example. In practical applications, the terminal device can also use other methods to embed the watermark information into the reference audio signal to obtain the target audio signal. This application does not limit the method of generating the target audio signal in any way.

[0097] Step 602: Play the target audio signal; and acquire the near-end audio signal.

[0098] The terminal device embeds watermark information into the reference audio signal to be played, obtains the target audio signal, and then plays the target audio signal. In application scenarios where echo cancellation is required, the playback and acquisition of the audio signal are usually performed simultaneously. Therefore, while playing the target audio signal, the terminal device also acquires the audio signal, which is the near-end audio signal.

[0099] For example, in real-time communication applications, the near-end audio signal is the audio signal collected by the terminal device itself. For instance, in a real-time voice call between user A and user B, the near-end audio signal is the audio signal collected by user A's terminal device from the environment in which user A is located, and vice versa. Similarly, in applications where a user converses with a terminal device, the near-end audio signal is also the audio signal collected by the terminal device itself. For example, in a scenario where a user converses with a smart speaker, the audio signal collected by the smart speaker in its environment is the near-end audio signal.

[0100] It should be understood that the near-end audio signal can include any audio signal collected by the terminal device in its environment. That is, the near-end audio signal can include not only the voice signal emitted by the user, but also the target audio signal played by the terminal device itself, and also the noise audio signal in the environment. This application does not limit the audio signals included in the near-end audio signal in any way.

[0101] Step 603: Perform watermark information parsing processing on the near-end audio signal.

[0102] After the terminal device collects the near-end audio signal, it needs to perform watermark information parsing processing on the near-end audio signal to determine whether the collected near-end audio signal includes the watermark information previously embedded in the reference audio signal.

[0103] In one possible implementation, when the terminal device performs watermark information parsing processing on the near-end audio signal, it can... Figure 10 The process shown is implemented. Figure 10 This is a schematic diagram illustrating the watermark information parsing process provided in an embodiment of this application. Figure 10 As shown, the watermark information parsing and processing flow includes the following steps:

[0104] Step 1001: Demodulate the near-end audio signal to obtain a binary bit stream.

[0105] After the terminal device acquires the near-end audio signal, in order to ensure that the near-end audio signal can be correctly demodulated and then analyzed to see if it contains watermark information, the terminal device needs to demodulate the audio signal of the corresponding frame length in the near-end audio signal according to the frame length of the channel-coded frame in the target audio signal, so as to obtain the corresponding binary bit stream.

[0106] Step 1002: Perform channel decoding on the binary bit stream to obtain a channel decoded stream.

[0107] After the terminal device demodulates the near-end audio signal to obtain a binary bitstream, it can further perform channel decoding on the binary bitstream to obtain a channel decoded stream.

[0108] In specific implementation, if the channel-coded frame of the target audio signal previously generated by the terminal device includes a synchronization code and an error correction code, the terminal device can first perform frame synchronization based on the synchronization code in the binary bit stream; then, during the channel decoding process of the binary bit stream, the error in the channel decoded stream is corrected based on the error correction code in the binary bit stream; if the number of corrected error bits does not exceed the preset number of error bits, the watermark demodulation algorithm is continued to perform watermark demodulation processing on the channel decoded stream; if the number of corrected error bits exceeds the preset number of error bits, the channel decoded stream is discarded.

[0109] For example, when a terminal device performs channel decoding on the binary bitstream obtained from demodulating the near-end audio signal, it can first perform frame synchronization based on the synchronization code in the binary bitstream to separate the channel decoded streams corresponding to each channel-coded frame. Then, the error correction code in the binary bitstream is used to correct errors in the separated channel decoded streams. It should be noted that during the process from when the target audio signal is played to when it is re-acquired, the target audio signal may be subject to interference due to factors such as the audio signal propagation environment, leading to errors in the target audio signal. To solve this problem, the terminal device can use the error correction code in the binary bitstream to correct errors in the channel decoded stream. That is, the terminal device can use the reverse algorithm as when generating the error correction code to reconstruct the previously generated channel-coded frame based on the error correction code, and then correct the errors in the channel decoded stream based on the reconstructed channel-coded frame. If the number of bit errors corrected by the terminal device based on the error correction code does not exceed the preset number of bit errors, it means that the target audio signal is still within the error correction capability range of the error correction code, and the terminal device can continue to perform watermark demodulation processing based on the channel decode stream that has been corrected for bit errors; if the number of bit errors corrected by the terminal device based on the error correction code exceeds the preset number of bit errors, it means that the target audio signal has exceeded the error correction capability range of the error correction code, and the information carried in the channel decode stream may have been distorted, so the channel decode stream can be discarded.

[0110] Step 1003: Perform watermark demodulation processing on the channel decoded stream using a watermark demodulation algorithm.

[0111] After the terminal device completes the channel decoding process for the binary bit stream and obtains the channel decoded stream, it can further use a watermark demodulation algorithm to perform watermark demodulation processing on the channel decoded stream in order to determine whether the channel decoded stream carries a hidden coded bit stream, that is, to determine whether the channel decoded stream contains a watermarked source coded frame.

[0112] It should be noted that the watermark demodulation algorithm used by the terminal device here should correspond to the watermark modulation algorithm used when embedding the watermark information. For example, if the terminal device uses an echo-hidden modulation algorithm to embed the watermark information into the reference audio signal, then the terminal device needs to use the corresponding cepstral method to perform watermark demodulation processing on the channel coded stream. Of course, if other watermark modulation algorithms are used when embedding the watermark information into the reference audio signal, the terminal device can also use other watermark demodulation algorithms to perform watermark demodulation processing on the channel coded stream accordingly. This application does not impose any limitations on the watermark demodulation algorithm used.

[0113] Step 1004: After the watermark source coding frame is demodulated from the channel decode stream through the watermark demodulation process, the watermark source coding frame is source decoded to obtain the watermark information.

[0114] If the terminal device demodulates the watermark source coding frame from the channel decoding stream through step 1003, the terminal device can continue to perform source decoding on the watermark source coding frame to obtain the watermark information carried in the watermark source coding frame.

[0115] In practice, the terminal device can first verify the watermark source encoding frame according to the check code in the watermark source encoding frame; if the verification passes, the watermark information is obtained from the watermark source encoding frame; otherwise, if the verification fails, the watermark source encoding frame can be discarded.

[0116] For example, assuming the checksum in the watermark source-encoded frame previously embedded in the reference audio signal by the terminal device is a CRC checksum, when the terminal device verifies the watermark source-encoded frame in the channel decoder stream, it divides it by the polynomial used when generating the watermark source-encoded frame. If the remainder is 0, it indicates that the codewords in the watermark source-encoded frame are correct, and the watermark source-encoded frame passes the verification; in this case, the watermark information can be obtained from the watermark source-encoded frame. Conversely, if the remainder is not 0, it indicates that the codewords in the watermark source-encoded frame are incorrect, and the watermark source-encoded frame fails the verification; in this case, the watermark source-encoded frame can be discarded. It should be understood that if the checksum in the watermark source-encoded frame previously embedded in the reference audio signal by the terminal device is another checksum, the terminal device can also use other methods to verify the watermark source-encoded frame in the channel decoder stream. This application does not impose any limitations on the method of verifying the watermark source-encoded frame.

[0117] It should be understood that if the terminal device previously performed source coding on the watermark information to generate a watermark source coding frame, and segmented the watermark information, and generated a watermark source coding frame based on the segmented sub-watermark information, then the terminal device now performs source decoding on the watermark source coding frame in the channel decoding stream to obtain the watermark information, which is actually a sub-watermark information obtained by segmenting the watermark information.

[0118] It should be understood that Figure 10 The watermark information parsing and processing implementation shown is only an example. In practical applications, terminal devices can also use other methods to perform watermark information parsing and processing on near-end audio signals. This application does not limit the implementation method of watermark information parsing and processing.

[0119] Step 604: After parsing the watermark information from the near-end audio signal through the watermark information parsing process, determine the echo delay based on the position of the watermark information in the target audio signal and the position of the watermark information in the near-end audio signal.

[0120] If the terminal device parses the watermark information it previously embedded in the reference audio signal in step 601 from the near-end audio signal it collected through step 603, the terminal device can determine the time lag of the echo in the near-end audio signal (corresponding to the target audio signal) relative to the target audio signal based on the embedding position of the watermark information in the target audio signal and the position of the watermark information parsed in the near-end audio signal, that is, determine the echo delay.

[0121] In one possible implementation, the terminal device can determine the time point at which the audio frame containing the watermark information is played through the audio playback channel as the first time point; it can determine the time point at which the audio frame containing the watermark information is acquired through the audio acquisition channel as the second time point; and then calculate the time difference between the second time point and the first time point, which is the echo delay.

[0122] Specifically, when the terminal device plays the target audio signal, it can record the time point when the audio frame containing the watermark information is played through the audio playback channel as the first time point. For example, assuming that the watermark information 'a' is embedded in the fifth audio frame of the target audio signal (this audio frame can be understood as the channel-coded frame mentioned above), the terminal device can record the time point when the fifth audio frame is played through the audio playback channel as the first time point. For example, assuming that the terminal device plays the fifth audio frame through the audio playback channel at 9:44:35, then 9:44:35 will be used as the first time point. When a terminal device acquires near-end audio signals, it can record the time points of each audio frame in the near-end audio signal acquired through the audio acquisition channel. If the terminal device determines, through watermark information parsing, that the tenth audio frame in the near-end audio signal includes watermark information 'a', then the terminal device can determine the time point at which the tenth audio frame was acquired through the audio acquisition channel as the second time point. For example, assuming the terminal device acquired the tenth audio frame at 9:44:36, then 9:44:36 is taken as the second time point. Furthermore, the terminal device can calculate the time difference between the second time point and the first time point as the echo delay. For example, if the first time point is 9:44:35 and the second time point is 9:44:36, the calculated echo delay is 1 second.

[0123] In another possible implementation, the terminal device can determine a first duration based on the number of the audio frame containing the watermark information in the target audio signal and the duration of the first frame; the first frame duration is the length of time of each audio frame played; the first duration is used to characterize the time interval between the time when the audio frame containing the watermark information is played through the audio playback channel and the start time of the audio signal playback. The terminal device can determine a second duration based on the number of the audio frame containing the watermark information in the near-end audio signal and the duration of the second frame; the second frame duration is the length of time of each audio frame acquired; the second duration is used to characterize the time interval between the time when the audio frame containing the watermark information is acquired through the audio acquisition channel and the start time of the audio signal acquisition, which is the same as the start time of the audio signal playback. Then, the difference between the second duration and the first duration is calculated to obtain the echo delay.

[0124] Specifically, in real-time voice call scenarios, the audio acquisition device and audio playback device of the terminal device typically operate simultaneously. That is, after the user initiates a voice call through the terminal device, the terminal device's speaker or earpiece begins playing an audio signal (which may include blank audio signals), while the terminal device's microphone also begins acquiring audio signals from the current environment (which may also include blank audio signals). In other words, for the terminal device, the audio playback start time (i.e., the time when the audio playback device begins playing the audio signal) and the audio acquisition start time (i.e., the time when the audio acquisition device begins acquiring the audio signal) are the same.

[0125] Based on this, the terminal device can determine the echo delay according to the duration between the start of audio playback and the playback time of the audio frame containing the watermark information, and the duration between the start of audio acquisition and the acquisition time of the audio frame including the watermark information. For example, the terminal device can start from the start of audio playback and sequentially number each audio frame in the target audio signal, beginning with 1. Then, the terminal device can calculate the duration between the start of audio playback and the playback time of the audio frame containing the watermark information b, i.e., the first duration, based on the number of the audio frame containing the watermark information b and the duration of the first frame (i.e., the duration of each played audio frame). For example, assuming the terminal device embeds the watermark information b in the fifth audio frame of the target audio signal, and the duration of each played audio frame is 100ms, the calculated first duration should be 5 * 100ms = 500ms. Accordingly, the terminal device can assign numbers starting from 1 to each audio frame in the near-end audio signal according to the acquisition time sequence, starting from the initial audio acquisition time. Then, the terminal device can calculate the duration from the initial audio acquisition time to the acquisition time of the audio frame containing watermark information b, i.e., the second duration, based on the number of the audio frame containing watermark information b and the second frame duration (i.e., the duration of each acquired audio frame). For example, assuming the terminal device parses the tenth audio frame in the near-end audio signal to contain watermark information b, and the duration of each acquired audio frame is 100ms, then the calculated second duration should be 10 * 100ms = 1000ms. Furthermore, the terminal device can calculate the difference between the second duration and the first duration to obtain the echo delay. For example, if the first duration is 500ms and the second duration is 1000ms, the calculated echo delay should be 500ms.

[0126] It should be understood that the above implementation of determining the echo delay based on the position of the watermark information in the target audio signal and the position of the watermark information in the near-end audio signal is only an example. In practical applications, other methods can also be used to determine the echo delay based on the position of the watermark information in the target audio signal and the position of the watermark information in the near-end audio signal. This application does not limit the method of determining the echo delay in any way.

[0127] After determining the echo delay, the terminal device can align the near-end audio signal and the target audio signal based on the echo delay. Then, based on the target audio signal, it can perform adaptive filtering and nonlinear processing on the near-end audio signal to eliminate the echo in the near-end audio signal corresponding to the target audio signal.

[0128] In practical implementation, the terminal device can shift the target audio signal backward along the time axis based on the echo delay, so that the starting time point of the target audio signal coincides with the starting time point of the near-end audio signal. Then, the terminal device can use the target audio signal as a reference signal for echo filtering, and perform adaptive filtering on the near-end audio signal based on the target audio signal to filter out the echo in the near-end audio signal corresponding to the target audio signal. Furthermore, the terminal device can perform nonlinear processing on the near-end audio signal obtained after adaptive filtering based on the target audio signal to filter out the nonlinear echo corresponding to the target audio signal. Thus, an echo-free near-end audio signal is obtained. In real-time communication applications, the terminal device can send the near-end audio signal obtained through the above processing to the terminal device of the communicating party. In applications where the user and the terminal device are interacting, the terminal device can perform subsequent analysis and processing based on the near-end audio signal obtained through the above processing to respond to the user's voice control signals.

[0129] The aforementioned echo delay determination method applies audio watermarking technology to determine echo delay. Based on the auditory masking mechanism of the human ear, the watermark information is embedded in the reference audio signal to obtain the target audio signal without affecting the audio playback quality and without being perceived by the human ear. Since the process from audio playback to echo acquisition is a closed loop, the near-end audio signal acquired when playing the target audio signal should also include the watermark information. Therefore, the echo delay can be determined based on the position of the watermark information in the target audio signal and the near-end audio signal. On the one hand, determining the echo delay based on the watermark information in the near-end audio signal does not have special requirements for the signal-to-noise ratio of the near-end audio signal. Therefore, even when the near-end audio signal is mixed with multiple audio signals, the method provided in this application embodiment can accurately determine the echo delay. On the other hand, for the device used to execute the method provided in this application embodiment, it can accurately determine the echo delay without consuming a large amount of computing resources. Furthermore, the method provided in this application has good compatibility and versatility with different hardware devices and software applications. That is, the echo delay can be accurately determined by the method provided in this application for different hardware devices and software applications.

[0130] To facilitate a further understanding of the echo delay determination method provided in this application embodiment, the following is an exemplary description of the echo delay determination method applied to a real-time communication scenario. The echo delay determination method mainly includes two stages: watermark information injection at the playback end and watermark information parsing at the recording end.

[0131] The implementation principle of watermark information injection on the playback end is as follows: Figure 11 As shown. Specifically, it includes the following three parts:

[0132] 1) Source Encoding: Obtaining the original watermark information (corresponding to the watermark information mentioned above). This original watermark information can typically be pre-defined text information or binary encoding. When performing source encoding on the original watermark information, it can first be divided into several sub-watermark information units, byte by byte. Then, source encoding is performed on each sub-watermark information to obtain the corresponding watermark source-encoded frame. This watermark source-encoded frame includes the byte length of the original watermark information, the sequence number of the sub-watermark information within the original watermark information, the content of the sub-watermark information itself, and a checksum. This checksum can be a CRC checksum or a checksum generated using other verification methods such as block check. The specific frame structure of the watermark source-encoded frame can be as follows: Figure 8 As shown.

[0133] 2) Audio Signal Preprocessing: The received far-end audio signal is preprocessed. The main purpose of preprocessing is to detect the energy spectrum envelope of the far-end audio signal and determine the positions in the far-end audio signal where the energy spectrum envelope exceeds a preset energy threshold. These positions are used as target positions for embedding watermark source-coded frames. A watermark loading enable flag is marked at these target positions. This prevents the watermark source-coded frames from being embedded in silent or low-energy audio signals, thus avoiding the loss of valid information at the decoding end.

[0134] 3) Channel Coding: The system acquires the watermarked source-coded frame generated through source coding and the far-end audio signal marked with a watermark enable flag obtained through audio signal preprocessing. Then, it performs channel coding on the far-end audio signal based on the acquired data. Specifically, when the current coding position in the far-end audio signal is marked with a watermark enable flag, a watermarked source-coded frame is added to the audio signal at the current coding position using a watermark modulation algorithm to obtain the first signal to be encoded. Then, channel coding is performed on this first signal to obtain the channel-coded frame corresponding to the current coding position. When the current coding position in the far-end audio signal is not marked with a watermark enable flag, the audio signal at the current coding position can be directly used as the second signal to be encoded, and channel coding is performed on this second signal to obtain the channel-coded frame corresponding to the current coding position.

[0135] In selecting the watermark modulation algorithm, based on considerations of the characteristics of the application scenario, and after evaluation and comparison, the time-domain bidirectional multi-core echo-hidden watermark modulation algorithm with strong robustness, small sound quality loss, and low complexity was finally adopted. The main principle of this algorithm is to use the time-series masking mechanism of human hearing to modulate the watermark information into early reflections that are indistinguishable to the human ear. The algorithm uses bidirectional echo to resist interference caused by spatial multipath reflections, and the use of multiple cores can enhance the data transmission rate.

[0136] To improve the recognition rate of watermark information parsing and its robustness to different transmission channels, a synchronization code can be added to the header and an error correction code to the tail of the generated channel-coded frame during channel coding. The specific frame structure of this channel-coded frame can be as follows: Figure 9 As shown. The synchronization code is a fixed string of codewords used for frame synchronization, and its specific length and content can be adjusted according to the actual channel conditions. The main function of the error correction code is to reduce the bit error rate at the receiver when the channel signal-to-noise ratio is poor. In this embodiment, a 31-bit BCH error correction code, which is more suitable for short codes, can be used.

[0137] The implementation principle of watermark information parsing at the recording end is as follows: Figure 12 As shown. Specifically, it includes the following four parts:

[0138] 1) Audio demodulation: According to the frame length of the channel-coded frame mentioned above, the audio signal of the corresponding frame length in the collected near-end audio signal is demodulated to obtain the corresponding binary bit stream.

[0139] 2) Channel Decoding: Channel decoding is performed on the binary bitstream demodulated from the audio. First, frame synchronization is performed using the synchronization code in the binary bitstream. Then, error correction codes in the binary bitstream are used to correct bit errors generated during channel transmission. If error correction is successful, i.e., the number of corrected bit errors does not exceed the preset number of bit errors, subsequent watermark demodulation processing is performed on the channel decoded stream obtained from channel decoding. If error correction fails, i.e., the number of corrected bit errors exceeds the preset number of bit errors, the audio data of that frame in the near-end audio signal is discarded, and the system waits to decode the next frame of audio data in the near-end audio signal.

[0140] 3) Watermark Demodulation: The hidden coded bit stream (i.e., the watermark source coded frame) is extracted from the channel decoding stream by using the watermark demodulation algorithm corresponding to the watermark modulation algorithm. For example, if the watermark modulation algorithm used previously was the echo hiding modulation algorithm, then the cepstral method can be used to demodulate the watermark source coded frame from the channel decoding stream.

[0141] 4) Source Decoding: Perform source decoding on the watermark source encoded frame, and perform source-side error checking based on the checksum in the watermark source encoded frame; if the check passes, parse the content in the watermark source encoded frame to obtain the byte length of the original watermark information, the sequence number of the sub-watermark information carried by the watermark source encoded frame in the original watermark information, and the content of the sub-watermark information, and mark the corresponding byte detection result position as 1; if the check fails, discard the audio data of that frame in the near-end audio signal, and wait to decode the next frame of audio data in the near-end audio signal.

[0142] In response to the echo delay determination method described above, this application also provides a corresponding echo delay determination device, so that the above echo delay determination method can be applied and implemented in practice.

[0143] See Figure 13 , Figure 13 The above text Figure 6 The diagram shows a structural schematic of an echo delay determination device 1300 corresponding to the echo delay determination method illustrated. Figure 13 As shown, the echo delay determining device 1300 includes:

[0144] The watermark embedding module 1301 is used to embed watermark information into the reference audio signal to be played to obtain the target audio signal.

[0145] Audio playback module 1302 is used to play the target audio signal;

[0146] Audio acquisition module 1303 is used to acquire near-end audio signals;

[0147] Watermark parsing module 1304 is used to perform watermark information parsing processing on the near-end audio signal;

[0148] The echo delay determination module 1305 is used to determine the echo delay based on the position of the watermark information in the target audio signal and the position of the watermark information in the near-end audio signal when the watermark information is parsed from the near-end audio signal through the watermark information parsing process.

[0149] Optional, in Figure 13 Based on the echo delay determination device shown, see Figure 14 , Figure 14 This is a schematic diagram of another echo delay determining device 1400 provided in an embodiment of this application. Figure 14 As shown, the watermark embedding module 1301 includes:

[0150] The source coding submodule 1401 is used to perform source coding on the watermark information to obtain a watermark source-coded frame;

[0151] Embedded position detection submodule 1402 is used to detect the target position for embedding watermark information in the reference audio signal;

[0152] The channel coding submodule 1403 is used to perform channel coding on the reference audio signal and the watermark source coding frame based on the target position in the reference audio signal to obtain the target audio signal.

[0153] Optional, in Figure 14 Based on the echo delay determination device shown, the source coding submodule 1401 is specifically used for:

[0154] The watermark information is divided into multiple sub-watermark information by dividing it into units of preset byte length;

[0155] For each of the sub-watermark information, source encoding is performed on the sub-watermark information to obtain the watermark source encoded frame corresponding to the sub-watermark information; the watermark source encoded frame includes the byte length of the watermark information, the sequence number of the sub-watermark information in the watermark information, the sub-watermark information, and a check code.

[0156] Optional, in Figure 14 Based on the echo delay determination device shown, the embedded position detection submodule 1402 is specifically used for:

[0157] Detect the energy spectral envelope of the reference audio signal;

[0158] The location in the reference audio signal whose energy spectrum envelope exceeds a preset energy threshold is determined as the target location, and a watermark is added to the target location to enable the watermark.

[0159] Optional, in Figure 14 Based on the echo delay determination device shown, the channel coding submodule 1403 is specifically used for:

[0160] Determine whether the current encoded position of the reference audio signal is the target position;

[0161] If so, the watermark source coding frame is embedded in the audio signal at the current coding position in the reference audio signal using a watermark modulation algorithm to obtain a first signal to be coded; channel coding is performed on the first signal to be coded to obtain the channel coding frame corresponding to the current coding position.

[0162] If not, the audio signal at the current encoding position in the reference audio signal is taken as the second signal to be encoded; channel coding is performed on the second signal to be encoded to obtain the channel-coded frame corresponding to the current encoding position;

[0163] The target audio signal is obtained by combining the channel-coded frames corresponding to each coding position in the reference audio signal.

[0164] Optional, in Figure 14 Based on the echo delay determination device shown, the frame header of the channel-coded frame is used to carry a synchronization code; the frame body of the channel-coded frame is used to carry a data packet; if the channel-coded frame corresponds to the target location, the data packet includes the first signal to be encoded; if the channel-coded frame does not correspond to the target location, the data packet includes the second signal to be encoded; the frame body of the channel-coded frame is used to carry an error correction code, which is generated based on the information carried by the frame header and frame body of the channel-coded frame.

[0165] Optional, in Figure 13 Based on the echo delay determination device shown, see Figure 15 , Figure 15 A schematic diagram of another echo delay determining device 1500 provided in an embodiment of this application. (See attached diagram.) Figure 15 As shown, the watermark parsing module 1304 includes:

[0166] The audio demodulation submodule 1501 is used to demodulate the near-end audio signal to obtain a binary bit stream;

[0167] The channel decoding submodule 1502 is used to perform channel decoding on the binary bit stream to obtain a channel decoded stream;

[0168] The watermark demodulation submodule 1503 is used to perform watermark demodulation processing on the channel decoded stream using a watermark demodulation algorithm.

[0169] The source decoding submodule 1504 is used to perform source decoding on the watermark source encoded frame to obtain the watermark information after the watermark source encoded frame is demodulated from the channel decode stream through the watermark demodulation process.

[0170] Optional, in Figure 15 Based on the echo delay determination device shown, when the channel coding frame in the target audio signal includes a synchronization code and an error correction code, the channel decoding submodule 1502 is specifically used for:

[0171] Frame synchronization is performed based on the synchronization code in the binary bit stream;

[0172] During the channel decoding process of the binary bit stream, the errors in the channel decoded stream are corrected based on the error correction code in the binary bit stream;

[0173] If the number of corrected bit errors does not exceed the preset number of bit errors, the channel decoded stream is watermarked and demodulated using the watermark demodulation algorithm; if the number of corrected bit errors exceeds the preset number of bit errors, the channel decoded stream is discarded.

[0174] Optional, in Figure 15 Based on the echo delay determination device shown, the source decoding submodule 1504 is specifically used for:

[0175] The watermark source encoding frame is verified according to the check code in the watermark source encoding frame;

[0176] If the verification passes, the watermark information is obtained from the watermark source encoding frame; if the verification fails, the watermark source encoding frame is discarded.

[0177] Optional, in Figure 13 Based on the echo delay determination device shown, the echo delay determination module 1305 is specifically used for:

[0178] The time point at which the audio frame containing the watermark information is played through the audio playback channel in the target audio signal is determined as the first time point;

[0179] The time point at which the audio frame containing the watermark information is acquired through the audio acquisition channel is determined as the second time point;

[0180] The echo delay is obtained by calculating the time difference between the second time point and the first time point.

[0181] Optional, in Figure 13 Based on the echo delay determination device shown, the echo delay determination module 1305 is specifically used for:

[0182] The first duration is determined based on the number of the audio frame containing the watermark information in the target audio signal and the duration of the first frame; the first frame duration is the duration of each audio frame played; the first duration is used to characterize the time interval between the time when the audio frame containing the watermark information is played through the audio playback channel and the start time of the audio signal playback.

[0183] The second duration is determined based on the number of the audio frame containing the watermark information in the near-end audio signal and the duration of the second frame; the second frame duration is the duration of each acquired audio frame; the second duration is used to characterize the time interval between the time when the audio frame containing the watermark information is acquired through the audio acquisition channel and the audio signal start acquisition time; the audio signal start playback time is the same as the audio signal start acquisition time;

[0184] The difference between the second duration and the first duration is calculated to obtain the echo delay.

[0185] Optional, in Figure 13 Based on the echo delay determination device shown, see Figure 16 , Figure 16 This is a schematic diagram of another echo delay determining device 1600 provided in an embodiment of this application. Figure 16 As shown, the device also includes:

[0186] The echo filtering module 1601 is used to align the near-end audio signal and the target audio signal based on the echo delay; and to perform adaptive filtering and nonlinear processing on the near-end audio signal based on the target audio signal to eliminate the echo in the near-end audio signal.

[0187] The aforementioned echo delay determination device applies audio watermarking technology to determine echo delay. Based on the auditory masking mechanism of the human ear, the watermark information is embedded in the reference audio signal to obtain the target audio signal without affecting the audio playback quality and without being perceived by the human ear. Since the process from audio playback to echo acquisition is a closed loop, the near-end audio signal acquired when playing the target audio signal should also include the watermark information. Therefore, the echo delay can be determined based on the position of the watermark information in the target audio signal and the near-end audio signal. On the one hand, determining the echo delay based on the watermark information in the near-end audio signal does not have special requirements for the signal-to-noise ratio of the near-end audio signal. Therefore, even when the near-end audio signal is mixed with multiple audio signals, the device provided in this application embodiment can accurately determine the echo delay. On the other hand, for the device running the device provided in this application embodiment, it can accurately determine the echo delay without consuming a large amount of computing resources. Furthermore, the device provided in this application embodiment has good compatibility and versatility across different hardware devices and software applications. That is, for different hardware devices and software applications, the echo delay can be accurately determined using the device provided in this application embodiment.

[0188] This application also provides a device for determining echo delay. Specifically, the device may be a terminal device. The terminal device provided in this application will be described below from the perspective of hardware implementation.

[0189] See Figure 17 , Figure 17 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. For example... Figure 17 As shown, for ease of explanation, only the parts related to the embodiments of this application are shown. For specific technical details not disclosed, please refer to the method section of the embodiments of this application. The terminal device can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc. Taking a smartphone as an example:

[0190] Figure 17 This is a block diagram illustrating a portion of the structure of a smartphone related to the terminal provided in the embodiments of this application. (Reference) Figure 17 The smartphone includes: a radio frequency (RF) circuit 1710, a memory 1720, an input unit 1730, a display unit 1740, a sensor 1750, an audio circuit 1760, a wireless fidelity (WiFi) module 1770, a processor 1780, and a power supply 1790, among other components. Those skilled in the art will understand that... Figure 17The smartphone structure shown does not constitute a limitation on smartphones and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0191] The memory 1720 can be used to store software programs and modules. The processor 1780 executes various functions and data processing of the smartphone by running the software programs and modules stored in the memory 1720. The memory 1720 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, phonebook, etc.). In addition, the memory 1720 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0192] The processor 1780 is the control center of the smartphone, connecting various parts of the smartphone via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 1720 and by accessing data stored in the memory 1720. Optionally, the processor 1780 may include one or more processing units; preferably, the processor 1780 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1780.

[0193] In this embodiment of the application, the processor 1780 included in the terminal also has the following functions:

[0194] Watermark information is embedded in the reference audio signal to be played to obtain the target audio signal;

[0195] Play the target audio signal; and acquire the near-end audio signal;

[0196] The near-end audio signal is processed by watermark information parsing;

[0197] When the watermark information is parsed from the near-end audio signal through the watermark information parsing process, the echo delay is determined based on the position of the watermark information in the target audio signal and the position of the watermark information in the near-end audio signal.

[0198] Optionally, the processor 1780 is further configured to perform steps of any implementation of the echo delay determination method provided in the embodiments of this application.

[0199] This application also provides a computer-readable storage medium for storing a computer program that executes any one of the implementation methods for determining echo delay described in the foregoing embodiments.

[0200] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any one of the implementation methods for determining echo delay described in the foregoing embodiments.

[0201] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0202] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0203] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0204] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0205] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing computer programs.

[0206] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0207] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method of echo delay determination, characterized by, The method comprises: source coding the watermark information to obtain a watermark source coding frame; detecting a target position in a reference audio signal to be played for embedding the watermark information, which comprises: detecting an energy spectrum envelope of the reference audio signal; determining a position where the energy spectrum envelope of the reference audio signal exceeds a preset energy threshold as the target position, and marking a watermark loading enabling flag bit for the target position; based on the target position in the reference audio signal, channel coding the reference audio signal and the watermark source coding frame to obtain a target audio signal, which comprises: judging whether a current coding position of the reference audio signal is the target position; if yes, embedding the watermark source coding frame in an audio signal at the current coding position of the reference audio signal by a time domain multi-core echo hiding watermark modulation algorithm to obtain a first to-be-coded signal, the time domain multi-core echo hiding watermark modulation algorithm being used to modulate the watermark information into an early reflection sound that cannot be distinguished by the human ear by using a heterochronic masking mechanism of human auditory perception; channel coding the first to-be-coded signal to obtain a channel coding frame corresponding to the current coding position; if not, taking the audio signal at the current coding position of the reference audio signal as a second to-be-coded signal; channel coding the second to-be-coded signal to obtain a channel coding frame corresponding to the current coding position; and combining channel coding frames corresponding to respective coding positions of the reference audio signal to obtain the target audio signal; playing the target audio signal; and collecting a near-end audio signal; performing watermark information analysis processing on the near-end audio signal; in a case where the watermark information is analyzed from the near-end audio signal through the watermark information analysis processing, determining an echo delay according to a position of the watermark information in the target audio signal and a position of the watermark information in the near-end audio signal; wherein the determining of the echo delay according to the position of the watermark information in the target audio signal and the position of the watermark information in the near-end audio signal comprises: determining a time point at which an audio frame in which the watermark information is embedded in the target audio signal is played through an audio playing channel as a first time point; determining a time point at which an audio frame including the watermark information in the near-end audio signal is collected through an audio collecting channel as a second time point; and calculating a time difference between the second time point and the first time point to obtain the echo delay; or, According to the number of the audio frame in which the watermark information is embedded in the target audio signal and a first frame duration, a first duration is determined; the first frame duration is a time length of each audio frame played; the first duration is used to represent a length of a time interval between a time of playing the audio frame in which the watermark information is embedded and a starting playing time of the audio signal; according to the number of the audio frame in which the watermark information is included in the near-end audio signal and a second frame duration, a second duration is determined; the second frame duration is a time length of each audio frame collected; the second duration is used to represent a length of a time interval between a time of collecting the audio frame in which the watermark information is included and a starting collecting time of the audio signal; the starting playing time of the audio signal is the same as the starting collecting time of the audio signal; a difference between the second duration and the first duration is calculated to obtain the echo delay.

2. The method of claim 1, wherein, The source encoding of the watermark information to obtain a watermark source encoding frame comprises: segmenting the watermark information in a preset byte length to obtain a plurality of sub-watermark information; for each of the sub-watermark information, source encoding is performed on the sub-watermark information to obtain a watermark source encoding frame corresponding to the sub-watermark information; the watermark source encoding frame includes byte length of the watermark information, arrangement serial number of the sub-watermark information in the watermark information, the sub-watermark information, and a check code.

3. The method of claim 1, wherein, The frame header of the channel encoding frame is used to carry a synchronization code; The frame body of the channel encoding frame is used to carry a data packet; if the channel encoding frame corresponds to the target position, the data packet includes the first to-be-encoded signal; if the channel encoding frame does not correspond to the target position, the data packet includes the second to-be-encoded signal; The frame body of the channel encoding frame is used to carry an error correction code, which is generated according to the frame header and the information carried by the frame body of the channel encoding frame.

4. The method of claim 1, wherein, The watermark information analysis processing of the near-end audio signal comprises: demodulation processing is performed on the near-end audio signal to obtain a binary bit stream; channel decoding is performed on the binary bit stream to obtain a channel decoding stream; watermark demodulation processing is performed on the channel decoding stream through a watermark demodulation algorithm; in a case where a watermark source encoding frame is demodulated from the channel decoding stream through the watermark demodulation processing, source decoding is performed on the watermark source encoding frame to obtain the watermark information.

5. The method of claim 4, wherein, In a case where the channel encoding frame in the target audio signal includes a synchronization code and an error correction code, the channel decoding of the binary bit stream to obtain a channel decoding stream comprises: frame synchronization is performed based on the synchronization code in the binary bit stream; in the process of channel decoding of the binary bit stream, error correction processing is performed on error codes in the channel decoding stream based on the error correction code in the binary bit stream; if the number of corrected error codes does not exceed a preset error code number, the watermark demodulation processing of the channel decoding stream through the watermark demodulation algorithm is performed; if the number of corrected error codes exceeds the preset error code number, the channel decoding stream is discarded.

6. The method according to claim 4 or 5, characterized in that, The source decoding of the watermark source coded frame comprises: checking the watermark source coded frame according to a check code in the watermark source coded frame; if the checking is passed, obtaining the watermark information from the watermark source coded frame; if the checking is not passed, discarding the watermark source coded frame.

7. The method of claim 1, wherein, The method further comprises: aligning the near-end audio signal and the target audio signal based on the echo delay; performing adaptive filtering processing and nonlinear processing on the near-end audio signal based on the target audio signal to eliminate the echo in the near-end audio signal.

8. An echo delay determination apparatus, characterized by The device comprises: a watermark embedding module configured to embed watermark information in a reference audio signal to be played to obtain a target audio signal; an audio playing module configured to play the target audio signal; an audio collecting module configured to collect a near-end audio signal; a watermark analyzing module configured to perform watermark information analyzing processing on the near-end audio signal; an echo delay determining module configured to, in a case where the watermark information is analyzed from the near-end audio signal through the watermark information analyzing processing, determine an echo delay according to a position of the watermark information in the target audio signal and a position of the watermark information in the near-end audio signal. The watermark embedding module comprises: a source coding sub-module configured to source encode the watermark information to obtain a watermark source coded frame; a embedding position detecting sub-module configured to detect a target position in the reference audio signal for embedding watermark information, which comprises: detecting an energy spectrum envelope of the reference audio signal; determining a position where the energy spectrum envelope of the reference audio signal exceeds a preset energy threshold as the target position, and marking a watermark loading enable flag bit for the target position; a channel coding sub-module configured to channel encode the reference audio signal and the watermark source coded frame based on the target position in the reference audio signal to obtain the target audio signal, which comprises: judging whether a current coding position of the reference audio signal is the target position; if yes, embedding the watermark source coded frame in an audio signal at the current coding position of the reference audio signal by a time domain multi-core echo hiding watermark modulation algorithm to obtain a first to-be-coded signal, the time domain multi-core echo hiding watermark modulation algorithm being configured to modulate the watermark information into an early reflection sound that cannot be distinguished by human ears by using a cross-time masking mechanism of human auditory perception; and channel encoding the first to-be-coded signal to obtain a channel coded frame corresponding to the current coding position; if not, taking the audio signal at the current coding position of the reference audio signal as a second to-be-coded signal; channel encoding the second to-be-coded signal to obtain a channel coded frame corresponding to the current coding position; and combining the channel coded frames corresponding to the respective coding positions of the reference audio signal to obtain the target audio signal. The echo delay determining module is specifically configured to: determining a time point at which an audio frame embedded with the watermark information in the target audio signal is played through an audio playing channel as a first time point; determining a time point at which an audio frame including the watermark information in the near-end audio signal is collected through an audio collecting channel as a second time point; calculating a time difference between the second time point and the first time point to obtain the echo delay; or, determining a first time length according to a number of the audio frame embedded with the watermark information in the target audio signal and a first frame time length; the first frame time length is a time length of each played audio frame; the first time length is used to represent a time interval length between a time at which the audio frame embedded with the watermark information is played through the audio playing channel and a starting playing time of the audio signal; determining a second time length according to a number of the audio frame including the watermark information in the near-end audio signal and a second frame time length; the second frame time length is a time length of each collected audio frame; the second time length is used to represent a time interval length between a time at which the audio frame including the watermark information is collected through the audio collecting channel and a starting collecting time of the audio signal; the starting playing time of the audio signal is the same as the starting collecting time of the audio signal; calculating a difference value between the second time length and the first time length to obtain the echo delay.

9. The apparatus of claim 8, wherein, The source encoding submodule is specifically configured to: divide the watermark information into a plurality of sub-watermark information in a preset byte length unit; perform source encoding on each of the sub-watermark information to obtain a watermark source encoding frame corresponding to the sub-watermark information; the watermark source encoding frame includes a byte length of the watermark information, an arrangement serial number of the sub-watermark information in the watermark information, the sub-watermark information, and a check code.

10. The apparatus of claim 8, wherein, The frame header of the channel encoding frame is used to carry a synchronization code; the frame body of the channel encoding frame is used to carry a data packet; if the channel encoding frame corresponds to the target position, the data packet includes the first to-be-encoded signal; if the channel encoding frame does not correspond to the target position, the data packet includes the second to-be-encoded signal; the frame body of the channel encoding frame is used to carry an error correction code, which is generated according to the information carried by the frame header and the frame body of the channel encoding frame.

11. The apparatus of claim 8, wherein, The watermark analysis module includes: an audio demodulation submodule configured to perform demodulation processing on the near-end audio signal to obtain a binary bit stream; a channel decoding submodule configured to perform channel decoding on the binary bit stream to obtain a channel decoding stream; a watermark demodulation submodule configured to perform watermark demodulation processing on the channel decoding stream through a watermark demodulation algorithm; a source decoding submodule configured to, in a case where a watermark source encoding frame is demodulated from the channel decoding stream through the watermark demodulation processing, perform source decoding on the watermark source encoding frame to obtain the watermark information.

12. The apparatus of claim 11, wherein, In a case where the channel encoding frame in the target audio signal includes a synchronization code and an error correction code, the channel decoding submodule is specifically configured to: perform frame synchronization based on the synchronization code in the binary bit stream; In the process of channel decoding the binary bit stream, error codes in the binary bit stream are used to correct errors in the channel decoding stream; If the number of corrected error bits does not exceed the preset number of error bits, the channel decoding stream is demodulated by the watermark demodulation algorithm; if the number of corrected error bits exceeds the preset number of error bits, the channel decoding stream is discarded.

13. The apparatus of claim 11 or 12, wherein, The source decoding submodule is specifically configured to: check the watermark source encoding frame according to the check code in the watermark source encoding frame; if the check is passed, the watermark information is obtained from the watermark source encoding frame; if the check is not passed, the watermark source encoding frame is discarded.

14. The apparatus of claim 8, wherein, The device further comprises: The echo filtering module is configured to align the near-end audio signal and the target audio signal based on the echo delay, and perform adaptive filtering and non-linear processing on the near-end audio signal based on the target audio signal to eliminate the echo in the near-end audio signal.

15. An electronic device, comprising: The electronic device comprises a processor and a memory; The memory is configured to store a computer program; The processor is configured to execute the echo delay determination method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that, The computer readable storage medium is configured to store a computer program, and the computer program is configured to execute the echo delay determination method according to any one of claims 1 to 7.

17. A computer program product, characterised in that, The computer program product comprises computer instructions, and the processor of the computer device executes the computer instructions, so that the computer device executes the echo delay determination method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Audible sound positioning method and system based on hidden channel

    CN105469799A

  • Echo delay determination method and device, and smart conference device

    CN106210371A

  • Digital watermark based echo inhibition method and system

    CN106601261A

  • Time delay estimation method and device and electronic equipment

    CN109727607A