Audio-based information hidden transmission method, audio decoding method and electronic equipment
By embedding audio information into the original audio carrier in the form of echo delay, the problem of difficulty in leaking and traceability of audio information in wireless communication is solved, and higher confidentiality and security are achieved.
Patent Information
- Application Number
- CN202510240047.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-24
AI Technical Summary
In the field of wireless communications, especially in the police and military industries, there is a risk of leaked audio information by secondary recording, and it is difficult to trace the source of the leak.
By obtaining the original information to be embedded, symbol information is generated, and the delay corresponding to each symbol in the symbol information is embedded in the original audio carrier to generate an echo frame, and finally combining the echo frame with the original audio carrier to form the target audio to achieve hidden transmission of information.
It improves the confidentiality and security of communication content. Even if the target audio is recorded by a third party, the embedded original information can be parsed from the recorded audio, traced the source of the leak, and improved the robustness of information embedding.
Smart Images

Figure CN120199258A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technologies, and in particular, to an audio-based information covert transmission method, an audio decoding method, and an electronic device. Background Art
[0002] In the field of wireless communication, especially for highly sensitive industries such as police and military, the confidentiality and security of call content are of vital importance. At present, in some occasions of critical communication or playing classified audio, there is a risk of secondary recording of the audio played by the terminal through an additional recording device, which may lead to information leakage. For the audio information leaked through secondary recording, it is currently difficult to track and determine the source of the leak. Summary of the Invention
[0003] The present application provides an audio-based information covert transmission method, an audio decoding method, and an electronic device to solve the above-mentioned technical problems existing in the related technologies.
[0004] According to an embodiment of the present application, there is provided an audio-based information covert transmission method, including: obtaining original information to be embedded; generating symbol information according to the original information; obtaining the time delay corresponding to each symbol in the symbol information, where the symbol information includes a plurality of symbols, and each type of symbol corresponds to at least one time delay; obtaining an original audio carrier, and sequentially embedding the time delay corresponding to each symbol into the original audio carrier to generate an echo frame; combining the echo frame and the original audio carrier to obtain a target audio, and transmitting or playing the target audio.
[0005] According to another embodiment of the present application, there is provided an audio-based information covert transmission device, including: an obtaining module, configured to obtain original information to be embedded; a generating module, configured to generate symbol information according to the original information; an embedding module, configured to obtain the time delay corresponding to each symbol in the symbol information, where the symbol information includes a plurality of symbols, and each type of symbol corresponds to at least one time delay; obtaining an original audio carrier, and sequentially embedding the time delay corresponding to each symbol into the original audio carrier to generate an echo frame; a combining module, configured to combine the echo frame and the original audio carrier to obtain a target audio, and transmit or play the target audio.
[0006] Optionally, the audio-based information covert transmission device further includes an updating module, configured to group the symbol information according to a preset length to obtain multiple groups of symbol sequences; add synchronization symbols at specified positions in each group of symbol sequences to obtain target symbol information, where the synchronization symbols are used to mark the positions of related symbol sequences, and the synchronization symbols correspond to specific time delays; update the symbol information according to the target symbol information.
[0007] Optionally, the audio-based information covert transmission device further includes a marking module, configured to obtain a marked audio signal for marking the start and end positions or the middle position of the audio; and insert the marked audio signal into the target audio as a synchronization signal.
[0008] Optionally, the embedding module further includes a first embedding unit, configured to group the symbol information according to a preset length to obtain multiple groups of symbol sequences; determine the first starting symbol group number embedded in the previous original audio carrier, and determine the number of symbols allowed to be embedded in the previous original audio carrier; determine the second starting symbol group number to be embedded in the current original audio carrier according to the first starting symbol group number and the number of symbols; and starting from the first symbol in the symbol sequence corresponding to the second starting symbol group number, sequentially embed the time delay corresponding to each symbol into the current original audio carrier to generate an echo frame.
[0009] Optionally, the embedding module further includes a second embedding unit, configured to determine the last symbol embedded in the previous original audio carrier; and starting from the next symbol of the last symbol, sequentially embed the time delay corresponding to each symbol information into the current original audio carrier to generate an echo frame.
[0010] Optionally, the generating module is further configured to frame the original audio carrier to obtain a sequence of short original audio frames; wherein, sequentially embed the time delay corresponding to each symbol into the original audio carrier to generate an echo frame; and combine the echo frame and the original audio carrier to obtain the target audio, including: sequentially embed the time delay corresponding to each symbol into the corresponding sequence of short original audio frames to generate a sequence of echo frames, and combine the sequence of echo frames and the sequence of short original audio frames to obtain the target audio.
[0011] Optionally, the audio-based information covert transmission device further includes a noise adding module, configured to generate comfort noise; add the comfort noise to the original audio carrier to obtain an optimized audio carrier, and update the original audio carrier with the optimized audio carrier.
[0012] Optionally, the embedding module further includes a third embedding unit, configured to sequentially embed the time delay corresponding to each symbol in the symbol information into the original audio carrier, and detect whether the last symbol in the symbol information has been embedded; if the last symbol in the symbol information has been embedded, start embedding from the first symbol in the symbol information again into the original audio carrier to generate an echo frame.
[0013] According to another embodiment of the present application, an audio decoding method is provided, including: collecting target audio; extracting each audio segment of the target audio; decoding the audio segments to obtain symbol information embedded in each audio segment; and obtaining original information according to the symbol information.
[0014] According to still another embodiment of the present application, a computer storage medium is further provided. A computer program is stored in the computer storage medium, wherein the computer program is configured to execute the steps in any one of the above device embodiments when running.
[0015] According to still another embodiment of the present application, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; wherein: the memory is used for storing a computer program; the processor is used for executing the steps in the above method by running the program stored on the memory.
[0016] According to still another embodiment of the present application, a computer program product containing instructions is further provided. When it runs on a computer, it causes the computer to execute the steps in the above method.
[0017] Through the embodiments of the present application, original information to be embedded is obtained; symbol information is generated according to the original information; time delays corresponding to each symbol in the symbol information are obtained, and an original audio carrier is obtained. The time delay corresponding to each symbol is sequentially embedded into the original audio carrier to generate an echo frame; the echo frame and the original audio carrier are combined to obtain target audio. By cleverly embedding the original information to be embedded in the form of echo time delay into the original audio carrier, the confidentiality and security of communication content are improved. At the same time, even if the target audio is accidentally recorded by a third-party recording device, the embedded original information can still be parsed from the recorded audio to trace the source of the leak. The solution of the present application can embed information to be embedded in audio with uncertain time, uncertain length, and unknown content, and can improve the robustness of information embedding in audio. Description of the Drawings
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0019] Figure 1 is a hardware structure block diagram of a walkie-talkie according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of an information hiding transmission method based on audio according to an embodiment of the present application;
[0021] Figure 3It is a schematic diagram of frame division of the original audio carrier in the embodiment of the present application;
[0022] Figure 4 It is a schematic diagram of symbol generation taking a specific echo as a synchronization feature in the embodiment of the present application;
[0023] Figure 5 It is a schematic diagram of an implementation example based on a specific audio as a synchronization feature in the embodiment of the present application;
[0024] Figure 6 It is a schematic diagram of an implementation example based on a specific echo as a synchronization feature in the embodiment of the present application;
[0025] Figure 7 It is a flowchart of an audio decoding method in the embodiment of the present application;
[0026] Figure 8 It is a structural block diagram of an information hiding transmission device based on audio in the embodiment of the present application. Detailed implementation manners
[0027] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] Embodiment 1
[0030] The method embodiment provided in Embodiment 1 of the present application can be executed in an intercom, a mobile phone, a computer, a tablet or a similar computing device. The solution is applicable to devices with audio playback functions, and is mainly aimed at embedding information into audio with unknown content and unknown length. Taking the operation on an intercom as an example, Figure 1 is a hardware structure block diagram of an intercom in an embodiment of the present application. As Figure 1 shown, the intercom may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above intercom may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above intercom. For example, the intercom may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0031] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to a method for audio-based information stealth transmission in an embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the intercom through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0032] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the intercom. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0033] In this embodiment, a method for audio-based information stealth transmission is provided. Figure 2It is a flowchart of an audio-based information hiding and transmission method according to an embodiment of the present application. As Figure 2 shown, the process includes the following steps:
[0034] Step S10, obtaining the original information to be embedded;
[0035] Step S20, generating symbol information according to the original information;
[0036] In this embodiment, the original information to be embedded is any information that needs to be embedded into the audio, such as the terminal identifier of the intercom itself, such as the Short User Single Call Identification Code (ISSI) or the Serial Number (SN). After embedding the original information into the audio, the subsequent parsing of the audio with the embedded information can determine which specific terminal device the audio originates from.
[0037] In this embodiment, the original information can be converted into bit information for transmission. Generating symbol information according to the original information includes: converting the original information into a bit information sequence, and mapping the bit information sequence into symbol information, where q bit information corresponds to one symbol, q≥1. In an example of this embodiment, the generated symbol information to be embedded is: s = {s(i)|i = 1, 2, …, m}, where m is the total number of symbols, and s(i) is a symbol. One symbol can represent q bit information, that is, multiple bit information can be mapped into one symbol for transmission.
[0038] Step S30, obtaining the time delay corresponding to each symbol in the symbol information, where the symbol information includes multiple symbols, and each type of symbol corresponds to at least one time delay;
[0039] In the embodiment, a mapping table of symbols and time delays can be created in advance, and the time delay corresponding to each symbol in the symbol information is matched in the mapping table. Among them, each type of symbol corresponds to p (p≥1) echo time delays. In an example of this embodiment, the time delay array is: d m*p , d m*p = {d(i, j)|i = 1, 2, …, m; j = 1, 2, …p}, and d(i, j) is the number of echo time delays corresponding to the i-th symbol in the echo time delay array, where the number is j.
[0040] In the embodiment, any symbol can be represented by one time delay or multiple time delays in combination. Optionally, multipath time delay and / or bipolar time delay can be adopted. Among them, multipath time delay means using multiple time delay values to represent the same symbol. For example: the time delays of 5ms and 10ms both represent symbol s(1). Bipolar time delay is using positive and negative bipolar time delay values to represent the same symbol. For example: using the original signal with a time delay of 5ms and the anti-phase signal with a time delay of 10ms to represent symbol s(1) at the same time. The specific time delay values and the mapping relationship with the symbols can be set according to the actual situation and are not limited in the embodiment.
[0041] Step S40: Obtain the original audio carrier, and sequentially embed the time delay corresponding to each of the symbols into the original audio carrier to generate echo frames;
[0042] Step S50: Combine the echo frames and the original audio carrier to obtain the target audio, and transmit or play the target audio.
[0043] In the process of sequentially embedding the time delay corresponding to each of the symbols into the original audio carrier, the time delay corresponding to each symbol can be directly embedded into each frame of the original audio carrier one by one; or the original audio can be framed first and then the time delay corresponding to each symbol can be embedded;
[0044] Preferably, in the first implementation manner of this embodiment, sequentially embedding the time delay corresponding to each of the symbols into the original audio carrier to generate echo frames includes: framing the original audio carrier to obtain a sequence of original audio short frames; sequentially embedding the time delay corresponding to each of the symbols into the corresponding original audio short frames to generate echo frames corresponding to each original audio short frame.
[0045] In this implementation manner, referring to Figure 3 , the original audio carrier is: v = {v(i)|i = 1, 2,..., N}, where N is the length of all sampling points of the original audio carrier, that is, the original audio carrier includes N audio data sampling points; framing the original audio carrier to obtain a sequence of original audio short frames includes: framing the original audio carrier according to a fixed sampling point length L to obtain N / L segments of original audio short frames. After the original audio carrier is framed, it can be expressed as: f i = {f i (j)|j = 1, 2,..., L}, where f i is each segment of the original audio short frame after the original audio carrier is framed, with a length of L; f i (j) is the j-th sampling point value in f i .
[0046] For the current symbol s(i) in the symbol information, match the time delay d(i, j) corresponding to the current symbol s(i) in the mapping table, (1 ≤ j ≤ r, r ≤ p), r is the number of time delays corresponding to the current symbol s(i), p is the number of time delay sequences, and embed the current symbol s(i) in the form of the time delay d(i, j) into the corresponding original audio short frame f l to obtain r echo frames f r*l `, and combine the r echo frames f r*l ` with the corresponding original audio short frame f l to obtain the intermediate audio short frame f l`。Embed all the symbols to be embedded into the original audio short frames of the original audio carrier in the same way as the current symbol s(i) above, according to a preset order or preset rules, and the target audio with symbol information added in the form of echo can be obtained. In this embodiment, after segmenting the original audio carrier, the symbol information is continuously embedded into the sequence of original audio short frames to generate the target audio. Optionally, one symbol information can be added to one original audio short frame, and the embedding position of the symbol information can be indicated by a synchronization signal to obtain the target audio.
[0047] In this embodiment, multiple time delays are used to represent the same symbol, which can improve the information concealment and enhance the decoding performance at the decoding end, thus enhancing the reliability of information transmission.
[0048] Optionally, if multiple devices use the same delay array for audio information embedding, there may be interference when collecting the audio played by multiple devices simultaneously. For this, in this embodiment, for different devices, different delay arrays can be set according to certain rules (such as the parity of the device number or other characteristics) to reduce the mutual interference between different devices.
[0049] In this embodiment, the original information to be embedded is obtained; symbol information is generated according to the original information; the time delay corresponding to each symbol in the symbol information is obtained, and the original audio carrier is obtained. The time delay corresponding to each symbol is sequentially embedded into the original audio carrier to generate echo frames; the echo frames and the original audio carrier are combined to obtain the target audio. By cleverly embedding the original information to be embedded in the form of echo time delay into the original audio carrier, the confidentiality and security of the communication content are improved. At the same time, even if the target audio is accidentally recorded by a third-party recording device, the embedded original information can still be parsed from the recorded audio to trace the source of the leak.
[0050] In the wireless communication intercom application scenario, the start and end times and the duration of the audio played by the intercom are relatively random. The start and end times, duration, and number of recordings of the played content are also random, and the recorded audio file may contain audio with and without concealed information; at the same time, there may also be a situation where the recorded audio spans two segments of audio, such as starting to record from the beginning of the first segment of audio, pausing the recording before finishing the first segment of audio, and continuing the recording halfway through the second segment of audio. Therefore, in order to accurately determine the position of the audio carrying concealed information and determine whether the carried concealed information is continuous, a synchronization scheme needs to be designed.
[0051] In the first synchronization scheme, before obtaining the time delay corresponding to each symbol in the symbol information, the method further includes:
[0052] A1. Group the symbol information according to a preset length to obtain multiple groups of symbol sequences;
[0053] The total number m of symbols for obtaining symbol information; the symbol information is grouped according to a fixed symbol length y to obtain m / y groups of symbol sequences: Y i , i = 1, 2, … m / y. Taking the case where the total number of symbols of the symbol information is 32 and is divided into 4 groups of symbol sequences as an example for illustration:
[0054] Y1 = Y11, Y12... Y1y; Y2 =
[0055] Y21, Y22... Y2y; Y3 = Y31, Y32... Y3y; Y4 = Y41, Y42... Y4y.
[0056] A2, add a synchronization symbol at a specified position in each group of symbol sequences to obtain target symbol information, where the synchronization symbol is used to mark the position of the associated symbol sequence, and the synchronization symbol corresponds to a specific time delay; in order to better determine which segment of symbol information is added to the audio carrier, different synchronization symbols and corresponding different synchronization time delays can be set;
[0057] A3, update the symbol information according to the target symbol information.
[0058] Obtain a synchronization symbol sequence, such as Figure 4 the synchronization symbols SYNC1, SYNC2 and SYNC3 in. Pre-create a mapping relationship between the synchronization symbol and the time delay. For example, the synchronization symbol SYNC1 corresponds to an echo time delay of 15 ms, the synchronization symbol SYNC2 corresponds to an echo time delay of 20 ms, etc.
[0059] Add a synchronization symbol at a specified position (the specified position can be the head, middle or tail of the symbol sequence to be sent, as long as the relative position is clear) in each group of symbol sequences to be sent to form target symbol information (such as Figure 4 the symbol sequence to be sent with the synchronization symbol added in). The synchronization symbol is used to mark the position of the associated symbol sequence to be sent. In this embodiment, the synchronization symbol corresponds to a specific time delay, and every other symbol sequence length, n known echo values are embedded. For example, every 10 symbols, 2 known echo values are embedded, and there can be multiple groups as the synchronization time delay.
[0060] In one example, there are 40 symbols of symbol information to be embedded. Every 10 symbols are embedded into a group of known synchronization symbols. Then there are a total of 4 symbol sequences with synchronization symbols, and the time delay corresponding to each symbol sequence represents 10 fixed symbols. When decoding the recorded audio, if the detected synchronization time delay is that of the fourth group, it means that the fourth symbol sequence has been received. In walkie-talkie communication, sometimes the voice information is short, perhaps only lasting for 1 or 2 seconds, but it may take 20 seconds to completely send the bit information. By grouping the symbol information in this embodiment, adding synchronization symbols to each symbol sequence to obtain the target symbol information, and segmentally embedding the target symbol information into the audio, the short voice information can be gradually accumulated. For example, the first voice accumulates the first and second symbols, and the second voice accumulates the third and fourth symbols. Eventually, all the information can be completely collected, improving the flexibility of information embedding based on audio.
[0061] In one example of this embodiment, the synchronization time delay (such as 15 ms) corresponds to the transmitted symbol sequence (0123456789). If the 15-ms synchronization time delay is lost at the decoding end, these 10 digits (0123456789) cannot be decoded. If the symbol information is subdivided, for example, the synchronization time delay (15 ms) corresponds to the transmitted symbol sequence (01), and the synchronization time delay (20 ms) corresponds to the transmitted symbol sequence (23). Even if the 15-ms data is lost, only the decoding of these two digits (01) is affected, and the transmitted symbol sequence (23) corresponding to 20 ms can still be correctly decoded. This embodiment maps the synchronization symbols with a fixed echo value, and for the information to be transmitted, a specific echo delay value is added as a synchronization means. Combining the synchronization scheme, continuously segmentally adding the symbol information, with each segment of symbol information corresponding to a different synchronization symbol, solves the problem that the audio segment duration is short and complete information cannot be added, resulting in wasted voice.
[0062] In the second synchronization scheme, after combining the echo frame and the original audio carrier to obtain the target audio, the method further includes:
[0063] B1. Obtain a marked audio signal, where the marked audio signal is used to mark the start and end positions or the middle position of the audio;
[0064] B2. Insert the marked audio signal as a synchronization signal into the target audio.
[0065] In this embodiment, a specific marked audio signal (such as a specific frequency or tone) is used as a synchronization feature, which is known to both the concealed information adding end and the concealed information decoding end. In each audio segment (which can be each original audio carrier segment, or sub-segments obtained by splitting each original audio carrier segment), the specific marked audio signal is added at specific positions according to a predetermined rule. For example, at the starting point, ending point of the audio segment, or a specific prompt tone is inserted at fixed intervals. Each prompt tone has different audio characteristics to mark the start, end, or intermediate position of the audio.
[0066] The marked audio signals corresponding to each audio segment can be correlated with each other to jointly represent a continuous audio segment, which is called an audio pair. In this embodiment, different audio pairs can be used for each audio segment to address the problems of possible recording interruptions and resumptions during the recording process, resulting in incorrect audio connections and thus incorrect information carrying sequences.
[0067] In an example of this embodiment, assume that there are three audio segments, and specific audio pairs (such as an audio pair includes a start audio signal and an end audio signal) are inserted at the start and end positions of each audio segment. For example, the first segment is a and b, the second segment is c and d, and the third segment is e and f. During the recording process, due to the randomness of the recording start and end times, there may be a situation where the start or end marker of a certain audio segment is not completely recorded (such as starting to record with a but not recording b, not recording c but starting to record from d). Then, by analyzing the corresponding relationships of these marked audio signals, the incorrect audio segments caused by connection errors can be identified and excluded. Specifically, if the start marker of a certain audio segment in the recording is a, but the end marker is not the expected b but d, it can be determined that there is an error in this segment of the recording, and corresponding strategies can be taken during the decoding process to exclude it. Through this embodiment, the uncertainty of the recording start and end time points, the uncertainty of the distance between the recording device and the audio playback device, the pause and resume scenarios during the recording process, and other sounds during the recording interval are resisted, improving the accuracy and reliability of information parsing in the audio.
[0068] In an embodiment of this embodiment, the original audio carrier includes multiple segments, and the time delay corresponding to each symbol is sequentially embedded into the original audio carrier. Generating an echo frame includes:
[0069] C41, grouping the symbol information according to a preset length to obtain multiple groups of symbol sequences;
[0070] Obtain the total number of symbols m of the symbol information; group the symbol information according to a fixed symbol length y to obtain m / y groups of symbol sequences: Y i , i = 1, 2,... m / y.
[0071] C42. Determine the first starting symbol group number embedded in the previous original audio carrier, and determine the number of symbols that the previous original audio carrier allows to be embedded;
[0072] C43. According to the first starting symbol group number and the number of symbols, determine the second starting symbol group number to be embedded in the current original audio carrier;
[0073] C44. Starting from the first symbol of the second starting symbol group number, sequentially embed the time delay corresponding to each symbol into the current original audio carrier to generate an echo frame.
[0074] In the case where the original audio carrier includes multiple segments, since the lengths of each segment of audio are not fixed, there may be a situation where the number of symbols that the audio length can embed is greater than or equal to the number of symbols to be embedded, and a situation where the number of symbols that the audio length can embed is less than the number of symbols to be embedded. Therefore, to ensure the complete embedding of symbols, the starting symbol value embedded at the starting point of each original audio carrier is not always 1, but the first symbol of the symbol sequence Y i where the value of the group number i is the group number to which the last symbol embedded in the previous segment of audio belongs.
[0075] In this embodiment, by obtaining the first starting symbol group number i n-1 embedded in the previous original audio carrier (the original audio of the (n - 1)th call, n > 1), n-1 and the number of symbols a that the previous original audio carrier allows to be embedded, n=1 according to the first starting symbol group number i n-1 and the number of symbols a, n determine the second starting symbol group number i to be embedded in the current original audio carrier (the original audio of the nth call). In this embodiment, the symbol information is grouped according to a preset length to obtain multiple groups of symbol sequences, each group of symbol sequences corresponding to a unique group number. The first starting symbol group number is the group number to which the first symbol embedded in the previous original audio carrier belongs, and the second starting symbol group number is the group number to which the starting symbol to be embedded in the current original audio carrier belongs.
[0076] Optionally, the second starting symbol group number i n can be calculated by the following formula:
[0077] where % is the modulo operation. When the current original audio carrier is the first segment of the audio carrier (such as the first call), i1 = 1.
[0078] In a specific example of this embodiment: Assume that the total number of symbols m of the symbol information is 32, and the symbol information is grouped according to a fixed symbol length y = 8, obtaining 4 groups of symbol sequences: Y1, Y2, Y3, Y4, and each group of sequences corresponds to 8 symbols. The walkie-talkie makes 5 calls corresponding to 5 segments of original audio carriers, and the number of symbols that can be embedded in each segment of the original audio carrier is a1 = 21, a2 = 11, a3 = 5, a4 = 8, a5 = 29 respectively. Then the starting sequence numbers of the symbol groups embedded each time are as follows, i1 = 1, i2 = 3, i3 = 4, i4 = 4, i5 = 1, i6 = 4.
[0079] In this embodiment, the calculation of the sequence number of the second starting symbol to be embedded in the current original audio carrier can also be carried out in the following manner: Calculate the ratio of the number of symbols allowed to be embedded in the previous original audio carrier to the preset length, and after rounding down, multiply it by the preset length to obtain the first sequence number, and use the value obtained by adding 1 to the first sequence number as the sequence number of the second starting symbol to be embedded in the current original audio carrier. For example, the calculation process of the sequence number of the second starting symbol to be embedded in the current original audio carrier is as follows: The number of symbols allowed to be embedded in the previous original audio carrier is 21, the preset length is 8, the ratio between the two is rounded down to 2, and 2 multiplied by 8 gives 16. Then the sequence number of the second starting symbol to be embedded in the current original audio carrier is 17, that is, the current original audio carrier starts to be embedded from the 17th symbol as the starting symbol.
[0080] This embodiment determines the starting embedding symbol of each segment of audio. During rerecording and decoding, the information embedded in each segment of audio can be accurately located and parsed. Against the randomness problems in the audio rerecording process, such as the uncertain start and end times of audio playback, the uncertain playback duration, the unknown playback content, etc.
[0081] In another embodiment of this example, the original audio carrier includes multiple segments, and the time delay corresponding to each symbol is sequentially embedded into the original audio carrier to generate an echo frame, including:
[0082] D41, determine the last symbol embedded in the previous original audio carrier;
[0083] D42, starting from the next symbol after the last symbol, sequentially embed the time delay corresponding to each symbol information into the original audio short frame corresponding to the current original audio carrier to generate an echo frame.
[0084] When processing multiple segments of original audio carriers, it is possible that the length of a single segment of speech is not sufficient to accommodate all the information to be embedded. In this case, the information may be partially embedded in the first segment of speech and then the remaining part is continued to be embedded in the subsequent segments.
[0085] In this embodiment, first, determine the last symbol already embedded in the previous segment of the original audio carrier. This can be done by maintaining a counter to track the number of symbols already embedded, or by using a specific time delay value to identify the last embedded symbol in each segment of the audio carrier. Then, starting from the symbol next to the last symbol, sequentially embed the time delay corresponding to each symbol into the original audio carrier to generate an echo frame. This embodiment can ensure the continuity and accuracy of the embedding process.
[0086] For the first synchronization scheme mentioned above, the symbols to be transmitted use the same embedding algorithm as the synchronization symbols. An embodiment of the first synchronization scheme includes: obtaining the original information to be embedded; generating symbol information according to the original information; grouping the symbol information according to a preset length to obtain multiple groups of symbol sequences; adding synchronization symbols at specified positions in each group of symbol sequences to obtain target symbol information, where the synchronization symbols are used to mark the positions of the associated symbol sequences, and the synchronization symbols correspond to specific time delays; updating the symbol information according to the target symbol information; obtaining the time delays corresponding to each symbol in the updated symbol information, where the symbol information includes multiple symbols, and each type of symbol corresponds to at least one time delay; obtaining the original audio carrier; grouping the symbol information according to a preset length to obtain multiple groups of symbol sequences; determining the first starting symbol group number embedded in the previous segment of the original audio carrier, and determining the number of symbols allowed to be embedded in the previous segment of the original audio carrier; determining the second starting symbol group number to be embedded in the current original audio carrier according to the first starting symbol group number and the number of symbols; starting from the first symbol of the second starting symbol group number, sequentially embed the time delay corresponding to each symbol into the current original audio carrier to generate an echo frame; combining the echo frame and the original audio carrier to obtain the target audio, and transmitting or playing the target audio.
[0087] Another embodiment of the first synchronization scheme includes: obtaining the original information to be embedded; generating symbol information according to the original information; grouping the symbol information according to a preset length to obtain multiple groups of symbol sequences; adding synchronization symbols at specified positions in each group of symbol sequences to obtain target symbol information, where the synchronization symbols are used to mark the positions of the associated symbol sequences, and the synchronization symbols correspond to specific time delays; updating the symbol information according to the target symbol information; obtaining the time delays corresponding to each symbol in the updated symbol information, where the symbol information includes multiple symbols, and each type of symbol corresponds to at least one time delay; obtaining the original audio carrier and determining the last symbol embedded in the previous segment of the original audio carrier; starting from the symbol next to the last symbol, sequentially embed the time delay corresponding to each symbol information into the original audio short frame corresponding to the current original audio carrier to generate an echo frame; combining the echo frame and the original audio carrier to obtain the target audio, and transmitting or playing the target audio.
[0088] An implementation manner for the second synchronization scheme includes: obtaining the original information to be embedded; generating symbol information according to the original information; obtaining the time delay corresponding to each symbol in the symbol information, where the symbol information includes multiple symbols, and each type of symbol corresponds to at least one time delay; obtaining the original audio carrier, grouping the symbol information according to a preset length to obtain multiple groups of symbol sequences; determining the first starting symbol group number embedded in the previous original audio carrier, and determining the number of symbols allowed to be embedded in the previous original audio carrier; determining the second starting symbol group number to be embedded in the current original audio carrier according to the first starting symbol group number and the number of symbols; starting from the first symbol of the second starting symbol group number, sequentially embedding the time delay corresponding to each symbol into the current original audio carrier to generate an echo frame; combining the echo frame and the original audio carrier to obtain a target audio, obtaining a marked audio signal, where the marked audio signal is used to mark the start and end positions or the middle position of the audio; inserting the marked audio signal as a synchronization signal into the target audio.
[0089] Another implementation manner for the second synchronization scheme includes: obtaining the original information to be embedded; generating symbol information according to the original information; obtaining the time delay corresponding to each symbol in the symbol information, where the symbol information includes multiple symbols, and each type of symbol corresponds to at least one time delay; obtaining the original audio carrier, determining the last symbol embedded in the previous original audio carrier; starting from the next symbol of the last symbol, sequentially embedding the time delay corresponding to each symbol information into the original audio short frame corresponding to the current original audio carrier to generate an echo frame; combining the echo frame and the original audio carrier to obtain a target audio, obtaining a marked audio signal, where the marked audio signal is used to mark the start and end positions or the middle position of the audio; inserting the marked audio signal as a synchronization signal into the target audio.
[0090] In any implementation manner of the present application, the original audio carrier can be framed to obtain a sequence of original audio short frames, and the time delay corresponding to each symbol in the symbol information is sequentially embedded into each short frame of the sequence of original audio carrier short frames one by one to obtain a sequence of echo frames, and then each echo frame and the corresponding original audio short frame are sequentially combined to obtain a target audio.
[0091] In another implementation manner of this embodiment, after obtaining the original audio carrier, the method further includes: generating comfort noise; adding the comfort noise to the original audio carrier to obtain an optimized audio carrier, and updating the original audio carrier with the optimized audio carrier.
[0092] There may be intermittent segments (i.e., silent segments) in the original audio signal. The audio amplitude of these intermittent segments is extremely small and even tends to 0. At this time, echo cannot be added to this segment of audio. In view of this situation, in this embodiment, comfort noise is added to avoid the problem that the hidden information in the audio intermittent segment cannot be added. In this embodiment, noise can be added to the overall original audio, or noise can be added at the speech intermittent positions. The method for adding comfort noise is as follows:
[0093] Generate Gaussian white noise, filter the Gaussian white noise to obtain comfort noise g = {g(i)|i = 1, 2, … N}, where N is the length of all sampling points of the original audio carrier. Weighted sum the comfort noise and the original audio carrier, and the final audio carrier is obtained. The specific weighting method is as follows v = α * v in + β * g; where v in is the original audio carrier sequence, g is the comfort noise sequence, α and β are the weighting values of the original audio carrier and the comfort noise respectively, and v is the audio carrier sequence after weighted summation of the original audio carrier and the comfort noise. Subsequently, v is used to add information.
[0094] In this embodiment, the time delay corresponding to each symbol is sequentially embedded into the original audio carrier. Generating an echo frame includes: sequentially embedding the time delays corresponding to the symbols in the symbol information into the original audio carrier, and detecting whether the last symbol in the symbol information has been embedded; if the last symbol in the symbol information has been embedded, start embedding from the first symbol in the symbol information into the original audio carrier again to generate an echo frame. In this embodiment, after each embedding of the last symbol in the symbol information, start circular embedding from the starting symbol to increase the information validity of multiple segments of audio and improve the parsing performance.
[0095] In this embodiment, an audio decoding method is also provided, as Figure 7 shown. The process includes the following steps:
[0096] S60, collect the target audio;
[0097] S70, extract each audio segment of the target audio;
[0098] S80, decode the audio segment to obtain the symbol information embedded in each audio segment;
[0099] S90, obtain the original information according to the symbol information.
[0100] In this embodiment, an audio decoding end (such as a recording device) collects target audio played by a walkie-talkie or other audio playback device through recording. During parsing, the decoding end first detects a synchronization signal in the target audio. According to the synchronization signal, each audio segment of the target audio is extracted, and the audio segment is decoded to obtain symbol information embedded in each audio segment; then the original information is obtained based on the symbol information. By detecting the synchronization signal, accurate extraction and parsing of information can be ensured.
[0101] Specifically, decoding the audio segment to obtain the symbol information embedded in each audio segment includes: parsing the delay embedded in the current audio segment, and based on the mapping relationship between the symbol and the delay, the symbol information embedded in the current audio segment can be parsed.
[0102] An implementation example of this embodiment based on a specific audio feature as a synchronization feature is as Figure 5 shown. Taking a walkie-talkie call as an example for description. It includes several modules such as voice input, comfort noise generation, comfort noise addition, frame segmentation, generation of symbol information to be embedded, information embedding, and synchronization addition. The functions of each module are described in detail as follows:
[0103] 1. Voice input: Refers to multiple segments of voice audio to be played.
[0104] 2. Comfort noise generation: For the input voice audio, there is an inter-call period, and the value in the processor is small, making it difficult to add information. Comfort noise is generated and added to the input voice audio for information addition. Gaussian noise can be used for filtering and then used as comfort noise.
[0105] 3. Comfort noise addition: Perform a weighted sum of the comfort noise and the input audio. The weight of the comfort noise can be 0, that is, no comfort noise is added.
[0106] 4. Frame segmentation module: Segment the input voice audio, which is a preprocessing for the information embedding module.
[0107] 5. Generation of symbol information to be embedded: This module can process the bit information to be embedded, including optional checking and encoding, and map the bit information to symbol information. Finally, the symbol information is embedded into the audio. One symbol can correspond to multiple bits.
[0108] 6. Information embedding: Use the echo algorithm for information embedding. One symbol can correspond to multiple echo delays, and a weighted sum is performed on multiple echo delays.
[0109] 7. Synchronization addition: Add synchronization audio to the voice after information embedding. In this implementation example, synchronization audio is only added at the beginning and end of the audio.
[0110] An implementation example of this embodiment based on a specific echo as a synchronization feature is asFigure 6 As shown below. Taking a walkie-talkie call as an example for description, it includes several modules such as voice input, comfort noise generation, comfort noise addition, framing, generation of embedded symbol information with added synchronization, synchronization addition, and information embedding. The functions of each module are described in detail as follows:
[0111] 1. Voice input: Refers to multiple segments of voice audio to be played.
[0112] 2. Comfort noise generation: For the input voice audio, there is an intermittent period during the call, and the value in the processor is relatively small, making it difficult to add information. Comfort noise is generated and added to the input voice audio for information addition. Gaussian noise can be used for filtering and then used as comfort noise.
[0113] 3. Comfort noise addition: Perform a weighted sum of the comfort noise and the input audio. The weight of the comfort noise can be 0, that is, no comfort noise is added.
[0114] 4. Framing module: Frame the input voice audio, which is a preprocessing for the information embedding module.
[0115] 5. Generate the symbol information to be embedded with added synchronization. This module can process the bit information to be embedded, including optional checking and encoding. And map the bit information to symbol information, where one symbol can correspond to multiple bits. And according to the rules, segment the symbol sequence, insert a synchronization symbol sequence into each segment, and then merge it as the symbol sequence to be embedded.
[0116] 6. Information embedding: Perform information embedding with an echo algorithm. One symbol can correspond to multiple echo delays, and a weighted sum is performed on multiple echo delays.
[0117] In this embodiment, by grouping symbol information, adding synchronization symbols, embedding target symbol information, and processing at the decoding end, information addition and extraction of audio signals can be achieved, while ensuring a certain degree of flexibility and robustness. It is applicable to the usage characteristics of walkie-talkie products and can meet the call usage scenarios that are indefinite in time, length, and unknown. It has no restrictions on the audio characteristics of the product (such as sampling rate, frequency response, equalization, etc.), and has little impact on the voice quality. Moreover, the method for concealing information in audio is simple, has a low coupling degree with the product, and can be integrated into any product with a voice playback function, improving the product security. This embodiment can also be applied to other scenarios. For example, information is added at the wireless communication sending end, and the receiving end can parse and know the information content hidden in the audio by the sending end, and multiple applications can be realized.
[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0119] Embodiment 2
[0120] In this embodiment, an audio watermark embedding device is also provided to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0121] Figure 8 is a structural block diagram of an information hiding and transmission device based on audio according to an embodiment of the present application, as Figure 8 shown, the device includes:
[0122] An acquisition module 81, configured to acquire the original information to be embedded;
[0123] A generation module 82, configured to generate symbol information according to the original information; an embedding module, configured to acquire the time delay corresponding to each symbol in the symbol information, where the symbol information includes multiple symbols, and each type of symbol corresponds to at least one time delay; acquire an original audio carrier, and sequentially embed the time delay corresponding to each symbol into the original audio carrier to generate an echo frame;
[0124] A combination module 83, configured to combine the echo frame and the original audio carrier to obtain a target audio, and transmit or play the target audio.
[0125] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.
[0126] Embodiment 3
[0127] An embodiment of the present application further provides a computer storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0128] Optionally, in this embodiment, the above computer storage medium may be configured to store a computer program for executing the following steps:
[0129] S1. Obtain the original information to be embedded;
[0130] S2. Generate symbol information according to the original information;
[0131] S3. Obtain the time delay corresponding to each symbol in the symbol information, where the symbol information includes a plurality of symbols, and each type of symbol corresponds to at least one time delay;
[0132] S4. Obtain the original audio carrier, and sequentially embed the time delay corresponding to each symbol into the original audio carrier to generate an echo frame;
[0133] S5. Combine the echo frame and the original audio carrier to obtain a target audio, and transmit or play the target audio.
[0134] Optionally, in this embodiment, the above computer storage medium may include but is not limited to: various media that can store computer programs such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs.
[0135] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0136] Optionally, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0137] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0138] S1. Obtain the original information to be embedded;
[0139] S2. Generate symbol information according to the original information;
[0140] S3. Obtain the time delay corresponding to each symbol in the symbol information, where the symbol information includes a plurality of symbols, and each type of symbol corresponds to at least one time delay;
[0141] S4. Obtain the original audio carrier, and sequentially embed the time delay corresponding to each symbol into the original audio carrier to generate an echo frame;
[0142] S5. Combine the echo frame and the original audio carrier to obtain a target audio, and transmit or play the target audio.
[0143] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0144] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0145] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0146] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the units or modules can be in an electrical or other form.
[0147] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0149] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned computer storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0150] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for concealed information transmission based on audio, characterized in that: The method comprises: Get the original information to be embedded; Generate symbol information according to the original information; Acquire a time delay corresponding to each symbol in the symbol information, wherein the symbol information includes a plurality of symbols, and each symbol corresponds to at least one time delay; Obtaining an original audio carrier, and sequentially embedding a time delay corresponding to each of the symbols into the original audio carrier to generate an echo frame; The echo frame and the original audio carrier are combined to obtain target audio, and the target audio is transmitted or played.
2. The method according to claim 1, characterized in that Before acquiring the delay corresponding to each symbol in the symbol information, the method further includes: Grouping the symbol information according to a preset length to obtain multiple groups of symbol sequences; Adding a synchronization symbol at a specified position of each group of symbol sequences to obtain target symbol information, wherein the synchronization symbol is used to mark the position of the associated symbol sequence, and the synchronization symbol corresponds to a specific time delay; The symbol information is updated according to the target symbol information.
3. The method according to claim 1, characterized in that After combining the echo frame and the original audio carrier to obtain the target audio, the method further includes: Acquire a marked audio signal, where the marked audio signal is used to mark the start and end positions or the middle position of the audio; The marker audio signal is inserted into the target audio as a synchronization signal.
4. The method according to any one of claims 1 to 3, characterized in that: The original audio carrier includes multiple segments, and the time delay corresponding to each symbol is embedded into the original audio carrier in sequence, and the echo frame is generated, comprising: Grouping the symbol information according to a preset length to obtain multiple groups of symbol sequences; Determine the first starting symbol group number embedded in the previous original audio carrier, and determine the number of symbols allowed to be embedded in the previous original audio carrier; Determine, according to the first starting symbol group number and the number of symbols, a second starting symbol group number to be embedded in the current original audio carrier; Starting from the first symbol in the symbol sequence corresponding to the second starting symbol group number, the time delay corresponding to each symbol is embedded into the current original audio carrier in turn to generate an echo frame.
5. The method according to any one of claims 1 to 3, characterized in that: The original audio carrier includes multiple segments, and the time delay corresponding to each symbol is embedded into the original audio carrier in sequence, and the echo frame is generated, comprising: Determine the last symbol embedded in the previous raw audio carrier; Starting from the next symbol of the last symbol, the time delay corresponding to each symbol is sequentially embedded into the current original audio carrier to generate an echo frame.
6. The method according to claim 1, characterized in that After obtaining the original audio carrier, the method further includes: Dividing the original audio carrier into frames to obtain an original audio short frame sequence; The method includes sequentially embedding the time delay corresponding to each of the symbols into the original audio carrier to generate an echo frame; and combining the echo frame and the original audio carrier to obtain the target audio, including: The time delay corresponding to each of the symbols is sequentially embedded into the corresponding original audio short frame sequence to generate an echo frame sequence, and the echo frame sequence and the original audio short frame sequence are combined to obtain the target audio.
7. The method according to claim 1, characterized in that After obtaining the original audio carrier, the method further includes: generating comfort noise; The comfort noise is added to the original audio carrier to obtain an optimized audio carrier, and the optimized audio carrier is used to update the original audio carrier.
8. The method according to claim 1, characterized in that The step of sequentially embedding the time delay corresponding to each of the symbols into the original audio carrier to generate an echo frame comprises: The time delay corresponding to each symbol in the symbol information is sequentially embedded into the original audio carrier, and it is detected whether the last symbol of the symbol information has been embedded; If it has been embedded into the last symbol of the symbol information, then it will be embedded into the original audio carrier again starting from the first symbol of the symbol information to generate an echo frame.
9. An audio decoding method, characterized in that: The method comprises: Collect target audio; Extracting each audio segment of the target audio; Decoding the audio segments to obtain symbol information embedded in each audio segment; The original information is obtained according to the symbol information.
10. An audio-based information concealment transmission device, characterized in that: The device comprises: An acquisition module, used to acquire original information to be embedded; A generating module, used for generating symbol information according to the original information; An embedding module is used to obtain a time delay corresponding to each symbol in the symbol information, wherein the symbol information includes multiple symbols and each symbol corresponds to at least one time delay; obtain an original audio carrier, and sequentially embed the time delay corresponding to each of the symbols into the original audio carrier to generate an echo frame; A combining module is used to combine the echo frame and the original audio carrier to obtain target audio, and transmit or play the target audio.
11. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; wherein: Memory, used to store computer programs; A processor, configured to execute the audio-based covert information transmission method according to any one of claims 1 to 8 by running a program stored in a memory.
12. A computer storage medium, characterized in that: The computer storage medium includes a stored program, wherein the program, when executed, executes the audio-based covert information transmission method according to any one of claims 1 to 8.
Citation Information
Cited By
Audio-based covert information transmission method, audio decoding method, and electronic device
WO2026179371A1