Information transmission method and device based on audio two-dimensional code, equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2026-08-11
AI Technical Summary
然而,不同读取场景中产生的噪音容易给信息传输带来干扰,进而使得音频二维码失真,最终造成基于音频二维码读取得到的目标信息不准确
[0049]本申请提出的音频二维码的信息传输方法、装置、设备及介质,其在应用于终端时,与客户端通信连接,通过获取预设的初始音频和待传输的目标信息;分别对初始音频和目标信息进行单独编码,得到音频编码特征和信息编码特征;对音频编码特征和信息编码特征进行混合编码,得到初始音频二维码;向客户端发送初始音频二维码,以使客户端接收终端发送的初始音频二维码,为初始音频二维码配置多种类型下的失真音频,并在初始音频二维码中引入多种失真音频,得到增强音频二维码,对增强音频二维码进行解码,得到目标信息;可以理解的是,通过在客户端向接收到初始音频二维码中引入多种失真音频,以使客户端从初始音频二维码的音频信号中剔除与引入的失真音频相同的或者相类似的信号,进而达到对初始音频二维码进行音频信号增强的目的,并对之后得到的增强音频二维码进行解码,得到具备较高准确度的目标信息。
Smart Images

Figure CN119446157B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to an information transmission method, apparatus, device and medium based on audio QR codes. Background Technology
[0002] An audio QR code is an audio message carrying QR code information, designed to enable the rapid transmission of information such as text, images, or files using sound.
[0003] In related technologies, target information is transmitted by reading the audio carrying QR code information from an audio QR code. However, noise generated in different reading scenarios can easily interfere with information transmission, causing distortion of the audio QR code and ultimately resulting in inaccurate target information obtained from reading the audio QR code. Summary of the Invention
[0004] The main objective of this application is to propose an information transmission method, apparatus, device, and medium based on audio QR codes, aiming to improve the accuracy of transmitting target information based on audio QR codes.
[0005] To achieve the above objectives, a first aspect of this application proposes an information transmission method based on audio QR codes, applied to a terminal, wherein the terminal and a client are connected in communication, and the method includes:
[0006] Obtain the preset initial audio and the target information to be transmitted;
[0007] The initial audio and target information are encoded separately to obtain audio coding features and information coding features;
[0008] The audio coding features and information coding features are mixed and encoded to obtain the initial audio QR code;
[0009] Send an initial audio QR code to the client so that the client can receive the initial audio QR code sent by the terminal. Configure multiple types of distortion audio for the initial audio QR code and introduce multiple types of distortion audio into the initial audio QR code to obtain an enhanced audio QR code. Decode the enhanced audio QR code to obtain the target information.
[0010] In some embodiments, the initial audio QR code is output by a pre-trained encoder, which is trained through the following steps:
[0011] Obtain the preset initial audio of the sample and the target information of the sample to be transmitted;
[0012] The initial audio of the sample and the target information of the sample are encoded separately to obtain the audio encoding features and the information encoding features of the sample.
[0013] The sample audio encoding features and sample information encoding features are mixed and encoded to obtain the initial audio QR code of the sample.
[0014] Based on the time domain difference between the initial audio sample and the initial audio sample QR code, the time domain loss value is calculated, and based on the frequency domain difference between the initial audio sample and the initial audio sample QR code, the frequency domain loss value is calculated.
[0015] Based on the time-domain loss value and the frequency-domain loss value, the encoder parameters are adjusted to obtain the trained encoder.
[0016] In some embodiments, applied to a client, the client communicating with a terminal, the method includes:
[0017] The receiving terminal sends an initial audio QR code, which is obtained by mixing audio encoding features and information encoding features; the audio encoding features and information encoding features are obtained by separately encoding the acquired preset initial audio and the target information to be transmitted.
[0018] Configure multiple types of distorted audio for the initial audio QR code, and introduce multiple distorted audio into the initial audio QR code to obtain an enhanced audio QR code;
[0019] The enhanced audio QR code is decoded to obtain the target information.
[0020] In some embodiments, multiple distorted audio values are introduced into the initial audio QR code to obtain an enhanced audio QR code, including:
[0021] Acquire ambient audio surrounding the receiving terminal when it sends the initial audio;
[0022] Select at least one target distortion audio that matches the scene audio from a plurality of configured distortion audios, and introduce at least one target distortion audio into the initial audio QR code to obtain an enhanced audio QR code;
[0023] Alternatively, an adjustable probability weight can be assigned to each distorted audio, and at least one target distorted audio can be randomly selected from multiple distorted audios based on the probability weight. This target distorted audio can then be introduced into the initial audio QR code to obtain an enhanced audio QR code.
[0024] In some embodiments, the enhanced audio QR code is decoded to obtain target information, including:
[0025] At least one sine wave signal is extracted from the enhanced audio QR code;
[0026] Select one of the preset decoders that matches the periodic characteristics of the sine signal to decode the enhanced audio QR code and obtain the target information;
[0027] When there are multiple sinusoidal signals, multiple decoders that match the periodic characteristics of each sinusoidal signal decode the audio QR code respectively to obtain multiple sub-target information;
[0028] Multiple sub-target information is fused to obtain target information.
[0029] In some embodiments, the target information is output by a pre-trained decoder, which is trained through the following steps:
[0030] The receiving terminal sends a sample initial audio QR code, which is obtained by mixing and encoding the sample audio encoding features and the sample information encoding features; the sample audio encoding features and the sample information encoding features are obtained by separately encoding the acquired preset sample initial audio and the target information of the sample to be transmitted.
[0031] Decode the initial audio QR code of the sample to obtain the target information of the first sample;
[0032] Configure sample distortion audio of various types for the initial sample audio QR code, and introduce various sample distortion audio into the initial sample audio QR code to obtain sample enhanced audio QR code;
[0033] Decode the enhanced audio QR code of the sample to obtain the target information of the second sample;
[0034] Based on the initial sample information and the first sample target information, the first reconstruction loss value is calculated, wherein the initial sample information is used to generate the initial audio QR code of the sample;
[0035] Based on the initial sample information and the target information of the second sample, the first distortion loss value is calculated, and based on the first sample target information and the target information of the second sample, the second distortion loss value is calculated. The first distortion loss value and the second distortion loss value are superimposed to obtain the second reconstruction loss value.
[0036] Based on the first and second reconstruction loss values, the decoder parameters are adjusted to obtain the trained decoder.
[0037] In some embodiments, the method further includes:
[0038] Based on the preset first balance coefficient, the result of superimposing the time domain loss value and the frequency domain loss value is updated to obtain the first updated loss value. The time domain loss value and the frequency domain loss value are obtained by the terminal when training the encoder that generates the initial audio QR code.
[0039] Based on a preset second balance coefficient, the result of superimposing the first reconstruction loss value and the second reconstruction loss value is updated to obtain the second updated loss value.
[0040] The total loss value is obtained by summing the first update loss value and the second update loss value.
[0041] Based on the total loss value, the encoder parameters of the encoder and the decoder parameters of the decoder are jointly adjusted to obtain the trained encoder and decoder.
[0042] To achieve the above objectives, a second aspect of this application provides an information transmission device based on an audio QR code, the device comprising:
[0043] The acquisition module is used to acquire the preset initial audio and the target information to be transmitted;
[0044] A separate encoding module is used to encode the initial audio and target information separately to obtain audio encoding features and information encoding features;
[0045] The hybrid encoding module is used to hybrid encode audio encoding features and information encoding features to obtain the initial audio QR code;
[0046] The target information determination module is used to send an initial audio QR code to the client so that the client can receive the initial audio QR code sent by the terminal, configure multiple types of distortion audio for the initial audio QR code, introduce multiple types of distortion audio into the initial audio QR code to obtain an enhanced audio QR code, and decode the enhanced audio QR code to obtain the target information.
[0047] To achieve the above objectives, a third aspect of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method of the first aspect described above.
[0048] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of the first aspect described above.
[0049] The information transmission method, apparatus, device, and medium for audio QR codes proposed in this application, when applied to a terminal, communicate with a client. They acquire a preset initial audio and target information to be transmitted; separately encode the initial audio and target information to obtain audio encoding features and information encoding features; perform mixed encoding of the audio encoding features and information encoding features to obtain an initial audio QR code; send the initial audio QR code to the client so that the client receives the initial audio QR code sent by the terminal; configure multiple types of distortion audio for the initial audio QR code; introduce multiple types of distortion audio into the initial audio QR code to obtain an enhanced audio QR code; decode the enhanced audio QR code to obtain the target information. It can be understood that by introducing multiple types of distortion audio into the received initial audio QR code, the client can remove signals from the audio signal of the initial audio QR code that are the same as or similar to the introduced distortion audio, thereby achieving the purpose of enhancing the audio signal of the initial audio QR code; and then decoding the subsequently obtained enhanced audio QR code to obtain target information with high accuracy. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of an optional application scenario of the information transmission device based on audio QR codes provided in this application embodiment;
[0051] Figure 2 This is an optional flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application;
[0052] Figure 3 This is an optional encoding flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application;
[0053] Figure 4 This is an optional audio encoder structure diagram of the information transmission method based on audio QR codes provided in the embodiments of this application;
[0054] Figure 5 This is another optional flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application;
[0055] Figure 6 This is another optional flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application;
[0056] Figure 7 This is an optional decoding flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application;
[0057] Figure 8 yes Figure 6 A flowchart of the implementation of step 302 in the process;
[0058] Figure 9 yes Figure 6 A flowchart of the implementation of step 303 in the process;
[0059] Figure 10 This is another optional flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application;
[0060] Figure 11 This is another optional flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application;
[0061] Figure 12 This is an optional flowchart of an information transmission device based on audio QR codes provided in the embodiments of this application;
[0062] Figure 13 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0064] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0066] An audio QR code is an audio message carrying QR code information, designed to enable the rapid transmission of information such as text, images, or files using sound.
[0067] In related technologies, target information is transmitted by reading the audio carrying QR code information from an audio QR code. However, noise generated in different reading scenarios can easily interfere with information transmission, causing distortion of the audio QR code and ultimately resulting in inaccurate target information obtained from reading the audio QR code.
[0068] Based on this, embodiments of this application provide an information transmission method, apparatus, device, and medium based on audio QR codes, aiming to improve the accuracy of transmitting target information based on audio QR codes.
[0069] It should be noted that in the embodiments of this application, when information related to user characteristics, such as basic user information or user identity, is required, the user's permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain sensitive personal information of a user, the user's individual permission or consent will be obtained first. Only after obtaining the user's individual permission or consent will the necessary data for the normal operation of the embodiments of this application be obtained. For example, when the target information to be transmitted is proposed in the embodiments of this application, the relevant user's authorization or consent will be obtained first; otherwise, the target information cannot be used in the embodiments of this application. Furthermore, other relevant data obtained by the information transmission device based on audio QR codes proposed in this application (hereinafter referred to as "information transmission device" for ease of description) are all legal data obtained after obtaining authorization or consent, and will not be elaborated upon here.
[0070] This application provides an embodiment of an information transmission method, apparatus, device, and medium based on audio QR codes, which is specifically described through the following embodiments. First, the application scenario of the information transmission apparatus in this application embodiment is described, such as... Figure 1 As shown, Figure 1 This is a schematic diagram of an optional application scenario of the information transmission device based on audio QR codes provided in this application embodiment. The information transmission device can be applied to a terminal or a client. In one example, the terminal is a computer and the client is a mobile phone. The terminal receives the target information to be transmitted and the initial audio used to embed the target information, and encodes the initial audio and the target information to generate an initial audio QR code. The initial audio QR code can be propagated through broadcasting or other means and received by the client. The client decodes the received initial audio QR code. In this way, the target information can be transmitted quickly without manually scanning the QR code with a mobile phone.
[0071] Having understood the example application scenarios of the information transmission device proposed in this application, in order to further understand the beneficial effects of the information transmission method based on audio QR codes (which can also be simply referred to as "information transmission method" for ease of description) proposed in the embodiments of this application on the information transmission device, the information transmission method proposed in the embodiments of this application will be described below.
[0072] In this embodiment, the description will focus on the information transmission device, which is applied to a terminal and communicates with a client. The information transmission device can be integrated into a computer device, such as a server. Figure 2 As shown, Figure 2 This is an optional flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application. Figure 2 The method may include, but is not limited to, the following steps 101 to 104. When the information transmission device executes the prediction method, the specific process is as follows. It should be noted first that this embodiment... Figure 2 The order of steps 101 to 104 is not specifically limited. The order of steps can be adjusted or some steps can be reduced or added according to actual needs.
[0073] Step 101: Obtain the preset initial audio and target information to be transmitted.
[0074] Step 101 will be described in detail below.
[0075] The target information to be transmitted refers to data in the form of text, images, or files that need to be transmitted to the client; or, the target information can also be QR code information in binary format converted from the data in the form of text, images, or files to be transmitted.
[0076] The initial audio is a pre-prepared audio segment before the information transmission device performs encoding processing. The initial audio has specific attributes (such as length, frequency, sound quality, etc.). This initial audio is usually imperceptible to the human ear. The initial audio will serve as a carrier for embedding the target information to be transmitted and will be encoded together with the target information.
[0077] Furthermore, since the information transmission device is located in the terminal, the acquired initial audio and the target information to be transmitted will be received and encoded by the terminal. The terminal can be a device capable of transmitting audio signals, such as a computer, server, or smart speaker. Of course, the selected terminal can be adapted to different application scenarios, and this embodiment does not impose any limitations on this.
[0078] Furthermore, the number of target information to be transmitted can be one or more. When there are multiple target information, the terminal can encode multiple target information into the initial audio at the same time, so that when the client receives and parses the corresponding initial audio QR code, the multiple target information can be presented in sequence.
[0079] Step 102: Encode the initial audio and target information separately to obtain audio coding features and information coding features.
[0080] Step 102 is described in detail below.
[0081] In some embodiments, such as Figure 3 As shown, Figure 3 This is an optional encoding flowchart of the information transmission method based on audio QR codes provided in this application embodiment. After obtaining the initial audio and target information, the initial audio can be input into a pre-trained encoder, which encodes the initial audio and target information. The encoder includes an audio encoder, an information encoder, and a hybrid encoder. The audio encoder encodes the initial audio to obtain audio encoding features; the information encoder encodes the target information to obtain information encoding features. In this way, the terminal can use different encoding methods to encode the initial audio and target information under different data formats separately. Furthermore, the hybrid encoding makes the encoded result structurally more compact, increasing the difficulty for malicious attackers to disrupt the encoding process, thereby improving the security of the encoded features.
[0082] Furthermore, the audio encoder can be a Pulse Code Modulation Encoder (PCM), an Advanced Audio Coding Encoder (AAC), etc. Of course, the audio encoder can also be selected according to the actual situation, and this application embodiment does not limit this. The information encoder can be a QR code encoder based on a convolutional neural network, a QR code encoder based on an autoencoder, a Quick Response Code encoder, etc. The information encoder can also be selected according to the actual situation, and this application embodiment does not limit this.
[0083] like Figure 4 As shown, Figure 4 This is an optional audio encoder structure diagram of the information transmission method based on audio QR codes provided in this application embodiment. The audio encoder consists of five one-dimensional convolutional layers (Conv1D), each followed by a Leaky ReLU activation function. After the multiple sequentially connected convolutional layers and Leaky ReLUs, a flattening layer and a multilayer perceptron (MLP) are also provided. The functions of each layer in the audio encoder are as follows:
[0084] (1) One-dimensional convolutional layer: Used to process one-dimensional data (such as time series data or audio signals), extracting local features from the data through convolution operations. Each convolutional kernel (or filter) slides on the input data, calculating the dot product between the input data and the convolutional kernel, thereby generating a feature map that captures information from different frequencies and locations in the input data. Figure 4 The numbers in parentheses after Conv1D are adjustable parameters for the convolutional layer, allowing the convolutional layer to process the input data based on pre-defined parameters.
[0085] (2) Leaky ReLU: Leaky ReLU is a non-linear activation function used to increase the non-linear expressive power of audio encoders. Compared with the traditional ReLU activation function, Leaky ReLU allows a small gradient (i.e. "leak") to pass through when the input is less than 0, which helps to solve the problem of neuron death caused by the input being less than 0. Multiple convolutional layers and the Leaky ReLU are connected in sequence to form a hierarchical feature extraction structure. Each convolutional layer can further extract higher-level features from the features extracted by the previous layer, thereby capturing more detailed features in the data.
[0086] (3) Flattening layer: used to convert multidimensional input data (such as feature maps output by convolutional layers) into one-dimensional vectors to obtain data that meets the input requirements of multilayer perceptron, so that the multilayer perceptron can process the converted one-dimensional vectors in the future.
[0087] (4) Multilayer perceptron: MLP is used to further learn the latent representation of data from the features extracted from the flattened layer. Then, the MLP layer can capture global information in the data and map it to the output space.
[0088] Furthermore, the audio encoder can utilize weight normalization techniques to decouple the magnitude of the weight tensor from the orientation of each one-dimensional convolutional layer, thereby accelerating the audio encoder's processing of input data. Ultimately, the audio encoder outputs audio encoded features.
[0089] Furthermore, the encoder structure of the information encoder is similar to that of the audio encoder, and will not be described in detail here. In one example, the initial audio x is input into the terminal's preset audio encoder, and the target information z (in this example, the acquired target information is set to QR code information) is input into the terminal's preset information encoder. Here, the dimensions of x and z are B·L and B·D, respectively, where L represents the length of the initial audio fill waveform, D represents the length of the target information, and B is the batch size. B, L, and D can be adaptively adjusted according to actual conditions. Next, the audio encoder encodes the initial audio x to obtain the audio encoding feature f. x The information encoder encodes the target information z to obtain the audio coding features f.x Information encoding features f with the same feature dimension z .
[0090] Step 103: Mix and encode the audio encoding features and information encoding features to obtain the initial audio QR code.
[0091] Step 103 will be described in detail below.
[0092] In some embodiments, such as Figure 3 As shown, the terminal can also be equipped with a hybrid encoder, which is connected to the audio encoder and the information encoder. When the audio encoder and the information encoder output the encoding results, the hybrid encoder can process the audio encoding features and the information encoding features (f) x +f z Hybrid encoding is performed to map the mixed representation of audio coding features and information coding features onto waveform perturbation, resulting in an initial audio QR code. The specific processing procedure is as follows: <1> As shown:
[0093] x'=EF(EA(x)⊕EQR(z)) <1>
[0094] Where x represents the initial audio; z represents the target information; x' represents the initial audio QR code output by the encoder; EA represents the audio encoder; EQR represents the information encoder; EF represents the hybrid encoder; and ⊕ represents the connection operation.
[0095] Furthermore, the hybrid encoder includes multiple coding processing structure layers, each of which includes a transposed convolutional layer (ConvTranspose). The transposed convolutional layer is used to upsample and restore the input feature map until the processed audio waveform δ is obtained. x When the audio waveform has the same dimension as the initial audio x, output the initial audio QR code.
[0096] Furthermore, to increase the encoder's versatility and adapt to target information of different dimensions, a linear projection layer can be added to the information encoder to map target information of different dimensions onto a unified feature dimension.
[0097] Step 104: Send an initial audio QR code to the client so that the client can receive the initial audio QR code sent by the terminal. Configure multiple types of distortion audio for the initial audio QR code and introduce multiple types of distortion audio into the initial audio QR code to obtain an enhanced audio QR code. Decode the enhanced audio QR code to obtain the target information.
[0098] Step 104 will be described in detail below.
[0099] The client can be a mobile communication device, such as a mobile phone, computer, or tablet. The terminal sends the initial audio QR code to the client by playing an audio signal carrying QR code information. Thus, the client can easily receive the audio QR code with the audio receiver turned on. For example, when a user wants to obtain target information (such as QR code information) using their mobile phone, the user can transmit information without scanning the QR code, greatly facilitating the daily life of visually impaired people, or enabling rapid information transmission in dimly lit environments.
[0100] Furthermore, during the process of the terminal sending the initial audio QR code to the client, it is subject to interference from the surrounding environment, causing other noises, such as but not limited to unwanted human voices, water sounds, and horn sounds, to be mixed into the audio signal of the initial audio QR code. To improve the accuracy of the target information ultimately decoded by the client, this application embodiment configures multiple types of distorted audio for the initial audio QR code. Distorted audio refers to noise audio in a preset scenario. By introducing multiple types of distorted audio into the received initial audio QR code, the client can remove signals from the initial audio QR code that are the same as or similar to the introduced distorted audio, thereby enhancing the initial audio QR code. This enhanced audio QR code can then be decoded to obtain target information with higher accuracy.
[0101] In some embodiments, such as Figure 5 As shown, Figure 5 This is another optional flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application. The initial audio QR code is output by a pre-trained encoder, which is trained through the following steps 201 to 205:
[0102] Step 201: Obtain the preset initial audio of the sample and the target information of the sample to be transmitted.
[0103] Step 202: Encode the initial audio of the sample and the target information of the sample separately to obtain the sample audio encoding features and the sample information encoding features.
[0104] Step 203: Mix the sample audio encoding features and sample information encoding features to obtain the initial audio QR code of the sample.
[0105] Step 204: Based on the time domain difference between the initial audio sample and the initial audio sample QR code, calculate the time domain loss value, and based on the frequency domain difference between the initial audio sample and the initial audio sample QR code, calculate the frequency domain loss value.
[0106] Step 205: Adjust the encoder parameters based on the time-domain loss value and the frequency-domain loss value to obtain the trained encoder.
[0107] Steps 201 to 205 are described in detail below.
[0108] In some embodiments, the encoder in the information transmission device needs to be trained before it is formally deployed to the terminal to ensure its accuracy when running on the terminal. Steps 101 to 103 for training the encoder are similar and will not be described again here. In particular, the initial audio sample and the target information of the sample to be transmitted can be obtained by manual input by the user or automatically imported from an open source database; this application does not impose specific restrictions on them.
[0109] Furthermore, after obtaining the initial audio QR code of the sample, the following formula is used... <2> The time-domain loss value L is calculated. ti :
[0110]
[0111] in, The L1 norm is the sum of the absolute values of all elements in a vector; x represents the initial audio sample; x ′ This represents the initial audio QR code output by the encoder; the following formula... <3> to <5> In the formula and <2> Parameters with the same characters have the same meaning, which will not be elaborated further.
[0112] Furthermore, through the following formula <3> The frequency domain loss value L is calculated. fi :
[0113]
[0114] Where φ(·) represents a preset function that converts the waveform into a spectrum.
[0115] Furthermore, after obtaining the temporal and frequency domain loss values, the encoder parameters are adjusted based on these two loss values. The encoder parameters can be network structure parameters, such as the number of encoder layers, the number of neurons, and activation functions; they can also be optimization parameters, such as the learning rate, batch size, and regularization parameters; or they can be feature parameters, such as feature vectors representing sparsity and feature vectors representing quantization accuracy in the encoder. Of course, the specific encoder parameters to be adjusted can be set according to the actual situation, and this application embodiment does not impose any restrictions on this.
[0116] Furthermore, since the encoder includes an audio encoder, an information encoder, and a hybrid encoder, the encoder parameters that are adjusted can be any one of the audio encoder, the information encoder, and the hybrid encoder, or the encoder parameters of the audio encoder, the information encoder, and the hybrid encoder can be adjusted together.
[0117] In another embodiment of this application, the description will focus on an information transmission device applied to a client, which is connected to a terminal for communication. The information transmission device can also be integrated into a computer device, such as a server. Figure 6 As shown, Figure 6 This is another optional flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application. Figure 6 The method may include, but is not limited to, the following steps 301 to 303. When the information transmission device executes the prediction method, the specific process is as follows. It should be noted first that this embodiment... Figure 6 The order of steps 301 to 303 is not specifically limited. The order of steps can be adjusted or some steps can be reduced or added according to actual needs.
[0118] Step 301: Receive the initial audio QR code sent by the receiving terminal, wherein the initial audio QR code is obtained by mixing audio encoding features and information encoding features; the audio encoding features and information encoding features are obtained by separately encoding the acquired preset initial audio and the target information to be transmitted.
[0119] In this embodiment, a communication connection is established before the client and terminal transmit information, so that the client and terminal can transmit information via audio QR codes. The communication connection between the client and terminal can be achieved through various methods, including but not limited to client-server (C / S) mode connection, communication protocol-based connection, Bluetooth connection, etc. This application embodiment does not limit the specific method of communication connection, and can be adapted to the actual situation.
[0120] Step 302: Configure multiple types of distorted audio for the initial audio QR code, and introduce multiple types of distorted audio into the initial audio QR code to obtain an enhanced audio QR code.
[0121] Step 303: Decode the enhanced audio QR code to obtain the target information.
[0122] Steps 301 to 303 are described in detail below.
[0123] In some embodiments, such as Figure 7 As shown, Figure 7This is an optional decoding flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application. When the decoder receives the initial audio QR code as input, it introduces a variety of distorted audio signals so that the client can remove signals that are the same as or similar to the introduced distorted audio signals from the audio signal of the initial audio QR code, thereby enhancing the initial audio QR code so that the enhanced audio QR code can be decoded later to obtain target information with higher accuracy. That is, the robustness of the information transmission device is improved.
[0124] For example, when a visually impaired person uses their mobile phone to receive an audio signal carrying QR code information played at a coffee shop counter, they may be affected by surrounding noise, machine noise, and music played in the shop. This application embodiment introduces distorted audio to denoise the audio signal of the initial audio QR code, so that the visually impaired person's mobile phone (client) can quickly and accurately identify the target information that the counter (terminal) wants to transmit based on the denoised audio signal.
[0125] The types of distorted audio can be: ① environmental background distorted audio (such as animal sounds, urban traffic noise, machine engine sounds, and external rain sounds); ② room impulse response distorted audio in different locations (such as shopping malls, living rooms, kitchens, coffee shops, offices, and hospitals); ③ music background distorted audio; ④ Gaussian noise distorted audio, etc.
[0126] It should be noted that the audio type of the distorted audio can be adapted to the actual situation. For example, when the location of the target information being transmitted is a school playground, the audio type of the distorted audio can also be reading aloud, birdsong, etc. This application embodiment does not impose specific limitations on this.
[0127] Furthermore, based on the following formula <4> and <5> Complete the enhancement and decoding of the initial audio QR code, and obtain the target information z′. aug :
[0128] δ aug =δ e +δ r +δ m +δ g <4>
[0129] z′ aug =D mp (x′+δ aug ) <5>
[0130] Where, δ e Indicates ambient background distortion audio, δ r Indicates room impulse response distortion audio, δ m Indicates distorted background audio, δ g D represents Gaussian noise-distorted audio;mp This is a pre-trained decoder.
[0131] In some embodiments, such as Figure 8 As shown, Figure 8 yes Figure 6 A flowchart of step 302 in the above steps describes the process of introducing multiple distorted audio values into the initial audio QR code to obtain an enhanced audio QR code, including the following steps 401 to 403:
[0132] Step 401: Obtain the ambient audio surrounding the receiving terminal when it sends the initial audio.
[0133] Step 402: Select at least one target distortion audio that matches the scene audio from the configured multiple distortion audios, and introduce at least one target distortion audio into the initial audio QR code to obtain an enhanced audio QR code.
[0134] Step 403, or, assign adjustable probability weights to each distorted audio, randomly select at least one target distorted audio from multiple distorted audios based on the probability weights, introduce at least one target distorted audio into the initial audio QR code, and obtain an enhanced audio QR code.
[0135] Steps 401 to 403 are described in detail below.
[0136] In some embodiments, to improve the efficiency of the decoder in generating enhanced audio QR codes, the ambient scene audio surrounding the client receiving terminal when the initial audio is sent can be obtained, and at least one target distorted audio that matches the ambient scene audio can be selected from multiple distorted audios to enhance the initial audio QR code.
[0137] For example, the distorted audio includes ambient background distortion audio, room impulse response distortion audio, background music distortion audio, and Gaussian noise distortion audio. When a visually impaired person uses their mobile phone to receive the initial audio QR code sent by the coffee shop counter, the coffee shop is playing music, and the background music interferes with the audio signal transmission of the initial audio QR code. At this time, the acquired scene audio includes music noise, so the information transmission device can determine that the background music distortion audio is the target distorted audio that matches the scene audio, and introduce the target distorted audio into the initial audio QR code to obtain an enhanced audio QR code. Of course, when multiple distorted audios matching the scene audio are identified, the number of target distorted audios is multiple.
[0138] In other embodiments, the scene audio obtained in practical applications typically includes multiple types and quantities of noise. In this case, multiple target distortion audios that match the scene audio can be selected from a configured set of distortion audios. These target distortion audios are then introduced into the initial audio QR code to obtain an enhanced audio QR code. Specifically, this is achieved through the following formula: <6> Select the target distortion audio δ from the configured multiple distortion audio samples. aug :
[0139]
[0140] in, Indicates the probability weight p e For distorted audio δ e Sampling is performed, and similarly... and Use a similar sampling method.
[0141] Furthermore, the probability weights p can be adaptively adjusted according to the actual situation. For example, the information transmission device can receive scene audio over a preset period of time and readjust the probability weights based on the duration of occurrence of each distorted audio element in the scene audio to obtain an updated formula. <6> Based on the updated formula <6> Randomly select at least one target distorted audio from multiple distorted audio samples; then, based on the formula... <5> A similar approach is used to enhance the initial audio QR code after introducing the target distorted audio, resulting in an enhanced audio QR code.
[0142] Understandably, randomly cascading different distorted audio signals can better remove noise from the real world, thereby enhancing the robustness of the information transmission device and improving its generalization ability in dealing with different initial audio QR codes.
[0143] In some embodiments, such as Figure 9 As shown, Figure 9 yes Figure 6 A flowchart of step 303 in the above steps, which decodes the enhanced audio QR code to obtain the target information, includes the following steps 501 to 504:
[0144] Step 501: Extract at least one sine wave signal from the enhanced audio QR code.
[0145] Step 502: Select one from multiple preset decoders that matches the periodic characteristics of the sine signal, decode the enhanced audio QR code, and obtain the target information.
[0146] Step 503: When there are multiple sinusoidal signals, the audio QR code is decoded by multiple decoders that match the periodic characteristics of each sinusoidal signal to obtain multiple sub-target information.
[0147] Step 504: The information of multiple sub-targets is fused to obtain the target information.
[0148] Steps 501 to 504 are described in detail below.
[0149] In some embodiments, since the audio signal of the enhanced audio QR code is composed of multiple sinusoidal signals with different frequencies, amplitudes, and phases superimposed, by first parsing the audio signal of the enhanced audio QR code to extract at least one sinusoidal signal, and then selecting one from multiple decoders that matches the periodic features of the extracted sinusoidal signal for decoding to obtain the corresponding target information, the accuracy and efficiency of decoding can be improved. Specifically, the process of decoding based on a pre-trained decoder to obtain the target information z′ is as follows: <7> As shown:
[0150] z′=D mp (x′) <7>
[0151] Where x′ represents the enhanced audio QR code; D mp This is a pre-trained decoder.
[0152] Furthermore, periodicity refers to the time period required for a sinusoidal signal to repeat a complete sine wave; that is, the time required for a certain phase point in a sinusoidal signal to reach the next identical phase point. The sinusoidal signals of different sound signals are usually different.
[0153] Furthermore, sine signals can be separated from the audio signal of the enhanced audio QR code using methods such as filtering and Fourier transform; correspondingly, the decoder for decoding the enhanced audio QR code can be a filtering decoder, a Fourier transform-based decoder, etc.; of course, the method for extracting sine signals and the decoder used for decoding can be adjusted according to the actual situation, and the embodiments of this application do not limit this.
[0154] In some embodiments, the pre-trained decoder consists of multiple sub-decoders, each of which consists of a set of hierarchical convolutional layers with Leaky ReLU activation. When the audio signal of the enhanced audio QR code includes multiple sine signals, each sub-decoder will process the audio containing different sine signals to obtain multiple sub-target information. Then, the multiple sub-target information will be stacked and subjected to mean pooling by a preset activation function (such as the sigmoid function) to finally obtain the fused and reconstructed target information z′.
[0155] In some embodiments, such as Figure 10 As shown, Figure 10This is another optional flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application. The target information is output by a pre-trained decoder, which is trained through the following steps 601 to 607:
[0156] Step 601: Receive the initial sample audio QR code sent by the receiving terminal. The initial sample audio QR code is obtained by hybrid encoding of sample audio encoding features and sample information encoding features. The sample audio encoding features and sample information encoding features are obtained by separately encoding the acquired preset initial sample audio and the target information of the sample to be transmitted.
[0157] Step 602: Decode the initial audio QR code of the sample to obtain the target information of the first sample.
[0158] Step 603: Configure various types of sample distortion audio for the initial sample audio QR code, and introduce various types of sample distortion audio into the initial sample audio QR code to obtain the sample enhanced audio QR code.
[0159] Step 604: Decode the sample-enhanced audio QR code to obtain the second sample target information.
[0160] Step 605: Based on the initial sample information and the first sample target information, calculate the first reconstruction loss value, wherein the initial sample information is used to generate the initial audio QR code of the sample.
[0161] Step 606: Calculate the first distortion loss value based on the initial sample information and the second sample target information, and calculate the second distortion loss value based on the first sample target information and the second sample target information. Superimpose the first distortion loss value and the second distortion loss value to obtain the second reconstruction loss value.
[0162] Step 607: Adjust the decoder parameters of the decoder according to the first reconstruction loss value and the second reconstruction loss value to obtain the trained decoder.
[0163] Steps 601 to 607 are described in detail below.
[0164] In some embodiments, the decoder in the information transmission device needs to be trained before it is formally deployed to the client to ensure its accuracy when running on the client. Steps 601 to 604 of training the decoder are similar to steps 301 to 303, and will not be described again here. In particular, to enhance the robustness of the decoder, the corresponding output results, namely the first sample target information and the second sample target information, are obtained by introducing distorted audio and not introducing distorted audio (step 602) to determine the loss value for adjusting the decoder parameters.
[0165] Furthermore, without introducing distorted audio, through the following formula <8> The first reconstruction loss value L is calculated. v :
[0166]
[0167] Among them, D mp For a pre-built decoder, x ′ z1 represents the initial audio QR code for the sample, and z1 represents the target information of the first sample. It is an L1 norm.
[0168] Furthermore, in the case of introducing distorted audio, the following formula is used: <9> The second reconstruction loss value L is calculated. a :
[0169]
[0170] Among them, D mp For a pre-built decoder, x ′ +δ aug To enhance the initial audio QR code, z2 represents the target information of the second sample. It is an L1 norm; This is the first distortion loss value. This is the second distortion loss value.
[0171] Furthermore, after obtaining the first reconstruction loss value and the second reconstruction loss value, the decoder parameters of the decoder are adjusted based on these two loss values. The decoder parameters can be network structure parameters, such as the number of decoder layers, the number of neurons, activation functions, etc.; they can also be optimization parameters, such as learning rate, batch size, and regularization parameters; or they can be feature parameters, such as feature vectors representing sparsity and feature vectors representing quantization accuracy in the decoder. Of course, the specific decoder parameters to be adjusted can be set according to the actual situation, and this application embodiment does not limit this.
[0172] Furthermore, different loss weights can be set for the first reconstruction loss value and the second reconstruction loss value, and the first reconstruction loss value and the second reconstruction loss value can be updated according to the loss weights. Then, the decoder parameters of the decoder can be adjusted based on the updated first reconstruction loss value and the second reconstruction loss value to obtain the trained decoder.
[0173] In some embodiments, such as Figure 11 As shown, Figure 11 This is another optional flowchart of the information transmission method based on audio QR codes provided in the embodiments of this application. The information transmission method further includes the following steps 701 to 704:
[0174] Step 701: Based on the preset first balance coefficient, update the result of superimposing the time domain loss value and the frequency domain loss value to obtain the first updated loss value. The time domain loss value and the frequency domain loss value are obtained by the terminal when training the encoder that generates the initial audio QR code.
[0175] Step 702: Based on the preset second balance coefficient, update the result of the superposition of the first reconstruction loss value and the second reconstruction loss value to obtain the second updated loss value.
[0176] Step 703: Add the first update loss value and the second update loss value to obtain the total loss value.
[0177] Step 704: Based on the total loss value, jointly adjust the encoder parameters of the encoder and the decoder parameters of the decoder to obtain the trained encoder and decoder.
[0178] Steps 701 to 704 are described in detail below.
[0179] In some embodiments, by the following formula <10> The total loss value L was calculated. total :
[0180] L total =λ imp (L ti +L fi )+λ mr (L v +L a ) <10>
[0181] Where, λ imp λ is the first balance coefficient. mr λ is the second balance coefficient. imp (L ti +L fi ) represents the first update loss value, λ mr (L v +L a The first and second balance coefficients can be adjusted adaptively according to the actual situation.
[0182] Thus, by jointly adjusting the correlated encoder and decoder parameters based on the total loss value, collaborative optimization of the encoder and decoder can be achieved, thereby improving the performance of the information transmission devices deployed on the terminal and client respectively. For example, the encoder and decoder parameters can be updated simultaneously in each iteration of the training process using gradient descent.
[0183] like Figure 12 As shown, Figure 12This is an optional flowchart of an information transmission device based on audio QR codes provided in this application embodiment. The information transmission device includes the following modules 801 to 804:
[0184] The acquisition module 801 is used to acquire the preset initial audio and the target information to be transmitted.
[0185] The separate encoding module 802 is used to encode the initial audio and target information separately to obtain audio encoding features and information encoding features.
[0186] The hybrid encoding module 803 is used to perform hybrid encoding of audio encoding features and information encoding features to obtain the initial audio QR code.
[0187] The target information determination module 804 is used to send an initial audio QR code to the client so that the client can receive the initial audio QR code sent by the terminal, configure multiple types of distortion audio for the initial audio QR code, introduce multiple types of distortion audio into the initial audio QR code to obtain an enhanced audio QR code, and decode the enhanced audio QR code to obtain the target information.
[0188] In some embodiments, the target information to be transmitted refers to data in the form of text, images, or files that need to be transmitted to the client; or, the target information may also be QR code information in binary format converted from the data in the form of text, images, or files to be transmitted.
[0189] The initial audio is a pre-prepared audio segment before the information transmission device performs encoding processing. The initial audio has specific attributes (such as length, frequency, sound quality, etc.). This initial audio is usually imperceptible to the human ear. The initial audio will serve as a carrier for embedding the target information to be transmitted and will be encoded together with the target information.
[0190] Furthermore, since the information transmission device is located in the terminal, the acquired initial audio and the target information to be transmitted will be received and encoded by the terminal. The terminal can be a device capable of transmitting audio signals, such as a computer, server, or smart speaker. Of course, the selected terminal can be adapted to different application scenarios, and this embodiment does not impose any limitations on this.
[0191] Furthermore, the number of target information to be transmitted can be one or more. When there are multiple target information, the terminal can encode multiple target information into the initial audio at the same time, so that when the client receives and parses the corresponding initial audio QR code, the multiple target information can be presented in sequence.
[0192] In some embodiments, after acquiring the initial audio and target information, the initial audio can be input into a pre-trained encoder, which encodes the initial audio and target information. The encoder includes an audio encoder, an information encoder, and a hybrid encoder. The audio encoder encodes the initial audio to obtain audio-coded features; the information encoder encodes the target information to obtain information-coded features. In this way, the terminal can use different encoding methods to encode the initial audio and target information under different data formats separately. Furthermore, the hybrid encoding makes the encoded result structurally more compact, increasing the difficulty for malicious attackers to disrupt the encoding process, thereby improving the security of the encoded features.
[0193] In some embodiments, the terminal may also be equipped with a hybrid encoder, which is connected to the audio encoder and the information encoder. When the audio encoder and the information encoder output the encoding result, the hybrid encoder can perform hybrid encoding of the audio encoding features and the information encoding features to map the hybrid representation of the audio encoding features and the information encoding features onto the waveform perturbation and obtain the initial audio QR code.
[0194] Furthermore, the hybrid encoder includes multiple encoding processing structure layers, each of which includes a transposed convolutional layer. The transposed convolutional layer is used to upsample and restore the input feature map until the processed audio waveform has the same dimension as the original audio waveform, at which point the original audio QR code is output.
[0195] Furthermore, to increase the encoder's versatility and adapt to target information of different dimensions, a linear projection layer can be added to the information encoder to map target information of different dimensions onto a unified feature dimension.
[0196] The client can be a mobile communication device, such as a mobile phone, computer, or tablet. The terminal sends the initial audio QR code to the client by playing an audio signal carrying QR code information. Thus, the client can easily receive the audio QR code with the audio receiver turned on. For example, when a user wants to obtain target information (such as QR code information) using their mobile phone, the user can transmit information without scanning the QR code, greatly facilitating the daily life of visually impaired people, or enabling rapid information transmission in dimly lit environments.
[0197] Furthermore, during the process of the terminal sending the initial audio QR code to the client, it is subject to interference from the surrounding environment, causing other noises, such as but not limited to unwanted human voices, water sounds, and horn sounds, to be mixed into the audio signal of the initial audio QR code. To improve the accuracy of the target information ultimately decoded by the client, this application embodiment configures multiple types of distorted audio for the initial audio QR code. Distorted audio refers to noise audio in a preset scenario. By introducing multiple types of distorted audio into the received initial audio QR code, the client can remove signals from the initial audio QR code that are the same as or similar to the introduced distorted audio, thereby enhancing the initial audio QR code. This enhanced audio QR code can then be decoded to obtain target information with higher accuracy.
[0198] The information transmission method, apparatus, device, and medium for audio QR codes proposed in this application, when applied to a terminal, communicate with a client. They acquire a preset initial audio and target information to be transmitted; separately encode the initial audio and target information to obtain audio encoding features and information encoding features; perform mixed encoding of the audio encoding features and information encoding features to obtain an initial audio QR code; send the initial audio QR code to the client so that the client receives the initial audio QR code sent by the terminal; configure multiple types of distortion audio for the initial audio QR code; introduce multiple types of distortion audio into the initial audio QR code to obtain an enhanced audio QR code; decode the enhanced audio QR code to obtain the target information. It can be understood that by introducing multiple types of distortion audio into the received initial audio QR code, the client can remove signals from the audio signal of the initial audio QR code that are the same as or similar to the introduced distortion audio, thereby achieving the purpose of enhancing the audio signal of the initial audio QR code; and then decoding the subsequently obtained enhanced audio QR code to obtain target information with high accuracy.
[0199] The specific implementation of this information transmission device is basically the same as the specific implementation of the information transmission method described above, and will not be repeated here.
[0200] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described information transmission method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0201] like Figure 13 As shown, Figure 13 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device includes:
[0202] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0203] The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the information transmission method of the embodiments of this application.
[0204] The input / output interface 903 is used to implement information input and output;
[0205] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0206] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);
[0207] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.
[0208] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described information transmission method.
[0209] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0210] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0211] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0212] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0213] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0214] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0215] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0216] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0217] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0218] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0219] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0220] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for information transmission based on audio QR codes, characterized in that, Applied to a terminal that communicates with a client, the method includes: Acquire a preset initial audio and target information to be transmitted, wherein the initial audio is the carrier of the target information; The initial audio and the target information are encoded separately to obtain audio encoding features and information encoding features; The audio encoding features and the information encoding features are mixed and encoded to obtain an initial audio QR code; The initial audio QR code is sent to the client so that the client receives the initial audio QR code sent by the terminal. Multiple types of distortion audio are configured for the initial audio QR code, and multiple types of distortion audio are introduced into the initial audio QR code to obtain an enhanced audio QR code. The enhanced audio QR code is decoded to obtain the target information. The process of introducing various distorted audio values into the initial audio QR code to obtain an enhanced audio QR code includes: Acquire the ambient audio surrounding the terminal when it sends the initial audio; Select at least one target distortion audio that matches the scene audio from a plurality of configured distortion audios, and introduce at least one target distortion audio into the initial audio QR code to obtain an enhanced audio QR code, wherein the distortion audio includes ambient background distortion audio, room impulse response distortion audio, music background distortion audio, and Gaussian noise distortion audio; Alternatively, an adjustable probability weight is assigned to each of the distorted audios, and at least one target distorted audio is randomly selected from the plurality of distorted audios based on the probability weight. At least one of the target distorted audios is introduced into the initial audio QR code to obtain an enhanced audio QR code. The initial audio QR code is output by a pre-trained encoder, which is trained through the following steps: Obtain the preset initial audio of the sample and the target information of the sample to be transmitted; The initial audio of the sample and the target information of the sample are encoded separately to obtain the sample audio encoding features and the sample information encoding features; The sample audio encoding features and the sample information encoding features are mixed and encoded to obtain the initial audio QR code of the sample. Based on the time domain difference between the initial audio sample and the initial audio sample QR code, a time domain loss value is calculated, and based on the frequency domain difference between the initial audio sample and the initial audio sample QR code, a frequency domain loss value is calculated. The encoder parameters are adjusted based on the time-domain loss value and the frequency-domain loss value to obtain the trained encoder.
2. A method for information transmission based on audio QR codes, characterized in that, Applied to a client, wherein the client communicates with a terminal, the method includes: The receiving terminal sends an initial audio QR code, wherein the initial audio QR code is obtained by mixing audio encoding features and information encoding features; the audio encoding features and the information encoding features are obtained by separately encoding the acquired preset initial audio and the target information to be transmitted, wherein the initial audio is the carrier of the target information; Configure the initial audio QR code with various types of distorted audio, and introduce various distorted audio into the initial audio QR code to obtain an enhanced audio QR code. The distorted audio includes ambient background distortion audio, room impulse response distortion audio, music background distortion audio, and Gaussian noise distortion audio. The enhanced audio QR code is decoded to obtain the target information; The process of introducing various distorted audio values into the initial audio QR code to obtain an enhanced audio QR code includes: Acquire the ambient audio surrounding the terminal when it sends the initial audio; Select at least one target distorted audio that matches the scene audio from a plurality of configured distorted audios, and introduce at least one of the target distorted audios into the initial audio QR code to obtain an enhanced audio QR code; Alternatively, an adjustable probability weight is assigned to each of the distorted audios, and at least one target distorted audio is randomly selected from the plurality of distorted audios based on the probability weight. At least one of the target distorted audios is introduced into the initial audio QR code to obtain an enhanced audio QR code. The initial audio QR code is output by a pre-trained encoder, which is trained through the following steps: Obtain the preset initial audio of the sample and the target information of the sample to be transmitted; The initial audio of the sample and the target information of the sample are encoded separately to obtain the sample audio encoding features and the sample information encoding features; The sample audio encoding features and the sample information encoding features are mixed and encoded to obtain the initial audio QR code of the sample. Based on the time domain difference between the initial audio sample and the initial audio sample QR code, a time domain loss value is calculated, and based on the frequency domain difference between the initial audio sample and the initial audio sample QR code, a frequency domain loss value is calculated. The encoder parameters are adjusted based on the time-domain loss value and the frequency-domain loss value to obtain the trained encoder.
3. The method according to claim 2, characterized in that, Decoding the enhanced audio QR code to obtain the target information includes: At least one sine signal is extracted from the enhanced audio QR code; Select one decoder from multiple preset decoders that matches the periodic characteristics of the sinusoidal signal, decode the enhanced audio QR code, and obtain the target information; When there are multiple sinusoidal signals, the audio QR code is decoded by multiple decoders that match the periodic characteristics of each sinusoidal signal to obtain multiple sub-target information. The target information is obtained by fusing multiple sub-target information.
4. The method according to claim 3, characterized in that, The target information is output by a pre-trained decoder, which is trained through the following steps: The terminal sends a sample initial audio QR code, wherein the sample initial audio QR code is obtained by hybrid encoding of sample audio encoding features and sample information encoding features; the sample audio encoding features and the sample information encoding features are obtained by separately encoding the acquired preset sample initial audio and the target information of the sample to be transmitted. The initial audio QR code of the sample is decoded to obtain the target information of the first sample; Configure various types of sample distortion audio for the initial sample audio QR code, and introduce various types of sample distortion audio into the initial sample audio QR code to obtain a sample enhanced audio QR code; The sample-enhanced audio QR code is decoded to obtain the second sample target information; Based on the initial sample information and the first sample target information, a first reconstruction loss value is calculated, wherein the initial sample information is used to generate the initial audio QR code of the sample; Based on the initial sample information and the target sample information, a first distortion loss value is calculated, and based on the first sample target information and the target sample information, a second distortion loss value is calculated. The first distortion loss value and the second distortion loss value are then superimposed to obtain a second reconstruction loss value. Based on the first reconstruction loss value and the second reconstruction loss value, the decoder parameters of the decoder are adjusted to obtain the trained decoder.
5. The method according to claim 4, characterized in that, The method further includes: Based on a preset first balance coefficient, the result of superimposing the time domain loss value and the frequency domain loss value is updated to obtain a first updated loss value, wherein the time domain loss value and the frequency domain loss value are obtained by the terminal when training the encoder that generates the initial audio QR code; Based on a preset second balance coefficient, the result of superimposing the first reconstruction loss value and the second reconstruction loss value is updated to obtain the second updated loss value. The total loss value is obtained by summing the first update loss value and the second update loss value. Based on the total loss value, the encoder parameters of the encoder and the decoder parameters of the decoder are jointly adjusted to obtain the trained decoder and encoder.
6. An information transmission device based on audio QR codes, characterized in that, The device includes: The acquisition module is used to acquire a preset initial audio and target information to be transmitted, wherein the initial audio is the carrier of the target information; A separate encoding module is used to encode the initial audio and the target information separately to obtain audio encoding features and information encoding features; A hybrid encoding module is used to hybrid encode the audio encoding features and the information encoding features to obtain an initial audio QR code; The target information determination module is used to send the initial audio QR code to the client so that the client receives the initial audio QR code sent by the terminal, configures multiple types of distortion audio for the initial audio QR code, introduces multiple types of distortion audio into the initial audio QR code to obtain an enhanced audio QR code, decodes the enhanced audio QR code to obtain the target information, wherein the distortion audio includes ambient background distortion audio, room impulse response distortion audio, music background distortion audio, and Gaussian noise distortion audio; The process of introducing various distorted audio values into the initial audio QR code to obtain an enhanced audio QR code includes: Acquire the ambient audio surrounding the terminal when it sends the initial audio; Select at least one target distorted audio that matches the scene audio from a plurality of configured distorted audios, and introduce at least one of the target distorted audios into the initial audio QR code to obtain an enhanced audio QR code; Alternatively, an adjustable probability weight is assigned to each of the distorted audios, and at least one target distorted audio is randomly selected from the plurality of distorted audios based on the probability weight. At least one of the target distorted audios is introduced into the initial audio QR code to obtain an enhanced audio QR code. The initial audio QR code is output by a pre-trained encoder, which is trained through the following steps: Obtain the preset initial audio of the sample and the target information of the sample to be transmitted; The initial audio of the sample and the target information of the sample are encoded separately to obtain the sample audio encoding features and the sample information encoding features; The sample audio encoding features and the sample information encoding features are mixed and encoded to obtain the initial audio QR code of the sample. Based on the time domain difference between the initial audio sample and the initial audio sample QR code, a time domain loss value is calculated, and based on the frequency domain difference between the initial audio sample and the initial audio sample QR code, a frequency domain loss value is calculated. The encoder parameters are adjusted based on the time-domain loss value and the frequency-domain loss value to obtain the trained encoder.
7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method of claim 1, or the method of any one of claims 2 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of claim 1, or the method of any one of claims 2 to 5.
Citation Information
Patent Citations
Audio frequency multi-code transmission method and corresponding device
CN103812824A
Speech enhancement system based on time modeling generative adversarial network
CN114495958A
Method for generating audio two-dimensional code and transmitting emergency information
CN117478239A