Data transmission methods, apparatus, electronic devices and computer-readable storage media

By acquiring and utilizing the types of environmental noise, the problem of susceptibility to interference in acoustic data transmission was solved, resulting in higher data transmission success rates and network transmission quality.

CN116189706BActive Publication Date: 2026-04-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing data transmission methods that transmit data via sound waves are easily affected by external environmental noise, causing the receiving terminal to be unable to accurately interpret the data, resulting in a low data transmission success rate.

Method used

The system acquires the audio characteristics of the current ambient noise, determines the type of ambient noise, adds it to the data to be transmitted, and converts it into audio data so that the receiving terminal can acquire the data according to the noise type.

Benefits of technology

By identifying and utilizing types of environmental noise, the success rate of data transmission is improved, enabling the receiving terminal to accurately extract the data to be transmitted and improving the quality of network transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189706B_ABST
    Figure CN116189706B_ABST
Patent Text Reader

Abstract

This invention discloses a data transmission method, apparatus, electronic device, and computer-readable storage medium. After acquiring the data to be transmitted and the current ambient noise, this invention extracts features from the current ambient noise to obtain its audio features. Then, based on the audio features, it determines the ambient noise type and adds this type to the data to be transmitted, obtaining the target data. The target data is then converted into audio data and played, allowing the receiving terminal to retrieve the data to be transmitted from the audio data according to the ambient noise type. This solution can improve the data transmission success rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and more specifically to a data transmission method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] In recent years, with the rapid development of internet technology, terminals can achieve fast data transmission through serial communication. However, if the serial cable is faulty or malfunctions, data transmission will fail. To address this issue, existing data transmission methods can send the data to be transmitted to the receiving terminal via sound waves.

[0003] In the process of researching and practicing existing technologies, the inventors of this invention discovered that transmitting data using sound waves is more susceptible to interference from external environmental noise, which makes the transmitted data obtained by the receiving terminal inaccurate or unable to be parsed from the collected audio, resulting in data transmission failure and thus a low data transmission success rate. Summary of the Invention

[0004] This invention provides a data transmission method, apparatus, electronic device, and computer-readable storage medium, which can improve the data transmission success rate.

[0005] A data transmission method, comprising:

[0006] Acquire the data to be transmitted and the current ambient noise, wherein the current ambient noise is the sound collected in real time under the current transmission environment;

[0007] The current environmental noise is subjected to feature extraction to obtain the audio features of the current environmental noise;

[0008] Based on the audio characteristics, the environmental noise type of the current environmental noise is determined;

[0009] The environmental noise type is added to the data to be transmitted to obtain the target transmission data;

[0010] The target transmission data is converted into audio data and played so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise.

[0011] Accordingly, embodiments of the present invention provide a data transmission device, including:

[0012] The acquisition unit is used to acquire the data to be transmitted and the current ambient noise, wherein the current ambient noise is the sound collected in real time under the current transmission environment;

[0013] An extraction unit is used to extract features from the current environmental noise to obtain the audio features of the current environmental noise;

[0014] The determining unit is configured to determine the type of ambient noise of the current ambient noise based on the audio features.

[0015] An adding unit is used to add the environmental noise type to the data to be transmitted to obtain the target transmission data;

[0016] The playback unit is used to convert the target transmission data into audio data and play the audio data so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise.

[0017] Optionally, in one embodiment, the data transmission device further includes an identification unit, which is specifically used to record the target audio data being played when audio data is detected to obtain recording data; to parse the recording data to obtain the current ambient noise type and initial audio data; and to identify the transmission data in the initial audio data according to the current ambient noise type.

[0018] Optionally, in some embodiments, the identification unit may be specifically used to evaluate the network transmission quality of the target audio data according to the current ambient noise type to obtain evaluation information; based on the evaluation information, determine the processing method for the target audio data; when the processing method is to receive the target audio data, identify the transmission data in the initial audio data.

[0019] Optionally, in some embodiments, the identification unit may be specifically used to extract baseband data from the initial audio data, and perform message header detection on the baseband data according to the data block identifier in the baseband data; when a message header exists in the baseband data, extract acoustic wave baseband data from the initial audio data; extract at least one data block identifier from the acoustic wave baseband data, and perform error correction on the extracted data block identifier to obtain a set of data block identifiers; obtain the target data block corresponding to each data block identifier in the set of data block identifiers, and fuse the target data blocks to obtain transmission data.

[0020] Optionally, in some embodiments, the identification unit may be specifically used to discard the initial audio data and generate a prompt message when the processing method is not to receive the target audio data; and send the prompt message to the sending terminal so that the sending terminal can replay the target audio data based on the prompt message.

[0021] Optionally, in some embodiments, the extraction unit may be specifically used to convert the current ambient noise channel to obtain the current ambient noise of the target channel; divide the current ambient noise of the target channel into blocks to obtain multiple audio data blocks; and extract features from the audio data blocks to obtain the audio features corresponding to the audio data blocks.

[0022] Optionally, in some embodiments, the extraction unit may be specifically used to extract spectral image information from the audio data block and identify the spectral image value corresponding to the audio data block in the spectral image information; obtain the data sampling amount corresponding to each audio feature dimension, and extract the initial audio feature corresponding to each audio feature dimension in the spectral image information based on the data sampling amount; and concatenate the initial audio feature with the spectral image value to obtain the audio feature corresponding to each audio data block.

[0023] Optionally, in some embodiments, the determining unit may be specifically used to identify candidate environmental noise types of corresponding audio data blocks in the audio features using a trained noise recognition model; count the number of types corresponding to each environmental noise type in the candidate environmental noise types; and filter out the environmental noise type of the current environmental noise from the candidate environmental noise types based on the number of types.

[0024] Optionally, in some embodiments, the determining unit may be specifically used to extract noise audio features from the audio features of each audio data block using a trained noise type recognition model to obtain a noise audio feature set of the current environmental noise; perform batch normalization processing on the noise audio features in the noise audio feature set to obtain a basic noise audio feature set; and identify the environmental noise type corresponding to each basic noise audio feature in the basic noise audio feature set to obtain the candidate environmental noise type of the audio data block.

[0025] Optionally, in some embodiments, the determining unit may be specifically used to obtain the time information of the audio data block in the current ambient noise, and sort the basic noise audio features in the basic noise audio feature set based on the time information; based on the sorting information, convert the basic noise audio features in the basic noise audio feature set into noise type features respectively; and filter out the ambient noise type corresponding to the noise type feature in the preset ambient noise type set to obtain the candidate ambient noise type of the audio data block.

[0026] Optionally, in some embodiments, the determining unit may be specifically used to determine the basic noise audio feature that needs to be converted in the basic noise audio feature set to obtain the current basic noise audio feature; based on the sorting information, query the target basic noise audio feature that ranks first in the basic noise audio feature set; and according to the query result of the target basic noise audio feature, convert the basic noise audio features in the basic noise audio feature set into noise type features respectively.

[0027] Optionally, in some embodiments, the determining unit may be specifically used to convert the current basic noise audio feature into a noise type feature based on the target basic noise audio feature when the target basic noise audio feature exists; when the target basic noise audio feature does not exist, the target basic noise audio feature is used as a noise type feature; and return to the step of determining the basic noise audio feature that needs to be converted in the set of basic noise audio features until all basic noise audio features in the set of basic noise audio features are converted into noise type features, thereby obtaining the noise type feature corresponding to each basic noise audio feature.

[0028] Optionally, in some embodiments, the determining unit may be specifically used to obtain the hidden state feature corresponding to the target basic noise audio feature, the hidden state feature being used to indicate the hidden state passed during the process of converting the target basic noise audio feature into a noise type feature; calculate the feature ratio of the current basic noise audio feature and the hidden state feature, and determine the target hidden state feature corresponding to the current basic noise audio feature based on the feature ratio; and perform a dimensional transformation on the target hidden state feature to obtain the noise type feature corresponding to the current basic noise audio feature.

[0029] Optionally, in some embodiments, the data transmission device may further include a training unit, which may be used to acquire audio data samples, the audio data samples including audio data labeled with environmental noise types; predict the environmental noise type of the audio data samples using a preset noise type recognition model to obtain a predicted environmental noise type; and converge the preset noise type recognition model based on the labeled environmental noise type and the predicted environmental noise type to obtain a trained noise type recognition model.

[0030] Optionally, in some embodiments, the playback unit may be used to encrypt the target transmission data and generate a frequency mapping table based on the encrypted transmission data; to perform audio encoding on the target transmission data based on the frequency mapping table to obtain the audio frequency points corresponding to the target transmission data; to map the frequency value of each audio frequency point in the frequency mapping table and generate an audio waveform based on the frequency value, and to use the audio waveform as audio data.

[0031] Optionally, in some embodiments, the playback unit may be specifically used to divide the target transmission data into multiple data blocks to obtain a first data block, and divide the data identifier code of the target transmission data generated based on the current time into multiple data blocks to obtain a second data block; filter at least one message header from a preset message header set and use the message header as a third data block; perform error correction encoding on the first data block, the second data block, and the third data block, and divide the error correction code into multiple data blocks to obtain a fourth data block; extract audio frequency points from the first data block, the second data block, the third data block, and the fourth data block to obtain the audio frequency points corresponding to the target transmission data.

[0032] Furthermore, embodiments of the present invention also provide an electronic device, including a processor and a memory, wherein the memory stores an application program, and the processor is used to run the application program in the memory to implement the data transmission method provided in embodiments of the present invention.

[0033] Furthermore, embodiments of the present invention also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the data transmission methods provided in embodiments of the present invention.

[0034] In this embodiment of the invention, after acquiring the data to be transmitted and the current ambient noise, feature extraction is performed on the current ambient noise to obtain its audio features. Then, based on the audio features, the ambient noise type is determined, and the ambient noise type is added to the data to be transmitted to obtain the target data. The target data is then converted into audio data and played, so that the receiving terminal can obtain the data to be transmitted from the audio data according to the ambient noise type. Since this scheme can acquire the current ambient noise, determine its ambient noise type, and add it to the data to be transmitted, the receiving terminal can extract the ambient noise type from the received audio data. Based on the ambient noise type, the current network transmission quality (QoS) of the data to be transmitted can be determined, and the data to be transmitted can be obtained based on this QoS. Therefore, the data transmission success rate can be improved. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a schematic diagram of a scenario for the data transmission method provided in an embodiment of the present invention;

[0037] Figure 2 This is a flowchart illustrating the data transmission method provided in an embodiment of the present invention;

[0038] Figure 3 This is a schematic diagram of the process for extracting audio features from current environmental noise provided in an embodiment of the present invention;

[0039] Figure 4 This is a schematic diagram of the current environmental noise format provided in an embodiment of the present invention;

[0040] Figure 5 This is a schematic diagram of the network structure of the noise type recognition model after training provided in an embodiment of the present invention;

[0041] Figure 6 This is a schematic diagram of the bidirectional BGRU network provided in an embodiment of the present invention;

[0042] Figure 7 This is a flowchart illustrating the process of determining the type of current ambient noise provided in an embodiment of the present invention.

[0043] Figure 8 This is a schematic diagram of the process for identifying candidate environmental noise types using a trained noise type identification model, provided in an embodiment of the present invention.

[0044] Figure 9 This is a schematic diagram of the message header format provided in an embodiment of the present invention;

[0045] Figure 10 This is a schematic diagram of the overall process of data transmission provided in the embodiments of the present invention;

[0046] Figure 11 This is another schematic diagram of the data transmission method provided in an embodiment of the present invention;

[0047] Figure 12 This is a schematic diagram of the data transmission device provided in an embodiment of the present invention;

[0048] Figure 13 This is another structural schematic diagram of the data transmission device provided in an embodiment of the present invention;

[0049] Figure 14 This is another structural schematic diagram of the data transmission device provided in an embodiment of the present invention;

[0050] Figure 15 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] This invention provides a data transmission method, apparatus, and computer-readable storage medium. The data transmission apparatus can be integrated into an electronic device, which may be a server or a terminal, etc.

[0053] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0054] For example, see Figure 1 Taking the integration of a data transmission device into an electronic device as an example, after acquiring the data to be transmitted and the current ambient noise, the electronic device extracts features from the current ambient noise to obtain the audio features of the current ambient noise. Then, based on the audio features, it determines the type of ambient noise and adds the type of ambient noise to the data to be transmitted to obtain the target data to be transmitted. The target data to be transmitted is then converted into audio data and played so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise, thereby improving the success rate of data transmission.

[0055] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0056] This embodiment will be described from the perspective of a data transmission device, which can be integrated into an electronic device, such as a server or a terminal. The terminal can include tablet computers, laptops, personal computers (PCs), wearable devices, virtual reality devices, or other smart devices capable of data transmission.

[0057] A data transmission method, comprising:

[0058] The system acquires the data to be transmitted and the current ambient noise, which is the sound collected in real time under the current transmission environment. It extracts features from the current ambient noise to obtain its audio features. Based on the audio features, it determines the type of ambient noise and adds the type of ambient noise to the data to be transmitted to obtain the target data to be transmitted. It then converts the target data to be transmitted into audio data and plays the audio data so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise.

[0059] like Figure 2 As shown, the specific process of this data transmission method is as follows:

[0060] 101. Obtain the data to be transmitted and the current ambient noise.

[0061] The current ambient noise refers to the sound collected in real time under the current transmission environment. For example, if the current transmission environment is a subway, then the current ambient noise could be the sound collected in real time during a preset time period on the subway, and so on.

[0062] The types of data to be transmitted can be varied. For example, it could be transaction data generated by facial recognition devices and desktop POS devices during transactions, or it could be other data that needs to be transmitted.

[0063] There are several ways to obtain the data to be transmitted and the current ambient noise, including the following:

[0064] For example, regarding the data to be transmitted, user input data can be directly obtained as the data to be transmitted, or business data returned by a business server or other terminals can be received and used as the data to be transmitted. Regarding the current ambient noise, after obtaining the data to be transmitted, the sound acquisition device of the data transmission device can collect the ambient sound at a preset time to obtain the current ambient noise. Alternatively, the ambient sound in the current environment can be collected in real time, and when the data to be transmitted is obtained, the ambient sound at the current moment and the preset time collected can be extracted from the collected ambient sound to obtain the current ambient noise.

[0065] 102. Extract features from the current ambient noise to obtain the audio features of the current ambient noise.

[0066] Among these, audio features can be characteristics that indicate audio information in the current ambient noise.

[0067] There are several ways to extract features from the current environmental noise, including the following:

[0068] For example, the current ambient noise channel can be converted to obtain the current ambient noise of the target channel. This target channel's current ambient noise can then be divided into blocks to obtain multiple audio data blocks. Feature extraction can then be performed on these audio data blocks to obtain the corresponding audio features. Specifically, this can be done as follows: Figure 3 As shown.

[0069] The current environmental noise can be in various formats, such as stereo, 16-bit and 48kbps, stereo, 16-bit and 32kbps, or stereo, 16-bit and 16kbps. Specific formats include... Figure 4 As shown. There are several ways to convert the audio channels of the current ambient noise. For example, the audio channels of the current ambient noise can be detected, and if the current ambient noise has two channels, it can be converted into a mono current ambient noise.

[0070] After channel conversion of the current ambient noise, the current ambient noise of the target channel can be divided into blocks. There are multiple ways to divide the blocks. For example, the current ambient noise of the target channel can be divided into blocks according to a preset time window to obtain multiple audio data blocks. The preset time window can be any time, such as 30 seconds or other times. The preset time window can be adjusted according to the actual application.

[0071] After dividing the current ambient noise of the target channel into blocks, feature extraction can be performed on the audio data blocks to obtain the audio features corresponding to the audio data blocks. There are several ways to extract features from audio data blocks. For example, extract the spectrum image information from the audio data blocks, identify the spectrum image value corresponding to the audio data blocks in the spectrum image information, obtain the data sampling amount corresponding to each audio feature dimension, and extract the initial audio features corresponding to each audio feature dimension in the spectrum image information based on the data sampling amount. Then, concatenate the initial audio features with the spectrum image to obtain the audio features corresponding to each audio data block.

[0072] In this context, the spectral image information can be understood as the spectral image of the audio data block, the audio feature dimension can be understood as the audio feature type, and the data sampling rate can be the number of data points collected when performing a Fourier transform on the audio data block. The data sampling rate corresponding to each audio feature dimension can be the number of data points in the audio data block that need to be collected for each type of audio feature during the Fourier transform. Based on the data sampling rate, there are multiple ways to extract the initial audio features corresponding to each audio feature dimension from the spectral image information. For example, based on the sampling rate, a Fourier transform can be used in the spectral image of the audio data block to extract the basic audio features for each audio feature dimension. Then, based on the type of the basic audio features, the basic audio features can be standardized to obtain the initial audio features. Specifically, this can be done as follows:

[0073] For example, with a data sampling size (fft size) of 4096, 128-dimensional Mel spectrogram features (mel128) are extracted and standardized to obtain the first standardized audio features. With an fft size of 4096, Mel cepstral features (mfcc) are extracted and standardized to obtain the second standardized audio features. With an fft size of 1024, zero-crossing rate features can be extracted and binary encoded to obtain the encoded audio features. Spectral flatness features and spectral centroid features can also be extracted and standardized to obtain the third and fourth standardized audio features. The first, second, third, and fourth standardized audio features and the encoded audio features are used as the initial audio features.

[0074] 103. Based on audio characteristics, determine the type of ambient noise in the current environment.

[0075] Among them, the environmental noise type can be understood as the scene type corresponding to the current transmission environment. For example, it can include various common life scenes such as normal environment, subway, supermarket, school and construction site.

[0076] There are several ways to determine the type of ambient noise based on audio characteristics, including the following:

[0077] For example, a trained noise recognition model can be used to identify candidate environmental noise types for corresponding audio data blocks in audio features. The number of types corresponding to each environmental noise type can be counted in the candidate environmental noise types. Based on the number of types, the environmental noise type of the current environmental noise can be selected from the candidate environmental noise types.

[0078] There are several ways to identify candidate environmental noise types for corresponding audio data blocks using a trained noise recognition model. For example, a trained noise type recognition model can be used to extract noise audio features from the audio features of each audio data block to obtain a set of noise audio features for the current environment. Batch normalization is then performed on the noise audio features in the noise audio feature set to obtain a set of basic noise audio features. The environmental noise type corresponding to each basic noise audio feature is then identified in the set of basic noise audio features to obtain the candidate environmental noise type for the audio data block.

[0079] There are several ways to extract noise audio features from the audio features of each audio data block using a trained noise type recognition model. For example, the audio features can be fused and converted into a spectrogram. The noise feature extraction network of the trained noise type recognition model can then be used to extract features from the spectrogram, thereby obtaining a set of noise audio features for the current environment. This set of noise audio features can include the noise audio features corresponding to each audio data block.

[0080] The network structure of the noise type recognition model after training can be varied. For example, it can be a residual network model that includes a residual network substructure and a batch normalization (BN) layer, such as... Figure 5 As shown, or, it can be any other network structure that includes both of these. The network structure of the noise feature extraction network of the noise type recognition model after training can consist of multiple network layers. For example, the second layer, excluding the input layer, can have 64 convolutional kernels, with a kernel size of 3x3, a stride of 1, and padding of 1. The third layer is a Maxpolling layer with windows of 2x2 and a stride of 2. The fourth layer uses 128 convolutional kernels, with a kernel size of 3x3, a stride of 1, and padding of 1. The fifth layer is a Maxpolling layer with windows of 2x2 and a stride of 2. The sixth layer uses 256 convolutional kernels, with a kernel size of 3x3, a stride of 1, and padding of 1. The seventh layer is a Maxpolling layer with windows of 2x2 and a stride of 2.

[0081] After extracting the noise audio features, batch normalization can be performed on the noise audio features in the noise audio feature set. There are several ways to perform batch normalization. For example, a BN layer can be used to perform batch normalization on the noise audio features in the noise audio feature set, and a Maxpolling layer can be used to perform pooling on the batch normalized noise audio features, thereby obtaining the basic noise audio features corresponding to each audio data block. The basic noise audio features are then combined to obtain the basic noise audio feature set.

[0082] After batch normalization of the noise audio features in the noise audio feature set, the environmental noise type corresponding to each basic noise audio feature can be identified in the basic noise audio feature set. There are multiple ways to identify the environmental noise type. For example, the time information of the audio data block in the current environmental noise can be obtained, and the basic noise audio features in the basic noise audio feature set can be sorted based on the time information. Based on the sorting information, the basic noise audio features in the basic noise audio feature set can be converted into noise type features respectively. The environmental noise type corresponding to the noise type feature can be selected from the preset environmental noise type set to obtain the candidate environmental noise type of the audio data block.

[0083] The time information can be understood as the temporal position of the audio data block within the current ambient noise, such as 5-6 seconds or a specific time interval. Based on this time information, there are multiple ways to sort the basic noise audio features in the basic noise audio feature set. For example, based on this time information, the audio data blocks in the current ambient noise can be sorted chronologically from front to back to obtain sorting information. Then, based on this sorting information, the corresponding basic noise audio features of the audio data blocks can be sorted to obtain the sorting information of the basic noise audio features.

[0084] After sorting the basic noise audio features, the basic audio features in the basic noise audio feature set can be converted into noise type features based on the sorting information. There are several ways to convert them. For example, the basic noise audio feature that needs to be converted can be determined in the basic noise audio feature set to obtain the current basic noise audio feature. Based on the sorting information, the target basic noise audio feature that ranks first in the basic noise audio feature set can be queried. Based on the query result of the target basic noise audio feature, the basic noise audio features in the basic noise audio feature set can be converted into noise type features.

[0085] Among them, the noise type feature can be the feature information indicating the environmental noise type of the audio data block. According to the target basic noise audio feature, there are multiple ways to convert the basic noise audio features in the basic noise audio feature set into noise type features. For example, when the target basic noise audio feature exists, the current basic noise audio feature is converted into a noise type feature according to the target basic noise audio feature. When the target basic noise audio feature does not exist, the target basic noise feature is used as the noise type feature, and the step of determining the basic noise audio feature that needs to be converted in the basic noise audio feature set is returned to be executed until all the basic noise audio features in the basic noise audio feature set are converted into noise type features, so as to obtain the noise type feature corresponding to each basic noise audio feature.

[0086] In this process, the transformation of basic noise audio features in the basic noise audio feature set relies on the latent state features generated during the processing of historical basic noise audio features. The so-called latent state features can be understood as intermediate features generated by the noise type feature transformation network when transforming the target basic noise audio features. These intermediate features are used to indicate the latent state passed during the transformation of the target basic noise features into noise types. Therefore, the existence of target basic noise audio features indicates that there were already transformed target basic noise audio features before the current basic noise audio features. There are multiple ways to transform the current basic noise into noise type features based on the target basic noise audio features. For example, the latent state features of the target basic noise audio features can be obtained, and the noise type feature transformation network can transform the current basic noise into noise type features based on the latent state features of the target basic noise audio features.

[0087] The noise type feature conversion network can be a bidirectional BGRU (gated recurrent network), with 256 hidden units, as shown in the example below. Figure 6 As shown, there are several ways to use a noise type feature transformation network to convert the current basic noise into noise type features. For example, the feature ratio of the current basic noise audio features and the hidden state features can be calculated, and based on the feature ratio, the target hidden state features corresponding to the current basic noise audio features can be determined. The target hidden state features can then be dimensionally transformed to obtain the noise type features corresponding to the basic noise audio features.

[0088] The dimensional transformation of the target hidden state features can be understood as converting the intermediate features obtained by the noise type feature transformation network based on the hidden state features into output features, thereby obtaining the output noise type features. There are various ways to perform dimensional transformation. For example, transformation parameters can be obtained and fused with the target hidden state features to obtain the noise type features corresponding to the basic noise audio features.

[0089] After converting the basic noise audio features into noise type features, the environmental noise type corresponding to the noise type feature can be selected from the preset set of environmental noise types. There are several ways to select the environmental noise type. For example, a softmax network can be used to map the noise environment type to the classification probability corresponding to each environmental noise type. Based on the classification probability, the environmental noise type corresponding to the noise type feature can be selected from the noise environment types, thereby obtaining the candidate environmental noise type of the audio data block.

[0090] After identifying the candidate ambient noise types for the audio data block, the number of each type can be counted. Then, based on the number of types, the ambient noise type of the current ambient noise can be selected from the candidate ambient noise types. There are several ways to select it. For example, the ambient noise type with the largest number of types in the candidate ambient noise types can be used as the ambient noise type of the current ambient noise. Alternatively, a weighting parameter for each type of ambient noise can be obtained. Based on this weighting parameter, the number of types of each ambient noise type can be weighted. Then, the ambient noise type with the largest number of weighted types can be selected from the candidate ambient noise types to obtain the ambient noise type of the current ambient noise.

[0091] The process of extracting features from the current environmental noise and determining the type of environmental noise can be described as follows: Figure 7 As shown, the current ambient noise is converted from stereo to mono, then the current ambient noise is divided into frames or blocks to obtain audio data blocks. Candidate ambient noise types are predicted for the audio data blocks, the prediction results are statistically analyzed, and the ambient noise type of the current ambient noise is selected.

[0092] The trained noise recognition model can be configured according to the needs of actual applications. Furthermore, it should be noted that the trained noise recognition model can be pre-configured by maintenance personnel or trained automatically by the data transmission device. That is, before the step "extracting noise audio features from the audio features of each audio data block using the trained noise type recognition model," the data transmission method may further include:

[0093] Obtain audio data samples, which include audio data with labeled environmental noise types. Use a preset noise type recognition model to predict the environmental noise type of the audio data samples to obtain the predicted environmental noise type. Converge the preset noise type recognition model based on the labeled environmental noise type and the predicted environmental noise type to obtain the trained noise type recognition model.

[0094] There are several ways to obtain audio data samples. For example, one approach is to determine the ambient noise audio waveforms for various scenarios, then collect audio data at different time intervals within these scenarios. The types of ambient noise in these audio data, labeled with the specific scenario, are used as positive samples. Conversely, negative sample audio data is collected in the same quantity as the positive samples, and the ambient noise types for each scenario are labeled to obtain the negative sample audio data. Finally, the positive and negative sample audio data are cleaned to obtain the final audio data samples.

[0095] After obtaining the audio data sample, a preset noise type identification model can be used to predict the environmental noise type of the audio data sample, thus obtaining the predicted environmental noise type. There are various prediction methods. For example, the preset noise type identification model can be used to extract features from the audio data sample to obtain the sample audio features. The sample environmental noise type of the data block of the audio data sample can be identified from the sample audio features. The number of sample types corresponding to each environmental noise type can be counted. Based on the number of sample types, the environmental noise type of the audio data sample can be selected from the sample environmental noise types to obtain the predicted environmental noise type.

[0096] After obtaining the predicted environmental noise type, the preset noise type recognition model can be converged based on the labeled environmental noise type and the predicted environmental noise type. There are several ways to converge. For example, the labeled environmental noise type and the predicted environmental noise type can be compared to obtain the loss information of the audio data sample. Based on this loss information, the gradient descent algorithm is used to update the network parameters of the preset noise type recognition model in order to converge the preset noise type recognition model and obtain the trained noise type recognition model.

[0097] The process of identifying candidate environmental noise types based on the audio features of audio data blocks using a trained noise type recognition model can be described as follows: Figure 8 As shown, audio data samples are collected, and then features are extracted from the collected audio data samples through audio feature engineering. Then, a preset noise type recognition model is designed. Then, the preset noise type recognition model is trained using the extracted sample audio features. Finally, the trained noise type recognition model is used to identify candidate environmental noise types in the audio data blocks from the audio features.

[0098] 104. Add the environmental noise type to the data to be transmitted to obtain the target data to be transmitted.

[0099] For example, the environmental noise type can be directly added to the data to be transmitted to obtain the target data to be transmitted. Alternatively, the data label corresponding to the environmental noise type can be selected from the preset data label set and added to the data to be transmitted to obtain the target data to be transmitted. Or, the environmental noise type can be converted into environmental noise type data and added to the data to be transmitted to obtain the target data to be transmitted.

[0100] 105. Convert the target transmission data into audio data and play the audio data so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise.

[0101] There are several ways to convert target transmission data into audio data, including the following:

[0102] For example, the target transmitted data can be encrypted, and a frequency mapping table can be generated based on the encrypted transmitted data. Based on the frequency mapping table, the target transmitted data can be audio encoded to obtain the audio frequency points corresponding to the target transmitted data. The frequency value of each audio frequency point can be mapped in the frequency mapping table, and an audio waveform can be generated based on the frequency value. The audio waveform is then used as audio data.

[0103] There are several ways to encrypt the target transmitted data. For example, an SE security chip can be used for encryption, or a hash algorithm or MD5 value can be used to encrypt the target transmitted data.

[0104] After encrypting the target transmission data, a frequency mapping table can be generated based on the encrypted data. A frequency mapping table can be understood as each data block's block identifier corresponding to a specific frequency; inputting a specific data block allows mapping to the corresponding frequency. There are several ways to generate frequencies based on encrypted transmission data. For example, the encrypted transmission data can be divided into frames, and each frame can be decomposed into multiple metadata blocks. Based on the numerical values ​​corresponding to the effective bits of the metadata, a mapping table between metadata and frequencies can be generated, and this mapping table can be used as the frequency mapping table.

[0105] The encrypted transmitted data is divided into frames. Each frame can include bytes of preset data. The frame length can be set according to the actual application, such as 14 bytes or other numbers of bytes. There are several ways to decompose each frame of data into multiple metadata blocks. For example, taking a frame of 14 bytes as an example, 5 bits of data are taken from each of the 14 bytes in sequence as metadata. If there are not enough 5 bits, they are padded with 0s from the adjacent bytes. In essence, this decomposes an 8-bit byte into 5-bit metadata. The maximum value can be 2^5, and the minimum value is 0. The metadata required for one frame is (N * 8 + 4 / 5), where N is the number of bytes in the frame. This metadata can be understood as a data block.

[0106] After decomposing the metadata, a mapping table between the metadata and the frequency can be generated based on the numerical values ​​corresponding to the significant digits of the metadata. There are several ways to generate the mapping table. For example, an arithmetic sequence or a geometric sequence can be used to generate the mapping table between the metadata and the frequency. Specifically, it can be shown in formula (1):

[0107] Freq metadata = Freq Base + metadata * θ f (1)

[0108] Among them, Freq metadata Freq represents the frequency of values ​​corresponding to metadata. Base The starting frequency is θ, metadata is the metadata, and θ is the starting frequency. f The frequency mapping table generated for the corresponding frequency difference between two adjacent metadata is shown in Table 1:

[0109] Table 1

[0110]

[0111] After generating the frequency mapping table, the target transmission data can be audio encoded to obtain the audio frequency points corresponding to the target transmission data. There are various ways to encode the target transmission data audio. For example, the target transmission data can be divided into multiple data blocks to obtain the first data block, and the data identifier code of the target transmission data generated based on the current time can be divided into multiple data blocks to obtain the second data block. At least one message header can be selected from the preset message header set and used as the third data block. Error correction encoding can be performed on the first, second, and third data blocks, and the error correction code can be divided into multiple data blocks to obtain the fourth data block. The audio frequency points can be extracted from the first, second, third, and fourth data blocks to obtain the audio frequency points corresponding to the target transmission data.

[0112] Here, data blocks can be contiguous or non-contiguous chunks, and the header is the structure added to the target data packet when transmitting the target data. The format of this header can be as follows: Figure 9 As shown.

[0113] The data identifier of the target transmitted data can be a universally unique identifier (UUID). There are several ways to generate the data identifier of the target transmitted data. For example, the current time can be obtained, the attribute information of the current time and the target transmitted data can be fused, the UUID algorithm can be used to generate a UUID from the fused data, and then the UUID can be divided into multiple chunks.

[0114] After dividing the data into the first, second, and third data blocks, error correction coding can be performed on them. There are various error correction coding methods. For example, Reed-Solomon (an error correction coding algorithm) can be used to encode the data to obtain the error correction (RS) code. The RS code can then be divided into multiple chunks to obtain the fourth data block.

[0115] After obtaining the first data block, the second data block, the third data block, and the fourth data block, the audio frequency points can be extracted from them. There are several ways to extract the audio frequency points. For example, the data block information of the first data block, the second data block, the third data block, and the fourth data block can be obtained, the data block identifier can be identified in the data block information, and the database identifier can be used as the audio frequency point of the target transmitted data.

[0116] After audio encoding of the target transmission data, the corresponding audio waveform can be generated. There are several ways to generate the audio waveform. For example, the frequency value of each audio frequency point in the frequency mapping table can be used to generate a single-frequency audio waveform for each audio frequency point. The single-frequency audio waveforms can be fused to obtain the audio waveform corresponding to the target transmission data, and the audio waveform can be used as the audio data of the target transmission data.

[0117] There are several ways to generate a single-frequency audio waveform for each audio frequency point based on the frequency value. For example, the sin function can be used to process the frequency value to obtain the single-frequency audio waveform for each audio frequency point.

[0118] After the target transmission data is converted into audio data, the audio data can be played. There are several ways to play the audio; for example, a corresponding audio signal can be generated based on the audio waveform of the target transmission data, and then the audio signal can be played to achieve audio data playback. When the receiving terminal detects the audio signal, it can acquire the audio data, extract the target transmission data from the audio data, identify the type of ambient noise in the target transmission data, and identify the data to be transmitted from the target transmission data based on the type of ambient noise.

[0119] Optionally, when audio data playback is detected, transmission data can also be obtained from the audio data. There are multiple ways to obtain the data, such as recording the target audio data being played, obtaining the recording data, parsing the recording data to obtain the current ambient noise type and the initial audio data, and identifying the transmission data from the initial audio data based on the current ambient noise type.

[0120] There are several ways to analyze the recording data. For example, the recording data (PCM) can be filtered to remove unnecessary signals and obtain filtered audio data. The type of ambient noise and the initial audio data can then be identified from the filtered audio data.

[0121] After analyzing the type of ambient noise and the initial audio data, the transmitted data can be identified in the initial audio data. There are several ways to identify the transmitted data. For example, based on the current type of ambient noise, the network transmission quality of the target audio data can be evaluated to obtain evaluation information. Based on the evaluation information, the processing method of the target audio data can be determined. When the processing method is to receive the target audio data, the transmitted data can be identified in the initial audio data.

[0122] Among them, Quality of Service (QoS) refers to the ability of a network to provide better service for specified network communications using various basic technologies. It is a network security mechanism. Evaluating QoS primarily involves determining the QoS value based on the current ambient noise type. There are several evaluation methods; for example, one can filter the QoS value corresponding to the current ambient noise type from a preset QoS value set and use that value as evaluation information. Regarding evaluating QoS based on the current ambient noise type, it's important to explain. For instance, the current ambient noise type could be subway noise or indoor noise. In subway noise conditions, recording audio data will capture more ambient noise than in indoor noise conditions, leading to inaccurate recording data. When the recording data is inaccurate, the accuracy of the transmitted data parsed from the recording will also decrease, resulting in data transmission failure. Therefore, by obtaining the ambient noise type in the current transmission environment, one can determine whether the received audio data is accurate, thereby improving the accuracy of data transmission.

[0123] After evaluating the network transmission instructions for the target audio data, the processing method for the target audio data can be determined based on the evaluation information. There are multiple ways to determine the processing method for the target audio data. For example, the QoS value can be compared with a preset evaluation threshold. When the QoS value exceeds the preset evaluation threshold, the processing method for the target audio data is determined to be receiving the target audio data. When the QoS value does not exceed the preset evaluation threshold, the processing method for the target audio data is determined to be not receiving the target audio data. Alternatively, when the QoS value is the preset QoS value, the processing method for the target audio data is determined to be receiving the target audio data. When the QoS value is not the preset QoS value, the processing method for the target audio data is determined to be not receiving the target audio data.

[0124] When the processing method is to receive target audio data, there are multiple ways to identify the transmitted data in the initial audio data. For example, the baseband data can be extracted from the initial audio data, and the baseband data can be detected by the data block identifier in the baseband data. If the baseband data contains a message header, the acoustic wave baseband data can be extracted from the initial audio data. At least one data block identifier can be extracted from the acoustic wave baseband data, and the extracted data block identifier can be corrected to obtain a set of data block identifiers. The target data block corresponding to each data block identifier in the set of data block identifiers can be obtained, and the target data blocks can be fused to obtain the transmitted data.

[0125] There are several ways to perform message header detection on the baseband data based on the data block identifiers in the baseband data. For example, the data block identifier corresponding to each frequency in the baseband data can be mapped in a frequency mapping table. This data block identifier is then matched with a preset data block identifier corresponding to the message header. When the match is successful, it can be determined that the baseband data contains a message header; when the match fails, it can be determined that the baseband data does not contain a message header. When the baseband data does not contain a message header, it means that the initial audio data does not contain the data to be transmitted or the transmitted data is incomplete, and the identification of transmitted data stops. When the baseband data contains a message header, it means that the initial audio data contains the data to be transmitted. In this case, the acoustic wave baseband data can be extracted from the initial audio data, and at least one data block identifier can be extracted from the acoustic wave baseband data. Reed-Solomon error correction is then applied to the extracted data block identifiers to obtain a set of data block identifiers. The method for extracting data block identifiers here is the same as the method for extracting data block identifiers from the baseband data, and will not be described in detail here.

[0126] After obtaining the set of data block identifiers, the target data block corresponding to each data block identifier in the set can be obtained and the target data blocks can be merged. There are multiple ways to merge them. For example, the target data block can be decrypted, and the decrypted data blocks can be spliced ​​or merged to obtain the merged data. The message header is then deleted from the merged data to obtain the transmitted data.

[0127] Optionally, when the processing mode is to not receive the target audio data, the initial audio data is discarded, a prompt message is generated, and the prompt message is sent to the sending terminal so that the sending terminal can replay the target audio data based on the prompt message. This determines that the current transmission network quality is poor when the target audio data is being played, requiring the sending terminal to replay the target audio data to resend the transmission data, thereby improving the data transmission success rate.

[0128] It should be noted that during the data transmission process, it can be as follows: Figure 10As shown, after obtaining the data to be transmitted and the current ambient noise, the QoS module can be used to determine the ambient noise type of the current ambient noise, add the ambient noise type to the data to be transmitted, and thus obtain the target data to be transmitted. The target transmission data is encoded using an acoustic encoding module. Then, a header is added by a header generation module, and a data identifier (UUID) is generated based on the current time. Error correction encoding is performed using a reed-solomom module. Finally, audio data (PCM) is generated and played. When the receiving terminal detects audio playback, it records the audio data and performs PCM parsing on the recording. The QoS module obtains the current ambient noise type and evaluates network transmission quality. If the network transmission quality meets the requirements, header detection is performed. If a header is present in the recording data, the acoustic fundamental frequency data is extracted, and the data block identifiers within the acoustic fundamental frequency data are obtained. Reed-solomom performs error correction to obtain a set of data block identifiers. The target data block corresponding to each data block identifier in the set is obtained, and the target data blocks are concatenated to obtain the data to be transmitted.

[0129] As can be seen from the above, in this embodiment, after acquiring the data to be transmitted and the current ambient noise, feature extraction is performed on the current ambient noise to obtain the audio features of the current ambient noise. Then, based on the audio features, the ambient noise type of the current ambient noise is determined, and the ambient noise type is added to the data to be transmitted to obtain the target data to be transmitted. The target data to be transmitted is then converted into audio data and played, so that the receiving terminal can obtain the data to be transmitted from the audio data according to the ambient noise type. Since this scheme can acquire the current ambient noise, determine the ambient noise type of the current ambient noise, and add the ambient noise type to the data to be transmitted, the receiving terminal can extract the ambient noise type from the received audio data. Thus, the current network transmission quality (QoS) of the data to be transmitted can be determined based on the ambient noise type, and the data to be transmitted can be obtained based on the current network transmission quality. Therefore, the transmission success rate of data transmission can be improved.

[0130] Based on the method described in the above embodiments, the following examples will provide further detailed explanations.

[0131] In this embodiment, the data transmission device is specifically integrated into an electronic device, which is the terminal. To distinguish them, the terminal can be divided into a sending terminal and a receiving terminal. The sending terminal can be a terminal that transmits the data to be transmitted, and the receiving terminal can be a terminal that receives the data to be transmitted.

[0132] (a) Training the noise type recognition model

[0133] (1) The sending terminal obtains audio data samples.

[0134] For example, the transmitting terminal determines the ambient noise audio waveforms for various scenarios. Then, it collects audio data for different time periods within these scenarios, labels the ambient noise type of the target scenario in this audio data as positive samples, determines the types of negative sample audio, collects the same number of negative sample audio data as the positive samples, and labels the ambient noise type of the scenario to obtain negative sample audio data. The positive and negative sample audio data are then cleaned to obtain audio data samples.

[0135] (2) The transmitting terminal uses a preset noise type identification model to predict the environmental noise type of the audio data sample and obtain the predicted environmental noise type.

[0136] For example, the transmitting terminal uses a preset noise type identification model to extract features from audio data samples, obtains the sample audio features of the audio data samples, identifies the sample environmental noise type of the data block of the audio data sample from the sample audio features, counts the number of sample types corresponding to each environmental noise type from the sample environmental noise types, and based on the number of sample types, filters out the environmental noise type of the audio data sample from the sample environmental noise types to obtain the predicted environmental noise type.

[0137] (3) The transmitting terminal converges the preset noise type identification model based on the labeled environmental noise type and the predicted environmental noise type to obtain the trained noise type identification model.

[0138] For example, the transmitting terminal compares the labeled environmental noise type with the predicted environmental noise type to obtain the loss information of the audio data sample. Based on this loss information, the gradient descent algorithm is used to update the network parameters of the preset noise type recognition model to converge the preset noise type recognition model, thereby obtaining the trained noise type recognition model.

[0139] (ii) The terminal uses a trained noise type identification model to determine the environmental noise type of the current environment.

[0140] The trained noise recognition model includes a noise feature extraction network and a noise type feature conversion network.

[0141] like Figure 11 As shown, a data transmission method has the following specific process:

[0142] 201. The sending terminal obtains the data to be transmitted and the current ambient noise.

[0143] For example, regarding the data to be transmitted, the sending terminal can directly obtain the user-input data as the data to be transmitted, or it can receive business data returned by a business server or other terminals and use that business data as the data to be transmitted. Regarding the current ambient noise, after obtaining the data to be transmitted, the sending terminal can use the sound acquisition device of the data transmission device to collect the ambient sound at a preset time to obtain the current ambient noise. Alternatively, it can collect the ambient sound in real time, and when the data to be transmitted is obtained, extract the ambient sound at the preset time from the collected ambient sound to obtain the current ambient noise.

[0144] 202. The transmitting terminal extracts features from the current ambient noise to obtain the audio features of the current ambient noise.

[0145] For example, the transmitting terminal detects the channels of the current ambient noise. When the current ambient noise has two channels, it converts the current ambient noise into a mono channel. The current ambient noise of the target channel is then divided into blocks according to a preset time window, resulting in multiple audio data blocks. The preset time window can be any time, such as 30 seconds or other times, and can be adjusted according to the actual application.

[0146] The transmitting terminal extracts spectral image information from the audio data block and identifies the spectral image value corresponding to the audio data block from the spectral image information. It obtains the data sampling amount corresponding to each audio feature dimension. With a data sampling amount (fftsize) of 4096, a 128-dimensional Mel spectrogram feature (mel128) is extracted and standardized to obtain the first standardized audio feature. With an fft size of 4096, Mel cepstral features (mfcc) are extracted and standardized to obtain the second standardized audio feature. With an fft size of 1024, zero-crossing rate features can be extracted and binary encoded to obtain the encoded audio feature. Spectral flatness and spectral centroid features can also be extracted and standardized to obtain the third and fourth standardized audio features. The first, second, third, and fourth standardized audio features, along with the encoded audio feature, are used as the initial audio features. The initial audio features are concatenated with the spectral image to obtain the audio feature corresponding to each audio data block.

[0147] 203. The transmitting terminal determines the type of ambient noise based on audio characteristics.

[0148] For example, the transmitting terminal fuses audio features and converts the fused audio features into a spectrogram. It then uses a noise feature extraction network of a trained noise type recognition model to extract features from the spectrogram, thereby obtaining a set of noise audio features for the current environmental noise. This set of noise audio features may include the noise audio features corresponding to each audio data block.

[0149] The noise feature extraction network of the noise type recognition model after training can consist of multiple network layers. For example, the second layer, excluding the input layer, can have 64 convolutional kernels with a kernel size of 3x3, a stride of 1, and padding of 1. The third layer is a Maxpolling layer with windows of 2x2 and a stride of 2. The fourth layer uses 128 convolutional kernels with a kernel size of 3x3, a stride of 1, and padding of 1. The fifth layer is a Maxpolling layer with windows of 2x2 and a stride of 2. The sixth layer uses 256 convolutional kernels with a kernel size of 3x3, a stride of 1, and padding of 1. The seventh layer is a Maxpolling layer with windows of 2x2 and a stride of 2.

[0150] The transmitting terminal uses a BN layer to perform batch normalization on the noise audio features in the noise audio feature set, and uses a Maxpolling layer to perform pooling on the batch normalized noise audio features, thereby obtaining the basic noise audio features corresponding to each audio data block. The basic noise audio features are then combined to obtain the basic noise audio feature set.

[0151] The transmitting terminal obtains the time information of the audio data block in the current ambient noise. Based on the time information, the audio data blocks in the current ambient noise are sorted from front to back according to the time order to obtain sorting information. Then, according to the sorting information, the basic noise audio features corresponding to the audio data blocks are sorted to obtain the sorting information of the basic noise audio features.

[0152] The process begins by identifying the current basic noise audio feature to be transformed from the set of basic noise audio features. Based on the ranking information, a target basic noise audio feature is searched for that preceding the current one. If a target basic noise audio feature exists, the current basic noise audio feature is transformed into a noise type feature. A bidirectional BGRU can be used to obtain the hidden state features of the target basic noise audio feature. A noise type feature transformation network is then used to transform the current basic noise audio feature into a noise type feature based on the hidden state features of the target basic noise audio feature. The feature ratio between the current basic noise audio feature and the hidden state feature is calculated. Based on this feature ratio, the target hidden state feature corresponding to the current basic noise audio feature is determined. A dimensionality transformation is then performed on the target hidden state feature to obtain the noise type feature corresponding to the basic noise audio feature. If no target basic noise audio feature exists, it is used as the noise type feature, and the process returns to the step of identifying the current basic noise audio feature to be transformed from the set of basic noise audio features. This process continues until all basic noise audio features in the set are transformed into noise type features, resulting in the noise type feature corresponding to each basic noise audio feature.

[0153] The transmitting terminal uses a softmax network to map the noise environment type to the classification probability corresponding to each noise type. Based on this classification probability, it filters out the environmental noise types corresponding to the noise type features from the noise environment types, thus obtaining candidate environmental noise types for the audio data block. The environmental noise type with the largest number of candidates is selected as the current environmental noise type. Alternatively, a weighting parameter can be obtained for each type of environmental noise. Based on this weighting parameter, the number of types for each environmental noise type is weighted, and then the environmental noise type with the largest number of weighted types is selected from the candidate environmental noise types, thus obtaining the current environmental noise type.

[0154] 204. The transmitting terminal adds the environmental noise type to the data to be transmitted to obtain the target data to be transmitted.

[0155] For example, the transmitting terminal can directly add the environmental noise type to the data to be transmitted to obtain the target data; or, it can filter out the data tag corresponding to the environmental noise type from the preset data tag set and add the data tag to the data to be transmitted to obtain the target data; or, it can convert the environmental noise type into environmental noise type data and add the environmental noise type data to the data to be transmitted to obtain the target data.

[0156] 205. Sending you will convert the target transmitted data into audio data and play the audio data.

[0157] For example, the sending terminal can use an SE security chip for encryption, or it can use a hash algorithm or MD5 value to encrypt the target transmitted data to obtain encrypted transmitted data. Taking each frame of data as 14 bytes as an example, take 5 bits of data from each of the 14 bytes in sequence as metadata. If there are not enough 5 bits of data, fill in the gaps or replace them with 0s from the adjacent bytes. What we need to do here is to decompose an 8-bit byte into a 5-bit metadata. The maximum value can be 2^5 and the minimum value can be 0. The metadata required for one frame is: (N * 8 + 4 / 5), where N is the number of bytes in the frame. The metadata here can be understood as a data block. Use an arithmetic sequence or a geometric sequence to generate a mapping table between metadata and frequency. Specifically, as shown in formula (1), use this mapping table as a frequency mapping table.

[0158] The transmitting terminal divides the target transmission data into multiple data blocks to obtain the first data block. It then obtains the current time and fuses the attribute information of the current time and the target transmission data. A UUID is generated from the fused data using a UUID algorithm. This UUID is then divided into multiple chunks to obtain the second data block. Reed-Solomon can be used to perform error correction coding on the first, second, and third data blocks to obtain RS codes. These RS codes are then divided into multiple chunks to obtain the fourth data block. The data block information of the first, second, third, and fourth data blocks is obtained, and the data block identifier is identified within this information. This identifier is used as the audio frequency point of the target transmission data.

[0159] The transmitting terminal processes the frequency value of each audio frequency point in the frequency mapping table using a sine function to obtain a single-frequency audio waveform for that audio frequency point. These single-frequency audio waveforms are then fused to obtain the audio waveform corresponding to the target transmitted data, which is used as the audio data for the target transmitted data. Based on the audio waveform of the target transmitted data, a corresponding audio signal is generated and then played, thereby realizing the playback of the audio data.

[0160] 206. When the receiving terminal detects that audio data is being played, the receiving terminal records the audio data being played and obtains the recording data.

[0161] For example, when the receiving terminal detects that audio data is being played, the receiving terminal starts the recording device to record the audio data being played, thereby obtaining the recording data.

[0162] 207. The receiving terminal parses the recording data to obtain the current ambient noise type and initial audio data.

[0163] For example, the receiving terminal can filter the recorded data (PCM) to remove unnecessary signals and obtain filtered audio data. In the filtered audio data, the current ambient noise type and the initial audio data can be identified.

[0164] 208. The receiving terminal identifies the transmitted data from the initial audio data based on the current ambient noise type.

[0165] For example, the receiving terminal filters the QoS value corresponding to the current ambient noise type from a preset QoS value set and uses this QoS value as evaluation information. The QoS value is compared with a preset evaluation threshold. If the QoS value exceeds the preset evaluation threshold, the processing method for the target audio data is determined to be receiving the target audio data; if the QoS value does not exceed the preset evaluation threshold, the processing method for the target audio data is determined to be not receiving the target audio data. Alternatively, if the QoS value is the preset QoS value, the processing method for the target audio data is determined to be receiving the target audio data; if the QoS value is not the preset QoS value, the processing method for the target audio data is determined to be not receiving the target audio data.

[0166] When the processing method is receiving audio data, the receiving terminal extracts the baseband data from the initial audio data, maps the data block identifier corresponding to each frequency in the baseband data to a frequency mapping table, and matches this data block identifier with the preset data block identifier corresponding to the message header. If the match is successful, it can be determined that the baseband data contains a message header; if the match fails, it can be determined that the baseband data does not contain a message header. If the baseband data does not contain a message header, it means that the initial audio data does not contain the data to be transmitted or the transmitted data is incomplete, and the identification of transmission data stops. If the baseband data contains a message header, the acoustic wave baseband data can be extracted from the initial audio data, and at least one data block identifier can be extracted from the acoustic wave baseband data. Reed-Solomon error correction is applied to the extracted data block identifiers to obtain a set of data block identifiers. The target data block corresponding to each data block identifier in the data block identifier set is obtained, the target data block is decrypted, and the decrypted data blocks are concatenated or merged to obtain merged data. The message header is then removed from the merged data to obtain the transmitted data.

[0167] When the processing mode is set to not receive target audio data, the receiving terminal discards the initial audio data and generates a prompt message, which is then sent to the sending terminal. The sending terminal then replays the target audio data based on the prompt. This indicates that the current network quality is poor when the target audio data is being played, requiring the sending terminal to replay the target audio data to resend the transmission, thereby improving the data transmission success rate.

[0168] As can be seen from the above, in this embodiment, after acquiring the data to be transmitted and the current ambient noise, the terminal extracts features from the current ambient noise to obtain the audio features of the current ambient noise. Then, based on the audio features, it determines the type of ambient noise and adds the type of ambient noise to the data to be transmitted to obtain the target data to be transmitted. The target data to be transmitted is then converted into audio data and played, so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise. Since this scheme can acquire the current ambient noise, determine the type of ambient noise, and add the type of ambient noise to the data to be transmitted, the receiving terminal can extract the type of ambient noise from the received audio data. Based on the type of ambient noise, it can determine the current network transmission quality (QoS) of the data to be transmitted and obtain the data to be transmitted based on the current network transmission quality. Therefore, the success rate of data transmission can be improved.

[0169] To better implement the above methods, embodiments of the present invention also provide a data transmission device, which can be integrated into an electronic device, such as a server or terminal, and the terminal may include a tablet computer, a laptop computer, and / or a personal computer.

[0170] For example, such as Figure 12 As shown, the data transmission device may include an acquisition unit 301, an extraction unit 302, a determination unit 303, an addition unit 304, and a playback unit 305, as follows:

[0171] (1) Obtain unit 301;

[0172] The acquisition unit 301 is used to acquire the data to be transmitted and the current ambient noise, wherein the current ambient noise is the sound in the current transmission environment collected in real time.

[0173] For example, the acquisition unit 301 can be used to directly acquire user-input data as data to be transmitted, or it can receive business data returned by a business server or other terminals and use that business data as data to be transmitted. After acquiring the data to be transmitted, the sound acquisition device of the data transmission device acquires the sound of the current environment at a preset time to obtain the current environmental noise. Alternatively, it can acquire the environmental sound of the current environment in real time. When the data to be transmitted is acquired, the environmental sound acquired at the preset time at the current moment is extracted from the acquired environmental sound to obtain the current environmental noise.

[0174] (2) Extraction unit 302;

[0175] The extraction unit 302 is used to extract features from the current environmental noise to obtain the audio features of the current environmental noise.

[0176] For example, the extraction unit 302 can be used to convert the current ambient noise channel to obtain the current ambient noise of the target channel, divide the current ambient noise of the target channel into blocks to obtain multiple audio data blocks, extract features from the audio data blocks to obtain the audio features corresponding to the audio data blocks.

[0177] (3) Determine unit 303;

[0178] The determination unit 303 is used to determine the type of ambient noise in the current environment based on audio characteristics.

[0179] For example, the determination unit 303 can be used to identify the candidate environmental noise type of the corresponding audio data block in the audio features using a trained noise recognition model, count the number of types corresponding to each environmental noise type in the candidate environmental noise types, and filter out the environmental noise type of the current environmental noise from the candidate environmental noise types based on the number of types.

[0180] (4) Add unit 304;

[0181] Adding unit 304 is used to add the environmental noise type to the data to be transmitted to obtain the target data to be transmitted.

[0182] For example, the addition unit 304 can be used to directly add the environmental noise type to the data to be transmitted to obtain the target data to be transmitted. Alternatively, it can filter out the data label corresponding to the environmental noise type from the preset data label set and add the data label to the data to be transmitted to obtain the target data to be transmitted. Or, it can convert the environmental noise type into environmental noise type data and add the environmental noise type data to the data to be transmitted to obtain the target data to be transmitted.

[0183] (5) Playback Unit 305;

[0184] The playback unit 305 is used to convert the target transmission data into audio data and play the audio data so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise.

[0185] For example, playback unit 305 can be used to encrypt the target transmitted data, generate a frequency mapping table based on the encrypted transmitted data, perform audio encoding on the target transmitted data based on the frequency mapping table to obtain the audio frequency points corresponding to the target transmitted data, map the frequency value of each audio frequency point in the frequency mapping table, generate an audio waveform based on the frequency value, and use the audio waveform as audio data. The audio data is then played so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise.

[0186] Optionally, the data transmission device may also include an identification unit 306, such as Figure 13 As shown, specifically, it can be done as follows:

[0187] The identification unit 306 is used to record the target audio data being played when audio data is detected to obtain transmission data.

[0188] For example, the identification unit 306 can be used to record the target audio data being played when audio data is detected to be playing, obtain the recording data, parse the recording data to obtain the current ambient noise type and the initial audio data, and identify the transmission data in the initial audio data according to the current ambient noise type.

[0189] Optionally, the data transmission device may also include a training unit 307, such as Figure 14 As shown, the specific details can be as follows:

[0190] Training unit 307 is used to train the preset noise type recognition model to obtain the trained noise type recognition model.

[0191] For example, training unit 307 can be used to acquire audio data samples, which include audio data labeled with environmental noise types. A preset noise type recognition model is used to predict the environmental noise type of the audio data samples to obtain the predicted environmental noise type. The preset noise type recognition model is then converged based on the labeled environmental noise type and the predicted environmental noise type to obtain the trained noise type recognition model.

[0192] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0193] As can be seen from the above, in this embodiment, after acquiring the data to be transmitted and the current ambient noise, feature extraction is performed on the current ambient noise to obtain the audio features of the current ambient noise. Then, based on the audio features, the ambient noise type of the current ambient noise is determined, and the ambient noise type is added to the data to be transmitted to obtain the target data to be transmitted. The target data to be transmitted is then converted into audio data and played, so that the receiving terminal can obtain the data to be transmitted from the audio data according to the ambient noise type. Since this scheme can acquire the current ambient noise, determine the ambient noise type of the current ambient noise, and add the ambient noise type to the data to be transmitted, the receiving terminal can extract the ambient noise type from the received audio data. Thus, the current network transmission quality (QoS) of the data to be transmitted can be determined based on the ambient noise type, and the data to be transmitted can be obtained based on the current network transmission quality. Therefore, the transmission success rate of data transmission can be improved.

[0194] This invention also provides an electronic device, such as... Figure 15 As shown, it illustrates a structural schematic diagram of the electronic device involved in an embodiment of the present invention, specifically:

[0195] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 15 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0196] The processor 401 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0197] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0198] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0199] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0200] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:

[0201] The system acquires the data to be transmitted and the current ambient noise, which is the sound collected in real time under the current transmission environment. It extracts features from the current ambient noise to obtain its audio features. Based on the audio features, it determines the type of ambient noise and adds the type of ambient noise to the data to be transmitted to obtain the target data to be transmitted. It then converts the target data to be transmitted into audio data and plays the audio data so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise.

[0202] For example, electronic devices can directly acquire user-input data as data to be transmitted, or they can receive business data returned by a business server or other terminals and use that business data as data to be transmitted. After acquiring the data to be transmitted, the sound acquisition device of the data transmission device acquires the sound of the current environment at a preset time to obtain the current environmental noise. Alternatively, the environmental sound of the current environment can be acquired in real time. When the data to be transmitted is acquired, the environmental sound acquired at the preset time at the current moment is extracted from the acquired environmental sound to obtain the current environmental noise. The current environmental noise is converted to obtain the current environmental noise of the target channel. The current environmental noise of the target channel is divided into blocks to obtain multiple audio data blocks. Feature extraction is performed on the audio data blocks to obtain the audio features corresponding to the audio data blocks. A trained noise recognition model is used to identify the candidate environmental noise types of the corresponding audio data blocks from the audio features. The number of types corresponding to each environmental noise type is counted in the candidate environmental noise types. Based on the number of types, the environmental noise type of the current environmental noise is selected from the candidate environmental noise types. The process involves several steps: First, adding environmental noise types to the data to be transmitted yields the target data. Alternatively, selecting data tags corresponding to the environmental noise type from a pre-defined data tag set and adding them to the data to be transmitted yields the target data. Second, converting the environmental noise type to environmental noise type data and adding it to the data to be transmitted yields the target data. The target data is then encrypted, and a frequency mapping table is generated based on this encrypted data. Audio encoding is performed on the target data based on the frequency mapping table to obtain the corresponding audio frequencies. The frequency value of each audio frequency is mapped in the frequency mapping table, and an audio waveform is generated based on this frequency value. This audio waveform is then used as the audio data. The audio data is played so that the receiving terminal can retrieve the data to be transmitted from the audio data based on the environmental noise type. When audio data playback is detected, the target audio data is recorded, and the recorded data is parsed to obtain the current environmental noise type and initial audio data. Based on the current environmental noise type, the transmitted data is identified from the initial audio data.

[0203] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0204] As can be seen from the above, in this embodiment of the invention, after acquiring the data to be transmitted and the current ambient noise, feature extraction is performed on the current ambient noise to obtain the audio features of the current ambient noise. Then, based on the audio features, the ambient noise type of the current ambient noise is determined, and the ambient noise type is added to the data to be transmitted to obtain the target data to be transmitted. The target data to be transmitted is then converted into audio data and played, so that the receiving terminal can obtain the data to be transmitted from the audio data according to the ambient noise type. Since this scheme can acquire the current ambient noise, determine the ambient noise type of the current ambient noise, and add the ambient noise type to the data to be transmitted, the receiving terminal can extract the ambient noise type from the received audio data. Thus, the current network transmission quality (QoS) of the data to be transmitted can be determined based on the ambient noise type, and the data to be transmitted can be obtained based on the current network transmission quality. Therefore, the transmission success rate of data transmission can be improved.

[0205] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0206] Therefore, embodiments of the present invention provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the data transmission methods provided in the embodiments of the present invention. For example, the instructions can execute the following steps:

[0207] The system acquires the data to be transmitted and the current ambient noise, which is the sound collected in real time under the current transmission environment. It extracts features from the current ambient noise to obtain its audio features. Based on the audio features, it determines the type of ambient noise and adds the type of ambient noise to the data to be transmitted to obtain the target data to be transmitted. It then converts the target data to be transmitted into audio data and plays the audio data so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise.

[0208] For example, user input data can be acquired as the data to be transmitted, or business data returned by a business server or other terminals can be received and used as the data to be transmitted. After acquiring the data to be transmitted, the sound acquisition device of the data transmission device collects the sound of the current environment at a preset time to obtain the current environmental noise. Alternatively, the environmental sound of the current environment can be collected in real time. When the data to be transmitted is acquired, the environmental sound collected at the preset time at the current moment is extracted from the collected environmental sound to obtain the current environmental noise. The current environmental noise is converted to obtain the current environmental noise of the target channel. The current environmental noise of the target channel is divided into blocks to obtain multiple audio data blocks. Feature extraction is performed on the audio data blocks to obtain the audio features corresponding to the audio data blocks. The trained noise recognition model is used to identify the candidate environmental noise types of the corresponding audio data blocks from the audio features. The number of types corresponding to each environmental noise type is counted in the candidate environmental noise types. Based on the number of types, the environmental noise type of the current environmental noise is selected from the candidate environmental noise types. The process involves several steps: First, adding environmental noise types to the data to be transmitted yields the target data. Alternatively, selecting data tags corresponding to the environmental noise type from a pre-defined data tag set and adding them to the data to be transmitted yields the target data. Second, converting the environmental noise type to environmental noise type data and adding it to the data to be transmitted yields the target data. The target data is then encrypted, and a frequency mapping table is generated based on this encrypted data. Audio encoding is performed on the target data based on the frequency mapping table to obtain the corresponding audio frequencies. The frequency value of each audio frequency is mapped in the frequency mapping table, and an audio waveform is generated based on this frequency value. This audio waveform is then used as the audio data. The audio data is played so that the receiving terminal can retrieve the data to be transmitted from the audio data based on the environmental noise type. When audio data playback is detected, the target audio data is recorded, and the recorded data is parsed to obtain the current environmental noise type and initial audio data. Based on the current environmental noise type, the transmitted data is identified from the initial audio data.

[0209] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0210] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0211] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the data transmission methods provided in the embodiments of the present invention, the beneficial effects that any of the data transmission methods provided in the embodiments of the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0212] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the data transmission or transaction data transmission aspects described above.

[0213] The present invention has provided a detailed description of a data transmission method, apparatus, electronic device, and computer-readable storage medium. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A data transmission method, characterized in that, include: Acquire the data to be transmitted and the current ambient noise, wherein the current ambient noise is the sound collected in real time under the current transmission environment; The current environmental noise is subjected to feature extraction to obtain the audio features of the current environmental noise; Based on the audio characteristics, the environmental noise type of the current ambient noise is determined; wherein, the environmental noise type refers to the scene type corresponding to the current transmission environment; The environmental noise type is added to the data to be transmitted to obtain the target transmission data; The target transmission data is converted into audio data and played so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise.

2. The data transmission method according to claim 1, characterized in that, Also includes: When audio data playback is detected, the target audio data being played is recorded to obtain the recording data; The recording data is parsed to obtain the current ambient noise type and initial audio data; Based on the current ambient noise type, the transmitted data is identified in the initial audio data.

3. The data transmission method according to claim 2, characterized in that, The step of identifying transmitted data from the initial audio data based on the current ambient noise type includes: Based on the current environmental noise type, the network transmission quality of the target audio data is evaluated to obtain evaluation information; Based on the evaluation information, determine the processing method for the target audio data; When the processing method is to receive the target audio data, the transmission data is identified in the initial audio data.

4. The data transmission method according to claim 3, characterized in that, The process of identifying transmitted data in the initial audio data includes: Baseband data is extracted from the initial audio data, and the baseband data is subjected to header detection based on the data block identifier in the baseband data; When a message header is present in the baseband data, the acoustic baseband data is extracted from the initial audio data; At least one data block identifier is extracted from the acoustic fundamental frequency data, and the extracted data block identifiers are corrected to obtain a set of data block identifiers; Obtain the target data block corresponding to each data block identifier in the data block identifier set, and merge the target data blocks to obtain the transmission data.

5. The data transmission method according to claim 3, characterized in that, After determining the processing method for the target audio data based on the evaluation information, the process further includes: When the processing method is to not receive the target audio data, the initial audio data is discarded and a prompt message is generated; The prompt message is sent to the sending terminal so that the sending terminal can replay the target audio data based on the prompt message.

6. The data transmission method according to any one of claims 1 to 5, characterized in that, The step of extracting features from the current environmental noise to obtain the audio features of the current environmental noise includes: The current ambient noise is converted into a channel to obtain the current ambient noise of the target channel; The current ambient noise of the target channel is divided into blocks to obtain multiple audio data blocks; Feature extraction is performed on the audio data block to obtain the audio features corresponding to the audio data block.

7. The data transmission method according to claim 6, characterized in that, The step of extracting features from the audio data blocks to obtain the audio features corresponding to each audio data block includes: Spectral image information is extracted from the audio data block, and the spectral image value corresponding to the audio data block is identified in the spectral image information; Obtain the data sampling amount corresponding to each audio feature dimension, and extract the initial audio feature corresponding to each audio feature dimension from the spectrum image information based on the data sampling amount. The initial audio features are concatenated with the spectral image values ​​to obtain the audio features corresponding to each audio data block.

8. The data transmission method according to claim 6, characterized in that, The step of determining the environmental noise type of the current environmental noise based on the audio features includes: The trained noise recognition model is used to identify candidate environmental noise types for the corresponding audio data blocks in the audio features; The number of types corresponding to each environmental noise type is counted among the candidate environmental noise types. Based on the number of types, the environmental noise type of the current environmental noise is selected from the candidate environmental noise types.

9. The data transmission method according to claim 8, characterized in that, The step of using a trained noise recognition model to identify candidate environmental noise types for corresponding audio data blocks in the audio features includes: The noise audio features are extracted from the audio features of each audio data block using a trained noise type recognition model to obtain the noise audio feature set of the current environmental noise. Batch normalization is performed on the noise audio features in the noise audio feature set to obtain the basic noise audio feature set; The environmental noise type corresponding to each basic noise audio feature is identified in the basic noise audio feature set to obtain the candidate environmental noise type of the audio data block.

10. The data transmission method according to claim 9, characterized in that, The step of identifying the environmental noise type of each basic noise audio feature set in the basic noise audio feature set to obtain the candidate environmental noise type of the audio data block includes: Obtain the time information of the audio data block in the current ambient noise, and sort the basic noise audio features in the basic noise audio feature set based on the time information; Based on the sorting information, the basic noise audio features in the basic noise audio feature set are converted into noise type features respectively; The environmental noise type corresponding to the noise type feature is selected from the preset set of environmental noise types to obtain the candidate environmental noise type of the audio data block.

11. The data transmission method according to claim 10, characterized in that, The step of converting the basic noise audio features in the basic noise audio feature set into noise type features based on the sorting information includes: The basic noise audio features that need to be converted are determined from the set of basic noise audio features to obtain the current basic noise audio features. Based on the sorting information, query the target basic noise audio feature that ranks before the current basic noise audio feature in the basic noise audio feature set; Based on the query results of the target basic noise audio features, the basic noise audio features in the basic noise audio feature set are converted into noise type features respectively.

12. The data transmission method according to claim 11, characterized in that, The step of converting the basic noise audio features in the basic noise audio feature set into noise type features based on the query results of the target basic noise audio features includes: When the target basic noise audio feature exists, the current basic noise audio feature is converted into a noise type feature based on the target basic noise audio feature; When the target basic noise audio feature does not exist, the target basic noise audio feature is used as the noise type feature; Return to the step of determining the basic noise audio feature that needs to be converted in the basic noise audio feature set, until all the basic noise audio features in the basic noise audio feature set are converted into noise type features, and obtain the noise type feature corresponding to each basic noise audio feature.

13. The data transmission method according to claim 12, characterized in that, The step of converting the current basic noise audio features into noise type features based on the target basic noise audio features includes: Obtain the hidden state features corresponding to the target basic noise audio features, and the hidden state features are used to indicate the hidden state passed in the process of converting the target basic noise audio features into noise type features; Calculate the feature ratio between the current basic noise audio features and the latent state features, and determine the target latent state features corresponding to the current basic noise audio features based on the feature ratio; The target hidden state features are transformed to obtain the noise type features corresponding to the current basic noise audio features.

14. The data transmission method according to claim 9, characterized in that, Before extracting noise audio features from the audio features of each audio data block using the trained noise type recognition model, the method further includes: Acquire audio data samples, which include audio data labeled with environmental noise types; The environmental noise type of the audio data sample is predicted using a preset noise type identification model to obtain the predicted environmental noise type; The preset noise type recognition model is converged based on the labeled environmental noise type and the predicted environmental noise type to obtain the trained noise type recognition model.

15. The data transmission method according to any one of claims 1 to 5, characterized in that, The step of converting the target transmission data into audio data includes: The target transmitted data is encrypted, and a frequency mapping table is generated based on the encrypted transmitted data; Based on the frequency mapping table, the target transmission data is audio encoded to obtain the audio frequency points corresponding to the target transmission data; The frequency value of each audio frequency point is mapped in the frequency mapping table, and an audio waveform is generated based on the frequency value, which is then used as audio data.

16. The data transmission method according to claim 15, characterized in that, The step of performing audio encoding on the target transmission data based on the frequency mapping table to obtain the audio frequency point corresponding to the target transmission data includes: The target transmission data is divided into multiple data blocks to obtain a first data block, and the data identifier code of the target transmission data generated based on the current time is divided into multiple data blocks to obtain a second data block; At least one message header is selected from a preset message header set, and the message header is used as the third data block; Error correction coding is performed on the first data block, the second data block, and the third data block, and the error correction coding is divided into multiple data blocks to obtain the fourth data block; Audio frequency points are extracted from the first data block, the second data block, the third data block, and the fourth data block to obtain the audio frequency points corresponding to the target transmission data.

17. A data transmission device, characterized in that, include: The acquisition unit is used to acquire the data to be transmitted and the current ambient noise, wherein the current ambient noise is the sound collected in real time under the current transmission environment; An extraction unit is used to extract features from the current environmental noise to obtain the audio features of the current environmental noise; The determining unit is configured to determine the environmental noise type of the current ambient noise based on the audio features; wherein, the environmental noise type refers to the scene type corresponding to the current transmission environment; An adding unit is used to add the environmental noise type to the data to be transmitted to obtain the target transmission data; The playback unit is used to convert the target transmission data into audio data and play the audio data so that the receiving terminal can obtain the data to be transmitted from the audio data according to the type of ambient noise.

18. The data transmission apparatus according to claim 17, characterized in that, The data transmission device further includes an identification unit. The identification unit is specifically used to record the target audio data being played when audio data is detected to obtain recording data; to parse the recording data to obtain the current ambient noise type and initial audio data; and to identify the transmission data in the initial audio data according to the current ambient noise type.

19. The data transmission apparatus according to claim 18, characterized in that, The identification unit is specifically used to evaluate the network transmission quality of the target audio data according to the current environmental noise type, and obtain evaluation information; based on the evaluation information, determine the processing method for the target audio data; When the processing method is to receive the target audio data, the transmission data is identified in the initial audio data.

20. The data transmission apparatus according to claim 19, characterized in that, The identification unit is specifically used to extract the baseband data from the initial audio data and perform message header detection on the baseband data according to the data block identifier in the baseband data; when there is a message header in the baseband data, the sound wave baseband data is extracted from the initial audio data. At least one data block identifier is extracted from the acoustic fundamental frequency data, and the extracted data block identifiers are corrected to obtain a set of data block identifiers; the target data block corresponding to each data block identifier in the set of data block identifiers is obtained, and the target data blocks are fused to obtain the transmission data.

21. The data transmission apparatus according to claim 19, characterized in that, The identification unit is specifically used to discard the initial audio data and generate a prompt message when the processing method is not to receive the target audio data; and to send the prompt message to the sending terminal so that the sending terminal can replay the target audio data based on the prompt message.

22. The data transmission apparatus according to any one of claims 17 to 21, characterized in that, The extraction unit is specifically used to convert the current ambient noise channel to obtain the current ambient noise of the target channel; divide the current ambient noise of the target channel into blocks to obtain multiple audio data blocks; and extract features from the audio data blocks to obtain the audio features corresponding to the audio data blocks.

23. The data transmission apparatus according to claim 22, characterized in that, The extraction unit is specifically used to extract spectral image information from the audio data block, and identify the spectral image value corresponding to the audio data block in the spectral image information; obtain the data sampling amount corresponding to each audio feature dimension, and extract the initial audio feature corresponding to each audio feature dimension in the spectral image information based on the data sampling amount. The initial audio features are concatenated with the spectral image values ​​to obtain the audio features corresponding to each audio data block.

24. The data transmission apparatus according to claim 22, characterized in that, The determining unit is specifically used to identify candidate environmental noise types of corresponding audio data blocks in the audio features using a trained noise recognition model; to count the number of types corresponding to each environmental noise type in the candidate environmental noise types; and to filter out the environmental noise type of the current environmental noise from the candidate environmental noise types based on the number of types.

25. The data transmission apparatus according to claim 24, characterized in that, The determining unit is specifically used to extract noise audio features from the audio features of each audio data block using a trained noise type recognition model, thereby obtaining a set of noise audio features for the current environmental noise. Batch normalization is performed on the noise audio features in the noise audio feature set to obtain the basic noise audio feature set; The environmental noise type corresponding to each basic noise audio feature is identified in the basic noise audio feature set to obtain the candidate environmental noise type of the audio data block.

26. The data transmission apparatus according to claim 25, characterized in that, The determining unit is specifically used to obtain the time information of the audio data block in the current environmental noise, and to sort the basic noise audio features in the basic noise audio feature set based on the time information; Based on the sorting information, the basic noise audio features in the basic noise audio feature set are converted into noise type features respectively; The environmental noise type corresponding to the noise type feature is selected from the preset set of environmental noise types to obtain the candidate environmental noise type of the audio data block.

27. The data transmission apparatus according to claim 26, characterized in that, The determining unit is specifically used to determine the basic noise audio features that need to be converted from the basic noise audio feature set, and to obtain the current basic noise audio features. Based on the sorting information, a target basic noise audio feature that ranks first among the current basic noise audio features is queried in the basic noise audio feature set; according to the query result of the target basic noise audio feature, the basic noise audio features in the basic noise audio feature set are converted into noise type features respectively.

28. The data transmission apparatus according to claim 27, characterized in that, The determining unit is specifically used to convert the current basic noise audio feature into a noise type feature based on the target basic noise audio feature when the target basic noise audio feature exists. When the target basic noise audio feature does not exist, the target basic noise audio feature is used as the noise type feature; Return to the step of determining the basic noise audio feature that needs to be converted in the basic noise audio feature set, until all the basic noise audio features in the basic noise audio feature set are converted into noise type features, and obtain the noise type feature corresponding to each basic noise audio feature.

29. The data transmission apparatus according to claim 28, characterized in that, The determining unit is specifically used to obtain the hidden state feature corresponding to the target basic noise audio feature, the hidden state feature being used to indicate the hidden state passed during the process of converting the target basic noise audio feature into a noise type feature; calculate the feature ratio between the current basic noise audio feature and the hidden state feature, and determine the target hidden state feature corresponding to the current basic noise audio feature based on the feature ratio; and perform a dimensional transformation on the target hidden state feature to obtain the noise type feature corresponding to the current basic noise audio feature.

30. The data transmission apparatus according to claim 24, characterized in that, The data transmission device may further include a training unit. The training unit is specifically used to acquire audio data samples, which include audio data labeled with environmental noise types; to predict the environmental noise type of the audio data samples using a preset noise type recognition model, thereby obtaining the predicted environmental noise type; and to converge the preset noise type recognition model based on the labeled environmental noise type and the predicted environmental noise type, thereby obtaining the trained noise type recognition model.

31. The data transmission apparatus according to any one of claims 17 to 21, characterized in that, The playback unit is specifically used to encrypt the target transmission data and generate a frequency mapping table based on the encrypted transmission data; and to perform audio encoding on the target transmission data based on the frequency mapping table to obtain the audio frequency points corresponding to the target transmission data. The frequency value of each audio frequency point is mapped in the frequency mapping table, and an audio waveform is generated based on the frequency value, which is then used as audio data.

32. The data transmission apparatus according to claim 31, characterized in that, The playback unit is specifically used to divide the target transmission data into multiple data blocks to obtain a first data block, and to divide the data identifier code of the target transmission data generated based on the current time into multiple data blocks to obtain a second data block; and to select at least one message header from a preset message header set and use the message header as a third data block; Error correction coding is performed on the first data block, the second data block, and the third data block, and the error correction coding is divided into multiple data blocks to obtain a fourth data block; audio frequency points are extracted from the first data block, the second data block, the third data block, and the fourth data block to obtain the audio frequency points corresponding to the target transmission data.

33. An electronic device, characterized in that, It includes a processor and a memory, the memory storing an application program, and the processor running the application program within the memory to perform the steps of the data transmission method according to any one of claims 1 to 16.

34. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the data transmission method according to any one of claims 1 to 16.

35. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the data transmission method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Audio optimization method and device

    CN109087659A