Audio transmission method and device, electronic equipment and readable storage medium

By introducing environment type identification and audio delivery level setting in traditional audio delivery technology, the problem that traditional technology cannot adapt to diverse environments is solved, and the efficient and accurate transmission of audio signals in complex environments is achieved, and the user experience is improved.

CN119946501APending Publication Date: 2025-05-06BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510096163.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional audio transmission technology cannot adapt to diverse environments, resulting in background noise covering up important audio information and cannot intelligently distinguish the attention of audio signals in different scenarios, affecting the user's auditory experience.

Method used

By determining the environment type based on the ambient audio signal, setting the audio transmission level of each ambient audio signal, and transmitting the target audio signal according to the level, accurately adapting to audio processing and scene requirements.

Benefits of technology

It realizes the best state of maintaining the quality of audio signal transmission in complex environments, ensuring that audio content is received and understood clearly and accurately in various environments, and improving the user's auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946501A_ABST
    Figure CN119946501A_ABST
Patent Text Reader

Abstract

The invention provides an audio transmission method and device, electronic equipment and a readable storage medium, and belongs to the technical field of audio transmission, and the method comprises the steps: determining an environment type based on a plurality of environment audio signals in an environment; determining an audio transmission level corresponding to each environment audio signal in the environment based on the environment type; and transmitting a target audio signal based on the audio transmission level, wherein the target audio signal is a signal in the plurality of environment audio signals. According to the audio transmission method and device, the electronic equipment and the readable storage medium provided by the invention, the auditory experience of the user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of audio transmission, and more specifically, relates to an audio transmission method and device, an electronic device, and a readable storage medium. Background Art

[0002] In today's digital age, audio is widely used in various fields, but traditional audio transmission technology has many defects. As people's environments become increasingly diverse, from crowded shopping malls, quiet and rigorous laboratories, to construction sites with complex sounds, a single fixed audio transmission mode can no longer meet the needs. In a noisy environment, background noise often masks important audio information, resulting in obstruction of information transmission. In addition, people pay different attention to various audio signals in different scenarios. Traditional technology cannot intelligently distinguish and balance the transmission proportion of each audio signal, affecting the user's auditory experience. Summary of the invention

[0003] The purpose of the present disclosure is to provide an audio transmission method and device, an electronic device, and a readable storage medium to enhance the user's auditory experience.

[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an audio transmission method, comprising: determining a type of environment based on a plurality of ambient audio signals in the environment; Determine, based on the environment type, an audio transmission level corresponding to each ambient audio signal in the environment; A target audio signal is delivered based on the audio delivery level, the target audio signal being a signal among a plurality of ambient audio signals.

[0005] According to a second aspect of the embodiments of the present disclosure, there is provided an audio transmission device, including: An environment type determination module, used to determine the environment type based on multiple ambient audio signals in the environment; An audio delivery level determination module, used to determine the audio delivery level corresponding to each ambient audio signal in the environment based on the environment type; The audio delivery module is configured to deliver a target audio signal based on the audio delivery level, wherein the target audio signal is a signal among a plurality of ambient audio signals.

[0006] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the above-mentioned audio transmission method when executing the computer program.

[0007] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned audio transmission method are implemented.

[0008] The beneficial effects of the audio transmission method and device, electronic device, and readable storage medium provided by the embodiments of the present disclosure are: the embodiments of the present disclosure can accurately determine the type of environment based on various types of ambient audio signals in the environment, so that the audio processing can be adapted to the actual scene; a corresponding audio transmission level is set for each ambient audio signal according to the environmental category, ensuring that the target audio signal can maintain the best transmission state in various complex environments. Transmitting the target audio signal according to the audio transmission level not only improves the transmission quality of the audio signal, but also ensures that the audio content can be clearly and accurately received and understood in various environments, thereby improving the user's auditory experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0010] Figure 1 A flowchart of an audio transmission method provided by an embodiment of the present disclosure; Figure 2 A structural block diagram of an audio transmission device provided by an embodiment of the present disclosure; Figure 3 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0011] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present disclosure. However, it should be clear to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present disclosure with unnecessary details.

[0012] In order to make the purpose, technical solutions and advantages of the present disclosure more clear, specific embodiments will be described below in conjunction with the accompanying drawings.

[0013] Please refer to Figure 1 , Figure 1 A flowchart of an audio transmission method provided by an embodiment of the present disclosure is provided, and the method includes: S101: Determine an environment type based on a plurality of ambient audio signals in the environment.

[0014] In this embodiment, feature extraction is performed on a plurality of ambient audio signals to obtain feature parameters corresponding to each ambient audio signal; and the feature parameters are classified to obtain environment types corresponding to the plurality of ambient audio signals.

[0015] In this embodiment, there are multiple different audio signals in each environment. For example, on a noisy street, there are the sounds of vehicles driving, people talking, and broadcasts from street shops; in a quiet library, there are the sounds of flipping books and occasional footsteps; and in the forest, there are the sounds of birds singing, wind, and rustling leaves.

[0016] In this embodiment, feature extraction is performed on multiple ambient audio signals to obtain frequency features, amplitude features, rhythm features, etc. corresponding to each ambient audio signal; the feature parameters can be classified based on a decision tree to obtain the environment types corresponding to the multiple ambient audio signals.

[0017] Decision trees can achieve classification by making conditional judgments on input feature parameters. The internal nodes of a decision tree represent the test conditions for a certain feature parameter. For example, for frequency features, a node can determine whether the frequency is greater than a set threshold. Branches can represent different test results (yes or no), and leaf nodes correspond to the final classification category, that is, the environment type.

[0018] For example, starting from the root node, it is assumed that the judgment is first made based on the frequency characteristics. If the frequency characteristics of an audio signal show that its frequency is low (for example, lower than a certain Hertz value), the corresponding branch is followed to the next layer of nodes, and the next layer of nodes can then be further judged based on the amplitude characteristics. If the amplitude is large (greater than a certain set amplitude standard), the next layer of nodes is entered to continue judging based on the rhythm characteristics. After layers of progressive conditional judgments, when the leaf node is finally reached, the type of environment to which the audio signal belongs can be determined, such as a "noisy street" environment.

[0019] S102: Determine, based on the environment type, the audio transmission level corresponding to each ambient audio signal in the environment.

[0020] In this embodiment, different environment types have different requirements for audio transmission. The importance of each environmental audio signal in the environment is graded based on the characteristics of the environment type.

[0021] For example, assuming that in a noisy factory workshop environment, the roar of the normal operation of the machine can be at a background low transmission level, because everyone is accustomed to the roar of the normal operation of the machine, and it is not a carrier of critical information transmission, but audio signals such as equipment failure alarm sounds that are related to production safety can be set to a high audio transmission level to facilitate staff to detect and deal with problems in a timely manner.

[0022] For example, assuming that there is a lot of background noise in a busy market, some relatively small and less important environmental audio signals (such as whispered conversations in the distance) can be assigned a lower audio transmission level, indicating that they can be appropriately weakened or ignored in subsequent processing. On the other hand, some more prominent audio signals that can represent the characteristics of the environment (such as shouting in the market) can be assigned a higher audio transmission level, and their presentation in the audio transmission process should be given special consideration.

[0023] S103: Deliver a target audio signal based on the audio delivery level, where the target audio signal is a signal among the multiple ambient audio signals.

[0024] In this embodiment, in response to the audio delivery level of the target audio signal being high, the signal strength of the target audio signal is enhanced; in response to the audio delivery level of the target audio signal being low, the signal strength of the target audio signal is reduced.

[0025] In this embodiment, if the audio transmission level of the target audio signal is high, measures such as enhancing the signal strength and optimizing the signal quality (such as improving clarity and reducing distortion) can be taken during the transmission process to enable it to be transmitted more clearly and effectively, so that the receiving end can better receive the important audio content. For those target audio signals with lower audio transmission levels, they can be suppressed to reduce the signal strength of the target audio signal, or reduce its proportion in the overall transmitted audio, or even choose not to transmit it to avoid interference with the transmission of more important audio signals.

[0026] For example, assuming that in a speech scene with background music, the speaker's voice is an important target audio signal with a high audio transmission level, which can be transmitted through audio equipment with appropriate volume and clear sound quality. If the background music is of a lower level, the volume can be appropriately turned down or faded so that the overall audio transmission effect meets the requirements of the speech scene.

[0027] From the above, it can be concluded that this embodiment can accurately determine the type of environment based on various types of ambient audio signals in the environment, so that the audio processing can be adapted to the actual scene; the corresponding audio transmission level is set for each ambient audio signal according to the environment category, ensuring that the target audio signal can maintain the best transmission state in various complex environments. Transmitting the target audio signal according to the audio transmission level not only improves the transmission quality of the audio signal, but also ensures that the audio content can be clearly and accurately received and understood in various environments, thereby improving the user's auditory experience.

[0028] In one embodiment of the present disclosure, before determining the environment type based on multiple ambient audio signals in the environment, the method further includes: Performing short-time Fourier transform on the mixed audio signal in the environment to obtain a plurality of time-frequency units, wherein any time-frequency unit includes an amplitude spectrum corresponding to the mixed audio signal; The magnitude spectrum is fed into a convolutional mental network to determine the cluster assignment probability of each time-frequency unit; Based on the clustering assignment probability of each time-frequency unit, multiple time-frequency units are classified to obtain multiple environmental audio signals.

[0029] In this embodiment, the short-time Fourier transform can convert the mixed audio signal that changes continuously over time into the time-frequency domain, and divide the mixed audio signal into a plurality of small time-frequency units.

[0030] By performing discrete Fourier transform on the audio signal within a certain time window, the time window can slide along the time axis to obtain the frequency component information corresponding to different moments, that is, the amplitude spectrum contained in each time-frequency unit.

[0031] The amplitude spectrum can reflect the amplitude of each frequency component in the corresponding time segment, and decompose the mixed audio signal from a complex waveform in the time domain into small units with clear frequency and corresponding amplitude characteristics in the time-frequency plane, laying the foundation for further analysis and distinction of different sound components.

[0032] In this embodiment, the amplitude spectrum of each time-frequency unit is input into a convolutional neural network, and the convolutional neural network can automatically learn the deep feature information contained in the amplitude spectrum.

[0033] The convolution layer in the convolutional neural network performs convolution operations by sliding the convolution kernel on the amplitude spectrum to extract the amplitude features of different local areas; the pooling layer performs dimensionality reduction and compression on the amplitude features to obtain the key amplitude features; after processing by multiple such layers, the fully connected layer integrates the amplitude features and the key amplitude features to obtain the probability that each time-frequency unit belongs to a different category, that is, the clustering assignment probability.

[0034] For example, in a street environment with the sounds of driving vehicles, pedestrians talking, and music from street shops, through short-time Fourier transform, it is possible to see the frequency characteristics of each of these different sounds and how their amplitudes change over time in different time-frequency units. For example, the amplitude corresponding to the low-frequency sound of a vehicle engine in certain time-frequency units is different from the amplitude corresponding to the high-frequency melody in the shop music.

[0035] For a certain time-frequency unit, the output is how likely it is to belong to the category of vehicle driving sound, how likely it is to belong to the category of people talking sound, etc. This probability reflects the degree of match between the feature pattern learned by the network and the different categories of sound of the time-frequency unit.

[0036] In this embodiment, the category with the highest probability of clustering assignment can be used as the category to which it belongs. In this way, the time-frequency units originally mixed together can be distinguished and correspond to different categories, which represent different environmental audio signals.

[0037] Exemplarily, in the above-mentioned street environment, the time-frequency units determined to belong to the category of vehicle driving sounds are grouped into one group, those belonging to the category of pedestrian conversation sounds are grouped into another group, and so on, ultimately achieving the separation of multiple different and relatively independent environmental audio signals from the mixed audio signal, providing a basic and clear source of audio data for subsequent operations such as determining the environment type based on these environmental audio signals, so that the entire audio transmission method can be carried out in an orderly manner according to the process.

[0038] From the above, it can be concluded that this embodiment processes the mixed audio signals in the environment through short-time Fourier transform, which can convert them into multiple time-frequency units with amplitude spectra, break the chaotic state of the time-domain audio signal, and make the hidden frequency and amplitude information intuitively presented, laying the foundation for subsequent precise analysis. The amplitude spectrum is input into the convolutional neural network, and its excellent feature extraction and pattern recognition expertise is used to accurately determine the clustering allocation probability of each time-frequency unit, and efficiently distinguish different sound categories. Finally, based on these probabilistic classifications, multiple environmental audio signals are successfully stripped out from the mixed audio, which greatly improves the accuracy of the preliminary work of audio processing and provides strong support for subsequent environmental judgment and audio transmission process optimization.

[0039] In one embodiment of the present disclosure, any time-frequency unit further includes a phase spectrum corresponding to the mixed audio signal; Based on the clustering assignment probability of each time-frequency unit, multiple time-frequency units are classified to obtain multiple environmental audio signals, including: Classifying multiple time-frequency units based on the clustering assignment probability of each time-frequency unit to obtain multiple initial environmental audio signals; Based on the phase spectrum corresponding to the mixed audio signal, an inverse short-time Fourier transform is performed on the multiple initial ambient audio signals to obtain multiple ambient audio signals.

[0040] In this embodiment, in addition to the amplitude spectrum corresponding to the mixed audio signal, the short-time Fourier transform can also obtain the phase spectrum. The amplitude spectrum reflects the amplitude of each frequency component, while the phase spectrum reflects the phase information of the signal at the corresponding frequency and time point. The phase information determines the relative position relationship and synthesis between different frequency components. For example, the amplitude spectrum is the "loudness description" of the sound, and the phase spectrum is the "position and mutual relationship description" of the sound. The combination of the two can more comprehensively reflect the characteristics of the audio signal in the time-frequency unit.

[0041] In this embodiment, the amplitude spectrum of the time-frequency unit is input into a convolutional neural network to obtain the clustering assignment probability of each time-frequency unit; then, a suitable classification strategy is adopted to classify each time-frequency unit based on this probability, and the sets of time-frequency units corresponding to these categories form multiple initial ambient audio signals.

[0042] The initial ambient audio signal is distinguished only based on the amplitude spectrum characteristics and clustering assignment probability, lacking comprehensive consideration of phase information.

[0043] In this embodiment, the inverse short-time Fourier transform is the inverse process of the short-time Fourier transform, and its purpose is to restore the information in the time-frequency domain back to the audio signal in the time domain.

[0044] Since the phase spectrum contains the phase relationship of different frequency components in the time-frequency unit, during the inverse short-time Fourier transform, the phase spectrum can be used together with the information of the amplitude spectrum to accurately combine the audio information carried by each time-frequency unit according to the correct phase relationship, frequency and amplitude, and restore multiple more accurate and complete ambient audio signals. The final ambient audio signal is not only accurately classified, but also can truly restore the actual situation of each ambient sound in the original mixed audio in the time domain.

[0045] From the above, it can be concluded that when processing the time-frequency unit, this embodiment adds a phase spectrum in addition to the amplitude spectrum, which can fully capture the characteristics of the audio signal. The relative position information of each frequency component contained in the phase spectrum complements the amplitude information of the amplitude spectrum, allowing the time-frequency unit to more accurately describe the audio. Then, multiple initial environmental audio signals are classified according to the clustering allocation probability to preliminarily achieve sound differentiation. Finally, the phase spectrum is used to perform an inverse short-time Fourier transform to restore the audio, so that the multiple environmental audio signals finally obtained are not only accurately classified, but also completely restore the environmental sounds in the original mixed audio in the time domain.

[0046] In one embodiment of the present disclosure, determining the environment type based on multiple ambient audio signals in the environment includes: The environment type is determined based on similarities between the plurality of environmental audio signals and the target environment type.

[0047] In this embodiment, the target environment type may be a pre-set or known environment type, such as a quiet library environment, a noisy factory workshop environment, a busy shopping mall environment, etc.

[0048] Calculate the similarity between multiple environmental audio signals and each target environment type. The similarity reflects the degree of fit between the current environmental audio signal and the typical audio features of a target environment type. Select the target environment type with the highest similarity as the actual type of the current environment. For example, after calculation, it is found that the similarity between the current environmental audio signal and the library environment is the highest among all target environment types, then the current environment is determined to be a library environment.

[0049] In this embodiment, feature matrices of multiple dimensions are extracted from each ambient audio signal, a relative deviation matrix on each dimension is calculated, a corresponding weight matrix is ​​constructed based on the feature matrices of multiple dimensions, and comprehensive deviation values ​​of multiple ambient audio signals and the target environment type are obtained based on the feature matrices of multiple dimensions and the corresponding weight matrix. Based on the comprehensive deviation values, the similarity between the multiple ambient audio signals and the target environment type is obtained.

[0050] In this embodiment, assuming that there are m ambient audio signals, for each ambient audio signal, a feature matrix of k dimensions is extracted, which is recorded as (where i = 1, 2, ...; m represents the i-th audio signal; j = 1, 2, ...; k represents the j-th feature dimension). For the target environment type, there is also a corresponding standard feature matrix T j .

[0051] First, calculate the relative deviation matrix on each feature dimension

[0052] in, Expressed as a positive coefficient to prevent the denominator from being zero.

[0053] Then, construct the weight matrix ,The weights can be pre-set according to the importance of different ,feature dimensions in a specific environment, and can be determined through a large amount of ,experimental data or the experience of domain experts.,For example, in a hospital environment, the frequency feature weight of the ,equipment operating sound is higher, and the amplitude feature weight of the ,human voice is relatively low.

[0054] Calculate the comprehensive deviation values ​​of multiple environmental audio signals from the target environment type:

[0055] in, It is expressed as a comprehensive deviation value between multiple environmental audio signals and the target environment type.

[0056] The similarity calculation formula is:

[0057] in, It is expressed as the similarity between the ambient audio signal and the target environment type. It is expressed as an adjustment coefficient, which is used to control the sensitivity of similarity to changes in deviation and can be adjusted according to the actual application scenario.

[0058] In this embodiment, compared with the traditional absolute difference calculation, the relative deviation can better adapt to the dynamic range of the characteristic values ​​in different environments, and can accurately reflect the degree of deviation from the actual audio signal even if the standard characteristic value of the target environment type changes greatly. Different weights are given to the characteristic dimensions according to different environments, so that the formula can flexibly and accurately measure the similarity in different scenarios. For example, in a factory workshop, the frequency and amplitude of the roar of the machine are considered, while in an office, the rhythm and amplitude of the human voice are more concerned, so as to meet the complex needs in actual applications and more accurately determine the environment type.

[0059] In one embodiment of the present disclosure, determining the audio delivery level corresponding to each ambient audio signal in the environment based on the environment type includes: Extract features from each ambient audio signal to obtain key feature information; The audio transmission level corresponding to each ambient audio signal in the environment is determined based on the correlation between the key feature information and the type of the environment.

[0060] In this embodiment, the ambient audio signal contains a lot of complex acoustic information. Through the feature extraction operation, key feature information that can reflect its essential characteristics can be mined.

[0061] In this embodiment, feature extraction can be performed on each ambient audio signal to obtain frequency features, amplitude features and rhythm features; frequency features may include main frequency, bandwidth, etc.; amplitude features may include average amplitude, peak amplitude, etc.; rhythm features may include beat interval duration, rhythm change rate, etc.

[0062] Frequency characteristics can reflect the pitch and frequency distribution of sound, amplitude characteristics can reflect the loudness of sound, and rhythm characteristics can characterize the regular characteristics of sound from the time dimension.

[0063] In this embodiment, different environment types have their own typical audio feature performances and people's expectations of audio characteristics.

[0064] Take a quiet library environment as an example. In this environment, audio signals such as faint sounds of turning pages and light footsteps have low frequencies, small amplitudes and slow rhythms. These key feature information are highly consistent with the quiet environment type of the library. Such audio signals can be judged to have a high correlation with the environment and can be assigned a lower audio transmission level accordingly, indicating that they should be weakened as much as possible during the audio transmission process to avoid interfering with the overall quiet atmosphere.

[0065] In this embodiment, each key feature is quantized to obtain a vector feature; a weight is assigned to each vector feature; then, the correlation between the key feature information and the type of environment is obtained based on the inverse of the weighted Euclidean distance; finally, the audio transmission level corresponding to each ambient audio signal in the environment is determined based on the correlation between the key feature information and the type of environment.

[0066] It can be concluded from the above that this embodiment can reasonably determine the transmission level based on the typical audio characteristics and expectations of different environments, combined with the reciprocal weighted Euclidean distance to calculate the correlation, and can effectively adapt to various environments and optimize the audio transmission effect.

[0067] In one embodiment of the present disclosure, it also includes: Decompose each ambient audio signal into initial wavelet coefficients of different frequencies; The wavelet coefficients are processed based on a preset threshold value to obtain target wavelet coefficients; Perform inverse wavelet transform on the target wavelet coefficients to obtain each ambient audio signal after denoising.

[0068] In this embodiment, each ambient audio signal is decomposed into initial wavelet coefficients of different frequencies based on wavelet transform. Wavelet transform can decompose the ambient audio signal into initial wavelet coefficients of different frequencies.

[0069] Through wavelet transform, the originally complex audio signal can be decomposed into initial wavelet coefficients corresponding to different frequency components. The initial wavelet coefficients carry the information of the audio signal in different frequency bands, so that the information of different frequencies can be processed separately later.

[0070] For example, for an ambient audio signal containing a mixture of multiple sounds, such as street ambient audio that contains both the low-frequency roar of driving vehicles and high-frequency sounds such as birdsong, the wavelet transform can clearly decompose it into the wavelet coefficients corresponding to the low-frequency part and the wavelet coefficients of the high-frequency part, etc., providing a detailed data basis at the frequency level for further processing.

[0071] In actual environmental audio signals, noise is inevitable, and this noise will also be reflected in the wavelet coefficients. The preset threshold is set to distinguish the wavelet coefficients that truly represent effective audio information from the wavelet coefficients that are mainly generated by noise.

[0072] The soft threshold or hard threshold processing method can be used. In hard threshold processing, if the absolute value of the wavelet coefficient is less than the preset threshold, it is directly set to zero, and if it is greater than the threshold, it is retained unchanged, which is equivalent to directly removing the wavelet coefficients that are considered to be noise. In soft threshold processing, when the absolute value of the wavelet coefficient is less than the threshold, it is set to zero. For the wavelet coefficients greater than the threshold, a certain shrinkage adjustment can be made to make it close to zero to a certain extent before being retained, which can minimize the impact on the effective audio information while removing the noise.

[0073] In this embodiment, the processed target wavelet coefficients are restored to a time-domain audio signal. Since the wavelet coefficients corresponding to the noise are removed or reasonably adjusted by threshold processing, the remaining valid target wavelet coefficients can be used to combine and restore them into a new audio signal according to the inverse operation rules of the wavelet transform when performing the inverse wavelet transform. Compared with the original ambient audio signal, the noise component of this audio signal is effectively suppressed, thereby achieving the denoising effect, so that when the environment type is determined based on these denoised ambient audio signals, the audio transmission level is set, etc., there can be purer and more accurate audio data as the basis, thereby improving the reliability and accuracy of the entire audio transmission method.

[0074] In this embodiment, the wavelet transform is expressed as:

[0075] in, represents the initial wavelet coefficient, j represents the scale parameter, k represents the displacement parameter, represents a discrete signal sequence (corresponding to the discretized ambient audio signal), represents discrete wavelet basis function. It is expressed as:

[0076] The discrete wavelet basis function is realized by discretely taking values ​​of scale parameters and displacement parameters according to specific rules. For example, a=2 j , b=k2 j (j, k∈Z, Z is an integer set). Where a is the scale factor and b is the translation factor.

[0077] By changing the scale parameter j and the displacement parameter k, the entire ambient audio signal is decomposed to obtain the initial wavelet coefficients at different scales (i.e. different frequencies) .

[0078] Taking the hard threshold method as an example to determine the preset threshold, the hard threshold function is defined as:

[0079] in, represents the target wavelet coefficient, and T represents the preset threshold.

[0080] Through threshold processing, the wavelet coefficients with smaller amplitudes, which are mainly generated by noise, can be processed to obtain target wavelet coefficients that can better represent the effective audio components.

[0081] The inverse wavelet transform of the target wavelet coefficients is expressed as:

[0082] The above formula can be used to restore the ambient audio signal after denoising. The denoised audio signal will be used as a purer and more reliable data basis in the subsequent steps of determining the environment type and audio transmission level based on the environmental audio signal, which will help improve the accuracy and performance of the entire audio transmission method.

[0083] From the above, it can be concluded that this embodiment decomposes the environmental audio signal into initial wavelet coefficients of different frequencies, accurately analyzes the signal frequency composition, and lays the foundation for subsequent fine processing. Then, the wavelet coefficients are processed according to the preset threshold, which can effectively identify and eliminate the noise components, and accurately screen out the target wavelet coefficients representing the effective audio. Finally, after the inverse wavelet transform is restored, the denoised audio signal is obtained, which provides pure data for the subsequent determination of the environment type and audio transmission level, greatly improving the accuracy of audio processing.

[0084] In one embodiment of the present disclosure, it also includes: The preset threshold is determined based on the audio delivery level corresponding to each environmental audio signal in the environment.

[0085] In this embodiment, different environments have their own corresponding audio transmission levels, and the audio transmission level reflects the importance attached to various audio signals and the expected audio presentation effect in the environment.

[0086] After the audio transmission level corresponding to each environmental audio signal in the environment is determined, the preset threshold value may be determined according to the audio transmission level.

[0087] In this embodiment, the preset threshold is determined based on a first formula, and the first formula is:

[0088] Where T represents the preset threshold, W represents the weight value of the corresponding environment, Represents the average amplitude of the ambient audio signal. Indicates the main frequency corresponding to the ambient audio signal, Indicates the energy concentration of the signal (the degree to which the signal energy is concentrated around the main frequency). Indicates the basic deviation value.

[0089] From the above, it can be concluded that this embodiment is closely integrated with environmental characteristics. The audio transmission levels in different environments are different, reflecting the different requirements for audio purity and prominence. For example, in a quiet conference room, a high level requirement corresponds to low noise tolerance, and a low threshold can be accurately set for deep denoising; while a high threshold can be set in a noisy market to avoid excessive removal of effective information. Accurately adapted thresholds can improve the denoising effect and provide a data foundation that better meets the needs for audio transmission.

[0090] Corresponding to the audio transmission method of the above embodiment, Figure 2 This is a structural block diagram of an audio transmission device provided by an embodiment of the present disclosure. For ease of description, only the parts related to the embodiment of the present disclosure are shown. Figure 2 The audio transmission device 20 includes: an environment type determination module 21, an audio transmission level determination module 22, and an audio transmission module 23.

[0091] Wherein, the environment type determination module 21 is used to determine the environment type based on multiple ambient audio signals in the environment; An audio delivery level determination module 22, configured to determine, based on the environment type, an audio delivery level corresponding to each ambient audio signal in the environment; The audio transmission module 23 is configured to transmit a target audio signal based on the audio transmission level, where the target audio signal is a signal among the multiple ambient audio signals.

[0092] In one embodiment of the present disclosure, the audio delivery device 20 further includes: an audio processing module; The audio processing module is specifically used for: Performing short-time Fourier transform on the mixed audio signal in the environment to obtain a plurality of time-frequency units, wherein any time-frequency unit includes an amplitude spectrum corresponding to the mixed audio signal; The magnitude spectrum is fed into a convolutional mental network to determine the cluster assignment probability of each time-frequency unit; Based on the clustering assignment probability of each time-frequency unit, multiple time-frequency units are classified to obtain multiple environmental audio signals.

[0093] In an embodiment of the present disclosure, any time-frequency unit further includes a phase spectrum corresponding to the mixed audio signal; and the audio processing module is further configured to: Classifying multiple time-frequency units based on the clustering assignment probability of each time-frequency unit to obtain multiple initial environmental audio signals; Based on the phase spectrum corresponding to the mixed audio signal, an inverse short-time Fourier transform is performed on the multiple initial ambient audio signals to obtain multiple ambient audio signals.

[0094] In one embodiment of the present disclosure, the environment type determination module 21 is specifically used to: The environment type is determined based on similarities between the plurality of environmental audio signals and the target environment type.

[0095] In one embodiment of the present disclosure, the audio delivery level determination module 22 is specifically configured to: Extract features from each ambient audio signal to obtain key feature information; The audio transmission level corresponding to each ambient audio signal in the environment is determined based on the correlation between the key feature information and the type of the environment.

[0096] In one embodiment of the present disclosure, the audio delivery device 20 further includes: a denoising module; The denoising module is specifically used to: decompose each ambient audio signal into initial wavelet coefficients of different frequencies; The wavelet coefficients are processed based on a preset threshold value to obtain target wavelet coefficients; Perform inverse wavelet transform on the target wavelet coefficients to obtain each ambient audio signal after denoising.

[0097] In one embodiment of the present disclosure, the denoising module is further configured to: The preset threshold is determined based on the audio delivery level corresponding to each environmental audio signal in the environment.

[0098] See also Figure 3 , Figure 3 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. Figure 3 The electronic device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303 and one or more memories 304. The processors 301, input devices 302, output devices 303 and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of the modules in the above-mentioned device embodiments, such as Figure 2 The functions of modules 21 to 23 are shown.

[0099] It should be understood that in the embodiment of the present disclosure, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0100] The input device 302 may include a touch panel, a fingerprint collection sensor (for collecting the user's fingerprint information and fingerprint direction information), a microphone, etc., and the output device 303 may include a display (LCD, etc.), a speaker, etc.

[0101] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A portion of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0102] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiments of the present disclosure can execute the implementation methods described in the first and second embodiments of the audio transmission method provided in the embodiments of the present disclosure, and can also execute the implementation methods of the electronic device described in the embodiments of the present disclosure, which will not be repeated here.

[0103] In another embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by the processor, all or part of the processes in the above-mentioned embodiment method are implemented, and the computer program can also be completed by instructing the relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0104] The computer-readable storage medium may be an internal storage unit of the electronic device of any of the aforementioned embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0105] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.

[0106] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0107] In the several embodiments provided in the present application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or it can be an electrical, mechanical or other form of connection.

[0108] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present disclosure.

[0109] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0110] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present disclosure, and these modifications or replacements should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.

Claims

1. An audio transmission method, characterized in that: include: determining a type of environment based on a plurality of ambient audio signals in the environment; Determine, based on the environment type, an audio transmission level corresponding to each ambient audio signal in the environment; A target audio signal is delivered based on the audio delivery level, the target audio signal being a signal among a plurality of ambient audio signals.

2. The audio transmission method according to claim 1, characterized in that: Before determining the environment type based on a plurality of ambient audio signals in the environment, the method further includes: Performing short-time Fourier transform on the mixed audio signal in the environment to obtain a plurality of time-frequency units, wherein any time-frequency unit includes an amplitude spectrum corresponding to the mixed audio signal; Inputting the amplitude spectrum into a convolutional neural network to determine the clustering probability of each time-frequency unit; The multiple time-frequency units are classified based on the clustering allocation probability of each time-frequency unit to obtain multiple environmental audio signals.

3. The audio transmission method according to claim 2, characterized in that: Any time-frequency unit also includes a phase spectrum corresponding to the mixed audio signal; The method of classifying the multiple time-frequency units based on the clustering allocation probability of each time-frequency unit to obtain multiple environmental audio signals includes: Classifying the multiple time-frequency units based on the clustering allocation probability of each time-frequency unit to obtain multiple initial environmental audio signals; The multiple initial ambient audio signals are subjected to inverse short-time Fourier transform based on the phase spectrum corresponding to the mixed audio signal to obtain multiple ambient audio signals.

4. The audio transmission method according to claim 1, characterized in that: The determining of the environment type based on a plurality of ambient audio signals in the environment comprises: The environment type is determined based on similarities between the plurality of ambient audio signals and a target environment type.

5. The audio transmission method according to claim 1, characterized in that: The step of determining the audio transmission level corresponding to each ambient audio signal in the environment based on the environment type includes: Extract features from each ambient audio signal to obtain key feature information; The audio transmission level corresponding to each ambient audio signal in the environment is determined based on the correlation between the key feature information and the type of the environment.

6. The audio transmission method according to claim 1, characterized in that: Also includes: Decompose each ambient audio signal into initial wavelet coefficients of different frequencies; Processing the wavelet coefficients based on a preset threshold to obtain target wavelet coefficients; Performing an inverse wavelet transform on the target wavelet coefficients to obtain each ambient audio signal after denoising.

7. The audio transmission method according to claim 6, characterized in that: Also includes: The preset threshold is determined based on the audio transmission level corresponding to each environmental audio signal in the environment.

8. An audio transmission device, characterized in that: include: An environment type determination module, used to determine the environment type based on multiple ambient audio signals in the environment; An audio delivery level determination module, used to determine the audio delivery level corresponding to each ambient audio signal in the environment based on the environment type; The audio delivery module is configured to deliver a target audio signal based on the audio delivery level, wherein the target audio signal is a signal among a plurality of ambient audio signals.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.