An intelligent audio processing method and related apparatus

By establishing an authenticated connection between the microphone and the TV and controlling sound effects and scenes, and by using a conversion network for time-frequency domain feature extraction and a linear transfer function for original sound suppression, the problem of insufficient separation accuracy between original audio and accompaniment audio in existing technologies is solved, thus achieving a highly reliable karaoke experience.

CN120812343BActive Publication Date: 2025-12-12GUANGZHOU CHANGJIA ELECTRONICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511303937.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-12-12
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing technologies that use Fast Fourier Transform and vocal separation models to separate original audio from accompaniment audio will lose phase information, resulting in insufficient separation accuracy and failing to meet users' karaoke experience needs.

Method used

The system receives and parses sound effect scene control commands through a microphone-to-TV authentication connection, performs original sound suppression analysis, extracts time-frequency domain features using a conversion network, and combines a linear transfer function to suppress original sound and generate the target mixed audio.

Benefits of technology

It achieves highly reliable original audio suppression, improves the separation accuracy of accompaniment audio and original audio, and makes the generated target mixed audio more in line with the user's actual needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812343B_ABST
    Figure CN120812343B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent audio processing method and related device, and relates to the technical field of audio processing.The method comprises the following steps: a microphone is authenticated with a receiver in a television set, and if the authentication is successful, the microphone and the television set form a connection state; a sound effect scene control instruction sent by the microphone is analyzed to obtain sound effect scene analysis information to enter a target sound effect scene mode; the current playing audio of the television set is analyzed based on an original sound suppression instruction; the current playing audio is subjected to time-frequency domain feature extraction based on a conversion network; the current playing audio is separated into accompaniment audio and original audio based on time-frequency domain feature information; the original audio is subjected to original sound suppression based on a target original sound suppression degree by using a linear transfer function, and target mixed audio is generated based on the suppressed original audio and the accompaniment audio.The application can realize high-reliability original sound suppression, and the generated target mixed audio is more in line with the actual needs of users.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio processing, in particular to an intelligent audio processing method and related device. BACKGROUND

[0002] With the improvement of people's living standards, people also pay more attention to the richness of the spiritual world, so various entertainment methods appear in people's daily life. Among various entertainment methods, K singing occupies a large proportion. Through K singing, not only can work pressure be released, but also the spiritual world can be enriched. And with the development of television technology, people can use a television to sing at home. When using a television to sing, people usually choose to suppress the original sound of the played audio to meet the singing experience. The separation of original sound and accompaniment audio is an important step of audio original sound suppression. Currently, fast Fourier transform and voice separation model are usually used to separate the original sound and accompaniment audio, but this way will lose phase information in spectrum processing, so that the separation accuracy of original sound and accompaniment is insufficient, resulting in that the suppression of original sound cannot achieve the expected effect, and cannot bring enough satisfactory singing experience to users. SUMMARY

[0003] The purpose of the present application is to overcome the shortcomings of the prior art, and the present application provides an intelligent audio processing method and related device, which can realize high-reliability original sound suppression, so that the generated target mixed audio is more in line with the actual needs of users.

[0004] In order to solve the above technical problems, the present application provides an intelligent audio processing method, which comprises:

[0005] The microphone is authenticated with the receiver inserted into the television based on a team matching mode, and if the authentication is successful, the microphone and the television form a connection state;

[0006] The television receives the sound effect scene control instruction issued by the microphone, analyzes the sound effect scene control instruction, obtains sound effect scene analysis information, and enters a target sound effect scene mode based on the sound effect scene analysis information;

[0007] After the television enters the target sound effect scene mode, the original sound suppression instruction issued by the microphone is received, the current played audio of the television is analyzed based on the original sound suppression instruction, and the target original sound suppression degree is obtained;

[0008] The time-frequency domain feature information is obtained by performing time-frequency domain feature extraction on the current played audio based on the conversion network;

[0009] The current played audio is separated into accompaniment audio and original sound based on the time-frequency domain feature information;

[0010] The original sound audio is suppressed based on the target original sound suppression degree using a linear transfer function to obtain suppressed original sound audio, and the target mixed audio is generated based on the suppressed original sound audio and the accompaniment audio.

[0011] Optionally, the microphone is authenticated with the receiver inserted into the television based on a team matching mode, and if the authentication is successful, the microphone and the television form a connection state, including:

[0012] The television and the receiver are connected based on a universal serial bus (USB) port;

[0013] The microphone searches based on a team matching mode, and when the receiver is searched, the microphone sends a pairing request instruction to the receiver, the receiver generates an authentication key based on the pairing request instruction, and sends the authentication key to the microphone;

[0014] The microphone authenticates the authentication key based on ObjectOutputStream sequences and ObjectInputStream reverse sequences, and if the authentication is successful, the microphone and the receiver form an interconnected pairing state;

[0015] The microphone sends a connection data packet to the receiver, creates a target link of the television and the microphone based on the connection data packet, and forms a connection state of the microphone and the television in the interconnected pairing state based on the target link.

[0016] Optionally, the sound effect scene control instruction is parsed to obtain sound effect scene parsing information, including:

[0017] The information format of the sound effect scene information is determined based on the sound effect scene control instruction;

[0018] The sound effect scene control instruction is parsed based on the information format of the sound effect scene information combined with the application layer to obtain the sound effect scene parsing information.

[0019] Optionally, the current playing audio of the television is analyzed for original sound suppression degree based on the original sound suppression instruction to obtain a target original sound suppression degree, including:

[0020] The corresponding application flag information is determined based on the original sound suppression instruction, and the corresponding application program module is matched based on the application flag information;

[0021] The current playing audio of the television is analyzed for original sound suppression degree based on the operation data in the original sound suppression instruction combined with the application program module to obtain the target original sound suppression degree.

[0022] Optionally, the time-frequency domain feature information is obtained by performing time-frequency domain feature extraction on the current playing audio based on the conversion network, including:

[0023] extract a feature sequence of the currently played audio, and perform time-domain autocorrelation processing on the feature sequence based on a time sequence correlation layer to obtain an autocorrelation feature vector sequence;

[0024] perform frequency-domain convolution processing on the autocorrelation feature vector sequence along a frequency domain direction based on a plurality of frequency-domain convolution kernels of different scales to obtain a frequency-domain convolution vector;

[0025] perform frequency-domain feature extraction based on a frequency-domain conversion network combined with a self-attention mechanism to obtain a frequency-domain feature map using the frequency-domain convolution vector;

[0026] perform attention fusion processing on the time-domain feature map and the frequency-domain feature map based on a time conversion network to obtain time-frequency domain feature information.

[0027] Optionally, the separating the currently played audio into accompaniment audio and original audio based on the time-frequency domain feature information comprises:

[0028] determining an accompaniment binary mask based on a mask analysis model combined with the time-frequency domain feature information;

[0029] determining initial accompaniment data and initial original data in the currently played audio based on the accompaniment binary mask;

[0030] performing background sound residue analysis on the initial original data based on an audio event classifier combined with a scoring mechanism to obtain background sound residue data;

[0031] determining the accompaniment audio and the original audio based on the background sound residue data, the initial accompaniment data and the initial original data using a codec network.

[0032] Optionally, the performing original sound suppression on the original audio based on the target original sound suppression degree using a linear transfer function to obtain suppressed original audio, and generating target mixed audio based on the suppressed original audio and the accompaniment audio comprises:

[0033] determining a first autocorrelation spectrum of the original audio and a second autocorrelation spectrum of the accompaniment audio, and determining a cross-correlation spectrum based on the target original sound suppression degree;

[0034] determining a linear transfer function based on the first autocorrelation spectrum, the second autocorrelation spectrum and the cross-correlation spectrum, and determining an original sound suppression frequency-domain signal based on the linear transfer function and the target original sound suppression degree;

[0035] performing original sound suppression on the original audio based on the original sound suppression frequency-domain signal to obtain suppressed original audio;

[0036] performing alignment and mixing processing based on the suppressed original audio and the accompaniment audio to obtain target mixed audio.

[0037] In addition, the application further provides an intelligent audio processing device, which comprises:

[0038] The device connection module is used for the microphone to authenticate with the receiver inserted into the TV based on the team matching mode, and if the authentication is successful, the microphone and the TV form a connection state.

[0039] The sound effect scene analysis module is used for the TV to receive the sound effect scene control instruction sent by the microphone, analyze the sound effect scene control instruction, obtain sound effect scene analysis information, and make the TV enter a target sound effect scene mode based on the sound effect scene analysis information.

[0040] The original sound suppression analysis module is used for receiving the original sound suppression instruction sent by the microphone after the TV enters the target sound effect scene mode, analyzing the current playing audio of the TV based on the original sound suppression instruction, and obtaining a target original sound suppression degree.

[0041] The time-frequency domain feature module is used for extracting time-frequency domain features of the current playing audio based on a conversion network, and obtaining time-frequency domain feature information.

[0042] The audio separation module is used for separating the current playing audio into accompaniment audio and original audio based on the time-frequency domain feature information.

[0043] The mixed audio generation module is used for suppressing the original audio based on the target original sound suppression degree by using a linear transfer function, obtaining suppressed original audio, and generating target mixed audio based on the suppressed original audio and the accompaniment audio.

[0044] In addition, the application further provides an electronic device, which comprises a processor and a memory, the memory is used for storing instructions, and the processor is used for calling the instructions in the memory, so that the electronic device executes the intelligent audio processing method.

[0045] In addition, the application further provides a computer readable storage medium, which stores computer instructions, when the computer instructions run on an electronic device, so that the electronic device executes the intelligent audio processing method.

[0046] In the embodiment of the present application, the microphone is authenticated with the receiver inserted into the TV set based on the team matching mode, so that the microphone and the TV set form a more stable connection state. The TV set receives the sound effect scene control instruction sent by the microphone, analyzes the sound effect scene control instruction, obtains sound effect scene analysis information, and enters a target sound effect scene mode based on the sound effect scene analysis information. In the sound effect scene mode, better sound effect scene can be provided for the user. The time-frequency domain feature information is obtained by performing time-frequency domain feature extraction on the currently played audio based on the conversion network, and the currently played audio is separated into accompaniment audio and original audio based on the time-frequency domain feature information, which effectively improves the separation precision of the accompaniment audio and the original audio. The original audio is suppressed based on the target original sound suppression degree and the linear transfer function, the suppressed original audio is obtained, and the target mixed audio is generated based on the suppressed original audio and the accompaniment audio, which can realize high-reliability original audio suppression and make the generated target mixed audio more meet the actual needs of the user. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0048] Figure 1 is a flowchart of the intelligent audio processing method in the embodiment of the present application;

[0049] Figure 2 is a flowchart of the intelligent audio processing method in another embodiment of the present application;

[0050] Figure 3 is a structural composition diagram of the intelligent audio processing device in the embodiment of the present application;

[0051] Figure 4 is a structural composition diagram of the electronic device in the embodiment of the present application. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0053] Embodiment one

[0054] Please refer toFigure 1 , Figure 1 is a flowchart of an intelligent audio processing method in the embodiment of the application, and the method comprises:

[0055] S11: the microphone is authenticated with the receiver inserted into the TV based on the team matching mode, and if the authentication is successful, the microphone and the TV form a connection state;

[0056] In the specific implementation process of the application, the microphone is authenticated with the receiver inserted into the TV based on the team matching mode, and if the authentication is successful, the microphone and the TV form a connection state, which comprises: the TV and the receiver are connected based on a universal serial bus (USB) port; the microphone searches based on the team matching mode, and when the receiver is searched, the microphone sends a pairing request instruction to the receiver; the receiver generates an authentication key based on the pairing request instruction and sends the authentication key to the microphone; the microphone authenticates the authentication key based on ObjectOutputStream sequence and ObjectInputStream reverse sequence, and if the authentication is successful, the microphone and the receiver form an interconnected pairing state; the microphone sends a connection data packet to the receiver, creates a target link of the TV and the microphone based on the connection data packet, and forms a connection state of the microphone and the TV in the interconnected pairing state based on the target link.

[0057] Specifically, the television is connected with the receiver based on a universal serial bus (USB) port, that is, the receiver is inserted into the USB port of the television, and the television is connected with the receiver, and it should be noted that the television is in a networked state. The microphone is in a team matching mode by pressing the microphone power key for three seconds, and the team matching mode is an active search state. In this state, the microphone actively searches for the corresponding receiver. The microphone searches based on the team matching mode, and when the receiver is searched, the microphone sends a pairing request instruction to the receiver. The receiver has a pairing program built-in, which is used for pairing with the microphone. The receiver generates an authentication key based on the pairing request instruction, obtains an initial key flag and a pairing mode flag in the pairing request instruction, generates an initial key and a random number according to the initial key flag and the pairing mode flag of the pairing request instruction, and performs exclusive or operation according to the device information and the random number of the microphone and the receiver to obtain first data. The device information includes device address type and device address. According to the first data, the communication parameters in the pairing request instruction are operated to obtain second data. According to the second data, the initial key is encrypted in the encryption chip by using a check tag to obtain an authentication key. The encryption operation includes exclusive or operation and symmetric encryption operation. The check tag is a hash sequence check value obtained by the receiver according to the device information and a preset key by hash operation, and the authentication key is sent to the microphone. The microphone authenticates the authentication key based on ObjectOutputStream sequence and ObjectInputStream reverse sequence. ObjectOutputStream sequence is an object stream function providing serialization function, which can convert objects into byte streams and realize persistent storage of objects. The authentication key is converted into a byte stream by ObjectOutputStream sequence. ObjectInputStream reverse sequence is an object stream function providing reverse serialization function. The obtained authentication key information can carry more resource quantity and resource domain through reverse sequence processing. The authentication key is converted into an original object by ObjectInputStream reverse sequence. The matching between the authentication key and the verification key is performed according to the original object. If the authentication is successful, that is, the authentication key and the verification key are matched successfully, the microphone and the receiver form an interconnected pairing state.The microphone sends a connection data packet to the receiver, the connection data packet includes information required for creating a link, such as a device address and a logical transmission address, and a target link of the television and the microphone is created based on the connection data packet, the television is in a connection state with the receiver, the connection data packet is transmitted to a controller of the television, and the television and the microphone in an interconnection pairing state establish a target link, the target link is a synchronous communication bridge between the microphone and the television, and the microphone and the television in the interconnection pairing state are connected based on the target link, that is, the microphone and the television in the interconnection pairing state are connected to the target link, the microphone and the television are connected, and the connection between the microphone and the television is more stable, so that the occurrence of a larger error in signal transmission is effectively avoided.

[0058] S12: The television receives the sound effect scene control instruction sent by the microphone, analyzes the sound effect scene control instruction, obtains sound effect scene analysis information, and enters a target sound effect scene mode based on the sound effect scene analysis information.

[0059] In the specific implementation process of the application, the analysis of the sound effect scene control instruction and the obtaining of the sound effect scene analysis information include: determining the information format of the sound effect scene information based on the sound effect scene control instruction; and analyzing the sound effect scene control instruction based on the application layer and the information format of the sound effect scene information to obtain the sound effect scene analysis information.

[0060] Specifically, the sound effect key in the microphone is double-clicked to generate a sound effect scene control instruction, the sound effect scene control instruction is transmitted to the television, the television receives the sound effect scene control instruction, determines the information format of the sound effect scene information based on the sound effect scene control instruction, and the sound effect scene control instruction contains the information format of the sound effect scene option, that is, the information format of the sound effect scene information. The sound effect scene control instruction is analyzed based on the application layer and the information format of the sound effect scene information, the sound effect scene control instruction is analyzed according to the information format in the application layer, the sound effect scene mode selected by the user is obtained, that is, the sound effect scene analysis information is obtained, and the sound effect scene mode is, for example, a music hall, a KTV, a concert, etc. The television enters a target sound effect scene mode based on the sound effect scene analysis information, that is, the television enters the sound effect scene mode selected by the user.

[0061] S13: After the television enters the target sound effect scene mode, a raw sound suppression instruction sent by the microphone is received, a raw sound suppression degree analysis of the current playing audio of the television is performed based on the raw sound suppression instruction, and a target raw sound suppression degree is obtained.

[0062] In the implementation of the present application, the original sound suppression degree of the current playing audio of the television is analyzed based on the original sound suppression instruction to obtain a target original sound suppression degree, comprising: determining corresponding application mark information based on the original sound suppression instruction, and matching the corresponding application program module based on the application mark information; analyzing the original sound suppression degree of the current playing audio of the television based on the application program module combined with the operation data in the original sound suppression instruction to obtain the target original sound suppression degree.

[0063] Specifically, after the television enters the target sound effect scene mode, the audio is played in the sound effect scene mode, the user short-presses the sound effect key in the microphone to generate an original sound suppression instruction, and sends the original sound suppression instruction to the television. The television receives the original sound suppression instruction issued by the microphone, determines the corresponding application mark information based on the original sound suppression instruction, different application mark information corresponds to different application program modules, determines the mark information of the original sound suppression application program according to the original sound suppression instruction, and matches the corresponding application program module based on the application mark information, that is, matches the original sound suppression application program module. The current playing audio of the television is analyzed based on the application program module combined with the operation data in the original sound suppression instruction. The operation data in the original sound suppression instruction includes the pressing time and number of the sound effect key. The original sound suppression application program module determines the suppression degree of the original sound of the current playing audio according to the operation data, that is, obtains the target original sound suppression degree. The original sound suppression degree can be divided into four levels, the first level has the smallest suppression degree, and the fourth level has the largest suppression degree. The original sound of the playing audio is eliminated in this level.

[0064] S14: performing time-frequency domain feature extraction on the current playing audio based on the conversion network to obtain time-frequency domain feature information;

[0065] In the implementation of the present application, the time-frequency domain feature information is obtained by performing time-frequency domain feature extraction on the current playing audio based on the conversion network, comprising: extracting a feature sequence of the current playing audio, and performing time domain autocorrelation processing on the feature sequence based on a time sequence correlation layer to obtain an autocorrelation feature vector sequence; performing frequency domain convolution processing on the autocorrelation feature vector sequence along the frequency domain direction based on a plurality of frequency domain convolution kernels of different scales to obtain a frequency domain convolution vector; performing frequency domain feature extraction based on the frequency domain conversion network combined with the self-attention mechanism using the frequency domain convolution vector to obtain a frequency domain feature map; and extracting a time feature map based on the time conversion network using the feature sequence, performing attention fusion processing on the time feature map and the frequency domain feature map to obtain the time-frequency domain feature information.

[0066] Specifically, a feature sequence of the currently played audio is extracted, the feature sequence includes a frequency domain amplitude spectrum of the currently played audio and a plurality of frequency domain vectors arranged in time sequence, and the feature sequence of the currently played audio can be extracted through a preset convolutional neural network. The feature sequence is processed by a time series correlation layer in a time domain autocorrelation manner. The time domain autocorrelation processing is an operation process for measuring the correlation of a selected frequency domain vector in the time domain direction with other frequency domain vectors in the feature sequence. One frequency domain vector is selected from the feature sequence, the correlation scores between the frequency domain vector and the remaining frequency domain vectors are calculated, the correlation scores are used as the correlation weights of the frequency domain vector, and the above processing is repeated until the correlation weights of all frequency domain vectors are calculated. After that, all frequency domain vectors and their corresponding correlation weights are arranged in time sequence to form a sequence, i.e., an autocorrelation feature vector sequence is obtained. The autocorrelation feature vector sequence is processed in a frequency domain convolution manner by a plurality of frequency domain convolution kernels of different scales along the frequency domain direction. The frequency domain direction can include a direction along the sampling frequency from small to large or a direction along the sampling frequency from large to small. Different scales of the convolution kernel can extract audio features at different levels to obtain a frequency domain convolution vector. The frequency domain convolution vector is used to extract frequency domain features based on a frequency domain conversion network combined with a self-attention mechanism. The frequency domain convolution vector is input into the frequency domain conversion network, and the weight of the frequency domain feature at each position on the output audio feature is determined by the self-attention mechanism as an attention result. The attention result is mapped to a separable space in the frequency domain. The self-attention mechanism can realize encoding of global input-output dependency relationship, effectively capture long-time dependency of the input audio, and obtain a frequency domain feature map. The feature sequence is used to extract a time feature map based on a time conversion network, i.e., the transposed amplitude spectrum corresponding to the frequency domain amplitude spectrum in the feature sequence is input into the time conversion network. The time conversion network is used to extract features of different time domain dimensions, and the time conversion network adopts a multi-head attention mechanism. The time feature map and the frequency domain feature map are processed by attention fusion, the time feature map and the frequency domain feature map are weighted according to the attention map corresponding to the time feature map and the frequency domain feature map, a weighted feature map corresponding to the time feature map and the frequency domain feature map is obtained, and the weighted feature maps corresponding to the time feature map and the frequency domain feature map are added to obtain a fusion feature map, i.e., the time-frequency domain feature information is obtained.

[0067] S15: separating the currently played audio into accompaniment audio and original audio based on the time-frequency domain feature information;

[0068] In the implementation of the present application, the current playing audio is separated into accompaniment audio and original audio based on the time-frequency domain feature information, which includes: determining accompaniment binary mask based on mask analysis model combined with time-frequency domain feature information; determining initial accompaniment data and initial original data in the current playing audio based on the accompaniment binary mask; performing background sound residue analysis on the initial original data based on audio event classifier combined with scoring mechanism to obtain background sound residue data; and determining accompaniment audio and original audio based on the background sound residue data, initial accompaniment data and initial original data using coding and decoding network.

[0069] Specifically, the accompaniment binary mask is determined based on the mask analysis model combined with the time-frequency domain feature information, and the time-frequency domain feature information is input into the mask analysis model. The mask analysis model is a converged model obtained by inputting a sample data set into a deep neural network for training, and the accompaniment binary mask is used for sound source separation. The initial accompaniment data and the initial original data in the current playing audio are determined based on the accompaniment binary mask. The original sound binary mask in the current playing audio is obtained, and the initial original spectrum and the initial accompaniment spectrum in the current playing audio are determined based on the original sound binary mask combined with the spectrum separation model. The initial original spectrum is multiplied by the accompaniment binary mask to obtain the accompaniment sub-spectrum. The initial original spectrum is subtracted from the accompaniment sub-spectrum to obtain the target original spectrum. The accompaniment sub-spectrum is added to the initial accompaniment spectrum to obtain the target accompaniment spectrum. The target original spectrum and the target accompaniment spectrum are subjected to inverse short-time Fourier transform to obtain the initial accompaniment data and the initial original data. The initial original data often contains some background sound residues, so it is necessary to analyze the background sound residues of the initial original data. The background sound residue data is obtained by inputting the initial original data into the audio event classifier, performing audio event scoring on each frame of signal in the initial original data using the audio event classifier, obtaining the audio event scoring score of each frame of signal, and regarding the frame of signal as the background sound residue signal in the initial original data if the audio event scoring score is greater than or equal to a preset threshold. The accompaniment audio and the original audio are determined based on the background sound residue data, the initial accompaniment data and the initial original data using the coding and decoding network. The time data of the background sound residue data is extracted from the initial original data, the background sound residue data and the initial accompaniment data are combined according to the time data to obtain the target accompaniment data, the background sound residue data in the initial original data is extracted to obtain the target original data, and the target original data and the target accompaniment data are input into the coding and decoding network to obtain the original frequency domain amplitude spectrum and the accompaniment frequency domain amplitude spectrum, i.e. the accompaniment audio and the original audio.

[0070] S16: performing original sound suppression on the original audio using a linear transfer function based on the target original sound suppression degree to obtain suppressed original audio, and generating target mixed audio based on the suppressed original audio and the accompaniment audio.

[0071] In the implementation of the present application, the original sound is suppressed based on the target original sound suppression degree and the linear transfer function to obtain the original sound suppression audio, and the target mixed audio is generated based on the original sound suppression audio and the accompaniment audio, which comprises: determining the first self-power spectrum of the original audio and the second self-power spectrum of the accompaniment audio, and determining the cross-power spectrum based on the target original sound suppression degree; determining the linear transfer function based on the first self-power spectrum, the second self-power spectrum and the cross-power spectrum, and determining the original sound suppression frequency domain signal based on the linear transfer function and the target original sound suppression degree; suppressing the original sound of the original audio based on the original sound suppression frequency domain signal to obtain the original sound suppression audio; and performing alignment and mixing processing based on the original sound suppression audio and the accompaniment audio to obtain the target mixed audio.

[0072] Specifically, the first self-power spectrum of the original audio and the second self-power spectrum of the accompaniment audio are determined, the frequency domain signal of the original audio is determined, the conjugate value of the corresponding complex number of the frequency domain signal is determined, the first self-power spectrum is determined according to the update rate of the power spectrum, the frequency domain signal and the conjugate value, the determination steps of the second power spectrum are the same as those of the first power spectrum, the cross-power spectrum is determined based on the target original sound suppression degree, the original sound suppression signal of the original audio is determined according to the target original sound suppression degree, the frequency domain signal of the original sound suppression signal is determined, and the cross-power spectrum is determined according to the frequency domain signal of the original sound suppression signal, the update rate of the power spectrum and the frequency domain signal of the accompaniment spectrum. The linear transfer function is determined based on the first self-power spectrum, the second self-power spectrum and the cross-power spectrum, the linear transfer function describes the relationship between the input signal and the output signal of the linear system, and the expression of the linear transfer function is:

[0073]

[0074] wherein F is the linear transfer function, H is the cross-power spectrum, G1 is the first power spectrum, G2 is the second power spectrum, is a parameter for controlling the suppression degree of the original sound signal, and the original sound suppression frequency domain signal is determined based on the linear transfer function and the target original sound suppression degree. The frequency domain signal of the original sound suppression audio is determined according to the linear transfer function and the target original sound suppression degree, i.e. the original sound suppression frequency domain signal is obtained. The original sound of the original audio is suppressed based on the original sound suppression frequency domain signal, the original sound of each frame of data of the original audio is suppressed according to the original sound suppression frequency domain signal, and the original sound suppression audio is obtained. The alignment and mixing processing is performed based on the original sound suppression audio and the accompaniment audio, the original sound suppression audio and the accompaniment audio are aligned according to the time axis of the original audio and the accompaniment audio, the mixed audio is obtained by mixing the aligned original sound suppression audio and the accompaniment audio, i.e. merging the aligned original sound suppression audio and the accompaniment audio, and the target mixed audio is output by the television set. At this time, the original sound of the target mixed audio output has been suppressed according to the set suppression level, which can meet the K song needs of the user.​

[0075] In the embodiment of the present application, the microphone is authenticated with the receiver inserted into the TV set based on the team matching mode, so that the microphone and the TV set form a more stable connection state. The TV set receives the sound effect scene control instruction issued by the microphone, analyzes the sound effect scene control instruction, obtains sound effect scene analysis information, and enters a target sound effect scene mode based on the sound effect scene analysis information. In the sound effect scene mode, better sound effect scenes can be provided for users. The time-frequency domain feature information is obtained by performing time-frequency domain feature extraction on the currently played audio based on the conversion network. The currently played audio is separated into accompaniment audio and original audio based on the time-frequency domain feature information, which effectively improves the separation accuracy of the accompaniment audio and the original audio. The linear transfer function is used to suppress the original audio based on the target original sound suppression degree, to obtain suppressed original audio, and the target mixed audio is generated based on the suppressed original audio and the accompaniment audio, which can realize high-reliability original audio suppression and make the generated target mixed audio more in line with the actual needs of users.

[0076] Embodiment two

[0077] Please refer to Figure 2 , Figure 2 is a flowchart of an intelligent audio processing method in another embodiment of the present application. The method comprises:

[0078] S201: The microphone is authenticated with the receiver inserted into the TV set based on the team matching mode. If the authentication is successful, the microphone and the TV set form a connection state.

[0079] S202: The TV set receives the sound effect scene control instruction issued by the microphone, analyzes the sound effect scene control instruction, obtains sound effect scene analysis information, and enters a target sound effect scene mode based on the sound effect scene analysis information.

[0080] S203: After the TV set enters the target sound effect scene mode, the original sound suppression instruction issued by the microphone is received, and the current played audio of the TV set is analyzed based on the original sound suppression degree to obtain a target original sound suppression degree.

[0081] S204: Time-frequency domain feature information is obtained by performing time-frequency domain feature extraction on the currently played audio based on the conversion network.

[0082] S205: The currently played audio is separated into accompaniment audio and original audio based on the time-frequency domain feature information.

[0083] S206: The first self-power spectrum of the original audio and the second self-power spectrum of the accompaniment audio are determined, and the mutual power spectrum is determined based on the target original sound suppression degree.

[0084] S207: determining a linear transfer function based on the first self-power spectrum, the second self-power spectrum and the cross-power spectrum, and determining a primary sound suppression frequency domain signal based on the linear transfer function and a target primary sound suppression degree;

[0085] S208: performing primary sound suppression on the primary audio based on the primary sound suppression frequency domain signal to obtain suppressed primary audio;

[0086] S209: performing alignment and mixing processing on the suppressed primary audio and the accompaniment audio to obtain target mixed audio.

[0087] In the embodiment of the present application, the microphone is authenticated with the receiver inserted into the TV based on the team matching mode, so that the microphone and the TV form a more stable connection state. The TV receives the sound effect scene control instruction sent by the microphone, analyzes the sound effect scene control instruction, obtains sound effect scene analysis information, and enters a target sound effect scene mode based on the sound effect scene analysis information. In the sound effect scene mode, better sound effect scene can be provided for the user. The time-frequency domain features of the currently played audio are extracted based on the conversion network to obtain time-frequency domain feature information, and the currently played audio is separated into accompaniment audio and primary audio based on the time-frequency domain feature information, which effectively improves the separation accuracy of the accompaniment audio and the primary audio. The primary audio is suppressed based on the target primary sound suppression degree using the linear transfer function to obtain suppressed primary audio, and the target mixed audio is generated based on the suppressed primary audio and the accompaniment audio, which can realize high-reliability primary audio suppression and make the generated target mixed audio more meet the actual needs of the user.

[0088] Embodiment three

[0089] Please refer to Figure 3 , Figure 3 is a structural composition diagram of the audio processing device based on intelligence in the embodiment of the present application, the device comprises:

[0090] The device connection module 31 is used for the microphone to authenticate with the receiver inserted into the TV based on the team matching mode, and if the authentication is successful, the microphone and the TV form a connection state.

[0091] The sound effect scene analysis module 32 is used for the TV to receive the sound effect scene control instruction sent by the microphone, analyze the sound effect scene control instruction, obtain sound effect scene analysis information, and enter a target sound effect scene mode based on the sound effect scene analysis information.

[0092] The primary sound suppression analysis module 33 is used for receiving the primary sound suppression instruction sent by the microphone after the TV enters the target sound effect scene mode, analyzing the primary sound suppression degree of the currently played audio of the TV based on the primary sound suppression instruction, and obtaining a target primary sound suppression degree.

[0093] The time-frequency domain feature module 34 is configured to perform time-frequency domain feature extraction on the currently played audio based on the conversion network to obtain time-frequency domain feature information.

[0094] The audio separation module 35 is configured to separate the currently played audio into accompaniment audio and original audio based on the time-frequency domain feature information.

[0095] The mixed audio generation module 36 is configured to perform original sound suppression on the original audio based on the target original sound suppression degree using a linear transfer function to obtain suppressed original audio, and generate target mixed audio based on the suppressed original audio and the accompaniment audio.

[0096] In the implementation of the application, the implementation of the device can refer to the implementation of the method described above, which will not be repeated here.

[0097] In the embodiment of the application, the microphone is authenticated with the receiver inserted into the TV based on the team matching mode, so that the microphone and the TV form a more stable connection state. The TV receives the sound effect scene control instruction sent by the microphone and analyzes the sound effect scene control instruction to obtain sound effect scene analysis information. The TV enters a target sound effect scene mode based on the sound effect scene analysis information. In the sound effect scene mode, the user can be provided with a better sound effect scene. The time-frequency domain feature information is obtained by performing time-frequency domain feature extraction on the currently played audio based on the conversion network. The currently played audio is separated into accompaniment audio and original audio based on the time-frequency domain feature information, which effectively improves the separation accuracy of the accompaniment audio and the original audio. The original sound suppression is performed on the original audio based on the target original sound suppression degree using a linear transfer function to obtain suppressed original audio, and the target mixed audio is generated based on the suppressed original audio and the accompaniment audio, which can realize high-reliability original sound suppression and make the generated target mixed audio more in line with the actual needs of the user.

[0098] The computer readable storage medium provided by the embodiment of the present application stores a computer program, and the program is executed by a processor to realize the intelligent audio processing method of any one of the above embodiments. The computer readable storage medium includes but is not limited to any type of disk (including a floppy disk, a hard disk, an optical disk, a CD-ROM, and a magneto-optical disk), a ROM (Read-Only Memory), a RAM (Random Access Memory), an EPROM (Erasable Programmable Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a flash memory, a magnetic card or an optical card. That is, the storage device includes any medium that stores or transmits information in a form capable of being read by a device (for example, a computer, a mobile phone), and can be a read-only memory, a magnetic disk or an optical disk, etc.

[0099] Embodiment Four

[0100] Please refer to Figure 4 , Figure 4 is a structural composition diagram of an electronic device in the embodiment of the present application.

[0101] The embodiment of the present application further provides an electronic device, as shown in Figure 4 , the electronic device includes a memory 41, a processor 43, and a computer program 42 stored in the memory 41 and executable on the processor 43. Those skilled in the art can understand that Figure 4The electronic device shown does not constitute a limitation on all devices, and can include more or fewer components than shown, or combine some components. The memory 41 can be used to store computer programs 42 and various functional modules, and the processor 43 runs the computer programs 42 stored in the memory 41 to perform various functional applications and data processing of the device. The memory can be an internal memory or an external memory, or include both the internal memory and the external memory. The internal memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, or random access memory. The external memory can include a hard disk, a floppy disk, a ZIP disk, a USB disk, a magnetic tape, etc. The processor 43 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, a single-chip processor, or the processor 43 can be any conventional processor, etc. The processor and the memory disclosed in the present application include but are not limited to these types of processors and memories. The processor and the memory disclosed in the present application are only examples and are not limited.

[0102] As an embodiment, the electronic device includes one or more processors 43, a memory 41, and one or more computer programs 42, wherein the one or more computer programs 42 are stored in the memory 41 and configured to be executed by the one or more processors 43, and the one or more computer programs 42 are configured to perform the intelligent audio processing method in any one of the above embodiments. For specific implementation process, please refer to the above embodiments, which will not be repeated here.

[0103] In the embodiment of the present application, the microphone is authenticated with the receiver inserted into the TV set based on the team matching mode, so that the microphone and the TV set form a more stable connection state. The TV set receives the sound effect scene control instruction sent by the microphone, analyzes the sound effect scene control instruction, obtains sound effect scene analysis information, and enters a target sound effect scene mode based on the sound effect scene analysis information. In the sound effect scene mode, better sound effect scenes can be provided for users. The time-frequency domain feature information is obtained by performing time-frequency domain feature extraction on the currently played audio based on the conversion network, the currently played audio is separated into accompaniment audio and original audio based on the time-frequency domain feature information, and the separation accuracy of the accompaniment audio and the original audio is effectively improved. The original audio is suppressed based on the target original sound suppression degree using the linear transfer function to obtain suppressed original audio, and the target mixed audio is generated based on the suppressed original audio and the accompaniment audio, which can realize high-reliability original audio suppression and make the generated target mixed audio more meet the actual needs of users.

[0104] In addition, the above describes in detail the intelligent audio processing method and related device provided by the embodiment of the present application. The principle and implementation mode of the present application are described by using specific examples in this paper. The above description of the embodiments is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. An intelligent audio processing method, characterized in that, The method comprises: The microphone is authenticated with the receiver inserted into the TV based on a team matching mode, and if the authentication is successful, the microphone and the TV form a connection state; The TV receives the sound effect scene control instruction sent by the microphone, analyzes the sound effect scene control instruction, obtains sound effect scene analysis information, and enters a target sound effect scene mode based on the sound effect scene analysis information; After the TV enters the target sound effect scene mode, the original sound suppression instruction sent by the microphone is received, the current playing audio of the TV is analyzed based on the original sound suppression instruction to obtain a target original sound suppression degree; The time-frequency domain feature information of the current playing audio is extracted based on a conversion network; The current playing audio is separated into accompaniment audio and original audio based on the time-frequency domain feature information; The original audio is suppressed based on the target original sound suppression degree using a linear transfer function to obtain suppressed original audio, and the target mixed audio is generated based on the suppressed original audio and the accompaniment audio.

2. The intelligent audio processing method of claim 1, wherein, The microphone is authenticated with the receiver inserted into the TV based on a team matching mode, and if the authentication is successful, the microphone and the TV form a connection state, comprising: The TV and the receiver are connected based on a universal serial bus (USB) port; The microphone searches based on the team matching mode, and when the receiver is searched, the microphone sends a pairing request instruction to the receiver, the receiver generates an authentication key based on the pairing request instruction, and sends the authentication key to the microphone; The microphone authenticates the authentication key based on ObjectOutputStream sequence and ObjectInputStream reverse sequence, and if the authentication is successful, the microphone and the receiver form an interconnected pairing state; The microphone sends a connection data packet to the receiver, creates a target link of the TV and the microphone based on the connection data packet, and forms a connection state of the microphone and the TV in the interconnected pairing state based on the target link.

3. The intelligent audio processing method of claim 1, wherein, The sound effect scene control instruction is analyzed to obtain sound effect scene analysis information, comprising: The information format of the sound effect scene information is determined based on the sound effect scene control instruction; The sound effect scene control instruction is analyzed based on the application layer to obtain the sound effect scene analysis information in combination with the information format of the sound effect scene information.

4. The intelligent audio processing method of claim 1, wherein, The current playing audio of the TV is analyzed based on the original sound suppression instruction to obtain a target original sound suppression degree, comprising: The corresponding application flag information is determined based on the original sound suppression instruction, and the corresponding application program module is matched based on the application flag information; The current playing audio of the TV is analyzed based on the application program module in combination with the operation data in the original sound suppression instruction to obtain the target original sound suppression degree.

5. The intelligent audio processing method of claim 1, wherein, The time-frequency domain feature information of the current playing audio is extracted based on a conversion network, comprising: The feature sequence of the current playing audio is extracted, and the time domain autocorrelation processing is performed on the feature sequence based on a time sequence correlation layer to obtain an autocorrelation feature vector sequence; The autocorrelation feature vector sequence is processed in the frequency domain by using several different scale frequency domain convolution kernels along the frequency domain direction to obtain a frequency domain convolution vector; The frequency domain convolution vector is used to extract frequency domain features by using a frequency domain conversion network combined with a self-attention mechanism to obtain a frequency domain feature map; The time feature map is obtained by using a feature sequence based on a time conversion network, and the time feature map and the frequency domain feature map are subjected to attention fusion processing to obtain time-frequency domain feature information.

6. The intelligent audio processing method of claim 1, wherein, The current playing audio is separated into accompaniment audio and original audio based on the time-frequency domain feature information, including: An accompaniment binary mask is determined based on a mask analysis model combined with the time-frequency domain feature information; Initial accompaniment data and initial original data in the current playing audio are determined based on the accompaniment binary mask; Background sound residual data is obtained by performing background sound residual analysis on the initial original data based on an audio event classifier combined with a scoring mechanism; The accompaniment audio and the original audio are determined based on the background sound residual data, the initial accompaniment data and the initial original data by using a coding and decoding network.

7. The intelligent audio processing method of claim 1, wherein, The original audio is subjected to original sound suppression based on a linear transfer function to obtain suppressed original audio, and a target mixed audio is generated based on the suppressed original audio and the accompaniment audio, including: A first autocorrelation spectrum of the original audio and a second autocorrelation spectrum of the accompaniment audio are determined, and a cross-correlation spectrum is determined based on a target original sound suppression degree; A linear transfer function is determined based on the first autocorrelation spectrum, the second autocorrelation spectrum and the cross-correlation spectrum, and an original sound suppression frequency domain signal is determined based on the linear transfer function and the target original sound suppression degree; The original audio is subjected to original sound suppression based on the original sound suppression frequency domain signal to obtain the suppressed original audio; The target mixed audio is obtained by performing alignment and mixing processing based on the suppressed original audio and the accompaniment audio.

8. An intelligent audio processing device, characterized by The device includes: A device connection module: for the microphone to authenticate with the receiver inserted into the TV based on the team matching mode, and if the authentication is successful, the microphone and the TV form a connection state; An audio scene analysis module: for the TV to receive the audio scene control instruction issued by the microphone, and to analyze the audio scene control instruction to obtain audio scene analysis information, and for the TV to enter a target audio scene mode based on the audio scene analysis information; An original sound suppression analysis module: for receiving the original sound suppression instruction issued by the microphone after the TV enters the target audio scene mode, and for analyzing the original sound suppression degree of the current playing audio of the TV based on the original sound suppression instruction to obtain a target original sound suppression degree; A time-frequency domain feature module: for extracting time-frequency domain features of the current playing audio based on a conversion network to obtain time-frequency domain feature information; An audio separation module: for separating the current playing audio into accompaniment audio and original audio based on the time-frequency domain feature information; A mixed audio generation module: for suppressing the original audio based on a linear transfer function based on the target original sound suppression degree to obtain suppressed original audio, and for generating a target mixed audio based on the suppressed original audio and the accompaniment audio. 9.An electronic device comprising a processor and a memory, wherein, The memory is configured to store instructions, and the processor is configured to invoke the instructions in the memory, so that the electronic device executes the intelligent audio processing method in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, when the computer instructions run on the electronic device, so that the electronic device executes the intelligent audio processing method in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent TV speech enhancement control method and system based on multi-microphone noise reduction

    CN110503975A

  • Local sound reinforcement method and device, terminal equipment, medium and product

    CN118900392A