Method and apparatus for acquiring symmetric audio, electronic device, and storage medium
By performing phase spectrum shifting and waveform symmetry detection on the original audio, the problem of waveform asymmetry in audio data is solved, the volume consistency of audio signals is achieved, and the deep learning network is trained stably, thus improving the user's listening experience.
Patent Information
- Application Number
- CN202111584642.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2041-12-22
AI Technical Summary
The existing audio data contains asymmetric speech signal waveforms, which leads to unstable training of deep learning networks and results in fluctuating audio volume, affecting the user's listening experience.
By performing phase spectrum shifting processing on the original audio, multiple offset audios with phase shifts are generated, and waveform symmetry detection is performed to obtain the target audio with waveform symmetry.
It significantly improves the volume consistency of audio signals, stabilizes the training of deep learning networks, reduces the probability of errors in speech synthesis, and enhances the user's listening experience.
Smart Images

Figure CN114495895B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of data processing, and in particular to the field of speech technology and deep learning. BACKGROUND
[0002] Many audio data exist the phenomenon of asymmetric speech signal waveform, which will cause the model training of deep learning network unstable, and the volume of generated audio up and down. SUMMARY
[0003] The present disclosure provides a symmetric audio acquisition method and device, electronic equipment, storage medium and computer program product.
[0004] According to an aspect of the present disclosure, a symmetric audio acquisition method is provided, comprising:
[0005] acquiring original audio to be processed;
[0006] performing phase spectrum offset processing on the original audio to generate a plurality of offset audios with phase offset;
[0007] performing waveform symmetry detection on the plurality of offset audios to acquire target audio with waveform symmetry from the plurality of offset audios.
[0008] In the embodiment of the present disclosure, based on the phase spectrum offset processing, the original audio is processed into the target audio with waveform symmetry, which greatly improves the volume consistency of the audio signal, is conducive to the stable training of the deep learning network, reduces the error probability of speech synthesis, and improves the listening experience of the user.
[0009] According to another aspect of the present disclosure, a symmetric audio acquisition device is provided, comprising:
[0010] an acquisition module configured to acquire original audio to be processed;
[0011] a generation module configured to perform phase spectrum offset processing on the original audio to generate a plurality of offset audios with phase offset;
[0012] a detection module configured to perform waveform symmetry detection on the plurality of offset audios to acquire target audio with waveform symmetry from the plurality of offset audios.
[0013] According to another aspect of the present disclosure, an electronic device is provided, comprising at least one processor and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the symmetric audio acquisition method of the first aspect of the present disclosure.
[0014] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method for acquiring symmetric audio according to the first aspect of the present disclosure.
[0015] According to another aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the method for acquiring symmetric audio according to the first aspect of the present disclosure.
[0016] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, nor are they used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:
[0018] Figure 1 is a flowchart of the method for acquiring symmetric audio according to one embodiment of the present disclosure;
[0019] Figure 2 is a flowchart of the method for acquiring symmetric audio according to one embodiment of the present disclosure;
[0020] Figure 3 is a flowchart of the method for acquiring symmetric audio according to one embodiment of the present disclosure;
[0021] Figure 4 is a flowchart of the method for acquiring symmetric audio according to one embodiment of the present disclosure;
[0022] Figure 5 is a flowchart of the method for acquiring symmetric audio according to one embodiment of the present disclosure;
[0023] Figure 6 is a technical effect display of the present disclosure;
[0024] Figure 7 is a structural diagram of the device for acquiring symmetric audio according to one embodiment of the present disclosure;
[0025] Figure 8 is a block diagram of an electronic device for implementing the method for acquiring symmetric audio according to the embodiments of the present disclosure. DETAILED DESCRIPTION
[0026] Exemplary embodiments of the present disclosure are described herein below with reference to the accompanying drawings, in which various details are set forth to facilitate an understanding of the present disclosure. It should be appreciated that various embodiments of the present disclosure can be practiced with variations of these details as would be obvious to one of ordinary skill in the art, and the scope of the present disclosure is not intended to be limited to the details shown and described herein. For the purpose of clarity, descriptions of well-known functions and constructions are omitted so as to not unnecessarily obscure the disclosure in detail.
[0027] For the convenience of understanding the present disclosure, the technical field to which the present disclosure pertains is first explained simply below.
[0028] Data processing is the collection, storage, retrieval, processing, transformation and transmission of data. Through data processing, valuable and meaningful data for certain people can be extracted and derived from a large amount of disorganized and difficult-to-understand data.
[0029] The key technologies of voice technology in the computer field are automatic speech recognition technology and speech synthesis technology. Let the computer be able to hear, see, speak and feel, which is the development direction of future human-computer interaction. Voice is the most promising human-computer interaction method in the future, and voice has more advantages than other interaction methods.
[0030] Deep learning is a new research direction in the field of machine learning, which is introduced into machine learning to make it closer to the original goal-artificial intelligence. Deep learning is to learn the internal law and representation hierarchy of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images and sounds. Its ultimate goal is to enable machines to have analysis and learning ability like people, and to be able to recognize text, images and sound data. Deep learning is a complex machine learning algorithm, and the effect achieved in speech and image recognition far exceeds that of previous related technologies.
[0031] Figure 1 is a flowchart of a method for acquiring symmetric audio according to an embodiment of the present disclosure, as shown in Figure 1 The method comprises the following steps:
[0032] S101, acquiring original audio to be processed.
[0033] Optionally, the original audio can be acquired from the network, or audio data collected by a real person can be uploaded. The network includes web pages, short video platforms, audio databases, etc.
[0034] In order to ensure the stability of the deep learning network training, the original audio needs to meet the following requirements: no continuous noise, standard speech speed, clear pronunciation, no repetition, smooth semantics, etc. Optionally, the audio data acquired from the network or collected by a real person can be screened to obtain original audio that meets the requirements.
[0035] The original audio can be male audio, female audio, Mandarin audio, or dialect audio.
[0036] S102, performing phase spectrum offset processing on the original audio to generate a plurality of offset audios with offset phases.
[0037] The phase spectrum and the amplitude spectrum of the original audio are extracted, the phase spectrum is subjected to phase offset processing to obtain a plurality of offset phase spectrums, and the plurality of offset phase spectrums are combined with the amplitude spectrum respectively to generate a plurality of offset audios with offset phases. Since only the phase spectrum is offset and the size of the amplitude spectrum is not affected, the offset audios generated have less difference in hearing sensation compared with the original audio, and the human ear is almost unable to distinguish them.
[0038] Optionally, the value of the phase offset can be obtained by an enumeration method or can be a pre-set value.
[0039] S103, performing waveform symmetry detection on the plurality of offset audios to obtain a target audio with a symmetrical waveform from the plurality of offset audios.
[0040] The waveform data of the plurality of offset audios are obtained, and the symmetry of the waveform is detected. Optionally, the positive and negative amplitude values of the offset audio at the same time can be compared to obtain the symmetry degree of the offset audio, and the target audio with a symmetrical waveform is selected from the plurality of offset audios according to the symmetry degree.
[0041] In the embodiments of the present disclosure, the original audio to be processed is obtained, the original audio is subjected to phase spectrum offset processing to generate a plurality of offset audios with offset phases, and waveform symmetry detection is performed on the plurality of offset audios to obtain a target audio with a symmetrical waveform from the plurality of offset audios. In the embodiments of the present disclosure, the original audio is processed into a target audio with a symmetrical waveform based on phase spectrum offset processing, which greatly improves the volume consistency of the audio signal, is conducive to the stable training of the deep learning network, reduces the error probability of speech synthesis, and improves the listening experience of the user.
[0042] Figure 2 is a flowchart of a method for obtaining a symmetrical audio according to an embodiment of the present disclosure, which is further combined with the above embodiments. Figure 2 The process of generating a plurality of offset audios with offset phases is explained and described, including the following steps:
[0043] S201, performing Fourier transform on the original audio to extract a first linear frequency spectrum, wherein the first linear frequency spectrum includes an original amplitude spectrum and an original phase spectrum.
[0044] In order to convert the original audio from a time domain signal to a frequency domain signal, a time-frequency conversion, i.e., a Fourier transform, needs to be performed on the original audio. Further, for a stationary signal, a Fourier transform can be used; for a non-stationary signal, a short-time Fourier transform can be used.
[0045] By the Fourier transform, a first linear spectrum of the original audio is extracted, and based on the first linear spectrum, an original amplitude spectrum and an original phase spectrum of the original audio can be calculated.
[0046] S202, performing phase shift processing on the original phase spectrum to obtain a plurality of offset phase spectrums corresponding to the original phase spectrum.
[0047] A plurality of phase shift amounts are selected from a set phase interval, and the original phase spectrum is subjected to phase shift processing based on the plurality of phase shift amounts to generate a plurality of offset phase spectrums.
[0048] The phase interval is an open interval from 0 to 180 degrees.
[0049] Optionally, a plurality of phase shift amounts can be enumerated from the set phase interval according to a set step size. For example, if the set step size is 10 degrees, the phase shift amounts are 10 degrees, 20 degrees,..., and 170 degrees, and after the phase shift processing, 17 offset phase spectrums offset by 10 degrees, 20 degrees,..., and 170 degrees are generated.
[0050] S203, generating a plurality of offset audios based on the original amplitude spectrum and the plurality of offset phase spectrums.
[0051] The original amplitude spectrum is combined with the plurality of offset phase spectrums to generate a plurality of second linear spectrums, and the plurality of second linear spectrums are subjected to inverse Fourier transform to obtain the plurality of offset audios.
[0052] Optionally, the inverse short-time Fourier transform, the inverse wavelet transform, or the inverse Hilbert transform can also be used to obtain the offset audios corresponding to the second linear spectrums.
[0053] In the embodiments of the present disclosure, the original audio is subjected to Fourier transform to extract a first linear spectrum, wherein the first linear spectrum includes an original amplitude spectrum and an original phase spectrum, the original phase spectrum is subjected to phase shift processing to obtain a plurality of offset phase spectrums corresponding to the original phase spectrum, and a plurality of offset audios are generated based on the original amplitude spectrum and the plurality of offset phase spectrums. In the embodiments of the present disclosure, the plurality of offset audios are generated based on the phase shift processing, which lays a foundation for obtaining the symmetric audio therefrom, and since the amplitude spectrum of the audio is not changed, the offset audio has little impact on the hearing sensation and is difficult to be distinguished by the human ear.
[0054] Figure 3 is a flowchart of a method for obtaining a symmetric audio according to an embodiment of the present disclosure, which is further combined with the above embodiments.Figure 3 The process of waveform symmetry detection is explained, including the following steps:
[0055] S301, extract the amplitude value of the offset audio, and determine the waveform symmetry parameter of the offset audio based on the amplitude value, wherein the waveform symmetry parameter is used to indicate the waveform symmetry degree of the offset audio.
[0056] Extract the amplitude value of the offset audio, input the amplitude value into the scoring function, and obtain the waveform symmetry parameter of the offset audio, wherein the waveform symmetry parameter is used to indicate the waveform symmetry degree of the offset audio.
[0057] The following three basic scoring functions are listed:
[0058] 1. score = -max(abs(x[n]))
[0059] 2. score = -abs(abs(max(x[n]))-abs(min(x[n]))).
[0060] 3. score = -abs(average(abs(pos_peaks(x[n])))-average(abs(neg_peaks(x[n]))))
[0061] average(abs(neg_peaks(x[n]))))
[0062] Wherein, score represents the final score, max represents the maximum value function, min represents the minimum value function, abs represents the absolute value function, average represents the average value function, pos_peaks represents obtaining the peak value sequence greater than 0 in a certain time sequence, and neg_peaks represents obtaining the peak value sequence less than 0 in a certain time sequence.
[0063] Based on the above three scoring functions, the greater the waveform symmetry parameter value, the more symmetrical the waveform of the offset audio.
[0064] Optionally, other scoring functions can be used, or the above scoring functions can be transformed. If the minus sign in the above scoring function is changed to a plus sign, the smaller the waveform symmetry parameter value, the more symmetrical the waveform of the offset audio.
[0065] S302, based on the waveform symmetry parameter, obtain the target audio with symmetrical waveform from the plurality of offset audios.
[0066] In some implementations, the greater the waveform symmetry parameter value, the more symmetrical the waveform of the offset audio, and the target audio is the offset audio with the maximum waveform symmetry parameter value.
[0067] In some other implementations, the smaller the waveform symmetry parameter value of the offset audio, the more symmetrical the waveform of the offset audio, and the target audio is the offset audio with the smallest waveform symmetry parameter value.
[0068] In the embodiments of the present disclosure, the amplitude value of the offset audio is extracted, and the waveform symmetry parameter of the offset audio is determined based on the amplitude value, wherein the waveform symmetry parameter is used to indicate the waveform symmetry degree of the offset audio, and the target audio with symmetrical waveform is obtained from the plurality of offset audios based on the waveform symmetry parameter. In the embodiments of the present disclosure, the target audio with symmetrical waveform is screened from the plurality of offset audios based on the waveform symmetry parameter, and the waveform symmetry degree of the offset audio can be clearly seen through this method, which is convenient for screening.
[0069] Figure 4 is a flowchart of a method for obtaining symmetrical audio according to an embodiment of the present disclosure, further combining Figure 4 Before determining the waveform symmetry parameter of the offset audio, the method further includes:
[0070] S401, determining the service type of the original audio, and determining the determination strategy of the waveform symmetry parameter according to the service type.
[0071] The three scoring functions in step S301 have no fixed advantages and disadvantages, and the most reasonable scoring function needs to be selected according to the actual situation.
[0072] For example, when the maximum amplitude value of the audio is close to 1, but the average volume is low, the average volume is amplified, and the amplitude overshoot phenomenon occurs. In order to avoid amplitude overshoot and increase the space of volume equalization, the first scoring function can be selected as the determination strategy of the waveform symmetry parameter. Through the first scoring function, the offset audio with smaller maximum value can be selected, and the space of volume adjustment can be increased.
[0073] When the second or third scoring function is selected as the determination strategy of the waveform symmetry parameter, the selected target audio can be more symmetrical. Among them, the third scoring function has the most stringent requirement on the waveform, and correspondingly, the calculation amount is also the largest.
[0074] S402, obtaining the waveform symmetry parameter according to the amplitude value and the determination strategy.
[0075] The amplitude value is input into the selected scoring function to obtain the waveform symmetry parameter.
[0076] In the embodiments of the present disclosure, the service type of the original audio is determined, the determination strategy of the waveform symmetry parameter is determined according to the service type, and the waveform symmetry parameter is obtained according to the amplitude value and the determination strategy. In the embodiments of the present disclosure, different waveform symmetry parameter determination strategies are implemented according to different service types, and through this method, the waveform symmetry parameter obtained can be more reasonable and more consistent with the actual situation.
[0077] Figure 5 is a flowchart of a symmetric audio acquisition method according to an embodiment of the present disclosure, as Figure 5 shown, based on the symmetric audio acquisition method provided by the present disclosure, the acquisition process of the symmetric audio in the actual application scene includes the following steps:
[0078] S501, acquiring a to-be-processed speech signal.
[0079] S502, performing short-time Fourier transform on the to-be-processed speech signal to extract a first linear spectrum.
[0080] S503, calculating an amplitude spectrum and a phase spectrum based on the first linear spectrum.
[0081] S504, enumerating a plurality of phase offsets, and generating a plurality of offset phase spectra based on the phase offsets.
[0082] S505, combining the amplitude spectrum with the plurality of offset phase spectra respectively to generate a plurality of second linear spectra.
[0083] S506, performing inverse short-time Fourier transform on the plurality of second linear spectra respectively to obtain a plurality of offset speech signals.
[0084] S507, scoring the plurality of offset speech signals by a scoring function to obtain a waveform symmetry score of each speech signal.
[0085] S508, screening out the speech signal with the highest score as the output waveform symmetric speech signal.
[0086] For specific implementation of the present embodiment, please refer to the relevant introduction in the embodiments of the present disclosure, which will not be repeated here.
[0087] Based on the symmetric audio acquisition method of the present disclosure, the original audio can be processed into the waveform symmetric target audio, and the technical effect of the present method is as Figure 6 shown.
[0088] In the present embodiment, the original speech signal is processed into the waveform symmetric speech signal based on the phase spectrum offset processing, which greatly improves the volume consistency of the speech signal, is conducive to the stable training of the deep learning network, reduces the error probability of speech synthesis, and improves the listening experience of the user.
[0089] Figure 7 is a structural diagram of a symmetric audio acquisition device according to an embodiment of the present disclosure, as Figure 7 shown, the symmetric audio acquisition device 700 includes:
[0090] The acquisition module 710 is configured to acquire a to-be-processed original audio.
[0091] The generating module 720 is configured to perform phase spectrum offset processing on the original audio to generate a plurality of offset audios with phase offsets.
[0092] The detecting module 730 is configured to perform waveform symmetry detection on the plurality of offset audios to obtain a target audio with waveform symmetry from the plurality of offset audios.
[0093] In the embodiments of the present disclosure, the original audio is processed into the target audio with waveform symmetry based on the phase spectrum offset processing, which greatly improves the volume consistency of the audio signal, is beneficial to stable training of the deep learning network, reduces the error probability of speech synthesis, and improves the listening experience of the user.
[0094] It should be noted that the foregoing explanation and description of the embodiment of the method for obtaining symmetric audio also apply to the embodiment of the device for obtaining symmetric audio, which will not be described here again.
[0095] Further, in a possible implementation manner of the embodiment of the present disclosure, the generating module 720 is further configured to perform Fourier transform on the original audio to extract a first linear spectrum, wherein the first linear spectrum includes an original amplitude spectrum and an original phase spectrum; perform phase offset processing on the original phase spectrum to obtain a plurality of offset phase spectrums corresponding to the original phase spectrum; and generate a plurality of offset audios based on the original amplitude spectrum and the plurality of offset phase spectrums.
[0096] Further, in a possible implementation manner of the embodiment of the present disclosure, the generating module 720 is further configured to combine the original amplitude spectrum with the plurality of offset phase spectrums respectively to generate a plurality of second linear spectrums; and perform inverse Fourier transform on the plurality of second linear spectrums respectively to obtain the plurality of offset audios.
[0097] Further, in a possible implementation manner of the embodiment of the present disclosure, the detecting module 730 is further configured to extract an amplitude value of the offset audio, and determine a waveform symmetry parameter of the offset audio based on the amplitude value, wherein the waveform symmetry parameter is used to indicate a waveform symmetry degree of the offset audio; and obtain the target audio with waveform symmetry from the plurality of offset audios based on the waveform symmetry parameter.
[0098] Further, in a possible implementation manner of the embodiment of the present disclosure, the detecting module 730 is further configured to determine a service type of the original audio, determine a determination strategy of the waveform symmetry parameter according to the service type; and obtain the waveform symmetry parameter according to the amplitude value and the determination strategy.
[0099] Further, in a possible implementation manner of the embodiment of the present disclosure, the generating module 720 is further configured to select a plurality of phase offsets from a set phase interval; and perform phase offset processing on the original phase spectrum based on the plurality of phase offsets to generate the plurality of offset phase spectrums.
[0100] Further, in a possible implementation of the embodiment of the present disclosure, the generating module 720 is further configured to enumerate a plurality of phase offsets from a set phase interval according to a set step length.
[0101] In the technical solution of the present disclosure, the acquisition, storage and application of user personal information are in line with relevant laws and regulations and do not violate public order and good customs.
[0102] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0103] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0104] As shown in Figure 8 The electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0105] Various components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc., an output unit 807, such as various types of displays, a speaker, etc., the storage unit 808, such as a magnetic disk, an optical disk, etc., and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0106] The computing unit 801 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs various methods and processes described above, such as the acquiring of symmetric audio method. For example, in some embodiments, the acquiring of symmetric audio method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded onto the RAM 803 and executed by the computing unit 801, one or more steps of the acquiring of symmetric audio method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the acquiring of symmetric audio method by other any suitable means, such as by means of firmware.
[0107] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0108] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0109] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0110] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0111] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0112] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server can arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0113] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.
[0114] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for obtaining symmetric audio, comprising: obtaining original audio to be processed; performing Fourier transform on the original audio to extract a first linear spectrum, wherein the first linear spectrum comprises an original amplitude spectrum and an original phase spectrum; performing phase offset processing on the original phase spectrum to obtain a plurality of offset phase spectrums corresponding to the original phase spectrum; combining the original amplitude spectrum with a plurality of the offset phase spectrums respectively to generate a plurality of second linear spectrums; performing inverse Fourier transform on the plurality of second linear spectrums respectively to generate a plurality of offset audios with phase offset; determining a service type of the original audio; determining a determination strategy of a waveform symmetry parameter according to the service type; extracting amplitude values of the offset audios; determining the waveform symmetry parameter of the offset audios according to the amplitude values and the determination strategy, wherein the waveform symmetry parameter is used to indicate a waveform symmetry degree of the offset audios; obtaining target audio with waveform symmetry from the plurality of offset audios based on the waveform symmetry parameter.
2. The method of claim 1, wherein, The phase offset processing on the original phase spectrum to obtain a plurality of offset phase spectrums corresponding to the original phase spectrum comprises: selecting a plurality of phase offsets from a set phase interval; performing phase offset processing on the original phase spectrum based on the plurality of phase offsets respectively to generate a plurality of the offset phase spectrums.
3. The method of claim 2, wherein, The selecting a plurality of phase offsets from a set phase interval comprises: enumerating a plurality of the phase offsets from the set phase interval according to a set step.
4. An apparatus for acquiring symmetrical audio, wherein, comprising: an obtaining module configured to obtain original audio to be processed; a generating module configured to perform Fourier transform on the original audio to extract a first linear spectrum, wherein the first linear spectrum comprises an original amplitude spectrum and an original phase spectrum; perform phase offset processing on the original phase spectrum to obtain a plurality of offset phase spectrums corresponding to the original phase spectrum; combine the original amplitude spectrum with a plurality of the offset phase spectrums respectively to generate a plurality of second linear spectrums; and perform inverse Fourier transform on the plurality of second linear spectrums respectively to generate a plurality of offset audios with phase offset; a detecting module configured to determine a service type of the original audio; determine a determination strategy of a waveform symmetry parameter according to the service type; extract amplitude values of the offset audios; determine the waveform symmetry parameter of the offset audios according to the amplitude values and the determination strategy, wherein the waveform symmetry parameter is used to indicate a waveform symmetry degree of the offset audios; and obtain target audio with waveform symmetry from the plurality of offset audios based on the waveform symmetry parameter.
5. The apparatus of claim 4, wherein, The generating module is further configured to: select a plurality of phase offsets from a set phase interval; perform phase offset processing on the original phase spectrum based on the plurality of phase offsets respectively to generate a plurality of the offset phase spectrums.
6. The apparatus of claim 5, wherein, The generating module is further configured to: enumerate a plurality of the phase offsets from the set phase interval according to a set step.
7. An electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-3.
8. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are for causing the computer to perform the method of any one of claims 1-3.
9. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-3.
Citation Information
Patent Citations
Bass enhancement for audio
CN101459865A
Method and device for detecting audio distortion
CN104167209A