Sound modification based on frequency composition
By determining the classification of sounds in the audio signal and selecting a specific frequency subband for modification, the problem that audio signal modification in the prior art affects the total energy level is solved, and a more natural and harmonious audio signal output is achieved.
Patent Information
- Application Number
- CN201980097055.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-05
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2039-06-05
AI Technical Summary
Existing audio signal modification techniques are prone to affect other sounds in the audio signal, resulting in significant changes in the total energy level, causing the modified audio signal to be overly loud or too gentle, sound harsh or unpleasant.
By determining the classification of each sound in the audio signal, selecting the frequency subband associated with the sound, and modifying it on the subband without modifying the other frequency subbands, the modified audio signal is generated.
Modification of the sound in the audio signal is achieved without significantly changing the total energy level, making the modified audio signal sound more natural and harmonious.
Smart Images

Figure CN113924620B_ABST
Abstract
Description
Background Art Technical Field
[0001] Various embodiments relate generally to audio processing and, more particularly, to sound modification based on frequency content.
[0002] Description of Related Technology
[0003] Augmented reality is an area of growing interest in which a real-world environment can be enhanced by computer-generated or computer-manipulated content. Computer-manipulated content may include audio signals that have been modified from original audio signals captured by an audio input device.
[0004] The audio signal captured by the audio input device for augmented reality processing may include multiple sounds. The multiple sounds may be generated by multiple sources in the environment (e.g., people, animals, objects). Audio processing of the original audio signal may include selectively modifying the audio signal to emphasize or weaken certain sounds, thereby enhancing the auditory realism that can be perceived by the user.
[0005] Conventional methods for selectively modifying sounds include selectively increasing or decreasing the energy level of certain sounds. For example, an inverse sound wave can be generated to cancel a certain sound in an audio signal. As another example, to enhance a certain sound, the amplitude of the channel corresponding to that sound can be increased.
[0006] One drawback of these conventional methods is that other sounds in the audio signal may be inadvertently affected by the modification. For example, the reverse sound wave may inadvertently eliminate all or part of the other sounds included in the audio signal. Another drawback is that the overall energy level of the audio signal may be significantly altered by the modification. As a result, the modified audio signal may be perceived by the user as too loud or too soft when output, making the sounds in the modified audio signal sound harsh and / or unpleasant.
[0007] As stated above, there is a need for more effective sound modification techniques. Summary of the Invention
[0008] One embodiment describes a method for modifying a sound included in an audio signal. The method includes: determining, for each sound included in a plurality of sounds included in the audio signal, one or more classifications associated with the sound; selecting a first frequency subband of a first sound included in the plurality of sounds based on a first classification associated with the first sound; and modifying the first frequency subband of the first sound without modifying at least a second frequency subband of the first sound to generate a modified audio signal.
[0009] Additionally, other embodiments provide a system and one or more computer-readable storage media configured to implement the above methods.
[0010] At least one advantage and technical improvement of the disclosed technology is that one or more sounds included in an audio signal can be modified without significantly changing the overall energy level of the audio signal. As a result, the modified audio signal sounds more natural and harmonious to the user than an audio signal modified using conventional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order that the above-described features of various embodiments may be understood in detail, the inventive concept briefly summarized above may be described in more detail by reference to various embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the drawings illustrate only typical embodiments of the inventive concept and are not to be considered as limiting its scope in any way, and that other equally effective embodiments may exist.
[0012] Figure 1 A sound modification system configured to implement one or more aspects of various embodiments is shown;
[0013] Figure 2A showing a graphical representation of sounds included in an audio signal prior to modification according to one or more aspects of various embodiments;
[0014] Figure 2B According to conventional technology, Figure 2A modification of sounds included in an audio signal;
[0015] Figure 2C Shows one or more aspects of various embodiments. Figure 2A Modification of sounds in an audio signal;
[0016] Figure 2D Showing one or more aspects of various embodiments in Figure 2C a graphical representation of the sound in the audio signal after the modification shown; and
[0017] Figure 3 A flow chart illustrating method steps for selectively modifying sounds in one or more audio signals according to one or more aspects of various embodiments. DETAILED DESCRIPTION
[0018] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that these inventive concepts may be practiced without one or more of these specific details.
[0019] Embodiments disclosed herein include a sound modification system comprising one or more audio input devices and one or more audio output devices, the one or more audio input devices being arranged to acquire one or more audio signals. The sound modification system also includes a processing unit coupled to the audio input device and the audio output device, wherein the processing unit operates to selectively modify one or more sounds included in the one or more audio signals and output the one or more modified audio signals via the one or more audio output devices. Sounds included in the one or more audio signals can be selected for modification based on user input. In various embodiments, sounds included in the one or more audio signals are modified by modifying frequency subbands of the sounds. The frequency subbands for modification can be selected based on one or more classifications associated with the sounds.
[0020] The sound modification system can be implemented in various forms of audio-based systems, such as personal headphones or other wearable audio devices, home audio systems, vehicle audio systems, and the like. The sound modification system can also be implemented in various forms of audio-enabled systems, such as smartphones, tablet computers, desktop computers, laptop computers, and the like. The sound modification system can determine multiple sounds included in an audio signal and selectively modify one or more sounds included in the multiple sounds. The sound modification system can perform its processing functions using a dedicated processing device and / or a separate computing device, such as a user's mobile computing device or a cloud computing system.
[0021] Figure 1 A sound modification system 100 is shown that is configured to implement one or more aspects of various embodiments. As shown, the sound modification system 100 includes a computing device 102, one or more input devices 152, one or more audio input devices 154, and one or more audio output devices 156. The computing device 102 includes one or more processing units 110, a memory 120, and input / output (I / O) 150. The sound modification system 100 may also include a display device 158.
[0022] The one or more processing units 110 may include any processing element capable of performing the functions described herein. Although depicted as a single element within the computing device 102, the one or more processing units 110 are intended to represent a single processor, multiple processors, one or more processors with multiple cores, and combinations thereof. The processing unit 110 may be any suitable processor, such as a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), any other type of processing unit, or a combination of different processing units, such as a CPU configured to operate in conjunction with a DSP. In general, the processing unit 110 may be any technically feasible hardware unit capable of processing data and / or executing software applications or modules, including the sound modification application 122.
[0023] Memory 120 can include a variety of computer-readable media selected based on their size, relative performance, or other capabilities: volatile and / or non-volatile media, removable and / or non-removable media, etc. Memory 120 can include cache, random access memory (RAM), storage devices, etc. Of course, various memory chips, memory bandwidths, and memory form factors can be selected interchangeably. The storage devices included as part of memory 120 generally provide non-volatile memory for computing device 102 and can include one or more different storage elements, such as flash memory, hard drives, solid-state drives, optical storage devices, and / or magnetic storage devices.
[0024] Memory 120 may include one or more applications or modules for performing the functions described herein. In various embodiments, any modules and / or applications included in memory 120 may be implemented locally by sound modification system 100 and / or may be implemented via a cloud-based architecture. For example, any modules and / or applications included in memory 120 may be implemented on a remote device (e.g., a computer) that communicates with sound modification system 100 via I / O 130 or network 160. For example , server systems, cloud computing platforms, etc.).
[0025] As shown, memory 120 includes a sound modification application 122 for determining one or more sounds included in an audio signal and selectively modifying the sounds in the audio signal. In various embodiments, sound modification application 122 selects certain sounds included in the audio signal, selects certain frequency subbands included in the selected sounds, and modifies the selected subbands. Memory 120 also includes a sound database 124 that stores information about sounds, including information about sound classifications and associated frequency ranges.
[0026] Processing unit 110 may communicate with other devices, such as peripheral devices or other networked computing devices, using input / output (I / O) 150. I / O 150 may include any number of different I / O adapters or interfaces for providing the functionality described herein. I / O 150 may include wired and / or wireless connections and may use a variety of formats or protocols ( For example , (a registered trademark of the Bluetooth Special Interest Group), (registered trademark of the Wi-Fi Alliance), Universal Serial Bus (USB), etc.).
[0027] I / O 150 may also include one or more network interfaces that couple processing unit 110 to one or more networked computing devices via network 160. Examples of networked computing devices include cloud computing systems 170, server systems, desktop computers, mobile computing devices such as smartphones or tablet computers, and wearable devices such as watches or headphones or head-mounted display devices. Of course, other types of computing devices may also be networked with processing unit 110. Network 160 may include one or more networks of various types, including a local area network or local access network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet). In some embodiments, a networked computing device may be used as an additional processing unit 110, additional memory 120, input device 152, audio input device 154, and / or audio output device 156.
[0028] The input device 152 is coupled to the processing device 110 and provides various inputs to the processing device 110. In some embodiments, the input device 152 may include a user interface for receiving user input, such as user selection of certain sounds that have been determined to be included in the audio signal and user adjustment of the energy levels of certain sounds. The user interface can take any feasible form to provide the functionality described herein, such as one or more buttons, switches, sliders, dials, knobs, touch-sensitive surfaces, etc. and / or a graphical user interface (GUI). The GUI can be provided by an optional display device 158. In some embodiments, the GUI may include a user interface object ( For example , slider control), where the slider control corresponds to the energy level of the sound ( For example , amplitude, perceived loudness).
[0029] The audio input device 154 includes one or more devices that capture sound waves appearing in the environment and generate audio signals from the captured sound waves. The audio input device 154 may include one or more microphones ( For example, omnidirectional microphones, microphone arrays) and / or one or more other transducers or sensors capable of converting sound waves into electrical audio signals. The audio input device 154 may include a sensor array including sensors of a single type or multiple different sensor types. The audio input device 154 may be worn by the user, or individually set in a fixed position or movable. The audio input device 154 may be set in the environment in any feasible manner. Additionally or alternatively, the audio input device 154 may include one or more devices or systems that can receive audio from recorded media ( For example , media playback device, media storage device) provides audio signals to the sound modification system 100. Additionally or alternatively, the audio input device 154 may include one or more intermediate devices ( For example , amplifier, mixer), the one or more intermediate devices can transmit audio signals from other audio input devices 154 to the sound modification system 100.
[0030] An audio output device 156 is included to output audio signals. The audio output device 156 may use any technically feasible audio output technology, such as a speaker or other suitable electroacoustic device. The audio output device 156 may be implemented using any number of form factors, such as a discrete speaker device, an on-device speaker, circumaural (circumaural), supra-aural (ear-hook), or in-ear headphones, hearing aids, wired or wireless headsets, body-worn ( For example , head-mounted, shoulder-mounted, arm-mounted, etc.) listening device, body-worn close-range directional speaker or speaker array, body-worn ultrasonic speaker array, etc. The audio output device 156 can be worn by the user, or separately set in a fixed position or movable.
[0031] It should be understood that various embodiments of the sound modification system 100 may have Figure 1 For example, in one embodiment, the computing device 102, input device 152, audio input device 154, audio output device 156, and optional display device 158 may be included in one device such as a smartphone, a tablet computer, a user-wearable headset, headphones, etc. In another embodiment, the audio input device 154 and the audio output device 156 may be separate from the computing device 102. For example, the audio input device 154 and the audio output device 156 may be included in a separate computing device 102 (e.g., a separate computer) connected via a wired or wireless connection. For example , smartphones, laptops, etc.) in the headset.
[0032] In operation, the sound modification application 122 obtains one or more audio signals ( For example, via the audio input device 154), and perform various operations on the one or more audio signals. In various embodiments, the sound modification application 122 can determine multiple sounds included in the one or more audio signals, thereby detecting and / or identifying sounds included in the one or more audio signals that are generated by different sources. The determination can include determining a corresponding audio channel, stream, and / or signal for each detected sound. The sound modification application 122 can use any technically feasible technique to determine the multiple sounds ( example like , sound spatialization, sound segmentation or decomposition, sound object detection, machine learning, etc.). In some embodiments, the function of determining multiple sounds can be performed by another module or application in the memory 120 that is separate from the sound modification application 122. It should be understood that although the embodiments described herein are described with respect to one audio signal and the sounds included in the audio signal for ease of understanding, the embodiments described are also applicable to multiple audio signals and the sounds included in the multiple audio signals. For example, the sound modification application 122 can obtain multiple audio signals ( For example , from multiple microphones and / or from a multi-channel playback device) and determine multiple sounds included in the multiple audio signals. In some embodiments, each audio signal can correspond to a different audio channel, each of the audio channels including different sounds.
[0033] The sound modification application 122 may also determine one or more classifications for each of the determined sounds included in the one or more audio signals. The classification associated with a sound may indicate the type of source producing the sound. For example, possible sound classifications include the type of person ( For example , male, female), type of animal ( For example , dog, cat, bird) and the type of object ( For example , vehicles, construction equipment). The sound modification application 122 may use any technically feasible technique to determine one or more classifications associated with the sound ( For example , machine learning, etc.). Possible classifications may include classifications at one level of granularity or at different levels of granularity. For example, at a lower, less detailed level of granularity, the classifications may include, for example, "person," "animal," and "object." At a higher, more detailed level of granularity, the classifications may include more specific classifications ( For example , “male voice”, “female voice”, “dog”, “cat”, “bird”, “car”, “traffic”, “airplane”, “construction equipment”, etc.).
[0034] In some embodiments, the determination of a sound in one or more audio signals and the determination of a classification of the sound may be performed together. For example, as described above, the processing of the audio signals for detecting and determining a sound in one or more audio signals may also include the processing of identifying the source of the sound and determining one or more classifications associated with the sound based on the identification. In some embodiments, the functionality of determining the classification of the sound may be performed by another module or application in memory 120 that is separate from the sound modification application 122 ( For example , as described above, in the same module or application as the module or application that performs the function of determining the sound).
[0035] The sound modification application 122 can receive user input that selects sounds included in one or more audio signals for modification. In some embodiments, the sound modification application 122 presents a user interface ( For example , GUI, audio prompts), which enable and / or prompt the user to adjust the desired energy level of a particular sound. For example, the GUI displayed by the sound modification application 122 on the display device 158 may include a slider or other control ( For example , dial, knob) for adjusting the level of each sound determined to be included in one or more audio signals. The slider control can be used with the energy level of the sound ( For example , an energy level relative to the original energy level in the audio signal). The user can move the slider of a sound via input device 152 to select the sound for modification and indicate the direction and approximate amount of modification to the sound. The user can move the slider of one or more sounds to select a sound for modification and indicate the direction and amount of modification to the selected sound.
[0036] In response to receiving user input selecting a sound for modification, the sound modification application 122 may determine frequency subbands within the selected sound for modification. In some embodiments, the sound modification application 122 determines the frequency subbands based at least on a classification associated with the sound. Specifically, the sound modification application 122 may determine one or more characteristic frequency subbands of the sound based on one or more classifications of the sound. In some embodiments, the characteristic frequency subbands of a sound classification indicate frequency ranges that are representative of the sound associated with that classification and contribute to the clarity of that sound. For example, if a sound is classified as a female voice, the sound modification application 122 may identify one or more frequency subbands that are unique to female voices and select one or more of these subbands for modification. Furthermore, in various embodiments, the characteristic frequency subbands of the sound may be frequency subbands that are typical for the sound and / or frequency subbands that significantly contribute to the discrimination, isolation, and / or perceived loudness of the user's voice. For example, the characteristic frequency subbands of a musical instrument may be important for the user to be able to hear the instrument among multiple sounds. As another example, amplifying the characteristic frequency subbands of the musical instrument may amplify the perceived loudness of the sound. As another example, modifying a frequency sub-band of an instrument sound outside of the sound's characteristic frequency sub-band may cause the sound to be perceived as unnatural or atypical for the instrument.
[0037] In some embodiments, the sound modification application 122 can identify characteristic frequency sub-bands for classification by referring to information stored in the sound database 124. For example, the sound modification application 122 can obtain information indicating the characteristic frequency sub-bands from the sound database 124. The sound modification application 122 can then modify the sound within the audio signal in the selected frequency sub-bands and generate an output audio signal including the modified sound, as will be described in further detail below.
[0038] In some embodiments, the sound database 124 includes a database of information about various sound categories. For a given sound category, the sound database 124 includes information about the characteristic frequency subbands associated with that category. Additionally, in some embodiments, the sound database 124 may include references from one category to one or more other categories. For example, a higher-level category may reference or point to one or more lower-level categories ( For example , animal sound classifications can reference classifications corresponding to specific types of animals, and vice versa. The information stored in sound database 124 can be stored and queried using any technically feasible database storage and query technology. In some embodiments, sound database 124 can be located in cloud computing system 170.
[0039] Figure 2AA graphical representation 202 of sounds included in an audio signal before modification, according to one or more aspects of various embodiments, is shown. As shown, graphical representation 202 includes a line graph representing a first sound 204 and a second sound 206 that have been determined to be included in the original audio signal, plotted against a frequency axis and an amplitude axis. Sound 204 has been determined to be associated with the "traffic" sound classification, and sound 206 has been determined to be associated with the "male voice" sound classification. As can be seen from representation 202, sounds 204 and 206 have different amplitudes at different frequency ranges.
[0040] Figure 2A Also shown are slider controls 208 and 210 for sounds 204 and 206, respectively. The sound modification application 122 may present ( For example , shown in the GUI) slider controls 208 and 210 for two sounds 204 and 206, thereby enabling a user to select and adjust one or more of the sounds 204 and 206. As shown, the slider controls 208 and 210 can start at a neutral position, where the neutral position can represent the energy level of the corresponding sound in the original audio signal.
[0041] Figure 2B According to conventional technology, Figure 2A If the user selects traffic sound 204 for adjustment, the traffic sound can be modified. For example, Figure 2B As shown, the amplitude of the traffic sound 204 is increased by a certain amount across the entire frequency band of the traffic sound 204. However, modifying the amplitude of the sound across the entire frequency band has the disadvantage of increasing the overall energy level of the output audio signal, which may be perceived by the user as significantly increasing the loudness of the output audio signal. This increased loudness may make the user feel unpleasant while listening to the output audio signal.
[0042] Figure 2C Shows one or more aspects of various embodiments. Figure 2A As described above, the sound modification application 122 can present ( For example , shown in the GUI) slider controls 208 and 210 for sounds 204 and 206. The user can manipulate the traffic sound slider 208 to select the traffic sound 204 for modification and indicate the direction and approximate amount of modification. For example, the traffic sound slider 208 is Figure 2C Shown in Figure 2A, the slider 208 shown in FIG has been moved upward, indicating that the user wishes to modify the traffic sound 204 upward. For example, the traffic sound 204 can be modified upward to make the sound 204 more prominent and / or more audible compared to other sounds. Additionally or alternatively, the sound can be modified downward to make the sound less prominent and / or less audible compared to other sounds. In response to the user selection, the sound modification application 122 increases the energy level of one or more frequency subbands of the traffic sound 204 ( For example , amplitude).
[0043] In some embodiments, the sound modification application 122 obtains information indicating frequency subbands from the sound database 124, where a frequency subband is a range of frequencies. In particular, with respect to the sound 204, the sound modification application 122 obtains information about the "traffic" sound classification from the sound database 124. Based on the obtained information (which may include information indicating characteristic frequency subbands of the "traffic" sound), the sound modification application 122 may select a frequency subband for modification among the characteristic subbands. For example, if the obtained information indicates that the sound associated with the "traffic" classification has a characteristic subband in the range of 2500 Hz to 3500 Hz, the sound modification application 122 may select a 2500 Hz to 3500 Hz subband, or a narrower subband in the range of 2500 Hz to 3500 Hz for modification. For example, in Figure 2C In FIG. 2 , a range centered around 3000 Hz has been selected for the traffic sounds 204 .
[0044] In some embodiments, the sound modification application 122 may analyze the selected sound in the one or more audio signals, and optionally also analyze other sounds in the one or more audio signals and / or analyze the one or more audio signals as a whole. The sound modification application 122 may use this analysis to determine and / or adjust the frequency sub-bands that will be selected for modification, in conjunction with or in lieu of determining the sub-bands based on information obtained from the sound database 124. In some embodiments, if the audio signal includes only one sound, that sound may be analyzed to determine the frequency sub-bands, rather than using information from the sound database 124 to determine the frequency sub-bands that are unique to the sound included in the audio signal. The sound modification application 122 may use any technically feasible technique ( For example , spectrogram analysis) to analyze the sound 204 to determine and select one or more characteristic frequency subbands in the sound 204. In some embodiments, the sound modification application 122 may first determine the frequency subbands in the sound 204 based on information from the sound database 124, and then adjust the determined subbands based on the analysis of the sound 204. The adjustment may include shifting the center of the subband and / or widening or narrowing the bandwidth of the subband.
[0045] After selecting the 3000 Hz sub-band for modification, the sound modification application 122 proceeds to modify the traffic sound 204 at the selected sub-band. In some embodiments, the sound modification application 122 uses parametric equalization techniques to modify the sound. The sound modification application 122 obtains the center frequency and bandwidth of the parametric equalization from the selected frequency sub-band and determines the amount of modification for the parametric equalization based on the slider control 208 manipulated by the user. In some embodiments, the amount of modification is based on the percentage of the upward or downward modification of the slider position ( For example , the percentage of increase or decrease in amplitude). In some other embodiments, the modification amount is based on the absolute modification amount of the slider position ( For example , the increase or decrease of the amplitude). Therefore, in Figure 2C , the amplitude of the sub-band portion 214 of the traffic sound 204 is increased to the modified portion 216 .
[0046] In some embodiments, the sound modification application 122 may also automatically modify one or more other sounds included in the one or more audio signals to make the sound modified according to the user's selection more prominent or less prominent. The automatic modification of the other sounds may occur without the user manually manipulating one or more sliders for the one or more other sounds. For example, Figure 2C As shown, the male voice sound 206 can be modified to make the upwardly modified traffic sound 204 more prominent. Therefore, the subbands of the male voice sound 206 can be selected in a manner similar to that described above. Furthermore, in some embodiments, the subbands in the male voice sound 206 can be selected based on proximity to the modified subband portion 216 in the traffic sound 204. For example, portion 218 in the male voice sound 206 can correspond to a subband that is close to the subband portion 216 in the traffic sound 204 and within the characteristic subband of the male voice sound 206. In some embodiments, a subband is close to another subband if the center frequencies of the two subbands differ by less than a predetermined amount and / or the two subbands overlap. In some other embodiments, the sound modification application 122 only modifies one or more sounds that have been affirmatively selected for modification by the user ( For example , the sound for which the user has manipulated the corresponding slider control).
[0047] More generally, in various embodiments, the sound modification application 122 can make a sound more or less prominent by any combination of: 1) modifying the sound across its entire frequency band ( For example , according to the above combination Figure 2B 2) Modify the frequency sub-bands of the sound ( For example , characteristic frequency subbands); and 3) modifying the frequency subbands of one or more other sounds ( For example, characteristic frequency sub-bands). For example, to make the first sound more prominent, the sound modification application 122 may simply amplify the first sound. As another example, the sound modification application 122 may amplify a specific frequency sub-band of the first sound ( For example , characteristic subbands). As a further example, the sound modification application 122 may attenuate one or more frequency subbands of one or more other sounds that perceptually compete with the first sound ( For example , characteristic subbands). As another example, sound modification application 122 may modify one or more frequency subbands of a first sound in one direction, and modify one or more frequency subbands of one or more other sounds that perceptually compete with the first sound in the opposite direction. Sound modification application 122 may select any combination of the above modifications based on the specific sounds included in the one or more audio signals. For example, if the sound to be made more prominent has a very low amplitude compared to other sounds, among the modifications that may be performed, sound modification application 122 may amplify the sound across its entire frequency band.
[0048] In some embodiments, a first sound is perceptually competitive with a second sound if the perceptual prominence of the second sound affects the perceptual prominence of the first sound based on any number of criteria, including but not limited to sound amplitude, characteristic frequency sub-band, etc. In some embodiments, the sound modification application 122 can analyze sounds in one or more audio signals and identify perceptually competitive sounds in any technically feasible manner, including but not limited to machine learning-based techniques and searching in a database ( For example , search in the sound database 124).
[0049] After selecting the subbands for the male voice sound 206, the sound modification application 122 modifies the subbands in the male voice sound 206 using, for example, parametric equalization techniques as described above. The modification to the male voice sound 206 may be in the opposite direction to the modification to the traffic sound 204, and in some embodiments, by approximately the same amount. Figure 2C , the amplitude of the sub-band portion 218 of the male speech sound 206 is reduced to a modified portion 220 .
[0050] Figure 2D Showing one or more aspects of various embodiments in Figure 2C A graphical representation of the sounds included in the audio signal after modification is shown. Figure 2D The graphical representation 202 in FIG. 1 shows a modified traffic sound 204 and a modified male voice sound 206. Figure 2A and Figure 2B The sounds shown are compared due to the modifications applied via the above techniques, Figure 2DThe traffic sound 204 and the male voice sound 206 shown have different peaks and valleys. Figure 2A 2. The overall energy level of the sound does not change significantly compared to the sound shown in FIG. 2. After completing the modification, the sound modification application 122 can generate a modified audio signal that includes the modified sounds 204 and 206. The sound modification application 122 can also cause the modified audio signal to be output to the user via one or more audio output devices 156. For example, the sound modification application 122 transmits the modified audio signal to the audio output device 156.
[0051] Figure 3 Flowchart showing method steps for selectively modifying sounds in one or more audio signals according to one or more aspects of various embodiments. Figures 1 to 2A to Figure 2D Although the method steps are described using a system, one skilled in the art will understand that any system configured to perform the methods in any order falls within the scope of the various embodiments.
[0052] like Figure 3 As shown, the method 300 begins at step 302, where the sound modification application 122 obtains one or more audio signals. The one or more audio signals can be input via one or more audio input devices 154 ( For example , a microphone that captures sound from the environment). In step 304, the sound modification application 122 determines a plurality of sounds ( For example , multiple sounds in an audio signal, sounds in each of one or more audio signals). The sound modification application 122 may use any technically feasible technique ( For example , machine learning, sound object detection, sound spatialization, sound segmentation or decomposition, etc.) to determine the sound.
[0053] At step 306, the sound modification application 122 determines the classification of the sounds included in the one or more audio signals. For each sound determined in step 304, the sound modification application 122 determines one or more classifications ( For example , whether the sound is that of a person, animal, or object, whether it is that of a specific type of person, animal, or object, etc.). One or more classifications associated with the sound may indicate the source (or type of source) of the sound. Sound modification application 122 may determine the classification using any technically feasible technique, which may include one or more of the same techniques used to determine the sound in step 304.
[0054] At step 308, the sound modification application 122 selects a sound from the one or more audio signals. The sound modification application 122 may select a sound based on user input. In various embodiments, the sound modification application 122 may present an interface ( For example , by displaying a graphical user interface in the application via the display device 158). The user can manipulate elements in the user interface ( For example , corresponding to a slider control of a sound, such as slider controls 208 and 210) to select a sound for modification. The sound modification application 122 can select a sound selected by the user via the user interface ( For example , based on a slider control manipulated by the user). In some embodiments, the sound modification application 122 can be configured by the user with rules specifying sounds to automatically modify upward or downward based on one or more specified categories. The sound modification application 122 can automatically select sounds based on those rules.
[0055] At step 310, the sound modification application 122 determines and selects frequency subbands in the sound selected in step 308. The sound modification application 122 determines characteristic frequency subbands in the sound and selects one or more of these subbands. The selection may be based on information obtained from the sound database 124 regarding one or more classifications associated with the sound and / or based on an analysis of the sound in the audio signal. For example , spectrogram analysis), to determine the characteristic sub-bands.
[0056] At step 312, the sound modification application 122 modifies the one or more frequency sub-bands selected in step 310. For example, Figure 2C As shown, the amplitude of portion 216 of sound 204 increases in accordance with the user's manipulation of slider control 208. The sound modification application may use any technically feasible technique ( For example , parametric equalization, graphic equalization) to modify the subband.
[0057] At step 314, the sound modification application 122 determines whether there are additional sounds to be modified in the one or more audio signals ( For example , by automatically modifying other sounds so that the modification of the modified sound is more prominent, based on a user selection via a user interface, by automatically modifying other sounds based on rules). If there are additional sounds to be modified, the method proceeds to step 308 and another sound to be modified is selected in the one or more audio signals. If there are no additional sounds to be modified, the method proceeds to step 316.
[0058] At step 316, the sound modification application 122 generates one or more audio signals ("modified audio signal(s)") that include the sound modified in step 312. At step 318, the sound modification application 122 causes the modified audio signal(s) to be output to the user as sound waves ( For example , via audio output device 156). For example, the sound modification application 122 may transmit one or more modified audio signals to one or more audio output devices 156. The sound modification application 122 may also cause one or more unmodified audio signals ( For example , including audio signals of sounds that have not been modified as described above) are output via the audio output device 156 together with one or more modified audio signals.
[0059] In some embodiments, one or more of the operations and techniques described above may be performed at the cloud computing system 170 in conjunction with the computing device 102. For example, the computing device 102 may transmit an audio signal obtained via the audio input device 154 for processing by the cloud computing system 170. In such embodiments, the processing at the cloud computing system 170 may include one or more operations performed by the sound modification application 122 described above ( For example , one or more of: determining a sound in the audio signal, determining a classification associated with the sound, and determining and selecting a frequency sub-band). Cloud computing system 170 may include one or more applications or modules that perform one or more of the same or similar operations as those performed by sound modification application 122, as described above. Additionally, in some embodiments, cloud computing system 170 may include historical and / or trained data ( For example , trained sound detection neural networks, etc.), the history and / or trained data can be used to assist in the above operations, and the cloud computing system 170 can further add to the history and / or trained data based on data received from the computing device 102.
[0060] In some embodiments, the operations and techniques described above are performed in real time or near real time, so that the time delay between capturing the sound via the audio input device 154 and outputting the modified sound via the audio output device 156 can be minimized. That is, the modified audio signal can be output in real time or near real time. Therefore, in some embodiments, the operations and techniques described above can be performed entirely locally on the computing device 102. Alternatively, the operations and techniques performed at the computing device 102 can be performed in conjunction with the cloud computing system 170 ( For example , querying the sound database 124 located at the cloud computing system 170 to determine and select frequency sub-bands, allowing the cloud computing system 170 to determine and classify the sounds in the audio signal, etc.).
[0061] In summary, the sound modification system determines a plurality of sounds included in one or more audio signals and determines, for each sound included in the plurality of sounds, one or more classifications associated with the sound. In some embodiments, the classifications may include a type of person ( For example , male, female), user identity ( For example , identity of a specific person), type of animal ( For example , cat, dog, bird) and the type of object ( For example , one or more of a car, traffic, construction equipment, and an airplane). The sound modification system selects a first sound from the plurality of sounds to modify. The sound modification system then determines a first frequency subband of the first sound based at least on the classification of the first sound, and modifies the first frequency subband. In some embodiments, the sound modification system modifies the first frequency subband by increasing or decreasing the amplitude of the first frequency subband of the first sound. The sound modification system may also select a second sound, determine a second frequency subband of the second sound, and modify the second frequency subband of the second sound.
[0062] At least one advantage and technical improvement of the disclosed technology is that one or more sounds included in one or more audio signals can be modified without significantly changing the overall energy level of the one or more audio signals. For example, the perceived loudness and / or timbre of the sounds in the modified audio signals can be substantially the same as before the modification. Modifying the sounds in the audio signals does not inadvertently cause the user to be oblivious to other sounds in the audio signals. As a result, the modified audio signals sound more natural and harmonious to the user than audio signals modified using conventional methods.
[0063] 1. In some embodiments, a computer-implemented method for modifying a sound included in an audio signal includes: determining one or more classifications associated with each sound included in a plurality of sounds included in at least one audio signal; selecting a first frequency subband of a first sound included in the plurality of sounds based on a first classification associated with the first sound; and modifying the first frequency subband of the first sound without modifying at least a second frequency subband of the first sound to generate a modified audio signal.
[0064] 2. The method according to clause 1 further includes selecting a third frequency subband of the second sound based on a second classification associated with the second sound included in the plurality of sounds, an analysis of the second sound, a frequency range of the first frequency subband, and at least one of a center frequency of the first frequency subband; and modifying the third frequency subband of the second sound without modifying at least the fourth frequency subband of the second sound.
[0065] 3. The method of clause 1 or 2, wherein modifying the first frequency subband of the first sound comprises performing parametric equalization on the first frequency subband.
[0066] 4. The method of any one of clauses 1 to 3, further comprising receiving user input, wherein the modifying is performed based on the user input.
[0067] 5. A method according to any one of clauses 1 to 4, further comprising displaying a user interface comprising a control object for each sound included in the plurality of sounds, wherein the user input is received via the control object corresponding to the first sound.
[0068] 6. A method according to any one of clauses 1 to 5, wherein selecting the first frequency sub-band of the first sound includes obtaining characteristic frequency information associated with the one or more classifications associated with the first sound from a database; and selecting the first frequency sub-band based on the information.
[0069] 7. A method according to any of clauses 1 to 6, wherein the modifying is performed in response to a determination that the first sound perceptually competes with a second sound included in the plurality of sounds.
[0070] 8. A method according to any one of clauses 1 to 7, further comprising generating at least a second audio signal comprising the modified audio signal, wherein the at least second audio signal comprises the plurality of sounds, and wherein the first frequency subband of the first sound included in the at least second audio signal is modified, and the second frequency subband of the first sound included in the at least second audio signal is not modified.
[0071] 9. In some embodiments, one or more non-transitory computer-readable storage media store instructions that, when executed by at least one processor, cause the at least one processor to perform the following steps: determining, for each sound included in a plurality of sounds included in at least one audio signal, one or more classifications associated with the sound; selecting a first frequency subband of a first sound included in the plurality of sounds based on a first classification associated with the first sound; modifying the first frequency subband of the first sound without modifying at least a second frequency subband of the first sound; selecting a third frequency subband of a second sound included in the plurality of sounds; and modifying the third frequency subband of the second sound without modifying at least a fourth frequency subband of the second sound to generate a modified audio signal.
[0072] 10. The one or more computer-readable storage media of clause 9, wherein modifying the first frequency subband of the first sound comprises performing parametric equalization on the first frequency subband.
[0073] 11. One or more computer-readable storage media according to clause 9 or 10, wherein the one or more computer-readable storage media also include receiving user input, wherein modifying the first frequency subband of the first sound includes increasing or decreasing the amplitude of the first frequency subband based on the user input.
[0074] 12. One or more computer-readable storage media according to any one of clauses 9 to 11, further comprising displaying a user interface comprising a control object for each sound included in the plurality of sounds.
[0075] 13. One or more computer-readable storage media according to any one of clauses 9 to 12, wherein selecting the third frequency subband of the second sound comprises selecting the third frequency subband based on at least one of a second classification associated with the second sound, an analysis of the second sound, a frequency range of the first frequency subband, and a center frequency of the first frequency subband.
[0076] 14. One or more computer-readable storage media according to any one of clauses 9 to 13, wherein selecting the first frequency sub-band of the first sound comprises obtaining characteristic frequency information associated with the one or more classifications associated with the first sound from a database; and selecting the first frequency sub-band based on the information.
[0077] 15. One or more computer-readable storage media according to any one of clauses 9 to 14, the one or more computer-readable storage media also comprising generating at least a second audio signal comprising the modified audio signal, wherein the at least second audio signal comprises the plurality of sounds, the first frequency sub-band of the first sound included in the at least second audio signal is modified, and the second frequency sub-band of the first sound included in the at least second audio signal is not modified, and wherein the third frequency sub-band of the second sound included in the at least second audio signal is modified, and the fourth frequency sub-band of the second sound included in the at least second audio signal is not modified.
[0078] 16. A system comprising: a memory; and at least one processor, the at least one processor being coupled to the memory and configured to: detect multiple sounds included in at least one audio signal; determine one or more classifications associated with each sound included in the multiple sounds; select a first frequency subband of a first sound included in the multiple sounds based on a first classification associated with the first sound; and modify the first frequency subband of the first sound without modifying at least a second frequency subband of the first sound to generate a modified audio signal.
[0079] 17. A system according to clause 16, wherein the at least one processor is further configured to select a third frequency subband of the second sound based on a second classification associated with a second sound included in the plurality of sounds, an analysis of the second sound, a frequency range of the first frequency subband, and at least one of a center frequency of the first frequency subband; and to modify the third frequency subband of the second sound without modifying at least a fourth frequency subband of the second sound.
[0080] 18. The system of clause 16 or 17, wherein the first classification is one of a human voice, an animal voice, or an object voice.
[0081] 19. A system according to any one of clauses 16 to 18, the system further comprising a database, wherein the database comprises at least one mapping of the first classification to one or more characteristic frequency sub-bands, and wherein the one or more characteristic frequency sub-bands comprises the first frequency sub-band.
[0082] 20. A system according to any one of clauses 16 to 19, wherein the at least one processor is further configured to generate at least a second audio signal comprising the modified audio signal, wherein the at least second audio signal comprises the plurality of sounds, and the first frequency sub-band of the first sound included in the at least second audio signal is modified, and the second frequency sub-band of the first sound included in the at least second audio signal is not modified; and cause the at least second audio signal to be output via an audio output device.
[0083] Any and all combinations of any one of the claim elements of any of the claims and / or any elements described in this application, in any form, are within the intended scope of the present embodiments and protection.
[0084] The description of the various embodiments has been presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
[0085] Aspects of the present embodiment may be embodied as a system, method or computer program product. Therefore, aspects of the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects, which may be collectively referred to herein as a "module" or "system". In addition, any hardware and / or software technology, process, function, component, engine, module or system described in the present disclosure may be implemented as a circuit or a collection of circuits. In addition, aspects of the present disclosure may take the form of a computer program product, which is embodied in one or more computer-readable media having computer-readable program code embodied thereon.
[0086] Any combination of one or more computer-readable media may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media would include the following media: an electrical connection having one or more conductors, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0087] Aspects of the present disclosure are described with reference to the flowchart illustrations and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each frame in the flowchart illustration and / or block diagram and the combination of frames in the flowchart illustration and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine. When the processor of a computer or other programmable data processing device executes instructions, it is possible to implement the function / action specified in one or more frames of the flowchart and / or block diagram. Such a processor may be, but is not limited to, a general-purpose processor, a special-purpose processor, an application-specific processor or a field programmable gate array.
[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functions and operations of possible implementations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, segment or portion of a code, and the code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions proposed in the box may not occur in the order proposed in the accompanying drawings. For example, depending on the functions involved, the two boxes shown in succession can be executed substantially simultaneously, or the boxes can sometimes be executed in the opposite order. It should also be noted that each box of the block diagram and / or flowchart illustration and the combination of boxes in the block diagram and / or flowchart illustration can be implemented by a system based on dedicated hardware that performs the specified function or action or a combination of dedicated hardware and computer instructions.
[0089] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, which is determined by the claims that follow.
Claims
1. A computer-implemented method for modifying a sound included in an audio signal, comprising: determining, for each sound included in a plurality of sounds included in at least one audio signal, one or more classifications associated with the sound; selecting a first frequency sub-band of a first sound included in the plurality of sounds based on a first classification associated with the first sound; modifying the first frequency sub-band of the first sound to generate a modified first frequency sub-band; and A second audio signal is generated by combining the modified first frequency sub-band and the unmodified second frequency sub-band of the first sound.
2. The method according to claim 1, further comprising: selecting a third frequency sub-band of a second sound included in the plurality of sounds based on at least one of a second classification associated with the second sound, an analysis of the second sound, a frequency range of the first frequency sub-band, and a center frequency of the first frequency sub-band; and The third frequency sub-band of the second sound is modified without modifying at least a fourth frequency sub-band of the second sound. 3 . The method of claim 1 , wherein modifying the first frequency subband of the first sound comprises performing parametric equalization on the first frequency subband. The method of claim 1 , further comprising receiving user input, wherein the modifying is performed based on the user input. 5 . The method of claim 4 , further comprising displaying a user interface including a control object for each sound included in the plurality of sounds, wherein the user input is received via the control object corresponding to the first sound.
6. The method of claim 1 , wherein selecting the first frequency sub-band of the first sound comprises: obtaining, from a database, characteristic frequency information associated with the one or more categories associated with the first sound; and The first frequency sub-band is selected based on the information. 7 . The method of claim 1 , wherein the modifying is performed in response to a determination that the first sound perceptually competes with a second sound included in the plurality of sounds.
8. One or more non-transitory computer-readable storage media storing instructions that, when executed by at least one processor, cause the at least one processor to perform the following steps: determining, for each sound included in a plurality of sounds included in at least one audio signal, one or more classifications associated with the sound; selecting a first frequency sub-band of a first sound included in the plurality of sounds based on a first classification associated with the first sound; modifying the first frequency subband of the first sound to generate a first modified frequency subband; selecting a second frequency sub-band of a second sound included in the plurality of sounds; modifying the second frequency subband of the second sound to generate a second modified frequency subband; as well as At least a second audio signal is generated by combining the first modified frequency subband and a first unmodified version of the third frequency subband of the first sound, and by combining the second modified frequency subband and a second unmodified version of the fourth frequency subband of the second sound.
9. The one or more non-transitory computer-readable storage media of claim 8, wherein modifying the first frequency subband of the first sound comprises performing parametric equalization on the first frequency subband.
10. The one or more non-transitory computer-readable storage media of claim 8, further comprising receiving user input, wherein modifying the first frequency subband of the first sound comprises increasing or decreasing an amplitude of the first frequency subband based on the user input.
11. The one or more non-transitory computer-readable storage media of claim 8, further comprising displaying a user interface comprising a control object for each sound included in the plurality of sounds.
12. One or more non-transitory computer-readable storage media according to claim 8, wherein selecting the third frequency sub-band of the second sound comprises selecting the third frequency sub-band based on at least one of a second classification associated with the second sound, an analysis of the second sound, a frequency range of the first frequency sub-band, and a center frequency of the first frequency sub-band.
13. The one or more non-transitory computer-readable storage media of claim 8, wherein selecting the first frequency sub-band of the first sound comprises: obtaining, from a database, characteristic frequency information associated with the one or more categories associated with the first sound; and The first frequency sub-band is selected based on the information.
14. A system for modifying a sound included in an audio signal, comprising: Memory; and at least one processor coupled to the memory and configured to: detecting a plurality of sounds included in at least one audio signal; determining, for each sound included in the plurality of sounds, one or more classifications associated with the sound; selecting a first frequency sub-band of a first sound included in the plurality of sounds based on a first classification associated with the first sound; modifying the first frequency sub-band of the first sound to generate a modified first frequency sub-band; as well as A second audio signal is generated by combining the modified first frequency sub-band and the unmodified second frequency sub-band of the first sound.
15. The system of claim 14, wherein the at least one processor is further configured to: selecting a third frequency sub-band of a second sound included in the plurality of sounds based on at least one of a second classification associated with the second sound, an analysis of the second sound, a frequency range of the first frequency sub-band, and a center frequency of the first frequency sub-band; and The third frequency sub-band of the second sound is modified without modifying at least a fourth frequency sub-band of the second sound. The system of claim 14 , wherein the first classification is one of a human voice, an animal voice, or an object voice.
17. The system of claim 14, further comprising a database, wherein the database comprises at least one mapping of the first classification to one or more characteristic frequency sub-bands, and wherein the one or more characteristic frequency sub-bands comprises the first frequency sub-band.
Citation Information
Patent Citations
Method and a hearing device for improved separability of target sounds
US20170374478A1
Real-time audio source separation using deep neural networks
US20180122403A1