Audio quality conversion device and control method thereof

The audio quality conversion device uses a trained neural network to enhance audio quality by adjusting sound characteristics based on environmental data, addressing the disparity between low- and high-performance recording equipment.

JP7787525B2Active Publication Date: 2025-12-17コッヘル インコーポレーテッド
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023575988
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-09
Filing Date
2022-06-09
Publication Date
2025-12-17
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

Existing audio data recorded in a first environment using low-performance recording equipment lacks the quality comparable to high-performance equipment, necessitating a method to enhance audio quality conversion.

Method used

An audio quality conversion device employing an artificial neural network trained with audio data from different recording environments, utilizing environmental data tags like distance, noise level, and spatial reverberation to convert audio data to match desired sound quality characteristics.

Benefits of technology

The device effectively converts audio data to resemble high-quality recordings, irrespective of the recording equipment's performance, by leveraging a trained neural network to adjust sound quality based on environmental data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007787525000001
    Figure 0007787525000001
  • Figure 0007787525000002
    Figure 0007787525000002
Patent Text Reader

Abstract

The audio quality conversion device according to the present invention includes a control unit having an artificial neural network that performs learning using a plurality of audio data recorded in different recording environments for a specific audio event and environmental data related to the recording environments corresponding to each of the audio data, and an audio input unit that receives external sound and generates audio recording data, and the control unit converts the audio recording data generated by the audio input unit based on the learning result of the artificial neural network.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus for correcting the quality of audio data and a control method thereof, and more particularly to an apparatus for improving the sound quality of audio data and a control method thereof. [Background technology]

[0002] Recently, artificial intelligence technologies such as deep learning have been applied to audio processing. Audio identification technology, one of the audio processing technologies, is developed to detect the subject from which the audio input originates and the situation in which the subject originates.

[0003] In order to realize an audio identification technology using artificial intelligence, a large number of audio inputs and corresponding pre-identified audio information, or audio analysis, are essential elements.

[0004] Meanwhile, as video platforms such as YouTube continue to develop, efforts are being made to improve audio quality using audio analysis technology. Content uploaded to video platforms is generally recorded using low-performance audio equipment, so the need for improved audio quality is gradually increasing. Summary of the Invention [Problem to be solved by the invention]

[0005] The technical object of the present invention is to provide an audio quality conversion device and a control method thereof that can convert recorded audio data using a pre-trained artificial intelligence model.

[0006] The technical object of the present invention is to provide an audio sound quality conversion device and a control method thereof that can convert audio data recorded in a first environment into audio data recorded in a second environment.

[0007] The technical objective of the present invention is to provide an artificial intelligence model that performs audio sound quality conversion so that audio data having a quality similar to that of high-performance recording equipment can be output even when using low-performance recording equipment.

[0008] The technical problem of the present invention is to propose a method for training an artificial intelligence model for performing audio conversion. [Means for solving the problem]

[0009] In order to solve the above problem, the audio quality conversion device of the present invention includes a control unit having an artificial neural network that performs learning using a plurality of audio data recorded in different recording environments for a predetermined audio event and environmental data related to the recording environments corresponding to each audio data, and an audio input unit that receives external sounds and generates audio recording data, and the control unit converts the audio recording data generated by the audio input unit based on the learning result of the artificial neural network. [Effects of the Invention]

[0010] According to the present invention, the sound quality of recorded audio data can be converted to suit various environments without being limited by the performance of recording equipment. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram showing components of an audio conversion device 100 according to an embodiment of the present invention. [Figure 2]FIG. 1 is a conceptual diagram illustrating an artificial neural network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0012] The objects and advantages of the present invention will become more apparent from the following detailed description, but the objects and advantages of the present invention are not limited to the following description alone. Furthermore, in the description of the present invention, if it is determined that a detailed description of a known technology related to the present invention may unnecessarily obscure the gist of the present invention, the detailed description will be omitted.

[0013] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily carry out the present invention. However, the present invention may be embodied in various different forms and is not limited to the embodiments disclosed below. Furthermore, in order to clearly disclose the present invention in the drawings, parts unrelated to the present invention are omitted, and the same or similar reference numerals in the drawings indicate the same or similar components.

[0014] In the following description, "one end" refers to the left side of FIG. 2, and "the other end" refers to the opposite side of "one end" and the right side of FIG.

[0015] The preferred embodiments of the present invention have been disclosed for illustrative purposes, and those skilled in the art may make various modifications, changes, and additions within the spirit and scope of the present invention, and such modifications, changes, and additions should be considered to fall within the scope of the claims. Furthermore, since those skilled in the art to which the present invention pertains may make various substitutions, modifications, and changes within the scope of the technical spirit of the present invention, the present invention is not limited by the above-described embodiments and the accompanying drawings.

[0016] In the exemplary systems described above, the methods are described as a series of steps or blocks and based on flowcharts, but the present invention is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and that other steps may be included or one or more steps of the flowcharts may be omitted without affecting the scope of the present invention.

[0017] In FIG. 1, components of an audio quality conversion device according to the present invention are explained.

[0018] As shown in FIG. 1, the audio quality converting device 100 may include an input unit 110, an output unit 120, a memory 130, a communication unit 140, a control unit 180, and a power supply unit 190.

[0019] More specifically, the communication unit 140 among the components may include one or more modules that enable wireless communication between the audio quality conversion device 100 and a wireless communication system, between the audio quality conversion device 100 and another audio quality conversion device 100, or between the audio quality conversion device 100 and an external server. The communication unit 140 may also include one or more modules that connect the audio quality conversion device 100 to one or more networks.

[0020] The input unit 110 may include a camera or video input unit for inputting a video signal, a microphone or audio input unit 111 for inputting an audio signal, and a user input unit (e.g., touch keys, mechanical keys, etc.) for receiving information input from a user. The voice data and image data collected by the input unit 110 may be analyzed and processed as a user control command.

[0021] The output unit 120 generates an output related to vision, hearing, or touch, and may include at least one of a display unit, an audio output unit, a haptic module, and an optical output unit. The display unit may be formed as an interlayer structure with a touch sensor or may be integrally formed with the touch sensor to implement a touch screen. Such a touch screen may function as a user input device that provides an input interface between the audio quality conversion device 100 and a user, and may also provide an output interface between the audio quality conversion device 100 and a user.

[0022] The memory 130 stores data supporting various functions of the audio sound quality conversion device 100. The memory 130 can store a number of application programs (or applications) run by the audio sound quality conversion device 100, as well as data and commands for the operation of the audio sound quality conversion device 100. At least some of these application programs can be downloaded from an external server via wireless communication. In addition, at least some of these application programs can be present in the audio sound quality conversion device 100 from the time of delivery for the basic functions of the audio sound quality conversion device 100 (e.g., incoming call, outgoing call, message receiving, outgoing call functions). Meanwhile, the application programs can be stored in the memory 130, installed in the audio sound quality conversion device 100, and driven by the control unit 180 to perform the operations (or functions) of the electronic device control device.

[0023] In addition to the operations related to the applications, the control unit 180 typically controls the overall operation of the audio quality conversion device 100. The control unit 180 processes signals, data, information, etc. input or output through the components detailed above, or runs application programs stored in the memory 130, thereby providing or processing appropriate information or functions to the user.

[0024] 1 to operate the application program stored in the memory 130. Furthermore, the control unit 180 can control at least some of the components detailed in Fig. 1 to operate the application program stored in the memory 130. Furthermore, the control unit 180 can operate at least two or more components included in the audio quality converting device 100 in combination with each other to operate the application program.

[0025] The power supply unit 190 receives an external power source and an internal power source under the control of the control unit 180 and supplies power to each component included in the audio quality conversion device 100. The power supply unit 190 includes a battery, which may be a built-in battery or a replaceable battery.

[0026] At least some of the components may cooperate with each other to implement the operation, control, or control method of the electronic device controller according to various embodiments described below. In addition, the operation, control, or control method of the electronic device controller may be implemented on the electronic device controller by running at least one application program stored in the memory 130.

[0027] For example, the audio quality converting device 100 may be implemented in the form of a separate terminal, such as a desktop computer or a digital TV, or a mobile terminal, such as a mobile phone, a notebook computer, a PDA, a tablet PC, a notebook computer, or a wearable device.

[0028] The following describes the training data used for learning the artificial neural network installed in the audio quality conversion device 100 according to the present invention.

[0029] In the following, "audio data" is defined as already recorded training data, which may have multiple tags.

[0030] For example, the tag may include environmental data related to the recording environment, such as the distance between the microphone and the source where the sound originates, the noise level of the recording location, and the spatial reverberation of the recording location.

[0031] The distance tag indicating the distance between the microphone and the sound source may be configured as a specific numerical value, and may be classified as a close distance, a medium distance, or a far distance.

[0032] A noise tag indicating the noise level of a recording location can be defined as a signal-to-noise ratio (SNR).

[0033] A spatial reverberation tag that indicates the spatial reverberation of a recording location can be defined as a Reverberation Time (RT) of 60 dB, where RT 60 dB means the time it takes for the measured sound pressure level to decrease by 60 dB after the sound source disappears.

[0034] The method for obtaining training data with the above tags is as follows.

[0035] For example, the training data set may consist of first audio data captured by a recording device with certain specifications in a location where the noise level is below a predetermined reference value, and second audio data in which noise data is added to the first audio data, which has the advantage that the training data set can be captured using only one recording device.

[0036] Alternatively, the training data sets may be acquired using different recording devices, which may improve the accuracy of the audio data included in the training data sets.

[0037] In this case, the recording device may be substantially the same device as the audio quality converting device 100 .

[0038] Meanwhile, each time a training data set is acquired, the recording device can assign a distance tag, a noise tag, and a spatial reverberation tag corresponding to the acquired training data set.

[0039] As an example, the recording device may generate a distance tag using an image associated with the sound source captured by a camera included in the recording device.

[0040] As another example, the distance tag may be set to a default value, in which case the recording device can control a display included in the recording device to output information related to the recording distance when collecting training data to guide the user to an appropriate distance value between the sound source and the microphone.

[0041] As described above, the artificial neural network according to the present invention can perform training using audio data having multiple tags.

[0042] In one embodiment, the control unit 180 may be equipped with an artificial neural network that performs learning using multiple audio data recorded in different recording environments for a given audio event and environmental data related to the recording environments corresponding to each audio data.

[0043] In this case, the environmental data may include at least one of the distance tag, noise tag, and spatial reverberation tag described above, i.e., the environmental data may correspond to parameters of a training data set applied to an artificial neural network.

[0044] For example, the artificial neural network described above may be trained using a training data set including first audio data corresponding to first environmental data and second audio data corresponding to second environmental data.

[0045] Meanwhile, the first audio data and the second audio data may be recordings of substantially the same audio event, i.e., the artificial neural network may perform training using the results of recording the same audio event in different recording environments in order to analyze differences in audio characteristics that occur when the same audio event is recorded in different recording environments.

[0046] In another embodiment, in order to reduce the cost required for training, a method of generating second audio data by adding additional noise to the first audio data may be considered. The second environment data may be set based on the added noise information. In this case, it is preferable to diversify the added noise to improve the training results of the artificial neural network.

[0047] That is, the artificial neural network can perform learning using first audio data recorded under environmental conditions where the noise level is below a predetermined value and second audio data obtained by combining the first audio data with pre-stored noise data.

[0048] Also, a first noise tag corresponding to the first audio data and a second noise tag corresponding to the second audio data may be set to different values.

[0049] In another embodiment, the first distance tag corresponding to the first audio data and the second distance tag corresponding to the second audio data may be set to different values. Any one training data set may include the first audio data and the second audio data recorded at locations spaced apart from a sound source by different distances.

[0050] Similarly, a first spatial reverberation tag corresponding to the first audio data and a second spatial reverberation tag corresponding to the second audio data may be set to different values. Any one training data set may include first audio data and second audio data recorded in spaces with different spatial reverberation values.

[0051] In addition, the audio input unit 111 can receive external sounds and generate audio recording data.

[0052] At this time, the audio recording data is defined as a different concept from the audio data described above in that it is newly acquired data to which no separate labeling or corresponding tag has been assigned.

[0053] Also, the control unit 180 can convert the audio recording data generated by the audio input unit based on the learning result of the artificial neural network.

[0054] As described above, an artificial neural network learns about a single audio event using a training data set that includes first and second audio data recorded in different recording environments.

[0055] In addition, when any newly recorded audio recording data is input, the control unit 180 can convert the sound quality characteristics of the input audio recording data using a pre-trained artificial neural network.

[0056] In one embodiment, the input unit 110 may include a conversion condition input unit (not shown) for receiving input of information related to the conversion of audio recording data.

[0057] Specifically, the control unit 180 can convert the audio recording data using an artificial neural network based on the information input to the conversion condition input unit.

[0058] At this time, the information input to the conversion condition input unit may include a variable related to environmental data, i.e., the information input to the conversion condition input unit may include at least one of a distance tag, a noise tag, and a spatial reverberation tag.

[0059] For example, if the information input to the conversion condition input unit includes a third distance tag, the control unit 180 can use a pre-trained artificial neural network to convert the newly recorded audio recording data into sound quality characteristics corresponding to the third distance tag.

[0060] As another example, if the information input to the conversion condition input unit includes a third noise tag, the control unit 180 can convert the newly recorded audio recording data into sound quality characteristics corresponding to the third noise tag using a pre-trained artificial neural network.

[0061] As another example, if the information input to the conversion condition input unit includes a third spatial reverberation tag, the control unit 180 can convert the newly recorded audio recording data into sound quality characteristics corresponding to the third spatial reverberation tag using a pre-trained artificial neural network.

[0062] On the other hand, if the environmental data corresponding to the newly recorded audio recording data can be identified, the control unit 180 can convert the sound quality characteristics of the audio recording data by taking into account the identified environmental data together with the information input to the conversion condition input unit.

[0063] For example, if the identified environmental data includes a first noise tag and the information input to the conversion condition input unit includes a third noise tag, the control unit 180 can convert the audio recording data into sound quality characteristics corresponding to the third noise tag by setting the input side variable of the artificial neural network to the first noise tag and the output side variable to the third noise tag.

[0064] That is, the conversion condition input unit can receive input of information related to criteria for changing the audio recording data.

[0065] When using such a trained artificial neural network, audio recording data recorded in a first recording environment can be converted to sound like it was recorded in a second recording environment. That is, an audio conversion device using an artificial neural network trained using audio data including the above-mentioned tags can convert the sound quality characteristics of audio recording data to sound like it was recorded in a recording environment desired by the user, regardless of the actual recording environment of the audio recording data.

[0066] Specifically, the environmental data may include at least one of information related to the distance between the location where the audio event occurs and a microphone that records the audio data, information related to the spatial reverberation at the location where the microphone is located, and information related to the noise at the location where the microphone is located, which may correspond to a distance tag, a noise tag, and a spatial reverberation tag, respectively.

[0067] For example, the control unit 180 may determine the noise-related information included in the environmental data by calculating a signal-to-noise ratio (SNR) of the audio data.

[0068] As another example, a camera of the audio conversion device may capture an image of the location where an audio event occurs, in which case the control unit 180 may use the image generated by the camera to calculate the distance between the location where the audio event occurs and the microphone.

[0069] As another example, the control unit 180 can determine information related to spatial reverberation included in the environmental data by measuring the reverberation time of the audio data.

[0070] Furthermore, the artificial neural network can perform learning using information related to the difference between the first audio data corresponding to the first environmental data and the second audio data corresponding to the second environmental data.

[0071] As another example, an artificial neural network may be trained using environmental data corresponding to a first training data set and environmental data corresponding to a second training data set.

[0072] Meanwhile, the audio input unit 111 may include a first microphone and a second microphone that are spaced a predetermined distance apart on the main body of the audio quality conversion device. In this case, the artificial neural network may perform learning using the first audio data acquired from the first microphone and the audio data acquired from the second microphone. In this way, when multiple microphones are installed at separate locations, audio data having different distance tag values ​​may be acquired when one audio event is recorded.

[0073] In the following, a microphone capability tag is defined as a new type of tag included in the aforementioned environmental data. The microphone capability tag can include information related to the capabilities of a microphone for recording audio data.

[0074] In one embodiment, the audio input unit may include a first microphone and a second microphone having different recording capabilities, in which case the artificial neural network may perform training using first audio data acquired from the first microphone and second audio data acquired from the second microphone.

[0075] As described above, the conversion condition input unit may receive input of information related to the microphone performance tag.

[0076] In addition, the control unit 180 may be equipped with an artificial neural network that performs learning using multiple audio data recorded from microphones with different performance for a specific audio event and microphone performance tags related to the performance of the microphones corresponding to each audio data.

[0077] Also, the control unit 180 can convert the audio recording data generated by the audio input unit using the trained artificial neural network as described above.

[0078] That is, the control unit 180 can convert the audio recording data acquired by the first microphone using the artificial neural network so that the audio recording data has sound quality characteristics corresponding to that acquired by the second microphone.

[0079] Specifically, the artificial neural network can be trained using information related to the difference between first audio data recorded by a first microphone and second audio data recorded by a second microphone.

[0080] On the other hand, it is preferable to set the performance of the first microphone and the second microphone to be meaningfully different.

[0081] In another embodiment, the artificial neural network can be trained using information related to the difference between the sound quality characteristics of first audio data recorded by a first microphone for an audio event and the sound quality characteristics of second audio data recorded by a second microphone having different performance than the first microphone.

[0082] In another embodiment, the artificial neural network may perform training using information related to the difference in sound quality between first and second audio data recorded in different recording environments and identified with the same label, and first environmental data corresponding to the first audio data and second environmental data corresponding to the second audio data.

[0083] The label may be preset by the user, such as "baby voice" or "siren sound."

[0084] In another embodiment, the artificial neural network may be trained using information related to at least one of the following: a label for an audio event, sound quality characteristics of audio data recording the audio event, performance characteristics of a microphone that captured the audio data, and the recording environment of the audio data.

[0085] Meanwhile, the control unit 180 may be equipped with a pre-trained artificial intelligence engine for identifying labels corresponding to audio data, i.e., the control unit 180 may use the artificial intelligence engine to identify the labels of audio recording data generated by the audio input unit.

Claims

1. An audio quality conversion device, a control unit including an artificial neural network that performs learning using a plurality of audio data recorded in different recording environments for a predetermined audio event and environmental data related to the recording environments corresponding to each of the audio data; and an audio input unit that receives external sound and generates audio recording data; The environmental data is the information includes at least one of information related to the distance between the location where the audio event occurs and a microphone that records the audio data, information related to spatial reverberation at a location where the microphone is located, and information related to noise at a location where the microphone is located; The control unit converting the audio recording data generated by the audio input unit based on the learning results of the artificial neural network; The audio input unit The audio quality converting device includes a first microphone and a second microphone that are spaced apart from each other by a predetermined distance and installed on the main body of the audio quality converting device; The artificial neural network performing learning using audio data acquired from the first microphone and audio data acquired from the second microphone; The artificial neural network An audio quality conversion device that performs learning using first audio data recorded under environmental conditions where the noise level is equal to or lower than a predetermined value and second audio data obtained by combining the first audio data with noise data previously stored.

2. a conversion condition input unit for receiving information related to the conversion of the audio recording data; The control unit 2. The audio quality conversion device according to claim 1, wherein the audio recording data is converted using the artificial neural network based on information input to the conversion condition input unit.

3. The artificial neural network 2. The audio quality conversion device according to claim 1, wherein the learning is performed using environmental data corresponding to the first training data and environmental data corresponding to the second training data.

Citation Information

Patent Citations

  • Data driven audio enhancement

    US20190392852A1

  • Electronic apparatus and controlling method thereof

    US20210021953A1

  • Method and system for processing sound characteristics based on deep learning

    WO2019233358A1

  • Improving audio quality of speech in sound systems

    WO2021070032A1