Voice processing unit, voice processing method and program

The audio processing apparatus addresses the labor-intensive task of volume adjustment for numerous audio files by utilizing a volume table to automate the process, thereby reducing man-hours and enhancing efficiency.

JP2025087467AActive Publication Date: 2025-06-10AZSTOKE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023202150
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-06-10
Estimated Expiration
2043-11-29

AI Technical Summary

Technical Problem

Adjusting the volume of a large number of audio files, such as those used in games, is labor-intensive and requires significant man-hours when done manually.

Method used

An audio processing apparatus that acquires audio files, searches a volume table for matching file names, and adjusts the volume of the audio files based on the corresponding volume values in the table.

Benefits of technology

This approach significantly reduces the labor required for volume adjustment of multiple audio files by automating the process through the use of a volume table, allowing for efficient handling of large numbers of files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087467000001_ABST
    Figure 2025087467000001_ABST
Patent Text Reader

Abstract

To provide a technique advantageous in lessening the burden of sound volume adjusting operation for a plurality of voice files.SOLUTION: A voice processing device has: acquisition means which acquires voice files; search means which searches a sound volume table including a pair of a character string and a sound volume value as one record for a record having a character string, partially matching the file name of an acquired voice file, as a registered character string; and adjustment means which adjusts the sound volume of the voice recorded in the voice file with a sound volume value described in the record obtained through the search.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio processing apparatus, an audio processing method, and a program.

Background Art

[0002] In an application that handles a plurality of audio files, in many cases, it is desirable that the volume of each file be adjusted to a specified volume. For example, in a game, if the volume of the operation sound (e.g., walking sound) of the same character varies greatly depending on the scene, it may give the user a sense of discomfort. Therefore, developers spend a great deal of effort in adjusting the volume of a plurality of audio files installed in the game.

[0003] Conventionally, volume adjustment for a plurality of audio files has been performed, for example, in the following procedure. (a) A plurality of delivered audio files are stored in a storage device. (b) Compare the reference audio file with one audio file selected from the plurality of audio files by listening. (c) Adjust the signal level of the audio file so that the perceived volume is the same. (d) Repeat (b) and (c) for the unprocessed audio files among the plurality of audio files.

[0004] Note that the adjustment of the signal level performed in step (c) is not limited to changing the audio data itself. For example, Patent Document 1 describes storing an automatic volume adjustment element in association with audio data and adjusting the volume using the automatic volume adjustment element when the audio data is played back. Patent Document 2 describes adding a playback control identifier related to the playback volume to the file name of a music file and adjusting the volume using the playback control identifier when the music file is played back.

Prior Art Documents

Patent Documents

[0005] Patent Document 1 Japanese Patent Application Laid-Open No. 2003-243952 Patent Document 2 Japanese Patent Application Laid-Open No. 2011-197664 Summary of the Invention Problems to be Solved by the Invention

[0006] However, for example, the number of audio files used in games may exceed tens of thousands. If the volume of such a large number of audio files is adjusted one by one, the man-hours required for the work will be enormous. Therefore, it is desired to reduce the labor involved in volume adjustment work for a plurality of audio files. Means for Solving the Problems

[0007] According to one aspect of the present invention, there is provided an audio processing apparatus comprising: acquisition means for acquiring an audio file; search means for searching a volume table including a pair of a character string and a volume value as one record for a record having a character string that partially matches the file name of the acquired audio file as a registered character string; and adjustment means for adjusting the volume of the audio recorded in the audio file according to the volume value described in the record obtained by the search. Effects of the Invention

[0008] According to the present invention, it is possible to provide a technique advantageous for reducing the labor involved in volume adjustment work for a plurality of audio files. Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiments for Carrying Out the Invention

[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. It should be noted that the following embodiments do not limit the invention according to the claims, and not all combinations of the features described in the embodiments are essential for the invention. Two or more of the features described in the embodiments may be arbitrarily combined. Also, the same or similar configurations are given the same reference numerals, and duplicate explanations are omitted.

[0011] FIG. 1 shows a block diagram showing the configuration of an audio processing apparatus C according to an embodiment. The audio processing apparatus C is an apparatus that displays an audio signal recorded in a file and performs various processes such as adjusting the signal level on the audio signal. In this specification, the term "audio" should be understood in a broad sense. "Audio" may include not only the voices of humans and animals but also musical sounds, computer-generated sound effects, etc. That is, in this specification, the term "audio" is intended to include "speech", "sound", and "audio (acoustics)".

[0012] The voice processing device C can be a computer device such as a personal computer or a workstation. The voice processing device C includes a CPU (Central Processing Unit) 101 that controls the entire device, a RAM 102 that functions as a main storage device and provides a work area for the CPU 101, and a ROM 103 that stores fixed data and programs. The voice processing device C also includes an audio interface (I / F) 104. A microphone M and a speaker S can be connected to the audio interface 104. A storage device (secondary storage device) 110 (storage unit) is connected to the voice processing device C via an interface (I / F) 105. The storage device 110 can be, for example, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof. Note that the storage device 110 may be configured inside the voice processing device C or outside it. The network interface 106 connects to the network N and communicates. The voice processing device C can be communicably connected to the server A via the network N, for example.

[0013] An input device K such as a keyboard and a mouse can be connected to the voice processing device C via an interface 107. An external media device F such as a CD-ROM drive and a DVD drive can be connected to the voice processing device C via an interface 108. Furthermore, the voice processing device C includes a video controller 109. The video controller 109 controls image display by a display device (display) D. A touch panel screen in which the input device K and the display D are integrated may be configured.

[0014] The boot program for starting the audio processing device C is stored in the ROM 103. Also, as shown in FIG. 1, the storage device 110 can install an operating system (OS) 111, signal processing programs 112 for performing audio signal processing, and one or more audio files 113. The audio file 113 may be supplied from an external device such as the server A via the network N, or may be supplied from a medium stored in the external media device F. Alternatively, the audio file 113 may be created from the sound picked up by the microphone M. Further, the storage device 110 also stores a loudness table 114 which will be described later.

[0015] The audio file 113 is an audio file in which audio content is recorded. In one example, the file format of the audio file 113 can be the WAVE file format commonly used in personal computers. The WAVE file can include a header and data of an audio signal. The header can include information such as the mono / stereo type, sampling frequency, quantization bit number, etc. Note that the file format of the audio file 113 is not limited to the WAVE file format. The file format of the audio file 113 may be a format other than the WAVE file format, for example, formats such as AIFF, MP3, AAC, etc.

[0016] The configuration of the audio processing device C in this embodiment is generally as described above. As an example, consider that this audio processing device C is used for game development. The number of audio files implemented in a game may reach tens of thousands or more. Since there are variations in the volume of the initial multiple audio files delivered, it is necessary to adjust the volume (adjustment of the signal level) for each audio file. However, if the volume of such a large number of audio files is adjusted one by one, the man-hours required will be enormous.

[0017] The sounds used in a game can include a variety of sounds such as character dialogue sounds, situation description (success, failure, etc.) sounds, sound effects, footstep sounds, explosion sounds, environmental sounds, BGM, and so on. The inventor of the present invention noticed that there is a relationship between the content of such sounds and appropriate volume values. In the present embodiment, the volume value is determined according to the content of the sound in the sound file.

[0018] In the field of game development, generally, each sound file is named so that the attributes of the sound can be understood to some extent. "Attributes" refer to things that can identify the content of the sound, such as character names, scene names, action names, the content of dialogue, and so on. The file name may include multiple attribute information, for example, in the form of "character name + action name". In game development, naming rules for sound files are usually determined so that they are not significantly changed during development. Therefore, it is possible to identify the content of the sound from the file name of the sound file and determine the volume value according to the identified content of the sound.

[0019] In this embodiment, when adjusting the volume of the voice recorded in each voice file, a volume table in which target volume values are described is referred to. Here, the volume value will be described. In this embodiment, as a measure (index) of the volume value, a loudness value considering human auditory characteristics is used. The loudness value is represented in units of, for example, LUFS (Loudness Units Full Scale) or LKFS (Loudness K-Weighted Full Scale). Therefore, in this embodiment, when performing loudness adjustment on the voice recorded in each voice file, the loudness table 114 in which target loudness values are described is referred to as the volume table. The loudness table 114 is a lookup table in which the correspondence between a character string that can be part of the file name of the voice file and the loudness value (target loudness value) which is the volume value is described. The loudness table may be called a "loudness list". FIG. 2 shows an example of the structure of the loudness table 114. The loudness table 114 includes a pair of a character string (registered character string) and a loudness value (target loudness value) as one record. The registered character string described in each record is a character string that can be part of the file name of the voice file. Note that the present invention is not limited to using the loudness value as the measure of the volume value. A measure other than the loudness value (for example, RMS) may be used as the measure of the volume value.

[0020] FIG. 3 shows an example of a setting screen 30 displayed on the display D. The CPU 101 as a display control unit displays each record of the loudness table on the setting screen 30 on the display D in an editable manner. The user can add and register records in the loudness table 114 via this setting screen 30. The number of records registered in the loudness table 114 is displayed in the record number display window 31. A record can be added in response to the addition button 32 being pressed (clicked by a mouse, tapped via a touch panel). The content of each registered record is displayed in the list 35. Each record in the list 35 has columns for "Search" and "Value". In the "Search" column, the registered string to be searched is displayed, and in the "Value" column, the loudness value corresponding to the registered string is displayed. When all records cannot be displayed within the display area of the list 35, it can be scrolled using the scroll bar 36.

[0021] The user can specify the loudness measurement method in the loudness setting column 33. Examples of the loudness measurement method include MaxMomentary, MaxShort-Term, and Integrated. Any one of these can be selected in the loudness setting column 33. MaxMomentary means that loudness calculation is performed for each of a plurality of measurement windows (400 msec in length) obtained by sliding the audio waveform on the time axis for a predetermined time, and the maximum value among them is adopted as the loudness value. MaxShort-Term means that loudness calculation is performed for each of a plurality of measurement windows (3 sec in length) obtained by sliding on the time axis for a predetermined time, and the maximum value among them is adopted as the loudness value. Integrated means that the loudness of the entire sound source (the entire sound of one audio file) is measured. In the example of FIG. 3, MaxMomentary is selected. Further, it may be possible to specify an arbitrary measurement window length instead of the specific measurement window lengths described above.

[0022] Before loudness adjustment is performed on the audio of an audio file, as an option, dynamic range compression may be performed. There may be a large variation in the playback volume between audio files. If the volume of the sound source is not adjusted as it is, the playback volume of a certain audio may be too low or too high, making it difficult to hear. Therefore, it is necessary to align the signal levels of each sound source. Dynamic range compression is performed to align the signal levels between such audios to be constant. Dynamic range compression generally includes a process of suppressing the part including the peak of the signal level and increasing the part with a low signal level. However, it is not simply necessary to make the signal level constant. In the case of human speech, if there is not enough intonation, the feeling of being compressed becomes stronger. Therefore, in dynamic range compression, it is necessary to appropriately set the signal level threshold for determining the compression target.

[0023] Dynamic range compression can also be manually performed by the user (manual compression) by moving any of the adjustment points arranged on the envelope among a plurality of adjustment points. However, it requires a great deal of effort to perform manual compression on all audios. Therefore, it is also possible to automatically perform dynamic range compression on the entire audio file. Automatically performing dynamic range compression is referred to here as "auto compression".

[0024] Automatic compression may include, for example, the following processing. The audio signal of the target audio file is composed of a plurality of frames. First, obtain the envelope of the audio signal. Next, detect the peak value of the envelope for each frame, and calculate the average value (the first average value) of the peak values for each detected frame. Next, detect peak values higher than the first average value, and calculate their average values (the second average value). Then, adjust the envelope so that at least a part of the peak values higher than the second average value is suppressed. For example, further detect peak values higher than the second average value, and calculate their average values (the third average value). Further, detect peak values higher than the third average value, and adjust them so that they approach the third average value. Note that such a processing method of automatic compression is only an example, and it may be realized by other processing methods.

[0025] In this embodiment, the user can specify whether to apply automatic compression to all the audio files stored in the working folder of the storage device 110. The setting screen 30 is provided with an automatic compression setting column 34 for instructing the execution of automatic compression. For example, radio buttons or check boxes are prepared in the automatic compression setting column 34, and the execution of automatic compression is specified by setting it to the selected state (ON). In the example of FIG. 3, the automatic compression setting column 34 is set to ON by radio buttons. In this case, after the dynamic range compression of the audio of the audio file is executed, loudness adjustment is performed.

[0026] The setting screen 30 further has a file name display column 37 for displaying the file names of one or more audio files to be subjected to loudness adjustment and stored in the working folder of the storage device 110.

[0027] FIG. 4 shows a flowchart of the audio processing method in the audio processing device C. The program corresponding to this flowchart is included in the signal processing program 112 and is executed by the CPU 101.

[0028] In step S11, the CPU 101 acquires one or more voice files 113 and stores them in a predetermined working folder of the storage device 110. In one example, the voice file 113 can be acquired from an external device such as server A via the network N. Alternatively, the voice file 113 may be acquired from a medium stored in the external media device F. Alternatively, the voice file 113 may be acquired by being created from the sound picked up by the microphone M.

[0029] In step S12, the CPU 101 acquires one voice file (target voice file) to be processed from the one or more voice files 113 stored in the working folder of the storage device 110 and loads it into the RAM 102. In step S13, the CPU 101 executes auto-completion on the target voice file. However, this step S13 is an option when the auto-completion setting field 34 in the setting screen 30 shown in FIG. 3 is in a selected state. When the auto-completion setting field 34 is not in a selected state, step S13 is skipped.

[0030] In step S14, the CPU 101 searches the loudness table 114 for a record having a character string that partially matches the file name of the target voice file as the registered character string.

[0031] In step S15, the CPU 101 performs loudness adjustment on the voice recorded in the target voice file according to the loudness value (target loudness value) described in the record obtained by the search in step S15. The loudness adjustment is performed, for example, by measuring the loudness value of the voice of the target voice file (when step S13 is executed, the voice of the target voice file after auto-completion is executed) according to the loudness measurement method specified in the loudness setting field 33, and based on the measurement result, adjusting the gain value of the voice so that the loudness value becomes the target loudness value.

[0032] In step S16, the CPU 101 as the display control unit causes the first waveform (the waveform before loudness adjustment), which is the voice of the audio file acquired in step S12 or the waveform of the voice with automatic comp applied in step S13, and the second waveform, which is the waveform of the voice after loudness adjustment, to be displayed in the display area of the display D. An example of waveform display will be described later.

[0033] In step S17, the CPU 101 determines whether there is an unprocessed audio file stored in the working folder of the storage device 110. If there is no unprocessed file, the process ends. If there is an unprocessed audio file, the process returns to step S12, and the process is repeated for the next audio file. Therefore, according to the present embodiment, when a plurality of audio files are acquired in step S11 and stored in the working folder of the storage device 110, for each of the plurality of audio files, the search in step S14 and the loudness adjustment in step S15 are sequentially performed.

[0034] FIG. 5 shows an example of the waveform display in step S16. Here, an example of waveform display when three audio files are processed is shown. The waveforms to be displayed are time-domain waveforms. Therefore, the horizontal axis of the waveform is the time axis, and the vertical axis indicates the signal level. In FIG. 5, above the display area, the waveform W11 after automatic comp (before loudness adjustment) of the voice of the first audio file, the waveform W12 after automatic comp (before loudness adjustment) of the voice of the second audio file, and the waveform W13 after automatic comp (before loudness adjustment) of the voice of the third audio file are arranged side by side along the time axis direction. A plurality of adjustment points P discretely arranged on the envelope obtained in automatic comp for adjusting the signal level may be displayed on each of the waveforms W11, W12, and W13. The user can manually adjust the signal level at that position, for example, by dragging an arbitrary adjustment point with the mouse.

[0035] In FIG. 5, at the lower part of the display area, the waveform W21 after loudness adjustment of the audio of the first audio file, the waveform W22 after loudness adjustment of the audio of the second audio file, and the waveform W23 after loudness adjustment of the audio of the third audio file are arranged side by side along the time axis direction. Each waveform after loudness adjustment is obtained by newly writing out the audio whose loudness has been adjusted in step S15 to a file. Note that the above-described waveform display mode is merely an example, and other display modes may be adopted.

[0036] (Other examples) As shown in FIG. 2, a plurality of records in the loudness table 114 are grouped according to the commonality of the prefixes of the registered strings. The prefix can be a string representing the attribute of the audio, which is defined by the naming rule. In that case, the fact that the prefixes are common means that the attributes of the audio are common. For example, the prefix "vo" represents the voice of a character, and the prefix "atk" represents the battle cry at the time of an attack (attack), etc. In the example of FIG. 2, the plurality of records are classified into group 1 with the prefix "vo_", group 2 with the prefix "vo_atk", group 3 with the prefix "vo_dmg", group 4 with the prefix "vo_move", and group 5 with the prefix "vo_cmm".

[0037] In step S14, the CPU 101 searches the loudness table 114 for a prefix that partially matches the file name of the target audio file to identify the group, and searches for a registered string that partially matches the file name from among the identified groups. In the example of FIG. 2, each group includes a representative record in which a pair of a string consisting only of the prefix and a loudness value is described. The representative record exists in the first row of each group.

[0038] In step S14, if no registered string that partially matches the file name is found among the identified groups other than the representative record as a result of the search, in step S15, loudness adjustment is performed based on the loudness value described in the representative record. Hereinafter, a specific example will be described. In step S14, first, a prefix that partially matches the file name of the target audio file is searched for from the representative record existing in the first line of each group. For example, assume that the prefix "vo_atk" partially matches the file name of the target audio file. In this case, the group to be searched is limited to group 2. Then, a registered string that partially matches the file name is searched for from group 2. In group 2, in addition to the representative record, records with registered strings such as "vo_atk_charge" and "vo_atk_s" are included. However, if no registered string that partially matches the file name is found among the records in group 2 other than the representative record, in step S15, loudness adjustment is performed based on the loudness value "-21" corresponding to the representative record (registered string "vo_atk").

[0039] According to the above processing, the search range of the loudness table can be limited, so the search speed is improved.

[0040] In the example of FIG. 2, group 1 with "vo_" as the prefix is positioned as the upper layer of other groups with "vo_" and other subsequent strings as the prefix. When the prefix that partially matches the file name of the target audio file is only "vo_", the loudness value is "-23" corresponding to "vo_" in group 1.

[0041] According to the embodiment described above, a search is performed for a record having a character string that partially matches the file name of an audio file as a registered character string from a volume table (loudness table) that includes a pair of a character string and a volume value (loudness value) as one record. Then, the volume of the audio recorded in the audio file is adjusted according to the volume value described in the record obtained by the search. If the volume table is created in advance, there is no need to separately perform settings for volume adjustment. Also, when processing a plurality of audio files, the above search and volume adjustment are sequentially performed for each audio file. In this way, volume adjustment is automatically performed for a plurality of audio files. Also, the number of records included in the volume table can be significantly less than the number of a plurality of audio files. Therefore, according to the present embodiment, compared with the prior art in which the audio of each of a plurality of audio files is adjusted one by one, the user's workload is significantly reduced.

[0042] Note that it is not essential that the loudness table 114 is stored in the storage device 110. For example, the loudness table 114 may be stored in an external device (for example, the server A in FIG. 1) connected via the network N, and the audio processing device C may refer to the loudness table 114 stored in the external device via the network N.

[0043] The present invention can also be implemented by causing a computer to execute a program for causing the computer to execute each step of the audio processing method described in the above embodiment.

[0044] The invention is not limited to the above embodiment, and various modifications and changes are possible within the scope of the gist of the invention.

Explanation of Signs

[0045] A: Server, C: Audio processing device, D: Display, K: Input device, 101: CPU, 112: Signal processing program, 114: Loudness table

Claims

1. An acquisition means for acquiring an audio file; A search means for searching a volume table that includes pairs of a character string and a volume value as one record for a record having a character string that partially matches the file name of the acquired audio file as a registered character string; An adjustment means for adjusting the volume of the audio recorded in the audio file according to the volume value described in the record obtained by the search; An audio processing apparatus characterized by comprising the above.

2. After a plurality of audio files are acquired by the acquisition means, the search and the volume adjustment are sequentially performed for each of the plurality of audio files. The audio processing apparatus according to claim 1, characterized in that.

3. The adjustment means performs dynamic range compression on the audio of each of the plurality of audio files, Based on the volume value recorded in the record obtained by the search, the volume value of the audio on which the dynamic range compression has been performed is adjusted. The audio processing apparatus according to claim 2, characterized in that.

4. The audio processing apparatus according to claim 3, further comprising a display control unit that displays on a display the waveform of the audio before adjustment by the adjustment means and the waveform of the audio after adjustment by the adjustment means.

5. The display control unit further displays each record of the volume table on the display so as to be editable. The audio processing apparatus according to claim 4, characterized in that.

6. A plurality of records in the volume table are grouped according to the commonality of the prefixes of the registered character strings, The search means searches the volume table for a prefix that partially matches the file name to identify a group, and searches for a registered character string that partially matches the file name from among the identified groups. The audio processing apparatus according to claim 1, characterized in that.

7. Each group includes a representative record in which a pair of a character string consisting only of the prefix and a volume value is described, As a result of the search by the search means, when no registered character string that partially matches the file name is found in the identified group other than the representative record, the adjustment means performs the volume adjustment according to the volume value described in the representative record. The audio processing apparatus according to claim 6, characterized in that.

8. ​ The voice processing apparatus according to claim 1, wherein the scale of the volume value is a loudness value. **Claim 9** A step in which an acquisition means acquires an audio file; A step in which a search means searches a volume table including a pair of a character string and a volume value as one record for a record having a character string that partially matches the file name of the acquired audio file as a registered character string; A step in which an adjustment means adjusts the volume of the audio of the audio file according to the volume value described in the record obtained by the search; A voice processing method characterized by comprising: **Claim 10** A program for causing a computer to function as each means of the voice processing apparatus according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Digital audio system, auto volume control factor generating method, auto volume control method, auto volume control factor generating program, auto volume control program, recording medium for recording the auto volume control factor generating program, and recording medium for recording the auto volume control program

    JP2003243952A

  • Music file reproduction device and system

    JP2011197664A