Sound processing device, voice processing method, and program
The integration of middleware and DAW in the audio processing apparatus enables automatic volume adjustment of multiple audio files by considering the volume adjustments made on the middleware, addressing the inefficiencies of manual adjustment and improving the overall efficiency and accuracy of the process.
Patent Information
- Application Number
- JP2023202151
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-11-29
AI Technical Summary
The existing methods for adjusting the volume of a large number of audio files, such as those used in games, are labor-intensive and inefficient, as they require manual adjustment of each file, and there is no way to track or consider the volume adjustments made by middleware in digital audio workstations (DAWs).
An audio processing apparatus and method that integrates middleware and a DAW, allowing the processor to acquire the change amount of a volume value based on routing information, enabling automatic volume adjustment of multiple audio files while considering the volume adjustments made on the middleware.
This solution significantly reduces the man-hours required for volume adjustment by automating the process and allows for volume adjustments to be performed in consideration of the volume adjustment history on the middleware, improving efficiency and accuracy.
Smart Images

Figure 2025087468000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an audio processing apparatus, an audio processing method, and a program.
Background Art
[0002] In applications that handle a plurality of audio files, in many cases, it is desirable that the volume of each file is adjusted to a specified volume. For example, in a game, if the volume of the movement sound (e.g., walking sound) of the same character varies greatly depending on the scene, it may give the user a sense of discomfort. Therefore, developers spend a great deal of effort in adjusting the volume of a plurality of audio files installed in the game.
[0003] Conventionally, volume adjustment for a plurality of audio files has been performed, for example, according to the following procedure. (a) A plurality of delivered audio files are stored in a storage device. (b) Compare the reference audio file with one audio file selected from the plurality of audio files by listening. (c) Adjust the signal level of the audio file so that the perceived volume is the same. (d) Repeat (b) and (c) for the unprocessed audio files among the plurality of audio files.
[0004] Note that the adjustment of the signal level performed in step (c) is not limited to changing the audio data itself. For example, Patent Document 1 describes storing an automatic volume adjustment element in association with audio data and adjusting the volume using the automatic volume adjustment element when the audio data is reproduced. Patent Document 2 describes adding a reproduction control identifier related to the reproduction volume to the file name of a music file and adjusting the volume using the reproduction control identifier when the music file is reproduced.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, for example, the number of audio files used in games may exceed tens of thousands. If the volume of such a large number of audio files is adjusted one by one, the man-hours required for the work will be enormous. Therefore, it is desired to reduce the labor required for volume adjustment by automating the volume adjustment for a plurality of audio files. In game development, two software programs, middleware (audio middleware) and a digital audio workstation (DAW), are used for the production and adjustment of audio files. However, on the DAW, it is impossible to grasp how the volume of each of a plurality of audio files has been adjusted by the middleware, and it has been impossible to perform volume adjustment taking into account the volume adjustment results on the middleware. The present invention provides a technique advantageous for automatic volume adjustment for a plurality of audio files.
Means for Solving the Problems
[0007] According to one aspect of the present invention, there is provided an audio processing apparatus for processing audio, including: a storage unit that stores middleware, which is software for performing audio processing, and a digital audio workstation (DAW), which is software different from the middleware for performing audio processing; and a processor that executes the middleware and the DAW, wherein the processor acquires a change amount of a volume value of the audio on the middleware based on routing information indicating a path until the audio recorded in the audio file set on the middleware reaches an output on the DAW.
Effects of the Invention
[0008] According to the present invention, it is possible to provide an advantageous technique for automatic volume adjustment for a plurality of audio files.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims, and not all combinations of the features described in the embodiments are essential to the invention. Two or more of the plurality of features described in the embodiments may be arbitrarily combined. Also, the same or similar configurations are assigned the same reference numerals, and duplicate explanations are omitted.
[0011] FIG. 1 shows a block diagram illustrating the configuration of an audio processing apparatus C according to an embodiment. The audio processing apparatus C is an apparatus that displays an audio signal recorded in a file and performs various processes such as adjusting the signal level on the audio signal. In this specification, the term "audio" should be understood in a broad sense. "Audio" may include not only the voices of humans and animals but also musical sounds, computer-generated sound effects, and the like. That is, in this specification, the term "audio" is intended to include "speech", "sound", and "audio (acoustics)".
[0012] The audio processing apparatus C can be a computer device such as a personal computer or a workstation. The audio processing apparatus C includes a CPU (Central Processing Unit) 101 that controls the entire apparatus, a RAM 102 that functions as a main storage device and provides a work area for the CPU 101, and a ROM 103 that stores fixed data and programs. The audio processing apparatus C also includes an audio interface (I / F) 104. A microphone M and a speaker S can be connected to the audio interface 104. A storage device (secondary storage device) 110 (storage unit) is connected to the audio processing apparatus C via an interface (I / F) 105. The storage device 110 can be, for example, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof. Note that the storage device 110 may be configured inside or outside the audio processing apparatus C. The network interface 106 is connected to the network N to perform communication. The audio processing apparatus C can be communicably connected to the server A via the network N, for example.
[0013] An input device K such as a keyboard and a mouse can be connected to the audio processing device C via the interface 107. Also, an external media device F such as a CD-ROM drive and a DVD drive can be connected to the audio processing device C via the interface 108. Furthermore, the audio processing device C includes a video controller 109. The video controller 109 controls image display by a display device (display) D. A touch panel screen in which the input device K and the display D are integrated may be configured.
[0014] The boot program for starting the audio processing device C is stored in the ROM 103. Also, as shown in FIG. 1, an operating system (OS) 111 and one or more audio files 113 can be installed in the storage device 110. The audio file 113 may be supplied from an external device such as the server A via the network N, or may be supplied from a medium stored in the external media device F. Alternatively, the audio file 113 may be created from sound picked up by the microphone M. Also, a loudness table 114 described later is stored in the storage device 110.
[0015] The audio file 113 is an audio file in which audio content is recorded. In one example, the file format of the audio file 113 can be the WAVE file format commonly used in personal computers. The WAVE file can include a header and data of an audio signal. The header can include information such as the mono / stereo type, sampling frequency, and quantization bit number. Note that the file format of the audio file 113 is not limited to the WAVE file format. The file format of the audio file 113 may be a format other than the WAVE file format, for example, formats such as AIFF, MP3, and AAC.
[0016] As an example, consider the case where the audio processing device C is used in game development. In implementing audio in game development, roughly speaking, a sound creator creates an audio file, and a programmer programs the game engine so that the created audio is played in accordance with user operations. In creating an audio file, with the growth and complexity of game development, two major tools have come to be used. One is a DAW for creating diverse audio files, and the other is middleware (audio middleware) for streamlining the effort of integrating audio into the game engine. DAW is an abbreviation for Digital Audio Workstation, and is software that enables recording / editing of audio for the purpose of audio production. Middleware is software that performs playback, processing, and management of audio passed to the game engine, and can create audio data to be played by the DAW. As such middleware, for example, there is Wwise manufactured by Audiokinetic. Therefore, middleware 115 and DAW 112 are also installed in the storage device 110 of the audio processing device C. The CPU 101 can function as a processor that executes the middleware and the DAW. The DAW and the middleware are configured to be able to perform cooperative processing such as performing transfer processing of an audio file between the two. For example, processing such as creating audio in the DAW, writing out the audio file, moving the audio file from the DAW to the middleware, and implementing the audio file in the game engine by the middleware can be performed. Also, when adjustment of the audio transferred from the DAW to the middleware is necessary, the audio file can be moved again from the middleware to the DAW, and the audio file can be adjusted in the DAW.
[0017] The number of audio files implemented in a game may reach tens of thousands or more. Since there are variations in the volume of a plurality of initially delivered audio files, it is necessary to perform volume adjustment (adjustment of the signal level) for each audio file. However, if the volume of such a large number of audio files is adjusted one by one, the man-hours required for the work will be enormous.
[0018] The sounds used in the game can include a variety of sounds such as character dialogue sounds, situation description (success, failure, etc.) sounds, sound effects, footstep sounds, explosion sounds, environmental sounds, BGM, etc. The inventor noticed that there is a relationship between the content of such sounds and appropriate volume values. In this embodiment, the volume value is determined according to the content of the sound in the sound file.
[0019] In the field of game development, generally, each sound file is named so that the attributes of the sound can be understood to a certain extent. "Attributes" refer to things that can identify the content of the sound, such as character names, scene names, action names, the content of the dialogue, etc. The file name may include multiple attribute information, for example, in the form of "character name + action name". In game development, the naming rules of sound files are determined so that they are not significantly changed during development. Therefore, it is possible to identify the content of the sound from the file name of the sound file and determine the volume value according to the identified content of the sound.
[0020] In this embodiment, when adjusting the volume of the audio recorded in each audio file on the DAW side, a volume table in which target volume values are described is referred to. Here, the volume value will be described. In this embodiment, as a measure (index) of the volume value, a loudness value considering human auditory characteristics is used. The loudness value is represented in units of, for example, LUFS (Loudness Units Full Scale) or LKFS (Loudness K-Weighted Full Scale). Therefore, in this embodiment, when performing loudness adjustment of the audio recorded in each audio file on the DAW side, the loudness table 114 in which target loudness values are described is referred to as the volume table. The loudness table 114 is a lookup table in which the correspondence between a character string that can be part of the file name of the audio file and a loudness value (target loudness value) which is the volume value is described. The loudness table may be called a "loudness list". FIG. 2 shows an example of the structure of the loudness table 114. The loudness table 114 includes a pair of a character string (registered character string) and a loudness value (target loudness value) as one record. The registered character string described in each record is a character string that can be part of the file name of the audio file. Note that the present invention is not limited to using the loudness value as the measure of the volume value. Other measures (for example, RMS) may be used as the measure of the volume value.
[0021] FIG. 3 shows an example of a loudness value setting screen 30 that is displayed on the display D when the DAW 112 is executed by the CPU 101. The CPU 101 as the display control unit displays each record of the loudness table on the setting screen 30 on the display D in an editable manner. The user can add and register records to the loudness table 114 via this setting screen 30. The number of records registered in the loudness table 114 is displayed in the record number display window 31. A record can be added in response to the addition button 32 being pressed (clicked by the mouse, tapped via the touch panel). The content of each registered record is displayed in the list 35. Each record in the list 35 has columns for "Search" and "Value". In the "Search" column, the registered string to be searched is displayed, and in the "Value" column, the loudness value corresponding to the registered string is displayed. If all the records cannot be displayed within the display area of the list 35, it can be scrolled using the scroll bar 36.
[0022] The user can specify the loudness measurement method in the loudness setting column 33. Examples of the loudness measurement method include MaxMomentary, MaxShort - Term, and Integrated. In the loudness setting column 33, any one of these can be selected. MaxMomentary means that loudness calculation is performed for each of a plurality of measurement windows (400 msec in length) obtained by sliding the audio waveform on the time axis for a predetermined time, and the maximum value among them is adopted as the loudness value. MaxShort - Term means that loudness calculation is performed for each of a plurality of measurement windows (3 sec in length) obtained by sliding on the time axis for a predetermined time, and the maximum value among them is adopted as the loudness value. Integrated means that the loudness of the entire sound source (the entire audio of one audio file) is measured. In the example of FIG. 3, MaxMomentary is selected. Further, it may be possible to specify an arbitrary measurement window length instead of the specific measurement window lengths described above.
[0023] Before loudness adjustment is performed on the audio of an audio file, as an option, dynamic range compression may be performed. There may be a large variation in the playback volume between audio files. If the volume of the sound source is not adjusted as it is, the playback volume of a certain audio may be too low or too high, making it difficult to hear. Therefore, it is necessary to align the signal levels of each sound source. Dynamic range compression is performed to align the signal levels between such audios to be constant. Dynamic range compression generally includes a process of suppressing the part including the peak of the signal level and increasing the part with a low signal level. However, it is not simply necessary to make the signal level constant. In the case of human speech, if there is no intonation to a certain extent, the feeling of being compressed becomes stronger. Therefore, in dynamic range compression, it is necessary to appropriately set the signal level threshold for determining the compression target.
[0024] Dynamic range compression can also be manually performed by the user by moving any adjustment point among a plurality of adjustment points arranged on the envelope (manual comp). However, it requires a great deal of labor to perform manual comp for all audios. Therefore, it is also possible to automatically perform dynamic range compression on the entire audio file. Automatically performing dynamic range compression is referred to here as "auto comp".
[0025] Automatic compression may include, for example, the following processing. The audio signal of the target audio file is composed of a plurality of frames. First, an envelope of the audio signal is obtained. Next, the peak value of the envelope for each frame is detected, and the average value (first average value) of the detected peak values for each frame is calculated. Next, peak values higher than the first average value are detected, and their average value (second average value) is calculated. Then, the envelope is adjusted so that at least a part of the peak values higher than the second average value is suppressed. For example, peak values higher than the second average value are further detected, and their average value (third average value) is calculated. Further, peak values higher than the third average value are detected and adjusted so that they approach the third average value. Note that such a method for automatic compression processing is only an example, and it may be realized by other processing methods.
[0026] In this embodiment, the user can specify whether to apply automatic compression to all the audio files stored in the working folder of the storage device 110. The setting screen 30 is provided with an automatic compression setting column 34 for instructing the execution of automatic compression. For example, radio buttons or check boxes are provided in the automatic compression setting column 34, and selecting it to the selected state (ON) specifies the execution of automatic compression. In the example of FIG. 3, the automatic compression setting column 34 is set to ON by radio buttons. In this case, after the dynamic range compression of the audio of the audio file is executed, loudness adjustment is performed.
[0027] The setting screen 30 further has a file name display column 37 that displays the file names of one or more audio files that are the targets of loudness adjustment and are stored in the working folder of the storage device 110.
[0028] Next, the management of audio files on middleware 115 will be described. Each individual audio file may contain one or more audio materials (recorded sound portions). One audio material is also called a "track". In the middleware, the audio (tracks) are classified in a hierarchical structure. FIG. 4 is a conceptual diagram showing an example of the hierarchical structure H of the audio managed on the middleware. Tracks are input to the input terminal, and all the tracks are collected into the master track MS and output from the output terminal. That is, all the audio played in the game is finally output after passing through the master track MS. Effects E and / or volume adjustment V are applied to the tracks IN1, IN2, and IN3 input to the input terminal individually. The hierarchical structure H may include bus tracks. A bus track is a track that combines one or more tracks. In FIG. 4, the bus track B1 combines the track IN1 and the track IN2 into one track and outputs it to the master track MS. By using the bus, effects and / or volume adjustment can be applied to multiple tracks collectively. Also, the hierarchical structure H may include auxiliary tracks. An auxiliary track is a copy of a certain track that is sent across. In FIG. 4, the auxiliary track AUX inputs the track IN3 and outputs it to the bus track B2. The bus track B2 inputs the track IN3 and also inputs the auxiliary track AUX. Auxiliary tracks are used, for example, when coexisting a track with effects (such as reverb and delay) and a track without effects.
[0029] The user can design a hierarchical structure through a hierarchical structure setting screen (not shown), and for any track of the audio files registered on the middleware, determine which input terminal in the hierarchical structure to place it on. As a result, for each track, a routing indicating the path from the input to the output in the hierarchical structure is determined. In this way, a volume adjustment unit is provided for each layer of the hierarchical structure H, and the track can be volume-adjusted each time it passes through each volume adjustment unit. The user can determine the effects to be applied to each layer of the designed hierarchical structure and the volume values of the volume adjustment units. The total change amount of the volume value of the audio on the middleware is calculated by summing the volume values at each volume adjustment unit on the path. For example, as shown in FIG. 4, assume that the loudness value on the master track is -6 dB, the loudness value on the bus track B1 is +4 dB, the loudness value on the track IN1 is +2 dB, and the loudness value on the track IN2 is set to -4 dB. In this case, the total change amount of the loudness value of the track IN1 on the middleware is (-6 dB)+(+4 dB)+(+2 dB)=0 dB That is. Also, the total change amount of the loudness value of the track IN2 on the middleware is (-6 dB)+(+4 dB)+(-4 dB)=-6 dB That is.
[0030] Through the design of the hierarchical structure on the middleware as described above, routing information indicating the path for the audio recorded in the audio file to reach the output is created. The routing information may include information on the path (hereinafter referred to as "routing") for the audio recorded in the audio file to reach the output and information on the volume values at each volume adjustment unit on the path. Note that in the above description, an example of one type of hierarchical structure H is shown, but it may be configured such that multiple types of hierarchical structures can be constructed on the middleware.
[0031] As described above, the loudness adjustment on the DAW is performed using the loudness value determined according to the file name of the audio file with reference to the loudness table. However, conventionally, on the DAW side, it has not been possible to grasp the volume adjustment history of each audio on the middleware. Therefore, on the DAW side, the loudness adjustment has been performed without considering the volume adjustment history on the middleware.
[0032] In the present embodiment, the CPU 101 acquires, on the DAW, the amount of change in the volume value in the routing of the audio on the middleware based on the routing information created by the middleware. Then, the CPU 101 performs volume adjustment on the DAW based on the acquired amount of change. FIG. 5 shows an example of a setting screen 40 for acquiring the amount of change in the volume value on the middleware, which is displayed on the display D when the DAW 112 is executed by the CPU 101. In the setting screen 40, the search check column 41 is a column for setting the connection of the middleware. As described above, there may be cases where a plurality of types of hierarchical structures are constructed on the middleware. The search path setting column 44 is a column for setting which hierarchical structure to be the search target when a plurality of types of hierarchical structures are constructed on the middleware. The detection path setting column 45 is a column for setting the hierarchy to be prioritized when a plurality of audio files with the same file name are found in the search path. The excluded routing setting column 47 is a column for setting the routing (excluded routing) to be excluded from the calculation target when obtaining the total amount of change in the volume value on the middleware, and the user can specify the information for specifying the excluded routing in the routing specification column 48. The addition routing setting field 49 is a field for setting routing (addition routing) in a hierarchical structure different from the hierarchical structure set in the search path setting field 44, which should be added to the calculation targets when obtaining the total change amount of the volume value on the middleware. The user can specify information for specifying the addition routing in the routing specification field 50. In addition to volume adjustment within the routing, there can be multiple exceptions, such as performing volume adjustment using effectors such as upmixing (e.g., 2.0ch → 4.0ch), downmixing (e.g., 4.0ch → 2.0ch), gain, limiter, and comp. The addition routing setting field 49 is prepared for such exceptional handling. When the routing specification fields 48 and 50 increase and it is not possible to display all of them in the display area, the scroll bar 51 can be used to scroll.
[0033] Figure 6 shows a flowchart of the audio processing method in the audio processing device C. The program corresponding to this flowchart is included in the DAW 112 and is performed by the CPU 101 during the execution of the DAW 112.
[0034] In step S11, the CPU 101 acquires the audio files registered in the middleware and stores them in a predetermined working folder of the storage device 110.
[0035] In step S12, the CPU 101 executes automatic comp on the acquired audio files. However, this step S12 is an option when the automatic comp setting field 34 on the setting screen 30 shown in Figure 3 is in the selected state. When the automatic comp setting field 34 is not in the selected state, step S12 is skipped.
[0036] In step S13, the CPU 101 searches the loudness table 114 for a record having a string that partially matches the file name of the acquired audio file as a registered string. The CPU 101 acquires the loudness value R described in the record obtained by this search.
[0037] In step S14, the CPU 101 determines whether there is an audio file below the search path set in the search path setting field 44 (that is, in at least any one routing in the hierarchical structure to be searched). If there is an audio file, the process proceeds to step S15. In step S15, the CPU 101 determines whether there are multiple audio files below the search path. If there are multiple audio files, the process proceeds to step S16, and if there is only one audio file, the process proceeds to step S18.
[0038] In step S16, the CPU 101 determines whether there is an audio file in the priority level set in the detection path setting field 45. If there is an audio file in the priority level set in the detection path setting field 45, the process proceeds to step S18, and if not, the process proceeds to step S17. In step S17, the routing information of the path of the audio file searched first among the multiple audio files identified in step S15 is acquired.
[0039] In step S18, the CPU 101 acquires all the routing information on the middleware including the routing information acquired in step S17. As described above, the routing information may include the information of the path until the audio recorded in the audio file set in the middleware reaches the output, and the information of the volume value at each volume adjustment unit on the path.
[0040] In step S19, the CPU 101 excludes the non-target routing set in the non-target routing setting field 47 from the calculation targets when obtaining the total change amount of the loudness value in step S21 described later.
[0041] There is only one routing that should be added to the calculation target when obtaining the total change amount of the volume value on the middleware. For example, when applying an effect, it is possible to set multiple routings. Fig. 4 shows an example where the routing in which Track IN3 is input to Bus B2 with an effect applied via the Auxiliary Track and the routing in which Track IN3 is input to Bus B2 in the dry state coexist. Among these two routings, the routing for which the volume value should be obtained is the routing in which Track IN3 is input to Bus B2 in the dry state. Therefore, in this case, the routing in which Track IN3 is input to Bus B2 via the Auxiliary Track should be set as an excluded routing. However, it is possible that the user may forget to set such a routing as an excluded routing. Therefore, in step S20, the CPU 101 determines whether the number of routings to be searched is one. If the number of routings to be searched is not one, the process proceeds to step S26, where the CPU 101 outputs an error and ends the process. If the number of routings to be searched is one, the process proceeds to step S21. In step S21, the CPU 101 calculates the total change amount T of the loudness value on the middleware. The total change amount T is obtained by adding the change amount of the loudness value in the routing where the audio file exists among the search paths set in the search path setting field 44 to the change amount of the loudness value in the routing set in the addition routing setting field 49. However, if no audio file exists below the search path set in the search path setting field 44 in step S14, the change amount of the loudness value of the search path is set to 0 in step S25.
[0042] In step S22, the CPU 101 calculates the final loudness value FR (final volume value). The final loudness value FR is obtained by calculating the difference between the loudness value R obtained from the loudness table in step S13 and the total change amount T of the loudness value obtained in step S21.
[0043] In step S23, the CPU 101 adjusts the loudness of the audio recorded in the target audio file according to the final loudness value FR calculated in step S22. The loudness adjustment is performed, for example, by measuring the loudness value of the audio in the target audio file (when step S12 is executed, the audio in the target audio file after auto-completion) according to the loudness measurement method specified in the loudness setting field 33, and based on the measurement result, adjusting the gain value of the audio so that the loudness value becomes the target loudness value.
[0044] In step S24, the CPU 101 as the display control unit causes the first waveform (the waveform before loudness adjustment), which is the waveform of the audio of the audio file acquired in step S11 or the audio to which auto-completion was applied in step S12, and the second waveform, which is the waveform of the audio after loudness adjustment, to be displayed in the display area of the display D. An example of waveform display will be described later.
[0045] If there are other unprocessed audio files registered in the middleware, the process returns to step S11, and the process is repeated for the next audio file. Therefore, if there are a plurality of audio files registered in the middleware, for each of the plurality of audio files, the search in step S13 to the loudness adjustment in step S23 is sequentially performed.
[0046] FIG. 7 shows an example of the waveform display in step S24. Here, an example of the waveform display when three audio files are processed is shown. The displayed waveform is a time-domain waveform. Therefore, the horizontal axis of the waveform is the time axis, and the vertical axis indicates the signal level. In FIG. 7, above the display area, the waveform W11 after automatic comp (before loudness adjustment) of the audio of the first audio file, the waveform W12 after automatic comp (before loudness adjustment) of the audio of the second audio file, and the waveform W13 after automatic comp (before loudness adjustment) of the audio of the third audio file are arranged side by side along the time axis direction. A plurality of adjustment points P discretely arranged on the envelope obtained in the automatic comp to adjust the signal level may be displayed on each of the waveforms W11, W12, and W13. The user can manually adjust the signal level at that position, for example, by dragging an arbitrary adjustment point with the mouse.
[0047] In FIG. 7, below the display area, the waveform W21 after loudness adjustment of the audio of the first audio file, the waveform W22 after loudness adjustment of the audio of the second audio file, and the waveform W23 after loudness adjustment of the audio of the third audio file are arranged side by side along the time axis direction. Each waveform after loudness adjustment is obtained by newly writing the audio for which loudness adjustment was performed in step S23 to a file.
[0048] Also, information on the loudness value of the audio file that is the target of the waveform display may be displayed here. For example, the file name of the audio file of each waveform, the total change amount T of the loudness value on the middleware, the loudness value R obtained by searching the loudness table, and the final loudness value FR are displayed. Thereby, the user can grasp how much the volume of each audio file has been adjusted in the middleware and the DAW. Also, according to the present embodiment, on the DAW, volume adjustment can be performed in consideration of the volume adjustment history in the middleware. Note that the above-described display modes of the waveform and the loudness value information are merely examples, and other display modes may be adopted.
[0049] (Other examples) As shown in FIG. 2, a plurality of records in the loudness table 114 are grouped by the commonality of the prefixes of the registered strings. The prefix can be a string representing the attribute of the voice, which is defined by the naming rule. In that case, the fact that the prefixes are common means that the attributes of the voices are common. For example, the prefix "vo" represents the voice of a character, and the prefix "atk" represents the battle cry during an attack, etc. In the example of FIG. 2, the plurality of records are classified into group 1 with the prefix "vo_", group 2 with the prefix "vo_atk", group 3 with the prefix "vo_dmg", group 4 with the prefix "vo_move", and group 5 with the prefix "vo_cmm".
[0050] In step S13, the CPU 101 searches the loudness table 114 for a prefix that partially matches the file name of the target audio file to identify the group, and then searches for a registered string that partially matches the file name from within the identified group. In the example of FIG. 2, each group includes a representative record in which a pair of a string consisting only of the prefix and the loudness value is described. The representative record exists in the first row of each group.
[0051] In step S13, if, as a result of the search, no registered string that partially matches the file name is found among the identified groups other than the representative record, then in step S23, loudness adjustment is performed based on the loudness value described in the representative record. A specific example will be described below. In step S13, first, a prefix that partially matches the file name of the target audio file is searched for from the representative records existing in the first line of each group. For example, assume that the prefix "vo_atk" partially matches the file name of the target audio file. In this case, the group to be searched is limited to group 2. Then, among group 2, a registered string that partially matches the file name is searched for. In group 2, in addition to the representative record, there are records with registered strings such as "vo_atk_charge" and "vo_atk_s", but if no registered string that partially matches the file name is found among group 2 other than the representative record, then in step S23, loudness adjustment is performed based on the loudness value "-21" corresponding to the representative record (registered string "vo_atk").
[0052] According to the above processing, since the search range of the loudness table can be limited, the search speed is improved.
[0053] Note that in the example of FIG. 2, group 1 with the prefix "vo_" is positioned as the upper layer of other groups with "vo_" and other subsequent strings as prefixes. When the prefix that partially matches the file name of the target audio file is only "vo_", the loudness value is "-23" corresponding to "vo_" in group 1.
[0054] According to the embodiment described above, a search is performed for a record having a character string that partially matches the file name of an audio file as a registered character string from a volume table (loudness table) that includes a pair of a character string and a volume value (loudness value) as one record. Then, the volume of the audio recorded in the audio file is adjusted according to the volume value described in the record obtained by the search. If the volume table has been created in advance, there is no need to separately perform settings for volume adjustment. Also, when processing a plurality of audio files, the above search and volume adjustment are sequentially performed for each audio file. In this way, volume adjustment is automatically performed for a plurality of audio files. Also, the number of records included in the volume table can be significantly less than the number of a plurality of audio files. Therefore, according to the present embodiment, compared with the conventional technique in which the audio of each of a plurality of audio files is adjusted one by one, the work man-hours of the user are significantly reduced. Furthermore, according to the present embodiment, as described above, the user can grasp how much volume adjustment has been performed on each audio file in the middleware and the DAW. Also, according to the present embodiment, on the DAW, volume adjustment can be performed in consideration of the volume adjustment history in the middleware.
[0055] Note that it is not essential that the loudness table 114 be stored in the storage device 110. For example, the loudness table 114 may be stored in an external device (for example, server A in FIG. 1) connected via the network N, and the audio processing device C may refer to the loudness table 114 stored in the external device via the network N.
[0056] The present invention can also be implemented by causing a computer to execute a program for causing the computer to execute each step of the audio processing method described in the above embodiment.
[0057] The invention is not limited to the above embodiments, and various modifications and changes are possible within the scope of the gist of the invention.
Explanation of Signs
[0058] A: Server, C: Audio processing device, D: Display, K: Input device, 101: CPU, 112: DAW, 114: Loudness table, 115: Middleware
Claims
1. An audio processing apparatus for processing audio, a storage unit storing middleware which is software for performing audio processing, and a digital audio workstation (DAW) which is software different from the middleware for performing audio processing, a processor for executing the middleware and the DAW, and having, on the DAW, the processor obtains a change amount of a volume value of the audio on the middleware based on routing information indicating a path until the audio recorded in the audio file set on the middleware reaches an output, characterized in that it is an audio processing apparatus.
2. On the DAW, the processor further performs volume adjustment of the audio based on the obtained change amount, characterized in that it is the audio processing apparatus according to Claim 1.
3. In the middleware, the audio is classified in a hierarchical structure, and a volume adjustment unit is provided for each layer, the routing information includes information on the path of the audio and information on volume values at each volume adjustment unit on the path, on the DAW, the processor sums up volume values at each volume adjustment unit on the path to calculate a total change amount of the volume value of the audio on the middleware, and performs volume adjustment of the audio based on the calculated total change amount, characterized in that it is the audio processing apparatus according to Claim 2.
4. On the DAW, the processor searches for a record having a character string that partially matches the file name of the audio file as a registered character string from a volume table including a pair of a character string and a volume value as one record, and performs volume adjustment of the audio recorded in the audio file with a final volume value which is a difference between the volume value described in the record obtained by the search and the total change amount, characterized in that it is the audio processing apparatus according to Claim 3.
5. characterized in that it further has setting means for setting an excluded routing which is a routing to be excluded from a calculation target when obtaining the total change amount, according to Claim 4.
6. characterized in that it further has setting means for setting an additional routing which is a routing in a hierarchical structure different from the hierarchical structure to be added to a calculation target when obtaining the total change amount, according to Claim 4.
7. The voice processing apparatus according to claim 1, wherein the scale of the volume value is a loudness value.
8. A voice processing method executed by a voice processing apparatus having a storage unit that stores middleware, which is software for performing voice processing, and a digital audio workstation (DAW), which is software different from the middleware for performing voice processing, and a processor that executes the middleware and the DAW, wherein, during execution of the DAW by the processor, a step of obtaining routing information indicating a path until the voice recorded in the voice file set in the middleware reaches the output; a step of obtaining a change amount of the volume value of the voice on the middleware; A voice processing method characterized by comprising:
9. A program for causing a computer to execute each step of the voice processing method according to claim 8.
Citation Information
Patent Citations
Digital audio system, auto volume control factor generating method, auto volume control method, auto volume control factor generating program, auto volume control program, recording medium for recording the auto volume control factor generating program, and recording medium for recording the auto volume control program
JP2003243952A
Music file reproduction device and system
JP2011197664A