Spot production data generation device and program
The spot production data generation device addresses the challenge of managing large audio file volumes by creating a single file with all replacement patterns, enabling dynamic audio adaptation based on playback time and environment, thus reducing management burden and enhancing flexibility.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2026-04-15
AI Technical Summary
The management of large volumes of audio files for spot announcements becomes burdensome due to the need for multiple replacements, and conventional object-based audio production systems are limited in flexibility and require pre-editing, preventing dynamic audio replacement during playback.
A spot production data generation device that combines audio data and generates acoustic metadata, allowing for a single file to contain all replacement patterns, enabling dynamic audio replacement based on playback time and environment.
Reduces data volume and management burden by generating a single file with all replacement patterns, facilitating flexible audio playback on standard renderers without the need for pre-generated files.
Smart Images

Figure 0007846559000001 
Figure 0007846559000002 
Figure 0007846559000003
Abstract
Description
Technical Field
[0001] The present invention relates to a spot control data generation device and a program thereof.
Background Art
[0002] Currently, in television and radio broadcasts and video and audio distribution services, spot announcements may be inserted as short videos or audio. A spot announcement is a short video or audio of about 15 to 30 seconds inserted during or between programs for the purpose of advertising products of program sponsors, promoting other programs of the broadcasting station, etc. Hereinafter, spot announcements are simply referred to as spots.
[0003] Some spots are further divided into multiple parts within that short time. In addition to the main promotional part, there is a part assuming that some voices are replaced according to the area or time where the spot is broadcast, such as "(broadcast by ○○ Broadcasting Station / △△ Broadcasting Station)", "○○[program name] is (tonight / tomorrow)". When producing a spot divided into such multiple parts, although only some voices are different, it is necessary to produce an audio file for the entire spot for all replacement patterns. In that case, a large number of completed audio files and video files incorporating the audio files are generated, resulting in a management burden to cope with the increase in data volume and prevention of file misplacement. In addition, when spots produced by the main broadcasting stations or key stations in large cities such as Tokyo are provided to regional broadcasting stations or affiliated broadcasting stations, the production and editing work of the replacement part is carried out at each regional broadcasting station.
[0004] Furthermore, in recent years, object-based audio systems have been increasingly adopted, primarily in the film industry, as a method for reproducing 3D sound. Object-based audio systems record and transmit audio objects and audio metadata that constitute the object-based sound, and a playback device called a renderer plays (renders) the content in a format appropriate to the playback environment. Since rendering is performed in the playback environment of each home, it is possible to replace the audio objects during rendering, enabling services such as dubbing from English to Japanese.
[0005] Currently, the introduction of object-based audio into broadcasting or video streaming services is being considered using audio encoding schemes such as MPEG-H and AC-4, which support object-based audio. Furthermore, the International Telecommunication Union Radiocommunication Sector (ITU-R) has defined an Audio Definition Model (ADM) as an international standard for audio metadata used in program production (see Non-Patent Document 1). The ITU-R also defines a standard renderer for ADMs that generates and makes program audio available for listening based on the ADM described in content production (see Non-Patent Document 2). [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] "Audio Definition Model", Recommendation ITU-R BS.2076-02 (10 / 2019) [Non-Patent Document 2] "Audio Definition Model renderer for advanced sound systems", Recommendation ITU-R BS.2127-0 (06 / 2019) [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] As mentioned above, when replacing some of the audio in a spot, the number of audio files created by the combination of replacements becomes enormous, increasing the size of the data to be managed and potentially leading to management problems such as misplacing files. Object-based audio is a technology that allows for audio replacement. However, conventional production equipment using object-based audio is designed for playback in movie theaters, and therefore, it is generally necessary to replace the audio beforehand using audio editing equipment, as in conventional methods. Furthermore, the few conventional production facilities that support audio replacement during playback are only designed for dubbing narration, limiting the types of replacements that can be performed. In addition, object-based audio content produced with conventional production equipment cannot have parts of the audio in the spot replaced on the playback device according to the playback time, etc.
[0008] The present invention has been made in view of these problems, and aims to provide a spot production data generation device and program therefor that generates spot production data that reduces the data size and number of audio files, and that allows a playback device to replace some of the audio in a spot according to the playback time. [Means for solving the problem]
[0009] To solve the aforementioned problems, the spot production data generation device according to the present invention is a spot production data generation device that generates data for spot production, and comprises an audio data combining unit, an acoustic metadata generation unit, and a data integration unit.
[0010] In this configuration, the spot production data generation device uses an audio data merging unit to input audio data consisting of one or more channels corresponding to each predetermined time interval of the spot, and generates combined audio data that is merged in the time direction. As a result, the audio data merging unit generates a single combined audio data that includes replaceable audio data within the time interval (part) and is merged in the time direction of the time interval.
[0011] Furthermore, the spot production data generation device uses an acoustic metadata generation unit to input a list of audio data channels for each time interval and generates acoustic metadata that identifies the playback position of each audio data channel in the combined audio data. As a result, the acoustic metadata generation unit generates acoustic metadata that allows a playback device to identify and play the target audio data channel from the combined audio data.
[0012] The spot production data generation device then uses a data integration unit to integrate the combined audio data and acoustic metadata to generate data for spot production. In this way, the data integration unit integrates the combined audio data and acoustic metadata into a single file. Furthermore, the spot production data generation device can be operated using a program that causes the computer to function as each of the aforementioned components. [Effects of the Invention]
[0013] According to the present invention, spot production data can be generated in a single file, allowing the playback device to replace some of the spot's audio according to the playback time. Furthermore, the present invention eliminates the need to generate multiple spots pre-generated for all possible time interval combinations. As a result, the present invention reduces the file management burden and generates spot production data with reduced data volume. [Brief explanation of the drawing]
[0014] [Figure 1] This is a block diagram showing the configuration of a data generation device for spot production according to an embodiment of the present invention. [Figure 2] This is an explanatory diagram illustrating the audio data and list input to a spot production data generation device according to an embodiment of the present invention. [Figure 3] This is an explanatory diagram illustrating a method for combining audio data by adjusting the number of channels in the audio data merging section. [Figure 4] It is an explanatory diagram for explaining the content of the object list generated by the object list generation unit of the acoustic metadata generation unit. [Figure 5] It is an explanatory diagram for explaining the content of the combination pattern list generated by the combination pattern generation unit of the acoustic metadata generation unit. [Figure 6] It is a flowchart showing the operation of the spot control data generation device according to an embodiment of the present invention. [Figure 7A] It is a diagram showing an example (1 / 2) of the acoustic metadata (ADM) generated by the metadata generation unit of the acoustic metadata generation unit. [Figure 7B] It is a diagram showing an example (2 / 2) of the acoustic metadata (ADM) generated by the metadata generation unit of the acoustic metadata generation unit. [Embodiments for Carrying Out the Invention]
[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings. [Configuration of Spot Control Data Generation Device] Referring to FIG. 1, the configuration of the spot control data generation device according to an embodiment of the present invention will be described.
[0016] The spot control data generation device 1 generates data for spot control. Here, the spot control data generation device 1 generates spot control data that includes, without duplication, all replacement patterns for producing spot audio from a plurality of audio data consisting of single or multiple channels and a list explaining the content of each channel of the audio data.
[0017] The audio data is audio data for each part (time interval) of the replacement unit of the spot, and there are a plurality (n: n is 2 or more). Of course, not all audio data needs to be a replacement target, and some parts may be fixed audio data. Here, n is the number of parts constituting the spot audio. Each channel of the audio data includes not only channels adapted to an audio format such as the left and right audio channels in stereo, but also a plurality of channels corresponding to the content such as the spot's performers and the uttered comments. The list is text data corresponding to one piece of audio data and including information for specifying the channels of the audio data and information for explaining the audio content of the channels.
[0018] Here, referring to FIG. 2, an example of audio data and a list will be described. In the following description, the case where the spot is composed of two parts (n = 2) will be described as an example. As shown in FIG. 2, the audio data A1 constituting the first half part is data with a time length of 10 seconds composed of four channels (ch1 to ch4). The list L1 corresponding to the audio data A1 indicates that the audio data A1 is a channel that can be replaced by different performers in this part. The list L1 indicates that channels 1-2 (ch1, ch2) of the audio data A1 are stereo audio where performer 1 appears. Also, the list L1 indicates that channels 3-4 (ch3, ch4) of the audio data A1 are stereo audio where performer 2 appears.
[0019] The audio data A2 constituting the second half part is data with a time length of 5 seconds composed of five channels (ch1 to ch5). The list L2 corresponding to the audio data A2 indicates that the audio data A2 is a channel that can be replaced by different comments in this part. The list L2 indicates that channel 1 (ch1) of the audio data A1 is monoral audio of "Broadcast next week!". Similarly, the list L2 indicates that channels 2-5 (ch2 to ch5) of the audio data A1 are monoral audio of "Broadcast the day after tomorrow!", "Broadcast tomorrow!", "Broadcast tonight!", and "Immediately after this!" respectively. Thus, the list L represents the channel configuration and content of the corresponding audio data A in text data. Returning to FIG. 1, the description will continue.
[0020] The spot production data generation device 1 includes an audio data combining unit 10, an acoustic metadata generation unit 20, and a data integration unit 30. The audio data combining unit 10 inputs audio data composed of one or more corresponding channels for each predetermined part (time interval) of the spot, and generates combined audio data combined in the time direction. The audio data combining unit 10 includes a channel number adjustment unit 11, a combining unit 12, and a combined audio data storage unit 13.
[0021] The channel number adjustment unit 11 inputs a plurality of audio data and adjusts them to the same number of channels. The channel number adjustment unit 11 inserts null data (silent data) into the audio data with the smaller number of channels between the combined audio data stored in the combined audio data storage unit 13 and the newly input audio data to equalize the number of channels. Note that for the audio data input for the first time, the channel number adjustment unit 11 outputs it directly to the combining unit 12.
[0022] Here, let the number of channels of the newly input audio data be m, and the number of channels of the audio data (combined audio data) stored in the combined audio data storage unit 13 be k. When m < k, the channel number adjustment unit 11 adds null data to the input audio data so that the number of channels becomes the same as that of the combined audio data (k). Then, the channel number adjustment unit 11 outputs the audio data with the equalized number of channels to the combining unit 12.
[0023] On the other hand, when m > k, the channel number adjustment unit 11 adds null data to the combined audio data stored in the combined audio data storage unit 13 so that the number of channels becomes the same as that of the input audio data (m). The channel adjustment unit 11 then outputs the input audio data directly to the merging unit 12. In the case where m=k, the channel adjustment unit 11 does not adjust the number of channels and outputs the input audio data directly to the merging unit 12.
[0024] The coupling unit 12 sequentially combines audio data that has been adjusted to the same number of channels by the channel number adjustment unit 11. The coupling unit 12 generates new combined audio data by arranging the audio data input from the channel number adjustment unit 11 seamlessly in the time direction with the combined audio data stored in the combined audio data storage unit 13. The coupling unit 12 then stores the generated combined audio data in the combined audio data storage unit 13. The coupling unit 12 stores the audio data input for the first time as combined audio data in the combined audio data storage unit 13. After the input of audio data is complete, the combining unit 12 outputs the combined audio data stored in the combined audio data storage unit 13, which contains all the combined audio data, to the data integration unit 30.
[0025] The processing in this audio data merging unit 10 will be schematically explained with reference to the diagram. For example, as shown in Figure 2, suppose the audio data A1 that makes up the first part consists of 4 channels (ch1 to ch4) and has a duration of 10 seconds, and the audio data A2 that makes up the second part consists of 5 channels (ch1 to ch5) and has a duration of 5 seconds.
[0026] In this case, as shown in Figure 3, the audio data merging unit 10 has 4 channels in the audio data for the first half that was input earlier, and 5 channels in the audio data for the second half of the heart that was input later. Therefore, the channel number adjustment unit 11 adds 1 channel of null data to the audio data A1 to make it 5 channels. Then, the audio data merging unit 10, using the merging unit 12, connects the audio data A1' and audio data A2 in the time direction to create a total of 15 seconds of audio data (merged audio data A C ) generates. Returning to Figure 1, we will continue our explanation of the configuration of the spot production data generation device 1.
[0027] The acoustic metadata generation unit 20 takes a list of audio data channels for each part (time interval) as input for each part and generates acoustic metadata that identifies the playback position of each audio data channel in the combined audio data. Here, the acoustic metadata generation unit 20 generates acoustic metadata from a list, using audio data with the number of channels constituting the audio as audio objects, and presetting (pre-set information) combinations of audio objects and playback times. For example, if the audio is stereo, the number of channels is "2", and if it is monaural, the number of channels is "1". The acoustic metadata generation unit 20 comprises an object list generation unit 21, a combination pattern generation unit 22, and a metadata generation unit 23.
[0028] The object list generation unit 21 takes multiple lists as input and generates an object list in which the data of the number of channels that make up the audio is assigned an ID (identifier) to each audio object. The object list includes at least the ID and the playback time of the audio object. The object list generation unit 21 assigns an ID to each audio object. In this case, when the list is first input, the object list generation unit 21 assigns IDs starting from the lowest number. Then, when subsequent lists are input, the object list generation unit 21 continues assigning IDs from where the previously assigned audio object IDs left off.
[0029] The ID should conform to the descriptive rules for the type of audio metadata that will ultimately be generated. For example, if the type of audio metadata is ADM (see Non-Patent Document 1), then in ADM, an audio object is represented by the descriptor "audioObject," and its ID is represented by AO_XXXX. XXXX is a four-digit hexadecimal number, where 0000 to 1000 are reserved numbers, and 1001 is the lowest number. Therefore, the object list generation unit 21 assigns an ID to each audio object, such as AO_1001, AO_1002, ...
[0030] Furthermore, the object list generation unit 21 records the playback time of the audio objects in the object list. In this case, when the list is first input, the object list generation unit 21 sets the playback time of the first part of the audio object to the time from 00:00:00 to the duration of that part. The duration may be obtained from the audio data merging unit 10 using the duration of the audio data corresponding to the list, or it may be pre-recorded in the list.
[0031] Furthermore, when the object list generation unit 21 receives a list for the second time, it sets the playback time to the time obtained by adding the duration of the audio data of the part to the end time of the audio data to be combined. The start time for playback of this audio object does not have to be 00:00:00; it may be set to another time, such as 10:00:00, to match the file management of broadcasters or other service providers. In that case, the object list generation unit 21 sequentially adds the duration of the audio data starting from that time to determine the playback time for each part.
[0032] The object list generation unit 21 may also add information indicating the audio content of the audio object (such as its name) to the object list, in addition to the ID and playback time of the audio object. When using ADM, this information can be used as the name of "audioObject". The object list generation unit 21 outputs the generated object list to the combination pattern generation unit 22.
[0033] Here, with reference to Figure 4, an example of an object list generated by the object list generation unit 21 will be described. The lists will be L1 and L2 as shown in Figure 2. List L1 is entered first, and the content of the first part has two variations (differences in performers). Therefore, the object list generation unit 21 assigns IDs starting from the lowest number and records AO_1001 and AO_1002 as IDs in the object list OL. In addition, the object list generation unit 21 records information indicating the content of the audio object (name N, audio format F) corresponding to each ID in the object list OL. Furthermore, since the duration of the first audio data is 10 seconds, the object list generation unit 21 records the playback time T of the audio object from 00:00:00 to 00:00:10 in the object list OL.
[0034] The next input, list L2, will have five different content variations for the second half (differences in spoken content). Therefore, the object list generation unit 21 assigns IDs from the continuation of the audio object IDs in the previous part, and records AO_1003 to AO_1007 as IDs in the object list OL. In addition, the object list generation unit 21 records information indicating the content of the audio object (name N, audio format F) corresponding to each ID in the object list OL. Furthermore, since the audio data for the second half is 5 seconds long, the object list generation unit 21 records the playback time T for the audio object as the time from 00:00:10 to 00:00:15, which is the end time of the previous part, in the object list OL. Returning to Figure 1, we will continue our explanation of the configuration of the spot production data generation device 1.
[0035] The combination pattern generation unit 22 generates a list of time-direction combination patterns (combination pattern list) of audio objects for each part of a spot, based on the object list generated by the object list generation unit 21. The combination pattern generation unit 22 generates a number of combination patterns equal to the number of audio objects in each part multiplied by the number of parts. For example, if a spot is divided into two parts, with K IDs in the first part and M IDs in the second part, the total number of possible combinations is K × M. This combination pattern directly becomes a preset (pre-configured ID, name, etc.) of audio object combinations described in the acoustic metadata.
[0036] Therefore, the combination pattern generation unit 22 assigns a preset ID corresponding to the combination. For example, if the type of audio metadata is ADM, then in ADM, the combination preset is represented by the descriptor "audioProgramme", and its ID is represented by APR_XXXX. The rules for the numbers that go into XXXX are the same as the rules for "audioObject". In other words, the combination pattern generation unit 22 assigns IDs such as APR_1001, APR_1002, ... to each combination of audio objects. Furthermore, the combination pattern generation unit 22 assigns a name to each combination of audio objects, which is a combination of the names of the audio objects.
[0037] The combination pattern generation unit 22 then generates a combination pattern list containing the ID, name, and other information for each combination pattern of audio objects. The combination pattern generation unit 22 outputs the object list and the combination pattern list to the metadata generation unit 23.
[0038] Here, with reference to Figure 5, an example of a combination pattern list generated by the combination pattern generation unit 22 will be described. The object list is as shown in Figure 4. The combination pattern list ML shown in Figure 5 is a list that associates the ID3 and name N3 of each combination with the ID3 and name N3 of all combinations of the two types of audio objects that make up the first part of the spot (ID1 and name N1) and the five types of audio objects that make up the second part (ID2 and name N2).
[0039] The combination pattern generation unit 22 generates a preset setting for sound metadata with ID3 set to APR_1001 and name N3 set to "Cast 1 × "Airing Next Week!"" by combining, for example, the audio object AO_1001 (Cast 1) included in the first part of the spot and the audio object AO_1003 ("Airing Next Week!") included in the second part. Similarly, the combination pattern generation unit 22 combines all the audio objects included in the first part and the audio objects included in the second part to generate a preset setting for sound metadata for each combination.
[0040] In this example, the name of the combined audio object (for example, "Cast 1 x 'Airing Next Week!'") was generated by simply combining the names of the original audio objects. However, the name of the combined object can be anything as long as it is identifiable among the audio objects. Returning to Figure 1, we will continue our explanation of the configuration of the spot production data generation device 1.
[0041] The metadata generation unit 23 generates acoustic metadata corresponding to the audio data combined by the audio data combining unit 10, based on the playback time of the audio object and the combination of audio objects. This acoustic metadata is data that indicates which audio channel is recorded at which playback time of the combined audio data. Here, the metadata generation unit 23 generates acoustic metadata using the IDs and playback times of audio objects identified in the object list generated by the object list generation unit 21, and the combinations of audio objects identified in the combination pattern list generated by the combination pattern generation unit 22, along with their IDs.
[0042] The metadata generation unit 23 generates acoustic metadata by integrating and formatting various information of the audio object according to the type of acoustic metadata to be generated. The acoustic metadata may be generated as XML text, or as binary data obtained by compressing XML text using a general-purpose compression method (e.g., gzip).
[0043] The type of acoustic metadata generated by the metadata generation unit 23 is not particularly limited. For example, the metadata generation unit 23 may generate acoustic metadata using ADM (Non-Patent Literature 1). Examples of the acoustic metadata of the ADM generated by the metadata generation unit 23 will be described later with reference to Figures 7A and 7B. The metadata generation unit 23 outputs the generated acoustic metadata to the data integration unit 30.
[0044] The data integration unit 30 integrates the audio data combined by the audio data merging unit 10 and the audio metadata generated by the audio metadata generation unit 20, and generates spot production data as a single file. In addition to integrating audio data and acoustic metadata to generate a single audio file, the data integration unit 30 may also generate a single video file by integrating audio data, acoustic metadata, and video data corresponding to the audio data.
[0045] To generate a single audio file, the data integration unit 30 can use a file format such as BW64 (BroadcastWave64), as shown in Reference 1 below. (Reference 1) “Long-form file format for the international exchange of audio programme materials with metadata”, Recommendation ITU-R BS.2088-1 (10 / 2019)
[0046] BW64 is an extension of WAVE and supports data sizes exceeding 4GB. Additionally, BW64 can write ADM. <axml>The chunk field is where the audio data is written. <data>In addition to chunks, there are other options available. Note that BW64 writes binary data. <bxml>Because a chunk is provided, acoustic metadata can be written even if it is generated as binary data.
[0047] Furthermore, in order for the data integration unit 30 to generate a single video file, it can use an audio encoding scheme that supports object-based audio such as MPEG-4 or AC-4, thereby recording audio data and an audio meter together with the video data in a single file. Alternatively, by using a container format that can store entire audio files that can be written to ADM, such as BW64, it is possible to record unencoded video and audio data in a single package. Examples of such container formats include MXF (Material Exchange Format) and IMF (Interoperable Mastering Format). With the configuration described above, the spot production data generation device 1 can generate a single audio or video file that includes all replacement patterns for the spot without duplication, and that can be replaced for each object according to the playback time.
[0048] [Operation of the data generation device for spot production] Next, referring to Figure 6 (and Figure 1 as appropriate for the configuration), the operation of the spot production data generation device according to an embodiment of the present invention will be described. In step S1, the spot production data generation device 1 receives audio data and a list corresponding to that audio data (see Figure 2). In step S2, the channel number adjustment unit 11 of the audio data merging unit 10 adjusts the number of channels between the audio data input in step S1 and the merged audio data stored in the merged audio data storage unit 13 to match the one with the larger number of channels (see Figure 3). Note that this step S2 process is omitted for the audio data input for the first time. In step S3, the merging unit 12 combines the audio data, whose number of channels was adjusted in step S2, in the time direction (see Figure 3). The merging unit 12 stores the combined audio data (combined audio data) in the combined audio data storage unit 13. For the audio data input initially, the merging unit 12 stores it as combined audio data in the combined audio data storage unit 13.
[0049] In step S4, the object list generation unit 21 of the acoustic metadata generation unit 20 assigns an ID (identifier) to the data of the number of channels constituting the audio as audio objects from the list input in step S1, and generates an object list that includes the playback time of the audio data (see Figure 4). Note that the order of processing in steps S2, S3 and step S4 may be reversed, or they may be processed in parallel.
[0050] In step S5, the spot production data generation device 1 determines whether all audio data and list inputs have been completed. If the input of all audio data and lists has not been completed (No in step S5), the spot production data generation device 1 returns to step S1 and performs its operation. On the other hand, if all audio data and list inputs have been completed (Yes in step S5), in step S6, the combination pattern generation unit 22 generates a list of time-direction combination patterns of audio objects for each part of the spot (combination pattern list) based on the object list generated in step S4 (see Figure 5).
[0051] In step S7, the metadata generation unit 23 generates acoustic metadata corresponding to the audio data combined by the audio data combining unit 10, based on the playback time of the audio object and the combination of audio objects. For example, the metadata generation unit 23 generates acoustic metadata using ADM (see Figures 7A and 7B below). In step S8, the data integration unit 30 integrates the combined audio data (combined audio data) obtained in step S3 by aligning the number of channels, with the acoustic metadata generated in step S7, and generates spot production data as a single file. Through the above operations, the spot production data generation device 1 can generate spot production data that includes all replacement patterns for spots without duplication and that can be replaced for each object according to the playback time.
[0052] [Example of acoustic metadata using ADM] The following describes an example of acoustic metadata generated by the acoustic metadata generation unit 20 of the spot production data generation device 1, with reference to Figures 7A and 7B.
[0053] The acoustic metadata shown in Figures 7A and 7B is an example (excerpt) of the ADM descriptors included in the lists of Figures 4 and 5 generated by the object list generation unit 21 and the combination pattern generation unit 22, and is written to a file for spot production data. As shown in Figures 7A and 7B, the ADM is written as XML text data in UTF-8 character encoding.
[0054] In Figures 7A and 7B, ACO_XXXX is the ID of the descriptor "audioContent" in ADM, AP_YYYYXXXX is the ID of "audioPack", ATU_ZZZZZZZZ is the ID of "audioTrackUID", and AC_00010001 is the ID of "audioChannelFormat". "audioContent" is a component of "audioProgramme" and can group multiple "audioObjects". In the spot production data generation device 1, grouping may or may not be performed. If grouping is performed, for example, multiple "audioObjects" for each input list will be made into one "audioContent", and an "audioContent" called ACO_1001 (performer) will be generated, which is a bundle of "audioObjects" AO_1001 (performer 1) and AO_1002 (performer 2) that were generated based on the list of performers. If grouping is not performed, an "audioContent" will be generated that has a one-to-one correspondence with the "audioObject", and the XXXX and names of ACO_XXXX and AO_XXXX will be common to both the "audioObject" and the "audioContent". Figure 7A shows an example where grouping is not performed.
[0055] "audioPack" is a descriptor that indicates the audio format (mono, stereo, etc.) of "audioObject". For common formats such as mono and stereo, IDs are defined in the common definitions shown in Reference 2 below. (Reference 2) "Common definitions for the Audio Definition Model", Recommendation ITU-R BS.2094-1 (06 / 2017) For example, in Figure 7B, AP_00010001 means mono, and AP_00010002 means stereo.
[0056] "audioTrackUID" is a unique ID used to uniquely identify an audio track in the audio metadata. Traditionally, in broadcasting and live streaming, the audio track and the playback format (playback position, etc.) of the audio signal on that track were fixed on a one-to-one basis from the start to the end of the content. In contrast, in object-based audio, as shown in Figure 2, the audio signal on a single audio track is played back in different formats depending on its playback time. Therefore, "audioTrackUID" is used as a virtual audio track that can be uniquely distinguished in the metadata according to the playback format, separate from the physical audio track. Since the ZZZZZZZZZ part of ATU_ZZZZZZZZZ requires a unique number that does not overlap with others, it is assigned sequentially starting from the lowest numbered "audioObject" and incrementing for each track associated with this "audioObject".
[0057] "audioChannelFormat" indicates the channel format (stereo L, R, etc.) of each audio track (virtual audio track), and its ID is defined in the common definition in Reference 2 mentioned above. Thus, the spot production data generation device 1 can describe acoustic metadata in a format that conforms to international standards such as ADM.
[0058] As explained above, the spot production data generation device 1 can generate spot production data as a single audio file, which is intended to generate spots in which some of the audio may be replaced depending on the main advertising content, or the time and region of broadcast. Furthermore, if the same data were generated using the conventional method, it would be necessary to produce 10 audio files for 15-second commercials. If this production were centralized at a specific headquarters or key station, it would create a burden in terms of managing those files and distributing them to broadcasting stations nationwide. On the other hand, the spot production data generation device 1 only needs to produce a single audio file and distribute that file (a single file) to broadcasting stations nationwide, making it easy to manage files and prevent mix-ups.
[0059] Furthermore, broadcasting stations nationwide can use a renderer compatible with this acoustic metadata to generate or play the necessary spot audio signals simply by selecting a preset described in the acoustic metadata. By using an international standard such as ADM for this acoustic metadata, it is possible to generate the necessary spot audio signals using a standard renderer compatible with ADM as defined by ITU-R. Therefore, the spot production data generated by the spot production data generation device 1 can generate or play audio using conventional playback devices.
[0060] Although embodiments of the present invention have been described above, the present invention is not limited to these embodiments and can be modified as appropriate within the scope of the technical idea of the invention. Furthermore, the spot production data generation device 1 can be operated with a program that causes the computer to function as each of the aforementioned components. In that case, the computer can store a program in its memory that describes the processing content for realizing the functions of each component of the spot production data generation device 1, and then have the computer's CPU read and execute this program.
[0061] This program can be recorded on a computer-readable recording medium. Furthermore, by recording the program on a computer-readable medium, it can be installed on a computer. Here, the computer-readable medium may be a non-transient recording medium. The non-transient recording medium is not particularly limited, but may include, for example, CD-ROMs or DVD-ROMs. [Explanation of Symbols]
[0062] 1. Data generation device for spot production. 10 Audio data merging section 11 Channel Count Adjustment Section 12 Joint 13 Combined audio data storage unit 20 Acoustic Metadata Generation Unit 21 Object List Generation Unit 22 Combination Pattern Generation Unit 23 Metadata Generation Unit 30 Data Integration Department< / bxml> < / data> < / axml>
Claims
1. A spot production data generation device that generates data for spot production, An audio data merging unit inputs audio data consisting of one or more channels corresponding to each predetermined time interval of a spot for each of the said time intervals, and generates combined audio data by combining them in the time direction. An acoustic metadata generation unit inputs a list of audio data channels for each time interval and generates acoustic metadata that identifies the playback position of each audio data channel in the combined audio data. A data integration unit that integrates the combined audio data and the acoustic metadata to generate the data for spot production, A data generation device for spot production, characterized by being equipped with the following features.
2. The audio data merging unit is characterized by inserting null data into the audio data with fewer channels than the already merged audio data and the newly input audio data to equalize the number of channels, and then merging the respective audio data, as described in claim 1.
3. The audio metadata generation unit generates audio metadata from the list, with the number of audio data constituting the number of channels constituting the audio being set as a single audio object, and the combination of the audio objects and the playback time being set as presets, as described in claim 1, for the spot production data generation device.
4. The spot production data generation apparatus according to claim 3, characterized in that the acoustic metadata generation unit generates the acoustic metadata using ADM as defined in ITR-R BS. 2076.
5. The data integration unit integrates the combined audio data and the acoustic metadata in the BW64 file format specified in ITR-R BS. 2088, as described in claim 1, for spot production data generation apparatus.
6. A program for causing a computer to function as a spot production data generation device according to any one of claims 1 to 5.
Citation Information
Patent Citations
ITRBS.2076-02
ITRBS.2127-0
Commercial material editing device
JP1996279983A
Acoustic signal reproducing device and acoustic signal preparation device
JP2014204323A
Acoustic signal auxiliary information conversion transmission apparatus and program
JP2019003185A