Data processing method and data processing device

The method generates multi-track audio data with separate motion channels to synchronize audio and video data efficiently, ensuring real-time performance and optimal playback across different environments.

JP2025124424APending Publication Date: 2025-08-26YAMAHA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024020476
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-14
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The synchronization of audio and video data is challenging due to the larger data volume of video, leading to delays and reduced real-time performance when transmitted separately, and playing back video in inappropriate environments can cause user discomfort.

Method used

A data processing method that generates multi-track audio data with audio signals in one channel and motion data associated with character movement in a separate channel, allowing synchronized playback without time codes, reducing data volume through efficient storage and transmission.

Benefits of technology

Enables optimal video playback in various environments without impeding real-time performance, providing an immersive experience by synchronizing audio and motion data efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025124424000001_ABST
    Figure 2025124424000001_ABST
Patent Text Reader

Abstract

To provide a data processing method capable of reproducing an optimum video image according to a reproduction environment without inhibiting a real-time performance.SOLUTION: A data processing method is a data processing method for generating multi-track audio data in which audio data on a plurality of channels including at least a first channel and a second channel is stored. A data string of a digital audio signal is stored in the first channel, and motion data which is motion information of a character related to the digital audio signal is stored in the second channel as a data string of the digital audio signal.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to a data processing method for processing multi-track audio data that stores audio data of multiple channels. [Background technology]

[0002] Patent document 1 discloses a transmitting means that associates the operation information and operation time stored in the second storage means with identification information that identifies the music information output by the first control means and transmits them to a specified server device.

[0003] Patent Document 2 describes generating a character motion image from the movement of the singer's head. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-47261 [Patent Document 2] Japanese Patent Application Laid-Open No. 2012-78526 Summary of the Invention [Problem to be solved by the invention]

[0005] The amount of video data is larger than the amount of audio data. Therefore, if audio data and video data are transmitted separately, the delay in the video data will be large. This makes it difficult to synchronize the audio data with the video data. Furthermore, if audio data and video data are transmitted in synchronization, real-time performance will be reduced (the transmission of the audio data will be delayed in order to synchronize the audio data with the video data).

[0006] Furthermore, when video data is distributed, the video may not be played back in a way that is appropriate for the environment of the playback venue. For example, if video data shot outdoors is played back in an indoor playback venue, the user may feel uncomfortable.

[0007] Therefore, an object of one embodiment of the present invention is to provide a data processing method that can reproduce an optimal video according to the reproduction environment without impeding real-time performance. [Means for solving the problem]

[0008] A data processing method according to one embodiment of the present invention is a data processing method for generating multi-track audio data that stores audio data of multiple channels including at least a first channel and a second channel, in which a data string of a digital audio signal is stored in the first channel, and motion data, which is information on character movement associated with the digital audio signal, is stored in the second channel as a data string of the digital audio signal. [Effects of the Invention]

[0009] According to one embodiment of the present invention, it is possible to play back an optimal video according to the playback environment without impeding real-time performance. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a block diagram showing a configuration of a data processing system 1. FIG. [Figure 2] 1 is a block diagram showing the main configuration of a data processing device 10. FIG. [Figure 3] 10 is a flowchart showing the operation of a CPU 104. [Figure 4] 10 is a flowchart showing the operation of a processing unit 154. [Figure 5] FIG. 2 is a diagram illustrating an example of the structure of motion data. [Figure 6] FIG. 10 is a diagram showing an example of a format when motion data is stored as a data string of a digital audio signal. [Figure 7] FIG. 1 is a block diagram showing a configuration of a data processing system 1A in a playback environment. [Figure 8]10 is a flowchart showing the operation of a processing unit 154 of the data processing device 10 during playback. DETAILED DESCRIPTION OF THE INVENTION

[0011] 1 is a block diagram showing the configuration of a data processing system 1. The data processing system 1 includes a data processing device 10, a mixer 11, and a motion capture device 12.

[0012] The devices are connected by a communication standard such as a USB cable, HDMI (registered trademark), Ethernet (registered trademark), or MIDI. The devices are installed in the room of a performer who is performing a live performance, for example.

[0013] The mixer 11 is connected to multiple audio devices such as microphones, musical instruments, and amplifiers. The mixer 11 receives digital or analog audio signals from the multiple audio devices. When the mixer 11 receives an analog audio signal, it converts the analog audio signal into a 24-bit digital audio signal with a sampling frequency of, for example, 48 kHz. The mixer 11 performs signal processing such as mixing, gain adjustment, equalization, and compression on the multiple digital audio signals. The mixer 11 transmits the processed digital audio signals to the data processing device 10.

[0014] The motion capture device 12 is a sensor for capturing the motion of an actor, and may be, for example, an optical, inertial, or image sensor. The data processing device 10 receives a sensor signal from the motion capture device 12. The data processing device 10 generates motion data based on the sensor signal from the motion capture device 12.

[0015] 2 is a block diagram showing the main configuration of the data processing device 10. The data processing device 10 is composed of a general personal computer or the like, and includes a display 101, a user interface (I / F) 102, a flash memory 103, a CPU 104, a RAM 105, and a communication interface (I / F) 106.

[0016] The display 101 is composed of, for example, an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode), and displays various information. The user I / F 102 is composed of a switch, a keyboard, a mouse, a trackball, a touch panel, or the like, and accepts user operations. When the user I / F 102 is a touch panel, the user I / F 102, together with the display 101, constitutes a GUI (Graphical User Interface, hereinafter abbreviated).

[0017] The communication I / F 106 is connected to the mixer 11 and the motion capture device 12 via a communication line such as a USB cable, HDMI (registered trademark), Ethernet (registered trademark), or MIDI. The communication I / F 106 receives a digital audio signal from the mixer 11. The communication I / F 106 also receives a sensor signal from the motion capture device 12. The mixer 11 may transmit the digital audio signal to the data processing device 10 via Ethernet (registered trademark) using a protocol such as Dante (registered trademark).

[0018] The CPU 104 corresponds to the processing unit of the present invention. The CPU 104 loads a program stored in a flash memory 103, which is a storage medium, into the RAM 105 to implement a predetermined function. For example, the CPU 104 displays an image for accepting user operations on the display 101 and accepts selection operations and the like for the image via the user I / F 102, thereby implementing a GUI. The CPU 104 receives a digital signal from another device via the communication I / F 106. The CPU 104 generates multi-track audio data based on the received digital signal. The CPU 104 also stores the generated multi-track audio data in the flash memory 103. Alternatively, the CPU 104 distributes the generated multi-track audio data.

[0019] The program read by CPU 104 does not need to be stored in flash memory 103 within the device itself. For example, the program may be stored in a storage medium of an external device such as a server. In this case, CPU 104 simply reads the program from the server into RAM 105 and executes it each time.

[0020] The multi-track audio data conforms to, for example, the Dante (registered trademark) protocol. For example, Dante (registered trademark) can transmit 64-channel audio signals over 100base-TX. Each channel stores, for example, a 24-bit digital audio signal with a sampling frequency of 48 kHz (uncompressed digital audio data in WAV format, etc.). However, in the present invention, the number of channels, sampling frequency, and bit depth of the multi-track audio data are not limited to this example. Furthermore, the multi-track audio data does not need to conform to the Dante (registered trademark) protocol.

[0021] Fig. 3 is a functional block diagram showing the minimum configuration of the present invention. The processing unit 154 shown in Fig. 3 is realized by software executed by the CPU 104. Fig. 4 is a flowchart showing the operation of the processing unit 154. The processing unit 154 receives a digital audio signal from the mixer 11 (S11) and a sensor signal from the motion capture device 12 (S12). The processes of S11 and S12 may be performed simultaneously, or either may be performed first.

[0022] Then, the processing unit 154 generates motion data based on the sensor signal received from the motion capture device 12 (S13).

[0023] FIG. 5 is a diagram showing an example of the structure of motion data. Motion data is information about the movement of a character, and is information for controlling the behavior of a 3D model of the character. The motion data in FIG. 5 includes description data and position information (bone data) for multiple bones. The multiple bones correspond to the movable parts of the character, such as the limbs and fingers. The multiple bones make up a skeleton that corresponds to the character's body.

[0024] Description data includes identification information for identifying bones (bone name information), information indicating the connection relationships with other bones, and skeleton identification information. Because description data is static information, it does not need to be included in motion data. For example, description data may be generated and output separately from motion data. Also, description data may be included once every few frames or every few tens of frames, rather than every frame.

[0025] Bone data includes position coordinate information and rotation coordinate information. Position coordinate information is expressed, for example, as a three-axis Cartesian coordinate system with a certain position as the origin. For example, if 32 bits of information are stored per frame for each of the three axes (X, Y, and Z), the position coordinate information will be 96 bits of data. For example, if 32 bits of information are stored per frame for each of the four axes (X, Y, Z, and W), the rotation coordinate information will be 128 bits. If one skeleton contains, for example, 30 bones, the motion data will be 6,720 bits of data per frame. If such bone data is driven at, for example, 60 fps, the motion data requires a data capacity of 403,200 bits (50,400 bytes) per second. Note that motion data does not need to be absolute coordinates. Motion data can also be relative coordinates with respect to a reference bone. Relative coordinates do not need to be included every frame; they can be included only in frames where there is a change relative to the reference bone.

[0026] After generating the motion data, the processing unit 154 generates multi-track audio data based on the digital audio signal and the motion data (S12). Specifically, the processing unit 154 stores the data string of the digital audio signal received from the mixer 11 in a first channel of the multi-track audio data. When the processing unit 154 receives audio data conforming to the Dante (registered trademark) protocol from the mixer 11, it stores the data string of the digital audio signal for each channel of the received audio data as is in the same channel of the multi-track audio data.

[0027] Furthermore, the processing unit 154 stores the motion data as a data string of a digital audio signal in the second channel of the multi-track audio data.

[0028] FIG. 6 shows an example of a format for storing motion data as a data sequence of digital audio signals. As described above, the digital audio signals of each channel of the multi-track audio data are, for example, 24-bit digital audio signals with a sampling frequency of 48 kHz. FIG. 5 shows a data sequence of a certain sample of a certain channel in the multi-track audio data. The multi-track audio data has a first channel in which the received digital audio signal is stored as a data sequence of digital audio signals as is, and a second channel consisting of a data sequence as shown in FIG. 6. There can be any number of first channels and any number of second channels. As an example, the processing unit 154 stores the digital audio signals of the audio data received from the mixer 11 in channels 1 to 32 as the first channels, and stores motion data in channels 33 to 64 as the second channels.

[0029] 6, each sample in the second channel consists of 8-bit header information and the motion data itself, which contains a 16-bit data string.

[0030] The 8-bit header information includes a start bit, a data size flag, and type data. The start bit is a 1-bit data that indicates whether the data of that sample is the first data or not. A start bit of 1 indicates that it is the first data. In other words, the samples from one sample whose start bit indicates 1 to the next sample whose start bit indicates 1 correspond to one piece of motion data.

[0031] The data size flag is 1 bit of data that indicates the data size. For example, if the data size flag is 1, it indicates that one sample contains 8 bits of data, and if the data size flag is 0, it indicates that one sample contains 16 bits of data. For example, if one motion data contains 24 bits of data, the start bit of the first sample will be 1 and the data size flag will be 0. The start bit of the second sample will be 0 and the data size flag will be 1. If one motion data is 8 bits of data, the start bit of each sample will be 0.

[0032] Next, type is identification information that indicates the type of data, and is 6-bit data. In this example, type is 6 bits, so it indicates 64 data types. Of course, the number of bits is not limited to 6 bits. The number of type bits can be set to the number of bits according to the number of types required.

[0033] For example, if the type is "01", it indicates motion data. If the type is "02", it indicates other data (e.g., MIDI data). If the type is "00", it indicates idle data (empty data). However, it is preferable not to use "3F", in which all bit data is 1. If the processing unit 154 does not use "3F", the device that plays back multi-track audio data can determine that some kind of failure has occurred when all bit data is 1. If all bit data is 1, the processing unit 154 does not play back the multi-track audio data and does not output a control signal. This prevents the processing unit 154 from damaging equipment such as lighting when an abnormal signal is output.

[0034] When the processing unit 154 sets the type to "00" and sets all bit data to 0, i.e., idle data (empty data), it is preferable that the start bit be set to 1 to prevent all bits from becoming 0. This allows the device that processes multi-track audio data to determine that some kind of failure has occurred when all bit data is 0. If all bit data is 0, the processing unit 154 does not play back the multi-track audio data. In this case, too, the processing unit 154 prevents an abnormal signal from being output, which may cause damage to the device.

[0035] Furthermore, the processing unit 154 may not play back the multi-track audio data if it contains data that does not match the data format shown in Fig. 6. For example, if the data size flag is 1, the lower 8 bits will all be 0's (or all 1's). Therefore, even if the data size flag is 1, the processing unit 154 will not play back the multi-track audio data if the lower 8 bits contain a mixture of 0's and 1's. In this case, too, the processing unit 154 will not cause damage to the device when an abnormal signal is output.

[0036] The processing unit 154 may include only data of the same type in the same channel, or may include data of different types in the same channel. That is, the processing unit 154 may include only one piece of identification information in the same channel among the second channels, or may include multiple pieces of identification information in one channel.

[0037] Furthermore, the processing unit 154 may store data of the same type across multiple predetermined channels. For example, the processing unit 154 stores data in three channels, 33, 34, and 35, in order. In this case, channels 33, 34, and 35 all contain the same type of bit data. If the data size flags of channels 33, 34, and 35 are all set to 0, the processing unit 154 can store three times as much data (48 bits) in one sample. Alternatively, if the processing unit 154 stores the same type of data in, for example, four channels, it can store four times as much data (64 bits) in one sample. In this way, by storing data across multiple channels, the processing unit 154 can also store data of a type that generates a large amount of data per unit time (e.g., video data) in one sample.

[0038] Furthermore, when storing data across multiple channels, processing unit 154 may add header information to only one representative channel (e.g., channel 33) and not to the other channels. In this case, processing unit 154 can store 16-bit data in the representative channel and up to 24-bit data in the other channels. That is, processing unit 154 can store up to 64-bit data across three channels. In this case, the data size flag for the representative channel may be, for example, 3 bits (bit data indicating 8 types of data sizes) instead of 1 bit.

[0039] The data processing device 10 may also receive signal processing parameters indicating the content of signal processing and data related to basic settings of the mixer 11 from audio equipment such as the mixer 11, and store them as a data string of digital audio signals in the second channel.

[0040] As mentioned above, in the present invention, the sampling frequency of multi-track audio data is not limited to 48 kHz, and the number of bits is not limited to 24. For example, if the sampling frequency of multi-track audio data is 96 kHz, the processing unit 154 may generate 48 kHz sampling data by invalidating one sample every two samples. Furthermore, if the number of bits is 32, the processing unit 154 may use only the most significant 24 bits, without using the least significant 8 bits.

[0041] For example, if the multi-track audio data is 48 kHz and 24 bits, one sample data string is 16 bits, and the multi-track audio data can store a capacity of 16 x 48000 = 768000 bits (96000 bytes) per second.

[0042] On the other hand, when bone data is driven at, for example, 60 fps, motion data requires a data capacity of 403,200 bits (50,400 bytes) per second. In other words, multi-track audio data has a bit rate approximately twice that of motion data. Therefore, the data processing device 10 can store motion data at a bit rate lower than that of multi-track audio data. The data processing device 10 may invalidate at least one sample among multiple samples and store data in the remaining sample. For example, the data processing device 10 may invalidate one sample among two samples and store motion data in the remaining sample. In this case, the motion data of the second channel may be out of sync by approximately one sample on the playback side. However, the time lag at a sampling frequency of 48 kHz is only 0.0208 msec, and even if there is an lag of several samples, the time lag will be less than 1 msec. Therefore, the viewer will not perceive a lag between the image and sound due to the motion data.

[0043] In this manner, multi-track data recording the live streaming performance is generated. The processing unit 154 outputs the multi-track audio data generated in this manner (S15). The multi-track audio data may be stored in the flash memory 103 of the device itself, or may be distributed to another device via the communication I / F 106.

[0044] As described above, the data processing system 1 is installed in a broadcaster's room where a performance such as a live performance is being performed. The data processing device 10 stores a digital audio signal received from the audio equipment during the live performance in a first channel, and stores motion data in a second channel. The digital audio signal and motion data are generated in accordance with the progress of the performance such as the live performance. The motion data is stored at the same sampling frequency (e.g., 48 kHz) as the digital audio signal.

[0045] Therefore, the data processing device 10 can synchronize audio signals and motion data without using time codes by generating multi-track audio data in which a digital audio signal with a predetermined sampling frequency (e.g., 48 kHz) is stored in a first channel and motion data with the same frequency (48 kHz) as the first channel is stored in a second channel. This eliminates the need for the data processing system 1 to prepare dedicated recording devices for each protocol used by multiple devices and record each piece of data individually. Furthermore, the data processing system 1 does not require a time code generator, cables, interfaces, etc. for synchronizing audio signals and motion data with time codes. Furthermore, the data processing system 1 does not require settings such as matching the frame rate of the time code for each device.

[0046] The second channel shown in this embodiment stores motion data as a data string of digital audio signals, and therefore conforms to a predetermined multi-track audio data protocol (e.g., the Dante (registered trademark) protocol). Therefore, the second channel can be distributed and played back as audio data, and can also be edited using an audio data editing application program such as a DAW (Digital Audio Workstation) to perform operations such as copying, cutting, pasting, and timing adjustment. For example, a user can use a DAW to cut and paste audio data from the first and second channels contained in a certain time period to a different time period, thereby moving not only the audio data but also the motion data to a different time period without disrupting synchronization.

[0047] 7 is a block diagram showing the configuration of a data processing system 1A in a reproduction environment. The data processing system 1A is installed in a venue for reproducing a performance such as a live musical performance at a remote location.

[0048] The same components as those in the data processing system 1 shown in Fig. 1 are denoted by the same reference numerals, and a description thereof will be omitted. The data processing system 1A includes a display 13. The display 13 may be a panel-type display device such as an LCD, or may be a video display device such as a projector.

[0049] The data processing device 10 at the playback venue does not need to have the same hardware configuration as the venue where the event, such as a live performance, was held.

[0050] 8 is a flowchart showing the operation of the processing unit 154 of the data processing device 10 during playback. First, the processing unit 154 accepts multi-track audio data (S21). The multi-track audio data is received from the data processing device 10 at the venue where the live performance is being held, or from a server. Alternatively, the data processing device 10 reads out multi-track audio data relating to a past live performance stored in the flash memory 103. Alternatively, the data processing device 10 reads out multi-track audio data relating to a past live performance stored in another device such as a server.

[0051] The processing unit 154 decodes the received multi-track audio data to reproduce digital audio signals and motion data (S22). For example, the processing unit 154 extracts digital audio signals from channels 1 to 32, which are the first channels. The processing unit 154 outputs the extracted digital audio signals to the mixer 11 (S23). The mixer 11 outputs the received digital audio signals to an audio device such as a speaker, and reproduces the singing sound or performance sound.

[0052] In addition, in the process of S22, the processing unit 154 reads 8-bit header information for each sample of channels 33 to 64, which are the second channel, and extracts the 8-bit or 16-bit motion data body. Note that the second channel may include data related to signal processing parameters and basic settings of the mixer 11. The processing unit may extract this signal processing parameters and data related to basic settings and output it to the mixer 11.

[0053] The mixer 11 receives signal processing parameters and setting data from the data processing device 10, and performs various signal processing on the received audio signal based on the signal processing parameters and setting data, thereby reproducing singing and performance sounds in the same condition as a live performance.

[0054] The processing unit 154 renders an image of the 3D model character based on the extracted motion data (S24). Here, rendering means controlling the 3D model of the character based on the motion data and converting the 3D model into a two-dimensional video signal. As shown in FIG. 5, the motion data includes bone data corresponding to the movable parts of the character, such as the limbs and fingers. The processing unit 154 determines the configuration of the 3D model of the character in each frame based on the position coordinate information and rotation coordinate information included in the bone data. The processing unit 154 generates an image of the character as viewed from a specified viewpoint using a predetermined method such as ray tracing or scan line.

[0055] The processing unit 154 outputs the generated video signal to the display unit 13 (S25). The display unit 13 displays an image based on the received video signal.

[0056] As described above, the data processing device 10 extracts and outputs the digital audio signal of the first channel of each sample at the playback site, and extracts the motion data of the second channel of each sample to render the character image, thereby enabling the digital audio signal and the video signal related to the character image to be played back in synchronization without using a time code.

[0057] Typically, the amount of video data is larger than the amount of audio data. Therefore, if audio data were delayed and synchronized with the video data, the delivery of the audio data would be delayed, resulting in a decrease in real-time performance. However, the motion data shown in this embodiment has a bit rate (403.2 kbps), which is half the bit rate (768 kbps) of 48 kHz, 24 bits. Therefore, the data processing system of this embodiment can store digital audio signals and motion data within the same frame, without degrading real-time performance. Furthermore, even if motion data is stored across multiple samples, the time lag for one sample at a sampling frequency of 48 kHz is less than 1 msec. Therefore, users of the data processing system of this embodiment can enjoy a new customer experience: they can perceive themselves as if they were participating in an event such as a live performance, even if they are in a different venue from the venue where the live performance is taking place.

[0058] Furthermore, for example, if the distribution venue is outdoors and video data shot outdoors is played back in an indoor playback venue, the user may feel uncomfortable. However, the data processing system of this embodiment can distribute motion data by, for example, capturing images of the venue using a camera installed at the destination venue and then synthesizing the character's image with the video of the venue. Therefore, the data processing system of this embodiment can play back the optimal video according to the playback environment. This allows users to enjoy a new customer experience, one that allows them to feel an immersive sensation that was previously not possible.

[0059] The process of outputting the digital audio signal in S23 and the process of outputting the video signal in S25 shown in FIG. 8 may be performed simultaneously, or either may be performed first.

[0060] (Variation 1) The motion data of the first modification includes main motion data, which is data on the movement of the main body part of the character, and sub-motion data, which is data on the movement of the body parts subordinate to the main body part.

[0061] The main parts of a character are at least the upper arms and forearms, which are important parts for conveying the performance. Subordinate parts are parts that are not directly necessary for conveying the performance, such as the torso, legs, hands, and fingers. However, the main and subordinate parts differ depending on the type of performance. For example, when playing the guitar, the movement of the forearm of the right arm and the movement of the forearm and fingers of the left arm are important parts for conveying the performance, while when playing the piano, the forearms and movements of both arms are important parts.

[0062] The multi-track audio data of the first modification has a first channel for storing a digital audio signal, a second channel for storing main motion data, and a third channel for storing a sub-motion channel.

[0063] The processing unit 154 of the data processing device 10 may determine whether or not to play back sub-motion data of the third channel during playback depending on the processing power, the usage rate of the CPU 104, or the free space of the RAM 105. As a result, even if the processing power, the usage rate of the CPU 104, or the free space of the RAM 105 is low, the data processing device 10 does not stop or delay playback, and real-time performance is not reduced. Furthermore, even if the data processing device 10 does not play back sub-motion data, it plays back main motion data, which is an important part for conveying the performance, so the user's sense of immersion is not impaired.

[0064] (Variation 2) The data processing device 10 according to the second modification accepts the designation of the character to be played back, and renders the video of the designated character using the extracted motion data.

[0065] For example, if the first venue to which distribution is to be made is in the Northern Hemisphere and it is winter, the data processing device 10 at the first venue will accept the designation of a character wearing winter clothing. On the other hand, if the second venue to which distribution is to be made is in the Southern Hemisphere and it is summer, the data processing device 10 at the second venue will accept the designation of a character wearing summer clothing.

[0066] This allows the data processing device 10 of the second modification to play back video that is more optimal for the playback environment.

[0067] (Variation 3) The data processing device 10 according to the third modification further corrects the extracted motion data based on the specified character in the second modification, and renders an image of the specified character using the corrected motion data.

[0068] The size of the 3D model data may vary depending on the character. Therefore, the data processing device 10 corrects the length of each bone data included in the motion data based on the data of the specified character (for example, data indicating height).

[0069] This allows the data processing device 10 of the third modification to reproduce more optimal video images.

[0070] (Variation 4) The data processing device 10 according to the fourth modification accepts the designation of a background image, and superimposes the rendered image of the character on the designated background image.

[0071] As described above, for example, if the venue from which the distribution originates is outdoors and video data shot outdoors is played back in an indoor playback venue, the user may feel uncomfortable. However, data processing device 10 of Modification 4 superimposes (combines) the rendered character video onto the specified background video. For example, if the playback venue is indoors, data processing device 10 of Modification 4 may superimpose an indoor background video. Alternatively, data processing device 10 of Modification 4 may superimpose a summer background video if the playback venue is summer, or a winter background video if the playback venue is winter.

[0072] Therefore, the data processing device 10 of the fourth modification can play back the most suitable video according to the playback environment, and the user can have a new customer experience of feeling immersed in the content, which was not possible with the conventional technology.

[0073] The description of the present embodiment should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, not by the above-described embodiments. Furthermore, the scope of the present invention includes a range equivalent to the claims. For example, in the above-described embodiment, the processing unit of the present invention is configured by a program read by the CPU 104, but it is also possible to realize the processing unit of the present invention by, for example, an FPGA (Field-Programmable Gate Array).

[0074] The data processing system shown in this embodiment can also be applied to systems that require a combination of video and audio, such as theme parks. Alternatively, the data processing system shown in this embodiment can also be applied to a remote session system. In a remote session, it is important to convey not only the sound but also the movements of the performer or singer. However, as described above, the data volume of video data is larger than the data volume of audio data, so transmitting the audio data and video data separately results in significant delays in the video data. Furthermore, transmitting the audio data and video data synchronously reduces real-time performance. In contrast, the data processing system shown in this embodiment transmits and receives motion data, which is smaller in data volume than video data, and can therefore transmit the movements of a remote performer or singer with low delay without reducing real-time performance. [Explanation of symbols]

[0075] 1: Data processing system, 1A: Data processing system, 10: Data processing device, 11: Mixer, 12: Motion capture, 13: Display, 101: Display, 102: User I / F, 103: Flash memory, 104: CPU, 105: RAM, 106: Communication I / F, 154: Processing unit

Claims

1. A data processing method for generating multi-track audio data storing audio data of multiple channels including at least a first channel and a second channel, comprising: storing a data string of a digital audio signal in the first channel; storing motion data, which is information on character movement and is associated with the digital audio signal, as a data string of the digital audio signal in the second channel; Data processing methods.

2. Distributing the multi-track audio data. The data processing method according to claim 1 .

3. A data processing method for reproducing multi-track audio data that stores audio data of multiple channels, comprising: Reproducing the data string stored in the first channel as a digital audio signal; extracting the data string stored in the second channel as motion data, which is information on character movement associated with the digital audio signal, and rendering an image of the character using the extracted motion data; Data processing methods.

4. the second channel includes identification information of the motion data; The data processing method according to any one of claims 1 to 3.

5. the motion data includes main motion data, which is data on the movement of a main body part of the character, and sub-motion data, which is data on the movement of a body part subordinate to the main body part; the audio data further comprises a third channel in which the sub-motion data is stored; The data processing method according to any one of claims 1 to 3.

6. Accepts the specification of the character to be played, Rendering an image of the specified character using the extracted motion data. The data processing method according to claim 3 .

7. correcting the extracted motion data based on the specified character; Rendering an image of the designated character using the corrected motion data; 7. The data processing method according to claim 6.

8. Accepts the background video specification, superimposing the rendered image of the character on the specified background image; The data processing method according to claim 3 .

9. A data processing method for generating multi-track audio data storing audio data of multiple channels including at least a first channel and a second channel, comprising: storing a data string of a digital audio signal in the first channel; storing motion data, which is information on character movement and is associated with the digital audio signal, as a data string of the digital audio signal in the second channel; A data processing device having a processing unit.

Citation Information

Patent Citations

  • Karaoke system

    JP2012078526A

  • Information processor and program

    JP2015047261A