Data processing method and data processing apparatus

CN122767005APending Publication Date: 2026-09-15YAMAHA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202580014258.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-14
Filing Date
2025-01-21
Publication Date
2026-09-15

AI Technical Summary

Benefits of technology

[0015] According to one embodiment of the present invention, preferred images can be regenerated without hindering real-time performance, depending on the regeneration environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122767005A_ABST
    Figure CN122767005A_ABST
Patent Text Reader

Abstract

The data processing method is a data processing method for generating multi-channel audio data, and the multi-channel audio data stores audio data of a plurality of channels including at least a first channel and a second channel. In the first channel, a data series of a digital audio signal is stored. In the second channel, action data of an action of a character associated with the digital audio signal is stored as a data series of a digital audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One embodiment of the present invention relates to a data processing method for processing multitrack audio data that stores audio data from multiple channels. Background Technology

[0002] Patent Document 1 discloses a sending unit that associates the action information and the action time stored in the second storage unit with identification information that identifies the music information output by the first control unit, and sends it to a designated server device.

[0003] Patent document 2 describes the generation of motion images of a character based on the movements of the singer's head.

[0004] Existing technical documents

[0005] Patent documents

[0006] Patent Document 1: Japanese Patent Application Publication No. 2015-47261

[0007] Patent Document 2: Japanese Patent Application Publication No. 2012-78526 Summary of the Invention

[0008] The problem that the invention aims to solve

[0009] Video data is much larger than audio data. Therefore, if audio and video data are transmitted separately, the latency of the video data will increase. Consequently, synchronizing audio and video data becomes difficult. Furthermore, if audio and video data are transmitted synchronously, real-time performance is reduced (audio data transmission is delayed to ensure consistency with video data).

[0010] Furthermore, sometimes the transmitted video data does not reproduce images that match the environment of the venue. For example, if video data taken outdoors is reproduced in an indoor venue, users will feel a sense of incongruity.

[0011] Therefore, one embodiment of the present invention aims to provide a data processing method that can regenerate preferred images without hindering real-time performance, based on the regeneration environment.

[0012] Methods for solving problems

[0013] One embodiment of the present invention relates to a data processing method for generating multitrack audio data, which stores audio data from multiple channels, including at least a first channel and a second channel. In the first channel, a data string of digital audio signals is stored, and in the second channel, motion data, which is associated with the digital audio signals and represents the character's actions, is stored as a data string of digital audio signals.

[0014] Invention Effects

[0015] According to one embodiment of the present invention, preferred images can be regenerated without hindering real-time performance, depending on the regeneration environment. Attached Figure Description

[0016] Figure 1 This is a block diagram showing the structure of data processing system 1.

[0017] Figure 2 This is a block diagram showing the main structure of the data processing device 10.

[0018] Figure 3 This is a flowchart illustrating the operation of CPU104.

[0019] Figure 4 This is a flowchart illustrating the operation of the processing unit 154.

[0020] Figure 5 This is a diagram illustrating a structural example of motion data.

[0021] Figure 6 This is a diagram illustrating an example of a format in which motion data is stored as a data column of digital audio signals.

[0022] Figure 7 This is a block diagram showing the structure of a data processing system 1A in a regenerative environment.

[0023] Figure 8 This is a flowchart illustrating the operation of the processing unit 154 of the data processing device 10 during regeneration. Detailed Implementation

[0024] Figure 1 This is a block diagram showing the structure of the data processing system 1. The data processing system 1 includes a data processing device 10, a mixer 11, and a motion capture device 12.

[0025] The devices are connected via communication standards such as USB cable, HDMI (registered trademark), Ethernet (registered trademark), or MIDI. The devices are installed, for example, in the room of a performer giving a live performance.

[0026] Mixer 11 connects to multiple audio devices such as microphones, musical instruments, or amplifiers. Mixer 11 receives digital or analog audio signals from these devices. Upon receiving an analog audio signal, mixer 11 converts it into a digital audio signal, for example, a 24-bit digital signal with a sampling frequency of 48kHz. Mixer 11 performs signal processing on the multiple digital audio signals, including mixing, gain adjustment, equalization, or compression. Mixer 11 then sends the processed digital audio signal to data processing unit 10.

[0027] The motion capture device 12 is a sensor used to capture the movements of a performer, such as an optical, inertial, or image-based sensor. The data processing device 10 receives sensor signals from the motion capture device 12. Based on the sensor signals from the motion capture device 12, the data processing device 10 generates motion data.

[0028] Figure 2 This is a block diagram showing the main structure of the data processing device 10. The data processing device 10 is composed of a general personal computer or the like, and includes a display 101, a user interface (I / F) 102, flash memory 103, a CPU 104, RAM 105, and a communication interface (I / F) 106.

[0029] The display 101 is composed of, for example, an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode), and displays various information. The user I / F 102 is composed of a switch, keyboard, mouse, trackball, or touch panel, and accepts user operations. When the user I / F 102 is a touch panel, it together with the display 101 constitutes a GUI (Graphical User Interface).

[0030] The communication I / F106 is connected to the mixer 11 and the motion capture device 12 via a communication cable such as a USB cable, HDMI (registered trademark), Ethernet (registered trademark), or MIDI. The communication I / F106 receives digital audio signals from the mixer 11. Additionally, the communication I / F106 receives sensor signals from the motion capture device 12. Furthermore, the mixer 11 can also transmit digital audio signals to the data processing device 10 via Ethernet (registered trademark) using protocols such as Dante (registered trademark).

[0031] CPU 104 corresponds to the processing unit of this invention. CPU 104 reads the program stored in flash memory 103 (which is a storage medium) into RAM 105 to implement a specified function. For example, CPU 104 displays an image on display 101 for accepting user operations, and accepts selection operations on the image via user I / F 102, thereby implementing a GUI. CPU 104 receives digital signals from other devices via communication I / F 106. CPU 104 generates multitrack audio data based on the received digital signals. Furthermore, CPU 104 stores the generated multitrack audio data in flash memory 103. Alternatively, CPU 104 transmits the generated multitrack audio data.

[0032] Furthermore, the program read by CPU 104 does not need to be stored in the flash memory 103 within this device. For example, the program can also be stored in the storage medium of an external device such as a server. In this case, CPU 104 only needs to read the program from the server into RAM 105 and execute it each time.

[0033] Multitrack audio data may comply with the Dante (registered trademark) protocol, for example. Dante (registered trademark) is capable of transmitting 64 channels of audio signals at 100base-TX. Each channel may store, for example, a 24-bit digital audio signal with a sampling frequency of 48kHz (uncompressed digital audio data such as WAV format). However, in this invention, the number of channels, sampling frequency, and bit count of the multitrack audio data are not limited to this example. Furthermore, the multitrack audio data need not comply with the Dante (registered trademark) protocol.

[0034] Figure 3 This is a functional block diagram illustrating the minimum structure of the present invention. Figure 3 The processing unit 154 shown is implemented by software executed by the CPU 104. Figure 4 This is a flowchart illustrating the operation of the processing unit 154. The processing unit 154 receives digital audio signals from the mixer 11 (S11) and sensor signals from the motion capture unit 12 (S12). The processing of S11 and S12 can be performed simultaneously or either one can be performed first.

[0035] Then, the processing unit 154 generates motion data based on the sensor signals received from the motion capture unit 12 (S13).

[0036] Figure 5 This is a diagram illustrating an example of the structure of motion data. Motion data is the character's motion information, which is used to control the movements of the character's 3D model. Figure 5The motion data includes description data and positional information (skeleton data) for multiple bones. These bones correspond to the character's movable parts such as hands, feet, and fingers. These bones together form the skeleton corresponding to the character's body.

[0037] Description data includes identification information for bones (bone name information), information indicating connections to other bones, and skeletal identification information. Since description data is static information, it does not need to be included in motion data. For example, description data can be generated and output separately from motion data. Furthermore, description data can be included every few frames or every tens of frames, rather than every frame.

[0038] Skeletal data includes positional coordinates and rotational coordinates. Positional coordinates are represented, for example, using orthogonal coordinates along three axes with a given origin. If each of the three axes (X, Y, Z) is stored, for example, 32 bits of information per frame, the positional coordinates become 96 bits of data. If each of the four axes (X, Y, Z, W) is stored, for example, 32 bits of information per frame, the rotational coordinates become 128 bits. Furthermore, if a skeleton includes, for example, 30 bones, the motion data becomes 6720 bits of data per frame. Driving such skeletal data at, for example, 60fps requires 403200 bits (50400 bytes) of data per second. Additionally, motion data does not need to be absolute coordinates. Motion data can also be relative coordinates to a reference bone. Relative coordinates are not required in every frame and can be included only in frames where there are changes relative to the reference bone.

[0039] After generating motion data, processing unit 154 generates multitrack audio data based on the digital audio signal and motion data (S12). Specifically, processing unit 154 stores the data column of the digital audio signal received from mixer 11 in the first channel of multitrack audio data. When processing unit 154 receives audio data conforming to the Dante (registered trademark) protocol from mixer 11, it directly stores the data column of the digital audio signal of each channel of the received audio data in the same channel of multitrack audio data.

[0040] In addition, the processing unit 154 stores motion data as a data column of digital audio signals in the second channel of the multi-track audio data.

[0041] Figure 6 This is a diagram illustrating an example of a format where motion data is stored as a data column of digital audio signals. As described above, as an example, the digital audio signals for each channel of the multitrack audio data are 24-bit digital audio signals with a sampling frequency of 48kHz. Figure 5 This shows a data column of a sample from a specific channel in a multitrack audio data set. The multitrack audio data set has a first channel and a second channel. The first channel directly stores the received digital audio signal as a data column of digital audio signals. The second channel consists of... Figure 6 The data is structured as shown. The number of first channels and second channels can also be several. As an example, the processing unit 154 stores the digital audio signal of the audio data received from the mixer 11 as the first channel in channels 1 to 32, and stores the motion data as the second channel in channels 33 to 64.

[0042] like Figure 6 As shown, each sample in the second channel consists of an 8-bit header and a data body of motion data. The data body includes a 16-bit data column.

[0043] The 8-bit header information includes a start bit, a data size flag, and type data. The start bit is a 1-bit data indicating whether the sample data is the first data. When the start bit is 1, it indicates that it is the first data. That is, the sequence from the sample with a start bit of 1 to the next sample with a start bit of 1 corresponds to one action data.

[0044] The data size flag is a 1-bit data symbol representing the data size. For example, a data size flag of 1 indicates that a sample contains 8 bits of data, while a data size flag of 0 indicates that a sample contains 16 bits of data. For instance, in an action data set containing 24 bits of data, the first sample has a starting bit of 1 and a data size flag of 0. The second sample has a starting bit of 0 and a data size flag of 1. In an action data set containing 8 bits of data, all samples have a starting bit of 0.

[0045] Next, the type is identification information indicating the category of data, and it is 6 bits. In this example, the type is 6 bits, thus representing 64 data categories. Of course, the number of bits is not limited to 6 bits. The number of bits for the type can be set to correspond to the number of categories required.

[0046] For example, data of type "01" represents motion data. Data of type "02" represents other data (e.g., MIDI data). Furthermore, data of type "00" represents idle data (empty data). However, it is preferable not to use "3F" where all bits are 1. When the processing unit 154 does not use "3F", the device for reproducing multitrack audio data can determine that a malfunction has occurred when all bits are 1. When all bits are 1, the processing unit 154 does not reproduce multitrack audio data and does not output control signals. Therefore, in the event of an abnormal signal output, the processing unit 154 will not cause damage to equipment such as lighting.

[0047] Furthermore, when the processing unit 154 sets the type to "00" and all bit data to 0, i.e., idle data (empty data), it is preferable to set the starting bit to 1 so that all bits are not 0. Therefore, the device processing multitrack audio data can determine that a certain fault has occurred when all bit data is 0. When all bit data is 0, the processing unit 154 will not reproduce multitrack audio data. In this case, the processing unit 154 will not output abnormal signals that could damage the equipment.

[0048] In addition, the processing unit 154 includes [missing information - likely related to...]. Figure 6 In cases where the data format is inconsistent, multitrack audio data may not be reproduced. For example, when the data size flag is 1, the lower 8 bits are all 0s (or all 1s). Therefore, even if the data size flag is 1, and the lower 8 bits contain a mixture of 0s and 1s, the processing unit 154 will not reproduce multitrack audio data. In this case, even if an abnormal signal is output, the processing unit 154 will not cause damage to the device.

[0049] Furthermore, the processing unit 154 may include only the same type of data in the same channel, or it may include different types of data in the same channel. That is, the processing unit 154 may include only one identification information in the same channel in the second channel, or it may include multiple identification information in one channel.

[0050] Furthermore, the processing unit 154 can also store data of the same category across multiple predetermined channels. For example, the processing unit 154 stores data sequentially in channels 33, 34, and 35. In this case, channels 33, 34, and 35 all contain bit data of the same type. If we assume that the data size flags of channels 33, 34, and 35 are all set to 0, the processing unit 154 can store three times (48 bits) of data in a single sample. Alternatively, for example, if the same category of data is stored in four channels, the processing unit 154 can store four times (64 bits) of data in a single sample. In this way, by storing data across multiple channels, the processing unit 154 can store data of the category with the largest data volume per unit time (e.g., image data) in a single sample.

[0051] Furthermore, when storing data across multiple channels, the processing unit 154 may representatively append header information to only one channel (e.g., channel 33), while leaving the other channels unattached. In this case, the processing unit 154 stores 16 bits of data in the representative channel, and the other channels can store a maximum of 24 bits of data. That is, the processing unit 154 can store a maximum of 64 bits of data across the three channels. Additionally, in this case, the data size flag for the representative channel may not be 1 bit, but rather, for example, 3 bits (representing 8 different data sizes).

[0052] In addition, the data processing device 10 can also receive signal processing parameters representing the content of signal processing and data related to the basic settings of the mixer 11 from audio equipment such as the mixer 11, and store them in the second channel as a data column of digital audio signals.

[0053] Furthermore, as described above, in this invention, the sampling frequency of the multitrack audio data is not limited to 48kHz, and the number of bits is not limited to 24 bits. For example, when the sampling frequency of the multitrack audio data is 96kHz, the processing unit 154 may also generate 48kHz sampling data as invalid data every two samples. Moreover, when the number of bits is 32 bits, the processing unit 154 may not use the lower 8 bits, but only the upper 24 bits.

[0054] In the case of multitrack audio data, for example, 48kHz, 24 bits, a single sample data column is, for example, 16 bits. In this case, the multitrack audio data can store a capacity of 16 × 48000 = 768000 bits (96000 bytes) of data per second.

[0055] On the other hand, for example, when driving skeletal data at 60fps, motion data requires a data capacity of 403,200 bits (50,400 bytes) per second. That is, multitrack audio data has approximately twice the bit rate of motion data. Therefore, the data processing device 10 is capable of storing motion data at a bit rate lower than that of multitrack audio data. Furthermore, the data processing device 10 can also set at least one sample out of multiple samples as invalid data and store data in the samples other than the invalid data. For example, the data processing device 10 can also set one of two samples as invalid data and store motion data in the remaining sample. In this case, on the playback side, the motion data of the second channel may deviate by about one sample. However, the time deviation at a sampling frequency of 48kHz is only 0.0208 msec, and even assuming a deviation of several samples, the time deviation converges to less than 1 msec. Therefore, the viewer will not perceive any deviation between the image and sound caused by the motion data.

[0056] As described above, multitrack data recording the live performance is generated. Processing unit 154 outputs the multitrack audio data generated as described above (S15). The multitrack audio data can be stored in the flash memory 103 of this device or transmitted to other devices via communication I / F 106.

[0057] As described above, the data processing system 1 is installed, for example, in the room of the transmitter performing a live performance. The data processing device 10 stores digital audio signals received from the sound equipment during the live performance in a first channel and motion data in a second channel. The digital audio signals and motion data are generated based on the same live performance. The motion data is stored at the same frequency as the sampling frequency of the digital audio signals (e.g., 48 kHz).

[0058] Therefore, by generating multitrack audio data that stores a digital audio signal with a predetermined sampling frequency (e.g., 48kHz) in the first channel and motion data with the same frequency (48kHz) in the second channel, the data processing device 10 can synchronize the audio signal and motion data without using timecode. Thus, the data processing system 1 does not need to prepare dedicated recording equipment to record each piece of data separately for each protocol used in multiple devices. Furthermore, the data processing system 1 does not need a timecode generator, cables, or interfaces for synchronizing the audio signal and motion data using timecode. Moreover, the data processing system 1 does not need to set frame rates for matching timecodes to each device.

[0059] The second channel shown in this embodiment stores motion data as a data column of digital audio signals, thus conforming to a defined multitrack audio data protocol (e.g., the Dante (registered trademark) protocol). Therefore, the second channel can be transmitted and played back as audio data, and can also be edited using audio data editing applications such as DAW (Digital Audio Workstation) by copying, cutting, pasting, or timing adjustments. For example, if a user uses DAW to cut and paste the audio data of the first and second channels included in a certain time period to different time periods, not only the audio data but also the motion data can be moved to different time periods without disrupting synchronization.

[0060] then, Figure 7 This is a block diagram illustrating the structure of a data processing system 1A in a reproducing environment. The data processing system 1A is, for example, installed in a venue used for remotely reproducing performances such as live concerts.

[0061] to and Figure 1 The data processing system 1A shown has the same structure and is given the same symbols, so the description is omitted. The data processing system 1A includes a display 13. The display 13 can be a panel-type display device such as an LCD, or an image display device such as a projector.

[0062] In addition, the data processing device 10 in the regenerated venue does not need to have the same hardware structure as the venue for events such as live performances.

[0063] Figure 8 This is a flowchart illustrating the operation of the processing unit 154 of the data processing apparatus 10 during playback. First, the processing unit 154 receives multitrack audio data (S21). The multitrack audio data is received from the data processing apparatus 10 at the venue where the live performance is taking place, or from a server. Alternatively, the data processing apparatus 10 reads multitrack audio data related to past live performances stored in the flash memory 103. Alternatively, the data processing apparatus 10 reads multitrack audio data related to past live performances stored in other devices such as a server.

[0064] The processing unit 154 decodes the received multitrack audio data and reproduces digital audio signals and motion data (S22). For example, the processing unit 154 extracts the digital audio signals of channels 1 to 32, which serve as the first channel. The processing unit 154 outputs the extracted digital audio signals to the mixer 11 (S23). The mixer 11 outputs the received digital audio signals to audio equipment such as speakers to reproduce singing and playing sounds.

[0065] Furthermore, the processing unit 154 reads 8 bits of header information from each sample of channels 33-64, which serve as the second channel in the processing of S22, and extracts the main body of the 8-bit or 16-bit motion data. Additionally, the second channel may also include data related to the signal processing parameters and basic settings of the mixer 11. The processing unit may also extract this data related to the signal processing parameters and basic settings and output it to the mixer 11.

[0066] The mixer 11 receives signal processing parameters and setting data from the data processing device 10, and performs various signal processing on the received audio signal based on these parameters and setting data. As a result, the mixer 11 reproduces singing and playing sounds in the same state as a live performance.

[0067] The processing unit 154 renders the image of the 3D model character based on the retrieved motion data (S24). Here, rendering means controlling the 3D model of the character based on the motion data and converting the 3D model into a two-dimensional image signal. Figure 5 As shown, the motion data includes skeletal data corresponding to the movable parts of the character, such as hands, feet, and fingers. The processing unit 154 determines the structure of the character's 3D model in each frame based on the position and rotation coordinates included in the skeletal data. The processing unit 154 generates an image of the character viewed from a specified perspective, based on methods such as ray tracing or scan lines.

[0068] The processing unit 154 outputs the generated image signal to the display 13 (S25). The display 13 displays an image based on the received image signal.

[0069] As described above, the data processing device 10 extracts and outputs the digital audio signal of the first channel of each sample in the regeneration venue, and extracts the motion data of the second channel of each sample to render the image of the character, so that even without using timecode, it can synchronously regenerate the image signal related to the digital audio signal and the image of the character.

[0070] Typically, video data has a larger data volume than audio data. Therefore, assuming that audio data is synchronized with video data by delaying them simultaneously, the audio data transmission will be delayed, thus reducing real-time performance. However, as an example, the motion data shown in this embodiment is at half the bit rate (403.2 kbps) of 48 kHz, 24-bit (768 kbps). Therefore, in the data processing system of this embodiment, digital audio signals and motion data can be stored within the same frame without reducing real-time performance. Furthermore, even assuming that motion data is stored across multiple samples, the time deviation of a single sample at a sampling frequency of 48 kHz converges to less than 1 msec. Therefore, users of the data processing system of this embodiment can obtain a new customer experience as described below: even if they are in a different venue than the live performance venue, they can perceive it as if they were attending a live performance.

[0071] Furthermore, for example, if the venue from which the transmission source is located is outdoors, and image data captured outdoors is then reproduced in an indoor venue, the user will experience a sense of incongruity. However, the data processing system of this embodiment transmits motion data; therefore, it is possible, for example, to use a camera located at the destination venue to capture images of the venue and composite the images of the characters into the images within that venue. Thus, the data processing system of this embodiment can reproduce images optimally suited to the reproduction environment. Consequently, the user can obtain a new customer experience as described below: a level of immersion that was previously impossible to achieve.

[0072] in addition, Figure 8 The processing of the output digital audio signal in S23 and the output processing of the video signal in S25 can be performed simultaneously or either one can be performed first.

[0073] (Variation Example 1)

[0074] The motion data in Variation Example 1 includes: motion data of the main body parts of the character, i.e., main motion data, and motion data of parts that are subordinate to the main body parts, i.e., sub-motion data.

[0075] The primary and secondary body parts of a character are at least the upper arm and forearm, which are the most important parts used to convey the performance. Secondary body parts, such as the torso, legs, hands, and fingers, are not directly required for conveying the performance. However, the distinction between primary and secondary body parts varies depending on the type of performance. For example, in guitar playing, the movements of the right forearm, left forearm, and fingers become important parts for conveying the performance; in piano playing, the forearms and movements of both arms become important parts.

[0076] The multitrack audio data of Variation 1 has: a first channel for storing digital audio signals, a second channel for storing main action data, and a third channel for storing sub-action channels.

[0077] The processing unit 154 of the data processing device 10 can also determine whether to regenerate the sub-motion data of the third channel during playback based on the processing capacity, the utilization rate of the CPU 104, or the available capacity of the RAM 105. Therefore, even when there is insufficient processing capacity, CPU 104 utilization rate, or available capacity of the RAM 105, the data processing device 10 will not stop or delay playback, thus not reducing real-time performance. Furthermore, even if the data processing device 10 does not regenerate the sub-motion data, it will still regenerate the main motion data of the reused parts used to convey the performance, thus not hindering the user's immersion.

[0078] (Variation Example 2)

[0079] The data processing device 10 involved in Variation Example 2 accepts the designation of the regenerated character and uses the retrieved motion data to render the image of the designated character.

[0080] For example, if the first venue, the destination, is in the Northern Hemisphere and it is winter, the data processing device 10 at the first venue will accept the designation of a character wearing winter clothing. On the other hand, if the second venue, the destination, is in the Southern Hemisphere and it is summer, the data processing device 10 at the second venue will accept the designation of a character wearing summer clothing.

[0081] Thus, the data processing apparatus 10 of Modified Example 2 is able to reproduce a more preferred image corresponding to the reproduction environment.

[0082] (Variation Example 3)

[0083] In the data processing apparatus 10 involved in Modification 3, in Modification 2, the extracted motion data is further corrected based on the specified character, and the corrected motion data is used to render the image of the specified character.

[0084] The size of 3D model data sometimes varies depending on the character. Therefore, the data processing device 10 corrects the length of each bone data included in the motion data based on the data of the specified character (e.g., data representing height).

[0085] Thus, the data processing apparatus 10 of Modified Example 3 is able to reproduce a more preferred image.

[0086] (Variation Example 4)

[0087] The data processing device 10 involved in Variation Example 4 accepts the designation of a background image and overlays the rendered image of the character onto the designated background image.

[0088] As described above, for example, if the venue from which the transmission source is located is outdoors, and the image data captured outdoors is reproduced in an indoor venue, the user will feel a sense of incongruity. However, in Modification 4, the data processing apparatus 10 overlays (composites) the rendered image of the character onto a specified background image. For example, if the venue for reproduction is indoors, the data processing apparatus 10 in Modification 4 overlays an indoor background image. Alternatively, the data processing apparatus 10 in Modification 4 can also overlay a summer background image when the venue for reproduction is in summer, and a winter background image when the venue for reproduction is in winter.

[0089] Therefore, the data processing apparatus 10 of Modification 4 is capable of reproducing preferred images corresponding to the reproducing environment. Thus, users can obtain a new customer experience as described below: they can actually feel an immersive experience that was previously impossible.

[0090] The description of this embodiment is illustrative in all respects and should not be considered limiting. The scope of the invention is not defined by the above embodiments, but by the claims. Furthermore, the scope of the invention includes the scope equivalent to the claims. For example, in the above embodiments, the program read by the CPU 104 constitutes the processing unit of the present invention, but the processing unit of the present invention can also be implemented by an FPGA (Field-Programmable Gate Array).

[0091] The data processing system shown in this embodiment can also be applied to systems that require combining images and sound, such as those from a theme park. Alternatively, the data processing system shown in this embodiment can also be applied to remote conversation systems. In remote conversations, transmitting not only sound but also the movements of performers or singers is crucial. However, as mentioned above, the amount of image data is greater than that of sound data; therefore, if sound and image data are transmitted separately, the latency of the image data will increase. Furthermore, if sound and image data are transmitted synchronously, real-time performance will be reduced. In contrast, the data processing system shown in this embodiment transmits and receives motion data that is smaller in size compared to image data, thus enabling the transmission of the movements of remote performers or singers with low latency without compromising real-time performance.

[0092] Explanation of reference numerals in the attached figures

[0093] 1: Data processing system; 1A: Data processing system; 10: Data processing device; 11: Mixer; 12: Motion capture device; 13: Display; 101: Display; 102: User I / F; 103: Flash memory; 104: CPU; 105: RAM; 106: Communication I / F; 154: Processing unit.

Claims

1. A data processing method for generating multitrack audio data, the multitrack audio data storing audio data from multiple channels, including at least a first channel and a second channel. The data column of the digital audio signal is stored in the first channel. In the second channel, information about the character's actions associated with the digital audio signal, i.e., action data, is stored as a data column of the digital audio signal.

2. The data processing method as described in claim 1, wherein, Transmit the multitrack audio data.

3. A data processing method that regenerates multitrack audio data containing multiple channels of audio data. The data stored in the first channel is reproduced as a digital audio signal. The data column stored in the second channel is extracted as motion data, which is information about the character's actions associated with the digital audio signal. The extracted motion data is then used to render the character's image.

4. The data processing method according to any one of claims 1 to 3, wherein, The second channel includes the identification information of the motion data.

5. The data processing method according to any one of claims 1 to 3, wherein, The motion data includes motion data for the main body parts of the character, i.e., main motion data, and motion data for parts that are subordinate to the main body parts, i.e., sub-motion data. The audio data also has a third channel, which stores the sub-action data.

6. The data processing method as described in claim 3, wherein, Accept the assignment of the role of regeneration. The extracted motion data is used to render the image of the specified character.

7. The data processing method as described in claim 6, wherein, Based on the specified role, the retrieved motion data is corrected. The image of the specified character is rendered using the corrected motion data.

8. The data processing method as described in claim 3, wherein, Accept the specified background image. The rendered image of the character is overlaid on the specified background image.

9. A data processing apparatus that generates multitrack audio data, the multitrack audio data storing audio data from multiple channels, including at least a first channel and a second channel. The data processing device includes a processing unit. The processing unit stores a data column of digital audio signals in the first channel. In the second channel, the processing unit stores the information about the character's actions, i.e., the action data, associated with the digital audio signal as a data column of the digital audio signal.

Citation Information

Patent Citations

  • Karaoke system

    JP2012078526A

  • Information processor and program

    JP2015047261A