Voice synchronization method and voice synchronization system
The voice synchronization method and system address synchronization errors due to processing time differences by calculating and adjusting playback times across multiple devices, ensuring synchronized audio output.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2026-03-18
AI Technical Summary
Existing voice synchronization methods fail to account for synchronization errors caused by differences in processing times between information terminals of different models, leading to deviations in audio playback across multiple devices.
A voice synchronization method and system that utilizes a master unit with a microphone to record and calculate time differences between slave units with speakers, adjusting playback times based on these differences to ensure synchronized audio output.
The system effectively suppresses synchronization discrepancies among multiple devices, enabling realistic and synchronized audio playback by correcting playback times based on calculated time lags.
Smart Images

Figure 2026049326000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a voice synchronization method and a voice synchronization system.
Background Art
[0002] Conventionally known music playback methods include a stereo playback method using 2ch speakers or a multi-channel playback method with multiple speakers arranged. Also, a streaming-type music playback method on an information terminal has been conventionally known. Furthermore, a music generation method in which a group is set for a plurality of speakers having a communication function to play the same music has been conventionally known.
[0003] For example, in a voice data playback system that transmits a playback command for playing voice data to each speaker, there is a possibility that a deviation may occur in the time when the playback command reaches each speaker device depending on the communication environment. Therefore, a technique for determining the playback start time in consideration of the arrival delay time has been conventionally known (see, for example, Patent Document 1).
[0004] Also, a technique for synthesizing a plurality of voice data without using a device capable of acquiring a reference time by calculating the time difference between devices and adjusting the time difference between the voice data to be synthesized based on the time difference has been conventionally known (see, for example, Patent Document 2).
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] When playing music using multiple information terminals, it is necessary to ensure synchronization of audio data across all terminals. For example, information terminals playing music will experience delays not only due to communication delays but also due to differences in processing time between terminals of different models. These differences in processing time between terminals are caused by differences in the processing time of the installed software, CPU (Central Processing Unit), or audio devices.
[0007] Thus, when playing music using multiple information terminals, synchronization errors occur not only due to communication delays but also due to delays resulting from differences in processing times between information terminals of different models. When playing music with a realistic feel using multiple information terminals, there is a problem in that synchronization errors in audio data due to differences in processing times between information terminals of different models cannot be ignored. It should be noted that the above-mentioned Patent Documents 1 and 2 do not solve these problems.
[0008] One embodiment of the present invention aims to provide a voice synchronization method and a voice synchronization system that can suppress synchronization errors when multiple information terminals are used to output voice data for playback. [Means for solving the problem]
[0009] One embodiment of the present invention is an audio synchronization method comprising: a first information terminal having a speaker emits synchronization audio data from the speaker; a first emission time recorded by the first information terminal using its own clock when it started the process of emitting the synchronization audio data from the speaker; a second information terminal having a microphone receives the synchronization audio data emitted from the speaker of the first information terminal with its microphone; a second emission time recorded by the second information terminal using its own clock when the speaker of the first information terminal emitted the synchronization audio data; a calculation of the amount of time difference from when the first information terminal started the process of emitting the synchronization audio data from the speaker to when the speaker emitted the synchronization audio data, based on the first emission time and the second emission time; and a synchronization of the playback audio data emitted from the speakers of a plurality of the first information terminals, each emitting playback audio data, based on the amount of time difference.
[0010] Furthermore, one embodiment of the present invention is a voice synchronization system having a plurality of first information terminals having speakers and a second information terminal having a microphone, comprising: a means for emitting synchronization voice data from the speaker of the first information terminal; a first recording means for recording a first emitting time when the first information terminal starts the process of emitting the synchronization voice data from the speaker using the clock of the first information terminal; a receiving means for receiving the synchronization voice data emitted from the speaker of the first information terminal with the microphone of the second information terminal; and the clock of the second information terminal The audio synchronization system includes: a second recording means that uses a buck to record a second speaking time when the speaker of the first information terminal speaks the synchronization audio data; a calculation means that calculates for each of the multiple first information terminals the amount of time difference from the start of the process of speaking the synchronization audio data from the speaker to the speaker speaking the synchronization audio data, based on the first speaking time and the second speaking time; and a synchronization means that synchronizes the playback audio data spoken by the multiple first information terminals that speak playback audio data from the speaker based on the amount of time difference. [Effects of the Invention]
[0011] This invention provides an audio synchronization method and system that can suppress synchronization discrepancies when multiple information terminals are used to output audio data for playback. [Brief explanation of the drawing]
[0012] [Figure 1] This is a diagram illustrating an example of the audio synchronization system according to this embodiment. [Figure 2] This is a functional configuration diagram of an example of the voice synchronization system according to this embodiment. [Figure 3] This flowchart shows an example of the processing procedure for the audio synchronization method according to this embodiment. [Figure 4] This is an explanatory diagram illustrating an example of the processing in steps S26 and S28. [Figure 5]This is an explanatory diagram illustrating an example of a multi-channel audio content distribution service to which the audio synchronization method according to this embodiment is applied. [Figure 6] This diagram illustrates an example of a new way to enjoy music, where individual handsets are assigned parts such as piano, vocals, drums, and bass, allowing users with handsets to play together as an ensemble. [Figure 7] This is an explanatory diagram illustrating an example of a large-scale event utilizing the audio synchronization method according to this embodiment. [Modes for carrying out the invention]
[0013] Next, embodiments of the present invention will be described in detail.
[0014] <System Configuration> The voice synchronization system 1 according to this embodiment has the configuration shown in Figure 1, for example. Figure 1 is a configuration diagram of an example of the voice synchronization system 1 according to this embodiment.
[0015] The voice synchronization system 1 shown in Figure 1 comprises a master unit 10, multiple slave units 12, a signaling server 14, and a sound source storage server 16. The master unit 10, the multiple slave units 12, the signaling server 14, and the sound source storage server 16 are connected via a communication network 18 to enable data communication. The communication network 18 is, for example, the Internet or a LAN (Local Area Network).
[0016] The master unit 10 is an example of a second information terminal having a microphone. The slave unit 12 is an example of a first information terminal having a speaker. The master unit 10 and the slave unit 12 are, for example, smartphones. The master unit 10 and the slave unit 12 may have an operation unit that receives operations from the user, a display unit that presents information to the user, and a reader unit that reads code information. For example, the code information is a barcode or a two-dimensional code. The reader unit is, for example, a camera. The camera may be a built-in camera or an external camera.
[0017] The master device 10 and the slave devices 12 may be, for example, tablet terminals, game terminals, or PCs (personal computers), etc. Further, the master device 10 and the slave devices 12 may be, for example, audio devices having a communication function, etc.
[0018] By performing the following-described synchronization process in cooperation, the master device 10 and the slave devices 12 suppress the synchronization deviation of the playback voice data caused by the difference in processing time among the slave devices 12 of different models, etc., and synchronize the playback voice data to be uttered from the speakers by the plurality of slave devices 12.
[0019] The signaling server 14 performs the process of P2P (peer-to-peer) communication connection for the master device 10 and the slave devices 12 to establish P2P communication. The master device 10 and the slave devices 12 that have established P2P communication perform messaging directly without going through a server such as the signaling server 14. The signaling server 14 is, for example, a PC or a workstation. The signaling server 14 may be realized by a cloud service.
[0020] The sound source storage server 16 stores the playback voice data to be uttered from the speakers by the plurality of slave devices 12. The sound source storage server 16 provides the playback voice data to the plurality of slave devices 12 (for example, the plurality of slave devices 12 participating in the utterance of the same playback voice data) that utter the same playback voice data from the speakers. The sound source storage server 16 is, for example, a PC or a workstation. The sound source storage server 16 may be realized by a cloud service.
[0021] Note that the configuration diagram of the voice synchronization system 1 shown in FIG. 1 is an example. For example, the classification of the signaling server 14 and the sound source storage server 16 is an example. The functions of the signaling server 14 and the sound source storage server 16 may be integrated or further subdivided.
[0022] <Hardware Configuration> 《Smartphone》 The master unit 10 and the slave unit 12 may be implemented by smartphones with the following hardware configuration. The smartphone has a display, a touch panel, a microphone, a speaker, and a camera. The smartphone also has a CPU, ROM (Read Only Memory), RAM (Random Access Memory), and a communication unit.
[0023] The CPU executes programs stored in ROM and controls various operations of the smartphone. RAM is used as the CPU's workspace. The CPU implements the various functions described below by executing programs.
[0024] The communication unit performs data communication via the communication network 18. The microphone converts received audio data, such as synchronization audio data, into electrical signals. The speaker converts audio data, such as synchronization audio data, into physical vibrations to produce (play) music or voice. The microphone and speaker process input and output of music or voice according to the control of the CPU.
[0025] A touch panel is an example of an operation unit that accepts various operations from the user. A display is an example of a display unit that presents various information to the user. A camera is an example of a reading unit that reads code information.
[0026] The above-described smartphone hardware configuration is merely an example, and it is not necessary to include all of the above-described components; it may also include components other than those described above. Furthermore, the hardware configurations of the master unit 10 and the slave unit 12 are not limited to the smartphone described above. The hardware configurations of the master unit 10 and the slave unit 12 may also include components of a tablet device, a game console, a PC, or an audio device with communication capabilities.
[0027] <Functional Configuration> The voice synchronization system 1 according to this embodiment is implemented with the functional configuration shown in Figure 2, for example. Figure 2 is a functional configuration diagram of an example of the voice synchronization system 1 according to this embodiment. The functional configuration of the master unit 10 and slave unit 12 of the voice synchronization system 1 shown in Figure 2 is implemented by the CPU executing programs such as the OS (operating system) and applications. The functional configuration diagram shown in Figure 2 omits components that are unnecessary for the explanation of this embodiment as appropriate.
[0028] The master unit 10 has a second messaging unit 20, a synchronization processing permission unit 22, a receiving unit 24, a second recording unit 26, and an instruction / command unit 28. The slave unit 12 has a first messaging unit 30, a synchronization processing request unit 32, a synchronization voice data generation unit 34, a first recording unit 36, a calculation unit 38, an instruction / command reception unit 40, and a synchronization unit 42.
[0029] The second messaging unit 20 of the master unit 10 accesses the signaling server 14 and processes a P2P communication connection to establish P2P communication with multiple slave units 12. For example, the second messaging unit 20 creates a group, such as a room, on the signaling server 14 to group and manage multiple slave units 12 that participate in the output of the same audio data for playback, according to the operation of the user operating the master unit 10.
[0030] The second messaging unit 20 displays code information for multiple slave units 12 to join the created group. The code information may be a two-dimensional code such as a QR code (registered trademark). The code information may be displayed on a screen of a large display or projector connected to the master unit 10. The code information may also be displayed on paper printed by a printing device.
[0031] The first messaging unit 30 of the slave unit 12 performs P2P communication connection processing to establish P2P communication with the master unit 10. For example, the first messaging unit 30 reads the code information of the group to join according to the operation of the user operating the slave unit 12. By reading the code information of the group to join, the first messaging unit 30 accesses the signaling server 14 and establishes P2P communication with the master unit 10 that created the group to join.
[0032] The master unit 10 that created the group and the multiple slave units 12 that join that group can communicate with each other by establishing P2P communication. Furthermore, a slave unit 12 that has established P2P communication with the master unit 10 will display, for example, a page for the group it is joining (e.g., a web page). This group page includes a button (e.g., a "Join" button) to request the start of the synchronization process described later.
[0033] The synchronization processing request unit 32 of the slave unit 12 requests the master unit 10 to start synchronization processing via messaging. Upon receiving the request to start synchronization processing, the synchronization processing permission unit 22 of the master unit 10 sends permission or denial of the start of synchronization processing to the slave unit 12 via messaging, according to the operation of the user operating the master unit 10. Upon receiving the request to start synchronization processing, the master unit 10 displays a group page that has buttons (for example, an "Approve" button or a "Do not approve" button) to select whether to allow or deny the start of synchronization processing from the slave unit 12.
[0034] The synchronization processing permission unit 22 of the master unit 10 sends a message to the slave unit 12 granting or denying permission to start synchronization processing. The synchronization processing permission unit 22 of the master unit 10 may also automatically send permission to start synchronization processing from the slave unit 12 if it is possible to start synchronization processing with the slave unit 12 that requested the start of synchronization processing. For example, if the synchronization processing permission unit 22 of the master unit 10 is currently in the process of synchronizing with another slave unit 12, it may send a message denying permission to start synchronization processing to a slave unit 12 that later requested the start of synchronization processing.
[0035] Upon receiving permission from the master unit 10 to begin the synchronization process, the synchronization voice data output unit 34 of the slave unit 12 outputs synchronization voice data with a predetermined signal from its own speaker. The first recording unit 36 of the slave unit 12 also records the first output time, which is the time when the process of outputting the synchronization voice data from its own speaker began, using its own clock. The first output time is the time (based on the slave unit 12's clock) when the synchronization voice data output unit 34 received the request to output the synchronization voice data from its own speaker.
[0036] The first voice output time is the time when the synchronization voice data output unit 34 receives a request to output the synchronization voice data from its own speaker, and therefore is not affected by the processing time of the process performed based on that request. In other words, the first voice output time is not affected by differences in processing times of slave units 12 of different models.
[0037] The receiver 24 of the master unit 10 receives the synchronization audio data emitted from the speaker of the slave unit 12 using its own microphone. The second recording unit 26 of the master unit 10 calculates the time when it received the synchronization audio data using its own clock and records it as the second emission time when the speaker of the slave unit 12 emitted the synchronization audio data. The second emission time is the time when the speaker of the slave unit 12 emitted the synchronization audio data (a time based on the clock of the master unit 10). Therefore, the second emission time is affected by the processing time of the slave unit 12, which is performed in response to the request to emit the synchronization audio data from the speaker. In other words, the second emission time is affected by the difference in processing time of slave units 12, such as different models. However, considering the purpose of this embodiment, which is to synchronize playback audio data emitted from speakers among multiple slave units 12, the processing time of the master unit 10 can be considered constant and can be ignored.
[0038] The calculation unit 38 of the slave unit 12 receives the second voice output time from the master unit 10 via messaging. Based on the first voice output time (the time based on the clock of the slave unit 12 when it started the process of outputting synchronization voice data from its own speaker) and the second voice output time (the time based on the clock of the master unit 10 corresponding to the timing when the synchronization voice data was actually output from the speaker of the slave unit 12), the calculation unit 38 calculates the difference (temporal correspondence) between the first voice output time and the second voice output time.
[0039] The difference between the first speaking time and the second speaking time is the amount of time lag from the start of the process of emitting the synchronization audio data from the speaker of the unit to the speaker emitting the synchronization audio data. Hereinafter, the difference between the first speaking time and the second speaking time will be referred to as the amount of time lag. The amount of time lag from the start of the process of emitting the synchronization audio data from the speaker of the unit to the speaker emitting the synchronization audio data represents the delay (delay amount) from the time the synchronization audio data emitting unit 34 receives the request to process the synchronization audio data from the speaker to the time the synchronization audio data is emitted from the speaker. In this embodiment, even if the clocks of the master unit 10 and the multiple slave units 12 are not in a temporal correspondence, and the clocks of the master unit 10 and the multiple slave units 12 are running in parallel at different times, the playback audio data to be emitted from the speaker can be synchronized among the multiple slave units 12.
[0040] The instruction and command unit 28 of the master unit 10 sends instructions to the multiple slave units 12 via messaging to play back audio data from the speakers of each slave unit 12, in accordance with the user's operation of the master unit 10. The instructions and commands that the instruction and command unit 28 sends to the multiple slave units 12 may specify the audio data to be played back that the slave units 12 should play back.
[0041] Furthermore, the instruction receiving unit 40 of the slave unit 12 receives instruction commands from the master unit 10 to play back audio data via messaging. When the playback audio data acquired from the sound source storage server 16 is played back by the unit's speaker, the synchronization unit 42 of the slave unit 12 synchronizes the playback audio data to be played back by the speaker among multiple slave units 12 based on the amount of time difference calculated by the calculation unit 38.
[0042] For example, the synchronization unit 42 synchronizes the playback audio data emitted from the speakers of multiple slave units 12 by correcting the start time of the playback audio data based on the amount of time difference calculated by the calculation unit 38. The synchronization unit 42 also corrects the start time of the playback audio data based on the amount of time difference calculated by the calculation unit 38 when the playback audio data is sound source data having multiple parts and the sound source data of the part assigned by the master unit 10 is emitted from the speaker, so that the sound source data of the part assigned by the master unit 10 is emitted from the speaker with the same sense of presence as an orchestra.
[0043] The synchronization unit 42 of the slave unit 12 can suppress synchronization errors when multiple slave units 12 are used to output audio data for playback from their speakers. As a result, the audio data for playback output from the speakers of multiple slave units 12 are synchronized, enabling realistic music playback.
[0044] <Processing> Figure 3 is a flowchart showing an example of the processing procedure for the audio synchronization method according to this embodiment.
[0045] In step S10, the master unit 10 and the slave unit 12 establish P2P communication via the signaling server 14. By establishing P2P communication, the master unit 10 and the slave unit 12 can communicate with each other.
[0046] In step S12, the slave unit 12 requests the master unit 10 to start the synchronization process via messaging.
[0047] In step S14, the master unit 10, via messaging, grants permission to the slave unit 12 that requested the start of the synchronization process in step S12 to begin the synchronization process. Note that the synchronization process performed by the voice synchronization system 1 according to this embodiment requires the slave unit 12 to emit synchronization voice data through its speaker and receive it with the master unit 10's microphone; therefore, permission to start the synchronization process is not granted to multiple slave units 12 simultaneously.
[0048] In step S16, the slave unit 12, having received permission to start the synchronization process, emits synchronization audio data with a predetermined signal from its own speaker. The synchronization audio data can be various audio data with a predetermined signal that can be identified by analysis by the master unit 10.
[0049] In step S18, the slave unit 12 records the first time it started the process of emitting synchronization audio data from its own speaker using its own clock.
[0050] In step S20, the master unit 10 receives the synchronization audio data emitted from the speaker of the slave unit 12 using its own microphone.
[0051] In step S22, the master unit 10 analyzes the audio data received by the microphone and calculates the time when it received the synchronization audio data with its own microphone (for example, the time when it began receiving the synchronization audio data with its own microphone) using its own clock, and records this as the second speaking time when the speaker of the slave unit 12 emitted the synchronization audio data.
[0052] In step S24, the slave unit 12 receives a second speech time from the master unit 10 via messaging.
[0053] In step S26, the slave unit 12 calculates the amount of time lag from when it starts the process of emitting the synchronization audio data from its own speaker until the speaker emits the synchronization audio data, based on the first utterance time recorded in step S18 and the second utterance time recorded by the master unit 10 in step S22.
[0054] The processing in steps S10 to S26 is performed for each of the multiple slave units 12 that participate in the output of the same audio data for playback.
[0055] In step S28, the master unit 10, in accordance with the user's operation, sends a command to multiple slave units 12 participating in the playback of the same audio data via messaging to instruct them to play the audio data. When the multiple slave units 12 receive the command to play the audio data from the master unit 10 and play the audio data from their own speakers, they synchronize the audio data to be played from their speakers based on the amount of time difference calculated by the processing in steps S10 to S26.
[0056] The processes in steps S26 and S28 will be further explained using Figure 4. Figure 4 is an explanatory diagram of an example of the processes in steps S26 and S28.
[0057] Figure 4(A) shows an explanatory diagram of an example of the process in step S26. The time difference between slave units A to C shown in Figure 4(A) represents the delay based on the difference in processing time between slave units 12, such as different models. The difference in processing time between slave units 12, such as different models, is caused by differences in the processing time of the software, CPU, or audio device installed in the slave unit 12.
[0058] Therefore, when slave units A to C begin the process of emitting synchronization audio data from their speakers in accordance with the instructions received from the master unit 10, a synchronization mismatch in the synchronization audio data may occur based on the time difference shown in Figure 4(A) (the difference in processing time between slave units A to C, such as different models).
[0059] Figure 4(B) shows an explanatory diagram of an example of the process in step S28. Sub-units A to C can suppress the synchronization delay of the playback audio data emitted from the speakers of sub-units A to C by adjusting the start time of the process of emitting synchronization audio data from the speakers of sub-units A to C based on the amount of time difference calculated by the processes in steps S10 to S26.
[0060] For example, the time difference between slave units A to C shown in Figure 4(A) increases in the order of slave unit A, slave unit C, and slave unit B. In the voice synchronization method according to this embodiment, by starting the processing time earlier for slave unit 12, which has a larger time difference, the start time of the playback audio data output from the speaker can be synchronized.
[0061] <Examples> The audio synchronization system 1 according to this embodiment can be applied to various fields such as multi-channel music playback, virtual concerts, or large-scale events.
[0062] Figure 5 is an explanatory diagram of an example of a multi-channel audio content distribution service to which the audio synchronization method according to this embodiment is applied. Figure 5 shows an implementation example that operates on the browsers of the master unit 10 and multiple slave units 12. In the music distribution platform, multi-channel audio data is stored in the audio storage server 16.
[0063] The audio synchronization method according to this embodiment allows multiple slave units 12 to suppress synchronization discrepancies when playing the same audio source data from their speakers. Multiple slave units 12 can synchronize and play the audio for each part of the audio source data distributed from the audio source storage server 16. Users can enjoy immersive audio played from the speakers of multiple slave units 12. Furthermore, as shown in Figure 5, by using streaming technology in conjunction, a performance environment like an online virtual concert can be virtually reproduced in each home.
[0064] Furthermore, in a multi-channel audio content distribution service to which the audio synchronization method according to this embodiment shown in Figure 5 is applied, instead of the existing music experience of simply playing synthesized stereo sound and listening alone, it is possible to provide a new way to enjoy music, as shown in Figure 6, for example, by assigning separate elements that make up the music, such as piano sound, vocal sound, drum sound, and bass sound to each individual slave unit 12, and allowing users with slave units 12 to play together in ensemble.
[0065] Furthermore, the audio synchronization method according to this embodiment may be used in large-scale events such as those shown in Figure 7. Figure 7 is an explanatory diagram of an example of a large-scale event utilizing the audio synchronization method according to this embodiment.
[0066] In large-scale events such as music events at live music venues, each audience member can use their smartphone or other device 12 to synchronize and play audio data, providing a musical experience in which the audience participates in the performance by the musicians.
[0067] For example, the audio synchronization system 1 according to this embodiment can be used for events using approximately 20 to 30 smartphones. By displaying code information, such as a QR code, on a large screen for smartphones to participate in the event, and allowing users to scan this code information, the audio synchronization method according to this embodiment can create a new type of music event in which many users can easily participate. The new type of music event created by the audio synchronization method according to this embodiment may be provided to event organizers as content for a new music experience mechanism. A large-scale event in which many users can easily participate may be a so-called flash mob. A flash mob is an act in which a group of people suddenly perform a dance or other performance without prior notice, attracting the attention of those around them before dispersing.
[0068] Furthermore, the audio synchronization system 1 according to this embodiment can synchronize audio data emitted from speakers of tablet terminals, PCs, or audio equipment with communication functions, similar to a smartphone. By utilizing this characteristic, the audio synchronization system 1 according to this embodiment can wirelessly realize a multi-channel audio system in large spaces such as outdoors or in complex facilities. Conventional multi-channel audio systems were based on the premise of high costs due to wired connections, associated special equipment, and wiring. The audio synchronization system 1 according to this embodiment can be replaced with an inexpensive wireless system (e.g., a wireless LAN system). It is also possible to introduce audio systems to sound events or complex facilities using a similar mechanism.
[0069] (summary) In this embodiment, the synchronization audio data emitted from the speaker of the slave unit 12 is received by the microphone of the master unit 10. The delay between the start of the process of emitting the synchronization audio data from the speaker of the slave unit 12 and the actual emission of the synchronization audio from the speaker is calculated as the time lag. This time lag will vary due to differences in the models or performance of the slave units 12. Therefore, in this embodiment, the playback audio data emitted from the speakers of multiple slave units 12 is synchronized based on the calculated time lag.
[0070] According to this embodiment, even if there are multiple slave units 12 with varying processing times for outputting audio data from the speaker due to differences in performance, for example, it is possible to suppress synchronization delays in music playback from the multiple slave units 12 and enable realistic music playback.
[0071] The present invention is not limited to the embodiments specifically disclosed above, and various modifications and changes are possible without departing from the scope of the claims. [Explanation of Symbols]
[0072] 1. Voice synchronization system 10 Master unit 12 Handset 14 Signaling Servers 16. Audio source storage server 24 Receiving Unit 26 Second Record Section 34. Voice output unit for synchronization audio data 36. First Record Section 38 Calculation Section 42 Classmates
Claims
1. The first information terminal having a speaker emits synchronization audio data from the speaker, The first information terminal records a first speaking time when it starts the process of speaking the synchronization audio data from the speaker using its own clock, The second information terminal having a microphone receives the synchronization audio data spoken from the speaker of the first information terminal using the microphone, The second information terminal uses its own clock to record the second speaking time when the speaker of the first information terminal spoke the synchronization audio data, A step of calculating the amount of time difference from when the first information terminal starts the process of outputting the synchronization audio data from the speaker based on the first and second output times until the speaker outputs the synchronization audio data, The steps include: multiple first information terminals that emit audio data for playback synchronize the audio data to be emitted from the speaker based on the amount of time difference; Cutter audio synchronization method.
2. The steps include establishing P2P (peer-to-peer) communication that enables messaging between the first information terminal and the second information terminal, The first information terminal requests the second information terminal to start the time difference calculation process via messaging, The second information terminal grants permission to the first information terminal to start the time difference calculation process via messaging, It further possesses, The first information terminal, after obtaining permission from the second information terminal, shall output the synchronization audio data from the speaker. The audio synchronization method according to claim 1, characterized by the following:
3. The second information terminal displays code information for participating in the output of the audio data for playback, The first information terminal reads the code information, It further possesses, The first information terminal establishes the P2P communication by reading the code information and participates in the output of the audio data for playback. The audio synchronization method according to claim 2, characterized by the following:
4. The aforementioned audio data for playback is sound source data having multiple parts, The second information terminal performs the step of assigning the part to be spoken from the speaker of the first information terminal, The audio synchronization method according to claim 1, further comprising:
5. When the multiple first information terminals that emit the audio data for playback emit the sound source data having multiple parts from their speakers based on a request from a second information terminal, they correct the start time of the sound source data emission based on the amount of time difference and emit the sound source data of the part assigned by the second information terminal from their speakers. The audio synchronization method according to claim 4.
6. The first and second information terminals are smartphones, tablet devices, PCs (personal computers), or audio devices with communication capabilities. The audio synchronization method according to any one of claims 1 to 5.
7. A voice synchronization system comprising a plurality of first information terminals having speakers and a second information terminal having a microphone, A means for outputting synchronization audio data from the speaker of the first information terminal, A first recording means that records the first time the first information terminal starts the process of emitting the synchronization audio data from the speaker using the clock of the first information terminal, A receiving means for receiving the synchronization audio data emitted from the speaker of the first information terminal with the microphone of the second information terminal, A second recording means that uses the clock of the second information terminal to record the second speaking time when the speaker of the first information terminal spoke the synchronization audio data, A calculation means for calculating, for each of the first information terminals, the amount of time difference from the start of the process of emitting the synchronization audio data from the speaker to the speaker emitting the synchronization audio data, based on the first emitting time and the second emitting time, Multiple first information terminals that emit audio data for playback synchronize the audio data to be emitted from the speaker based on the amount of time difference, A voice synchronization system having
Citation Information
Patent Citations
Shank of rotary tool
JP1986050707A
Voice data reproduction system
JP2022041959A