Display method and program

The display method in DAWs, which includes showing character images and controlling sound generation parameters, addresses the lack of user presence in music production, achieving enhanced user engagement and immersion by synchronizing character operations with sound production.

JP2025084513APending Publication Date: 2025-06-03YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023198472
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing Digital Audio Workstations (DAWs) lack a method to effectively provide users with a sense of presence of characters during music production, especially when using sound synthesis techniques with external plug-ins.

Method used

A display method that includes showing a screen with operators for controlling sound generation parameters and a character image corresponding to a set character, starting sound signal generation in a timbre matching the character when a predetermined instruction is received, based on preset pitch information, character information, and sound generation parameter values.

Benefits of technology

This approach allows users to experience a heightened sense of presence during music production by displaying character images that perform operations corresponding to the production situation, enhancing user engagement and immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025084513000001_ABST
    Figure 2025084513000001_ABST
Patent Text Reader

Abstract

To give users a sense of realism of characters by displaying character images corresponding to tones so that they operate according to a production status when music is produced using sound generation software.SOLUTION: A display method includes displaying a screen including one or more operators for controlling sound generation parameters and a character image corresponding to a set character, and initiating generation of a sound signal in a tone corresponding to the character when a predetermined instruction is accepted. The generation of the sound signal is based on pre-set pitch information, character information, and values of the sound generation parameters. Displaying the screen includes displaying the character image to perform a first action while the sound signal is not being generated.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for displaying images.

Background Art

[0002] There is known a technique for improving expressiveness such as emotional expression by an object by synchronizing voice and the movement of the object.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] A DAW (Digital Audio Workstation) used in music production has a function as software for editing sound data and generating sounds according to the playback position.

[0005] When using a sound synthesis technique in a DAW, external functions such as plug-ins may be used. A character image is displayed on the screen for operating the plug-in to give the user a sense of presence.

[0006] One of the objects of the present invention is to give the user a sense of presence of a character by displaying a character image corresponding to a timbre so as to operate according to the production situation when producing music using software for generating sounds.

Means for Solving the Problems

[0007] A display method according to an embodiment includes displaying a screen including one or more operators for controlling sound generation parameters and a character image corresponding to a set character, and starting generation of a sound signal in a timbre corresponding to the character when a predetermined instruction is received. Generation of the sound signal is based on preset pitch information and character information and values of the sound generation parameters. Displaying the screen includes displaying the character image so as to perform a first operation while the sound signal is not being generated.

Advantages of the Invention

[0008] According to the present invention, when making music using software that generates sound, by displaying a character image corresponding to the timbre so as to operate according to the production situation, a sense of presence of the character can be given to the user.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Embodiments for Carrying Out the Invention

[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. The embodiments shown below are examples, and the present invention is not construed as being limited to these embodiments. In the drawings referred to in this embodiment, the same parts or parts having the same function are denoted by the same reference numerals or similar reference numerals, and the repeated description thereof may be omitted. The drawings may be schematically described with the dimensional ratios being different from the actual ratios or a part of the configuration being omitted from the drawings in order to clarify the description.

[0011] <Data Processing Device> The data processing device in one embodiment of the present invention is a device equipped with a computer, such as a desktop terminal, a notebook terminal, a smartphone, or a tablet terminal. The data processing device provides a DAW for the user to execute music editing and the like by executing the installed application program. The DAW generates a sound signal based on data set corresponding to a plurality of tracks, for example, information for generating sound such as sound control data and waveform data.

[0012] The sound control data is data including information for controlling the generated sound at a plurality of time steps according to the passage of time, for example, MIDI data and the like. The waveform data is data obtained by sampling the sound signal waveform, for example, WAV data, MP3 data, and the like. The DAW reads out the information corresponding to the playback position among the data of each track and generates a sound signal corresponding to the playback position. By synchronously generating sound signals corresponding to a plurality of tracks, the sound corresponding to each track is output as a mixed sound (mixed sound).

[0013] FIG. 1 is a block diagram showing a data processing apparatus according to an embodiment. The data processing apparatus 1 includes a control unit 11, a storage unit 13, a display unit 15, an operation unit 17, an interface 21, and a communication unit 23. The control unit 11, the storage unit 13, the display unit 15, the operation unit 17, the interface 21, and the communication unit 23 are connected to each other via a bus 19.

[0014] The control unit 11 is an example of a computer including a processor such as a CPU and a storage device such as a RAM. The control unit 11 executes the program 13a stored in the storage unit 13 using the CPU, and realizes functions related to a DAW such as processing related to a signal generation method and processing related to a display method in the data processing apparatus 1.

[0015] The display unit 15 is a display that displays a music editing screen. The operation unit 17 is a device for inputting a user's instruction, and is an operation device including an operator that outputs an operation signal corresponding to the input operation to the control unit 11. The operator is, for example, a slider or a knob button. The operator is provided as an operator image displayed on the display unit 15. The operation signal includes information regarding the instruction value of the operator.

[0016] The interface 21 includes a module for communicating with a microphone, a musical instrument, a speaker, etc. by wired communication or wireless communication such as infrared communication or short-range wireless communication. The communication unit 23 is a communication module that connects to a network under the control of the control unit 11 and communicates with an external device connected to the network.

[0017] The storage unit 13 is a storage device such as a non-volatile memory or a hard disk drive. The storage unit 13 stores the program 13a executed in the control unit 11, various data including music data 13b, voice synthesis data 13c, character image data 13d required in the DAW realized by this program 13a, and a learned model 13e.

[0018] Program 13a is downloaded from an external device via a network such as the Internet and installed in the data processing device 1 by being stored in the storage unit 13.

[0019] Music data 13b is stored in the storage unit 13. The music data 13b is data stored in the storage unit 13 for each piece of music and includes sound control data CDa. The music data 13b may further include sound control data CDb, waveform data WD, etc. These data may be generated in the process of music production by a DAW or may be provided to the data processing device 1 in advance in the same manner as the program 13a.

[0020] The sound control data CDa and CDb are, for example, data in MIDI format and include a plurality of types of parameters for controlling generated sounds at a plurality of time steps according to the passage of time. In this example, the sound control data CDa is data used when generating a singing synthesis sound. The sound control data CDb is data used when generating an instrument sound. The plurality of parameters in the sound control data CDa are classified into at least a parameter (first parameter P1) belonging to a first group and a parameter (second parameter P2) belonging to a second group. The first parameter P1 is treated as a parameter corresponding to future information, which is information read after a read position TR described later, and the second parameter P2 is treated as a parameter corresponding to current information, which is information read at the read position TR. The second parameter P2 may include parameters related to past information, which is information read before the read position TR.

[0021] The first parameter P1 is information determined for each sound defined in time series, and includes, for example, character information, pitch information, duration information, intensity information, etc. The first parameter P1 includes at least character information and pitch information. The character information includes information for generating characters such as lyrics for singing as sounds corresponding to each sound. When the sound control data CDa is data for generating instrument sounds, the character information may not be used, or instead of lyrics or the like, other information, for example, phoneme information may be used as information for reproducing the playing method of the instrument. The pitch information includes information for specifying the pitch corresponding to each sound. The duration information includes information for specifying the start position and the duration (note value) of the sound corresponding to each sound. The duration information may be defined by a plurality of information such as note on and note off. The intensity information includes information for specifying the velocity corresponding to each sound.

[0022] The second parameter P2 includes information for controlling the generated sound at each time step, and is also referred to as a sound generation parameter. The second parameter P2 is, for example, power information, transpose information, formant information, etc. The power information includes information for specifying the dynamics of the sound. For example, when the set value becomes larger, a shouting-like sound of the voice is reproduced, and when the set value becomes smaller, a whispering-like sound of the voice is reproduced. The transpose information includes information for raising or lowering the pitch of the sound. The formant information includes information for specifying the timbre. For example, when the set value becomes larger, a sound closer to a male voice is reproduced, and when it becomes smaller, a sound closer to a female voice is reproduced. The second parameter P2 may be information for specifying the transparency, brightness, breath, mouth opening method, etc. of the voice. In other words, the second parameter P2 is a parameter for adjusting the timbre.

[0023] The set value of the information included in the second parameter P2 may be directly used as a value (control value) for controlling the generated sound, or may be updated or corrected according to the instruction information based on the instruction value of the operator and used as the control value.

[0024] A plurality of parameters in the sound control data CDb includes at least the parameters belonging to the third group. The parameters of the third group include the parameters corresponding to the current information, and hereinafter are referred to as the third parameter P3. The third parameter P3 may include the parameters related to the past information.

[0025] The third parameter P3 includes the information determined for each sound defined in the time series and the information for controlling the generated sound at each time step. The former information is the information used as the current information or the past information. The latter information is the same as the second parameter P2 described above. The third parameter P3 is, for example, pitch information, duration information, intensity information, transpose information, etc.

[0026] The waveform data WD is the data obtained by sampling the sound signal waveform and indicates the waveform data corresponding to each time step.

[0027] The voice synthesis data 13c may be downloaded from an external device via a network and stored in the storage unit 13, or may be provided in a state recorded on a non-transitory computer-readable recording medium. The voice synthesis data 13c includes a plurality of types of parameters for generating a singing synthesis sound of a desired timbre. The voice synthesis data 13c is provided to the learned model 13e described later.

[0028] Similarly, the character image data 13d may be downloaded from an external device via a network and stored in the storage unit 13, or may be provided in a state of being recorded on a non-transitory computer-readable recording medium. The character image data 13d is data for displaying a character image 900, which will be described later, on the screen displayed on the display unit 15. The character image data 13d includes still images and a plurality of moving images corresponding to one or more characters, and each of the plurality of moving images corresponds to a plurality of operation patterns of the character. The character corresponds to a predetermined timbre generated based on the above-described speech synthesis data 13c. In other words, the character image data 13d is stored in association with the speech synthesis data 13c.

[0029] The learned model 13e is a statistical estimation model used when generating a singing synthesis sound. For example, a known machine learning model using a DNN can be used. The machine learning model may be a machine learning model using, for example, a CNN (Convolutional Neural Network), an RNN (Recurrent Neural Network), or the like. When the speech synthesis data 13c and the sound control data CDa are input, the learned model 13e outputs a sound signal according to the values of a plurality of parameters included in the speech synthesis data 13c and the values of the first parameter P1 and the second parameter P2 included in the sound control data CDa. Note that the sound synthesis process using such a learned model 31e can be realized, for example, by the configuration described in International Publication No. 2021 / 251364 (particularly the second embodiment), and can be implemented by those skilled in the art.

[0030] <Music Editing Screen> The music editing screen will be described. The music editing screen is a screen provided by a DAW realized in the data processing device 1.

[0031] FIG. 2 is an example of a music editing screen in an embodiment. In the example shown in FIG. 2, the music editing screen WS includes a track area AT, an editing area AE, and an operation image area AS. In all cases, the horizontal direction indicates the temporal position, i.e., the time step. When the playback process is being executed, an image GT indicating the playback position is displayed in the track area AT and the editing area AE.

[0032] The track area AT is an area for displaying an image corresponding to data for each track. In this example, it includes areas AT1, AT2, AT3, and AT4 corresponding to the first to fourth tracks. Sound control data CDa is assigned to the first track, sound control data CDb is assigned to the second and third tracks, and waveform data WD is assigned to the fourth track. Accordingly, an image corresponding to the first parameter P1 in the sound control data CDa is displayed in the area AT1 (the first area). Images corresponding to the third parameter P3 in the sound control data CDb are displayed in the areas AT2 and AT3 (the third area).

[0033] The editing area AE includes, in this example, a piano roll area AP and a set value area AC. The piano roll area AP is an area (the first area) where information on the first parameter P1 is displayed when the sound control data CDa is assigned to the track selected in the track area AT, and a sound control image corresponding to the first parameter P1 is displayed. The piano roll area AP is an area (the third area) where information on the third parameter P3 is displayed when the sound control data CDb is assigned to the track selected in the track area AT, and a sound control image corresponding to the third parameter P3 is displayed.

[0034] The sound control image is, for example, an image GL indicating character information and an image GR indicating pitch information and duration information. Other information such as intensity information may be further represented by the color of the image GR, the width of the band, etc. The image GR includes information determined for each sound defined in time series, includes the information of the first parameter P1 if it is the sound control data CDa, and includes the information of the third parameter P3 if it is the sound control data CDb. When waveform data WD is assigned to the track selected in the track area AT, the piano roll area AP is replaced by an area for displaying the waveform, which is the editing area AE.

[0035] The setting value area AC (second area) displays an image GP indicating the setting value of the information for controlling the generated sound at each time step. That is, the image GP includes the information of the second parameter P2 if it is the sound control data CDa, and includes the information of the third parameter P3 if it is the sound control data CDb. In the setting value area AC, an image related to the information of the first parameter P1 that is not displayed in the editing area AE may be displayed in conjunction with the editing area AE.

[0036] When the user instructs a change operation on the image via the operation unit 17 or the interface 21 and changes the position of the image according to the information of each data such as the image GR and the image GP on the music editing screen WS, the information of the corresponding data is changed. Therefore, the image displayed in the piano roll area AP can also be said to be an image related to the change of the first parameter P1 or the third parameter P3. The image displayed in the setting value area AC can also be said to be an image related to the change of the second parameter P2 or the third parameter P3.

[0037] In this example, the operation image area AS includes operation images GC1, GC2, and GC3 for receiving the user's instructions. The operation image GC1 is an image imitating a button for receiving instructions such as play start and play end. The operation image GC2 is an image imitating a button for receiving the setting of the moving speed of the play position, that is, the play speed. The operation image GC3 is an image imitating a slider for receiving the setting of the volume.

[0038] <Sound output unit> The function of the sound output unit realized by the control unit 11 executing the program 13a will be described.

[0039] FIG. 3 is a functional block diagram showing the sound output unit in one embodiment. The sound output unit 1000 includes a first sound generation unit 100a, a second sound generation unit 100b, a third sound generation unit 100c, a data editing unit 200, a read position determination unit 300, and a synthesis unit 500.

[0040] The data editing unit 200 edits the information of the sound control data CDa and CDb according to the user's instruction and updates the data. The read position determination unit 300 generates a read position TR according to the reproduction instruction and the reproduction speed. The reproduction speed may be a value set by the operation image GC2 or a value set in the sound control data CDa and CDb. In this example, the read position TR is represented by the position of the time step.

[0041] The first sound generation unit 100a generates a sound signal Wa corresponding to the track to which the sound control data CDa is assigned. The first sound generation unit 100a includes the learned model 13e shown in FIG. 1. The second sound generation unit 100b generates a sound signal Wb corresponding to the track to which the sound control data CDb is assigned. The third sound generation unit 100c generates a sound signal Wc corresponding to the track to which the waveform data WD is assigned. The first to third sound generation units 100a, 100b, and 100c read data according to the read position TR and generate the respective sound signals Wa, Wb, and Wc. In this way, as the reproduction position advances and the read position TR also advances, the sound signals Wa, Wb, and Wc corresponding to the reproduction position are output. When the reproduction position advances according to the read position TR, the position of the image GT displayed on the music editing screen WS described above also moves according to the read position TR.

[0042] Since the audio signal Wa contains information on the first parameter P1, it contains information after the playback position, that is, future information. Since the audio signal Wb contains information on the third parameter P3, it contains information at the playback position or before the playback position, that is, current information or past information. The synthesizing unit 500 mixes the audio signals Wa, Wb, and Wc generated in each track and outputs a synthesized audio signal Wm.

[0043] When generating only the singing voice synthesis sound, in the sound output unit 1000, the second sound generation unit 100b and the third sound generation unit 100c may be omitted.

[0044] <Audio Signal Generation Method> The audio signal generation method executed in the control unit 11 will be described. The audio signal generation method described here starts when the program 13a is executed. Here, the overall flow of the DAW will be described.

[0045] FIG. 4 is a flowchart showing a signal generation method in an embodiment. The control unit 11 displays a music editing screen WS (step S10). As described above, the music editing screen WS includes a track area AT and an editing area AE. The track area AT displays an image corresponding to the first parameter for controlling the generated sound if the track is assigned the sound control data CDa, and displays an image corresponding to the third parameter for controlling the generated sound if the track is assigned the sound control data CDb.

[0046] As described above, the editing area AE includes a piano roll area AP and a set value area AC. If it is a track to which the sound control data CDa is assigned, the piano roll area AP is used as an area for changing the information of the first parameter P1 by displaying an image related to the change of the first parameter P1. If it is a track to which the sound control data CDa is assigned, the set value area AC is used as an area for changing the information of the second parameter P2 by displaying an image related to the change of the second parameter P2. The editing window described later is also an example of displaying an image related to the change of the second parameter P2. If it is a track to which the sound control data CDb is assigned, the piano roll area AP and the set value area AC are used as areas for changing the information of the third parameter P3 by displaying an image related to the change of the third parameter P3.

[0047] The control unit 11 generates the music data 13b by editing the information of various parameters according to the user's instruction (step S20). The control unit 11 waits for a reproduction start instruction from the user (step S30; No). When receiving the reproduction start instruction (step S30; Yes), the control unit 11 determines the reading position TR so as to advance the reproduction position for each time step (step S40). When the reading position TR reaches the end position of the music data 13b (step S50; Yes), the control unit 11 waits for the reproduction start instruction again (step S30; No).

[0048] When the reading position TR has not reached the end position of the music data 13b (step S50; No), the control unit 11 generates a sound signal corresponding to each track (step S60). In step S60, the sound signal is generated based on the preset pitch information and character information (first parameter P1) and the sound generation parameter, that is, the value of the second parameter P2. Next, the sound signals corresponding to each track are mixed to generate and output a mixed sound signal Wm (step S70), and the process returns to step S40 to determine the reading position TR with the reproduction position advanced.

[0049] <Editing Window> In the music editing screen WS shown in FIG. 2, an editing window that can be displayed by inputting a predetermined operation will be described. The editing window is a screen provided by the DAW as a user interface.

[0050] FIG. 5 is an example of a music editing screen in an embodiment. In the music editing screen WS shown in FIG. 5, an editing window WE is displayed under the control of the control unit 11. The editing window WE can be displayed for each track. Therefore, a plurality of editing windows WE may be displayed on the music editing screen WS.

[0051] In this example, the editing window WE includes operation images GC4, GC5, GC6, GC7, GC8, and GC9. The operation image GC4 is an image imitating a button for transposing keys. The operation image GC5 is an image imitating a slider for changing various set values from a reference value. The operation image GC6 is an image imitating a slider for finely adjusting the sounding timing. The operation image GC7 is an image imitating a button for adjusting the pitch (tuning). The operation image GC8 is an image imitating a button for setting the tone color. For example, by operating the button of the operation image GC8, a desired voice bank can be set. The operation image GC9 is an image imitating a button for loading an external sequence file. The images imitating the buttons or sliders of each of the operation images GC4, GC5, and GC7 function as operators for controlling the second parameter P2 when the user clicks the button or slides the slider on the editing window WE. These operation images GC4, GC5, and GC7 are images related to the change of the second parameter P2 and are used as areas for changing the information of the second parameter P2.

[0052] The editing window WE further includes a setting image 800 and a character image 900. The setting image 800 includes an image related to the second parameter P2. The character image 900 is an image corresponding to a predetermined character, and the character is determined based on the set tone color. The tone color can be set by operating the button of the operation image GC8.

[0053] In the editing window WE, while no sound signal is being generated, a character image 900 performing a first operation is displayed, and while a sound signal is being generated, a character image 900 performing a second operation different from the first operation is displayed. Each of the first operation and the second operation includes one or more preset operations, and each operation includes one or more preset operation patterns.

[0054] The first operation is an operation performed while no sound signal is being generated, for example, an operation indicating that the character is in a waiting state. The state where no sound signal is being generated includes cases where the playback state is paused, stopped, or the playback has ended, or cases where the sound signal Wa generated from the information included in the first parameter P1 input to the first sound generation unit is silent while the playback state is in progress. Examples of the first operation include an operation of blinking, an operation of moving the face, an operation of changing the body orientation, an operation of moving a part of the body, etc. When the first operation is an "operation of blinking", the operation pattern includes, for example, blinking once, blinking twice continuously, etc. Also, when the first operation is an "operation of moving the face", the operation pattern includes nodding the head up and down, shaking the head left and right, tilting the face, etc. Also, when the first operation is an "operation of changing the body orientation", the operation pattern includes turning to the right from the front, turning to the front from the right, making a turn, etc. Also, when the first operation is an "operation of moving at least a part of the body", the operation pattern includes raising both hands, floating in place, jumping, etc. One or more operation patterns are selected as the first operation from one or more operation patterns.

[0055] In FIG. 6, as a first operation, a character image 900 that blinks is shown as an example. The first operation and / or the operation pattern of the first operation are determined based on the value of a second parameter P2 (sound generation parameter) while no sound signal is being generated. For example, the first operation may be determined based on a preset predetermined value of the second parameter P2, or may be determined based on the value of each piece of information of the second parameter. The first operation and / or the operation pattern of the first operation can be switched by the user changing the value of the second parameter P2 while no sound signal is being generated. Also, a character image 900 that performs one or more first operations may be displayed in the editing window WE. For example, while no sound signal is being generated, a character image 900 that performs one first operation may be displayed in the editing window WE, or a character image 900 that performs a plurality of first operations simultaneously may be displayed in the editing window WE.

[0056] The second operation is an operation that is performed while a sound signal is being generated. For example, it indicates that the character is in a state of performing or singing, and is an operation that changes over time. The period while a sound signal is being generated is a case where at least the playback state is in playback or audition, and there is a sound signal Wa corresponding to the read position TR. The second operation includes, for example, an operation of singing, an operation of dancing, an operation of changing facial expressions, a moving operation, and the like. When the second operation is an "operation of singing", the operation pattern includes opening the mouth to sing, singing while closing the mouth, and the like. When the second operation is an "operation of dancing", the operation pattern includes stepping, jumping, and the like. When the second operation pattern is an "operation of changing facial expressions", the operation pattern includes looking happy, looking sad, looking painful, and the like. When the second operation is a "moving operation", the operation pattern includes a running operation, a walking operation, and the like.

[0057] The second operation and / or the operation pattern of the second operation can be determined based on the value of a predetermined parameter at the reading position TR, the set values of the first parameter P1 and the second parameter P2, and the like. For example, when there is character information at the reading position TR, an operation of singing may be selected as the second operation. Also, the second operation and the operation pattern of the second operation may be selected by the user. In this case, the user may set the conditions for displaying the character image 900 that performs the desired operation pattern by setting each parameter value when the character image 900 performs a predetermined operation.

[0058] The character image 900 that performs the second operation can be displayed based on the first parameter P1. For example, the character image 900 that performs the second operation is displayed based on the character information included in the first parameter P1. Specifically, when the character image 900 that performs the singing operation is displayed, as shown in FIG. 7, the shape of the character's mouth is controlled based on the character information. That is, the shape of the character's mouth changes based on the character information. In FIG. 7, the shape of the mouth of the character pronouncing the vowels (a, e, i, o, u) is shown as an example. In FIG. 7, only the upper body of the character image 900 is shown, and the lower body side is omitted.

[0059] Also, for example, the character image 900 that performs the second operation is displayed based on the pitch information included in the first parameter P1. Specifically, when the character image 900 performs an operation of changing the facial expression, the face of the character image 900 is displayed based on the set character and the pitch information included in the first parameter P1. For example, when an appropriate pitch range for the set character is preset, the singing operation is selected as the second operation, and when the pitch information moves from an appropriate pitch range for the character to a pitch range outside the appropriate pitch range (inappropriate pitch range), the character image 900 changes from a happy expression to a sad expression as shown in FIG. 8. At this time, the face color of the character image 900 may change. Conversely, when the pitch information moves from an inappropriate pitch range for the character to an appropriate pitch range, the character image 900 changes from a sad expression to a happy expression. Note that the appropriate pitch range is a pitch range in which a sound signal generated using the character sounds natural or comfortable to hear, for example, in terms of sound perception, and may also be expressed as a recommended pitch range or a proficient pitch range. On the other hand, the inappropriate pitch range is a pitch range in which a sound signal generated using the character sounds unnatural, uncomfortable, or stimulating to hear, and may also be expressed as a non-recommended pitch range or an unproficient pitch range.

[0060] Also, the character image 900 that performs the second operation can be displayed based on the value of a second parameter P2 (sound generation parameter), which can be changed while a sound signal is being generated. For example, when the user makes the set value of the power information larger than a predetermined value, the size of the mouth of the character image 900 performing the singing operation is displayed larger than a preset size. Conversely, when the user makes the set value of the power information smaller than a predetermined value, the size of the mouth of the character image 900 performing the singing operation is displayed smaller than a preset size.

[0061] In addition, the character image 900 that performs the second operation can be displayed based on the playback speed. Specifically, when the character image 900 performing a dancing motion is displayed, the dancing speed of the character may change based on the playback speed of the sound signal. Also, when the character image 900 performing a moving motion is displayed, the speed of the moving motion of the character image 900 may change based on the playback speed of the sound signal. Further, the operation pattern of the second operation may be switched based on the playback speed of the sound signal. For example, when the character image 900 performing a moving motion is displayed, as shown in FIG. 9, the running motion and the walking motion of the character may be switched based on the playback speed of the sound signal. That is, when the playback speed of the sound signal is equal to or lower than a predetermined speed, the character image 900 performing a walking motion is displayed, and when the playback speed of the sound signal exceeds the predetermined speed, the character image 900 performing a running motion is displayed. Also, the speed of the running motion or the walking motion of the character image 900 may change according to the change in the playback speed.

[0062] As described above, in this embodiment, even while the sound signal is not being generated, the character image 900 that performs the first operation is displayed in the editing window WE. Further, the character image 900 that performs the first operation changes the first operation and / or the operation pattern of the first operation in response to a change in the value of the second parameter P2 while the sound signal is not being generated. By displaying the character image 900 that performs the first operation not only while the sound signal is being generated but also while the sound signal is not being generated, a sense of presence of the character can be given to the user.

[0063] <Display method> A method for displaying the character image 900, which is realized by the control unit 11 executing the program 13a, will be described. The display method described here starts when the editing window WE is displayed.

[0064] FIG. 10 is a flowchart showing a display method in one embodiment. The control unit 11 determines whether or not a sound signal is being generated (step S80). Here, the process of generating the sound signal corresponds to step S60 shown in FIG. 3. When no sound signal is being generated (step 80; No), the control unit 11 acquires the second parameter P2 (step S90), and displays a character image 900 that performs a first operation in the editing window WE based on the value of the acquired second parameter P2 (step S100). After the process of step S100, the process returns to the process of step 80.

[0065] When a sound signal is being generated (step 80; Yes), the control unit 11 acquires the first parameter P1, the second parameter P2, and the set playback speed (step 110), and displays a character image 900 that performs a second operation in the editing window WE based on the acquired first parameter P1, second parameter P2, and playback speed (step S120). After the process of step S120, the process returns to the process of step 80.

[0066] <Modification Example> The present disclosure is not limited to the above-described embodiments, and includes various other modification examples. For example, the above-described embodiments have been described in detail for easy understanding of the present disclosure, and are not necessarily limited to those having all the configurations described. Also, other configurations may be added to the configuration of one embodiment, some configurations may be deleted, or some configurations may be replaced with other configurations. Some modification examples will be described below.

[0067] (1) In the editing window WE, the user can select whether to display or hide the character image 900. For example, when the character image 900 is set to be displayed in the editing window WE, the user can hide the character image 900 in the editing window WE by performing a predetermined operation on the character image 900 displayed in the editing window WE. Here, the predetermined operation is not limited, but includes clicking, long-pressing, and swiping the character image 900 on the editing window WE, etc. Conversely, when the hidden character image 900 is to be displayed again on the editing window WE, the character image 900 may be displayed again in the editing window WE by performing a predetermined operation in the space where the character image 900 was originally displayed.

[0068] (2) In the above-described embodiment, the character image 900 is set by operating the button of the operation image GC8 to set a predetermined tone color, and the character corresponding to the set tone color is selected. However, according to the music data, the character automatically displayed as the character image 900 may be set. For example, the character corresponding to the tone color set in a predetermined track of the music data may be selected. Also, a corresponding character may be set for the music data. Further, the character may be selected based on a predetermined algorithm at a predetermined timing (for example, when the editing window WE is displayed, when the program is started, when a playback instruction is received, etc.).

[0069] (3) In the above-described embodiment, the character image 900 is displayed in the editing window WE. However, the screen on which the character image 900 is displayed is not limited to the editing window WE. For example, the character image 900 may be displayed on the music editing screen WS. Also, a screen on which only the character image 900 is displayed may be provided.

[0070] (4) In the above-described embodiment, regarding the character image 900 that performs the first operation and is displayed in the editing window WE while no sound signal is being generated, it has been described that the first operation and / or the operation pattern of the first operation are determined based on the value of the second parameter P2 while no sound signal is being generated. However, the pattern of the first operation may be determined based on the first parameter P1. For example, when the first parameter P1 is acquired while no sound signal is being generated and the pitch information included in the acquired first parameter P1 includes a pitch in a sound range inappropriate for the character, the facial expression of the displayed character image 900 may be a painful expression.

Description of Reference Numerals

[0071] 1: Data processing device, 11: Control unit, 13: Storage unit, 13a: Program, 13b: Music data, 13c: Sound synthesis data, 13d: Character image data, 13e: Trained model, 15: Display unit, 17: Operation unit, 19: Bus, 21: Interface, 23: Communication unit, 200: Data editing unit, 300: Reading position determination unit, 500: Synthesis unit, 800: Setting image, 900: Character image, 1000: Sound output unit

Claims

1. displaying a screen including one or more operators for controlling sound generation parameters and a character image corresponding to a set character; starting generation of a sound signal in a timbre corresponding to the character upon receiving a predetermined instruction; comprising; the generation of the sound signal is based on preset pitch information and character information and values of the sound generation parameters; the displaying of the screen includes, while the sound signal is not being generated, displaying the character image so as to perform a first operation, a display method.

2. the first operation includes a plurality of patterns that perform different operations; the displaying of the screen includes controlling to switch the pattern of the first operation in response to a change in the sound generation parameters, the display method according to claim 1.

3. the displaying of the screen further includes, while the sound signal is being generated, displaying the character image so as to perform a second operation different from the first operation, the display method according to claim 1.

4. the character image performing the second operation is displayed based on the character information, the display method according to claim 3.

5. the character image performing the second operation is displayed based on the character and the pitch information, the display method according to claim 3.

6. the character image performing the second operation is displayed based on values of the sound generation parameters that can be changed during the generation of the sound signal, the display method according to claim 3.

7. the sound signal is further generated based on a predetermined playback speed; the character image performing the second operation is displayed based on the playback speed, the display method according to claim 3.

8. the display method according to claim 1 further includes switching the character image on the screen to be non-displayed upon receiving a predetermined operation on the character image.

9. the display method according to claim 1 further includes providing an interface for setting the character.

10. a program for causing a computer to execute the display method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Voice synchronization processor, voice synchronization processing program, voice synchronization processing method, and voice synchronization system

    JP2015148932A