System for producing content including sound, computer-implemented method, and computer program
The system addresses the challenge of interpreting timbre by visually representing it with graphics, facilitating easier sound creation and communication by allowing users to understand and compare timbre spatially.
Patent Information
- Application Number
- JP2024118433
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Existing sound creation systems struggle with the difficulty of interpreting and comparing timbre due to its transient nature, making it challenging for users, especially beginners, to achieve a desired timbre without a clear spatial representation.
A system that visually represents timbre using graphics, such as speech bubbles, allowing users to associate sound data with corresponding graphics, enabling visual understanding and comparison of timbre through spatial representation.
Enables users to grasp and compare timbre visually, facilitating easier sound creation and communication of timbre impressions, even for those who cannot hear the sound, by providing a common understanding through graphic representation.
Smart Images

Figure 2026017618000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a production system, a computer-implemented method, and a computer program for content including sound such as music. [Background technology]
[0002] Patent Document 1 discloses a message processing device that processes the display of speech bubbles containing messages. This message processing device displays speech bubbles in a form corresponding to the message, and expresses information related to emotions and voice expressions accompanying the message in the speech bubbles. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6434799 Summary of the Invention
[0004] The inventors came up with the idea that if tone could be visually understood and interpreted in a production environment for content containing sound such as music (for example, music pieces, videos containing sound, etc.), it would be easier to handle sound data.
[0005] Therefore, technology that allows us to visually understand and interpret timbre is desirable.
[0006] One aspect of the present disclosure is a system for producing content including sound data, the system comprising a user interface displayed to a user to produce content including the sound data, the user interface configured to display a graphic having a form corresponding to the tone of the sound data.
[0007] Another aspect of the present disclosure is a method, the method being a computer-implemented method executed by a computer, that includes displaying, in a user interface for content including sound data, a graphic having a form corresponding to a tone of the sound data.
[0008] Another aspect of the present disclosure is a computer program that causes a computer to execute processing including displaying, in a user interface for content including sound data, a graphic having a form corresponding to the tone of the sound data.
[0009] Further details will be described in the following embodiments. [Brief explanation of the drawings]
[0010] [Figure 1] Figure 1 is a diagram showing the configuration of the production system. [Figure 2] FIG. 2 is a diagram illustrating an example of a user interface. [Figure 3] Figure 3 shows an activity diagram of the user and the production system. [Figure 4] FIG. 4 is a diagram showing another example of a user interface. [Figure 5] FIG. 5 is a diagram showing another example of a user interface. [Figure 6] FIG. 6 is a diagram showing another example of a user interface. [Figure 7] FIG. 7 is a flowchart of the graphic generation process. [Figure 8] FIG. 8 is a diagram showing an estimation model. [Figure 9] FIG. 9 is a diagram showing a group of models for estimating a figure. [Figure 10] FIG. 10 is a diagram showing another example of the user interface. DETAILED DESCRIPTION OF THE INVENTION
[0011] <1. Overview of the system for producing sound content, computer implementation method, and computer program>
[0012] (1) A system according to an embodiment may be a system for producing content including sound data. The system according to an embodiment may include a user interface that is displayed to a user to allow the user to produce content including sound data. The user interface may be configured to display a graphic having a shape corresponding to the timbre of the sound data. In this case, a user using the system for producing content including sound data can visually understand and interpret the timbre of the sound data by using the graphic having a shape corresponding to the timbre.
[0013] (2) It is preferable that the graphic includes a speech bubble. Speech bubbles have long been used in manga, and for many people, they visualize sound symbols and allow for interpretation of impressions from visual information. Therefore, by including a speech bubble in the graphic, the tone can be appropriately expressed.
[0014] (3) The sound data may include instrument sound data. The system according to the embodiment can, for example, graphically indicate the difference in timbre between a plurality of different instrument sound data for the same instrument, making it easier for the user to visually understand the difference in timbre between the plurality of different instrument sound data for the same instrument.
[0015] The user interface may be configured to display the graphic together with text associated with the sound data, such as the name of the sound data, such that the combination of text and graphic allows a user to visually understand the sound data more easily.
[0016] (4) The user interface may be configured to display characters related to the sound data in the graphic.
[0017] (5) A system according to an embodiment may include a processor that generates the graphic from the sound data and stores the sound data and the generated graphic in a memory in association with each other. The user interface may be configured to display the graphic associated with the sound data in the memory as a graphic indicating the tone of the sound data.
[0018] (6) It is preferable that the graphics generated by the processor be editable by the user, so that the graphics generated by the processor can reflect the user's sense of tone.
[0019] (7) Generating the figure from the sound data may include using an estimation model to estimate the figure from the sound data.
[0020] (8) Preferably, the estimation model is provided for each type of sound, which can improve the accuracy of estimation.
[0021] (9) Generating the figure from the sound data may include estimating the type of the figure and the parameters of the figure using a first estimation model that estimates the type of the figure from the sound data and a second estimation model that estimates parameters of the figure for the type of the figure estimated by the first estimation model from the sound data. Such two-stage estimation can improve the accuracy of estimation.
[0022] A method according to an embodiment may be a computer-implemented method executed by a computer, and may include displaying, in a user interface relating to content including sound data, a graphic having a shape corresponding to a timbre of the sound data.
[0023] A computer program according to an embodiment can cause a computer to execute processing including displaying, in a user interface relating to content including sound data, a graphic having a form corresponding to the tone of the sound data.
[0024] <2. Examples of a system for producing sound content, a computer implementation method, and a computer program>
[0025] Hereinafter, the embodiments will be described in more detail with reference to the drawings.
[0026] FIG. 1 shows a production system 100 according to an embodiment. The production system 100 provides a user with a production environment for content including sound data. The production system 100 according to an embodiment displays a graphic representing the timbre of the sound data handled in the content production environment. This graphic can visualize and represent the timbre of the sound data. The production system 100 according to an embodiment allows a user to visually understand and interpret the timbre without hearing the sound.
[0027] Here, sound can be divided into, for example, human voices and sounds other than human voices. Sounds other than human voices include, for example, the sounds of musical instruments, sounds emitted by other devices, or sounds generated in nature (environmental sounds). Content that includes sound is, for example, music or videos that include sound. Music is, for example, vocal music, instrumental music, or orchestral music. Videos that include sound are, for example, movies.
[0028] The production system 100 according to the embodiment preferably handles sound data that includes at least sounds other than human voices, such as sound data of musical instruments or environmental sounds.
[0029] The production system 100 according to the embodiment may be, for example, a music production system such as a digital audio workstation. The music production system is used for recording, editing, and playing music. The music production system allows editing (sound creation) by selecting multiple sound sources and adjusting the acoustic characteristics of each sound. In many cases, a piece of music is created by combining multiple sound data, such as sound data for multiple instruments, and by performing editing tasks such as adjusting the acoustic characteristics of each sound.
[0030] The production system 100 according to the embodiment may be a video production system that includes sound. In such a video production system, editing can be performed, such as adjusting the acoustic characteristics of one or more sounds included in a video.
[0031] The production system 100 can be configured with one or more computers. The production system 100 may be configured with a stand-alone computer, or may be configured with computers that configure a server and a client connected via a network.
[0032] 1, the computer constituting the production system 100 may include a processor 110 and a memory 120 connected to the processor 110. The computer constituting the production system 100 includes an input / output interface .
[0033] The processor 110 is, for example, a CPU, a GPU, or another type of processor. The memory 120 may include, for example, a primary storage device and a secondary storage device. The primary storage device is, for example, a RAM. The secondary storage device is, for example, a hard disk drive (HDD) or a solid-state drive (SSD). The memory 120 may include a computer program 125 executed by the processor 110. The processor 110 reads and executes the computer program 125 stored in the memory 120. The computer program 125 has program code representing instructions for causing a computer to function as the production system 100 according to the embodiment.
[0034] The music production system 100 according to the embodiment has a function for handling graphics that represent the timbre of sound data. This function may include, for example, displaying, editing, or storing graphics. This function may be realized by a computer executing a computer program 125. A program module for realizing the function for handling graphics that represent the timbre of sound data may be pre-installed in the computer program 125, or may be provided as a plug-in program for expanding the functionality of a computer program 125 that does not already have this function.
[0035] Furthermore, various data 121, 122, 123, and 124 may be stored in the memory 120. Fig. 1 exemplarily shows sound data 121, graphics 122 (graphic data 122), an inference model group 123 (parameters of each model), and training data 124 as data that may be stored in the memory 120. Details of these data 121, 122, 123, and 124 will be described later.
[0036] The production system 100 may include a display 140. The display 140 may be connected to the processor 110 via the input / output interface 130. The display 140 outputs, for example, a screen that serves as a user interface 200. FIG. 2 shows an example of the screen that serves as the user interface 200. The user interface 200 is displayed so that a user can create content that includes sound data. The screen that serves as the user interface 200 displays data necessary for editing sound, etc., and can accept editing operations by the user. In other words, the sound data is displayed on the user interface 200 and edited on the user interface 200.
[0037] The production system 100 may include an input device 150. The input device 150 may be connected to the processor 110 via the input / output interface 130. The input device 150 may be, for example, a mouse, a keyboard, a microphone, or other input device. A user of the production system 100 (e.g., a sound creator) can use the input device 150 to perform editing (sound creation) operations such as adjusting the acoustic characteristics of each sound while referring to the screen of the user interface 200.
[0038] In the production environment provided by the production system 100, the user expresses and produces a tone color through the process of sound source selection and sound creation. Here, editing acoustic features (also called "acoustic parameters") to express and produce a tone color is called "sound creation."
[0039] When creating a sound, for example, various types of effectors are combined with the tone of the instrument being used, and the task is to come closer to the desired tone from among countless combinations.
[0040] Here, sound is made up of three elements: "pitch (= interval)," "loudness," and "timbre." Pitch is determined by the frequency of the sound. Loudness is determined by the amplitude of the sound wave (sound pressure). Timbre refers to the feeling and color of sound that is unique to a sound-producing body, such as a musical instrument or the human voice, and is determined by the waveform of the sound wave. Even if the sound pressure and frequency are the same, if the waveform is different, the timbre will be different.
[0041] In sound creation, achieving a desired timbre is a difficult task for beginner musicians who do not have a clear impression of timbre. This problem arises from the fact that timbre is a temporal medium and has a transient nature. Because of the transient nature of timbre, it is difficult to evaluate the relative information, and it is difficult to interpret timbre by listening alone. This problem makes it difficult for people to acquire a sense of timbre.
[0042] In the sound creation process, to bring a transient sound closer to the desired timbre, the sound creator (user) must continually memorize the sound information created through trial and error, such as adjusting acoustic parameters. Unless users associate and memorize the timbre with the acoustic parameter information, it is difficult to reproduce previously created timbres. However, because there is no conventional way to represent timbre itself as spatial information like musical notes on a staff, it is not possible to compare multiple sound data created during the sound creation process at a glance. This means that users must memorize and process a lot of information during the trial and error process. While adjusting acoustic parameters, clarifying the user's desired timbre requires the user to mentally model the relationships between multiple sound data that are successively created and disappear. This makes sound creation a difficult task for music beginners who lack sufficient metacognitive awareness of timbre.
[0043] Therefore, in this embodiment, tone, which is a temporal medium, is expressed as a spatial medium, making it possible to visually grasp the meaning and impression of tone (figuration of tone; visualization of tone). In this embodiment, by using a figure representing tone (a figure that visualizes tone), it becomes possible to visually grasp the meaning and impression of tone.
[0044] The bouba-kiki effect is known as a problem that deals with the relationship between speech sounds and shapes. To verify the bouba-kiki effect, an experiment was conducted to demonstrate the perceived similarity between specific sounds or words and the shapes of figures. In this experiment, participants were presented with two randomly selected shapes and asked to choose the name they felt best associated with each, either "bouba" or "kiki." The results confirmed that many people tended to associate rounded shapes with "bouba" and angular shapes with "kiki." This is thought to be due to the fact that the shape of the mouth during articulation and vocalization evokes the meaning and image of the spoken word.
[0045] This phenomenon, in which a sound evokes a specific image or sensation, is called sound symbolism. Onomatopoeia is a typical example of sound symbolism. Onomatopoeia is a way of imitating and verbalizing sounds, noises, actions, emotions, and other objects. For example, the sound of a dog barking is expressed as "woof woof," and the sound of a quiet, sad cry is expressed as "sob sob." Onomatopoeia is commonly understood by speakers of the same language and is frequently used in everyday conversation.
[0046] Manga also uses a variety of expressive techniques to communicate the emotions of characters and the situation within the manga to readers. Speech bubbles, one of the expressive techniques used in manga, allow the shape of the emotions and impressions contained in the characters' lines to be expressed. In other words, the different shapes of speech bubbles are used to express the intentions and impressions contained in the acoustic information of the speech (i.e., paralanguage). The use of speech bubbles in manga can be seen as an example of visualizing the sound symbols of speech sounds and is commonly understood by the majority of readers.
[0047] The example of speech bubbles in comics, which visualize sound symbols and allow for interpretation of impressions from visual information, gave us the idea that it might be possible to visualize not only speech sounds but also the timbre of other sounds, such as musical instruments. If timbre information, which has a transient nature, could be visualized and spatially recorded, it could be used as a symbol for timbre retrieval. However, it has not been clear whether a common understanding exists among many people between specific acoustic features related to timbre and the shape of a figure, as in the bouba-kiki effect, even for timbres such as musical instruments that do not involve the shape of the mouth.
[0048] Therefore, the inventors hypothesized that "visualizing timbre using a speech bubble-like graphic will enable many people to gain a common understanding of timbre through visual information," and conducted an experiment to verify the validity of recording the timbre of musical instruments and the like using graphics such as speech bubbles. The results of the experiment suggested that the visibility of the graphic affected the interpretation of the impression of timbre. Therefore, it became clear that by using a graphic representing timbre (a graphic visualizing timbre) as in this embodiment, it is possible to visually grasp the meaning and impression of timbre.
[0049] The experiment investigated the correspondence between the subjects' impressions of four types of tones, including three types of effectors with typical roles in tones (Overdrive (role: distortion), Chorus (modulation), and Feedback Delay (spatial)), as well as a sound source with no effects. Four types of figures were used, including spiky speech bubbles representing shouts and anger. Results showed that almost all subjects showed a common tendency in their responses for the combinations of figures and sound sources. For example, it was shown that there is a strong correlation between the impression of a "distorted" tone, such as Overdrive, and spiky figures (see Nagata, Taishi, and Yamanishi, Yoshinori, "Verifying the Possibility of Visualizing Tones Using Speech Bubbles," Information Processing Society of Japan Research Report, Vol. 2024-MUS-139 No. 28, March 10, 2024).
[0050] The above experiment revealed new insight: many people share a common understanding of the relationship between timbre and graphics, for example, through their experiences with speech bubbles in comic books. This embodiment utilizes this new insight to enable visual understanding of the meaning and impression of a timbre using graphics representing the timbre. As a result, it becomes possible to understand and evaluate the meaning and impression of a timbre without listening to the sound. Representing timbre as a spatial medium makes it possible to list information about the timbre, such as the meaning and impression of the timbre, as well as the organization and parameter information of the effectors used. This solves challenges that are difficult for beginner musicians, such as comparing timbre information and understanding the relative characteristics of timbres. Therefore, sound creation becomes easier, for example, in music and other sound creation situations.
[0051] Furthermore, if the meaning and impression of a tone can be grasped visually, when content including sound is provided to a user, even if the user cannot hear the sound (for example, if the user is hearing impaired) or the sound is not played, the user can still recognize the tone by visually presenting the tone to the user. Therefore, if a graphic representing the tone is associated with sound data and the graphic corresponding to the sound data is presented to the content user (content listener), the user can grasp the tone without hearing the sound data.
[0052] In this way, this embodiment makes it possible to visualize timbre information in speech bubble shapes during sound creation, performance, and the like. Furthermore, by making it possible to save and visualize timbre impressions as symbols, it becomes easier to compare timbres and use different timbres for different pieces of music. Furthermore, by allowing the shape of the speech bubble to be freely manipulated using parameters, it becomes possible to visualize the impressions each user has of a timbre as graphical information. This not only serves as an interface for searching sound data, but also contributes to embodying the user's own sense of timbre, serving as a trigger for recalling timbres, and supporting communication between performers.
[0053] In this embodiment, the graphics mainly reflect timbre information, but they may also reflect sound information other than timbre information (for example, sound repetition such as delay).Furthermore, they may also reflect sound elements other than timbre, such as "pitch" or "loudness."
[0054] Fig. 2 shows an example of a user interface 200 that visualizes tones. As described above, the screen of the user interface 200 shown in Fig. 2 may be displayed on the display 140 of the production system 100, for example, to allow a user to produce content that includes sound data.
[0055] The screen of the user interface 200 shown in Fig. 2 includes a list display section 210 of one or more pieces of sound data that make up a certain sound (for example, one piece of music). The user selects sound data of one or more instruments, etc. that make up the music to be created, and displays the selected sound data in the list display section 210. The user can edit the sound by adjusting the acoustic parameters (such as effector parameters) of each piece of sound data displayed in the list display section 210.
[0056] The list display unit 210 may include, for example, an identification display unit 210A for each piece of sound data. The identification display unit 210A is a display that allows the user to visually identify each piece of sound data. As an example, the identification display unit 210A in FIG. 2 may display characters indicating the name of an instrument to indicate which instrument the sound data belongs to. In FIG. 2, the characters indicating the name of an instrument are, for example, "Drums," "Guitar A," "Guitar B," and "Bass."
[0057] The identification display unit 210A in FIG. 2 may display graphics indicating the timbre of each piece of sound data. In FIG. 2, graphics 211 and 212 indicating the timbre of the sound data are displayed for "Guitar A" and "Guitar B," respectively. In FIG. 2, graphics 211 and 212 indicating the timbre are displayed along with characters indicating the sound data, allowing the user to visually easily grasp the difference in timbre. For example, in FIG. 2, a sharp, thorn-shaped graphic 211 is displayed in association with "Guitar A," and a rounded, cloud-shaped graphic 212 is displayed in association with "Guitar B." This allows the user to easily understand that although "Guitar A" and "Guitar B" are both sound data of the same instrument, a guitar, they have different timbres. For example, the sharp graphic 211 of "Guitar A" gives the impression of an intense timbre, while the rounded graphic 212 of "Guitar B" gives the impression of a gentle timbre.
[0058] In the identification display section 210A of FIG. 2, no figures are associated with "Drums" and "Bass," but this simply indicates that the sound data has not yet been associated with figures; figures can also be associated with "Drums" and "Bass" by operating the interface 200.
[0059] The list display section 210 includes, for example, a piano roll display section 210B (sound information time change display section) for each sound data. The piano roll display section 210B displays the time change of sound information (performance information) of each sound data in a piano roll format. The piano roll display section 210B only shows the time change of pitch and tone, and does not visualize the tone color.
[0060] In the interface 200 of Fig. 2, areas 220, 230, 240, and 250 outside the list display section 210 are areas for displaying and operating selected sound data from among the sound data displayed in the list display section 210, and are operated for sound creation. In Fig. 2, "Guitar A" is selected, and in this case, areas 220, 230, 240, and 250 are areas for displaying and operating the sound data of "Guitar A."
[0061] 2 may include a designation unit 220 that designates the type of effector, and a setting unit 230 that sets the parameters (acoustic parameters) of the designated effector. The user can adjust the timbre of each sound data by operating the designation unit 220 and the setting unit 230.
[0062] The interface 200 of FIG. 2 includes a figure list display section 240 that displays a list of candidate figures (figure types) that can be associated with sound data. Here, multiple types of speech bubble figures 231, 232, and 233 are displayed as examples of candidate figures. The user can select a figure (figure type) that can be associated with sound data (here, the selected sound data of "Guitar A") from the figure list display section 240. The selected figure is stored in the sound production system 100 in association with the sound data (sound data of "Guitar A"). For example, as shown in the figure, sound data 121 and figure 122 are stored in memory 120 in association with each other.
[0063] The graphics associated with the sound data may include graphics other than speech bubbles. Speech bubbles are used to express dialogue in manga, and many people have a common understanding of the relationship between tone and graphics through their experience with speech bubbles in manga. Therefore, it is preferable that the graphics include speech bubbles.
[0064] The speech balloons 231, 232, 233 may include, for example, an oval or balloon-shaped speech balloon 231, a thorn-shaped speech balloon 232, and a cloud-shaped speech balloon 233. Note that speech balloons used in cartoons sometimes have protrusions called "horns" or "tails" that extend toward the speaker, but the figures associated with sound data may or may not have "horns" or "tails."
[0065] The interface 200 of Fig. 2 includes a graphic display unit 250 that displays a graphic 250A associated with sound data. The graphic display unit 250 displays a graphic 250A selected from the graphic list display unit 240 and associated with sound data. In Fig. 2, a thorn-shaped graphic 250A (corresponding to graphic 232) associated with the sound data of "Guitar A" is displayed. The system 100 may automatically display text related to the associated sound data (for example, "Guitar A" indicating the name of the selected sound data) within the displayed graphic 250A.
[0066] Graphic display unit 250 includes operation unit 250B for editing (adjusting) the shape of graphic 250A. With operation unit 250B in FIG. 2, the user can freely adjust the kurtosis and number of spikes in spike-shaped graphic 250A to obtain a graphic that appropriately expresses the tone color. Furthermore, graphic display unit 250 may also allow the user to enter and edit text within graphic 250A.
[0067] The information written in the figure associated with the sound data may be information such as the name of an instrument or singer (see figure 250A in FIG. 2), or may be the name of a part or the name of a piece of music. In other words, the figure may be a figure with a label such as text. Non-verbal information such as a photograph, illustration, or color may also be added to the figure. Note that the information does not have to be written inside the figure. The information may also be written outside the figure, as in the identification display section 210A in FIG. 2.
[0068] The setting and editing of a figure to be associated with sound data may be performed by a more detailed user interface than the figure display unit 250 shown in Fig. 2. Figs. 4 to 6 show examples of screens that serve as a user interface 300 for setting and editing a figure to be associated with sound data. The user interface 300 is also provided by the production system 100. Fig. 3 shows an example of the production system 100 and an activity between users that uses this user interface 300.
[0069] 3, a user can set effector parameters for sound data (step S301) and test play the sound data (step S302) when creating a sound using the production system 100. The system 100 stores the parameter-set sound data 121 in memory 120 (see FIG. 1) (step S303).
[0070] When the sound data 121 is saved, the system 100 displays a list of candidate graphic types associated with the sound data (step S304). For example, in the graphic type display section 310 shown on the screen of the user interface 300 in Fig. 4, three graphic types, "regular polygon," "spiky speech bubble," and "fluffy speech bubble," are displayed as candidates for selection by the user. Here, "regular polygon" indicates that the graphic type is a regular polygon, "spiky speech bubble" indicates that the graphic type is a thorn-shaped speech bubble, and "fluffy speech bubble" indicates that the graphic type is a cloud-shaped speech bubble.
[0071] The user selects a graphic type on interface 300 that gives the user an impression that matches the tone, and fine-tunes the shape of the graphic by setting the graphic parameters of the selected graphic type according to the tone (step S305). The processes from step S301 to step S305 can be executed repeatedly as needed for trial and error and re-creation by the user.
[0072] The user interface 300 includes a first editing unit 320 shown in FIG. 4, a second editing unit 330 shown in FIG. 5, and a third editing unit 340 shown in FIG.
[0073] When "regular polygon" is selected in the shape type display unit 310, the first editing unit 320 (regular polygon editing unit 320) in FIG. 4 is used by the user to freely adjust (adjust parameters) the shape of the "regular polygon" to match the tone. As an example, the first editing unit 320 in FIG. 4 includes a first adjustment unit 321 for the number of corners of the "regular polygon" (regular polygon) (e.g., 3 to 100), a second adjustment unit 322 for the color of the shape, and an input unit 323 for the text to be displayed in the shape. The first editing unit 320 also includes a display unit 324 for the edited shape.
[0074] When the "Speech bubble figure spiky" (spiky shape) is selected in the figure type display unit 310, the second editing unit 330 (spiky editing unit 330) in FIG. 5 is used by the user to freely adjust (adjust parameters) the shape of the spiky shape to match the tone. For example, the second editing unit 330 in FIG. 5 includes a first adjustment unit 331 for adjusting the length of the major axis (horizontal radius) of the entire figure, a second adjustment unit 332 for adjusting the length of the minor axis (vertical radius), a third adjustment unit 333 for adjusting the number of spiky points, a fourth adjustment unit 334 for adjusting the length of spiky points, a fifth adjustment unit 335 for adjusting the kurtosis of spiky points, a sixth adjustment unit 336 for adjusting the color of the figure, and an input unit 337 for inputting characters to be displayed in the figure. The second editing unit 330 also includes a display unit 338 for the edited figure.
[0075] When the "Speech bubble shape spiky" (cloud shape) is selected in the shape type display unit 310, the third editing unit 340 (fluffy editing unit 340) in FIG. 6 is used by the user to freely adjust (parameter-adjust) the shape of the cloud shape to match the tone. For example, the third editing unit 340 in FIG. 6 includes a first adjustment unit 341 for adjusting the length of the major axis of a large ellipse corresponding to the entire cloud shape, a second adjustment unit 342 for adjusting the length of the minor axis of the large ellipse, a third adjustment unit 343 for adjusting the number of elliptical points, a fourth adjustment unit 344 for adjusting the length of the major axis of a small ellipse included in the cloud shape, a fifth adjustment unit 345 for adjusting the length of the minor axis of the small ellipse, a sixth adjustment unit 346 for adjusting the color of the shape, and an input unit 347 for inputting characters to be displayed within the shape. The third editing unit 340 also includes a display unit 348 for the edited shape.
[0076] 4 to 6, the editing units 320, 330, and 340 can adjust not only the shape of a figure but also the form including color, so that a wider variety of emotions can be expressed by the figure. In addition, it may be possible to adjust the thickness of the lines that represent the shape of the figure and the line shape (solid line / dotted line, single line / multiple line, etc.).
[0077] Once the timbre of the sound data 121 and the corresponding graphic 122 are completed through trial and error or re-creation, the graphic 122 is stored in the memory 120 in association with the sound data 121 as an indication of the timbre of the sound data 121 (step S306; see FIG. 1). This enables the system 100 to display the graphic 122 indicating the timbre of the stored sound data 121.
[0078] In the examples of Figures 3 to 6, the type of figure and the parameters of the figure associated with the sound data are generated by the user's judgment. However, the generation of the figure associated with the sound data (e.g., the determination of the type of figure and the parameters of the figure) may be performed automatically by the system 100. Figure 7 shows an example of the procedure of the figure generation process executed by the system 100.
[0079] For automatic generation of graphics by the system 100, the system 100 uses a model 123 (see FIG. 1) for estimating an appropriate graphic from sound data. The model 123 is, for example, a machine learning model, but may also be a mathematical model that obtains an appropriate graphic from sound data using a preset calculation formula, or a table that associates acoustic features related to the timbre of sound data with graphics.
[0080] If the model 123 is a machine learning model, the machine learning model is obtained by machine learning using a combination of sound data (or its acoustic features) and a figure corresponding to the timbre of the sound data as training data 124 (see Figure 1).
[0081] 7, as shown in Figures 8(A) and 8(B), two types of estimation models, a first estimation model 123A and a second estimation model 123B, may be used as the model 123. Each of the first estimation model 123A and the second estimation model 123B may be a machine learning model.
[0082] As shown in Figure 8(A), when the first estimation model 123A receives the acoustic features of sound data (or the sound data itself), it estimates and outputs a graphic type (e.g., "regular polygon," "spiky," or "fluffy") corresponding to the timbre of the sound data.
[0083] As shown in Fig. 8(B), when the second estimation model 123B receives the acoustic features of sound data (or the sound data itself), it outputs graphic parameters corresponding to the timbre of the sound data. As shown in Fig. 9, in the system 100, a second estimation model is provided for each graphic type (e.g., "regular polygon," "spiky," and "fluffy"). The system 100 selects a second estimation model corresponding to the graphic type estimated by the first estimation model 123A, and estimates graphic parameters of that graphic type using the selected second estimation model.
[0084] The parameters for adjusting the shape of a figure may differ for each figure type. Therefore, the estimation model 123 is divided into a first estimation model 123A for estimating the figure type and a second estimation model 123B for each figure type. By performing a first-stage estimation of the figure type and a second-stage estimation of the parameters of that figure type, it is possible to appropriately estimate the figure parameters, which may differ for each figure type.
[0085] Furthermore, the estimation model 123 (for example, the first estimation model 123A and the second estimation model 123B) may be provided for each type of sound (for example, the type of musical instrument). By using an estimation model according to the type of sound, it is possible to estimate an appropriate figure according to the type of sound. FIG. 9 shows an example in which the first estimation model 123A and the second estimation model 123B are provided for each type of sound.
[0086] In FIG. 9, examples of sound types are shown as "guitar," "bass," "drums," "(human) singing," and "environmental sounds (telephone ringing, car sounds, etc.)," and a first estimation model and multiple second estimation models are associated with each sound type. For example, "guitar" is associated with "first estimation model 1A," and three second estimation models 2A-1, 2A-2, and 2A-3. Similarly, "bass," "drums," "singing," and "environmental sounds" are associated with first estimation models and second estimation models.
[0087] It should be noted that a common estimation model may be used regardless of the type of sound.
[0088] 7, in the figure generation process, acoustic features are extracted from sound data to which a figure is associated (step S701). The acoustic features are input to a first estimation model (step S702). The first estimation model outputs a figure type according to the acoustic features (step S703).
[0089] If a first estimation model or the like is provided for each type of sound, prior to step S702, the system 100 identifies the type of sound (e.g., type of musical instrument) from the sound data (or its acoustic features), or the user selects the type of sound. The system 100 identifies the type of sound by using, for example, a model that estimates the type of sound from the sound data (or its acoustic features). This model may also be a machine learning model. The system 100 selects a first estimation model according to the identified or selected type of sound (see FIG. 9).
[0090] The system 100 presents the graphic type estimated by the first estimation model as a candidate to the user (step S704). The presentation is performed, for example, on the user interface 200, 300. At this time, the system 100 may also present other graphic types as other candidates.
[0091] The user accepts the candidate graphic type presented by the system 100 or separately selects another candidate, and the system 100 then selects a second estimation model according to the graphic type accepted or separately selected by the user (step S705).
[0092] In the system 100, the acoustic features calculated in step S701 are input to the selected second estimation model (step S706). The second estimation model outputs graphic parameters corresponding to the acoustic features (step S707). The graphic parameters are graphic parameters of a graphic type approved or separately selected by the user.
[0093] The system 100 draws a graphic based on the graphic type approved or separately selected by the user and its graphic parameters, and presents it to the user (step S708). The presentation is performed, for example, on the user interface 200, 300. This allows the system 100 to automatically generate a graphic representing the timbre of the sound data and present it to the user.
[0094] The user can also fine-tune the shape of the presented figure by adjusting the figure parameters (step S709). By allowing the user to adjust the figure generated by system 100, the user's own sense of tone can be reflected based on the figure generated by the system.
[0095] The system 100 records the generated graphic (or the adjusted graphic if the user has made any adjustments) in memory in association with the sound data (step S710). The graphic may be recorded, for example, as image data representing the graphic, or as a graphic pattern and its graphic parameters. By recording the graphic in association with the sound data, the system 100 can display a graphic representing the tone color of the sound data on the interface 200.
[0096] Furthermore, the system 100 may record the correspondence between the sound data and the figure adjusted by the user as training data for machine learning of the inference model (step S711). For example, the system 100 may record as training data a correspondence between a figure type and a figure parameter of the figure type and an acoustic feature of the corresponding sound data. The system 100 may perform machine learning of an inference model based on the training data (step S712). This allows the user's sensitivity to timbre to be reflected in the inference model.
[0097] Machine learning of the inference model may be performed for each user. In this case, machine learning is performed to obtain an inference model for each user using learning data for each user. In this case, each inference model is personalized for the user, and the figure associated with the sound data becomes more appropriate from the perspective of each user.
[0098] The machine learning of the inference model may be performed by integrating training data from multiple users, resulting in a generalized inference model that can be used by multiple users. The generalized inference model may also be fine-tuned using training data from each user.
[0099] 10(A) and 10(B) show other examples of user interfaces for the system 100. The user interface 400 shown in FIG. 10(A) is a user interface for a music production system and has similar functions to the interface 200 shown in FIG. 2, but differs in its arrangement and layout. In the user interface shown in FIG. 10(A), the timbres of three types of guitar sound data ("other than lead solo," "lead guitar solo," and "backing") are represented by balloons 401, 402, and 403. These are sound data for the same instrument, the guitar. In the user interface 400 shown in FIG. 10(A), the differences in the timbres of the sound data for "other than lead solo," "lead guitar solo," and "backing" can be visually grasped by the differences in the balloons. Note that the user interface 400 also allows selection of a shape type and editing (adjustment) of the shape of the shape.
[0100] User interface 500 shown in FIG. 10(B) is a user interface for a video production system (video editing system). This user interface 500 has a function 510 for playing video and a function 520 for displaying a seek bar. The seek bar is a function for displaying and specifying a part of a video. The seek bar function 520 displays images of each point in time of the video in chronological order, and also displays graphics 521, 522, 523, and 524 that indicate the timbre of the sound included in the video at each point in time, corresponding to the chronological order. These graphics 521, 522, 523, and 524 allow the system user to visually grasp the timbre of the sound at each point in time, without having to listen to the sound.
[0101] The present invention is not limited to the above-described embodiment, and various modifications are possible. [Explanation of symbols]
[0102] 100: Production System 110: Processor 120: Memory 121: Sound data 122: Shapes 123: Estimation model 123A: First estimation model 123B: Second estimation model 124: Training data 125: Computer Programs 130: Input / output interface 140: Display 150: Input device 200: User Interface 210: List display section 210A:Identification display section 210B: Piano roll display 211: Shapes 212: Shapes 220:Specified part 230: Setting section 231: Speech bubble shape 232: Speech bubble shape 233: Speech bubble shape 240: Figure list display section 250: Graphic display section 250A: Shape 250B:Operation unit 300: User Interface 310: Graphic type display section 320: 1st Editorial Department 321: 1st adjustment section 322:Second adjustment section 323: Input section 324:Display section 330: 2nd Editorial Department 331: 1st adjustment section 332:Second adjustment section 333:Third adjustment section 334: 4th adjustment section 335: 5th adjustment section 336: 6th adjustment section 337: Input section 338: Display section 340: 3rd Editorial Department 341: 1st adjustment section 342:Second adjustment section 343:Third adjustment section 344: 4th adjustment section 345: 5th adjustment section 346: 6th adjustment section 347: Input section 348:Display section 400: User Interface 401: Speech bubble shape 402: Speech bubble shape 403: Speech bubble shape 500: User Interface 520: Seek bar function 520: Video display function 521: Shapes 522: Shapes 523: Shapes 524: Shapes
Claims
1. A production system for content including sound data, a user interface displayed for a user to create content including sound data; the user interface is configured to display a graphic having a form corresponding to the timbre of the sound data. Production system.
2. The graphic includes a speech bubble graphic. The production system of claim 1 .
3. The sound data includes musical instrument sound data. The production system of claim 1 .
4. The user interface is configured to display characters related to the sound data in the graphic. The production system of claim 1 .
5. a processor that generates the figure from the sound data and stores the sound data and the generated figure in a memory in association with each other; the user interface is configured to display the graphic associated with the sound data in the memory as a graphic indicating the tone color of the sound data. The production system of claim 1 .
6. The graphics generated by the processor are editable by a user. The production system according to claim 5 .
7. Generating the figure from the sound data includes using an estimation model to estimate the figure from the sound data. The production system according to claim 5 .
8. The estimation model is provided for each type of sound. The production system of claim 7 .
9. generating the figure from the sound data, a first estimation model that estimates the type of the figure from the sound data; a second estimation model that estimates, from the sound data, parameters of the figure of the type estimated by the first estimation model; estimating the type of the shape and the parameters of the shape using The production system according to claim 5 .
10. 1. A computer-implemented method performed by a computer, comprising: In a user interface related to content including sound data, displaying a graphic having a form corresponding to the tone of the sound data. Computer-implemented methods.
11. A computer program for causing a computer to execute processing including displaying a graphic having a form corresponding to the tone of sound data in a user interface relating to content including sound data.
Citation Information
Patent Citations
Electronic blackboard system
JP1989034799A