Audio control device and audio control method

The voice control device addresses the limitation of conventional systems by using a natural language analysis unit to calculate multiple evaluation axes, enabling detailed and flexible musical performance adjustments, thus delivering music in the desired style of the user.

WO2026013821A1PCT designated stage Publication Date: 2026-01-15NT T INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/025023
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Conventional techniques fail to provide music in the manner desired by the user, as they often rely on a single adjective to express complex performance preferences, which may not capture the user's nuanced interpretation.

Method used

A voice control device with a natural language analysis unit that calculates values corresponding to multiple evaluation axes and a performance information modification unit to update musical performance parameters based on these values, allowing for more detailed and flexible control of musical output.

Benefits of technology

Enables music to be provided in a format desired by the user, accommodating diverse and complex performance styles through the analysis of natural language input and adjustment of performance parameters such as tempo, loudness, and timing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024025023_15012026_PF_FP_ABST
    Figure JP2024025023_15012026_PF_FP_ABST
Patent Text Reader

Abstract

An audio control device (10) has a natural language analysis unit (151) and a rendition information modification unit (153). The natural language analysis unit (151) analyzes inputted natural language text and calculates values that correspond to each of a plurality of evaluation axes. On the basis of the plurality of calculated values, the rendition information modification unit (153) updates rendition information (141) that defines how music is rendered. An output control unit (152) outputs audio that renders music in accordance with the rendition information (141).
Need to check novelty before this filing date? Find Prior Art

Description

Audio control device and audio control method

[0001] The present invention relates to a voice control device and a voice control method.

[0002] Since not all performance information is specified in a musical score, it is known that even the same score can be performed in a variety of ways depending on the performer's interpretation or preference (see, for example, Non-Patent Document 1). In order to realize the diverse performances desired by listeners in a system, there is a technology that allows the performance information to be changed by inputting one of six adjectives that represent the preferred style of the performance (see, for example, Non-Patent Document 2).

[0003] Carlos E. C. et al., “Computational Models of Expressive Music Performance: A Comprehensive and Critical Review”, Front. Digit. Humanit., 2018Canazza, Sergio, Giovanni De Poli, and Antonio Roda. “Caro 2.0: an interactive system for expressive music rendering.” Advances in Human-Computer Interaction 2015Cancino-Chacon, Carlos, et al. “On the characterization of expressive performance in classical music: First results of the con espressione game.” arXiv preprint(2020).

[0004] However, conventional techniques may not be able to provide music in the manner desired by the user.

[0005] For example, in the technology of Non-Patent Document 2, the user's preferred style is expressed in one word, so it may not be possible to express the complex aspects desired by the user.

[0006] The present invention has been made in view of the above, and has as its object to output music in a manner desired by a user.

[0007] In order to solve the above-mentioned problems and achieve the objectives, the voice control device of the present invention is characterized by having a natural language analysis unit that analyzes input natural language text and calculates values ​​corresponding to each of multiple evaluation axes, and a performance information modification unit that updates performance information that defines the manner of musical performance based on the values.

[0008] According to the present invention, it is possible to provide music in a format desired by the user.

[0009] Fig. 1 is a diagram showing an example of the configuration of a voice control system according to a first embodiment. Fig. 2 is a flowchart showing the processing flow of a voice control device according to the first embodiment. Fig. 3 is a diagram explaining prompts and weighted vectors. Fig. 4 is a diagram showing an example of the configuration of a voice control device according to a second embodiment. Fig. 5 is a flowchart showing the processing flow of the voice control device according to the second embodiment. Fig. 6 is a diagram showing an example of a computer that executes a voice control program.

[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are designated by the same reference numerals.

[0011] [First embodiment] [Configuration of voice control device] Fig. 1 is a diagram showing an example of the configuration of a voice control system according to a first embodiment. The voice control system 1 shown in Fig. 1 can provide music in a format desired by a user. Even for the same piece of music, different users may have different preferred performance styles. The voice control system 1 accepts performance instructions from the user in intuitive natural language and reflects the accepted performance instructions in the performance of the music to be output.

[0012] 1, the voice control system 1 includes a voice control device 10, a voice output device 20, and an input device 30. The voice control device 10 is connected to an LLM server 40 via a network N.

[0013] The voice control device 10 is a general-purpose computer such as a PC. The voice output device 20 is a speaker, a headphone, etc. The input device 30 is a keyboard, a microphone, etc. The LLM server 40 generates a response based on the input prompt using a large language model (LLM).

[0014] A user inputs performance instructions to the audio control device 10 via the input device 30. The performance instructions may be input as text or as audio. Audio input performance instructions are converted into text using voice recognition or the like. The user can also listen to the audio output by the audio control device 10 via the audio output device 20.

[0015] As shown in FIG. 1 , the voice control device 10 includes a communication unit 11 , an input unit 12 , an output unit 13 , a storage unit 14 , and a control unit 15 .

[0016] The communication unit 11 performs data communication with other devices via a network. For example, the communication unit 11 is a network interface card (NIC). The input unit 12 is an interface connected to the input device 30. The output unit 13 is an interface connected to the audio output device 20.

[0017] The storage unit 14 is a storage device such as a hard disk drive (HDD), a solid state drive (SSD), an optical disk, etc. Note that the storage unit 14 may also be a data-rewritable semiconductor memory such as a random access memory (RAM), a flash memory, or a non-volatile static random access memory (NVSRAM). The storage unit 14 stores an operating system (OS) and various programs executed by the voice control device 10.

[0018] The storage unit 14 stores performance information 141, which is information that defines the manner in which music is played. For example, the performance information 141 may be MIDI data. Alternatively, the performance information 141 may be an audio file in a format such as mp3 or wav.

[0019] The control unit 15 controls the entire voice control device 10. The control unit 15 is, for example, an electronic circuit such as a central processing unit (CPU), a micro processing unit (MPU), or a graphics processing unit (GPU), or an integrated circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).

[0020] The control unit 15 also has an internal memory for storing programs that define various processing procedures and control data, and executes each process using the internal memory. The control unit 15 also functions as various processing units when various programs are run. For example, the control unit 15 has a natural language analysis unit 151, an output control unit 152, and a performance information modification unit 153.

[0021] The natural language analysis unit 151 analyzes the input natural language text and calculates values ​​corresponding to each of the multiple evaluation axes. A vector containing the values ​​calculated by the natural language analysis unit 151 as elements is called a weighted vector.

[0022] The output control unit 152 outputs sounds that play music in accordance with the performance information 141. For example, the output control unit 152 outputs a signal generated based on the performance information 141, which is MIDI data, to the audio output device 20 via the output unit 13.

[0023] The performance information modifying unit 153 updates the performance information 141 based on the multiple values, i.e., the weighted vector, calculated by the natural language analyzing unit 151. As a result, the performance information 141 changes in accordance with the weighted vector.

[0024] The processing of the voice control device 10 will be described in detail with reference to Figures 2 and 3. Figure 2 is a flowchart showing the flow of processing of the voice control device according to the first embodiment. Figure 3 is a diagram illustrating prompts and weighted vectors.

[0025] 2, first, the natural language analysis unit 151 receives input of a performance instruction in natural language (step S101). The performance instruction may be a free description in natural language of the user's desired performance.

[0026] As shown in FIG. 3, for example, a performance instruction 510 such as "I want the performance to give a soft impression" is input to the sound control device 10.

[0027] Next, the natural language analysis unit 151 analyzes the performance instructions (step S102). Furthermore, the natural language analysis unit 151 calculates values ​​corresponding to each of the multiple evaluation axes based on the analysis results (step S103). The natural language analysis unit 151 then maps a weighted vector based on the calculated values ​​(step S104).

[0028] Here, the natural language analysis unit 151 uses LLM to analyze performance instructions, calculate values ​​corresponding to evaluation axes, and map weighted vectors. However, in addition to LLM, the natural language analysis unit 151 can also perform analysis using deep learning models that handle natural language, such as BART (Bidirectional Encoder Representations from Transformers) and Transformer.

[0029] 1 also shows, as an example, a configuration in which the voice control device 10 uses an external LLM server 40 that provides services related to LLM (e.g., ChatGPT (registered trademark)). On the other hand, the voice control device 10 may have a deep learning model such as LLM internally.

[0030] For example, if a virtual LLM server is built inside the voice control device 10, the natural language analysis unit 151 can request the LLM server to generate a response by executing an API (Application Programming Interface). In this case, the natural language analysis unit 151 can communicate with the LLM server via the communication unit 11 through a virtual network built inside the voice control device 10.

[0031] The multiple evaluation axes may be predetermined. For example, Non-Patent Document 2 lists adjectives such as "Hard," "Heavy," "Dark," "Soft," "Light," and "Bright" as words that describe playing styles. Similarly, Non-Patent Document 3 lists words such as "agitated" and "gentle."

[0032] Each evaluation axis in the first embodiment may be associated with one or more of the words shown in Non-Patent Document 2 and Non-Patent Document 3.

[0033] Here, two words that can have opposing meanings are associated with each evaluation axis. Specifically, four evaluation axes are used: "agitated⇔gentle axis," "nervous⇔hard axis," "warm⇔graceful axis," and "happy⇔cold axis."

[0034] For example, "agitated" and "gentle" can have opposing meanings. A user might imagine an "agitated" style of performance as a fast tempo with emphasis on accents, and a "gentle" style as a slow tempo with less emphasis on accents.

[0035] The natural language analysis unit 151 creates the prompt 520 shown in FIG. 3 in the following procedure. First, a sentence expressing the system's role is added to the prompt 520. The sentence expressing the system's role includes an instruction to analyze the content of the performance and an instruction to output a value in an appropriate domain for the evaluation axis. For example, the sentence expressing the system's role is "Analyze the content of the instructions for the classical music performance and assign a score between -1 and 1 for each axis defined in the evaluation axis."

[0036] Next, the natural language analysis unit 151 adds a sentence explaining each evaluation axis to the prompt 520. For example, the natural language analysis unit 151 adds the following sentence to explain the "agitated⇔gentle axis", "agitated⇔gentle axis: an axis for evaluating the busyness of a performance, and a large evaluation value as a positive value is assigned to a performance that gives an agitated feeling, and a large evaluation value as a negative value is assigned to a performance that gives a gentle feeling." Furthermore, the natural language analysis unit 151 adds sentences explaining each of the "nervous⇔hard axis", "warm⇔graceful axis", and "happy⇔cold axis".

[0037] The natural language analysis unit 151 inputs the prompt 520 to the LLM server 40 and obtains a response 530. The response 530 describes the value of each element of the weighted vector. In the example of FIG. 3, it is shown that the value of the "agitated⇔gentle axis" is "-0.5", the value of the "nervous⇔hard axis" is "0.6", the value of the "warm⇔graceful axis" is "0.0", and the value of the "happy⇔cold axis" is "0.7".

[0038] The performance information alteration unit 153 updates the performance information 141 based on the weighted vector (step S105). The performance information alteration unit 153 can update parameters such as tempo, loudness (volume), timing, and note value in the performance information 141, which is MIDI data, according to a predetermined rule, depending on the value of each element of the weighted vector. For example, the performance information alteration unit 153 performs the update according to the rule "increase the tempo by 10% times the value of the 'agitated⇔gentle axis'." Furthermore, the parameters to be updated may be deviations from the timing and note value written in the musical score (for example, offset values ​​from the note value as written in the musical score).

[0039] The performance information alteration unit 153 may use a machine learning technique such as a recurrent neural network or a transformer to update the performance information 141. In this case, the performance information alteration unit 153 may use a model that has been trained using, as training data, a combination of a weighted vector and the update content of the performance information 141 performed by a person in accordance with the weighted vector.

[0040] Returning to Fig. 2, the output control unit 152 outputs music based on the performance information 141 (step S106). This allows the user to listen to audio corresponding to the performance information 141. The user can then input additional performance instructions based on the results of listening. If additional performance instructions have been input (step S107, Yes), the audio control device 10 returns to step S101 and repeats the process. This allows the audio control device 10 to receive feedback from the user regarding the updated performance information 141 and modify the performance information 141 to meet the user's wishes.

[0041] If no additional performance instruction has been input (No in step S107), the sound control device 10 ends the process, although sound may continue to be output.

[0042] The natural language analysis unit 151 may automatically generate an explanation of the evaluation axis in the prompt 520. In this case, the voice control device 10 stores a history of performance instructions for each user, and when similar performance instructions are repeatedly input for the same performance information 141, the voice control device 10 can use the contents of the weighted vector and the expression vector corresponding to the previous performance instructions as input.

[0043] For example, consider a case where the previous performance instruction was "I want the performance to give a soft impression," and the current performance instruction is "Even softer." The performance instruction "Even softer" corresponds to a performance instruction that makes the previous performance instruction more noticeable.

[0044] In this case, the natural language analysis unit 151 creates a weighted vector by increasing each value without changing the ratio of the previous weighted vector. For example, if the previous weighted vector was (-0.5, 0.6, 0.0, 0.7) as shown in response 530 in Figure 3, the natural language analysis unit 151 creates a weighted vector (-0.6, 0.72, 0.0, 0.84) by multiplying each value by 1.2. Note that the multiplication factor, such as 1.2, may be set in advance.

[0045] Furthermore, the natural language analysis unit 151 may add a supplementary explanation to the prompt 520 by using meta information for the performance information 141. For example, if the meta information indicates that the composer of the score corresponding to the performance information 141 is Bach, the natural language analysis unit 151 adds the following sentence to the prompt 520: "The score in question was composed by Bach, and it is said that it is desirable to use expressions that give an overall majestic impression, while romantic expressions are undesirable." In this case, the voice control device 10 may store supplementary explanations for each composer (for example, the portion of the sentence, "It is said that it is desirable to use expressions that give an overall majestic impression, while romantic expressions are undesirable.") in a database.

[0046] Furthermore, the meta information is not limited to the composer, but may also include the genre, the period in which the music was composed, the performer, the conductor, etc. The meta information may be stored in advance in the audio control device 10 together with the performance information 141, or may be input to the audio control device 10 together with performance instructions. The audio control device 10 may also acquire the meta information from a website (Reference 1: https: / / theclassicreview.com / ).

[0047] [Effects of the First Embodiment] As described above, the natural language analysis unit 151 analyzes the input natural language text and calculates values ​​corresponding to each of the multiple evaluation axes. The performance information modification unit 153 updates the performance information 141 that defines the musical performance aspect based on the multiple calculated values. Furthermore, the output control unit 152 outputs audio that plays music in accordance with the performance information 141.

[0048] For example, in the technology described in Non-Patent Document 2, performance information changes depending on a single adjective. However, a single adjective may not be able to express the user's desired performance style, etc. In contrast, the voice control device 10 of the first embodiment calculates values ​​according to multiple evaluation axes, thereby diversifying the performance information 141 in accordance with the user's wishes and enabling more flexible and detailed control of the performance information 141. As a result, according to the first embodiment, music is provided in the style desired by the user.

[0049] For example, in the technology described in Non-Patent Document 2, only one of the words "dark" or "soft" can be reflected in the performance information, whereas in the first embodiment, the words "dark and soft" can be reflected in the performance information.

[0050] Second Embodiment A second embodiment will be described below. In the description of the second embodiment, the description of the parts common to the first embodiment will be omitted as appropriate.

[0051] First, quantitative instructions and qualitative instructions will be explained. A performance instruction is quantitative if the operation of the performance information in response to the instruction is clear. For example, "Please turn up the volume of the XX part" and "Please make it quieter overall" are quantitative instructions, and it is clear that they are instructions regarding volume. Also, for example, "Please play this at a constant tempo" and "Please speed up the tempo a little more" are quantitative instructions, and it is clear that they are instructions regarding tempo. Also, for example, "Please align the timing of the notes in the chords at the beginning of the measure" is a quantitative instruction, and it is clear that it is an instruction regarding timing. Also, for example, "Please shorten the notes overall" is a quantitative instruction, and it is clear that it is an instruction regarding the length of the notes.

[0052] For example, "I want the performance to give a soft impression" is a qualitative instruction, and it is not clear what parameter the instruction is directed to.

[0053] The configuration of the voice control device 10 according to the second embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of the configuration of the voice control device according to the second embodiment. Fig. 4 shows the configurations of a natural language analysis unit 151 and a performance information modification unit 153.

[0054] 4, the natural language analysis unit 151 includes a performance instruction determination unit 1511 and an impression extraction unit 1512. The performance information modification unit 153 includes an impression conversion unit 1531, a parameter replacement unit 1532, and a parameter extraction unit 1533.

[0055] The performance instruction determination unit 1511 determines whether the performance instruction text includes quantitative instructions for musical performance. The performance instruction determination unit 1511 may perform this determination using keyword matching, regular expressions, or the like, or may perform this determination using a machine learning technique such as LLM.

[0056] The impression extracting unit 1512 extracts information relating to impressions from the qualitative instructions included in the performance instructions. Here, the impression extracting unit 1512 acquires the weighted vector in the first embodiment as information relating to impressions.

[0057] The parameter extraction unit 1533 identifies and extracts parameters to be updated from quantitative instructions. Note that the parameters of the performance information 141 are divided into categories such as tempo, loudness (volume), timing, and sound duration. The parameter extraction unit 1533 identifies which category the quantitative instruction corresponds to. For example, the parameter extraction unit 1533 may extract a certain category and its state from a qualitative instruction using a Japanese dependency analyzer such as CaboCha (Reference 2: https: / / taku910.github.io / cabocha / ). For example, the parameter extraction unit 1533 extracts the parameter category "tempo" and the state "raise" from the quantitative instruction "I'd like the tempo to be increased a little."

[0058] If the text includes quantitative instructions, the parameter replacing unit 1532 updates the parameters included in the performance information 141 in accordance with the quantitative instructions, based on the extraction result by the parameter extracting unit 1533. Furthermore, the impression converting unit 1531 updates the performance information 141 based on the weighted vector acquired by the impression extracting unit 1512. Furthermore, the impression converting unit 1531 can use music information (meta information), previous performance instructions, etc.

[0059] That is, quantitative instructions are processed by the parameter extraction unit 1533 and the parameter replacement unit 1532. On the other hand, qualitative instructions are processed by the impression extraction unit 1512 and the impression conversion unit 1513. The impression extraction unit 1512 and the impression conversion unit 1513 can process qualitative instructions in the same manner as in the first embodiment.

[0060] 5 is a flowchart showing the flow of processing of the voice control device according to the second embodiment. As shown in FIG. 5, first, the natural language analysis unit 151 accepts input of a performance instruction in natural language (step S201).

[0061] If the performance instruction includes a qualitative instruction (step S202, Yes), the natural language analysis unit 151 maps a weighted vector based on the qualitative instruction (step S203).The performance information modification unit 153 updates the performance information based on the weighted vector (step S204).

[0062] If the performance instruction does not include a qualitative instruction (No at step S202), the performance information altering unit 153 proceeds to step S205.

[0063] If the performance instructions include quantitative instructions (Yes in step S205), the performance information modifying unit 153 converts the quantitative instructions into parameters (step S206). For example, the performance information modifying unit 153 converts the quantitative instructions into parameter categories and states. The performance information modifying unit 153 replaces the parameters of the performance information based on the converted parameters (step S207).

[0064] If the performance instruction does not include a quantitative instruction (No at step S205), the performance information altering section 153 proceeds to step S208.

[0065] The output control unit 152 outputs music based on the performance information (step S208). If an additional performance instruction is input (step S209, Yes), the audio control device 10 returns to step S201 and repeats the process. If an additional performance instruction is not input (step S209, No), the audio control device 10 ends the process. However, audio may continue to be output.

[0066] If the performance instruction includes both qualitative and quantitative instructions, part of the performance information updated in step S204 may be overwritten in step S207. For example, consider the case where a performance instruction such as "Please speed up the tempo but play in a way that gives a soft impression" is input. This performance instruction includes both the quantitative instruction of "speed up the tempo" and the qualitative instruction of "Please play in a way that gives a soft impression."

[0067] In this case, in step S204, the voice control device 10 makes changes to the performance information, such as slowing down the tempo or lowering the volume, based on the qualitative instruction, "Please make the performance give a soft impression." Then, in step S207, the voice control device 10 makes changes to speed up the tempo based on the quantitative instruction, "Increase the tempo." As a result, the tempo parameter in the performance information is overwritten with a faster one. As a result, the voice control device 10 can achieve both a soft impression and an increased tempo based on the performance instruction.

[0068] According to the second embodiment, quantitative instructions can be more clearly recognized, and therefore the user's wishes can be reflected in the performance information 141 with higher accuracy.

[0069] [Program] In one embodiment, the voice control device 10 can be implemented by installing a voice control program that executes the above-described processes as package software or online software on a desired computer. For example, by having an information processing device execute the above-described voice control program, the information processing device can function as the voice control device 10. The information processing device referred to here includes desktop and notebook personal computers. Other examples of information processing devices include smartphones, tablet terminals, and the like.

[0070] 6 is a diagram showing an example of a computer that executes a voice control program. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0071] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM (Random Access Memory) 1012. The ROM 1011 stores, for example, a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.

[0072] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the programs that define each process of the voice control device 10 are implemented as program modules 1093 in which computer-executable code is written. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, the program modules 1093 for executing processes similar to those of the functional configuration of the voice control device 10 are stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced with an SSD.

[0073] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. The CPU 1020 then reads the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes the processing of the above-described embodiment.

[0074] The program module 1093 and program data 1094 may not necessarily be stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0075] Although the present invention has been described above as an embodiment, the present invention is not limited to the description and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention.

[0076] REFERENCE SIGNS LIST 10 Audio control device 11 Communication unit 12 Input unit 13 Output unit 14 Memory unit 15 Control unit 20 Audio output device 30 Input device 40 LLM server 151 Natural language analysis unit 152 Output control unit 153 Performance information modification unit 510 Performance instruction 520 Prompt 530 Answer 1511 Performance instruction determination unit 1512 Impression extraction unit 1531 Impression conversion unit 1532 Parameter replacement unit 1533 Parameter extraction unit

Claims

1. A voice control device comprising: a natural language analysis unit that analyzes input natural language text and calculates values ​​corresponding to each of multiple evaluation axes; and a performance information modification unit that updates performance information that defines the manner of musical performance based on the values.

2. The audio control device described in claim 1, further comprising an output control unit that outputs audio that plays music in accordance with the performance information, wherein the natural language analysis unit further analyzes natural language text input after the audio is output and calculates values ​​corresponding to each of a plurality of evaluation axes, the performance information modification unit updates performance information that defines the manner of musical performance based on the values, and the output control unit outputs audio that plays music in accordance with the performance information.

3. The voice control device described in claim 1, characterized in that the natural language analysis unit determines whether the text contains quantitative instructions regarding musical performance, and the performance information modification unit, if the text contains the quantitative instructions, updates the parameters contained in the performance information in accordance with the quantitative instructions.

4. A voice control method executed by a voice control device, comprising: a natural language analysis step of analyzing input natural language text and calculating values ​​corresponding to each of a plurality of evaluation axes; and a performance information modification step of updating performance information defining the manner of musical performance based on the values.

Citation Information

Patent Citations

  • Mobile terminal device

    JP2004163511A

  • Information processing device, and its control method

    JP2022135126A

  • Sound generation method, sound generation system and program

    JP2023131494A