Device and method

The apparatus and method facilitate the creation of facility-specific music by using facility and user information to generate prompts for AI models, addressing the complexity of prompt generation and ensuring the music aligns with the facility's purpose and user preferences, thus enhancing the atmosphere and user experience.

WO2025262804A1PCT designated stage Publication Date: 2025-12-26NTT DOCOMO INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/022085
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-18
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Creating music suitable for a facility using AI models is challenging due to the complexity of generating appropriate prompts, which requires advanced skills and time, and often results in unintended outputs.

Method used

An apparatus and method that includes an acquisition unit for facility and user information, a determination unit for composition information, and a generation unit to create prompts for an AI model to generate songs based on this information, ensuring the music aligns with the facility's purpose and user preferences.

Benefits of technology

Enables easy creation of music tailored to facilities by generating prompts that reflect facility and user information, resulting in songs that enhance the facility's atmosphere and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024022085_26122025_PF_FP_ABST
    Figure JP2024022085_26122025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure makes it possible to easily create a song suitable for a facility. A device 10 according to the present disclosure comprises: an acquisition unit 12 that acquires at least one of facility information pertaining to a facility and user information pertaining to a user of the facility; a determination unit 13 that determines, on the basis of at least one of the facility information and the user information, feature information pertaining to the feature of a song to be played in the facility; and a generation unit 14 that generates, on the basis of the feature information, a prompt for instructing the AI model 41 to create a song.
Need to check novelty before this filing date? Find Prior Art

Description

Apparatus and method

[0001] One aspect of the present disclosure relates to an apparatus and method for generating a prompt.

[0002] Music has traditionally been played as background music (BGM) in various facilities, such as commercial facilities. Music played in facilities can influence the atmosphere of the facility, guide the behavior of facility users, and encourage users to make purchases, so it is desirable to play music that is appropriate for the purpose of the facility. Furthermore, with the recent advancement of AI (artificial intelligence) technology, music composition techniques using AI models have been developed (see, for example, Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2023-129639

[0004] In the composition work using the AI ​​model described above, it is not easy for a user to make the AI ​​model output the song that the user wants. For example, in order to make the AI ​​model output the song that the user wants, the user needs to create appropriate prompts to input to the AI ​​model, but generating the prompts can require advanced skills and a lot of time.

[0005] An object of the present disclosure is to provide an apparatus and method that can easily create songs that are suitable for a facility.

[0006] An apparatus according to one aspect of the present disclosure includes an acquisition unit that acquires at least one of facility information regarding a facility and user information regarding users of the facility, a determination unit that determines composition information regarding the composition of a song to be played at the facility based on at least one of the facility information and the user information, and a generation unit that generates a prompt to instruct an AI model to create a song based on the composition information.

[0007] A method according to one aspect of the present disclosure includes steps of acquiring at least one of facility information regarding a facility and user information regarding users of the facility, determining composition information regarding the composition of a song to be played at the facility based on at least one of the facility information and the user information, and generating prompts to instruct an AI model to create a song based on the composition information.

[0008] According to one aspect of the present disclosure, music suitable for a facility can be easily created.

[0009] Fig. 1 is a diagram illustrating an overview of the processing of a composition system according to an embodiment of the present disclosure. Fig. 2 is a diagram illustrating a device configuration of the composition system. Fig. 3 is a flowchart illustrating an example of the operation of the composition system. Fig. 4 is a diagram illustrating an example of configuration information. Fig. 5 is a diagram illustrating an example of a prompt. Fig. 6 is a diagram illustrating a device configuration of a composition system according to a modified example. Fig. 7 is a diagram illustrating a device configuration of a composition system according to a modified example. Fig. 8 is a diagram illustrating an example of a hardware configuration of a composition system according to an embodiment of the present disclosure.

[0010] The present disclosure will be described with reference to the accompanying drawings. Whenever possible, the same parts are designated by the same reference numerals and redundant description will be omitted.

[0011] The composition system according to this embodiment is a computer system for creating music to be performed in a facility. "Performing" refers to generating sounds using devices such as speakers or musical instruments, and includes playing back recorded sounds. A song is a combination of sounds, including not only artificial sounds such as those emitted by musical instruments, but also natural sounds such as the murmuring of a river. Music created by the composition system according to this embodiment may be played back as background music (BGM) in the facility.

[0012] Facilities include buildings, places, and equipment that provide users with specific purposes and functions. Facilities may be private facilities owned or operated by individuals or private companies, or public facilities owned or operated by national or local governments. Facilities may be, for example, commercial facilities, cultural facilities, entertainment facilities, medical facilities, or transportation facilities. Commercial facilities may be, for example, retail stores, shopping malls, department stores, supermarkets, convenience stores, offices, cafes, restaurants, or hotels. Cultural facilities may be, for example, art galleries, museums, theaters, concert halls, libraries, or science museums. Entertainment facilities may be, for example, amusement parks, theme parks, movie theaters, or sports facilities. Medical facilities may be, for example, hospitals, clinics, health centers, or pharmacies. Transportation facilities may be, for example, train stations, airports, or bus terminals.

[0013] An overview of the processing of the composition system according to this embodiment will be described with reference to FIG. 1 . The processing shown below is an example of processing performed by the composition system. First, the composition system determines configuration information for a song to be composed based on at least one of facility information and user information. Next, the composition system generates prompts for instructing the AI ​​model to create a song based on the configuration information. Next, the composition system inputs the generated prompts into the AI ​​model and causes the AI ​​model to output a song, thereby creating a song to be performed at the facility. As described above, the composition system according to this embodiment has the function of generating prompts for instructing the AI ​​model to create a song. Therefore, the composition system also functions as a prompt generation system for generating prompts.

[0014] Facility information is information about the facility where the music will be played. The facility information may include, for example, at least one of the following information: the name, type, atmosphere, purpose of use, location (location), surrounding environmental conditions, business hours, structure, texture, site area, functions, service content, available facilities, volume of ambient sound, seating, and image of the facility.

[0015] The information about the type of facility may be, for example, information indicating that the facility is a large commercial facility for families or a commercial complex. The information about the texture of the facility may include information about the materials (concrete, wood, etc.) of the building or structure of the facility.

[0016] Information about the loudness of environmental sounds in a facility is information about the loudness of sounds occurring in the facility. Environmental sounds in a facility include users' voices, noise, background noise, natural sounds, etc. Information about the loudness of environmental sounds may be obtained based on data measured by a measuring device installed in the facility. Information about the loudness of environmental sounds may be information expressed using abstract expressions such as loud or quiet, or may be information expressed using specific numerical values ​​in decibels.

[0017] The information about seats may be information about the type of seats (counter seats, table seats, etc.) or information about the number of seats. The information about the image of the facility may be information that verbalizes the impression that the facility gives to users. For example, the information about the image of the facility may be information such as "a facility where families can have fun" or "a facility that provides a relaxing space in an desirable location."

[0018] The user information is information about users of the facility. The user information may include information about at least one of the number of users at the facility, the degree of congestion, attributes, consumption behavior, and behavior patterns within the facility.

[0019] Information related to the degree of user congestion is, for example, information indicating the degree of congestion of users at a facility, and may be information expressed by the number of users or density (congestion) at the facility. The degree of congestion may be information expressed by abstract expressions such as crowded or empty, or may be information expressed by concrete numerical expressions such as the number of users per unit space or density. The degree of congestion may be a numerical value calculated, for example, by dividing the number of seats in use at a facility by the total number of seats.

[0020] The user attributes may be, for example, information about the user's age group (child or elderly), composition (family or single), occupation (student or working person), or nationality (foreigner or not). The consumption behavior may be information about the amount of money the user spends at the facility (e.g., the price range of the goods or services purchased). The behavioral pattern may be information about the user's average stay time at the facility or the user's travel route.

[0021] The composition information is information about the composition of a song. The composition information may include elements and parameters used by the AI ​​model to create a song. The composition information is incorporated, for example, as composition conditions in prompts that the AI ​​model follows when composing a song. The composition information may include, for example, information about at least one of the tempo, melody, volume, genre, image, and length of the song.

[0022] Information about the tempo of a song may be expressed by an abstract expression such as "fast" or "slow," or may be expressed by a specific numerical value in units of BPM (Beats Per Minute). Information about melody is information indicating the characteristics of a melody or melodic style, and may include information such as the scale, types of chords, and tempo used in the song. Information about volume may be expressed by an abstract expression such as "loud" or "quiet," or may be expressed by a specific numerical value in units of decibels. Genre may be information about a music category, such as classical, techno, jazz, rock, pop, or easy listening. Information about image may be information indicating the impression the song gives to the listener. Information about image may be, for example, "fun pop music," "fast-tempo music," or "luxurious music." Information about length may be information expressed in seconds, minutes, or hours.

[0023] [System Configuration] Fig. 2 is a diagram showing the device configuration of a composition system 1 according to this embodiment. The composition system 1 includes a device 10. The device 10 according to this embodiment is, for example, a Retrieval-Augmented Generation (RAG) system. The device 10 is connected to a user terminal 20, a database 30, and a server device 40 via a communication network N. The configuration of the communication network N is not limited. For example, the communication network N may be configured to include the Internet or an intranet.

[0024] The device 10 includes, as functional components, a reception unit 11, an acquisition unit 12, a determination unit 13, a generation unit 14, a creation unit 15, and a control unit 16. The reception unit 11 is a functional element that receives information regarding various requests and instructions from, for example, a user terminal 20. The acquisition unit 12 is a functional element that acquires various data and information from, for example, the user terminal 20 and a database 30. The determination unit 13 is a functional element that determines configuration information based on facility information and user information. The generation unit 14 is a functional element that generates, based on the configuration information, a prompt for instructing an AI model 41 (described later) to create a song. The creation unit 15 is a functional element that creates a song. The control unit 16 is a functional element that controls a song playback device (such as a speaker) installed in the facility to play a song.

[0025] The user terminal 20 is a computer used by a user. The user uses the composition system 1 to create a piece of music to be performed at a facility. The user may be, for example, a facility manager, facility operator, composer, music producer, or music distribution service provider. There is no limitation on the type of computer used as the user terminal 20. The user terminal 20 may be, for example, a personal computer, a high-function mobile phone (smartphone), a mobile phone, a personal digital assistant (PDA), a tablet terminal, a wearable terminal, or other portable terminal. There is no limitation on the number of user terminals 20.

[0026] The database 30 is a non-transitory storage device that stores data and information used in the composition system 1. The database 30 stores, for example, facility information, user information, etc. The database 30 may be constructed as a single database or may be a collection of multiple databases. The location of the database 30 is not limited. The database 30 may be provided, for example, in a computer system separate from the composition system 1.

[0027] The server device 40 includes an AI model 41. The AI ​​model 41 according to this embodiment is a generative artificial intelligence (AI) model. The generative AI model is a model that, in response to a prompt including input information, creates content according to any one or a combination of instructions, context, questions, and output formats indicated by the prompt, and returns the content as output information.

[0028] A prompt is a set of instructions or information input to a generative AI model. This prompt can include initial information, parameters, questions, etc. for the generative AI model to perform a specific task. In the prompt, the text expresses, for example, the instructions to be executed by the generative AI model, the task to be executed by the generative AI model, the background / context (e.g., role, condition) that the generative AI model should take into consideration, the question that the generative AI model should answer, the output format of the output information from the generative AI model, etc.

[0029] The prompt may be input to the generative AI model along with input information that serves as the target or reference for the command or task to be executed by the generative AI model. Such input information may include data files with filenames that include a predetermined extension, such as audio data, text data, image data, application-related data, voice data, video data, and still image data. Application-related data is data such as document data, table data, and graph data that can be processed by a default application program.

[0030] The AI ​​model 41 creates a song according to the given prompts, which may include, for example, creating an audio file such as an MP3, creating a Musical Instrument Digital Interface (MIDI) file, creating sheet music, and / or creating a chord progression.

[0031] As another example, the AI ​​model 41 may be a large language model (LLM). In this case, the AI ​​model 41 may be, for example, an interactive AI configured to include a large language model and a user interface (UI) for interacting with a user, enabling text chat or voice chat with the user. Examples of such an AI model 41 include ChatGPT, GPT (registered trademark)-3.5, GPT-4V, PaLM2, etc. When the AI ​​model 41 is a large language model, the AI ​​model 41 may create text information such as sheet music or chord progressions to create a song.

[0032] The AI ​​model 41 according to this embodiment accepts prompts containing instructions (tasks) for creating a song as input information, creates a song in accordance with the prompts, and outputs the created song as output information.

[0033] The operation of the composition system 1 will be described with reference to Figure 3. Specifically, a prompt generation method according to this embodiment and a method for creating a song using a prompt generated by this method will be described. The process of the prompt generation method can be considered a part of the process of the song creation method. Figure 3 is a flowchart showing an example of the operation of the composition system 1.

[0034] The process of the following song generation method may be initiated by receiving song creation request information (triggered by the creation request information) by the device 10. The creation request information may be request information notified to the receiving unit 11 of the device 10 from the user terminal 20 when, for example, a creation request button displayed on a web page or application page displayed on the display of the user terminal 20 is selected (e.g., pressed by the user).

[0035] In step S11, the acquisition unit 12 acquires facility information. The acquisition unit 12 may acquire the facility information from the user terminal 20. For example, a user operates the user terminal 20 to input facility information about a facility where the user wishes to play (reproduce) a song into the user terminal 20. The acquisition unit 12 acquires the input facility information from the user terminal 20. The acquisition unit 12 may acquire the facility information from the database 30. The facility information may be stored in advance in the database 30 by the user or a person other than the user.

[0036] The acquisition unit 12 may acquire facility information from both the user terminal 20 and the database 30. The acquisition unit 12 may acquire other facility information from the database 30 based on specific facility information acquired from the user terminal 20. For example, the acquisition unit 12 may acquire a facility name from the user terminal 20 as facility information, and acquire other facility information related to the facility with that name (such as the facility's location, type, and business hours) from the database 30. The acquisition unit 12 may acquire information on the Internet as facility information. For example, the acquisition unit 12 may search for information on the Internet using the facility name, etc., and acquire facility information based on the search results.

[0037] In this embodiment, the acquisition unit 12 acquires the facility name "Shopping Mall A" and the facility type "Family-oriented daily necessities store" as facility information.

[0038] In step S12, the acquisition unit 12 acquires user information. The acquisition unit 12 may acquire the user information from the user terminal 20. For example, the user operates the user terminal 20 to input user information of a facility where the user wishes to play (reproduce) a song into the user terminal 20. The acquisition unit 12 acquires the input user information from the user terminal 20.

[0039] The acquisition unit 12 may acquire user information from the database 30. The user information may be stored in advance in the database 30 by the user or a person other than the user. The user information may be stored in the database 30 in association with facility information. In this case, the acquisition unit 12 may acquire user information corresponding to the facility information acquired in step S11 from the database 30. The acquisition unit 12 may acquire user information from both the user terminal 20 and the database 30.

[0040] The acquisition unit 12 may acquire information on the Internet as user information. For example, the acquisition unit 12 may search for information on the Internet using facility information (such as the name of the facility) and acquire user information based on the search results.

[0041] In this embodiment, the acquisition unit 12 acquires the average stay time of users in the facility, "2 hours," as user information.

[0042] In step S13, the acquisition unit 12 acquires timing information. The timing information is information indicating the timing at which the music will be played at the facility. The timing information may include information regarding at least one of the time of day, date, season, and day of the week at which the music will be played. The timing information may be, for example, information indicating that the music will be played during the day, or information indicating that the music will be played in the spring.

[0043] The acquisition unit 12 may acquire timing information from the user terminal 20. For example, when a user wishes to create a song to be played during the day, the user operates the user terminal 20 to input information indicating that the song will be played during the day as timing information into the user terminal 20. The acquisition unit 12 acquires the input timing information from the user terminal 20.

[0044] The acquisition unit 12 may acquire timing information from the database 30. The acquisition unit 12 may acquire (determine) timing information based on various information. For example, the acquisition unit 12 may acquire information about business hours of a facility stored as facility information in the database 30, and acquire information indicating a specific time period included in the business hours as timing information.

[0045] In this embodiment, the acquisition unit 12 acquires, as timing information, the time period in which the song is played, "daytime," and the season in which the song is played, "spring."

[0046] In step S14, the determination unit 13 determines the configuration information based on at least one of the facility information and the user information. In this embodiment, the determination unit 13 determines information related to the image, volume, and length of the song as the configuration information. In this example, the information related to the image and volume of the song among the configuration information is stored in the database 30 in association with the facility information.

[0047] FIG. 4 is a diagram showing an example of configuration information stored in the database 30. In the example of FIG. 4, the facility name and type, which are facility information, are associated with the image and volume of the music, which are configuration information. Specifically, the facility name "Shopping Mall A" and the type "Family General Goods Store" are associated with the image "fun pop music" and the volume "slightly loud," which are configuration information. The name "Shopping Mall A" and the type "Family General Goods Store" are associated with the image "fast tempo music" and the volume "high." The name "Department Store C" and the type "luxury department store" are associated with the image "luxury music" and the volume "medium." The name "Art Museum D" and the type "large art museum" are associated with the image "music that matches the museum's image (nature)" and the volume "low." In the example shown in FIG. 4, time zone information and seasonal information are also associated with the configuration information.

[0048] The determination unit 13 identifies information related to the image and volume that corresponds to the facility information acquired in step S11 and determines it as the configuration information. In this example, the determination unit 13 determines the configuration information based not only on the facility information but also on the timing information acquired in step S13. In the example shown in FIG. 4 , information related to two different images ("fun pop music" and "fast-tempo music") and information related to two different volume levels ("slightly loud" and "loud") are associated with facility information of the same content (name "Shopping Mall A" and type "Family General Goods Store"), and the facility information alone does not identify a single combination of the configuration information. Therefore, the determination unit 13 determines the configuration information by taking timing information into consideration as well.

[0049] In this embodiment, the facility information acquired by the acquisition unit 12 in step S11 is information on the facility name "Shopping Mall A" and the type "Family General Goods Store." Furthermore, the timing information acquired by the acquisition unit 12 in step S13 is information on the time period "daytime" and the season "spring." The determination unit 13 identifies a combination of configuration information corresponding to the facility information and timing information (in the example of FIG. 4 , the image "fun pop music" and the volume "slightly loud") and determines it as the configuration information.

[0050] Furthermore, in this embodiment, the determination unit 13 determines the length of the song as the configuration information. In this embodiment, the determination unit 13 determines the length of the song based on timing information. Specifically, the timing information acquired by the acquisition unit 12 in step S13 includes information on the time period "daytime." Therefore, the determination unit 13 determines, for example, the number of hours of a time period corresponding to daytime as the length of the song. As an example, the determination unit 13 may acquire information such as sunrise and sunset times and facility opening hours from the database 30 or information on the Internet, and determine the length of the song (the number of hours of a time period corresponding to daytime) based on this information. In this embodiment, the determination unit 13 determines the length of the song to be "10 hours" as the configuration information.

[0051] In step S15, the generation unit 14 generates a prompt for instructing the AI ​​model 41 to create a song based on the configuration information. Generating a prompt based on the configuration information may mean generating a prompt including at least one piece of configuration information. The AI ​​model 41 creates a song based on the input prompt. The generation unit 14 generates a prompt so as to satisfy conditions for a prompt that can be input to the AI ​​model 41. The conditions for the prompt may be, for example, the number of tokens or a format (language, etc.). The generation unit 14 may generate a prompt using a generative AI model or another trained model. The generation unit 14 may acquire a prompt format stored in advance in the database 30, etc., and generate a prompt by incorporating the configuration information and a user's request regarding the prompt acquired via the user terminal 20 into the format. The method for generating the prompt is not limited.

[0052] 5 shows an example of a prompt generated by the generation unit 14. The prompt includes a task (instructions for the AI ​​model 41) to be executed by the AI ​​model 41, conditions, etc. In this example, the prompt includes a background that the AI ​​model 41 should take into consideration (the role of the AI ​​model 41), a task to be executed by the AI ​​model 41, and conditions under which the AI ​​model 41 executes the task.

[0053] The description of the background (role of the AI ​​model 41) that the AI ​​model 41 should take into consideration, for example, improves the expertise of the answer (output information) by the AI ​​model 41. In this example, the prompt includes the sentence "You are a musician" as a sentence specifying the role of the AI ​​model 41. This causes the AI ​​model 41 to attempt to create an answer that reflects the position and knowledge of the AI ​​model 41 as a musician.

[0054] The task to be performed by the AI ​​model 41 is to create a song, and in this example, the prompt contains the instruction "Create a song that meets the conditions" as a statement instructing the task.

[0055] The conditions under which the AI ​​model 41 executes the task are the conditions under which the AI ​​model 41 composes a song. The generation unit 14 includes the configuration information determined in step S14 as conditions in the prompt. Specifically, the generation unit 14 includes the configuration information determined in step S14, such as the image of "fun pop music," the volume of "slightly loud," and the length of the song of "10 hours," in the prompt.

[0056] The generation unit 14 may include the user information acquired in step S12 in the prompt. In this example, the generation unit 14 includes the average stay time of "2 hours" as a condition in the prompt. The generation unit 14 may include the timing information acquired in step S13 in the prompt. In this example, the generation unit 14 includes the season of "spring" as a condition in the prompt.

[0057] As described above, information that serves as a reference for the task to be performed by the AI ​​model 41 can be input to the AI ​​model 41 along with a prompt. In this example, a song data file that serves as a reference for the task of composing music is input to the AI ​​model 41 along with the prompt. The generation unit 14 includes information that enables the AI ​​model 41 to identify the reference song data file in the prompt. In this example, the generation unit 14 includes the file names of the reference song data files ("Song Title A", "Song Title B") in the prompt.

[0058] In step S16, the creation unit 15 creates a song using the AI ​​model 41. The creation unit 15 inputs the prompt generated in step S15 and a reference song data file as input information to the AI ​​model 41. The AI ​​model 41 creates a song based on the input prompt. The AI ​​model 41 creates a song in accordance with the roles, tasks, and conditions included in the prompt.

[0059] In this example, the AI ​​model 41 creates a song that meets the requirements of a musician. As an example, the AI ​​model 41 creates a song with a melody and tempo that matches the image of "fun pop music" (a song that gives the user the impression of a fun pop song). The AI ​​model 41 determines a volume value (e.g., information expressed by a specific numerical value in decibels) corresponding to a volume of "slightly loud." The AI ​​model 41 creates a song with a melody and tempo that will not tire the user during the average stay time (2 hours). The AI ​​model 41 creates a song that is 10 hours long (a continuous song without any breaks over the course of 10 hours). The AI ​​model 41 creates a song that allows the user to imagine a landscape of the season "spring." The AI ​​model 41 creates a song similar to the song "Song Title A" and song "Song Title B" input along with the prompt, for example, by referring to these songs. The creation unit 15 acquires the created song as output information from the AI ​​model 41. That is, the creation unit 15 creates a song using the AI ​​model 41.

[0060] In step S17, the control unit 16 controls a music playback device (such as a speaker) installed in the facility to play the music created by the creation unit 15. At this time, the control unit 16 may play the music according to the volume determined by the AI ​​model 41.

[0061] At least one of the processes in steps S11 to S16 described above may be repeatedly executed. For example, the acquisition unit 12 may newly acquire at least one of facility information and user information after a predetermined time has elapsed since the previous acquisition, and the determination unit 13 may determine the configuration information based on at least one of the newly acquired facility information and user information. Furthermore, the generation unit 14 may generate a prompt based on the newly determined configuration information, and the creation unit 15 may create a song using the newly generated prompt. The control unit 16 may play the newly created song consecutively (without a break) with a song that is already being played.

[0062] As described above, the device 10 according to one aspect of the present disclosure includes an acquisition unit 12 that acquires facility information related to the facility and user information related to users of the facility, a determination unit 13 that determines composition information related to the composition of a song to be played at the facility based on at least one of the facility information and the user information, and a generation unit 14 that generates a prompt to instruct the AI ​​model 41 to create a song based on the composition information.

[0063] A prompt generation method according to one aspect of the present disclosure includes steps of acquiring facility information about a facility and user information about users of the facility, determining composition information about the composition of a song to be played at the facility based on at least one of the facility information and the user information, and generating a prompt for instructing an AI model 41 to create a song based on the composition information.

[0064] For example, when creating a song using an AI model (such as a generative AI model), in order for the AI ​​model to output the song the user desires, the user needs to appropriately create prompts to input to the AI ​​model. However, generating prompts can require advanced skills and a lot of time. For example, if a prompt that does not include clear instructions is used, the model may generate an unintended song (a song that does not reflect the user's wishes). In addition, prompts that can be input to the model may have conditions (such as the number of tokens and format), and it is not easy to generate prompts that take these conditions into account.

[0065] However, in the above-described aspects of the present disclosure, configuration information regarding the composition of a song to be played at a facility is determined based on at least one of facility information and user information, and a prompt for instructing the AI ​​model 41 to create a song is generated based on the configuration information. Therefore, a prompt that reflects at least one of the facility information and user information is generated by the device 10 without the user having to generate the prompt themselves. By inputting the generated prompt into the AI ​​model 41 to create a song, it is possible to create a song that reflects information regarding at least one of the facility and the users of the facility, for example. In other words, according to the aspects of the present disclosure, a song suitable for a facility can be easily created.

[0066] The composition information includes information on at least one of the tempo, melody, volume, genre, image, and length of the song, which allows a song to be created that has at least one of the tempo, melody, volume, genre, image, and length that is suitable for the facility.

[0067] The acquisition unit 12 acquires timing information regarding the timing at which the song will be played, and the determination unit 13 determines the composition information based on the timing information, thereby enabling the creation of a song suited to the timing (time period, season, etc.) at which the song will be played.

[0068] [Modification] The present disclosure is not limited to the above embodiment. For example, in step S14, the determination unit 13 may determine the configuration information using an AI model (such as a generative AI model). For example, the determination unit 13 may input at least one of the acquired facility information and user information as input information to the AI ​​model and acquire information corresponding to the configuration information from the AI ​​model. The AI ​​model may be a trained model that has learned the correspondence between at least one of the facility information and user information and the configuration information.

[0069] In step S14, the determination unit 13 may determine the composition information based only on the facility information, or may determine the composition information based only on the user information. The determination unit 13 may determine the composition information based on the attributes of the majority (attributes that include many users) among the attributes of users present at the facility. For example, if most of the users present at the facility are children, the determination unit 13 may determine the composition information so that a song suitable for children is created.

[0070] The determination unit 13 may determine the configuration information based on, for example, environmental conditions in the vicinity of the facility, or the type and atmosphere of the facility, etc. As an example, if it is preferable to lower the volume, such as when a house is located near the facility or when the facility is an art museum, the determination unit 13 may determine the volume included in the configuration information to be "low."

[0071] In step S11, the acquisition unit 12 may acquire information about the business hours of the facility as facility information, and in step S14, the determination unit 13 may determine the composition information based on the information about the business hours. For example, the determination unit 13 may determine, based on the information about the business hours of the facility and the timing information, during which time period the song will be played during the business hours of the facility. If the determination unit 13 determines that the song will be played just before the end of the business hours of the facility, the determination unit 13 may determine, as the composition information, a tempo or melody of the song that encourages facility users to change their behavior (to leave the facility). This makes it possible to create a song that takes the business hours of the facility into consideration.

[0072] In step S12, the acquisition unit 12 acquires information regarding the average length of stay of users at the facility as user information, and in step S14, the determination unit 13 may determine the composition information based on the information regarding the average length of stay of users. For example, the determination unit 13 may determine a length of music (composition information) that is longer than the average length of stay of users. This makes it possible to create music that takes into account the average length of stay of users at the facility. For example, it is possible to provide many users with music without interruption throughout their stay at the facility.

[0073] In step S12, the acquisition unit 12 may acquire information regarding the degree of congestion of users in the facility as user information, and in step S14, the determination unit 13 may determine the configuration information based on the information regarding the degree of congestion. For example, when the degree of congestion is high, the determination unit 13 may determine the volume of the music included in the configuration information to be higher than a predetermined value. A case where the degree of congestion is high may be, for example, when the density of users in the facility exceeds a predetermined value or when the number of users in the facility exceeds a predetermined number. When the degree of congestion is low, the determination unit 13 may determine the volume of the music included in the configuration information to be lower than a predetermined value. A case where the degree of congestion is low may be, for example, when the density of users in the facility is lower than a predetermined value or when the number of users in the facility is lower than a predetermined number. In a facility with a high degree of congestion (a large number of users), the density of users per unit space is high, and it may be difficult to hear the music being played. However, according to this aspect, music at a volume appropriate for the situation in the facility can be provided to users.

[0074] In step S11, the acquisition unit 12 may acquire information regarding the volume of environmental sound in the facility as facility information, and in step S14, the determination unit 13 may determine the configuration information based on the information regarding the volume of environmental sound. For example, if the environmental sound is loud, the determination unit 13 may determine the volume of the music included in the configuration information to be higher than a predetermined value. A case where the environmental sound is loud may be, for example, a case where the volume of the environmental sound in the facility exceeds a predetermined value. If the environmental sound is quiet, the determination unit 13 may determine the volume of the music included in the configuration information to be lower than a predetermined value. A case where the environmental sound is quiet may be, for example, a case where the volume of the environmental sound in the facility is lower than a predetermined value. In a facility with loud environmental sound, it may be difficult to hear the music being played, but according to this aspect, it is possible to provide the user with music at a volume appropriate to the situation in the facility.

[0075] The acquisition unit 12 may acquire weather information regarding the weather in an area where the facility is located. The determination unit 13 may determine the configuration information based on the weather information. The acquisition unit 12 may acquire weather information regarding the time when the song will be played based on the timing information. The weather information may be information regarding at least one of the weather (sunny, cloudy, rainy, snowy, etc.) and the temperature. The acquisition unit 12 may acquire weather information from the user terminal 20. For example, if a user wishes to create a song to be played when it is raining, the user operates the user terminal 20 to input information indicating that the weather is raining into the user terminal 20 as weather information. The acquisition unit 12 may acquire the input weather information from the user terminal 20. The acquisition unit 12 may acquire weather information from the database 30. The acquisition unit 12 may acquire information on the Internet as weather information. For example, the acquisition unit 12 may acquire weather information (weather forecast) corresponding to the time period input as timing information from the Internet. This makes it possible to create a song suited to the weather.

[0076] If a device (such as a camera) that captures images or videos of the facility is installed in the facility, the acquisition unit 12 may acquire at least one of the facility information and user information based on at least one of the captured images and videos of the facility. For example, the acquisition unit 12 may acquire images or videos of the facility in real time from a camera installed in the facility and analyze the images or videos to acquire the degree of user congestion, the average length of time users stay, etc. This makes it possible to acquire more accurate facility information and user information.

[0077] The composition system 1 is not limited to the configuration shown in FIG. 2 . In the configuration shown in FIG. 2 , the AI ​​model 41 may be located in the device 10. As another example, as shown in FIG. 6 , at least one of the functional elements of the device 10 may be located in the user terminal 20. The user terminal 20 may have the functions of the device 10 and function as the device 10. That is, the device 10 may be included in the user terminal 20. This configuration can be realized, for example, by installing an application that executes the functions of the device 10 in the user terminal 20. In this configuration, the AI ​​model 41 is located on a network (e.g., a cloud) like ChatGPT. In this configuration, the RAG app may be provided in the user terminal 20. Information (knowledge DB) accessed by the RAG may be located on the network.

[0078] As another configuration, as shown in FIG. 7 , not only the functional elements of the device 10 but also the AI ​​model 41 may be located in the user terminal 20. That is, the device 10 and the AI ​​model 41 may be configured to be included in the user terminal 20. This configuration can be realized, for example, by installing an application that executes the functions of the device 10 and an application that executes the functions of the AI ​​model 41 in the user terminal 20. In this configuration, the AI ​​model 41 is located inside the user terminal 20, such as tsuzumi, a type of large-scale language model (LLM). In this configuration, the RAG app may be provided in the user terminal 20. The information (knowledge DB) accessed by the RAG may be located inside the user terminal 20 and / or on a network. In any of the configurations shown in FIGS. 2 , 6 , and 7 , an external server (e.g., an internal server of a company) that can be used to obtain information related to business content may be located outside (e.g., on a network).

[0079] The generation unit 14 may include multiple pieces of the same type of information (e.g., composition information such as the tempo, melody, genre, and image of the song) in the prompt. When multiple pieces of information are included in the prompt, the generation unit 14 may include at least one piece of degree information in the prompt to determine the extent to which the AI ​​model 41 should consider each piece of information (the degree to which it relies or emphasizes it). For example, when the generation unit 14 includes two pieces of information about the song genre, "techno" and "jazz," the generation unit 14 may include in the prompt as degree information a percentage indicating the degree to which each genre should be considered. Specifically, the degree information may be expressed as a numerical value in percentages, such as "techno 70%" and "jazz 30%." This allows the AI ​​model 41 to create a techno song incorporating jazz elements. By including the degree information in the prompt in this way, the generation unit 14 can generate a prompt that includes instructions and conditions for creating a complex song that sounds as if it were conceived by a human. The degree information may be information received from the user (for example, information received by the reception unit 11 from the user terminal 20), or may be information determined by the generation unit 14 based on information previously registered in the database 30 by the operator of the composition system 1, etc.

[0080] The device and method of the present disclosure have the following configuration.

[0081] [1] An apparatus comprising: an acquisition unit that acquires at least one of facility information regarding a facility and user information regarding users of the facility; a determination unit that determines composition information regarding the composition of a song to be played at the facility based on at least one of the facility information and the user information; and a generation unit that generates a prompt to instruct an AI model to create the song based on the composition information.

[0082] [2] The device according to [1], wherein the composition information includes information regarding at least one of the tempo, melody, volume, genre, image, and length of the song.

[0083] [3] The device according to [1] or [2], wherein the acquisition unit further acquires timing information regarding a timing at which the song is played, and the determination unit determines the composition information based on the timing information.

[0084] [4] The device according to any one of [1] to [3], wherein the facility information includes information about business hours of the facility, and the determination unit determines the configuration information based on the information about business hours.

[0085] [5] The device described in any one of [1] to [4], wherein the user information includes information regarding an average stay time of the user at the facility, and the determination unit determines the configuration information based on the information regarding the average stay time.

[0086] [6] The device described in any one of [1] to [5], wherein the user information includes information regarding the degree of congestion of the users in the facility, and the determination unit determines the configuration information based on the information regarding the degree of congestion.

[0087] [7] The device according to any one of [1] to [6], wherein the facility information includes information about the volume of environmental sound in the facility, and the determination unit determines the configuration information based on the information about the volume of environmental sound.

[0088] [8] The device according to any one of [1] to [7], wherein the acquisition unit further acquires weather information relating to weather in an area where the facility is located, and the determination unit determines the configuration information based on the weather information.

[0089] [9] The device according to any one of [1] to [8], wherein the acquisition unit acquires at least one of the facility information and the user information based on at least one of an image and a video taken at the facility.

[0090]

[10] A method comprising: acquiring at least one of facility information regarding a facility and user information regarding users of the facility; determining composition information regarding the composition of a song to be performed at the facility based on at least one of the facility information and the user information; and generating prompts to instruct an AI model to create the song based on the composition information.

[0091] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (e.g., via wire, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or multiple devices with software.

[0092] Functions include, but are not limited to, judgment, determination, discrimination, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, regard, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.

[0093] 8 is a diagram showing an example of the hardware configuration of a composition system 1 (prompt generation system) according to this embodiment. The composition system 1 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage device 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.

[0094] In the following description, the term "apparatus" may be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the composition system 1 may be configured to include one or more of the apparatuses shown in the drawings, or may be configured to exclude some of the apparatuses.

[0095] Each function in the composition system 1 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.

[0096] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, at least one of the functional units of the composition system 1 described above may be realized by the processor 1001.

[0097] The processor 1001 also reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with the programs. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, at least one of the functional units of the composition system 1 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be used for other functional blocks. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.

[0098] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be referred to as a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing the accompanying determination method according to one embodiment of the present disclosure.

[0099] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.

[0100] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, network controller, network card, communication module, etc. The communication device 1004 may include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to implement at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, at least one of the functional units of the composition system 1 described above may be implemented by the communication device 1004. The communication device 1004 may be implemented with a transmitter and a receiver that are physically or logically separated.

[0101] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that accepts input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. Note that the input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).

[0102] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.

[0103] The composition system 1 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.

[0104] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.

[0105] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0106] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.

[0107] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0108] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).

[0109] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.

[0110] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0111] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.

[0112] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0113] Note that terms described in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Furthermore, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.

[0114] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed using absolute values, may be expressed using relative values ​​from a predetermined value, or may be expressed using other corresponding information. For example, a radio resource may be indicated by an index.

[0115] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.

[0116] In this disclosure, the terms "Mobile Station (MS)," "user terminal," "User Equipment (UE)," "terminal," and the like may be used interchangeably.

[0117] A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terminology.

[0118] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.

[0119] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.

[0120] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0121] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.

[0122] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.

[0123] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.

[0124] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."

[0125] 1...composition system, 10...device, 11...reception unit, 12...acquisition unit, 13...determination unit, 14...generation unit, 15...creation unit, 16...control unit, 20...user terminal, 30...database, 40...server device, 41...AI model, 1001...processor, 1002...memory, 1003...storage, 1004...communication device, 1005...input device, 1006...output device, 1007...bus, N...communication network.

Claims

1. An apparatus comprising: an acquisition unit that acquires at least one of facility information regarding a facility and user information regarding users of the facility; a determination unit that determines composition information regarding the composition of a song to be played at the facility based on at least one of the facility information and the user information; and a generation unit that generates prompts to instruct an AI model to create the song based on the composition information.

2. The device according to claim 1, wherein the composition information includes information regarding at least one of the tempo, melody, volume, genre, image, and length of the song.

3. The device according to claim 1, wherein the acquisition unit further acquires timing information regarding a timing at which the song is played, and the determination unit determines the configuration information based on the timing information.

4. The device according to claim 1, wherein the facility information includes information about business hours of the facility, and the determination unit determines the configuration information based on the information about business hours.

5. The device according to claim 1, wherein the user information includes information relating to an average length of stay of the user at the facility, and the determination unit determines the configuration information based on the information relating to the average length of stay.

6. The device according to claim 1, wherein the user information includes information relating to the degree of congestion of the users in the facility, and the determination unit determines the configuration information based on the information relating to the degree of congestion.

7. The device according to claim 1, wherein the facility information includes information relating to the volume of environmental sound in the facility, and the determination unit determines the configuration information based on the information relating to the volume of environmental sound.

8. The device according to claim 1, wherein the acquisition unit further acquires weather information relating to the weather in an area where the facility is located, and the determination unit determines the configuration information based on the weather information.

9. The device according to claim 1, wherein the acquisition unit acquires at least one of the facility information and the user information based on at least one of an image and a video taken at the facility.

10. A method comprising the steps of: acquiring at least one of facility information regarding a facility and user information regarding users of the facility; determining composition information regarding the composition of a song to be performed at the facility based on at least one of the facility information and the user information; and generating prompts to instruct an AI model to create the song based on the composition information.

Citation Information

Patent Citations

  • Electronic musical instrument

    JP1991265899A

  • Remote controller with interphone function

    JP2002058078A

  • Remote control unit

    JP2006029634A

  • On-vehicle music producing device, and on-vehicle entertainment system

    JP2006069288A

  • Content generation device and content generation method

    JP2006084749A