Method, system, computer-readable recording medium, and program
By identifying motion features in a video and using AI to generate music with a beat-based unit of time, the method addresses the lack of sound creation from moving images, resulting in synchronized and industrially applicable music.
Patent Information
- Application Number
- JP2025084183
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2044-11-26
AI Technical Summary
There is no method for creating sound based on moving images.
A method is provided for identifying motion features in a video and using artificial intelligence to generate music with the time between similar motion features as a unit of time, or beat, to synchronize sound with the video.
This configuration allows for the generation of industrially applicable music that matches the video, reducing the workload of sound creators and enabling automatic generation of highly usable sound.
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for creating sound. [Background technology]
[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.
[0003] Patent Document 1 discloses a music generation system. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-154777 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the inventors have recognized that at least the above embodiment has a drawback in that there is no method for creating sound based on moving images. [Means for solving the problem]
[0006] At least one aspect of the present disclosure provides a method for manufacturing a semiconductor device, comprising: 1. A method of producing sound, comprising: The device, Identifying motion features that characterize the behavior of objects in a video; The AI is required to generate music with the time between similar motion features as a unit of time, with the unit time being a beat. method to provide. [Effects of the Invention]
[0007] This configuration has the advantage of being able to generate industrially applicable music based on video.
[0008] These and other aspects, features, and advantages of the present disclosure will become apparent from the following detailed written description of the preferred embodiments and aspects taken in conjunction with the following drawings, variations and modifications of which may be made without departing from the spirit and scope of the novel concepts of the present disclosure. Aspects of one embodiment of the present disclosure may be combined with or substituted for one or more aspects of another embodiment of the present disclosure, unless inconsistent. DETAILED DESCRIPTION OF THE INVENTION
[0009] The following disclosure provides many different embodiments and examples for implementing different features of the presented subject matter. To simplify the disclosure, specific examples of components and arrangements are disclosed below. Of course, these are merely examples and are not intended to be limiting. For example, a structure in which a first feature is covered by or in contact with a subsequently disclosed second feature may include embodiments in which the first and second features are formed in direct contact, as well as embodiments in which an additional feature is formed between the first and second features to prevent direct contact between the first and second features. Furthermore, the disclosure may repeat reference numbers and / or letters in various examples. Such repetition is for the purposes of brevity and clarity and does not, in itself, require a relationship between the various embodiments and / or configurations described. Furthermore, when a first element is described as being "coupled" or "coupled" to a second element, such description includes embodiments in which the first and second elements are directly coupled or coupled to each other, as well as embodiments in which the first and second elements are indirectly coupled or coupled to each other with one or more other intervening elements therebetween.
[0010] As used herein, the phrase "at least one of" encompasses all exemplified variations. For example, the phrase "comprises at least one of A, B, or C" is equivalent to "consisting of A, B, and C and combinations thereof," and encompasses all possible variations of A, B, C, A+B, A+C, B+C, and A+B+C.
[0011] In this disclosure, disclosures using a machine, an electronic operator, or a computer may include embodiments of a method, a recording medium, an apparatus, or a program. As used herein, the statement "A is B" can be replaced with "A includes B" unless there is a contradiction or unless otherwise stated in this specification.
[0012] The terms in this disclosure, including the terms set forth in the claims, may be interpreted in light of the descriptions and drawings set forth in the specification, and further, unless otherwise indicated inconsistently, by what one or more citizens, past, present, or future, have so called, so designated, so understood, or so performed, or may so do, unless otherwise indicated in the present disclosure. The operating method used in at least one or more embodiments can take the following embodiments: The following description will be made with reference to JP6456303 (the following reference begins), which clearly explains at least one or more embodiments.
[0013] As used herein, the term "computer" generally includes, as known in the art, a processor; memory; at least one information storage / retrieval device, such as a hard drive, disk drive, or flash drive or memory stick, or other non-transitory computer-readable medium or non-transitory storage device; at least one input device, such as a keyboard, mouse, point-and-touch device, touch screen, or microphone; and a display structure, such as a well-known computer screen. In addition, a computer may include one or more network connections, such as wired or wireless connections. As known in the art, such a computer or computer system may include more or less of the above, including, for example, but not limited to, tablet computers and smart devices, as well as other electronic media and devices.
[0014] As used herein, the terms "cloud" or "cloud computing" refer to a centralized and virtualized computing facility in which all computing resources are shared. Application systems and subsystems can no longer be referred to as specific machines because they are all in the "cloud."
[0015] As used herein, the term "distributed Internet service system" refers to a distributed Internet service platform that transforms Internet applications to run in various computing environments. The DIS system distributes Internet applications, including content, data, and logic, to whatever extent appropriate and along the network to any number and type of devices via a Component Distribution Server / Asset Distribution Server. Through the DIS, Internet applications can be hosted and centrally managed, with services based on each user's needs, and cached and executed locally on the user's device or nearby locations while maintaining their integrity. Web-enabled computing devices can be upgraded with DIS software to become DIS-enabled, enjoying and running distributed Internet services. The distributed internet services system is more fully described in any one of the following patent families: U.S. Patent Nos. 7,136,857, 7,150,015, 7,181,731, 7,209,921, 7,430,610, 7,685,183, 7,685,577, 7,752,214, 8,326,883, 8,386,525, 8,443,035, 8,458,142, 8,458,222, 8,473,468, 8,527,545, and 8,650,226, and U.S. Patent Publication Nos. 20120005205, and 20130091252, all of which, like the present invention, are commonly owned by OP40 Holdings, Inc., and all of which are incorporated by reference. (End of quote)
[0016] The operating method used in at least one embodiment can be implemented in the following manner using a conventional Internet system that does not use a distributed Internet. Reference is made to JP7113047 (the following reference begins), which clearly explains at least one embodiment.
[0017] Embodiments including those specifically disclosed in this specification can provide an automated response system that is based on artificial intelligence and is implemented in a manner that resembles a real conversation with a human, thereby enabling more natural conversations with users and quickly and conveniently handling inquiries, reservations, delivery orders, etc.
[0018] The electronic devices 110, 120, 130, and 140 may be fixed or mobile terminals implemented by computer systems. Examples of the electronic devices 110, 120, 130, and 140 include AI speakers, smartphones, mobile phones, navigation systems, personal computers (PCs), laptop PCs, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), tablets, game consoles, wearable devices, internet of things (IoT) devices, virtual reality (VR) devices, and augmented reality (AR) devices. While FIG. 1 illustrates an AI speaker as the electronic device 110, in embodiments of the present invention, the electronic device 110 may represent one of a variety of physical computer systems capable of communicating with other electronic devices 120, 130, and 140 and / or servers 150 and 160 via a network 170 using a substantially wireless or wired communication method.
[0019] The communication method is not limited, and may include not only communication methods using communication networks (such as a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, and a satellite network) that can be included in network 170, but also short-range wireless communication between devices. For example, network 170 may include any one or more of networks such as a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a broadband network (BBN), and the Internet. Furthermore, network 170 may include any one or more of network topologies including, but not limited to, a bus network, a star network, a ring network, a mesh network, a star-bus network, a tree network, or a hierarchical network.
[0020] The servers 150 and 160 may each be realized by one or more computer devices that communicate with the multiple electronic devices 110, 120, 130, and 140 via the network 170 and provide instructions, codes, files, content, services, etc. For example, the server 150 may be a system that provides a first service to the multiple electronic devices 110, 120, 130, and 140 connected via the network 170, and the server 160 may be a system that provides a second service to the multiple electronic devices 110, 120, 130, and 140 connected via the network 170. As a more specific example, the server 150 may provide a service (such as an auto-answer service, for example) targeted by an application, which is a computer program installed and executed in the multiple electronic devices 110, 120, 130, and 140, as a first service to the multiple electronic devices 110, 120, 130, and 140. As another example, the server 160 may provide, as a second service, a service of distributing files for installing and executing the above-mentioned application to the multiple electronic devices 110, 120, 130, and 140.
[0021] 2 is a block diagram illustrating the internal configuration of an electronic device and a server according to an embodiment of the present invention. In FIG. 2, the internal configuration of electronic device 110 and the internal configuration of server 150 are described as examples of electronic devices. Furthermore, other electronic devices 120, 130, 140 and server 160 may also have the same or similar internal configuration as electronic device 110 or server 150 described above.
[0022] The electronic device 110 and the server 150 may include memories 211 and 221, processors 212 and 222, communication modules 213 and 223, and input / output interfaces 214 and 224. The memories 211 and 221 may be non-transitory computer-readable recording media and may include non-transitory mass storage devices such as random access memory (RAM), read-only memory (ROM), a disk drive, a solid state drive (SSD), and flash memory. The non-transitory mass storage devices such as ROM, SSD, flash memory, and disk drive may be included in the electronic device 110 and the server 150 as separate non-transitory storage devices distinct from the memories 211 and 221. The memories 211 and 221 may also store an operating system and at least one program code (e.g., code for a browser installed and executed on the electronic device 110, or code for an application installed on the electronic device 110 to provide a particular service). Such software components may be loaded from a computer-readable recording medium separate from the memories 211 and 221. Such other computer-readable recording media may include computer-readable recording media such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. In other embodiments, software components may be loaded into memory 211, 221 through communication modules 213, 223 that are not computer-readable recording media. For example, at least one program may be loaded into memory 211, 221 based on a computer program (such as the above-mentioned application) being installed by a file provided over network 170 by a developer or a file distribution system that distributes application installation files (such as the above-mentioned server 160, for example).
[0023] The processors 212, 222 may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 212, 222 by the memories 211, 221 or the communication modules 213, 223. For example, the processors 212, 222 may be configured to execute instructions received according to program code stored in a storage device such as the memories 211, 221.
[0024] The communication modules 213 and 223 may provide a function for the electronic device 110 and the server 150 to communicate with each other via the network 170, or may provide a function for the electronic device 110 and / or the server 150 to communicate with other electronic devices (for example, the electronic device 120) or other servers (for example, the server 160). For example, a request generated by the processor 212 of the electronic device 110 in accordance with program code recorded in a recording device such as the memory 211 may be transmitted to the server 150 via the network 170 under the control of the communication module 213. Conversely, a control signal, instruction, content, file, etc. provided under the control of the processor 222 of the server 150 may be received by the electronic device 110 via the communication module 213 of the electronic device 110 via the communication module 223 and the network 170. For example, control signals, instructions, content, files, etc. from the server 150 received through the communication module 213 may be transmitted to the processor 212 or memory 211, and the content, files, etc. may be recorded on a recording medium (the non-transitory recording device described above) that the electronic device 110 may further include.
[0025] The input / output interface 214 may be a means for interfacing with the input / output device 215. For example, the input device may include a keyboard, a mouse, a microphone, a camera, etc., and the output device may include a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface 214 may be a means for interfacing with a device that integrates input and output functions into one, such as a touchscreen. The input / output device 215 may be configured as a single device together with the electronic device 110. Furthermore, the input / output interface 224 of the server 150 may be a means for interfacing with an input or output device (not shown) that may be connected to or included in the server 150. As a more specific example, when the processor 212 of the electronic device 110 processes instructions of a computer program loaded in the memory 211, a service screen or content configured using data provided by the server 150 or the electronic device 120 may be displayed on a display via the input / output interface 214.
[0026] In other embodiments, the electronic device 110 and the server 150 may include more components than those shown in FIG. 2 . However, it is not necessary to explicitly illustrate most of the conventional components. For example, the electronic device 110 may be implemented to include at least some of the input / output devices 215 described above, and may further include other components such as a transceiver, a camera, various sensors, a database, etc. As a more specific example, if the electronic device 110 is an AI speaker, the electronic device 110 may be implemented to further include various components typically included in AI speakers, such as various sensors, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration. (End of quote)
[0027] According to at least one embodiment, a user terminal includes a control unit, a RAM, a storage unit, a graphics processing unit, a communication interface, and an interface unit, each connected by an internal bus.
[0028] According to at least one embodiment, the control unit is composed of a CPU and a ROM. The control unit executes programs stored in the storage unit and controls the user terminal. The RAM is the work area of the control unit. The storage unit is a memory area for saving programs and data. The control unit reads and processes the programs and data from the RAM. The control unit processes the programs and data loaded into the RAM and outputs drawing commands to the graphics processing unit.
[0029] According to at least one embodiment, the graphics processing unit is connected to a display unit. The display unit has a display screen. When the control unit outputs a drawing command to the graphics processing unit, the graphics processing unit outputs a video signal for displaying an image on the display screen. Here, the display unit may be a touch panel equipped with a touch sensor. The touch panel of the display unit functions as an input unit.
[0030] According to at least one embodiment, the communication interface can be connected to a communication network wirelessly or via a wire, and can transmit and receive data to and from a server device via the communication network. Data received via the communication interface is loaded into RAM, and processed by the control unit. An external memory (e.g., an SD card) is connected to the interface unit.
[0031] According to at least one embodiment, the user terminal is a computing device having a display screen and an input section, but is not limited thereto. Examples of user terminals include conventional mobile phones, tablet devices, smartphones, and desktop or laptop personal computers. VR goggles may also be configured with a screen (or two display panels, one for each eye) attached to a frame (or headset) strapped or attached to the head. The user terminal also has an audio output section.
[0032] According to at least one embodiment, the user terminal is communicatively connected to the server device via a communications network, and is capable of transmitting or receiving information via the communications network.
[0033] According to at least one embodiment, the server device includes at least a control unit, a RAM, a storage unit, and a communication interface, which are connected to each other by an internal bus.
[0034] According to at least one embodiment, the control unit is composed of a CPU and a ROM, executes a program stored in the storage unit, and controls the server device. The control unit also has an internal timer for measuring time. The RAM is the work area of the control unit. The storage unit is a memory area for saving programs and data. The control unit reads the program and data from the RAM and performs program execution processing based on information received from the user terminal, etc.
[0035] This document discloses AI. According to at least one embodiment, artificial intelligence includes machine learning, deep learning, generative AI, large-scale language models, LLMs, foundational models, and generative AI. Generative AI uses transformers and multiple mechanisms called attention. It employs self-supervised learning and extract prediction. In this case, the AI can guess the next word. Given a sentence, it guesses the next word based on the sentence up to that point. A large number of supervised learning problems are created. These techniques result in an AI that can guess the next word. Generative AI can predict grammatical structure, topic connections, and the type of sentence a person with a certain writing style is likely to write. Furthermore, simply by guessing the next sentence, generative AI can learn the underlying structure, causal relationships, and knowledge. Generative AI is scalable, and the larger the number of parameters, the higher its accuracy. In conventional statistics and machine learning, overfitting can occur if the model parameters are too large compared to the sample size of the data. In LLM, the higher the number of parameters, the higher its accuracy. One generative AI has 175 billion parameters and uses supervised learning to ensure smooth dialogue. They are instructed not to say anything strange. They write reviews and act as call center operators.
[0036] According to at least one embodiment, a large language model (LLM) is a machine learning natural language processing model built non-comprehensively using large datasets and deep learning techniques. Typically, a task-specific training technique known as "fine-tuning" is used to adapt LLMs to various natural language processing (NLP) tasks, such as text classification and generation, sentiment analysis, text summarization, and question answering. According to at least one embodiment, self-supervised learning is similar to intrinsic human intelligence. When humans perform actions, they constantly predict the next event and the next input. In the process, they learn the structure of the external world. Predicting the next word is considered intrinsic intelligence and is similar to what the cerebral cortex does. According to at least one embodiment, a large language model memorizes all input information but generalizes only to the extent necessary to predict the next word. It does not attempt to generalize all information from the beginning. A large language model requires capacity to memorize information. This also requires parameters. According to at least one embodiment, the large-scale language model includes eight models with 175 billion parameters and eight models with 220 billion parameters.
[0037] According to at least one embodiment, we represent videos and images as a collection of visual patches, which are small units of data similar to text tokens in LLMs. Patches effectively represent models of visual data and serve as a highly scalable and effective representation for training generative models on a wide variety of videos and images. We convert videos into patches by first compressing the video into a low-dimensional latent space and then decomposing the representation into spatiotemporal patches.
[0038] According to at least one embodiment, a video compression network is a network that reduces the dimensionality of visual data, taking raw video as input and outputting a temporally and spatially compressed latent representation. An AI is trained on this compressed latent space and then generates video within this compressed latent space.
[0039] According to at least one embodiment, Spacetime Latent Patches extracts a set of spacetime patches that act as Transformer tokens given a compressed input video. The patch-based representation allows Sora to be trained on videos and images of various resolutions, lengths, and aspect ratios, and controls the size of the generated video by placing randomly initialized patches on an appropriately sized grid during inference.
[0040] According to at least one embodiment, the AI is a diffusion model, trained to predict the original "clean" patch when fed a noisy patch (and conditioning information such as a text prompt). The AI is a diffusion transformer, which exhibits remarkable scaling properties in a variety of domains, including language modeling, computer vision, and image generation. Diffusion transformers are also effective as video generation models. The AI significantly improves the quality of the samples as the training computational effort increases.
[0041] According to at least one embodiment, the AI applies caption regeneration techniques to train a highly descriptive caption model and then use it to generate text captions for all videos in a training set. Training highly descriptive captions improves the fidelity of the text as well as the overall quality of the generated videos. GPT is leveraged to convert short user prompts into long, detailed captions that are sent to the model. This allows the AI to generate high-quality videos that accurately follow the user prompts.
[0042] According to at least one embodiment, AI can perform vectorization in natural language processing according to the following process. First, a preprocessing step is performed on the given text. This preprocessing step involves removing unnecessary words, such as JavaScript code and HTML tags, from the text. These codes are used for displaying information on the Internet and are therefore not generally used in natural language processing. Next, the text is divided into words using morphological analysis. Morphological analysis is the process of classifying natural language sentences written in text into the smallest meaningful linguistic units. Morphological analysis tools available include "MeCab," "JUMAN," and "JANOME." Normalization involves unifying words with the same meaning, such as spelling variations, into a single word. Stop words are words that are not processed because they cannot be used in natural language processing. Examples of stop words include particles and auxiliary verbs, which have no meaning on their own. When calculating vectors, these words may be removed, leaving only meaningful words. Vectorization may also be performed without removing these stop words. Vectorization is the process of converting words, which are strings of characters, into vectors. Vectorization converts word data into numerical data. Converting words into vectors is done using methods called bag of words or distributed representations. Bag of words is a method of vectorizing a sentence using the number of words that appear in a given sentence. It focuses on how often a word appears in a sentence, and does not take into account the order of the words or sentences. Distributed representation is a vectorization method that focuses on the meaning of words. By vectorizing the meaning of words, it is possible to assign vectors that are similar to words with similar meanings or usages, and the relationships between words can also be expressed as vectors. Representation as vectors makes it possible to add and subtract word meanings from each other. In applied processing, natural language converted into numerical data can be used as input for machine learning. Specifically, vectorized natural language is input into a classifier to classify the sentences.Tools used here include "TensorFlow," "scikit-learn," and "PyTorch."
[0043] This disclosure relates to sound. In at least one embodiment, sound includes music or sound effects. Music expresses various emotions and stories by combining the duration, pitch, intensity, and timbre of sounds, and includes singing, musical instrument performances, and natural sounds. Music is composed by determining a scale, chords, or melody. In at least one embodiment, once a scale is determined, chords are also determined based on its constituent notes. There are theories about chord progressions, and there are patterns that have been considered good progressions in songs to date. Chords greatly influence the mood of a song and can have a significant impact on a person's emotions and perceptions at that time. Similarly, melodies do not deviate significantly if based on the constituent notes of the chords, but tend to become monotonous, so a certain amount of randomness may be necessary. Varying the chords within a range that does not deviate significantly can sometimes prevent monotony.
[0044] A method for creating sound, in which an apparatus identifies motion features characteristic of the motion of an object in a moving image, is disclosed (in this disclosure, the term "method" may be interpreted interchangeably with the term "steps" unless expressly stated to the contrary). In at least one embodiment, a moving image is data or data that is created by sequentially displaying still images. A moving image may include audio, and synchronization of the video and audio may enhance the expression. An object is data or related processes manipulated within the moving image. Examples of objects include moving objects in the moving image (including living and non-living objects; non-exhaustive examples of the former include people, people dancing, animals, and plants swaying in the wind; non-exhaustive examples of the latter include moving objects (e.g., cars), the ebb and flow of waves, flashing lights, changes in the wavelength or intensity of light, and digital representations drawn using computer graphics). In this disclosure, objects that a person perceives as moving when watching a moving image are considered to be "objects" or data corresponding to such objects. In this disclosure, "motion features" may refer to "motion features" (this applies throughout this disclosure unless explicitly stated to the contrary). Features are numerical values contained in a dataset that represent the characteristics of that data. For example, when representing a person's face in a video, features include "eye size," "nose height," and "skin color." In addition, when a person is dancing, the way their organs (including hands, feet, hips, etc.) move can also be used as features. By providing these numerical values to a computer, the computer can distinguish between people's faces, recognize specific people, and recognize dance choreography. In addition, in the ebb and flow of the ocean, the way the waves rise and fall can naturally be used as features. In computer graphics, the way a line or surface moves and the way its color changes can also naturally be used as features. Types of features include, non-inclusively: Numerical data is data that can be expressed numerically, such as height, weight, and age. Categorical data is data that can be expressed in categories, such as gender, nationality, and occupation. Text data is data that can be expressed as text, such as sentences, words, and keywords.Image data (including still images in video) are features extracted from images, such as color, shape, and texture. Video data are features extracted from video, such as color, shape, texture, and their presence or change over time. Audio data are features extracted from audio, such as pitch, frequency, and timbre.
[0045] In at least one embodiment, the device can be created by any of the methods, machines, or AI described above. In this disclosure, the term "device" may be replaced with "AI" unless expressly stated to the contrary. The device analyzes data from a target video. The device identifies an object in the video. The target object can be identified by a user. In another embodiment, the device identifies an object that appears relatively frequently among objects present in the video. The device can request an AI to perform such identification. (Throughout this disclosure, unless expressly stated to the contrary, the term "the device performs a certain process" may be replaced with "the device requests the AI to perform the process." In this case, the AI can be an AI that has learned the process or judgment using the learning method described above (including supervised learning).) In such cases, the target object is automatically identified by the device, rather than specified by a user. The device identifies a motion characteristic of the object. As an example, in a video of a dancer dancing, the dancer's hand waving is a motion characteristic of the object. In this embodiment, the movement of the dancer's head, hips, foot movement, and finger movement can also be set as movement features. As already explained, the movement features can be specified by the user. Furthermore, the device can also specify, as movement features, a movement that appears relatively frequently from among the movements of the object.
[0046] This disclosure discloses a method for requesting an AI to generate music with the time between similar movement features as a unit of time, where the unit time is a beat. In this disclosure, unless explicitly stated otherwise, the term "generate music" can be replaced with "generate sound." In at least one embodiment, a video contains similar movement features. As a non-exhaustive example, a dancer may repeatedly perform the same dance choreography. In this case, a device compares the movement features of each movement and recognizes them as similar. The criteria for determining similarity can be determined by a user or automatically by a machine. In the former case, the device displays videos or still images of two or more movement features on the screen of a device operated by the user. The user instructs the device as to whether two or more movement features are similar or not similar. Based on this instruction, the device determines whether the movement features are similar. In the latter case, the device determines whether two or more movement features contain clearly different movement features. If no obvious differences in features or quantities can be detected, the videos can be judged as similar. In this judgment, the AI can learn about the similarity through supervised learning. The device can utilize the functions of the AI or request the AI to judge the similarity and obtain the judgment result.
[0047] The device identifies the time between movement features. As a non-exhaustive example, in the case of a dance, if a dancer repeats the same dance choreography (similar movement features), the device identifies the difference (or distance) between the appearance (or application) times of each movement feature in the video. As a non-exhaustive example, assume that a video of a dance performance is three minutes long. If the dancer repeats a movement feature at 1 minute 10 seconds and 1 minute 15 seconds, the time between the movement features is five seconds, and the unit time is five seconds. The device recognizes the time between the movement features as a unit time. The device can transmit information about the unit time to another device or AI. In other embodiments, the device can use AI that has learned to recognize unit time through supervised learning, as described above, or request that AI to recognize unit time.
[0048] In at least one embodiment, a beat in music includes a unit of rhythm. As a non-exhaustive example, a beat includes a fixed time interval (a basic unit in music that repeats at regular intervals), a contrast of dynamics (strong and weak beats, which create rhythm), or the basis of tempo (the speed of the beat determines the tempo and the speed of a song). A beat may be a time signature. A time signature is the basis of the rhythm of an entire song and indicates the number of beats in a measure. For example, in a 4 / 4 time signature, there are four beats in a measure, each of which is the same length. In this disclosure, the term "beat" can be replaced with "meter" unless explicitly stated to the contrary. The device requests the AI to generate music with a unit of time being the beat (meter). As a non-exhaustive example, if a video of a dance performance is three minutes long and the unit of time is five seconds, the device requests the AI to generate music with five seconds being one beat. In other embodiments, the device may generate the music itself using the AI, as described above.
[0049] According to the above configuration, sound can be automatically created based on video. The created sound or music has beats based on the features of the video, so there is a high degree of match between the video and the generated sound. This reduces the workload of the sound creator, and has industrial applicability in that it can automatically generate highly usable sound that matches the video and the generated sound to a high degree.
[0050] When requesting an AI to generate music with a unit time of beats, a method is disclosed for requesting the AI to generate music of the same length as the video duration (including the time for which movement features should be identified. In this disclosure, unless explicitly stated to the contrary, "movement duration" can be understood interchangeably with "video length" or "video duration"). The device requests the AI to generate music of the same length as the video duration. As a non-exhaustive example, if a video of a dance performance is three minutes long, the device requests that it generate three minutes of music because the video duration is three minutes. In other embodiments, the device sets the time for which movement features should be identified as the video duration. In at least one embodiment, the time for which movement features should be identified refers to all or part of the time period within the video during which the device or AI is allowed to identify movement features. The time for which movement features should be identified can be determined by a user or automatically by the device. In the former case, the user can specify the time for which movement features should be identified, as long as it is within the video duration. As a non-exhaustive example, if a video of a dance performance is three minutes long, the time for identifying movement features can be determined to be between 0 minutes 0 seconds and 1 minute 30 seconds. The device identifies movement features only within the determined time using the method described above. As an example of the latter, the device can also determine a predetermined portion of the video time as the time for identifying movement features. Naturally, the device can also determine the entire video time as the time for identifying movement features. The device is requested to generate music of the same length as the time for identifying movement features.
[0051] A method is disclosed in which a device requests that music be generated so that the beat of the music is the same as the unit time in a video. The above term "becomes the same" can be replaced with "synchronize." In at least one embodiment, the device generates music with the same duration as the video, and requests that the music be generated so that the unit time in the video and the unit time of the music are the same. As a non-exhaustive example, assume that a video of a dance performance is three minutes long. If the dancer repeats a movement feature at 1 minute 12 seconds and 1 minute 17 seconds, the unit time is five seconds. The device further requests the AI to generate music with five seconds as one beat and to generate music that is three minutes long. Then, because the device requests that the music be generated so that the unit time in the video and the unit time of the music are the same, the generated sound generates beats that are the same (synchronized) as the beats at 1 minute 12 seconds and 1 minute 17 seconds in the video. In this case, if the generated music has beats every 5 seconds starting from 0 minutes 2 seconds, it will be identical (synchronized) with the beats at 1 minute 12 seconds and 1 minute 17 seconds in the video. In other words, the device recognizes the beats of the music that correspond to the unit time in the video as specific beats, and requests that music be generated with the overall beat determined based on the specific beats (or so as not to contradict the specific beats). Taking the previous example, the beats at 1 minute 12 seconds and 1 minute 17 seconds correspond to specific beats, and music with the overall beat determined based on the specific beats is generated.
[0052] According to the above configuration, it is possible to automatically create sound synchronized with the beat of a video. The created sound or music not only has a beat based on the features of the video, but is also synchronized with at least the beat of the video, so there is a high degree of match between the video and the generated sound. In this case, the generated sound can be used industrially as is by playing the video and the generated sound simultaneously. This reduces the workload of the sound creator and has industrial applicability in that it is possible to automatically generate highly usable sound that matches the video and the generated sound to a high degree.
[0053] A method is disclosed for requesting an AI to generate music with a unit time as a beat, requesting a relatively short unit time as a beat from two or more unit times resulting from similar movement features. In at least one embodiment, a device may recognize three or more similar movement features from a video. In this case, the device may recognize two or more unit times. The device recognizes the shortest unit time from the two or more unit times resulting from similar movement features as the unit time to be used as the beat. The device can transmit this unit time to another device or AI, or request the AI to recognize or identify this unit time to be used as the beat. As previously described, the AI learns to use the shortest unit time as the beat through supervised learning using a video with two or more unit times resulting from similar movement features. As a non-exhaustive example, consider a three-minute video of a dance performance. For example, the A-melody, B-melody, and chorus are commonly used in popular songs such as pop and rock, and each consists of a distinct melody. A music video may have a structure of an A-verse, a chorus, a B-verse, and another chorus, and the choreography of the chorus may be similar. Suppose a dancer repeats the same movement characteristics at 1 minute 12 seconds, 1 minute 17 seconds, 2 minutes 12 seconds, and 2 minutes 17 seconds. The device will recognize four unit times: 5 seconds, 55 seconds, 1 minute, and 1 minute 5 seconds. In this case, the device will recognize 5 seconds as the shortest unit time. The device will recognize this shortest unit time as a beat, or request the device or AI to recognize the shortest unit time as a beat.
[0054] With the above configuration, it is possible to automatically create sound synchronized with the beat of a video, even for complex video in which similar motion features are scattered throughout the video. The created sound or music not only has a beat based on the features of the video, but is also synchronized with at least the beat of the video, resulting in a high degree of match between the video and the generated sound. In this case, the generated sound can be used industrially as is by playing the video and the generated sound simultaneously. This has industrial applicability in that it reduces the workload of the sound creator and enables the automatic generation of highly usable sound that closely matches the video and the generated sound.
[0055] A method is disclosed in which, when requesting an AI to generate music using a unit time as a beat, the AI is requested to select a unit time with a relatively high occurrence frequency from two or more unit times resulting from similar motion features as the beat. In at least one embodiment, the device calculates the occurrence frequency of the recognized unit time. The occurrence frequency includes a numerical value indicating how many times a certain event or object appears within a specific period or range. For example, the occurrence frequency includes an event (e.g., a specific word, number, action, phenomenon), a range (e.g., a specific sentence, data set, time, space), and the frequency of a word in a sentence (e.g., how many times the word "cat" appears in a sentence). The device calculates the occurrence frequency of the unit time within the time range of the video or the time range in which motion features are to be identified. As a non-exhaustive example, suppose a video of a dance performance is three minutes long and the dancer's choreography is recognized as a motion feature. The device then recognizes that there are 50 five-second unit times, 10 ten-second unit times, 3 twenty-second unit times, and 1 one-minute unit time. In this case, the device recognizes the most frequently occurring 5 seconds as the unit of time to be used as a beat. The device then generates music using this unit of time as a beat, or requests the AI to do so.
[0056] As already explained, the above configuration allows for automatic creation of sounds synchronized with the beat of a video, even for complex videos in which similar motion features are scattered throughout the video. Furthermore, since the most frequently occurring unit time is the beat, the beat of the generated music is unlikely to mismatch the beat of the video. Therefore, the created sound or music not only has a beat based on the features of the video, but is also synchronized with at least the beat of the video, resulting in a high degree of match between the video and the generated sound. In this case, the generated sound can be used industrially as is by playing the video and the generated sound simultaneously. This has industrial applicability in that it reduces the workload of the sound creator and automatically generates highly usable sounds that closely match the video and the generated sound.
[0057] The present disclosure discloses a method for requesting an AI to generate music with a unit time as a beat, in which the AI generates music of the same duration as the video or the time for which action features are to be identified. When two or more unit times based on dissimilar action features are identified, the AI generates two or more pieces of music and requests that the pieces be connected. In at least one embodiment, the device recognizes unit times based on similar action features, as already described. In this case, the device recognizes a unit time based on action feature A and a unit time based on action feature B. These are two or more unit times based on dissimilar action features. When two or more unit times based on dissimilar action features are identified, the device generates two or more pieces of music. In other embodiments, the AI or another device is requested to generate two or more pieces of music in this manner. The device connects the generated music. Connecting includes connecting multiple melodies. The connection may be smooth and natural. As a non-exhaustive example, the device generates music A generated from a video (including a music video or movie) in a unit time based on action feature A and music B generated in a unit time based on action feature B. The device concatenates music A and music B. One concatenation method is to concatenate music data so that one music can be played as another music ends. Another method is to concatenate a portion of music data with a portion of other music data so that one music ends midway and another music starts. In at least one embodiment, when two or more pieces of music are generated, the second piece of music can be a modulated piece of music from one piece of music.
[0058] With the above configuration, even for complex videos in which two or more unit times are identified based on dissimilar motion features, it is possible to automatically create sound synchronized with the beat of the video. Furthermore, because the sound is based on two or more unit times, the possibility of it not matching the beat of the video is further reduced. Therefore, the created sound or music is largely synchronized with the beat of the video, so there is a high degree of match between the video and the generated sound. In this case, the generated sound can be used industrially as is by playing the video and the generated sound simultaneously. This has industrial applicability in that it reduces the workload of the sound creator and automatically generates highly usable sound that matches the video and the generated sound to a high degree.
[0059] The method also discloses a method for requesting that music be concatenated at a time point identified based on the frequency of occurrence of two or more action features. In at least one embodiment, the device generates music for a time period identical to the time of the video or the time period for which action features should be identified, and identifies a time point at which the unit time changes for two or more units of time. As a non-exhaustive example, the device calculates the frequency or probability of occurrence of each unit of time based on each action feature within a time period of the video. For example, the device identifies, within the time period of the video or the time period for which action features should be identified, a time period in which unit times based on action feature A occur relatively frequently and a time period in which unit times based on action feature B occur relatively frequently. The device generates music based on the unit time with the highest frequency of occurrence based on these frequencies of occurrence. If the unit time with the highest frequency of occurrence changes, the device generates music based on the new unit time. The device concatenates these two pieces of music. In another embodiment, the device generates music for a time period identical to the time period of the video (the time period for which action features should be identified), and generates two or more pieces of music when two or more dissimilar action features are identified. The device identifies a time point (in this paragraph, unless explicitly stated to the contrary, the terms "time point" and "time" can be replaced with "time period") identified based on the frequency of occurrence of two or more dissimilar action features as the change time. For example, within the time range of the video or the time for which action features are to be identified, the device identifies a time period in which unit times based on action feature A occur relatively frequently and a time period in which unit times based on action feature B occur relatively frequently. The device can identify a time point at which the most frequently occurring action feature changes, or a time point between or approximately midway between time periods in which different action features occur frequently, as the change time. As a non-exhaustive example, consider a three-minute video of a dance performance. If unit times based on action feature A occur relatively frequently in the time period from 0:00 to 1:30, and unit times based on action feature B occur relatively frequently in the time period from 1:30 to 3, the device identifies 1:30, which is approximately midway between the time periods, based on the frequency of occurrence of two or more action features. Once identified, the device requests that the music be linked.That is, the generated music is music A based on movement feature A from 0 minutes 0 seconds to 1 minute 30 seconds, and music B based on movement feature B from 1 minute 30 seconds to 3 minutes. These are connected at the 1 minute 30 second point.
[0060] As already explained, the above configuration makes it possible to automatically create sound synchronized with the beat of a video, even for complex video in which two or more unit times are identified by dissimilar motion features. Furthermore, because the sound is based on two or more unit times, the possibility of the sound not matching the beat of the video is further reduced. Therefore, the created sound or music is largely synchronized with the beat of the video, resulting in a higher degree of match between the video and the generated sound. In this case, the generated sound can be used industrially as is by playing the video and the generated sound simultaneously. This has industrial applicability in that it reduces the workload of the sound creator and automatically generates highly usable sound that matches the video and the generated sound to a high degree.
[0061] In at least one embodiment, the video includes a movie, dance, or choreography. An example includes a video published on YouTube (registered trademark). In this case, music can be automatically generated to match the published video. The video includes a live-streamed video. In this case, the device generates music according to any of the conditions or configurations described herein simultaneously with the live streaming. The device can simultaneously broadcast the generated music simultaneously with the live-streamed video. With the above configuration, music that matches the beat of the video can be provided simultaneously with the live streaming. Such a highly flexible generation method is not possible with a human composer, and therefore has superior advantages over conventional techniques and has industrial applicability. In at least one embodiment, the device can analyze environmental sounds and audience voices during the streaming in real time, recognize one or more of these as specific sounds, features, or characteristic quantities, and request AI to generate sounds and sound effects accordingly.
[0062] In at least one embodiment, the video includes video used in VR or AR. VR (Virtual Reality) is a technology that provides a completely immersive experience in a virtual world. By wearing a device such as VR goggles, a user can feel as if they are in another world. AR (Augmented Reality) is a technology that overlays digital information on the real world. Virtual objects can be displayed on real scenery and information can be added through a smartphone camera. The device can generate music for a video presented to a user through VR or AR in accordance with any of the conditions or configurations described in this disclosure. The device can output the generated music as audio at the same time as the video is presented to the user through VR or AR. Audio output includes an embodiment in which audio is output from a user terminal. The above configuration allows music to be provided that matches the beat of the VR or AR video. This has industrial applicability in that it reduces the workload of a sound creator and automatically generates highly usable sound that closely matches the video and the generated sound.
[0063] In at least one embodiment, the AI repeatedly learns to predict the next note through unsupervised learning. As already explained, the flow of a composition is hierarchically structured according to its structure. The elements that make up music are said to be at least scales, chords, and melodies. Once a scale is decided, chords are also decided based on the notes that make up the scale. There are theories about chord progressions, and many patterns that have been considered good progressions in songs to date have been accumulated. Chords greatly influence the mood of a song and are said to have a significant impact on people's emotions and perceptions at that moment. The AI repeatedly learns to predict the next note through unsupervised learning using music data. Furthermore, the AI performs supervised learning on the music data to determine what emotions the music evokes.
[0064] In at least one embodiment, the device or AI generates music based on music or related information that has already been uploaded to the internet or recorded in a database as completed, regardless of whether the author is registered. As a non-exhaustive example, when a unit time based on motion feature A is 5 seconds, the device identifies music with a beat length of 5 seconds that has already been uploaded to the internet as completed. If multiple pieces of music are identified, the device or AI identifies music that meets any of the conditions described in this disclosure. The device or AI then modifies the identified music to generate music.
[0065] In at least one embodiment, the device or AI identifies music based on music or related information already uploaded to the internet or stored in a database as completed works, regardless of whether the author is registered. In this embodiment, the device or AI is configured to identify existing music that meets certain criteria, rather than generate music. The device presents the identified music or information about the music to the user.
[0066] In at least one embodiment, the device or AI generates music using music or related information that has already been uploaded to the internet or stored in a database as completed work, regardless of whether the work has been registered by the author.
[0067] In at least one embodiment, the device adds data indicating that the AI-generated music was generated by AI to the data of the AI-generated music. When providing a video to a user with the generated music, if the device recognizes that the music contains such data, it notifies the user that the music was generated by AI, or displays such information on the user's device. In another embodiment, when the device generates music using music that has already been uploaded to the Internet or recorded in a database as a completed work, it adds information about the name of the music or the copyright holder to the generated music data. If the device recognizes that the music contains such information, it notifies the user of the name of the music, that the copyright holder created it, or that the AI generated it based on that music, or displays such information on the user's device. In at least one embodiment, the device generates data or prompts displaying copyright notices in a format compatible with the distribution platform. As an example, in the case of YouTube (registered trademark), the device generates data or prompts displaying copyright notices specified by the platform, or generates such data to accompany the generated video. These are sometimes referred to as copyright management tools. The above configuration has industrial applicability in that it can reduce the risk of copyright infringement in music data added or generated by AI.
[0068] In at least one embodiment, a user can specify to the device which sounds to rely on when generating music. In one example, the device accepts the sound data to be relied upon. The device contains two or more user-specified sound sources, allowing the user to select any of the sound sources. "Containing two or more sound sources" refers not only to embodiments in which the device itself stores these sound sources, but also to embodiments in which the device obtains these sound sources from an external storage medium. For example, these sound sources are pre-composed music. The user selects one or more sound data to rely upon, and the device generates music based on the selected music. As a non-exhaustive example, if a video has a unit time of five seconds and the selected sound data is pre-composed music with a beat of seven seconds, the device adjusts the playback speed of the pre-composed music to change it to music with a beat of five seconds. In this example, the changed music corresponds to the "generated music" in the above embodiment. According to the above configuration, pre-composed music that synchronizes with the beat of the video can be automatically identified. Furthermore, since the identified pre-composed music is based on a unit time, it matches the beat of the video. In this case, the presented sound can be used industrially as is by playing the video and the generated sound simultaneously. This has industrial applicability in that it reduces the workload of the sound creator, reduces the workload of searching for music with a matching beat, and automatically prepares highly usable sound that closely matches the video and the generated sound. In these embodiments, the device can generate data or prompts for license indications for the used pre-made sound within the generated music data or store them accompanying the data. An advantage of such an embodiment is at least that it can provide a system that smoothly manages copyright management and licensing.
[0069] In at least one embodiment, the device analyzes motion features (e.g., dance moves) in the video in more detail and generates different musical elements (e.g., beats and melodies) for each motion feature (e.g., dance steps). These musical elements may be adapted or combined according to any of the methods described herein. These musical elements may be added to the generated music as long as they do not contradict the beat of the generated music.
[0070] In at least one embodiment, the device recognizes a predetermined motion feature of an object in a video as a specific motion feature and generates a specific sound at the time the specific motion feature occurs. The device can generate the specific sound in addition to the generated music. In another embodiment, the device can generate the specific sound independently of the generated music. The user registers information about the specific motion feature or specific sound with the device. As a non-exhaustive example, the user registers a person's surprised expression or a predetermined reaction as the specific motion feature, the surprised expression, or a prompt as information about the specific sound. The specific sound includes pre-made music or sound effects specified by the user. The device can store two or more of such pre-made music or sound effects internally or receive them from an external storage device. The device can present two or more of such available specific sounds to the user. The user can specify the specific sound to be used from among them. If two or more specific sounds are specified, the device will use the specific sounds alternately, randomly, or in accordance with an embodiment of the present disclosure as the specific sound to be generated. If the video is a two-hour movie, and a character in the movie shows signs of surprise at the 1 hour, 20 minutes, and 43 second mark, the device generates a specific sound representing surprise at that time. According to the above configuration, a specific sound based on a specific action feature can be automatically generated. Furthermore, the generated specific sound is based on the time at which the specific action feature occurs, and therefore coincides with the timing of the specific action feature in the video. In this case, the presented sound can be used industrially as is by simultaneously playing the video and the generated sound. This reduces the workload of the specific sound creator and automatically prepares highly usable sounds that closely match the video and the generated sound, making it industrially applicable. As already explained, this configuration can be combined with or replaced with one or more aspects of other embodiments disclosed herein, unless inconsistent. That is, it is also possible to analyze environmental sounds and audience voices during streaming in real time and generate background music and sound effects (specific sounds) accordingly. In another embodiment, the present invention can be implemented as an application that automatically adds a surprise sound effect to a surprised action.
[0071] In at least one embodiment, the moving image includes a movie. According to this embodiment, optimal music or specific sounds can be generated separately for each sequence in the movie.
[0072] An embodiment different from the above will be disclosed.
[0073] A method is disclosed in which a device identifies features of objects in a video and requests an AI to generate music based on the features. This method includes identifying motion features that are characteristic of a movement and requesting an AI to generate music based on the motion features. According to at least one embodiment, the device or AI represents a video or image as a collection of visual patches, which are small data units similar to text tokens in an LLM. The device combines a video recognition AI with an LLM to understand each scene in the video. Specifically, the video recognition AI individually recognizes various objects and environments that make up a scene, such as people, cars, buildings, animals, trees, and other natural objects, as well as weather and changes therein. The AI learns emotions corresponding to the visual patches using the learning method described above (including supervised learning). Furthermore, the AI may pre-find the LLM using sample videos in the target domain. For example, the AI learns that a particular visual patch corresponds to a mood, atmosphere, or emotion (collectively referred to as a "concept" in this disclosure). The device requests the AI to generate music based on the motion features. As an example, if the video is a three-minute music video and the dancer's facial expression is smiling, the device identifies motion characteristics such as cheerfulness and joy, and requests the AI to generate music based on those characteristics. In another embodiment, the device also requests the AI to identify motion characteristics that characterize the movements of objects in the video. As another example, if the video is a three-minute music video and ruins are depicted behind the dancer as a dramatic effect, the device identifies auxiliary characteristics such as ruins and darkness as auxiliary characteristics, and requests the AI to generate music based on those auxiliary characteristics. In another example, the video is a movie, and the video's scenes and emotional flow can be analyzed to generate a corresponding musical composition (intro, climax, ending). Furthermore, the music corresponding to a concept includes an embodiment corresponding to a musical genre (classical, EDM, hip hop, etc.). According to this embodiment, music that matches the style or concept of the video can be automatically generated.This has industrial applicability in that it reduces the workload of sound creators and automatically prepares highly usable sounds that closely match the concept of the video and the generated sound.
[0074] In at least one embodiment, the device can also use user-specified features as action features or auxiliary features. The user inputs text into the device the concept of the music to be generated. These texts may be prompts. The device requests the AI to generate music based on the concept input by the user. Furthermore, the device may request the AI to generate music of the same duration as the video or the time for which the action feature is to be identified. In one example, the user can select a mood or atmosphere and generate music that matches the selected mood. This embodiment allows the user to specify their preferences, making it possible to automatically generate music that matches the user's preferences. This has industrial applicability in that it reduces the workload of the sound creator and automatically prepares highly usable sounds that closely match the concept of the video and the generated sound.
[0075] In at least one embodiment, the device analyzes lyrics as features specified by the user. The device or AI recognizes concepts contained in the lyrics as features. For example, the device or AI extracts concepts such as love, heartbreak, and sadness based on the lyrics. The device or AI generates sound based on the extracted concepts. When two or more concepts are recognized based on the lyrics, the AI is requested to base the sound on the feature with the highest frequency of occurrence, as described below. This method allows music to be generated based on the main concept of the lyrics, even for complex lyrics. This has industrial applicability in that it reduces the workload of the sound creator and automatically prepares highly usable sound that closely matches the concept of the video and the generated sound.
[0076] In at least one embodiment, the device requests an AI to convert objects in a video into text. As already described, the AI represents videos and images as a collection of visual patches, which are small data units similar to text tokens in an LLM. The device combines a video recognition AI with an LLM to understand each scene in the video. Specifically, the video recognition AI individually recognizes various objects and environments that make up a scene, such as people, cars, buildings, animals, trees, and other natural objects, as well as weather and changes therein. The AI learns characters corresponding to the visual patches using the learning method described above (including supervised learning). This method allows the AI to describe a given video in text (convert it into text). The AI generates music based on the converted text. Specifically, the device identifies features based on the converted text and requests the AI to generate music based on the features. The method of generating music based on features is disclosed elsewhere in this disclosure.
[0077] A method is disclosed in which, when requesting an AI to generate music based on features, if two or more features are present, the AI is requested to base the music on at least the feature that is recognized for a relatively long time. In at least one embodiment, the device may recognize two or more features. In this case, the AI is requested to base the music on the feature that is recognized for the longest time in the video. The device identifies the time at which each feature is recognized. As a non-exhaustive example, if a video is a three-minute movie, of which two minutes and 40 seconds depict a heartbreak, and 20 seconds depict a happy and uplifting romantic relationship as a flashback, the characteristic of "heartbreak" is 2 minutes and 40 seconds, and the characteristic of "fun" is 20 seconds. In other words, the device requests the AI to create music based on the characteristic of "heartbreak," which is the longest feature. Note that this method is not intended to be limited to one feature; it is also possible to generate music based on two or more features. With the above configuration, music can be generated based on a main concept even for a video with multiple concepts. Therefore, even if a video contains contradictory concepts, the music is based on the main concept, so there is a high possibility that the concept of the video will match. This has industrial applicability in that it reduces the workload of sound creators and automatically prepares highly usable sounds that closely match the concept of the video and the generated sound.
[0078] This disclosure discloses a method for requesting an AI to generate music based on features. When two or more features are present, the method requests the AI to generate music based on at least the feature with a relatively high frequency of recognizable features. As described above, in at least one embodiment, the device may recognize two or more features. In this case, the AI is requested to generate music based on the feature with the highest frequency of recognizable features in the video. The device identifies the number of times each feature is present. As a non-exhaustive example, if a video is a three-minute movie in which the actors laugh 15 times, cry once, and are angry once, the most frequently occurring feature is "laughter." In other words, the device requests the AI to create music based on the feature "laughter," which is the longest feature. Note that this method is not limited to a single feature; it is also possible to generate music based on two or more features. With the above configuration, music can be generated based on a main concept even for a video with multiple concepts. Therefore, even if a video contains contradictory concepts, the music is based on the main concept, so there is a high possibility that the concept of the video will match. This has industrial applicability in that it reduces the workload of sound creators and automatically prepares highly usable sounds that closely match the concept of the video and the generated sound.
[0079] This disclosure discloses a method for requesting AI to generate music based on features, generating music of the same duration as the time of the video or the time at which action features should be identified. When two or more dissimilar features are identified, the method requests the AI to generate two or more pieces of music and connect the pieces of music. As previously described, the device or AI can generate two or more pieces of music when two or more features are identified. Furthermore, as previously described, the device or AI can connect two or more pieces of music. With the above configuration, even for complex videos with two or more dissimilar features, it is possible to automatically create sound that is similar to the concept of the video. Furthermore, because the sound is based on two or more concepts, the possibility of it not matching the concept of the video is further reduced. Therefore, the created sound or music is largely synchronized with the concept of the video, resulting in a high degree of match between the video and the generated sound. In this case, the generated sound can be used industrially as is by simultaneously playing the video and the generated sound. This has industrial applicability in that it reduces the workload of sound creators and automatically generates highly usable sound that closely matches the concept of the video and the generated sound.
[0080] The method further discloses a method for requesting that the music be concatenated at a time point identified based on the frequency of appearance of two or more features. As described above, the device can calculate the frequency or probability of appearance of each feature within a time period of a video. For example, within a range of the video time or the time period for which action features are to be identified, the device identifies a time period in which unit times based on feature A occur relatively frequently and a time period in which unit times based on feature B occur relatively frequently. The device generates music based on the unit time with the highest frequency of appearance based on these frequencies of appearance. If the unit time with the highest frequency of appearance changes, the device generates music based on the new unit time. The device concatenates these two pieces of music. In another embodiment, the device generates music for the same time as the video time (the time for which action features are to be identified), and generates two or more pieces of music when two or more dissimilar action features are identified. The device identifies the time point identified based on the frequency of appearance of two or more dissimilar action features as the change time. For example, within a time range for which a video or action feature is to be identified, the device may identify a time period in which unit times based on feature A occur relatively frequently and a time period in which unit times based on feature B occur relatively frequently. The device may identify a time when the most frequently occurring action feature changes, or a time between or approximately midway between time periods with high occurrence frequencies based on different action features, as the change time. As a non-exhaustive example, suppose a movie video is 10 minutes long. If the time period from 0:00 to 1:30 shows that unit times based on feature A (e.g., the setting is in ruins) occur relatively frequently, and the time period from 1:30 to 3:00 shows that the time period with unit times based on feature B (e.g., the setting is in an ornate building (such as a castle)) occurs relatively frequently, the device may identify 1:30, which is approximately midway between the two time periods, based on the frequency of occurrence of two or more action features. At the identified time, the device requests that the music be combined. That is, the generated music is music A based on feature A from 0:00 to 1:30, and music B based on feature B from 1:30 to 3:00. These are connected at the 1:30 point.
[0081] As already explained, the above configuration allows for the automatic creation of sound similar to the concept of a complex video, even for videos in which two or more unit times are identified based on dissimilar features. Furthermore, because the sound is based on two or more unit times, the possibility of the sound not matching the concept of the video is further reduced. Therefore, the created sound or music is largely synchronized with the concept of the video, resulting in a high degree of match between the video and the generated sound. In this case, the generated sound can be used industrially as is by simultaneously playing the video and the generated sound. This has industrial applicability in that it reduces the workload of sound creators and automatically generates highly usable sound that closely matches the concept of the video and the generated sound.
[0082] In at least one embodiment, when generating music based on lyrics or characters, the device recognizes characters or concepts in a sentence that indicate rhythm and requests AI to generate music with a rhythm based on those characters or concepts. The above configuration allows for the generation of music with a specified rhythm, further reducing the likelihood that the generated music will not match the concept of the video. Therefore, the generated sound can be used industrially as is by simultaneously playing the video and the generated sound. This has industrial applicability in that it reduces the workload of the sound creator and automatically generates highly usable sound that closely matches the concept of the video and the generated sound.
[0083] This disclosure provides an embodiment different from the above. The device requests an AI to create lyrics based on music data. The device or AI breaks down the music data into units such as scales, chords, or melodies. The AI uses supervised learning to learn the mood, atmosphere, or emotion represented by each unit. Furthermore, the AI learns and fine-tunes using combinations of previously published musical scales, chords, or melodies and lyrics. This allows the AI to predict what characters and lyrics will appear when given a scale, chord, or melody. It can also predict what words will appear next given the characters or lyrics. In this way, the AI can create lyrics based on music data. The device requests the AI to create lyrics based on music data. The above configuration allows lyrics to be generated based on traditional musical concepts, reducing the likelihood that the generated lyrics will be inconsistent with the musical concept. Furthermore, because AI uses natural language processing and learning, the generated lyrics are unlikely to contain grammatical errors. Therefore, the generated lyrics, along with the music, can be used industrially as is. This reduces the workload of the lyricist and allows for the automatic generation of highly usable lyrics that closely match the concept of the music and the generated lyrics, making it industrially applicable. In at least one embodiment, the device further requests the AI to generate a singing voice based on the generated lyrics. In at least one embodiment, the device requests the AI to generate music, requests the AI to write lyrics based on the generated music, requests the AI to generate a singing voice based on the generated lyrics, and requests the AI to match (integrate or combine) the generated singing voice with music data (or the device itself matches the data). Examples of AI applications for generating singing voices include Vocaloid (registered trademark). This method has industrial applicability in that it can automatically generate a singing voice that matches the concept of a video.
[0084] The following will disclose an outline of the above-described embodiment.
[0085] 1. A method of producing sound, comprising: The device, Identifying motion features that characterize the behavior of objects in a video; The AI is required to generate music with the time between similar motion features as a unit of time, with the unit time being a beat. method.
[0086] When requesting AI to generate music with a unit time being a beat, Generate music of the same duration as the time of the video or the time at which the action feature is to be identified; The music must be generated so that the unit time in the video and the beat of the music are the same. method.
[0087] When requesting AI to generate music with a unit time being a beat, Among two or more units of time resulting from similar movement characteristics, a relatively short unit of time is required to be defined as a beat. method.
[0088] When requesting AI to generate music with a unit time being a beat, A request is made to select a unit time having a relatively high occurrence frequency from two or more unit times resulting from similar motion characteristics as a beat. method.
[0089] When requesting AI to generate music with a unit time being a beat, Generate music of the same duration as the time of the video or the time at which the action feature is to be identified; When two or more time units with dissimilar motion features are identified, two or more pieces of music are generated; Request that the music be linked; method.
[0090] In the method, requesting that the music be linked when identified based on the frequency of occurrence of two or more motion characteristics; method.
[0091] 1. A method of producing sound, comprising: The device, Identify the features (feature values) of objects in the video, Ask the AI to generate music based on features, method.
[0092] When asking an AI to generate music based on features, When there are two or more characteristics, the AI must at least base its selection on the characteristic that has been recognized for a relatively long time. method
[0093] When asking an AI to generate music based on features, When there are two or more features, the AI must at least base its selection on the feature with the highest relative frequency of occurrence. method
[0094] When asking an AI to generate music based on features, Generate music of the same duration as the time or feature of the video to be identified; When two or more dissimilar features are identified, two or more pieces of music are generated; Request that the music be linked; method.
[0095] In the method, requesting that the music be linked once identified based on the frequency of occurrence of two or more characteristics; method.
[0096] In at least one embodiment, the phrase "requesting the AI to determine a relatively short unit time from among two or more unit times resulting from similar movement features as a beat" can be replaced with "requesting the AI to determine the shortest unit time from among two or more unit times resulting from similar movement features as a beat." The phrase "requesting the AI to determine a unit time with a relatively high occurrence frequency of a specified unit time from among two or more unit times resulting from similar movement features as a beat" can be replaced with "requesting the AI to determine a unit time with a highest occurrence frequency of a specified unit time from among two or more unit times resulting from similar movement features as a beat." The phrase "when two or more features are present, requesting the AI to base the AI on at least a feature with a relatively long duration in which the feature is recognized" can be replaced with "when two or more features are present, requesting the AI to base the AI on at least a feature with the longest duration in which the feature is recognized." The phrase "when two or more features are present, requesting the AI to base the AI on at least a feature with a relatively high occurrence frequency in which the feature is recognized" can be replaced with "when two or more features are present, requesting the AI to base the AI on at least a feature with a highest occurrence frequency in which the feature is recognized."
[0097] 1. A method of producing sound, comprising: The device, Generate music with the same duration as the video or a specified duration. Asking AI to generate music based on the characteristics of lyrics corresponding to the video, method.
[0098] In the above configuration, the specified time includes a time specified by the user. The device requests the AI to generate music within the time range specified by the user. The method for generating music based on the characteristics of lyrics and its advantages are as described above.
[0099] In at least one embodiment, the user can make detailed musical adjustments (tempo, key, effects) to the AI-generated music or vocals through the device. The device changes the tempo, key, effects, etc. of the generated music or vocals in response to the user's adjustments. This adjustment can be made even if the video is a real-time video or a streaming video.
[0100] In at least one embodiment, two or more users can collaborate through the device to input rhythms and concepts to be generated and fine-tune musical adjustments (tempo, key, effects) for AI-generated music or vocals. The device accepts information input from two or more users. Two or more users can collaborate on editing a single piece of data through the device. This editing can be done even if the video is real-time video or streaming video. As a non-exhaustive example, this editing function is referred to as "collaboration mode" or "real-time mode."
[0101] In at least one embodiment, the device requests the AI to create lyrics in one or more foreign languages based on lyrics or a concept. The device requests the AI to create music based on the input foreign language lyrics or concept. The AI fine-tunes the concept and music based on the foreign language. As a non-exhaustive example, the AI fine-tunes ethnic music concepts, such as the koto for the Japanese concept and the bagpipes for the European or Irish concepts. The device also requests the AI to generate music for the foreign concept when lyrics in a foreign language are input.
[0102] In at least one embodiment, the device requests the AI to generate music or video in a format compatible with major distribution platforms (non-inclusively, YouTube, TikTok, Instagram, etc., all trademarks). In another embodiment, the device converts the AI-generated music or video into a format compatible with major distribution platforms. An advantage of such an embodiment is at least that it creates platform-compatible data, improving user convenience.
[0103] The invention according to the present disclosure may have at least one of the above-described effects.
Claims
1. 1. A method of producing sound, comprising: The device, Automatically identify objects or environmental features in the video, When there are two or more characteristics, music is generated based on the characteristic that is recognized for a relatively long time, corresponding to the mood, atmosphere or emotion of said characteristic. method.
2. 1. A method of producing sound, comprising: The device, Automatically identify objects or environmental features in the video, If there are two or more features, generate music corresponding to the mood, atmosphere or emotion of the feature based on the feature with a relatively high frequency of occurrence in which the feature is recognized. method.
3. 1. A method of producing sound, comprising: The device, Automatically identify objects or environmental features in the video, generating music corresponding to the mood, atmosphere or emotion of said characteristic; generating two or more pieces of music when two or more dissimilar features are identified; concatenating the music once identified based on the frequency of occurrence of two or more features; method.
4. A system for creating sound, comprising: The device, Automatically identify objects or environmental features in the video, When there are two or more characteristics, music is generated based on the characteristic that is recognized for a relatively long time, corresponding to the mood, atmosphere or emotion of said characteristic. system.
5. A system for creating sound, comprising: The device, Automatically identify objects or environmental features in the video, If there are two or more features, generate music corresponding to the mood, atmosphere or emotion of the feature based on the feature with a relatively high frequency of occurrence in which the feature is recognized. system.
6. A system for creating sound, comprising: The device, Automatically identify objects or environmental features in the video, generating music corresponding to the mood, atmosphere or emotion of said characteristic; generating two or more pieces of music when two or more dissimilar features are identified; concatenating the music once identified based on the frequency of occurrence of two or more features; system.
7. A computer-readable recording medium in an apparatus having a processor, having recorded thereon a program that causes the processor to execute a method according to any one of claims 1 to 3.
8. A program comprising instructions for causing a processor to execute a method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Automatic musical composition device, automatic composition method, automatic composition program and memory medium
JP2002287746A
Music generation system
JP2006154777A
Automated Music Composition and Generation Machines, Systems and Processes Employing Language and / or Graphical Icon-Based Music Experience Descriptors
JP2018537727A
Music content generation
JP2023513586A
System for providing generative artificial intelligence based music making service using photo
KR102703767B1