Method and apparatus for generating sound
By identifying operation characteristics in moving images and using AI to generate music with a beat based on these characteristics, the method addresses the lack of synchronization in existing sound creation techniques, resulting in music that closely matches the visual content and has industrial applicability.
Patent Information
- Application Number
- JP2025016751
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-02-04
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2044-11-26
AI Technical Summary
Existing methods lack a systematic approach to create sound based on moving images, limiting the ability to generate music that synchronizes with the visual content.
A method where a device identifies operation characteristics from a moving image and requests an AI to generate music with a unit time as a beat, using the time between similar operation characteristics as the unit time.
This configuration enables the generation of music that is synchronized with the moving image, achieving a high degree of coincidence between the visual content and the generated sound, which is industrially applicable.
Abstract
Description
Technical Field
[0001] The present invention relates to a method for creating sound.
Background Art
[0002] The statements in this section only provide background information related to the present disclosure and do not necessarily constitute prior art.
[0003] Patent Document 1 discloses a music generation system.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, the inventor recognized that at least in the above-described embodiment, there is a disadvantage that there is no method for creating sound based on a moving image.
Means for Solving the Problems
[0006] At least one aspect of the present disclosure is a method for creating sound, wherein a device identifies operation characteristics that are characteristics of the operation of an object in a moving image, and requests an AI to generate music with a unit time being a beat, where the time between similar operation characteristics is used as the unit time. Method is provided.
Effects of the Invention
[0007] In this configuration, there is at least the usefulness that industrially applicable music can be generated based on a moving image.
[0008] These and other aspects, features, and advantages of the present disclosure will become apparent from the following detailed written description of the preferred embodiments and aspects taken in conjunction with the following drawings, but modifications and variations thereof may be practiced without departing from the spirit and scope of the novel concepts of the present disclosure. Aspects in one embodiment of the present disclosure can be combined with, or replaced by, one or more of the aspects in another embodiment of the present disclosure, as long as there is no contradiction.
Best Mode for Carrying Out the Invention
[0009] In the following disclosure, numerous different embodiments and examples are provided for implementing different features of the presented subject matter. To simplify the present disclosure, specific examples of components and arrangements are disclosed below. Of course, these are merely examples and are not intended to be limiting. For example, a structure in which a first feature is covered by, or in contact with, a second feature subsequently disclosed may include embodiments in which the first feature and the second feature are formed so as to be in direct contact, as well as embodiments in which additional features are formed between the first feature and the second feature so that the first feature and the second feature are not in direct contact. Further, in the present disclosure, reference numerals and / or letters may be repeated in various examples. Such repetition is for the sake of brevity and clarity and does not in itself require that there be a relationship between the various embodiments and / or the configurations being described. Further, when a first element is described as being "connected" or "coupled" to a second element, such description includes embodiments in which the first element and the second element are directly connected or coupled to each other, as well as embodiments in which the first element and the second element are indirectly connected or coupled to each other with one or more other elements intervening therebetween.
[0010] As used herein, the recitation "at least one of" includes all variations exemplified. For example, the recitation "comprises at least one of A, B, or C" is synonymous with "consisting of A, B, C and combinations thereof", and encompasses all possible variations of A, B, C, A + B, A + C, B + C, and A + B + C.
[0011] In this disclosure, the disclosure of using a machine, an electronic operator, or a computer can include embodiments of a method, a recording medium, an apparatus, or a program. The description "A is B" used herein can be replaced with "A includes B" as long as there is no contradiction or unless otherwise stated in this specification.
[0012] The terms in this disclosure, including those recited in the claims, can be interpreted in consideration of the descriptions and drawings described in the specification, and further, as long as there is no contradiction with the suggestions in this disclosure, based on matters that one or more members of the public have so named, indicated, understood, or practiced, or that are possible in the past, present, or future. Regarding the operation method used in at least one or more embodiments, the following embodiments can be adopted. The description of JP6456303, which well explains at least one or more embodiments, is cited for explanation (hereinafter, citation begins).
[0013] As used herein, the term "computer", as known in the art, generally includes a processor, a memory such as a hard drive, disk drive or flash drive or memory stick, or other non-transitory computer-readable medium or non-transitory storage device, at least one information storage / search device, such as a keyboard, mouse, point and touch device, touch screen, or microphone, at least one input device, and a display structure such as a well-known computer screen. Additionally, a computer may include one or more network connections, such as a wired or wireless connection. As known in the art, such a computer or computer system may include more or less of the items listed above, and is not limited to, for example, tablet computers or smart devices, but includes other electronic media and electronic devices.
[0014] As used herein, the term "cloud" or "cloud computing" refers to a centralized and virtualized computing facility where all computing resources are shared. For application systems and subsystems, since they are all within the "cloud", it is no longer possible to refer to a specific machine.
[0015] As used herein, the term "Distributed Internet Service System" refers to a distributed Internet service platform that transforms Internet applications for execution in various computing environments. The DIS system distributes Internet applications, including content, data, and logic, via a Component Distribution Server / Asset Distribution Server, to any number and any type of device, to whatever extent appropriate, and along the network. Through DIS, Internet applications can be hosted and centrally managed as services based on each user's needs, locally cached and executed at the user's device or nearby location while maintaining their integrity. Web-enabled computing devices can be upgraded with DIS software to become DIS-compliant for enjoying and executing distributed Internet services. The distributed Internet service system is fully described in any one of the patent families of U.S. Patent Nos. 7,136,857, 7,150,015, 7,181,731, 7,209,921, 7,430,610, 7,685,183, 7,685,577, 7,752,214, 8,326,883, 8,386,525, 8,443,035, 8,458,142, 8,458,222, 8,473,468, 8,527,545, and 8,650,226, and U.S. Patent Publications 2012 / 0005205 and 2013 / 0091252, all of which are jointly owned by OPI40, Holdings, Inc. and are hereby incorporated by reference. (End of citation)
[0016] Regarding the operation method used in at least one or more embodiments, the following embodiments can be taken for the conventional Internet method that does not use a distributed Internet. The description of JP7113047, which well explains at least one or more embodiments, will be cited for explanation (hereinafter, citation starts).
[0017] Embodiments including the matters specifically disclosed in this specification can provide an automatic response system realized in a form that actually converses with humans based on artificial intelligence, thereby realizing a more natural conversation with the user while quickly and conveniently processing inquiries, reservations, delivery orders, etc.
[0018] The plurality of electronic devices 110, 120, 130, 140 may be fixed terminals or mobile terminals realized by a computer system. Examples of the plurality of electronic devices 110, 120, 130, 140 include AI speakers, smartphones, mobile phones, navigation devices, PCs (personal computers), notebook PCs, digital broadcast terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablets, game consoles, wearable devices, IoT (internet of things) devices, VR (virtual reality) devices, AR (augmented reality) devices, etc. As an example, in FIG. 1, an AI speaker is shown as the electronic device 110. However, in the embodiments of the present invention, the electronic device 110 may mean one of various physical computer systems that can communicate with other electronic devices 120, 130, 140 and / or servers 150, 160 via the network 170 using substantially wireless or wired communication methods.
[0019] The communication method is not limited, and it may include not only communication methods using communication networks that the network 170 can include (for example, mobile communication networks, wired Internet, wireless Internet, broadcast networks, satellite networks, etc.), but also short-range wireless communication between devices. For example, the network 170 may include any one or more of networks such as PAN (personal area network), LAN (local area network), CAN (campus area network), MAN (metropolitan area network), WAN (wide area network), BBN (broadband network), and the Internet. Further, the network 170 may include any one or more of network topologies including bus network, star network, ring network, mesh network, star-bus network, tree or hierarchical network, etc., but is not limited thereto.
[0020] Servers 150 and 160 may each be implemented by one or more computer devices that communicate with a plurality of electronic devices 110, 120, 130, 140 via a network 170 to provide instructions, code, files, content, services, etc. For example, server 150 may be a system that provides a first service to a plurality of electronic devices 110, 120, 130, 140 connected via network 170, and server 160 may also be a system that provides a second service to a plurality of electronic devices 110, 120, 130, 140 connected via network 170. As a more specific example, server 150 may provide, as the first service, a service (such as an automatic response service, for example) targeted by a corresponding application to a plurality of electronic devices 110, 120, 130, 140 through an application that is a computer program installed and executed on the plurality of electronic devices 110, 120, 130, 140. As another example, server 160 may provide, as the second service, a service that distributes files for installation and execution of the above-described application to a plurality of electronic devices 110, 120, 130, 140.
[0021] FIG. 2 is a block diagram for explaining the internal configurations of an electronic device and a server in an embodiment of the present invention. In FIG. 2, the internal configuration of electronic device 110 and the internal configuration of server 150 are described as examples for an electronic device. Also, the other electronic devices 120, 130, 140 and server 160 may have the same or similar internal configurations as the above-described electronic device 110 or server 150.
[0022] The electronic device 110 and the server 150 may include memories 211 and 221, processors 212 and 222, communication modules 213 and 223, and input / output interfaces 214 and 224. The memories 211 and 221 may be non-transitory computer-readable recording media, and may include non-transitory mass storage devices such as RAM (random access memory), ROM (read only memory), disk drives, SSDs (solid state drives), flash memories, and the like. Here, non-transitory mass storage devices such as ROM, SSD, flash memory, and disk drives may be included in the electronic device 110 or the server 150 as separate non-transitory recording devices distinct from the memories 211 and 221. Also, the memories 211 and 221 may record an operating system and at least one program code (for example, code for a browser installed and executed in the electronic device 110, an application installed in the electronic device 110 for providing a specific service, etc.). Such software components may be loaded from a computer-readable recording medium different from the memories 211 and 221. Such another computer-readable recording medium may include computer-readable recording media such as floppy (registered trademark) drives, disks, tapes, DVD / CD-ROM drives, memory cards, and the like. In other embodiments, the software components may be loaded into the memories 211 and 221 through the communication modules 213 and 223 which are not computer-readable recording media. For example, at least one program may be loaded into the memories 211 and 221 based on a computer program (for example, the above-described application) installed by a file distributed by a file distribution system (for example, the above-described server 160) that distributes developer or application installation files via the network 170.
[0023] Processors 212 and 222 may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to processors 212 and 222 by memory 211, 221 or communication modules 213, 223. For example, processors 212 and 222 may be configured to execute instructions received according to program code recorded in a recording device such as memory 211, 221.
[0024] Communication modules 213 and 223 may provide functions for the electronic device 110 and the server 150 to communicate with each other via the network 170, or may provide functions for the electronic device 110 and / or the server 150 to communicate with other electronic devices (for example, the electronic device 120) or other servers (for example, the server 160). As an example, a request generated by the processor 212 of the electronic device 110 according to program code recorded in a recording device such as the memory 211 may be transmitted to the server 150 via the network 170 under the control of the communication module 213. Conversely, control signals, instructions, contents, files, etc. provided under the control of the processor 222 of the server 150 may be received by the electronic device 110 through the communication module 223 of the server 150 and the communication module 213 of the electronic device 110 via the network 170. For example, control signals, instructions, contents, files, etc. received through the communication module 213 may be transmitted to the processor 212 and the memory 211, and the contents and files may be recorded in a recording medium (the non-transitory recording device described above) that the electronic device 110 may further include.
[0025] The input / output interface 214 may be means for interfacing with an input / output device 215. For example, the input device may include devices such as a keyboard, a mouse, a microphone, a camera, etc., and the output device may include devices such as a display, a speaker, a tactile feedback device, etc. As another example, the input / output interface 214 may be means for interfacing with a device in which functions for input and output are integrated into one, such as a touch screen. The input / output device 215 may be composed of the electronic device 110 and one device. Also, the input / output interface 224 of the server 150 may be means for interfacing with a device (not shown) for input or output that can be connected to or included in the server 150. As a more specific example, when the processor 212 of the electronic device 110 processes the instructions of a computer program loaded in the memory 211, a service screen or content configured using data provided by the server 150 or the electronic device 120 may be displayed on the display through the input / output interface 214.
[0026] Also, in other embodiments, the electronic device 110 and the server 150 may include more components than the components shown in FIG. 2. However, it is not necessary to clearly show most of the conventional components in the figure. For example, the electronic device 110 may be implemented to include at least a part of the input / output device 215 described above, or may further include other components such as a transceiver, a camera, various sensors, a database, etc. As a more specific example, when the electronic device 110 is an AI speaker, various components such as various sensors, a camera module, physical buttons, buttons using a touch panel, an input / output port, a vibrator for vibration, etc., generally included in an AI speaker may be implemented to be further included in the electronic device 110. (End of citation)
[0027] A machine is disclosed. According to at least one embodiment, the user terminal consists of a control unit, a RAM, a storage unit, a graphics processing unit, a communication interface, and an interface unit, which are respectively connected by an internal bus.
[0028] According to at least one embodiment, the control unit is composed of a CPU and a ROM. The control unit executes the program stored in the storage unit to control the user terminal. The RAM is the work area of the control unit. The storage unit is a storage area for storing programs and data. The control unit reads the program and data from the RAM for processing. The control unit processes the program and data loaded in the RAM and outputs a drawing command to the graphics processing unit.
[0029] According to at least one embodiment, the graphics processing unit is connected to the display unit. The display unit has a display screen. When the control unit outputs a drawing command to the graphics processing unit, the graphics processing unit outputs a video signal for displaying an image on the display screen. Here, the display unit may be a touch panel equipped with a touch sensor. The touch panel of this display unit functions as an input unit.
[0030] According to at least one embodiment, the communication interface can be connected to a communication network wirelessly or by wire, and it is possible to transmit and receive data with a server device via the communication network. The data received via the communication interface is loaded into the RAM and processed by the control unit. An external memory (e.g., SD card, etc.) is connected to the interface unit.
[0031] According to at least one embodiment, the user terminal is not particularly limited as long as it is a computer device having a display screen and an input unit. Examples of the user terminal include a conventional mobile phone, a tablet terminal, a smartphone, a desktop or laptop personal computer, etc. It may also be composed of a VR goggle, that is, a screen (or two display panels, one for each eye) attached to a frame (or headset) fixed or attached to the head with a strap. The user terminal has an audio output unit.
[0032] According to at least one embodiment, the user terminal can be communicatively connected to the server device via a communication network. It can communicate via the communication network to send or receive information.
[0033] According to at least one embodiment, the server device includes at least a control unit, a RAM, a storage unit, and a communication interface, which are respectively connected by an internal bus.
[0034] According to at least one embodiment, the control unit is composed of a CPU and a ROM, executes a program stored in the storage unit, and controls the server device. The control unit also has an internal timer for measuring time. The RAM is a work area of the control unit. The storage unit is a storage area for storing programs and data. The control unit reads programs and data from the RAM and performs program execution processing based on information received from the user terminal, etc.
[0035] Disclosed is AI. According to at least one embodiment, artificial intelligence includes machine learning, deep learning, generative AI, large language models, LLM, foundation models, generative AI. Generative AI uses transformers and employs a number of mechanisms called attention. Self-supervised learning and Extract Prediction are used. In this case, the AI can guess the next word. Given a sentence, it guesses the next word from the text up to that point. A large number of supervised learning problems are created. As a result, an AI that can guess the next word can be created. Generative AI can predict grammar structures, topic connections, and that a person of such a style is likely to write such a sentence. Furthermore, generative AI can learn the structure, causal relationships, and knowledge behind just guessing the next sentence. Generative AI has a high speed of scaling, and the larger the number of parameters, the higher the accuracy. Ordinary statistics and machine learning will overfit if the model parameters are made too large compared to the data sample size. For LLM, the larger the number of parameters, the higher the accuracy. A certain generative AI has 175 billion parameters. Generative AI is covered with supervised learning to make conversations smooth. It is taught not to say strange things. It writes impression essays or acts as a call center operator.
[0036] According to at least one embodiment, a large language model (LLM) is, non-exclusively, a natural language processing model of machine learning constructed using a large dataset and deep learning techniques. Generally, a technique called "fine-tuning" for training on a specific task is used to adapt it to various natural language processing (NLP) tasks such as text classification, generation, sentiment analysis, text summarization, and question answering. According to at least one embodiment, self-supervised learning is close to human essential intelligence. When a person acts, they always predict the next event that will occur and predict the next input. In that process, they can learn the structure of the external world. Predicting the next word is an essential intelligence and is considered to be similar to what the cerebral cortex does. According to at least one embodiment, a large language model memorizes the input information but generalizes to the extent necessary to predict the next word. It does not generalize all the information from the beginning. A large language model requires capacity to memorize information. Also, parameters are required for that purpose. According to at least one embodiment, a large language model is equipped with 175 billion parameters or eight models with 220 billion parameters.
[0037] According to at least one embodiment, videos and images are represented as a set of visual patches, which are small data units similar to the text tokens of an LLM. Patches effectively represent a model of visual data and are used as a very scalable and effective representation for training generative models with various types of videos and images. First, a video is compressed into a low-dimensional latent space, and then the representation is decomposed into spatio-temporal patches to convert the video into patches.
[0038] According to at least one embodiment, a Video compression network is a network that reduces the dimension of visual data, receives raw videos as input, and outputs a temporally and spatially compressed latent representation. An AI is trained in this compressed latent space and then generates videos within this compressed latent space.
[0039] According to at least one embodiment, given a compressed input video, Spacetime Latent Patches extract a series of spatio-temporal patches that function as transformer tokens. With the patch-based representation, Sora can be trained on videos and images of various resolutions, lengths, and aspect ratios, and during inference, by arranging randomly initialized patches into a grid of appropriate size, it controls the size of the generated video.
[0040] According to at least one embodiment, the AI is a diffusion model and is trained to predict the original "clean" patches when noisy patches (and conditional information such as text prompts) are input. The AI is a diffusion transformer, which exhibits remarkable scaling properties in various areas such as language modeling, computer vision, and image generation. The diffusion transformer is also effective as a video generation model. As the computational cost of training increases, the quality of the samples improves significantly for the AI.
[0041] According to at least one embodiment, the AI applies caption regeneration technology to train a highly explanatory caption model and then uses it to generate text captions for all videos in the training set. Training highly explanatory captions improves not only the overall quality of the generated videos but also the faithfulness of the text. Utilize GPT to convert short user prompts into long detailed captions and send them to the model. This enables the AI to generate high-quality videos that exactly follow the user's prompts.
[0042] According to at least one embodiment, in natural language processing, vectorization can be performed by AI along the following process. First, cleaning processing of the given text is performed as preprocessing. In the cleaning process, unnecessary words such as JavaScript code and HTML tags included in the text are removed. Since these codes are used for display on the Internet, they are generally not used information in natural language processing. Subsequently, the text is segmented into word level by morphological analysis. Morphological analysis is to classify into the smallest language units with meaning in a sentence of natural language written in characters. As morphological analysis tools, "MeCab", "JUMAN", and "JANOME" can be used. In normalization, words with the same meaning such as writing variations are unified into one word. Stop words are words that are excluded from processing for reasons such as being unusable in natural language processing. Examples of stop words include those that do not have meaning alone, such as particles and auxiliary verbs among words. When calculating vectors, these may be removed and only meaningful words may be targeted. Vectorization may also be performed without removing these stop words. Vectorization is a process of converting a word, which is a character string, into a vector. By vectorization, word data is converted into numerical data. When converting a word into a vector, it is performed by a method called Bag of Words or distributed representation. Bag of Words is a method of vectorizing a text using the number of occurrences of words that appear in the given text. Since it focuses on how many words appear in the text, the order of words and text is not considered. Distributed representation is a method of vectorizing by focusing on the meaning of a word. By vectorizing the meaning of a word, it is possible to give vectors close to words with similar meanings and usage, and the relationship between words can also be expressed by vectors. By expressing in vectors, addition and subtraction of word meanings are possible. The application process can utilize the natural language converted into numerical data as input for machine learning. Specifically, the vectorized natural language is input into a classifier to perform text classification.Examples of tools used here include "TensorFlow", "scikit-learn", "PyTorch", etc.
[0043] Disclose about sound. In at least one embodiment, the sound includes music or sound effects. Music combines aspects such as the length, pitch, intensity, and timbre of sound to express various emotions and stories, and includes singing, musical instrument performances, and natural sounds. Music composition can be performed by determining a scale (key), chord (harmony), or melody. In at least one embodiment, once the scale is determined, the chord is also determined according to its constituent notes. There is a theory for chord progressions, and there are patterns that have been considered good progressions in past songs. Chords greatly influence the mood of a song and can have a significant impact on a person's emotions and feelings at that time. Similarly, the melody is generally based on the constituent notes of the chord and does not deviate significantly, but it tends to become monotonous, so some randomness may be required. It may be possible to prevent monotony by varying it within a range that does not deviate significantly from the chord.
[0044] A method for generating sound, which discloses a method for a device to identify an operation feature that is a feature of the operation of an object in a video (in this disclosure, the description of "method" can also be interpreted as "step" unless there is an explicitly contradictory description). In at least one embodiment, a video is a creation or data that creates movement by continuously displaying still images one by one. A video may include audio, and when the video and audio are synchronized, it may have a richer expression. An object is data operated in a video or a process related thereto. Examples of objects include moving bodies in a video (including living and non-living things. Non-exhaustive examples of the former include people, people dancing, animals, plants swaying in the wind, etc. Non-exhaustive examples of the latter include moving bodies (such as cars), the ebb and flow of waves, the blinking of lights, changes in the wavelength or intensity of light, digital expressions drawn by computer graphics). In this disclosure, when a person watches a video, an object that can be recognized as moving is regarded as this "object" or data corresponding to the "object". In this disclosure, the "operation feature" may be an "operation feature quantity" (the same applies throughout this disclosure unless there is an explicitly contradictory description). A feature quantity is a numerical value representing the feature of the data included in a certain data set. For example, when representing a person's face in a video, examples of feature quantities include "eye size", "nose height", "skin color", etc. In addition, when a person is dancing, the way the organs (including hands, feet, waist, etc.) move can also be a feature quantity. By giving these numerical values to a computer, the computer can distinguish people's faces, recognize specific people, or recognize dance choreography. In addition, in the ebb and flow of waves, the way the waves rise and the way the waves are low can naturally be feature quantities. In addition, in the case of computer graphics, the way a certain line or surface moves and the way the color changes can naturally become feature quantities. Non-exhaustive examples of the types of feature quantities are as follows. Numerical data is data that can be represented by numerical values such as height, weight, age, etc. Categorical data is data that can be represented by categories such as gender, nationality, occupation, etc. Text data is data that can be represented by text such as articles, words, keywords, etc.Image data (including still images in videos) are features extracted from images, such as color, shape, texture, etc. Video data are features extracted from videos, such as color, shape, texture, the presence or change of these over the playback time, etc. Audio data are features extracted from audio, such as pitch, frequency, timbre, etc.
[0045] In at least one embodiment, the apparatus can be created by any of the methods, machines, or AIs described above. In the present disclosure, the description of "apparatus" can be replaced with "AI" unless there is an explicitly contradictory description. The apparatus analyzes the data of the target video. The apparatus identifies the objects in the video. The user can identify the target object. As another embodiment, the apparatus identifies the object with a relatively high appearance frequency from among the objects present in the video. The apparatus can also request such identification from the AI (throughout the present disclosure, unless there is an explicitly contradictory description, the description "the apparatus performs a certain process" can be replaced with "the apparatus requests the AI to perform that process". In this case, the AI can use an AI that has learned its process and judgment by the learning methods (including supervised learning) as described above). In such a case, the target object is automatically identified by the apparatus rather than being specified by the user. The apparatus identifies the motion features of the object. As an example, for a video of a dancer dancing, the way the dancer waves their hand is a motion feature of the object. In this embodiment, it is of course also possible to take as motion features the way the dancer shakes their head, sways their waist, moves their feet, and moves their fingers. As already explained, the motion features can be identified by the user. Furthermore, the apparatus can also identify as motion features the motions with a relatively high appearance frequency from among the motions of the object.
[0046] Disclosed is a method of requiring AI to generate music with the time between similar motion characteristics as the unit time and the unit time as the beat. In the present disclosure, unless there is an explicitly contradictory description, the description of "generating music" can be replaced with "generating sound". In at least one embodiment, there are similar motion characteristics in the video. As a non-exhaustive example, regarding dance, a dancer may repeat the choreography of the same dance. In this case, the device compares the motion characteristics of each motion and recognizes them as similar motion characteristics. Regarding the criteria for determining similarity, there are a method determined by the user and a method automatically determined by the machine. In the former case, the device displays videos or still images of two or more motion characteristics on the screen of the terminal operated by the user. The user designates to the device whether the videos or still images of two or more motion characteristics are to be regarded as similar motion characteristics or not recognized as similar. The device determines whether the motion characteristics are similar or not based on the designation. In the latter case, the device determines whether the videos or still images of two or more motion characteristics contain clearly different motion characteristics. If no obvious difference in characteristics or feature amounts can be detected, these videos can be determined to be similar. In this determination, AI can learn through supervised learning about whether these are similar or not. The device can also utilize the function of the AI or require the AI to determine similarity or dissimilarity and obtain the determination result.
[0047] The device identifies the time between operation characteristics. As a non-exhaustive example, regarding dance, when a dancer repeats the choreography of the same dance (similar operation characteristics), the difference (or distance) in the appearance time (or application time) of each operation characteristic in the video is identified. As a non-exhaustive example, assume that a video of a certain dance performance is a 3-minute video. And when the dancer repeats the operation characteristics at the 1 minute 10 seconds point and the 1 minute 15 seconds point, the time between the operation characteristics is 5 seconds, and the unit time is 5 seconds. The device recognizes the time between operation characteristics as the unit time. The device can transmit the information of the unit time to other devices or AI. In other embodiments, the device can also use an AI that has learned to recognize the unit time through supervised learning as described above, or request the AI to recognize the unit time.
[0048] In at least one embodiment, the beat in sound includes a unit of rhythm. As a non-exhaustive example, the beat includes a certain time interval (a basic unit that is repeated at a certain period in music), a contrast of strong and weak (there are strong beats and weak beats, and this contrast generates rhythm), or a basis of tempo (the speed of the beat becomes the tempo, which determines the speed of the song). The beat may be a meter. The meter is the basis of the rhythm of the whole song and indicates how many beats are in a measure. For example, in 4 / 4 time, there are 4 beats in one measure, and the length of each beat is the same. In the present disclosure, the description of "beat" can be replaced by "meter" unless there is an explicitly contradictory description. The device requests the AI to generate music with the unit time as the beat (meter). As a non-exhaustive example, if a video of a certain dance performance is a 3-minute video and the unit time is 5 seconds, the device requests the AI to generate music with 5 seconds as one beat. In other embodiments, the device may generate the music itself using the AI as described above.
[0049] According to the above configuration, sound can be automatically created based on a video. Since the created sound or music has a beat based on the feature amount of the video, the degree of coincidence between the video and the generated sound is high. This can reduce the workload of the person creating the sound or automatically generate highly available sound with a high degree of coincidence between the video and the generated sound, which has industrial applicability.
[0050] When requesting an AI to generate music with a unit time as a beat, a method of requesting to generate music with the same length as the time of the video (including the time for specifying the motion feature. In the present disclosure, unless there is an explicitly contradictory description, the "time of the video" can be understood as replaced by the "length of the video" or the "length of the time of the video") will be disclosed. The device requests the AI to generate music with the same length as the length of the time of the video. As a non-exclusive example, when a video of a dance performance is a 3-minute video, since the time of the video is 3 minutes, the device requests to generate 3-minute music. In other embodiments, the device sets the time for specifying the motion feature as the length of the video. The time for specifying the motion feature is, in at least one embodiment, the entire or a part of the time zone within the video where the device or the AI is permitted to specify the motion feature. There are a method determined by the user and a method automatically determined by the device for the time for specifying the motion feature. In the former, the user can determine the time for specifying the motion feature as long as it is within the time of the video. As a non-exclusive example, when a video of a dance performance is a 3-minute video, the time from 0 minutes 0 seconds to 1 minute 30 seconds can be determined as the time for specifying the motion feature. The device specifies the motion feature only within the determined time by the above method. As an example of the latter, the device can also determine a predetermined part of the time of the video as the time for specifying the motion feature. Naturally, the entire time of the video can also be determined as the time for specifying the motion feature. The device requests to generate music with the same length as the time for specifying the motion feature.
[0051] Disclosed is a method that requires a device to generate music such that the unit time in a video is the same as the beat of the music. The description of "the same" above can be replaced with "synchronized". In at least one embodiment, the device is required to generate music at the same time as the time of the video and to generate music such that the unit time in the video is the same as the unit time of the music. As a non-exclusive example, assume that a video of a dance performance is a 3-minute video. And if the dancer repeats the movement characteristics at the 1 minute 12 second mark and the 1 minute 17 second mark, the unit time is 5 seconds. Further, the device requires the AI to generate music with a beat of 1 beat per 5 seconds and to generate music with a length of 3 minutes. And since the device requires the music to be generated such that the unit time in the video is the same as the unit time of the music, the generated sound generates beats so as to be the same as (synchronized with) the beats at the 1 minute 12 second mark and the 1 minute 17 second mark in the video. In this case, if the generated music has beats in unit times of 5 seconds starting from the 0 minute 2 second mark, it is the same as (synchronized with) the beats at the 1 minute 12 second mark and the 1 minute 17 second mark in the video. In other words, the device is required to recognize the beats of the music corresponding to the unit time in the video as specific beats and to generate music that determines the overall beats based on (or not conflicting with) the specific beats. Taking the previous example, the beats at the 1 minute 12 second mark and the 1 minute 17 second mark correspond to the specific beats, and based on the specific beats, music that determines the overall beats is generated.
[0052] According to the above configuration, it is possible to automatically create a sound synchronized with the beat of the video. The created sound or music not only has beats based on the feature amount of the video, but also is at least synchronized with the beat of the video, so the degree of coincidence between the video and the generated sound is higher. In this case, if the generated sound and the video are played simultaneously, the generated sound can be directly used industrially. This reduces the workload burden on the person creating the sound, or has industrial applicability in that it can automatically generate highly usable sounds with a high degree of coincidence between the video and the generated sound.
[0053] When asking an AI to generate music with beats per unit time, a method is disclosed that requests, from among two or more units of time resulting from similar motion characteristics, a relatively short unit of time to be used as the beat. In at least one embodiment, the device may recognize three or more similar motion characteristics from within a video. In this case, the device may recognize two or more units of time. The device recognizes the shortest unit of time from among two or more units of time resulting from similar motion characteristics as the unit of time to be used as the beat. The device can transmit such a unit of time to be used as the beat to other devices or the AI, or request the AI to recognize or identify such a unit of time to be used as the beat. As already explained, the AI learns, through supervised learning using a video having two or more units of time resulting from similar motion characteristics, to use the shortest unit of time as the beat. As a non-exhaustive example, suppose a dance performance video is three minutes long. For example, A melody, B melody, and refrain are often used mainly in pop and rock ballads and are each composed of a catchy melody. In a music video, it may be a combination of the structure of A melody, refrain, B melody, refrain, and the choreography of the refrain may be similar. And suppose the dancer repeats the motion characteristics at the 1 minute 12 second, 1 minute 17 second, 2 minute 12 second, and 2 minute 17 second time points. The device will recognize four types of units of time: 5 seconds, 55 seconds, 1 minute, and 1 minute 5 seconds. In such a case, the device recognizes 5 seconds as the shortest unit of time. The device recognizes that shortest unit of time as the beat, or requests the device or AI to recognize the shortest unit of time as the beat.
[0054] According to the above configuration, even for a complex video in which similar motion features are scattered in the video, it is possible to automatically create a sound synchronized with the rhythm of the video. The created sound or music not only has a rhythm based on the feature amount of the video, but also is at least synchronized with the rhythm of the video, so the degree of coincidence between the video and the generated sound is higher. In this case, the generated sound can be directly used industrially if the video and the generated sound are played simultaneously. This reduces the workload of the person creating the sound, or has industrial applicability in that it can automatically generate a highly available sound with a high degree of coincidence between the video and the generated sound.
[0055] When requesting AI to generate music with a unit time as the rhythm, a method is disclosed in which, from two or more unit times resulting from similar motion features, a unit time with a relatively high appearance frequency among the specified unit times is requested to be used as the rhythm. In at least one embodiment, the device calculates the appearance frequency for the recognized unit time. The appearance frequency includes a numerical value indicating how many times an event or object appears within a specific period or range. As an example, the appearance frequency includes an event (a specific word, number, action, phenomenon, etc.), a range (a specific text, dataset, time, space, etc.), and the appearance frequency of a word in a text (how many times the word "cat" appears in a certain text, etc.). The device calculates the appearance frequency of the unit time within the range of the time of the video or the time range for which the motion feature is to be specified. As a non-exhaustive example, assume that a video of a dance performance is 3 minutes long and the choreography of the dancer is recognized as the motion feature. And assume that the device recognizes that the unit time of 5 seconds appears 50 times, the unit time of 10 seconds appears 10 times, the unit time of 20 seconds appears 3 times, and the unit time of 1 minute appears 1 time. In this case, the device recognizes the 5 seconds with the highest appearance frequency as the unit time to be used as the rhythm. The device either generates the music with this unit time as the rhythm itself or requests AI to generate it in this way.
[0056] According to the above configuration, as already described, even for a complex video in which similar motion features are scattered in the video, it is possible to automatically create a sound synchronized with the rhythm of the video. Furthermore, since the unit time with the highest appearance frequency is used as the rhythm, the rhythm of the generated music is less likely to be inconsistent with the rhythm of the video. Therefore, the created sound or music not only has a rhythm based on the feature amount of the video, but also is at least synchronized with the rhythm of the video, so the degree of coincidence between the video and the generated sound is higher. In this case, if the generated sound and the video are played simultaneously, the generated sound can be directly used industrially. This reduces the workload burden on the person creating the sound, or has industrial applicability in that it can automatically generate a highly usable sound with a high degree of coincidence between the video and the generated sound.
[0057] When requesting an AI to generate music with beats per unit time, music of the same time as the time of the video or the time to identify the motion characteristics should be generated. When two or more unit times with dissimilar motion characteristics are identified, two or more pieces of music should be generated and requested to be concatenated. Disclosed is a method. In at least one embodiment, the device recognizes unit times based on similar motion characteristics as described above. In this case, the device will recognize the unit time based on the motion characteristic A and the unit time based on the motion characteristic B. These are two or more unit times with dissimilar motion characteristics. When two or more unit times with dissimilar motion characteristics are identified, the device generates two or more pieces of music. In other embodiments, an AI or other device is requested to generate two or more pieces of music in this way. The device concatenates this generated music. Concatenation includes a state where a plurality of melodies are connected. The concatenation may be connected in a smooth and natural flow. As a non-exhaustive example, the device generates music A generated at a unit time based on the motion characteristic A and music B generated at a unit time based on the motion characteristic B from a video (including music videos and movies). The device concatenates music A and music B. There is a method of concatenating music data so that another piece of music can be played simultaneously with the end of a certain piece of music. There is also a method of concatenating a part of the music data and a part of the data of another piece of music so that a certain piece of music ends halfway and another piece of music is played. In at least one embodiment, when generating two or more pieces of music, the music obtained by modulating one piece of music can be used as the two pieces of music.
[0058] According to the above configuration, even for a complex video in which two or more unit times with dissimilar motion characteristics are specified, a sound synchronized with the rhythm of the video can be automatically created. Furthermore, since it is a sound based on two or more unit times, the possibility of being inconsistent with the rhythm of the video is further reduced. Therefore, since most of the created sound or music is synchronized with the rhythm of the video, the degree of coincidence between the video and the generated sound is higher. In this case, the generated sound can be directly industrially utilized if the video and the generated sound are played simultaneously. This reduces the workload burden on the person creating the sound or can automatically generate a highly usable sound with a high degree of coincidence between the video and the generated sound, which has industrial applicability.
[0059] Disclosed is a method of requesting to concatenate the music at a time point specified based on the occurrence frequencies of two or more operation characteristics in the method. In at least one embodiment, the apparatus generates music at the same time as the time of the video or the time for which the operation characteristics are to be specified, and specifies the time points at which the unit time changes for two or more unit times. As a non-exhaustive example, the apparatus calculates the occurrence frequency or occurrence probability in the time zone of the video for each unit time based on each operation characteristic. For example, the apparatus specifies a time zone in which the occurrence of the unit time based on the operation characteristic A is relatively high and a time zone in which the occurrence of the unit time based on the operation characteristic B is relatively high within the time of the video or the range of the time for which the operation characteristics are to be specified. The apparatus generates music based on the unit time with the highest occurrence frequency based on these occurrence frequencies. When the unit time with the highest occurrence frequency changes, the apparatus generates music based on the new unit time. The apparatus concatenates these two pieces of music. In another embodiment, the apparatus generates music at the same time as the time of the video (the time for which the operation characteristics are to be specified), and generates two or more pieces of music when two or more dissimilar operation characteristics are specified. The apparatus specifies, as the change time, the time point specified based on the occurrence frequencies of two or more dissimilar operation characteristics (in this paragraph, unless explicitly stated to the contrary, the descriptions of "time point" and "time" can be replaced with "time zone"). For example, the apparatus specifies a time zone in which the occurrence of the unit time based on the operation characteristic A is relatively high and a time zone in which the occurrence of the unit time based on the operation characteristic B is relatively high within the time of the video or the range of the time for which the operation characteristics are to be specified. The apparatus can specify, as the change time, the time point at which the operation characteristic with the highest occurrence frequency changes, or the time between or approximately the midpoint of the time zones with high occurrence frequencies based on different operation characteristics. As a non-exhaustive example, assume that a video of a dance performance is 3 minutes long. If the occurrence of the unit time based on the operation characteristic A is relatively high in the time zone from 0 minutes 0 seconds to 1 minute 30 seconds and the occurrence of the unit time based on the operation characteristic B is relatively high in the time zone from 1 minute 30 seconds to 3 minutes, the apparatus specifies 1 minute 30 seconds, which is approximately the midpoint, based on the occurrence frequencies of the two or more operation characteristics. The apparatus requests to concatenate the music at the specified time point.That is, the generated music is music A based on the operation feature A from 0 minutes and 0 seconds to 1 minute and 30 seconds, and music B based on the operation feature B from 1 minute and 30 seconds to 3 minutes. These are connected at the 1 minute and 30 seconds mark.
[0060] According to the above configuration, as already explained, even for a complex video in which two or more unit times based on dissimilar operation features are specified, it is possible to automatically create a sound synchronized with the rhythm of the video. Furthermore, since it is a sound based on two or more unit times, the possibility of being inconsistent with the rhythm of the video is further reduced. Therefore, the created sound or music is mostly synchronized with the rhythm of the video, so the degree of coincidence between the video and the generated sound is higher. In this case, the generated sound can be directly used industrially if the video and the generated sound are played simultaneously. By this, it is possible to reduce the workload of the person creating the sound, or there is industrial applicability in that it is possible to automatically generate a highly available sound with a high degree of coincidence between the video and the generated sound.
[0061] In at least one embodiment, the video includes movies, dances, and choreographies. As an example, it includes videos published on Youtube (registered trademark). In this case, it is possible to automatically generate music suitable for the published video. The video includes live-streamed videos. In this case, the device can generate music according to the conditions or configurations described in any of the present disclosures simultaneously with the live stream. The device can distribute this generated music simultaneously with the live-streamed video. According to the above configuration, it is possible to provide music that matches the rhythm of the video simultaneously with the live stream. Such a highly flexible generation method cannot be achieved by human composers, so it has an excellent effect compared to conventional technologies and has industrial applicability. In at least one embodiment, the device can request the AI to analyze the ambient sound and the voices of the audience during distribution in real time, recognize one or more of them as specific sounds, features, and feature quantities, and generate sounds and sound effects corresponding thereto.
[0062] In at least one embodiment, the video includes videos used in VR or AR. VR (Virtual Reality) is a technology that provides an experience of fully immersing in a virtual world. By wearing a device such as a VR headset, users can feel as if they are in another world. AR (Augmented Reality) is a technology that superimposes digital information on the real world. Through the camera of a smartphone, virtual objects can be displayed or information can be added to the real scenery. The device can generate music according to the conditions or configurations described in any of the present disclosures for the videos presented to the user through VR or AR. The device can output the generated music as audio at the same time as presenting it to the user through VR or AR. The audio output includes embodiments of outputting audio from the user terminal. According to the above configuration, music that matches the rhythm of the VR or AR video can be provided. This can reduce the workload of the person creating the sound or automatically generate highly available sounds with a high degree of match between the video and the generated sound, which has industrial applicability.
[0063] In at least one embodiment, the AI repeats learning to predict the next sound through unsupervised learning. As already explained, the composition process is hierarchical according to its structure. The elements that make up music are at least scale (key), chord (harmony), and melody. Once the scale is determined, the chord is also determined according to the constituent notes. There is a theory for chord progressions, and there are also many patterns that have been considered good progressions in past songs and have been accumulated. It is said that chords greatly influence the mood of a song and have a very significant impact on people's emotions and feelings at that time. The AI repeats learning to predict the next sound through unsupervised learning using music data. Furthermore, the AI performs supervised learning on the music data regarding what kind of emotions the music exhibits.
[0064] In at least one embodiment, the device or AI generates music based on music or related information that has already been completed and uploaded on the Internet or recorded in a database, regardless of whether the author is registered. As a non-exhaustive example, when the unit time based on the operation feature A is 5 seconds, the device identifies music with a tempo of 5 seconds that has already been completed and uploaded on the Internet. When multiple pieces of music are identified, the music that meets any of the conditions described in this disclosure is identified. The device or AI modifies the identified music or generates music.
[0065] In at least one embodiment, the device or AI identifies one piece of music based on music or related information that has already been completed and uploaded on the Internet or recorded in a database, regardless of whether the author is registered. In this embodiment, the device or AI is configured to identify existing music that meets the conditions, rather than generating music. The device presents the identified music or information about the music to the user.
[0066] In at least one embodiment, the device or AI uses music or related information that has already been completed and uploaded on the Internet or recorded in a database, regardless of whether the author is registered, to generate music.
[0067] In at least one embodiment, the device adds data indicating that the music was generated by AI to the data of the music generated by AI. When providing the user with the music that generated the video, if the device recognizes that the music has such data, it presents to the user that the music was generated by AI, or displays that fact on the user terminal. In another embodiment, when the device generates music using music that has already been uploaded on the Internet or recorded in a database as a completed work, it adds information about the name of the original music or the copyright holder to the data of the generated music. If the device recognizes that the music contains such information, it presents to the user the name of the music, that its copyright holder created it, or that AI was generated based on that music, or displays that fact on the user terminal. In at least one embodiment, the device generates data or a prompt for copyright display in a format compatible with the distribution platform. As an example, on YouTube (registered trademark), data or a prompt for copyright display defined by the platform is generated, or generated along with the video for which such data was generated. These may be referred to as copyright management tools. According to the above configuration, there is industrial applicability in that the risk of copyright infringement can be reduced in the data of music added or generated by AI.
[0068] In at least one embodiment, the user can specify the sounds to be relied on in generating music for the device. As an example, the device receives the sound data to be relied on. The device has recorded two or more sound sources that the user can specify, and the user can select any sound source. "Two or more sound sources are recorded" includes not only embodiments in which the device itself stores these sound sources, but also embodiments in which the device obtains these sound sources from an external storage medium. For example, these sound sources are existing music. The user selects one or more sound data to be relied on, and the device generates music based on the selected music. As a non-exhaustive example, when the video has a unit time of 5 seconds and the selected sound data is existing music with a tempo of 7 seconds, the device adjusts the playback speed of the existing music to be relied on to change it to music with a tempo of 5 seconds. In this example, the changed music corresponds to the "generated music" in the above embodiment. According to the above configuration, it is possible to automatically identify existing music that synchronizes with the tempo of the video. Furthermore, since the identified existing music is sound based on the unit time, it matches the tempo of the video. In this case, if the presented sound is played simultaneously with the video and the generated sound, it can be directly used industrially. This reduces the workload of the person creating the sound, reduces the workload of searching for music with a matching tempo, or can automatically prepare highly available sounds with a high degree of matching between the video and the generated sound, which has industrial applicability. In these embodiments, the device can generate or save, within the data of the generated music or attached to the data, data or a prompt indicating the license for the existing sound used. The advantage of such an embodiment is at least that it can provide a system for smoothly managing copyrights and granting licenses.
[0069] In at least one embodiment, the device analyzes the motion characteristics in the video (for example, dance movements) in more detail and generates different music elements (beats and melodies) for each motion characteristic (for example, dance steps). These music elements are adopted or concatenated based on any of the methods in the present disclosure. These music elements may be added to the generated music as long as they do not conflict with the beats of the generated music.
[0070] In at least one embodiment, the device recognizes a predetermined motion characteristic of an object in a video as a specific motion characteristic, and generates a specific sound, which is a predetermined sound, at the time when the specific motion characteristic occurs. The device can generate the specific sound in addition to the generated music. In another embodiment, the device can generate the specific sound independently of the generated music. The user registers information regarding the specific motion characteristic or the specific sound with the device. As a non-exhaustive example, the user registers with the device a state where a person is surprised or a predetermined reaction as information regarding the specific motion characteristic, and a state of being surprised or a prompt as information regarding the specific sound. The specific sound includes existing music and sound effects specified by the user. The device can store two or more such existing music and sound effects by itself, or receive a transfer from an external storage device. The device can present two or more such available specific sounds to the user. The user can specify the specific sound to be used from among them. When two or more specific sounds are specified, the device uses, as the specific sound to be generated, the specific sounds alternately, randomly, or in accordance with the embodiments of the present disclosure. When the video is a two-hour movie and at 1 hour 20 minutes and 43 seconds, a person in the movie shows a surprised state, the device generates a specific sound representing a surprised sound at that time. According to the above configuration, a specific sound based on a specific motion characteristic can be automatically generated. Further, since the generated specific sound is a sound based on the time when the specific motion characteristic occurs, it coincides with the timing when the specific motion characteristic of the video occurs. In this case, if the presented sound and the generated sound are played simultaneously with the video, it can be directly industrially applicable. This reduces the workload of the person creating the specific sound, or automatically prepares a highly available sound with a high degree of coincidence between the video and the generated sound, which has industrial applicability. As already described, this configuration can be combined with or replaced by one or more of the aspects in other embodiments of the present disclosure as long as there is no contradiction. That is, it is also possible to analyze the ambient sound and the voices of the audience during distribution in real time and generate corresponding BGM and sound effects (specific sounds). In other embodiments, it can be implemented as an application that automatically inserts a surprised sound effect for a surprised motion.
[0071] In at least one embodiment, the video includes a movie. According to this embodiment, for each sequence of the movie, an optimal music or a specific sound can be generated separately.
[0072] Disclosed are embodiments different from the above.
[0073] Disclosed is a method that requires an AI for a device to identify features of an object in a video and generate music based on the features. This method includes a method that requires an AI to identify motion features that are features of motion and generate music based on the motion features. According to at least one embodiment, the device or AI represents videos or images as a set of visual patches, which are small data units similar to text tokens of an LLM. The device combines video recognition AI and an LLM to understand each scene of the video. Specifically, it utilizes video recognition AI to separately recognize various objects and environments such as people, cars, buildings, animals, trees, etc. that make up the scene, and weather, as well as their changes. The AI learns the emotions corresponding to the visual patches through the learning methods (including supervised learning) as described above. Further, the AI may be pre-finetuned with the LLM using sample videos in the target field. For example, for a specific visual patch, the AI learns to correspond to a mood or atmosphere, emotion (collectively referred to as "concept" in this disclosure). The device requires the AI to generate music based on the features based on the motion features. As an example, when the video is a 3-minute music video and the dancer's expression is a smiling face, the device identifies motion features such as bright and happy, and requires the AI to generate music based on those features. In other embodiments, the device also requires the AI to identify motion features that are features of the motion of an object in the video. As another example, when the video is a 3-minute music video and there is an artistic ruin behind the dancer, the device identifies accessory features such as ruin and dark as accessory features, and requires the AI to generate music based on those accessory features. In other examples, the video is a movie, and it is possible to analyze the scenes and emotional flow of the video and generate a corresponding music composition (intro, climax, ending). Also, the music corresponding to the concept includes embodiments corresponding to music genres (classical, EDM, hip-hop, etc.). According to this embodiment, it is possible to automatically generate music along with the taste or concept of the video.This can reduce the workload of the person creating the sound, or has industrial applicability in that it can automatically prepare highly available sounds with a high degree of concept match between the video and the generated sound.
[0074] In at least one embodiment, the device can also use features specified by the user as operation features or accessory features. The user inputs, in text, to the device the concept of the music to be generated. These texts may be prompts. The device requests the AI to generate music based on the concept input by the user. Further, the device may request the AI to generate music at the same time as the time of the video or the time to identify the operation feature. In one example, the user can select a mood or atmosphere and generate music that matches the selected mood. According to this embodiment, since the user can specify their preferred taste, music can be automatically generated in line with the user's taste. This can reduce the workload of the person creating the sound, or has industrial applicability in that it can automatically prepare highly available sounds with a high degree of concept match between the video and the generated sound.
[0075] In at least one embodiment, the device analyzes the lyrics as a feature specified by the user. The device or the AI recognizes the concept included in the lyrics as a feature. For example, the device or the AI extracts concepts such as love, heartbreak, sadness, etc. based on the lyrics. The device or the AI generates sound based on the extracted concept. When two or more concepts are recognized based on the lyrics, as will be described later, the AI is requested to be based on the feature with the highest appearance frequency for which the feature is recognized. According to this method, even if the lyrics are complex, music can be generated based on the main concept of the lyrics. This can reduce the workload of the person creating the sound, or has industrial applicability in that it can automatically prepare highly available sounds with a high degree of concept match between the video and the generated sound.
[0076] In at least one embodiment, the device requests the AI to convert the objects in the video into text. As already explained, the AI represents the video or image as a set of visual patches, which are small data units similar to the text tokens of the LLM. The device combines the video recognition AI and the LLM to understand each scene of the video. Specifically, the video recognition AI is utilized to individually recognize various objects and environments in the scene, such as people, cars, buildings, animals, natural objects like trees, and weather, as well as their changes. The AI learns the text corresponding to the visual patches by the learning method (including supervised learning) as described above. By this method, the AI can describe (convert into text) the given video in text. The AI generates music based on the converted text. Specifically, the device requests the AI to identify features based on the converted text and generate music based on the features. The method of generating music based on features is as disclosed elsewhere in the present disclosure.
[0077] When asking an AI to generate music based on features, if there are two or more features, a method is disclosed that requires the AI to at least base it on the feature with a relatively longer time during which the feature is recognized. In at least one embodiment, the device may recognize two or more features. In this case, the AI is required to base it on the feature with the longest time during which the feature is recognized in the video. The device identifies the time during which each feature is recognized. As a non-exclusive example, if the video is a 3-minute movie, 2 minutes and 40 seconds of which depict a lovelorn state, while 20 seconds depict a happy and excited state of being in love as a past recollection, the feature of "lovelorn" is 2 minutes and 40 seconds, and the feature of "happy" is 20 seconds. That is, the device asks the AI to create music based on the feature of "lovelorn", which is the longest feature. Note that this method is not intended to be limited to one feature, and it is also possible to generate music based on two or more features. According to the above configuration, even for a video having multiple concepts, music based on the main concept can be generated. Therefore, even for a video containing conflicting concepts, since it is music based on the main concept, there is a high possibility that the concept will match the video. This can reduce the workload burden on the person creating the sound, or automatically prepare a highly available sound with a high degree of concept match between the video and the generated sound, which has industrial applicability.
[0078] When asking an AI to generate music based on features, when there are two or more features, at least disclose a method that requires the AI to be based on features with a relatively high frequency of appearance where the features are recognized. As already explained, in at least one embodiment, the device may recognize two or more features. In this case, ask the AI to be based on the feature with the highest frequency of appearance in the video. The device identifies the number of times each feature is recognized. As a non-exhaustive example, if the video is a 3-minute movie, and the number of times the actor laughs is 15, the number of times crying is 1, and the number of times angry is 1, the feature with the highest frequency is "laughing". That is, the device asks the AI to create music based on the feature of "laughing", which is the longest feature. Note that this method is not intended to be limited to one feature, and it is also possible to generate music based on two or more features. According to the above configuration, even for a video having multiple concepts, music based on the main concept can be generated. Therefore, even for a video containing conflicting concepts, since it is music based on the main concept, there is a high possibility that the concept will match the video. This reduces the workload burden on the person creating the sound, or automatically prepares a highly available sound with a high degree of concept match between the video and the generated sound, which has industrial applicability.
[0079] When requesting an AI to generate music based on features, generate music at the same time as the time for identifying the time or motion features of the video, and when two or more dissimilar features are identified, request to generate two or more pieces of music and concatenate the music. As already explained, the device or AI can generate two or more pieces of music when two or more features are identified. Also, as already explained, the device or AI can concatenate two or more pieces of music. According to the above configuration, even for a complex video having two or more dissimilar features, a sound similar to the concept of the video can be automatically created. Furthermore, since it is a sound based on two or more concepts, the possibility of not matching the concept of the video is further reduced. Therefore, the created sound or music is mostly synchronized with the concept of the video, so the degree of coincidence between the video and the generated sound is higher. In this case, the generated sound can be directly used industrially if the video and the generated sound are played simultaneously. This reduces the workload burden on the person creating the sound, or has industrial applicability in that it can automatically generate a highly available sound with a high degree of concept match between the video and the generated sound.
[0080] Disclosed is a method that, at a time point specified based on the occurrence frequencies of two or more features, requests that the music be concatenated. As already explained, the device can calculate the occurrence frequency or occurrence probability of each feature in the time period of the video. For example, the device identifies a time period in which the occurrence of the unit time based on feature A is relatively high and a time period in which the occurrence of the unit time based on feature B is relatively high within the time of the video or the range of time for which the motion features are to be identified. The device generates music based on the unit time with the highest occurrence frequency based on these occurrence frequencies. When the unit time with the highest occurrence frequency changes, the device generates music based on the new unit time. The device concatenates these two pieces of music. In another embodiment, the device generates music at the same time as the time of the video (the time for which the motion features are to be identified), and when two or more dissimilar motion features are identified, generates two or more pieces of music. The device specifies, as the change time, the time point specified based on the occurrence frequencies of two or more dissimilar motion features. For example, the device identifies a time period in which the occurrence of the unit time based on feature A is relatively high and a time period in which the occurrence of the unit time based on feature B is relatively high within the time of the video or the range of time for which the motion features are to be identified. The device can specify, as the change time, the time point when the motion feature with the highest occurrence frequency changes, or the time period between or the approximate midpoint of the time periods with high occurrence frequencies based on different motion features. As a non-exclusive example, assume that the video of a certain movie is 10 minutes long. If in the time period from 0 minutes 0 seconds to 1 minute 30 seconds, the occurrence of the unit time based on feature A (for example, the feature that the stage is in a ruin) is relatively high, and in the time period from 1 minute 30 seconds to 3 minutes, the occurrence of the unit time based on feature B (for example, the feature that the stage is in a magnificent building (such as a castle)) is relatively high, the device specifies 1 minute 30 seconds, which is the approximate midpoint, based on the occurrence frequencies of the two or more motion features. The device requests that the music be concatenated at the specified time point. That is, the generated music is music A based on feature A from 0 minutes 0 seconds to 1 minute 30 seconds and music B based on feature B from 1 minute 30 seconds to 3 minutes. These are concatenated at the time point of 1 minute 30 seconds.
[0081] According to the above configuration, as already explained, even for a complex video in which two or more unit times with dissimilar features are specified, a sound similar to the concept of the video can be automatically created. Further, since it is a sound based on two or more unit times, the possibility of being inconsistent with the concept of the video is further reduced. Therefore, since the created sound or music is mostly synchronized with the concept of the video, the degree of coincidence between the video and the generated sound is higher. In this case, the generated sound can be directly used industrially if the video and the generated sound are played simultaneously. This reduces the workload burden on the person creating the sound, or has industrial applicability in that it can automatically generate a highly usable sound with a high degree of concept match between the video and the generated sound.
[0082] In at least one embodiment, when generating music based on lyrics or text, the device recognizes the characters or concepts indicating rhythm in the text and requests the AI to generate music with a rhythm based on the characters or concepts. According to the above configuration, since music with a specified rhythm can be generated, the possibility of the generated music being inconsistent with the concept of the video is further reduced. Therefore, the generated sound can be directly used industrially if the video and the generated sound are played simultaneously. This reduces the workload burden on the person creating the sound, or has industrial applicability in that it can automatically generate a highly usable sound with a high degree of concept match between the video and the generated sound.
[0083] Disclosed are embodiments different from the above. The device requests the AI to create lyrics based on music data. The device or the AI decomposes the music data into units such as scale, chord, or melody. The AI learns the mood, atmosphere, or emotion indicated by each unit through supervised learning. Further, the AI performs learning and fine-tuning based on the combination of scale, chord, or melody and lyrics of music that has already been published. As a result, when given a scale, chord, or melody, the AI can predict what kind of characters or lyrics will come. Furthermore, for the characters or lyrics that come next, it is also possible to predict what kind of words will come next. In this way, the AI can create lyrics based on music data. The device requests the AI to create lyrics based on music data. According to the above configuration, since lyrics can be generated based on traditional concepts of music, the possibility that the generated lyrics will not match the concept of the music is reduced. Further, since the AI performs natural language processing and learning, the possibility that the generated lyrics will contain grammatical errors is also surely low. Therefore, the generated lyrics can be directly used industrially along with the music. This reduces the workload of the person creating the lyrics or has industrial applicability in that it can automatically generate highly usable lyrics with a high degree of concept matching between the music and the generated lyrics. In at least one embodiment, the device further requests the AI to generate a singing voice based on the generated lyrics. In at least one embodiment, the device requests the AI to generate music, requests the AI to create lyrics based on the generated music, requests the AI to generate a singing voice based on the generated lyrics, and requests the AI to combine (integrate, combine) the generated singing voice and music data (or the device itself combines the data). Applications of the AI for generating a singing voice include Vocaloid (registered trademark). This method has industrial applicability in that it can automatically generate a singing voice that matches the concept of the video.
[0084] The following discloses the outline of the embodiments described above.
[0085] A method of creating sound, wherein a device identifies operation characteristics that are characteristics of the operation of an object in a video, and requests an AI to generate music with the unit time being the beat, where the unit time is the time between similar operation characteristics. Method.
[0086] When requesting an AI to generate music with the unit time being the beat, generate music at the same time as the time of the video or the time for which the operation characteristics are to be identified, and request to generate music such that the unit time in the video is the same as the beat of the music. Method.
[0087] When requesting an AI to generate music with the unit time being the beat, request to use, as the beat, a relatively short unit time from among two or more unit times resulting from similar operation characteristics. Method.
[0088] When requesting an AI to generate music with the unit time being the beat, request to use, as the beat, a unit time with a relatively high appearance frequency among two or more unit times resulting from similar operation characteristics. Method.
[0089] When requesting an AI to generate music with the unit time being the beat, generate music at the same time as the time of the video or the time for which the operation characteristics are to be identified, and when two or more unit times resulting from dissimilar operation characteristics are identified, generate two or more pieces of music, and request to concatenate the music. Method.
[0090] In the above method, At a time point specified based on the occurrence frequencies of two or more operation characteristics, request to concatenate the music. Method.
[0091] A method for creating sound, wherein the device identifies the characteristics (feature quantities) of the objects in the video, and requests the AI to generate music based on the characteristics. Method.
[0092] When requesting the AI to generate music based on the characteristics, if there are two or more characteristics, request the AI to at least base on the characteristics for which the recognized time is relatively long. Method
[0093] When requesting the AI to generate music based on the characteristics, if there are two or more characteristics, request the AI to at least base on the characteristics for which the recognized occurrence frequency is relatively high. Method
[0094] When requesting the AI to generate music based on the characteristics, generate music at the same time as the time of the video or the time for which the characteristics should be identified, if two or more non-similar characteristics are identified, generate two or more pieces of music, and request to concatenate the music. Method.
[0095] In the above method, at a time point specified based on the occurrence frequencies of two or more characteristics, request to concatenate the music. Method.
[0096] In at least one embodiment, the description "requiring that, among two or more unit times resulting from similar operation characteristics, a relatively short unit time be taken as a beat" can be replaced with "requiring that, among two or more unit times resulting from similar operation characteristics, the shortest unit time be taken as a beat". The description "requiring that, among two or more unit times resulting from similar operation characteristics, a unit time with a relatively high occurrence frequency of a specified unit time be taken as a beat" can be replaced with "requiring that, among two or more unit times resulting from similar operation characteristics, a unit time with the highest occurrence frequency of a specified unit time be taken as a beat". The description "when there are two or more features, requiring the AI to be based at least on a feature with a relatively long time during which the feature is recognized" can be replaced with "when there are two or more features, requiring the AI to be based at least on a feature with the longest time during which the feature is recognized". The description "when there are two or more features, requiring the AI to be based at least on a feature with a relatively high occurrence frequency of the recognized feature" can be replaced with "when there are two or more features, requiring the AI to be based at least on a feature with the highest occurrence frequency of the recognized feature".
[0097] A method of creating sound, wherein the device, generates music at the same time as the time of the video or the specified time, and requires the AI to generate music based on the characteristics of the lyrics corresponding to the video. Method.
[0098] In the above configuration, the specified time includes the time specified by the user. The device requires the AI to generate music within the range of the time specified by the user. The method of generating music based on the characteristics of the lyrics and its advantages are as described above.
[0099] In at least one embodiment, the user can make fine adjustments to the music or vocals generated by the AI (tempo, key, effects) through the device. The device changes the tempo, key, effects, etc. of the generated music or vocal data according to the user's adjustments. Such adjustments can be made even if the video is a live video or a distributed video.
[0100] In at least one embodiment, two or more users can jointly make inputs to the device regarding the rhythm and concept to be generated, and make fine adjustments to the music (tempo, key, effects) for the music or vocals generated by the AI through the device. The device receives information inputs from two or more users. Two or more users can jointly incorporate one piece of data through the device. Such editing can be done even if the video is a live video or a distributed video. As a non-exhaustive example, this editing function is referred to as "collaboration mode" or "real-time mode".
[0101] In at least one embodiment, the device requests the AI to create lyrics in one or more foreign languages according to the lyrics and concept. The device requests the AI to create music according to the input foreign language lyrics or concept. The AI performs fine-tuning on the concept and music corresponding to the foreign language. As a non-exhaustive example, for the concept of Japan, koto is used for fine-tuning, and for the concepts of Europe and Ireland, bagpipes are used for fine-tuning for concepts related to ethnic music. Also, when foreign language lyrics are input, the device requests the AI to generate music with the concept of that foreign country.
[0102] In at least one embodiment, the device requires the AI to generate music or video in a compatible format such that the generated music or video operates on major distribution platforms (non-exhaustively, such as YouTube, TikTok, Instagram, etc., all trademarks). In other embodiments, the device converts the music or video generated by the AI into a compatible format such that it operates on major distribution platforms. The advantage of such embodiments is, at least, to create platform-compatible data and improve user convenience.
[0103] The invention according to the present disclosure only needs to be able to exhibit at least one of the above-described effects.
Claims
1. 1. A method of producing a sound, comprising: The device, Identifying motion features that are characteristic of the motion of a moving object in a video; The AI is required to generate music with the time between similar motion features as a unit of time, with the unit time being one beat. method.
2. 1. A method of producing a sound, comprising: The device, Identifying motion features that are characteristic of the motion of an object in a video; The AI is required to generate music with the time between similar motion features as one beat, method.
3. 3. A method according to claim 1 or 2, comprising: When you ask an AI to generate music with a unit time being a beat, Generate music of the same duration as the time of the video or the time at which the action feature is to be identified; Requires that music be generated so that the unit time in the video and the beat of the music are the same. method.
4. 3. A method according to claim 1 or 2, comprising: Video includes movies, Generate music for each sequence or scene of a movie; method.
5. An apparatus for producing sound, comprising: Identifying motion features that are characteristic of the motion of a moving object in a video; The AI is required to generate music with the time between similar motion features as a unit of time, with the unit time being one beat. Device.
6. An apparatus for producing sound, comprising: Identifying motion features that are characteristic of the motion of an object in a video; The AI is required to generate music with the time between similar motion features as one beat, Device.
7. An apparatus as claimed in claim 5 or 6, comprising: When you ask an AI to generate music with a unit time being a beat, Generate music of the same duration as the time of the video or the time at which the action feature is to be identified; Requires that music be generated so that the unit time in the video and the beat of the music are the same. Device.
8. An apparatus as claimed in claim 5 or 6, comprising: Video includes movies, Generate music for each sequence or scene of a movie; Device.
Citation Information
Patent Citations
Automatic musical composition device, automatic composition method, automatic composition program and memory medium
JP2002287746A
Music generation system
JP2006154777A
Automated Music Composition and Generation Machines, Systems and Processes Employing Language and / or Graphical Icon-Based Music Experience Descriptors
JP2018537727A
Music content generation
JP2023513586A
JPP3578464B