A method for displaying and generating sound sequences

EP4744040A1Pending Publication Date: 2026-05-20WETWEAK SA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
WETWEAK SA
Filing Date
2024-07-10
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Existing generative artificial intelligence systems for sound sequences face limitations in accurately defining the style of sound sequences using text prompts, leading to inefficient and imprecise audio content generation.

Method used

A method involving the retrieval of sound sequences, display of an interactive 2D or 3D sound map where sound sequences are represented as points with distances inversely proportional to their similarity, allowing users to select zones for generating new sound sequences that share similarities with neighboring sounds, utilizing self-learning machine learning and dimensionality reduction techniques like t-SNE for personalized and high-quality audio creation.

Benefits of technology

Enables faster and more precise generation of user-specific, high-quality audio content, revolutionizing user interaction with AI, and providing robust traceability and copyright protection through watermarked audio sequences, enhancing creativity and transparency in audio content creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024056710_16012025_PF_FP_ABST
    Figure IB2024056710_16012025_PF_FP_ABST
Patent Text Reader

Abstract

A method for generating sound sequences, comprising the steps of: retrieving (100) a plurality of sound sequences; displaying (102) on a display an interactive 2D or 3D sound map (4) where said sound sequences are represented with a point so that the distance between said points is inversely proportional to a similarity between the associated sound sequences; selecting (103) a zone on said interactive 2D or 3D sound map to trigger generation (104) of a new sound sequence with a generative artificial intelligence system, so that the new sound sequence shares similarities with all the neighbour sound sequences.
Need to check novelty before this filing date? Find Prior Art

Description

A method for displaying and generating sound sequencesTechnical domain

[0001] The present invention concerns an improved method for displaying and generating sound sequences using a generative artificial intelligence system.Related art

[0002] Generative artificial intelligence has already been used for generating sound sequences, Including lyrics or complete songs.

[0003] One difficulty with existing systems is to enter prompts, i.e., to define with words or otherwise the type of sound sequence one wants to generate. Text prompts like "please generate a romantic love song with violins" or "I need an Intro with lot of bass and 128 bpm for a techno song" offer only limited ways In defining the style of sound sequence one wants to generate.Short disclosure of the invention

[0004] An aim of the present invention is to provide a new method for telling a generative artificial intelligence system what type of sound sequence a user wants to generate.

[0005] Those aims are reached with a method for generating sound sequences, comprising the steps of: retrieving a plurality of sound sequences; displaying an interactive 2D or 3D sound map where said sound sequences are represented with a point so that a distance between said points is inversely proportional to a similarity between the associatedsound sequences; selecting a zone on said interactive sound map to trigger generation of a new sound sequence with a generative artificial intelligence system, so that the new sound sequence shares similarities with all the neighbour sound sequences.

[0006] Preferred embodiments are indicated in the dependent claims.

[0007] The distance may be explicitly computed.

[0008] The method may include a step of determining a similarity between said sound sequences, prior to said displaying a 2D or 3D sound map. The similarity is inversely proportional to the distance between vectors assigned to the sound sequence.

[0009] The vectors are assigned to the sound sequences such that their distance is inversely proportional to the similarity between the sound sequences. This assignment is possible without explicitly determining this distance.

[0010] The invention thus provides a solution for audio creators and users to let generate audio content in a faster and more precise.

[0011] The method not only enables the creation of high-quality audio content tailored to user needs but also revolutionizes the interaction between users and Al.

[0012] The system autonomously learns from user interactions and generates new sounds based on their preferences, offering a personalized, unique experience.

[0013] Some aspects of the invention lie in the combination of autonomous learning capabilities, innovative audio mapping applications, and robust traceability features.

[0014] The audio sequences generated are both user-specific and high- quality, going beyond traditional keyword-based systems.

[0015] Additionally, each audio sequence may be watermarked. The watermark offers copyright protection with an audio token that can be tracked. This ensures transparency, traceability, and rightful remuneration for content creators.

[0016] The method is based on a self-learning machine learning technology that, based on user preferences and behaviors, can create new sounds tailored to the specific needs of the user. This ability to autonomously learn from user interactions and generate new sounds represents a significant advancement in the field of audio content creation.

[0017] Moreover, the system leverages dimensionality reduction techniques, such as t-SNE, to create an interactive sound map where similar sounds are grouped. By hovering over the map with the mouse, the user can listen to each sound. The sound map evolves autonomously to improve sound search, increase similarity between sounds, and generate new sounds by combining a wide catalogue to complete the "missing" sound between two sounds.

[0018] The invention also offers advanced techniques to ensure the traceability and transparency of each audio content used within the system. Each audio sequence integrated into the Al's generative content database is watermarked, thereby creating a unique identifier for each audio sequence, preserving its originality, and protecting it from unauthorized use. This watermark serves as a digital signature, providing traceability within the ecosystem and enabling transparent marketplace distribution. It also provides content creators with the ability to track in real-time theusage and distribution of their audio sequences, thereby ensuring the integrity of their work and enabling them to receive rightful recognition and remuneration.

[0019] The generative artificial intelligence system can integrate features such as lyric generation, mood-based audio creation, sound quality adaptation, music teaching assistance, community functions, automatic mixing and mastering, and voice-activated commands. These evolutions aim to protect the rights of artists, enhance music quality, promote creativity and other possibilities like the blockchain technologies.

[0020] According to one aspect, the invention is also related to a method for displaying an interactive 2D or 3D sound map where sound sequences are represented with a point so that the distance between said points is inversely proportional to the similarity between the associated sound sequences, said distance being user dependent.

[0021] The distance is user-dependent, since the similarities between sound sequences is subjective; a first users may find two sound sequences to be very similar while another user may find those to be quite different.

[0022] The 2D or 3D sound map is thus user dependent. The collection of sound sequences in the sound map may be user dependent. The distance between sound sequences on the map may be user dependent.

[0023] According to one aspect, the invention is also related to a method for entering as prompt at least one initial sound sequence to a generative Al system, generating with this system a new sound sequence, and including in this new sound sequence a watermark with an identifier of the initial sound sequence(s).

[0024] The watermark may be directly generated by the generative Al system, for example based on an identifier watermarked or otherwise associated with the initial sound sequence.

[0025] Alternatively, the watermark may be added to the sound sequence after its generation by the generative Al system, for example based on an identifier watermarked or otherwise associated with the initial sound sequence.Figures

[0026] Some embodiments of the present invention are illustrated with the following figures:• Figure 1 illustrates a flowchart of a computer-implemented method according to the invention• Figure 2 illustrates an example of 2D sound map.• Figure 3 illustrates a display with an example of 2D sound map.Preferred embodiment

[0027] The system according to the present invention comprises a database which stores a collection of sound sequences. The database may be stored on a user's computer or server, or in the cloud. It may be a private database, i.e. accessible only by the user or by people to whom that user has given access, for example within a family or organisation. It can also be a public database, for example a web database, such as Spotify or Tidal, etc. It can contain a large number of sound sequences, for example more than 50 sound sequences, possible more than 1000 sound sequences, possibly more than 10'000 sound sequences, or more.

[0028] The database may consist of a single relational database, or several databases with relationships between them. Some of the data in the database may be public, i.e. it may be shared by several users without specific authorisation for that data. Some data in the database may be private, i.e. accessible only by one user and other people authorised by that user. Private data can include, for example, notes, classification tags, comments, etc. Private data can be stored in the same database as public data, or in another database, for example as a file on the user's computer or cloud space.

[0029] The sound sequences saved in the database can comprise pieces of music (e.g. songs), sounds, gimmicks, presets, etc. The sequences can be stored in an audio format (e.g. .wav, .jpeg, etc.), or as midi files. Sequences can be stored in audio format (e.g. .wav, .jpeg, etc), or as midi files.

[0030] Some or all sound sequences can be associated with metadata, for example the name of the artist, the title of the sequence, the publication date, classification tags, etc. Some sound sequences can be associated with still images, for example an album cover, or moving images, for example a video clip.

[0031] Some metadata can be present for all users, and introduced for example by the person who loaded the sound sequence into the database. This data may include, for example, the name of the artist, album, title, genre classification, lyrics, etc.

[0032] Some metadata can be introduced by the user, for example text tags, notes, classifications, etc. Some metadata may depend on the user's interactions with the sound sequence, e.g. number of listens, number of recent listens, number of downloads, number of times the sequence has been mixed with another, etc.

[0033] Some metadata can be determined by a computer program that has access to the database. This automatic determined metadata mayinclude, for example, an automatic classification according to musical style. This classification can be carried out by a classifier, for example based on an audio analysis of the sound sequence and / or on a self-learning system, for example a neural network.

[0034] The determined metadata may also include, for example, the energy of the song, the energy in different frequency ranges, the subjective loudness in predefined frequency ranges, the panning balance, the peak level, a music scale analysis, the tonality, a detection of discontinuities in the song; and / or a monaurality-compatibility. The determined metadata may include the language of the lyrics.

[0035] The subjective loudness may be measured in LUFS (loudness units relative to full scale) or any other unit that indicate the subjective, i.e., perceived loudness of a monochannel or multichannel signal.

[0036] The various metadata gathered in this way can form a multidimensional vector (i.e., a n-tuple) associated with each sound sequence. Examples of two such vectors A, B are illustrated on Figure 2 in a two-dimensional space. The two vectors A and B have similar valued along the Y axis but very different values along the X axis. A third vector C is closer to B than two A, indicating that the corresponding sound sequence C is closer to the sound sequence B than two in the X, Y plane.

[0037] Sound sequences are preferably associated with n-dimensional vectors with n larger than 2.

[0038] At least some sound sequences in the database may include or be associated with watermarks, for example an encoding of digital information in the time and / or frequency domain. The watermark may be used to ensure traceability of the sound sequence. The watermark may include, for example, a unique identifier of the sound sequence. The watermark may include a copyright indication.

[0039] The system may include a module for automatically adding watermarks to sound sequences in the database. The system may include a module for editing and / or displaying watermarks to sound sequences in the database.

[0040] The vector associated with each sound sequence can be userdependent, if it contains components specific to each user.

[0041] The method of the invention may include a step 100 of retrieving some sound sequences. Retrieving may include downloading, uploading, copying, transferring etc.

[0042] The user can filter among the collection of sound sequences available in the collection, in order to limit the number of retrieved sound sequences, or to retain only sound sequences corresponding to user- defined criteria. Filtering can be carried out using classification tags, for example according to the type of sequences, the style of music, the notes of the user or a community, according to duration, quality, the name of the artist, the name of the album, the title of the song, the date, and so on. Filtering can correspond to the selection of a collection of sound sequences assembled by the user, or made available by another user.

[0043] The vectors are assigned to each sound sequence such that a distance d between two vectors is inversely proportional to the similitude between the corresponding sound sequences.

[0044] The method of the invention may include a step 101 of calculating this distance d between pairs of sound sequences from the collection of sound sequences. This distance calculation can be carried out using a distance calculation software module. This distance can be calculated, for example, as a distance between vectors associated with each sound sequence, or as a weighted distance between vectors associated with each sound sequence. Certain components of the vector can be ignored inthe calculation of this distance. Some vector components can be weighted differently to other components.

[0045] The assignment of vectors to sound sequences such that a distance d between two vectors is inversely proportional to the similitude between the corresponding sound sequences is also possible without explicitly determining the distance d.

[0046] A self-learning classifier system, for example a neural network based classifier system, may be used for the assignment of vectors to sound sequence such that this distance d between two vectors is inversely proportional to the similitude between the corresponding sound sequences.

[0047] The classifier system may be trained based on previous interactions of the user, and progressively learn which sound sequences will be considered to be close from each other by each user.

[0048] The self-learning classifier system, for example a neural network based classifier system, may compute this distance d between two sound sequences, or some components of the vector associated with each sound sequence.

[0049] The method of the invention may include a step of reducing the dimensionality of said vectors, using a dimension reduction algorithm. To this end, an algorithm can be implemented which allows each multidimensional vector to be modelled by a two- or three-dimensional vector, so that similar sound sequences are modelled by close 2D or 3D vectors with a high probability. The non-linear dimensionality reduction technique may use a t-distributed stochastic neighbor embedding (t-SNE), or a Uniform Manifold Approximation and Projection (UMAP).

[0050] The method of the invention can include a step 102 of displaying said vectors on a 2D or 3D sound map 4, which thus makes it possible to represent the collection of sound sequences in 2 or 3 dimensions. Similar sound sequences are thus represented by similar points on this sound map.

[0051] The colour of each point can represent an additional dimension.

[0052] The different sound sequences on the map can be marked with an icon, an image (such as an image corresponding to the cover album), an image generated on the fly by a generative Al depending on the content of the sound sequence, and / or with a text, such as a contextual window.

[0053] The 2D or 3D sound map is user-dependent, since the distance between each pair of points depends on user criteria. Thus, two users who wish to display a sound map corresponding to a same collection of sound sequences will have a different representation.

[0054] The 2D sound map 4 can be displayed on a screen, for example on the user's computer or smartphone. An example of representation of such a 2D sound map 4 is illustrated on Figure 3. The sound sequences A and B may be represented with points at position corresponding to the extremity of their vectors pointing from the origin.

[0055] A 2D or 3D sound map can be displayed with augmented or virtual reality glasses, and / or in a metaverse displayed with such glasses or on a display.

[0056] The sound map may be zoomable. The sound map may be scrollable.

[0057] The sound map is interactive. The user can hover a cursor over the interactive sound map, or click on the map, to listen to the sound sequences, or extracts from those sound sequences.

[0058] Users can select a sound sequence directly from the sound map, and load it into sound design software or DAWs (Digital Audio Workstations) to combine it with other sound sequences, arrange it or apply presets.

[0059] Preferably, the user can modify the metadata of the sound sequences represented on the sound map directly from this sound map. For example, they can add tags, notes, etc with a right-click or other interaction with a sound sequence selected on the sound map. Interactions with a sound sequence, for example when he listens to it, downloads it, transfers it to his DAW, and / or when he combines or mixes it with another sound sequence etc, can also modify the metadata associated with this sound sequence.

[0060] Modifying the metadata associated with a sound sequence automatically moves the point associated with that sound sequence, according to the new distances from other points.

[0061] A computer module allows the user to move a sound sequence on the sound map, for example by drag and drop, in order to position it closer to other sound sequences that the user considers to be close according to subjective criteria. This movement can cause an automatic update of the metadata associated with the moved sound sequence.

[0062] A computer module allows the user to import a new sound sequence 40, 41, etc into a collection or list, by placing it by drag and drop at a location C they deem appropriate on the sound map 4, close to other sound sequences they deem close according to their subjective criteria. This positioning can automatically update the metadata associated with the moved sound sequence.

[0063] A computer module allows the user to delete a new sound sequence directly from the sound map, close to other sound sequences that they judge to be close according to their subjective criteria. This positioningcan cause an automatic update of the metadata associated with the displaced sound sequence.

[0064] The user preferably has the option of modifying the sound map displayed by selecting proximity criteria to include and / or exclude. For example, a user may choose to use only one component of the vector associated with each sound sequence (e.g. sequence style, BPMs, etc), or a limited number of criteria, to determine distance.

[0065] The user can preferably modify the sound map 4 displayed by filtering the sound sequences displayed directly from the sound map, according to one or more criteria. For example, a user can choose to display on the sound map only those sequences corresponding to a given style, BPMs, etc.

[0066] A computer module allows the user to trigger the automatic generation 104 of a new sound sequence, by selecting during step 103 an appropriate zone (location) on the sound map, and / or by selecting several existing sound sequences from this sound map. These selected sound sequences, or sound sequences close to the selected location, are used as input by a generative artificial intelligence system, and thus form a kind of prompt for this generative Al system.

[0067] A user can also use additional prompts to the generative Al system, including text prompts, colours, still images and / or animated images. These additional prompts may be entered by the user and / or combined with the selected location on the sound map. For example, a user may enter a zone (location) on the sound map, and simultaneously upload an image corresponding to the style of sound sequence he wants to generate; both inputs will be used as a prompt to the generative Al system for triggering the generation of a new sound sequence.

[0068] The generative artificial intelligence system can use a latent space, i.e. a compressed representation of the sound sequences andassociated metadata. The latent space representation may be based on the same data dimensionality reduction method used for generating the sound map, or on another dimensionality reduction method.

[0069] The generative artificial intelligence system can use a foundation model, i.e., a large machine learning model trained on a vast quantity of data, preferably including sound sequences data. Existing foundation models, such as Transformers, BERT, RoBERTa, variants of GPT, and so on. This foundation model may be adapted to the task of generating sound sequences based on different types of prompts.

[0070] The generative artificial intelligence system generates during step 104 a new sound sequence on the basis of this input. The system can use the input sound sequence(s), features extracted from these input sound sequences and / or metadata associated with these sound sequences to generate one or more new sound sequences close to the input sound sequences.

[0071] The sound sequences generated in this way can be listened to by the user. They can be automatically positioned on the sound map at the chosen location. Metadata can be associated with this new sound sequence, for example metadata entered by the user or calculated automatically. The introduction or automatic determination of this metadata can influence the position of the new sound sequence on the 2D or 3D sound map.

[0072] Generating a new sound sequence using this method can also have an impact on the positioning of existing sound sequences on the 2D or 3D sound map. For example, selecting two sound sequences to generate a new one may imply that the user considers that these sound sequences are close, and therefore tends to bring them closer together.

[0073] The watermark associated with the sound sequences input to the generative artificial intelligence system are advantageously directly included in the sound sequences generated by this Al system. Sufficientlyrobust watermarks may be chosen for this purpose. The neural network at the heart of the generative artificial intelligence system can be trained specifically to generate sound sequences that are different from the input sound sequences, but which include the same watermark.

[0074] It is also possible to read and save the information associated with the watermarks of the input sound sequences, and then integrate this information into a watermark integrated into the sound sequences generated from these inputs.

[0075] A watermark is preferably automatically added to the sound sequences generated by the artificial intelligence system, in order to include, for example, a unique identifier and / or copyright indications, and thus guarantee the traceability of the sequences generated in this way.

[0076] A computer module calculates a hash of the sound sequence thus generated and, optionally, of the associated metadata. This hash is preferably stored in a blockchain to prove the creation date of the sound sequence.

[0077] The invention is also related to a computer program product storing a computer program arranged for performing the steps of one of the preceding claims when executed by a processor.Additional Features and Terminology

[0078] Many other variations than those described herein will be apparent from this disclosure. For example, depending on the embodiment, certain acts, events, or functions of any of the algorithms described herein can be performed in a different sequence, can be added, merged, or left out altogether (for example, not all described acts or events are necessary for the practice of the algorithms). Moreover, in certain embodiments, acts or events can be performed concurrently, for instance, through multi-threaded processing, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially. In addition, different tasks or processes can be performed by different machines or computing systems that can function together.

[0079] Unless otherwise specified, the various illustrative logical blocks, modules, engines and algorithm steps described herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, engines, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. The described functionality can be implemented in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the disclosure.

[0080] Unless otherwise specified, the various illustrative logical blocks engines and modules described in connection with the embodiments disclosed herein can be implemented or performed by a machine, a microprocessor, a computer, a smartphone, a tablet, a server, a plurality of computers and / or servers connected through a network, for example in a LAN or in a cloud, a state machine, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A hardware processor can include electrical circuitry or digital logic circuitry configured to process computer-executable instructions. In another embodiment, a processor includes an FPGA or other programmable device that performs logic operations without processing computer-executable instructions. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other suchconfiguration. A computing environment can include any type of computer system, including, but not limited to, a computer system based on a microprocessor, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a computational engine within an appliance, to name a few.

[0081] Unless otherwise specified, the steps of a method, process, or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module stored in one or more memory devices and executed by one or more processors, or in a combination of the two. A software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of non-transitory computer-readable storage medium, media, or physical computer storage known in the art. An example storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The storage medium can be volatile or nonvolatile. The processor and the storage medium can reside in an ASIC.

[0082] Conditional language used herein, such as, among others, "can," "might," "may," "e.g.," and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements or states. Thus, such conditional language is not generally intended to imply that features, elements or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements or states are included or are to be performed in any particular embodiment. The terms "comprising," "including," "having," and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term "or" is used in its inclusive sense (and not in its exclusive sense) so that when used,for example, to connect a list of elements, the term "or" means one, some, or all the elements in the list. Further, the term "each," as used herein, in addition to having its ordinary meaning, can mean any subset of a set of elements to which the term "each" is applied.

Claims

Claims1. A method for generating sound sequences, comprising the steps of: retrieving a plurality of sound sequences (100); displaying (102) an interactive 2D or 3D sound map (4) where said sound sequences are represented with a point (A, B) so that a distance (d) between said points is inversely proportional to a similarity between the associated sound sequences; selecting (103) a zone (20) on said interactive sound map to trigger generation of a new sound sequence (104) with a generative artificial intelligence system, so that the new sound sequence shares similarities with all the neighbour sound sequences.

2. The method of claim 1, comprising a step (101) of computing a distance between sound sequences.

3. The method of one of the claims 1 or 2, wherein a n-dimensional vector is assigned to each sound sequence, n being an integer larger than 2, wherein a similarity between two sound sequences depends on the distance (d) between their corresponding vectors, wherein said step of displaying an interactive sound map includes a nonlinear dimensionality reduction technique for modelling each n- dimensional vector by a two or three dimensional vector in such a way that similar sound sequences are modeled by nearby points on said map and dissimilar sequences are modeled by distant points on said map with high probability.

4. The method of claim 3, said non-linear dimensionality reduction technique using a t-distributed stochastic neighbor embedding (t-SNE) or a Uniform Manifold Approximation and Projection (UMAP).

5. The method of one of the claims 1 to 4, wherein said distance (d) depends on user-dependent subjective criteria.

6. The method of one of the claims 1 to 5, wherein said distance (d) depends on user-dependent subjective criteria depending on the user interaction with said sound map.

7. The method of one of the claims 1 to 6, comprising the steps of: hovering a cursor over said interactive sound map; playing the sound sequences corresponding to the points of said interactive sound map.

8. The method of one of the claims 1 to 7, comprising the steps of: entering user feedback regarding a subjective similarity between sound sequences; reorganizing said interactive sound map (4) in response to said user feedback.

9. The method of one of the claims 1 to 8, comprising a step of using a neural network for determining a subjective distance (d) between sound sequences.

10. The method of one of the claims 1 to 9, comprising a step of training said generative artificial intelligence system with watermarked sound sequences, and retrieving said watermarks in sound sequences generated by said generative artificial intelligence system.

11. The method of one of the claims 1 to 10, comprising a step of mixing one said sound sequence.

12. The method of one of the claims 1 to 11, comprising a step of computing a hash of one said sound sequence, and storing said hash in a blockchain.

13. A computer program product storing a computer program arranged for performing the steps of one of the preceding claims when executed by a processor.