Circuitry, electronic device and method

WO2026202223A1PCT designated stage Publication Date: 2026-10-01SONY GROUP CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/058700
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-26
Publication Date
2026-10-01

Smart Images

  • Figure EP2026058700_01102026_PF_FP_ABST
    Figure EP2026058700_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a circuitry configured to acquire captured audio data corresponding to environmental music, to recognize a song based on the captured audio data, to determine a current song location of the environmental music by performing a machine learning model, based on an input of the captured audio data and song audio data corresponding to the recognized song, the machine learning model being configured to output the current song location, and to output current lyrics corresponding to recognized song and being synchronized with the music output by the speaker.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CIRCUITRY, ELECTRONIC DEVICE AND METHOD TECHNICAL FIELD

[0002] The present disclosure generally pertains to the field of providing musical information, in particular to a circuitry, an electronic device and a method.

[0003] TECHNICAL BACKGROUND

[0004] Wearable devices such as smart glasses and head-mounted displays are increasingly being used to enhance user interactions and provide immersive experiences.

[0005] In the realm of music and entertainment, there has been a growing demand for systems that can provide real-time musical information such as display of lyrics for music songs being played. Traditional methods of lyric display often require direct connection to the audio source, limiting the flexibility and convenience for users. Moreover, these systems typically rely on pre-loaded lyrics and do not dynamically adapt to live music playback.

[0006] Although there exist techniques for providing musical information such as lyrics, it is generally desirable to improve on existing techniques.

[0007] SUMMARY

[0008] According to a first aspect, the present disclosure provides a circuitry configured to acquire captured audio data corresponding to environmental music, recognize a song based on the captured audio data, determine a current song location of the environmental music by performing a machine learning model, based on an input of the captured audio data and song audio data corresponding to the recognized song, the machine learning model being configured to output the current song location, and output current lyrics corresponding to recognized song and being synchronized with the music output by the speaker.

[0009] According to a second aspect, the present disclosure provides an electronic device comprising the circuity of the first aspect and configured to capture the audio data corresponding to the environmental music and to display the current lyrics on a display unit.

[0010] According to a third aspect, the present disclosure provides a method comprising the steps of acquiring captured audio data corresponding to environmental music, recognizing a song based on the captured audio data, determining a current song location of the environmental music by performing a machine learning model, based on an input of the captured audio data and songaudio data corresponding to the recognized song, the machine learning model being configured to output the current song location, and outputting current lyrics corresponding to the recognized song and being synchronized with the music output by the speaker.

[0011] Further aspects are set forth in the dependent claims, the drawings and the following description.

[0012] BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Embodiments are explained by way of example with respect to the accompanying drawings, in which:

[0014] Fig. 1 illustrates an electronic device wearable by a user;

[0015] Fig. 2 illustrates a processor configuration for outputting synchronized lyrics;

[0016] Fig. 3 illustrates an example of a song location determination;

[0017] Fig. 4 illustrates an example of a machine learning model for determining a current song location;

[0018] Fig. 5 illustrates a further example of the processor configuration, including lyric generation and music instruction generation;

[0019] Fig. 6 illustrates the lyric generation using a neural network model;

[0020] Fig. 7 illustrates a music instruction generation using a further neural network model;

[0021] Fig. 8 illustrates a method for displaying synchronized lyrics;

[0022] Fig. 9 illustrates a method for determining a current song location; and

[0023] Fig. 1 Oillustrates a configuration of the electronic device.

[0024] DETAILED DESCRIPTION OF EMBODIMENTS

[0025] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.

[0026] In the following, it is explained how wearable technology and artificial intelligence (Al) may by leveraged to improve on providing lyrics synchronized with music playback.

[0027] In this way, it is possible to enhance a concert experience by allowing attendees to sing along with live performances, providing a hands-free karaoke experience in karaoke bars by displaying lyrics directly on smart glasses, and transforming parties and gatherings into karaoke sessions, making it easy for guests to join in the fun. Additionally, it assists musicians and singers inpracticing their songs by displaying lyrics and helping them stay in sync with the music, and aids users in learning new languages by displaying lyrics in both the original language and the user’s native language, fostering improved pronunciation and comprehension.

[0028] Therefore, some embodiments are directed to a circuitry configured to acquire captured audio data corresponding to environmental music, to recognize a song based on the captured audio data, to determine a current song location of the environmental music by performing a machine learning model, based on an input of the captured audio data and song audio data corresponding to the recognized song, the machine learning model being configured to output the current song location, and to output current lyrics corresponding to recognized song and being synchronized with the music output by the speaker.

[0029] The circuitry may include a processor, a memory (RAM, ROM or the like), a storage, input means (mouse, keyboard, camera, etc.), output means (display (e.g., liquid crystal, (organic) light emitting diode, etc.), loudspeakers, etc.) and a (wireless) interface, etc., as it is generally known for electronic devices (computers, smartphones or the like). Moreover, it may include sensors for sensing still image or video image data (image sensor, camera sensor, video sensor, etc.), and for sensing sound (microphones).

[0030] The environmental music refers to music that is being played or performed. For example, the music may be output from a speaker. Alternatively or additionally, the music may be performed live by musicians, such as a singer, a band, or an orchestra without being output by a speaker. The environmental music is captured by a microphone which is configured to output an audio signal corresponding to the environmental music. This audio signal is then converted into the captured audio data. In some examples, the microphone is a digital microphone which is configured to convert audio signals into audio data. In other examples, the circuitry may be configured to receive audio signals and to convert them into audio data.

[0031] The environmental music may be captured for a predetermined time such that the captured audio data represents only a snippet (brief segment) of the environmental music. Alternatively, the environmental music may be captured continuously.

[0032] The circuitry is configured to recognize or identify a song from the captured audio data.

[0033] Techniques for recognizing a song based on a captured audio data are known to the skilled person and include, e.g., audio fingerprinting which is used by services like Shazam.

[0034] When recognizing a song, a corresponding song ID may be determined or retrieved.The circuitry uses a machine learning model to determine the current song location which refers to a current (time) position or timestamp within the song as it is currently being played or performed in the environment (in other words, the current song location within the current environmental music).

[0035] By performing or applying the machine learning model, the captured audio data and the song audio data are used as inputs for the model and the current song location is the resulting output. The song audio data includes sound information of the entire recognized song. That is, the sound audio data represents the entire song. The song audio data is typically stored as digital samples that represent the sound wave at various points in time. In other words, the song audio data is time-series data. The song audio data can be stored in formats such as MP3, WAV, FLAC, etc. The song audio data can be retrieved, using the song ID, from online music databases, streaming services, local storage, or cloud storage.

[0036] To determine the current song location within the environmental music based on the captured audio data and the song audio data, several machine learning models can be employed. Using such machine learning models improves the accuracy of the timestamp identification

[0037] For example, a Convolutional Neural Networks (CNNs) may be used which analyzes the audio data by extracting features from spectrograms or other representations of the audio data (or audio signal). These features help identify patterns and match segments of the captured audio to the corresponding segments in the song audio data. Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) networks may be used, which are suited for analyzing sequential data like audio signals. These models learn temporal dependencies and patterns in the audio data, allowing them to accurately determine the current position within the song. Further, transformer models can also be applied to audio data. They can process long sequences of audio data and capture relationships between different parts of the audio signal to determine the current song location. Furthermore, custom neural network architectures designed for sequence alignment can be used to match the captured audio data with the song audio data, providing an accurate determination of the current song location.

[0038] These machine learning models are trained on large datasets of audio data to learn the patterns and features that correspond to different positions within songs, enabling them to perform realtime processing and deliver accurate results.Based on the current song location, the current lyrics are output. The current lyrics may be taken from the original lyrics of the recognized song. Alternatively, the current lyrics may be taken from lyrics suggestions which are generated / created based on the song audio data.

[0039] The current lyrics are output in synchronization with the environmental music. This means that the lyrics displayed (on a display unit) are timed to match the music being played or performed in the environment. The synchronization ensures that the lyrics correspond precisely to the part of the song that is currently being heard by a user, allowing the user to follow along with the song in real-time.

[0040] Thus, the user can sing along, read the lyrics, or understand the song better without any lag or mismatch between the audio and the text. This precise timing is particularly useful in settings such as concerts, karaoke sessions, and language learning.

[0041] In conclusion, the circuity provides more accurate synchronization of lyrics with the environmental music.

[0042] In some embodiments, the circuitry may be further configured to determine the current lyrics based on the current long location and lyrics information corresponding to the recognized song. The lyrics information may include time-synced data regarding the lyrics. In this data, the lyrics are pre-associated with specific timestamps within the song.

[0043] The current song location, which indicates the timestamp (i.e., position within the song), may be used with time-synced lyrics data to determine the current lyrics by matching the timestamp of the song location with the corresponding timestamp in the time-synced lyrics data. The circuitry may be configured to extract the lyrics associated with that specific timestamp from the time-synced lyrics data.

[0044] In some embodiments, the circuitry may be further configured to retrieve time-synced (original) lyrics data as the lyrics information. Here, the time-synced (original) lyrics correspond to the original lyrics of the recognized song.

[0045] Time-synced lyrics data can be retrieved from various sources, such as online music databases and lyric services. Platforms like Musixmatch, LyricFind, and Genius offer extensive collections of lyrics, with Musixmatch and LyricFind providing APIs that allow for easy integration into applications. Further, streaming services (e.g., Spotify, Apple Music, and Amazon Music) are another resource for retrieving time-synced lyrics. Also, local storage on devices or applications can house music libraries with time-synced lyricsIn some embodiments, the circuitry may be further configured to generate time-synced lyrics suggestion data as the lyrics information.

[0046] That is, the circuitry is capable of generating the time-synced lyrics suggestion data as alternative lyrics that are different from the original lyrics of the recognized song. These suggestions are also time-synced, meaning they are associated with specific timestamps within the song, ensuring accurate synchronization with the environmental music being played.

[0047] In some embodiments, the circuitry may be further configured to perform a neural network model, based on an input of the song audio data, the neural network model being configured to output the time-synced lyrics suggestion data.

[0048] The neural network model is used to perform the task of generating time-synced lyrics suggestion data. The neural network model may be an artificial intelligence (Al) model, it may include a convolutional neural network (CNN) or multimodal Al model, e.g., transformer model, or the like.

[0049] The neural network model uses the song audio data as input. Since the song audio data includes sound information of the entire recognized song, providing the neural network model with the necessary context to generate relevant lyrics suggestions.

[0050] Optionally, the neural network model may additionally use the original lyrics, i.e., time-synced (original) lyrics data, the song as part of its input. This allows the model to generate lyrics suggestions that are closely related to the original lyrics.

[0051] In some embodiments, the neural network model may be performed based on a system prompt used as further input, the system prompt including instructions for generating the time-synced lyrics suggestion data .

[0052] The system prompt may be used as an additional input to the neural network model for generating time-synced lyrics suggestion data. The system prompt may specify instructions in regarding various aspects of the lyrics, such as tone, style, mood, genre, and other characteristics. The system prompt may be in text form or include text information.

[0053] The use of the system prompt enhances the flexibility and precision of the neural network model by providing specific guidelines that influence the lyrics generation process.

[0054] In some embodiments, the neural network model may include a Convolutional Autoencoder (CAE) encoder-decoder architecture. A CAE is a type of neural network model used for unsupervised learning.The CAE encoder may receive the song audio data corresponding to the song and, optionally, the time-synced (original) lyrics data as input, apply a series of convolutional transformations, and encode the input into a compressed latent representation while preserving essential features. The CAE encoder may utilize convolutional layers and pooling mechanisms to achieve dimensionality reduction and feature extraction of the song audio data and, optionally, the time-synced (original) lyrics data.

[0055] The corresponding decoder module may be configured to generate, as output, the alternative lyrics, i.e., the time-synced lyrics suggestion data, from latent representation using transposed convolutional or upsampling layers, thereby restoring spatial dimensions and recovering structural details.

[0056] In some embodiments, the circuitry may be further configured to apply an audio encoder to the captured audio data and the song audio data, wherein the audio encoder is configured to extract features from the captured audio data and the song audio data and the machine learning model is performed based on the extracted features as input.

[0057] The audio encoder is used to process the captured audio data and song audio data. Specifically, the audio encoder is configured to extract features (feature representations) from audio data (audio signals). The audio encoder may be an encoder of a CAE encoder-decoder architecture. In some examples, a single audio encoder may be used. When utilizing a single audio encoder for both the captured audio data and the song audio data, the audio encoder may be configured to process the input sequentially or in parallel. In other examples, two audio encoders may be employed, each dedicated to the respective audio data. The first audio encoder may be used for processing the captured audio data, while the second audio encoder may be used for processing the song audio data.

[0058] The machine learning model uses the extracted features from the audio encoder as input. The machine learning model processes these features to determine the current song location within the song played in the environmental music.

[0059] In some embodiments, the circuitry may be further configured to output music instructions based on the song audio data, the music instructions including instructions for playing an instrument according to the environmental music.

[0060] The music instructions may include information such as chord progressions, notes, harmonies, and rhythms. These instructions can be displayed visually on a display unit or throughaugmented reality (AR) overlays on wearable devices like smart glasses. The AR overlays may indicate notes or chords on the user’s instrument, providing real-time guidance on where to place fingers, strum, or press keys.

[0061] In some embodiments, the circuitry may be further configured to perform a further neural network model, based on an input of the song audio data, the further neural network model being configured to output the music instructions.

[0062] The further neural network model may be configured similarly to the neural network model used for generating alternative lyrics. Thus, the above description with respect to the configuration of the neural network model for generating alternative lyrics applies correspondingly to the further neural network.

[0063] The further neural network model differs from the neural network model for generating alternative lyrics by way of the input used and the output generated.

[0064] The further neural network model may be configured to analyze the song audio data as input to understand its structure, melody, rhythm, harmony, and other musical attributes. Using this analysis, the model generates the music instructions.

[0065] The music instructions may guide the user on how to play an instrument in synchronization with the music. For this synchronization, the further neural network model may use, as further input, the current song location previously determined.

[0066] In some embodiments, the circuitry is further configured to acquire a translation of the current lyrics from the original language of the song and to output the translation.

[0067] The translation of the current lyrics may be acquired through various methods and sources. For example, the circuitry may be configured to interface with online translation services such as Google Translate, Microsoft Translator, or other APIs that provide language translation capabilities. These services can translate the lyrics in real-time based on the input text (i.e., the text / lyrics included in time-synced lyrics data or corresponding to the current lyrics), allowing the circuitry to send the original lyrics to the translation API and receive the translated lyrics in the user’s native language.

[0068] Further, the circuity may be configured to access a database that contains pre-stored translations of song lyrics. The circuitry may query this database using the song ID to retrieve the translation of the song lyrics.Furthermore, the circuitry may include local translation modules that are part of the device’s software. These modules may translate lyrics based on built-in dictionaries and language rules, allowing the circuitry to use an onboard language translation module to translate the lyrics. Also furthermore, the circuitry may be configured to perform a machine learning model trained for language translation. This model may process the original lyrics and generate translations internally without relying on external services. This machine learning model processes the original lyrics and outputs translated lyrics based on its training data.

[0069] The current lyrics may be displayed simultaneously in both the original language and the user’s native language on a display unit. This dual-language display aids the user in learning new languages by fostering improved pronunciation and comprehension.

[0070] Some embodiments pertain to an electronic device including the circuitry explained in this specification. The electronic device may include any feature of the circuitry discussed in this specification.

[0071] The electronic device may be further configured to capture the audio data corresponding to the environmental music and to display the current lyrics on a display unit.

[0072] Alternatively, the electronic device may acquire the captured audio data, for example from a connected device which may include at least one microphone for capturing the audio data corresponding to the environmental music.

[0073] Further, the electronic device may acquire the captured audio data from or via a storage (e.g., a database, cloud server, etc.).

[0074] Some embodiments pertain to a method comprising the step of acquiring captured audio data corresponding to environmental music, recognizing a song based on the captured audio data, determining a current song location of the environmental music by performing a machine learning model, based on an input of the captured audio data and song audio data corresponding to the recognized song, the machine learning model being configured to output the current song location, and outputting current lyrics corresponding to the recognized song and being synchronized with the music output by the speaker.

[0075] The method may include any feature of the of the circuitry and / or the electronic device discussed in this specification.

[0076] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried outon the computer and / or processor. In some embodiments, also a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.

[0077] Referring to Fig. 1, an electronic device 1 wearable by a user U are shown. In this example, the electronic device 1 is configured as smart glasses (or a head mounted display).

[0078] The electronic device 1 comprises a display unit 2 which is configured to display content. In this example, the display unit 2 is semi-transparent.

[0079] Further, the electronic device 1 comprise a microphone (not shown) which is configured to capture music output by a speaker 3 (exemplary environmental music). The music corresponds to a song.

[0080] The electronic device 1 is configured to display, on its display unit 2, lyrics corresponding to and synchronized with the music in real-time, without being connected to the audio source. To this end, the electronic device 1 is configured as shown in Fig. 10 including a processor (CPU) 1001. The processor 1001 will now be described with reference to Fig. 2 which illustrates components / units and data flow involved in the process of displaying synchronized lyrics on smart glasses.

[0081] A song recognition 100 is configured to receive captured audio data as captured by the microphone of the electronic device 1. The captured audio data 10 corresponds to the environmental music. The song recognition 100 is configured to determine a song ID 12 based on the captured audio data 10. The song ID 12 includes information regarding a name of the song and the artis. Methods for performing the song recognition 100 are known to the skilled person.

[0082] A song acquisition 200 is configured to acquire or retrieve song audio data 14 corresponding to the song ID. For example, the song acquisition 200 may be configured to access online music databases or streaming services such as Spotify, Apple Music, Deezer, or other music libraries. Using the Song ID 12, the song acquisition 200 may be configured to query these services to retrieve the corresponding song audio data 14. Additionally or alternatively, the electronic device 1 may have a local storage with a music library and the song acquisition 100 may be configured to search this storage to find and retrieve the song audio data 14 using the Song ID. Additionally or alternatively, the electronic device 1 may be configured to access cloud storageservices where the user U might store their music collections. The song acquisition 200 may be configured to request the song audio data 14 from the cloud using APIs and the song ID 12. A song location determination 300 is configured to determine a current song location 16 in the recognized song by analyzing the song audio data 14 and to output the current song location 16. A lyrics retrieval 400 is configured to retrieve time-synced (original) lyrics data 18 based on the song ID 12. The time-synced lyrics data 18 includes lyrics that correspond to song according to the song ID 12 and are associated with specific timestamps in the song. The lyrics retrieval 400 may be configured to access a database / service like SyncedLyrics, Musixmatch, and LyricFind for retrieving the time-synced lyrics data 18.

[0083] A lyric extraction 500 is configured to determine the current lyrics 20 of the environmental music based on the time-synced lyrics data 18 and the current song location 16. That is, the lyric extraction 500 is configured to extract, from the time-synced lyrics data 18, the current lyrics 20 which is synchronized with the environmental music.

[0084] The lyric extraction 500 is configured to use the current song location 16 to find a corresponding timestamp in the time-synced lyrics data 18. The lyric extraction 500 is configured to look for the portion of the lyrics that matches the current timestamp.

[0085] Once the matching timestamp is found, the lyric extraction 500 is configured to extract the portion of the lyrics that corresponds to the current song location 16. This might involve extracting a single line, multiple lines, or even individual words.

[0086] The current lyrics 20 are displayed on the display unit 2. This might occur in a karaoke-like fashion, where the electronic device 1 continuously display the current lyrics in synchronization the song within the environmental music. As the song progresses, the lyric extraction 500 is configured to dynamically update the current lyrics 20 to be displayed based on the current song location 16. The current lyrics 20 may be highlighted, using effects such as color changes, underlining, or bold text to emphasize, on the display unit 2, the words currently being sung. This visual cue helps the user U stay on track with the song.

[0087] Fig. 3 shows an example of the song location determination 300. The song location determination 300 includes a machine learning model 301 for determining the current song location 16.

[0088] The captured audio data 10 and the song audio data 14 are fed into the machine learning model 301.The machine learning model 301 is configured to output (predict), based on both data inputs, the current song location 16. For example, the machine learning model 301 may use algorithm to identifies patterns and matches segments of the captured audio data 10 with the corresponding segments of the song audio data 14.

[0089] The machine learning model 301 outputs the current song location 16, which indicates the precise timestamp or position in the song currently output by the speaker 3.

[0090] Fig. 4 shows an example of the machine learning model 301. The machine learning model 301 includes a song audio segmentation 301, a captured audio segmentation 302, a first CAE encoder 303, a second CAE encoder 304, a similarity matching 305 and a location determination 306. The song audio data 14 and the captured audio data 10 are processed by the song audio segmentation 301 and the captured audio segmentation 302, respectively. The song audio data segmentation 301 is configured to divide the song audio data 14 into a plurality of overlapping song audio segments 31. The overlap of two consecutive segments may be, for example, 0,5 seconds. Further, a start time (within the song) for each segment (e.g., 0s, 0,5s, Is, 1.5s, etc.) is also included (stored / associated) with the respective segment. A segment can be, e.g., 0,5 seconds to 5 seconds long. The captured audio data segmentation 302 is configured to extract only the latest segment (captured audio segment 32) of the captured audio data 10, e.g., the last 0,5 seconds to 5 seconds of the captured audio data 10. The segment length of the song audio segments 31 and the captured audio segment 32 may be the same. It is possible to simultaneously operate on different segment lengths, but only the latest segment of the captured audio data 10 is required for further processing.

[0091] The first and second CAE encoder 303, 304 are configured to process the song audio segments 31 and the captured audio segment 32, respectively. The first CAE encoder 303 is configured to extract, on a segment-by-segment basis, features from the song audio segments 31 and to compress these features into a latent representation, resulting in song audio representations 33. Each segment is compressed into a respective representation. Key features to be extracted include, e.g., spectral patterns, amplitude variations, temporal dynamics, etc.

[0092] Similarly, the second CAE encoder 304 is configured to extract features from the captured audio segment 32 and to compress these features into a latent representation, resulting in a captured audio representation 34.The representations 33, 34 also includes information regarding the start time for the respective segments.

[0093] The similarity matching 305 is configured to use the song audio representations 33 and the captured audio representation 34 as input and to identify, based on the input, a segment from the song audio segments 31 that has the highest similarity with the captured audio segment 32. The similarity matching 305 may use distance metrics to measure a similarity between latent representations, i.e., the song audio representation 33 of a respective song audio segment 31 and the captured audio representation 34 of the captured audio segment 32.

[0094] The similarity matching 305 is configured to output the identified song audio segment 35 that has the highest similarity.

[0095] As mentioned before, each song audio segment 35 includes information about the start time (within the song). The song location extraction 306 is configured to extract, from the identified song audio segment 35, the start time associated with the identified song audio segment 35 and to determine, based on the start time, the current song location within the song in the environmental music as currently heard by the user U.

[0096] In some examples, a delay or offset may have to be considered which occurs due to the time needed to capture the environmental music, process, and identify the song audio segment 35 with the highest similarity. To account for this, a processing delay time for capturing the environmental music, processing the segments, and identifying the best match is measured. The song location extraction 306 is then configured to determine the current song location 16 based on the start time of the identified song audio segment 35 and the processing delay time. In some examples, the processing delay time may be added to the start time. The adjusted start time will reflect the current song location of the song within the environmental music as currently heard by the user U.

[0097] The first and second CAE encoders 303, 304 are part of respective CAE encoder-decoder architectures. Generally, a CAE consists of an encoder and a decoder. The encoder is responsible for compressing input audio features into a lower-dimensional latent representation. It typically includes convolutional layers that extract meaningful features from audio data, i.e., the song segments 31 or captured audio segment 32. The decoder reconstructs the audio features from the latent representation, ensuring that the compressed data retains essential information. However, since the decoder is not required for determining the current song location 16, it is not shown in the machine learning model 301 of Fig. 4.Still, the decoder is required during training of a CAE. The training process involves both the encoder and the decoder to ensure that the latent representations learned by the encoder retain the essential features of the input data, which can be accurately reconstructed by the decoder.

[0098] Training data comprises segmented audio features of audio data. A diverse dataset of segmented audio features from various songs and environmental recordings is created / provided. This diverse dataset helps the CAE to learn generalized representations that can be applied to different audio inputs, enhancing its robustness and accuracy. The CAE is trained using a loss function like Mean Squared Error (MSE), which measures the difference between the original audio features (representations) generated by the CAE encoder and the reconstructed features output by the decoder. The goal is to minimize this error during training. Optimization algorithms such as Adam or Stochastic Gradient Descent (SGD) are applied to update the weights of the CAE based on the loss function. The CAE is trained for several epochs, iterating over the dataset multiple times to improve the model’s accuracy.

[0099] The first and second CAE encoder 303, 304 are the result of respective trainings. In some examples, the first and second CAE encoder 303, 304 are trained based on the same training data. In other examples, the first CAE encoder 303 may use a first set of training data specifically tailored for generating latent representations of song audio data and the second CAE encoder 304 a second set of training data specifically tailored for generating latent representations of captured audio data.

[0100] Fig. 5 illustrates a further example of processor 1001. The configuration of the processor 1001 shown in Fig. 5 is the same as shown in Fig. 2 but additionally includes a lyric generation 700 and a music instruction generation 800. In some examples (now shown), only one of the lyric generation 700 and the music instruction generation 800 may be present.

[0101] The lyric generation 700 is configured to generate time-synced lyrics suggestion data 22 based on the song audio data 14 and the time-synced lyrics data 18. Specifically, the lyric generation 700 is configured to receive the song audio data 14 (as indicated by arrows connected via reference sign “a”) and, optionally, the time-synced lyrics data 18 and to use these data as a basis to generate alternative lyrics suggestions (included in the time-synced lyrics suggestion data 22) for the song identified in the environmental music.

[0102] The music instruction generation 800 is configured to generate music instructions 24 in real-time (i.e., instrumental suggestions such as chords, notes, harmonies, etc.) based on the song audio data 14 (as indicated by the arrows connected via reference sign “a”). The electronic device 1may be configured to display, based on the music instructions 24, augmented reality (AR) overlays on the display unit 2 to visually indicate notes or chords on the user’s instrument. Fig. 6 illustrates an example of the lyric generation 700 for generating the time-synced lyrics suggestion data 22 based on the song audio data 14 (and, optionally, the time-synced lyrics data 18) and a system prompt 71. In this example, the lyric generation 700 is a neural network model that includes a CAE encoder-decoder architecture with an CAE encoder 701 and a CAE decoder 708, a text encoder 706, and a Diffusion UNet Latent Diffusion Model (LDM) 703 to process inputs and generate, as output, the time-synced lyrics suggestion data 22.

[0103] The song audio data 14 and, optionally, the time-synced lyrics data 18 corresponding to the song output by the speaker 3 are provided as input data for the CAE Encoder 701. The CAE Encoder 701 is configured to the extract features from the song audio data 14 and the time-synced lyrics data 18 and to compress these features into a latent representation.

[0104] Further, the system prompt 71 is provided as input into the text encoder 706, where the text encoder 706 is configured to extract text features from the system prompt 71. The output of text encoder 706 is synchronized at combination node 707 with a time signal t.

[0105] The (lyric) system prompt 71 corresponds to a set of rules, guidelines, or parameters for the lyrics generation 700. It can define tone, scope, and limitations. For example, the system prompt 71 can include information that provides background and structural details about the music, for example metadata about music, such as genre, key signature, chord progression, tempo time signature, rhythmic patterns, instrumentation, mood and emotion, structure and / or dynamics. Alternatively or additionally, the system prompt 71 may include information regarding the lyrics to be suggested, such as topic of the lyrics, the mood (e.g., sad, happy, moody, etc.), language of the lyrics, etc.

[0106] The output from CAE Encoder 701 is combined at combination node 702 with noise(t) to introduce variability and robustness in the timed-synced lyrics suggestion data 22 to be output. This combined output from node 702 is then fed into an encoder 704 of the Diffusion UNet LDM 17. Still further, the output from text encoder 706 synchronized to the time signal t (i.e., the output from node 707) and fed into both the encoder 704 and a decoder 705 of the Diffusion UNet LDM 703. The time signal t of node 707 is time-synced with the signal representing the noise (t). The time synchronization at node 707 is performed since the Diffusion UNet LDM 703 requires both the encoded song audio data 14 (and, optionally, the time-synced lyrics data 18) (i.e., the output of the CAE Encoder 701) and the encoded system prompt 71 (i.e., the output ofthe text encoder 706) at each denoising step. The time synchronization at node 707 ensures these two inputs are properly combined and injected into the Diffusion UNet LDM 703.

[0107] The Diffusion UNet LDM 703 is a Latent Diffusion Model using a UNet architecture. The role of the Diffusion UNet Latent Diffusion Model (LDM) 700 is configured to refine, enhance and then transform the latent representations provided by the CAE Encoder 701, including the introduction of variability and robustness through the addition of noise. The Diffusion UNet LDM 703 employs a generative modeling technique known as the denoising diffusion process which iteratively transforms noisy data into high-quality, realistic samples. The Diffusion UNet LDM 703 iteratively refines the inputs, leveraging its generative capabilities to produce higher quality latent representations. That is, the refined representations influenced by the features extracted from the system prompt 71 help ensure that the final timed-synced lyrics suggestion data 22 output by the CAE Decoder 708 retains the desired characteristics and style according to the system prompt 71.

[0108] The Diffusion UNet LDM 703 works by encoding the input data into a compressed representation, refining it through a bottleneck layer, and then decoding it to reconstruct the original data. The encoder 704 captures abstract features at multiple scales, while the decoder 705 uses skip connections to retain high-resolution features. The latent diffusion model aspect ensures that the input data is iteratively refined, producing the higher-quality outputs.

[0109] The output of the Diffusion UNet LDM 703 is fed into the CAE Decoder 708 which is configured to generate the final timed-synced lyrics suggestion data 22 from the refined latent representation.

[0110] The time-synced lyrics suggestion data 22 includes (suggested) lyrics that correspond to the song according to the song ID 12 and are associated with specific timestamps in the song.

[0111] The time-synced lyrics suggestion data 22 can be used by the lyrics extraction 500 to provide current (suggested) lyrics in synchronizations with the song present in the environmental music. The training of a CAE encoder-decoder architecture as described with respect to Fig. 4 applies correspondingly to the training of the CAE encoder-decoder architecture 700 shown in Fig. 6. Fig. 7 illustrates an example of the music instruction generation 800 for generating the music instructions 24 based on the song audio data 14, (optionally) the current song location 16, and a system prompt 81. In this example, the music instruction generation 800 is a (further) neural network model with a configuration corresponding to the configuration of the neural networkmodel as shown in Fig 6. That is, the music instruction generation 800 includes a CAE encoderdecoder architecture with a CAE encoder 802 and a CAE decoder 808, a text encoder 806, and a Diffusion UNet Latent Diffusion Model (LDM) 803 to process inputs and generate, as output, the music instructions 24.

[0112] The song audio data 14 and the current song location 16 are provided as input data for the CAE encoder 802. The CAE encoder 802 is configured to extract features from the song audio data 14 and the current song location 16 and to compress these features into a latent representation. Further, the system prompt 81 is provided as input into the text encoder 806, where the system prompt 81 is processed to extract text features. The output of the text encoder 806 is synchronized at combination node 807 with a time signal t.

[0113] The (music) system prompt 81 corresponds to a set of rules, guidelines, or parameters for the music instruction generation 800. It the system prompt 81 may include information that determines the type of music instructions 24 to be output. For example, the system prompt 81 may include an indication of an instrumentation, specifying the types of musical instruments that should be used (e.g., piano, guitar, drums), the skill level of the user U (e.g., beginner, intermediate, expert) to provide music instructions which adapted to the skill level, etc.

[0114] The output from CAE encoder 802 is combined at combination node 801 with noise(t) to introduce variability and robustness in the music instructions 24 to be output. This combined output from node 801 is then fed into an encoder 804 of the Diffusion UNet LDM 803. Still further, the output from text encoder 806 synchronized to time signal t (i.e., the output from node 807) is fed into both the encoder 804 and a decoder 805 of the Diffusion UNet LDM 803.

[0115] The Diffusion UNet LDM 803 is a Latent Diffusion Model using a UNet architecture as described with respect to Fig. 6, to which it is referred to avoid repetitions.

[0116] The output of the Diffusion UNet LDM 803 is then fed into the CAE decoder 808 which is configured to generate the final music instructions 24 from the refined latent representation. The music instructions 24 may include information such as a chord progression details about the sequence of chords that should be used (e.g., I-IV-V, ii-V-I), notes or a music sheet, harmonies, etc. The music instructions 24 enable the user U to play along the song output by the speaker 3. The music instructions 24 may further include information to generate AR overlays to be displayed on the display unit 2 of the smart glasses 1, the AR overlays indicating notes or chords on a user’s instrument which are to be played.Fig. 8 illustrates a method 80 for displaying the current lyrics 20.

[0117] At 81, audio data corresponding to the environmental music is captured. The audio data may be captured by the microphone of the electronic device 1.

[0118] At 82, the song is recognized based on the captured audio data 10. That is, the song ID 12 of the song is determined.

[0119] At 83, the song audio data 14 corresponding to the recognized song is acquired.

[0120] At 84, the time-synced lyrics data 18 corresponding to the recognized song are retrieved.

[0121] Additionally or optionally, at 84, the time-synced lyrics suggestion data 22 may be generated. At 85, a machine learning model (e.g., machine learning model 301) is performed to determine a current song location 16 of the song as the environmental music.

[0122] At 86, the current lyrics 20 are extracted from the time-synced lyrics data 18. The current lyrics 20 are synchronized with the current song location 16. If previously generated at 84, the current lyrics 20 may be also extracted from the time-synced lyrics suggestion data 22.

[0123] At 87, the current lyrics 20 are displayed on the display unit 2 of electronic device 1.

[0124] Optionally, at 88, a further machine learning model (e.g., machine learning 801) is performed to generate the music instructions 24. The music instructions 24 may be displayed on the display unit 2 of the smart glasses 1.

[0125] Fig. 9 illustrates a method 90 for determining the current song location 16 according to step 85. Method 90 may be performed by the machine learning model of Fig. 4.

[0126] At 91, the captured audio data 10 is segmented by extracting the latest segment (captured audio segment 32) from the captured audio data 10.

[0127] At 92, the song audio data 14 is segmented into a plurality of overlapping song audio segments 31. Each segment represents a brief snippet of the song audio data, and the segments overlap to ensure continuity and accuracy in the later feature extraction.

[0128] At 93, generating the song audio representations 33 based on the song audio segments 31.

[0129] At 94, generating the captured audio representation 32 based on the captured audio segment 32. At 95, identifying the song audio segment 35 (from the song audio segments 31) that has the highest similarity with the latest captured audio segment 31.At 96, the current song location is determined based on the identified song audio segment 35. This may involve extracting the start time from the identified song audio segment 35 and considering any processing delay time to accurately reflect the current position within the song. Fig. 10 illustrates a configuration of the electronic device 1.

[0130] The electronic device 1 includes a CPU 10001 as processor. Additionally or alternatively, other computation hardware, such as GPU, TPU, DSP etc. may be used.

[0131] The electronic device 1 further includes camera(s) 1106, microphone(s) 1007 and an audio interface / loudspeaker(s) 1008 that are connected to the processor 1001.

[0132] The processor 1001 may for example perform the machine learning models / neural network models of Figs. 3, 4, 6 and 7. Furthermore, the processor 1001 may for example implement the methods of Fig. 8 and 9.

[0133] The microphone 107 may be configured to receive any kind of audio signal.

[0134] The audio interface 1008 may be configured to emit audio signals to the user U. The audio interface 1008 may be headphones, e.g., on-ear, in-ear, over-ear, wireless headphones and the like, or may consist of one or more loudspeakers that are distributed over a predefined space and are configured to render any kind of audio, such as 3D audio.

[0135] The camera 1006 may be one or more cameras, such as an RGB camera, a ToF camera, for example, an iToF or dTof or the like.

[0136] The electronic device 1 further includes a user interface 1009 (e.g., the display unit 2 of Fig. 1) that is connected to the processor 1001. This user interface 1009 acts as a human-machine interface and enables a dialogue between a user and the electronic device 1.

[0137] The electronic device 1 further includes a Bluetooth interface 1004, and a WLAN interface 1005. These units 1004, 1005 act as I / O interfaces for data communication with external devices. For example, additional loudspeakers, microphones, and cameras with WLAN or Bluetooth connection may be coupled to the processor 1001 via these interfaces 1004 and 1005.

[0138] The electronic device 1 further includes a data storage 1002 and a data memory 1003 (here a RAM). The data memory 1003 is arranged to temporarily store or cache computer instructions for processing by the processor 1001 (e.g., song recognition 100, song acquisition 200, song location determination 300, lyrics retrieval 400, lyrics extraction 500, lyrics generation 700, music instruction generation 800, and / or data (e.g., captured audio data 10, song ID 12, songaudio data 14, current song location 16, time-synced lyrics data 18, current lyrics 20, time-synced lyrics suggestion data 22, music instructions 24, song audio segments 31, captured audio segment 32, song audio representations 33, captured audio representation 34, system prompts 71, 81). The data storage 1002 may be arranged as a long-term storage.

[0139] The connection between the processor 1001 and the camera 1006 may include a camera serial interface (CSI). The CSI is an interface between a camera 1006 and a host processor 1001. Thus, control signals and data from the processor 1001 to the camera 1006 as well as from the camera 1006 to the processor 1001 may be sent.

[0140] Furthermore, the electronic device 1 includes a neural network processor 1010. The neural network processor 1010 may include a graphics processing unit (GPU) and / or a tensor processing unit (TPU). The neural network processor 1010 may be configured to perform the machine learning model and the neural network model (e.g., an artificial neural network), for example, the machine learning models 301 of Fig. 3 and 3, and the neural network models 700, 800 of Fig. 6 and 7.

[0141] The electronic device 1 may be a stationary electronic device, such as a laptop computer, a personal computer, or a mobile device of any other kind of portable or wearable device, for example, a smartphone, a tablet computer, smart glasses, the smart glasses 1 or head mounted displays (HMDs) of Fig. 1, or other types of smart wearable devices, or the like.

[0142] It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding.

[0143] Please note that the division of the processor 1001 into a unit is only made for illustration purposes and that the present disclosure is not limited to any specific division of functions in specific units. For instance, the control of processor 1001 could be implemented by a respective programmed processor, field programmable gate array (FPGA) and the like.

[0144] Methods for controlling an electronic device, such electronic device 1 discussed above, is described above and under reference of Fig. 8 and 9. The methods can also be implemented as a computer program causing a computer and / or a processor, such as processor 1001 discussed above, to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer-readable recording medium is provided that storestherein a computer program product, which, when executed by a processor, such as the processor described above, causes the method described to be performed.

[0145] All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.

[0146] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.

[0147] Note that the present technology can also be configured as described below.

[0148] (1) Circuitry configured to:

[0149] acquire captured audio data corresponding to environmental music;

[0150] recognize a song based on the captured audio data;

[0151] determine a current song location of the environmental music by performing a machine learning model, based on an input of the captured audio data and song audio data corresponding to the recognized song, the machine learning model being configured to output the current song location; and

[0152] output current lyrics corresponding to recognized song and being synchronized with the environmental music.

[0153] (2) The circuitry according to (1), further configured to determine the current lyrics based on the current long location and lyrics information corresponding to the recognized song.

[0154] (3) The circuitry according to (2), further configured to retrieve time-synced lyrics data as the lyrics information.

[0155] (4) The circuitry according to (2), further configured to generate time-synced lyrics suggestion data as the lyrics information.

[0156] (5) The circuitry according to (4), further configured to perform a neural network model, based on an input of the song audio data, the neural network model being configured to output the time-synced lyrics suggestion data.(6) The circuitry according to (5), wherein the neural network model is performed based on a system prompt used as further input, the system prompt including instructions for generating the time-synced lyrics suggestion data.

[0157] (7) The circuitry according to (5) or (6), wherein the neural network model includes a Convolutional Autoencoder, CAE, encoder-decoder architecture.

[0158] (8) The circuitry according to anyone of (1) to (7), further configured to apply an audio encoder to the captured audio data and the song audio data, wherein the audio encoder is configured to extract features from the captured audio data and the song audio data and the machine learning model is performed based on the extracted features as input.

[0159] (9) The circuitry according to anyone of (1) to (8), further configured to output music instructions based on the song audio data, the music instructions including instructions for playing an instrument according to the environmental music.

[0160] (10) The circuitry according to (9), further configured to perform a further neural network model, based on an input of the song audio data, the further neural network model being configured to output the music instructions.

[0161] (11) An electronic device comprising the circuity according to (1).

[0162] (12) The electronic device according to (11), further configured to capture the audio data corresponding to the environmental music and to display the current lyrics on a display unit. (13) A method, comprising:

[0163] acquiring captured audio data corresponding to environmental music;

[0164] recognizing a song based on the captured audio data;

[0165] determining a current song location of the music output by the environmental music by performing a machine learning model, based on an input of the captured audio data and song audio data corresponding to the recognized song, the machine learning model being configured to output the current song location; and

[0166] outputting current lyrics corresponding to the recognized song and being synchronized with the music output by the speaker.

[0167] (14) The method according to (13), further comprising determining the current lyrics based on the current song location and lyrics information corresponding to the recognized song.(15) The method according to (14), further comprising retrieving time-synced lyrics data as the lyrics information.

[0168] (16) The method according to (14), further comprising generating time-synced lyrics suggestion data as the lyrics information.

[0169] (17) The method according to (15), further comprising performing a neural network model, based on an input of the song audio data, the neural network model being configured to output the time-synced lyrics suggestion data.

[0170] (18) The method according to (17), wherein the neural network model is performed based on a system prompt used as further input, the system prompt including instructions for generating the time-synced lyrics suggestion data.

[0171] (19) The method according to (17) or (18), wherein the neural network model includes a Convolutional Autoencoder, CAE, encoder-decoder architecture.

[0172] (20) The method according to anyone of (13) to (19), further comprising applying an audio encoder to the captured audio data and the song audio data, wherein the audio encoder is configured to extract features from the captured audio data and the song audio data and the machine learning model is performed based on the extracted features as input.

[0173] (21) The method according to anyone of (13) to (20), further comprising outputting music instructions based on the song audio data, the music instructions including instructions for playing an instrument according to the environmental music.

[0174] (22) The method according to (21), further comprising performing a further neural network model, based on an input of the song audio data, the further neural network model being configured to output the music instructions.

[0175] (23) The method according to (13), further comprising capturing the audio data corresponding to the environmental music and displaying the current lyrics on a display unit.

[0176] (24) A computer program comprising program code causing a computer to perform the method according to anyone of (13) to (23), when being carried out on a computer.

[0177] (25) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (13) to (23) to be performed.

Claims

CLAIMS1. Circuitry configured to:acquire captured audio data corresponding to environmental music;recognize a song based on the captured audio data;determine a current song location of the environmental music by performing a machine learning model, based on an input of the captured audio data and song audio data corresponding to the recognized song, the machine learning model being configured to output the current song location; andoutput current lyrics corresponding to recognized song and being synchronized with the environmental music.

2. The circuitry according to claim 1, further configured to determine the current lyrics based on the current long location and lyrics information corresponding to the recognized song.

3. The circuitry according to claim 2, further configured to retrieve time-synced lyrics data as the lyrics information.

4. The circuitry according to claim 2, further configured to generate time-synced lyrics suggestion data as the lyrics information.

5. The circuitry according to claim 4, further configured to perform a neural network model, based on an input of the song audio data, the neural network model being configured to output the time-synced lyrics suggestion data.

6. The circuitry according to claim 5, wherein the neural network model is performed based on a system prompt used as further input, the system prompt including instructions for generating the time-synced lyrics suggestion data.

7. The circuitry according to claim 5, wherein the neural network model includes a Convolutional Autoencoder, CAE, encoder-decoder architecture.

8. The circuitry according to claim 1, further configured to apply an audio encoder to the captured audio data and the song audio data, wherein the audio encoder is configured to extract features from the captured audio data and the song audio data and the machine learning model (301a) is performed based on the extracted features as input.

9. The circuitry according to claim 1, further configured to output music instructions based on the song audio data, the music instructions including instructions for playing an instrument according to the environmental music.

10. The circuitry according to claim 9, further configured to perform a further neural network model, based on an input of the song audio data, the further neural network model being configured to output the music instructions.

11. An electronic device comprising the circuity according to claim 1.

12. The electronic device according to claim 11, further configured to capture the audio data corresponding to the environmental music and to display the current lyrics on a display unit.

13. A method, comprising:acquiring captured audio data corresponding to environmental music;recognizing a song based on the captured audio data;determining a current song location of the music output by the environmental music by performing a machine learning model, based on an input of the captured audio data and song audio data corresponding to the recognized song, the machine learning model being configured to output the current song location; andoutputting current lyrics corresponding to the recognized song and being synchronized with the music output by the speaker.

14. The method according to claim 13, further comprising determining the current lyrics based on the current song location and lyrics information corresponding to the recognized song.

15. The method according to claim 14, further comprising retrieving time-synced lyrics data as the lyrics information.

16. The method according to claim 14, further comprising generating time-synced lyrics suggestion data as the lyrics information.

17. The method according to claim 16, further comprising performing a neural network model, based on an input of the song audio data, the neural network model being configured to output the time-synced lyrics suggestion data.

18. The method according to claim 17, wherein the neural network model is performed based on a system prompt used as further input, the system prompt including instructions for generating the time-synced lyrics suggestion data.

19. The method according to claim 17, wherein the neural network model includes a Convolutional Autoencoder, CAE, encoder-decoder architecture.

20. The method according to claim 13, further comprising applying an audio encoder to the captured audio data and the song audio data, wherein the audio encoder is configured to extract features from the captured audio data and the song audio data and the machine learning model is performed based on the extracted features as input.

21. The method according to claim 13, further comprising outputting music instructions based on the song audio data, the music instructions including instructions for playing an instrument according to the environmental music.

22. The method according to claim 21, further comprising performing a further neural network model, based on an input of the song audio data, the further neural network model being configured to output the music instructions.

23. The method according to claim 13, further comprising capturing the audio data corresponding to the environmental music and displaying the current lyrics on a display unit.