System for the automatic enrichment of audio streams with dynamically generated video content
The advanced system addresses scalability, flexibility, and personalization challenges in audio stream enrichment by integrating AI-driven data processing and adaptive video compression, ensuring optimal quality and secure user interactions.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2026-03-12
AI Technical Summary
Existing systems for enriching audio streams with video content face limitations in scalability, flexibility, personalization, adaptive video compression, dynamic application of video filters and transitions, accurate subtitle generation, and secure digital ID management, particularly in handling unstructured data and real-time network conditions.
An advanced system integrating AI-based data processing algorithms for aggregating and analyzing data from structured and unstructured sources, applying adaptive video compression, dynamic video filters, and secure digital ID management, with modules for metadata analysis, subtitle generation, and real-time video transition, enabling personalized and efficient multimedia content production.
The system provides scalable, flexible, and personalized multimedia content production with optimal video quality under varying network conditions, dynamic visual effects, accurate subtitles, and secure user interactions, overcoming limitations of conventional technologies.
Smart Images

Figure IB2025058220_12032026_PF_FP_ABST
Abstract
Description
[0001] LEIBO. / 20e2025
[0002] “System for the automatic enrichment of audio streams with dynamically generated video content”
[0003] Description
[0004] Field of the invention
[0005] The present invention operates in the radio and radiovision fields, integrating completion systems between structured and unstructured data. In particular, the invention falls within the field of multimedia processing and transmission systems, with particular reference to technologies for the automatic enrichment of audio content with dynamically generated visual elements. The system integrates audio metadata analysis algorithms, Al-based video processing, and advanced data compression and transmission techniques.
[0006] Prior art
[0007] The present patent application relates to a system for the automatic enrichment of audio streams with dynamically generated video content, an innovation that aims to address the growing need for rich and engaging multimedia content in the digital age. The evolution of communication technologies and the increasing demand for high-quality audiovisual content have created a significant challenge for content creators, broadcasters and streaming platforms. Until now, producing rich audiovisual content has required a significant investment of time, and human and financial resources, limiting the ability of many organizations to provide a complete and personalized multimedia experience. Existing solutions have often involved manual editing and post-production processes, which are not only expensive but also poorly scalable to meet the demand for real-time content. Furthermore, traditional audiovisual production methods have failed to fully exploit the potential of data and metadata associated with audio content, thus missing opportunities for contextual enrichment and personalization. The systems currently used for generating video content from audio streams are typically limited in their ability to dynamically adapt to network conditions and user preferences, and LEIBO. / 20e2025 often lack sophisticated mechanisms and automation for integrating external data sources as well as real-time analysis of social media trends. Previous solutions also struggled to effectively handle adaptive video compression, which is crucial for ensuring smooth transmission over networks with variable bandwidth, and offered limited options for applying video filters and contextual transitions. Automatically generating accurate subtitles and seamlessly integrating external guests through digital ID management systems were additional challenges that existing systems addressed with mixed results.
[0008] In the field of audio and video system integration, and especially in the field of connecting video images to an audio input, there are numerous patents that have addressed providing solutions to various problems.
[0009] An example is the subject of patent US2012232681A1 by M.L. STARLIGHT et al. The invention provides a method for creating audiovisual content starting from audio content. The patent uses software that extracts metadata from an audio file and searches a database of visual elements to then produce the audiovisual content. The aforementioned patent presents the technical problem of being exclusively referred to structured databases. In fact, the patent envisages of video content from one or more databases that are structured. However, today the most quantitatively available data is unstructured data which also provides the advantage of being continuously updated / updatable by means of various different sources.
[0010] Some technologies for collecting unstructured data are known to date, however, the present invention is aimed at providing a system capable of also keeping track of the people who use it, thus allowing sending personalized and specialized advertising messages between audio / video based on the use the user is making thereof. The system also allows synchronizing different videos and audios, delaying scenes, or even preparing them in advance so as to always provide multimedia content smoothly and always ready to be played.
[0011] Another example of a document that provides a system for offering real-time information and providing video, audio, text and other output is patent application KR20050012358A by HONG SOON HO. The patent application aims to provide a system for offering personalized real-time multimedia traffic information combining a moving image, text and voice, and a LEIBO. / 20e2025 related service method. The patent application describes a traffic information preprocessor that converts a basic signal into video output on a digital display with text. A basic traffic information database stores still or continuous images of multiple locations, traffic detection data, and incident information. A traffic information speech generator converts traffic information and accident information loaded from a traffic detection database and an accident information database into speech. A moving image traffic information generator comprises an image-text combination module, an image size control module, and a moving image conversion module.
[0012] The patent application, in addition to being exclusively related to traffic information, presents a series of fundamental problems related to the impossibility for system users to personalize their experience and to do so in real time, furthermore the system described is limited only to the use of data extracted from structured sources.
[0013] Continuing the analysis of the state of the art, it is therefore important to note that the traditional solutions for enriching audio streams with video content have often presented significant limitations in terms of scalability, flexibility and personalization. As seen, the conventional systems, by way of non-limiting example, typically rely on static approaches that are unable to dynamically adapt to changing network conditions or user preferences in real time.
[0014] One critical area where the existing systems are lacking is the ability to efficiently integrate and analyze data coming from different sources. While some solutions offer basic metadata incorporation capabilities, they lack a complete, integrated system for aggregating and analyzing data from structured and unstructured sources. This limitation prevents the creation of truly contextual and personalized video content.
[0015] Furthermore, the current systems often have a limited ability to handle video compression adaptively. The conventional compression techniques apply uniform algorithms to the entire video stream, without considering the relative importance of different areas of the image or the context of the content. This can lead to suboptimal video quality, especially under less than ideal network conditions. LEIBO. / 20e2025
[0016] Another area where existing solutions show shortcomings is the dynamic application of video filters and transitions. While some systems offer a selection of predefined filters, they lack a sophisticated mechanism for contextually applying these visual effects based on the real-time analysis of audio content and associated metadata.
[0017] Automatically generating accurate and synchronized subtitles represents a further challenge for the current systems. Many existing solutions rely on basic speech recognition technologies that often produce inaccurate results, especially in the presence of different accents or specialized terminology.
[0018] One critical area where the conventional systems show significant limitations is the management of digital IDs and the integration of external guests. The existing solutions often lack a robust and secure system for verifying digital identity and managing user preferences, limiting the possibilities for personalization and interaction.
[0019] The present invention stands out from the existing solutions for its integrated and highly flexible approach to the enrichment of audio streams aimed at providing solutions to the aforementioned technical problems. The present invention comprises advanced data processing algorithms that offer the possibility of aggregating, analyzing and structuring data coming from a plurality of sources. By integrating advanced data processing technologies, adaptive video compression, contextual visual effects application and secure digital ID management, the present invention provides a comprehensive and highly flexible solution that overcomes the limitations of existing systems. This innovation opens up new possibilities in the field of multimedia content production and distribution, offering a level of personalization, efficiency and quality previously unattainable with the conventional technologies.
[0020] The advantages offered by the present invention will be clearer in light of the detailed descriptions which follow.
[0021] Description of the invention
[0022] The system according to the present invention is an advanced system for the automatic enrichment of audio streams with dynamically generated video content. The heart of the LEIBO. / 20e2025 system is a data processing algorithm that receives an audio stream, analyzes its metadata, searches unstructured sources, and organizes the information into a structured database. This algorithm automatically generates an enriched video stream by combining the original audio with the obtained visual information. The system also includes audio / video composition software to integrate the audio stream with the generated video content, a camera for capturing images and videos in real time, and a data transmission device for transmitting the enriched video stream over a communication network. A key element is the artificial intelligence-based video compression module, which applies adaptive compression to the video stream based on network conditions, preferably compressing the parts that are less important for perceptual quality. The system also includes a dynamic video filter application module, a subtitle generation and application module, an advanced metadata analysis module to identify musical genre, specific themes and moods, a video transition management module to apply context- aware transitions, and a verified digital ID management module to allow the connection of external guests. The data processing algorithm also handles selective camera activation based on predefined criteria, such as identifying peaks of interest on social media. The system offers a preview module with multi-angle display and prediction of the impact of settings on the final content. Users can personalize camera activation rules, and the digital ID management system associates each ID with a specific user profile with personalized preferences. The video compression module can selectively compress parts of the video that are less important for perceptual quality, such as background areas or areas with less motion. The preview module includes a multi-angle display and an impact prediction system, as well as modules for consumption estimation and delay prediction. The dynamic video filter application module comprises a library of predefined filters and an automatic selection module. The data processing algorithm includes a social media monitoring module for analyzing trends and discussions in real time. The advanced metadata analysis module comprises modules for musical genre identification, specific theme recognition and mood analysis. The video transition management module includes one or more predefined transition libraries and a contextual analysis module. Finally, the verified digital ID management module comprises a LEIBO. / 20e2025 user profile database and a personalization module for automatically applying user preferences. The system also includes a user interface to allow users to customize the rules and criteria for camera activation.
[0023] The advantages offered by the present invention are evident in the light of the description presented thus far and will be even clearer thanks to the attached figures and the related detailed description.
[0024] Description of the figures
[0025] The invention will be described hereinafter in at least a preferred embodiment by way of nonlimiting example with the aid of the appended figures, in which:
[0026] - FIGURA 1 it shows a general view of a system for the automatic enrichment of audio streams with dynamically generated video content according to the present invention;
[0027] - FIGURA 2 it shows the use of cameras 104 communicating with a data transmission device 105 that transmits video streams through a communication network 106 in accordance with the present invention;
[0028] - FIGURA 3 it shows how the communication network 106 works synergistically with a video compression module 107, with a dynamic video filter application module 108, with a subtitle generation and application module 109 and with a video transition management module 111 to generate an enriched video stream that an audio / video composition software 103 integrates into the audio stream;
[0029] - FIGURA 4 it shows a user with a personal device using an interface from which to access the streams of the cameras 104 and from which to manage the modules 107, 108, 109 and 111.
[0030] Detailed description of the invention
[0031] The system of the present invention has a number of advantages, the main ones being summarized below: the system allows enriching an audio stream with visual material (photos, videos, icons, text) extracted from both structured and unstructured sources; LEIBO. / 20e2025
[0032] - the system enables Al-based adaptive video compression that overcomes the limitations of conventional compression systems. This approach, which intelligently analyzes the content of a video stream to apply selective compression, ensures optimal video quality even in non-ideal network conditions (slow data transmission network);
[0033] - the system allows you applying video filters and transitions (as well as other effects such as subtitles) dynamically, offering a level of sophistication and contextualization capable of adapting the visual appearance of the content in real time based on an in-depth analysis of the audio context;
[0034] - the system allows a high degree of personalization and storage of personalized features by means of digital IDs.
[0035] The advantages just described will be even more apparent in light of the descriptions which follow.
[0036] The present invention will now be illustrated by way of a purely non-limiting or binding example, resorting to the figures which illustrate some embodiments with respect to the present inventive concept.
[0037] With reference to FIG. 1, it shows a general view of a system for the automatic enrichment of audio streams with dynamically generated video content according to the present invention. In FIG. 1 as in the following description, the embodiment of the present invention currently considered the best is illustrated.
[0038] The present invention describes a system for the automatic enrichment of audio streams having metadata, with dynamically generated video content, which comprises at least:
[0039] - a data processing algorithm 101 capable of receiving an audio stream, analyzing its metadata, performing searches on structured and / or unstructured sources based on such metadata, organizing information obtained from the analysis of the metadata and said searches in a structured database 102 and automatically generating an enriched video stream by combining the audio stream with the visual information obtained. Said video stream is generated by acquiring videos, when available, related to said audio stream (such as music videos), from one or more cameras 104 when audio and video are transmitted LEIBO. / 20e2025 simultaneously (for example during live streaming), and / or further, by acquiring videos from parts of videos, the video stream can in fact be generated by combining together various parts of videos, various images and / or various texts which are shown to the public on backgrounds of various nature;
[0040] - an audio / video composition software 103 configured to integrate the audio stream with the generated video stream;
[0041] - a data transmission device 105 (shown in FIG. 2) configured to transmit the enriched video stream through a communication network 106;
[0042] - an artificial intelligence-based video compression module 107 configured to apply adaptive compression to the video stream in an interval preferably comprised between 10% and 90% based on the conditions of the communication network 106, preferably compressing the parts of the video that are less important for the perceptual quality. Here, “adaptive compression” means compression that is performed in different ways. The adaptive compression module 107 is configured to apply different compression levels to different parts;
[0043] - a dynamic application module 108 of video filters configured to apply specific filters to the video stream in real time. As just mentioned, said module 108, in preferred embodiments of the present invention, operates synergistically with the video compression module 107 and the related artificial intelligence, to apply video filters to said video stream, so as to produce a digital compression (weight reduction) and facilitate (make it faster and smoother) its sending. The dynamic video filter application module 108 comprises at least a library of predefined filters and an automatic filter selection module. The predefined filter library contains a wide range of visual effects, including color, contrast, saturation filters, as well as more complex effects such as simulating cinematic or artistic styles. Each filter in the library is associated with metadata describing its characteristics and optimal usage scenarios;
[0044] - a subtitle generation and application module 109 configured to create and overlay custom subtitles on the video stream. In some embodiments of the present invention, the subtitles LEIBO. / 20e2025 may be automatically generated using speech transcription and translation algorithms;
[0045] - a metadata analysis module 110 configured to identify, through the analysis of the audio stream metadata, at least the musical genre, specific themes and mood. In some preferred embodiments of the present invention, said advanced analysis module 110 is capable of analyzing said metadata through a musical genre identification module. The advanced metadata analysis module 110 therefore comprises, in some embodiments of the present invention, a musical genre identification module, a theme recognition module, and a mood analysis module. The musical genre identification module uses machine learning algorithms, preferably convolutional neural networks, trained on large music datasets to accurately classify the genre of the audio track being played. The musical genre identification module is able to recognize a wide range of musical genres and sub-genres, even adapting to fusion or experimental styles. The topic-specific recognition module analyzes song lyrics (if available) and associated metadata to identify the main topics covered in the song. This topic recognition module uses natural language processing and semantic analysis techniques to extract key concepts and categorize them into broader themes. The mood analysis module is configured to determine the emotional atmosphere of the audio content. This mood analysis module analyzes acoustic characteristics such as tempo, pitch, dynamics and timbre of the music, combining them with lyric analysis (if present) to infer the overall mood of the song. The system uses a multidimensional scale to classify mood, considering dimensions such as, but not limited to, energy, emotional valence and tension.
[0046] In other preferred embodiments of the present invention, the metadata contains information already indicative of the mood, musical genre and tone of the audio track. In fact, the metadata in these embodiments are numeric values that can refer to entries present in the structured database 102. By way of non-limiting or non-binding example, the structured database 102 comprises a list of moods and the metadata can have a numeric or other identifier value that references the location of a specific mood in the structured database 102; LEIBO. / 20e2025
[0047] - a video transition management module 111 configured to apply context-aware transitions appropriate to the video context. The video transition management module 111, in some embodiments of the present invention, comprises a library of predefined transitions and a contextual analysis module. The predefined transition library contains a wide range of transition effects, from simple fades to more complex effects such as 3D rotations, morphing or particle-based transitions. Each transition in the library is associated with metadata that describes its visual characteristics and emotional impact. The contextual analysis module is configured to select and apply appropriate transitions based on the context of the video and the characteristics of the audio stream. This contextual analysis module analyzes the visual content immediately before and after the transition point, considering factors such as dominant color, motion, and image composition. Furthermore, the contextual analysis module takes into account audio characteristics such as rhythm, intensity and mood of the music to select a transition that harmoniously integrates with the audiovisual stream. The system applies the transitions in real time, ensuring perfect synchronization with the audio and maintaining the visual coherence of the video stream;
[0048] - a verified digital ID management module 112 configured to allow the connection of external guests via preset and verified digital IDs; said verified digital ID management module 112 is capable of associating each ID with a specific user profile, preferably including preferences related to musical genres and video quality settings;
[0049] - a preview module capable of providing a forecast of the influence of the settings on the enriched video stream, preferably including information on possible consumption and potential delays in transmission due to settings chosen by a user and / or a system administrator user. Where “settings” here means the entire set of additions, video filters, compressions, subtitles, and unions of various still image and / or video components, related to the enriched video stream. Thereby, the preview module can show how a different type of filter or compression can result in a more or less smooth video experience (for example, a video that continuously stops due to excessive weight with respect to the available network). The preview module, in some preferred embodiments of the present invention, LEIBO. / 20e2025 further comprises a consumption estimation module and a delay prediction module. The consumption estimation module is configured to calculate and display the expected resource consumption based on the selected settings. Said consumption estimation module analyzes parameters such as video resolution, frame rate, compression level and applied effects to estimate CPU, GPU and memory usage. The estimates are presented to the user in graphical form, highlighting any configurations that could overload the system.
[0050] The preview module, as seen, also comprises said delay prediction module which is configured to estimate and display potential transmission delays based on the complexity of the chosen settings. Said preview module takes into account the available bandwidth, the applied compression level and the complexity of the video effects to calculate the expected delay in transmitting the video stream. Information about potential delays is presented to the user in a very short time (within a few milliseconds), allowing settings to be optimized to ensure smooth transmission.
[0051] With specific reference to FIG. 2, it shows a data transmission device 105 which, comprising appropriate transmission antennas, sends the audio streams acquired by two cameras 104 through said network 106. FIG. 2 shows a camera 104 physically connected (by means of wired connection) to said device 105 and a second camera 104 which is connected to the device 105 by means of wireless connection which may be, by way of non-limiting or binding example, a connection via Bluetooth, via local Wi-Fi network and / or other.
[0052] In some embodiments of the present invention such as that shown in FIG. 2 and 4, the system further comprises one or more cameras 104 for capturing images and videos in real time; said cameras 104 being connected to the system and generating a video stream available for live streaming and / or to be saved and enjoyed at a later time.
[0053] FIG. 3 shows a communication network 106 which, in synergy with the video compression modules 107, the dynamic application 108 of video filters, the generation and application 109 of subtitles and the management 111 of video transitions, generates an enriched video stream which an audio / video composition software 103 integrates into the audio stream. The modules 107, 108, 109 and 111 are shown as being mutually interconnected to each other, as changes LEIBO. / 20e2025 made by means of any of the modules 108, 109 and 111 will be reflected in the video compression managed by the module 107 and vice versa.
[0054] In some embodiments of the present invention, the system provides graphical interfaces to users to enable the activation and / or deactivation of said cameras 104; said graphical interfaces being configured to enable users to customize the rules and criteria for activating the camera 104, preferably including options for defining social media interest thresholds, specific keywords, or time events. The user interface then allows users to customize the rules and criteria for the activation of the camera. The interface offers advanced options for defining social media interest thresholds, allowing users to specify the level of engagement or number of mentions needed to automatically activate the camera. Users can also define specific keywords or hashtags that, if detected in social media discussions, will trigger the camera to activate. The interface also includes the ability to set time events, allowing users to schedule the camera to activate at specific times during audio stream playback. The user interface is preferably implemented as a responsive web application, accessible from different devices and optimized for use on both desktop and mobile devices.
[0055] In other embodiments of the present invention such as that shown in FIG. 4, said graphical interfaces communicate with said preview module to show said users multi-angle previews of the shots from said cameras 104. Here, “multi-angle previews” are intended to mean previews in which the screens related to the shots of each camera are accessible and are preferably aggregated in a single GUI (Graphic User Interface) from which the user can choose to view the shots of a camera 104 by clicking on the relevant preview and / or can choose a sequence and / or a loop of cameras 104 whose content is to be shown. By way of non-limiting or binding example, a sequence can be defined by setting that the footage from a first camera 104 is shown for a time X (for example equal to 2 minutes) and that the footage from a second camera 104 is shown, after said time X, until the end of the transmissions. Another loop of cameras 104 can be defined by specifying to the system to show the footage from a first camera 104 for a certain time T1 (for example equal to 30 seconds), then show the footage from a second camera 104 for a time T2 (for example equal to 1 minute) and then start the cycle again from LEIBO. / 20e2025 the first camera 104. As seen, the multi -angle display is configured to simultaneously show different shooting perspectives, allowing the user to evaluate the overall visual effect of the generated content. This feature preferably uses the said graphical interface divided into multiple panes, each showing a different camera angle. The user can select and zoom into a specific view for a more detailed analysis. The influence forecasting system is configured to simulate the effect of the selected settings on the final content. This system uses real-time rendering algorithms to generate an accurate preview of the final video, taking into account all the parameters set, including filters, transitions and special effects.
[0056] In some embodiments of the present invention, the preview module comprises a delay prediction module capable of performing automatic network data transmission speed tests; said preview module is adapted to perform at least an automatic test when a user requests to view a preview. By way of non-limiting or binding example, the automatic speed test is a PING (Packet INtemet Groper) test conducted by applying filters, video effects, automatic subtitles, compressions and / or other transformations to the video stream chosen by the users. The delay prediction module thereby estimates and displays potential transmission delays based on the complexity of the chosen settings.
[0057] In some embodiments of the present invention, the data processing algorithm 101 comprises a social media monitoring module configured for:
[0058] - analyzing in real time trends and discussions on social networks related to the topics present in the audio stream. Said social media monitoring module uses natural language processing and sentiment analysis techniques to identify and track relevant conversations on platforms such as but not limited to Twitter, Facebook and Instagram. The social media monitoring module is able to recognize hashtags, mentions and keywords related to the audio content being played. The social media monitoring module is also configured to selectively activate the camera when a spike in interest in a specific topic is detected. The system uses trend detection algorithms to identify sudden increases in social activity related to audio content;
[0059] - automatically activate a camera 104 when a peak of interest on a specific topic is detected. When a peak in interest is detected, the system automatically activates the camera to capture LEIBO. / 20e2025 relevant visual content, thus integrating the video stream with topical elements and increasing audience engagement.
[0060] In some embodiments of the present invention, said compression module 107 comprises at least a video content analysis module. Said video content analysis module being capable of examining said video stream at any moment and identifying the areas of greater and lesser importance for the perceptual quality. Said video content analysis module uses deep learning algorithms, preferably convolutional neural networks, trained on large image and video datasets to recognize and classify elements in the video. The parts (meaning both parts on a timeline in the video stream, and parts at different positions within a frame) identified as less important, such as background areas or parts with less motion, are compressed with a higher compression level, preferably in a range comprised between 60% and 90% and even more preferably between 70% and 80%. Conversely, parts of greater visual interest, such as faces, foreground objects or areas with significant movement, parts related for example to a refrain, are compressed with a lower compression level, preferably in a range comprised between 10% and 40% and even more preferably between 15% and 25%. The result related to "compression" or the reduction of the "weight" / "memory space" occupied by one or more frames of a video is preferably achieved by using video filters which, by operating blurs / posterizations / desaturations and / or other graphic and computer transformations (for example digital coding) on one or more frames and on their totality or only in specific areas of them automatically, reduce their weight without, however, affecting the experience of enjoying the enriched video stream. The compression module 107 is configured to constantly monitor the conditions of the communication network and dynamically adapt the overall compression level. In the event of slow or congested connections, the system increases the overall compression level, while maintaining the differentiation between areas of greater and lesser importance. When the available bandwidth is higher, the system reduces the compression level to improve the overall video quality. The video compression module is preferably implemented using dedicated hardware, such as GPUs or specialized video processing ASICs , to ensure real-time performance even on high-resolution video streams. LEIBO. / 20e2025
[0061] In some embodiments of the present invention, said dynamic filter application module 108 comprises automation algorithms for choosing and applying filters and video effects based on audio and video stream parameters. These automation algorithms operate these choices and these applications at every instant during the stream, obtaining data relating to the available connection bandwidth at every instant. These automation algorithms include machine learning algorithms, preferably recurrent neural networks, for analyzing the audio and video content and select the most appropriate filter (within the available bandwidth, so as to generate a seamless video stream as well as most appropriate to the context of the video) at each instant. As seen, the appropriateness is also evaluated based on the context of the video; in fact, the system takes into account factors such as the musical genre, the mood of the audio track, the visual content and the user's preferences to make the selection. Filters are applied to the video stream in real time, with smooth transitions between different effects to maintain visual consistency. The analysis algorithms are such as to evaluate the time of a transition so that “fluidity” is guaranteed, for example by preventing transitions having a certain duration in time (for example 0.3 seconds) from being placed next to transitions with a completely different duration (for example transitions of 4 seconds). The analysis algorithms also evaluate the audio and video context and adjust the succession of said transitions instant by instant, even being able to choose a sudden change in the type of transitions if this is linked to a sudden change in the audio / video stream.
[0062] In some embodiments of the present invention, the verified digital ID management module 112 comprises:
[0063] - a database of user profiles, where each profile is associated with a specific digital ID. The profiles include musical genre preferences, video quality settings, favorite filters, and other user experience personalizations. The database is preferably implemented using a distributed database management system to ensure high availability and scalability. By way of non-limiting or binding example, the database management system may be an RDBMS (Relational DataBase Management System);
[0064] - a personalization module configured to automatically apply the user’s preferences LEIBO. / 20e2025 regarding musical genre and video quality settings based on the digital ID used for access. When a user logs into the system using their verified ID, the module instantly retrieves the associated profile and applies the saved preferences to the generated video stream. This includes adapting video filters, selecting visual content related to preferred musical genres, and applying desired video quality settings. The system is able to dynamically update the user profile based on the interactions and preferences expressed during use, constantly improving the personalization of the experience.
[0065] In some embodiments of the present invention, the system comprises one or more digital libraries of preset digital IDs associated with a short description; a user external to the system, by connecting to it, can choose a digital ID from said libraries based on the description, to have a personalized user experience. The digital IDs have descriptions that can be related to the preferred musical genres for that ID, the preferred type of transitions, filters, video compression, and even possible camera change presets 104 when available. As an example, an ID description can describe a preset that is such as to show camera shots 104 that are designed to always film the close-up of the people speaking, for example, in a talk / an interview / a meeting or other.
[0066] Finally, it is clear that modifications, additions or variations that are obvious to a person skilled in the art can be made to the invention described so far, without thereby departing from the scope of protection provided by the attached claims.
Claims
LEIBO. / 20e2025Claims1. System for the automatic enrichment of audio streams having metadata, with dynamically generated video contents, characterised in that it comprises at least:- a data processing algorithm (101) capable of receiving an audio stream, analysing its metadata, performing searches on structured and / or unstructured sources based on such metadata, organizing information obtained from the analysis of the metadata and said searches in a structured database (102) and automatically generating an enriched video stream by combining the audio stream with the visual information obtained;- an audio / video composition software (103) configured to integrate the audio stream with the generated video stream;- a data transmission device (105) configured to transmit the enriched video stream through a communication network (106);- an artificial intelligence-based video compression module (107) configured to apply adaptive compression to the video stream based on the conditions of the communication network (106), preferably compressing the parts of the video that are less important for the perceptual quality;- a dynamic application module (108) of video filters configured to apply specific filters to the video stream in real time;- a subtitle generation and application module (109) configured to create and overlay custom subtitles on the video stream;- a metadata analysis module (110) configured to identify, through the analysis of the audio stream metadata, at least the musical genre, specific themes and mood;- a video transition management module (111) configured to apply transitions appropriate to the video context;- a verified digital ID management module (112) configured to allow the connection of external guests via preset and verified digital IDs; said verified digital ID management module (112) is capable of associating each ID with a specific user profile, including preferences related to musical genres and video quality settings;- a preview module capable of providing a forecast of the influence of the settings on the enriched video stream, including information on possible consumption and potential delays in transmission due to settings chosen by a user and / or a system administrator user.
2. System for automatically enriching audio streams with dynamically generated videoLEIBO. / 20e2025 content, according to claim 1, characterized in that it further comprises one or more cameras (104) for capturing images and videos in real time; said cameras (104) being connected to the system and generating a video stream available for live streaming and / or to be saved and enjoyed at a later time.
3. System for automatically enriching audio streams with dynamically generated video content, according to the preceding claim 2, characterized in that it provides graphical interfaces to users to enable the activation and / or deactivation of said cameras (104); said graphical interfaces being configured to enable users to customize the rules and criteria for activating the camera (104), preferably including options for defining thresholds of interest on social media, specific keywords or time events.
4. System for the automatic enrichment of audio streams with dynamically generated video content, according to the preceding claim 3, characterized in that said graphic interfaces communicate with said preview module to show said users multi-angle previews of the shots from said cameras (104).
5. System for the automatic enrichment of audio streams with dynamically generated video content, according to any of the preceding claims 2-4, characterized in that the preview module comprises a delay prediction module capable of carrying out automatic tests of data transmission speed on the network; said preview module is capable of carrying out at least an automatic test when a user requests to view a preview.
6. System for the automatic enrichment of audio streams with dynamically generated video content, according to any of the preceding claims 2-5, characterized in that the data processing algorithm (101) comprises a social media monitoring module configured to:- analyse in real time trends and discussions on social networks related to the topics present in the audio stream;- automatically activate a camera (104) when a peak of interest on a specific topic is detected.
7. System for the automatic enrichment of audio streams with dynamically generated video content, according to any of the preceding claims, characterized in that said compression module (107) comprises at least a video content analysis module; said video content analysis module being able to examine said video stream at any time and identify the areas of greater and lesser importance for perceptual quality.
8. System for the automatic enrichment of audio streams with dynamically generated video content, according to any of the preceding claims, characterized in that said dynamicLEIBO. / 20e2025 filter application module (108) comprises automation algorithms able to choose and apply filters and video effects based on parameters of the audio and video stream; said automation algorithms operate said choices and said applications at any time during the stream, obtaining, at any time, data relating to the available connection bandwidth.
9. System for the automatic enrichment of audio streams with dynamically generated video content, according to any of the preceding claims, characterised in that the verified digital ID management module (112) comprises:- a database of user profiles, where each profile is associated with a specific digital ID;- a personalization module configured to automatically apply the user’s preferences regarding music genre and video quality settings based on the digital ID used for access.
10. System for the automatic enrichment of audio streams with dynamically generated video content, according to any of the preceding claims, characterized in that it comprises one or more digital libraries of preset digital IDs associated with a short description; a user external to the system, by connecting to it, can choose a digital ID from said libraries based on the description, to have a personalized user experience.
Citation Information
Patent Citations
The configuration and service method of real-timemultimedia traffic information system which combinesmoving image, character and voice
KR1020050012358A
System and method for using a list of audio media to create a list of audiovisual media
US20120232681A1
Creating a video for an audio file
US10122983B1
System, method, and program product for generating graphical video clip representations associated with video clips correlated to electronic audio files
US10123068B1
Method, System, and Apparatus for Generating Video Content
US20180075879A1