Audio and video processing for identification and synchronization

CN122847872APending Publication Date: 2026-09-29HULSEY LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202580016002.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-30
Filing Date
2025-02-27
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

此外,当前方法缺乏动态连接管理,通常需要人工干预来维持或终止连接,这降低了系统的可靠性和鲁棒性

Benefits of technology

[0037]根据本发明的一个方面,提供了一种存储用于执行上述任何方面的方法的指令的瞬态或非瞬态计算机可读介质。有益地,该介质提供了一种方便高效的方式来实现所述方法,确保配对和同步过程执行的一致性和可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122847872A_ABST
    Figure CN122847872A_ABST
Patent Text Reader

Abstract

A method for media content identification and content information retrieval on a user device is disclosed. The method includes monitoring a portion of an audio track associated with media content and detecting whether an audio tag is embedded in the portion. The audio tag contains identification data associated with known media content at a known time point. If the audio tag is detected, the identification data is extracted. If the audio tag is not detected, the identification data is received based on a best match of an audio fingerprint. The media content associated with the audio track is then identifiable based on the known media content of the identification data. The method also includes retrieving content information associated with the identification data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various aspects of this invention relate to establishing secure communication between devices. Specifically, but not exclusively, various aspects of this invention relate to establishing secure communication for output synchronization between devices using audio or video signals. Specifically, but not exclusively, various aspects of this invention relate to audio or video signal processing for relevant information retrieval and / or establishing bidirectional communication. Background Technology

[0002] The current media consumption landscape is characterized by access to a vast amount of content through various channels such as television, radio, and streaming services. However, despite this abundance of content, viewers often encounter difficulties identifying specific content, especially when they lack prior information such as titles or actors. Traditional search methods relying on text input are inadequate in terms of real-time content identification and contextual understanding. Current technologies primarily focus on using metadata or manual search input for content identification. This approach is not only time-consuming but also inefficient when viewers lack knowledge of content details. Furthermore, these methods fail to provide real-time synchronization with the content, limiting the depth of user interaction. Multiple search interactions between users and their devices are frequently required, resulting in a poor human-computer interaction experience, increased battery consumption and computational resource requirements, and significantly reduced efficiency.

[0003] Examples of existing technologies include Amazon X-Ray™ functionality and other similar technologies, in which topical information is displayed to the user based on streaming content shown at a specific point in time. However, there is no disclosure of displaying relevant topical information based on content broadcast from devices, apps, or services not associated with the user's device. At best, external devices like mobile phones must maintain continuous communication with display devices such as televisions to provide time-related information, which is inefficient, requires additional computing resources, and has limited functionality. Another example, US9837127B2, discloses inserting audio cues in the post-production of media files, marking specific items and placing them on the timeline, and inserting the cues into the marked audio tracks. Synchronization is impossible when the audio cues are undetectable, such as in noisy environments or when listening to media files without inserted cues. Furthermore, there is no solution to the problem of identifying, extracting, and generating relevant content information associated with media files, resulting in long search times for users, lengthy / inaccurate data retrieval, and generally poor interaction between the user's device and the media content.

[0004] Furthermore, the interconnected devices are complex. The field of secure communication and output synchronization focuses on ensuring that data exchanged between devices is protected from unauthorized access, tampering, and eavesdropping, while maintaining consistency and synchronicity of outputs between devices. This involves encryption protocols, authentication mechanisms, and synchronization technologies to achieve real-time or near real-time synchronization in distributed systems, the Internet of Things (IoT), and collaborative environments. However, challenges remain. Traditional methods often rely on visual or manual input for device pairing and synchronization, which can be cumbersome and error-prone. Moreover, ensuring secure communication between devices while maintaining real-time output synchronization remains a significant technical hurdle. These issues are particularly pronounced in environments where multiple devices need to interact seamlessly and securely, such as smart homes, collaborative workspaces, and multimedia streaming services.

[0005] One major issue is secure and efficient device pairing. Existing systems often suffer from security limitations due to reliance on traditional methods such as Bluetooth pairing or QR code scanning, which are vulnerable to man-in-the-middle attacks and unauthorized access. These methods are easily detected and intercepted, compromising overall system security. Furthermore, the user experience is cumbersome and error-prone, requiring manual input of PIN codes or QR code scanning. This complexity makes the setup process less intuitive and more frustrating, increasing the likelihood of user errors. Another significant issue is output synchronization between devices. For example, ensuring perfect synchronization of audio and video outputs across different devices is challenging in multimedia applications. Existing systems struggle to maintain real-time synchronization between device outputs. Factors such as latency, network jitter, and varying processing times can all lead to desynchronization, especially in multimedia applications requiring perfect audio-video alignment. Moreover, current methods lack dynamic connection management, often requiring manual intervention to maintain or terminate connections, which reduces system reliability and robustness. The lack of comprehensive tools like Software Development Kits (SDKs) and predefined instructions for secure communication and synchronization also presents challenges for developers, hindering the widespread adoption and integration of advanced technologies, especially for older media devices. These issues highlight the need to optimize secure communication and synchronization technologies. Summary of the Invention

[0006] One aspect of this invention proposes an innovative audio analysis system for improving content discovery and interaction, such as for television and radio. Unlike conventional methods, the disclosed system allows users to identify and engage with media content without prior knowledge of its title or details. This novel technology aims not only to identify the content being played but also to synchronize or link content discovery and interaction with the content being played in real time, thereby providing improvements in efficiency, accuracy, and reliability.

[0007] According to one aspect of the present invention, a system for media content interaction is provided. The system is configured to process known media content on a workstation device; interact with the media content on a user device; and transmit content information to the user device.

[0008] According to another aspect of the present invention, a method for media content interaction is provided. The method includes processing known media content on a workstation device, wherein processing the known media content includes: processing the known media content on the workstation device; interacting with the media content on a user device; and transmitting content information to the user device.

[0009] Processing known media content involves acquiring the known media content, which includes a known audio track. A timeline is then generated from the known media content, comprising multiple time points. For each time point in the timeline, an audio tag containing identification data is embedded in the known audio track at a frequency detectable by a microphone and imperceptible to the user on the user's device. The audio track with the embedded audio tag is called the embedded audio track. By processing one or more topic-related information extracted from the known media content at each time point, content information associated with the identification data is generated. The embedded audio track and the content information associated with each time point in the timeline are then stored.

[0010] Interacting with media content involves using a user device's microphone to listen to a portion of an unknown audio track, which is associated with media content emitted from a speaker. The media content is identified by acquiring identification data. This identification data is associated with known media content at a known point in time. If an audio tag is detected within the portion of the unknown audio track, identification data is extracted from that audio tag. If no audio tag is detected within the portion of the unknown audio track, identification data is received from a server device.

[0011] Based on the interaction between the user and the user device, a request for content information is initiated to the server device. Upon receiving the request, the content information associated with the identification data is transmitted from the server device to the user device. The transmission of content information includes: receiving a request for content information from the user device, wherein the request includes the identification data; using the identification data to retrieve stored content information associated with the identification data; and outputting the content information to the user device.

[0012] According to another aspect of the invention, a method for media content identification and content information retrieval on a user device is provided. The method includes listening to a portion of an audio track using a microphone on the user device, wherein the audio track is associated with media content emitted from a speaker. An audio tag is then detected embedded in this portion of the audio track, wherein the audio tag is embedded at a frequency detectable by the microphone and imperceptible to the user of the user device. The audio tag contains identification data associated with known media content at a known point in time.

[0013] If an audio tag is detected, identification data is extracted. If no audio tag is detected, identification data is received based on the best match of an audio fingerprint. This audio fingerprint includes one or more audio features extracted from the portion of the audio track, and the best match is the audio fingerprint most similar to this fingerprint among multiple audio fingerprints. Then, based on the known media content of the identification data, the media content associated with the audio track can be identified. The method also includes retrieving content information associated with the identification data.

[0014] According to another aspect of the present invention, a method for performing audio track identification and content information transmission on a server device is provided. The method includes generating a first audio fingerprint, the first audio fingerprint comprising one or more audio features extracted from a first portion of an audio track. The one or more audio features are extracted from the first portion obtained from a user device, or the one or more audio features are obtained from a user device. The method further includes comparing the first audio fingerprint with a plurality of audio fingerprints. Each of the plurality of audio fingerprints is associated with identification data, wherein the identification data is associated with a known audio track at a known point in time.

[0015] The best match is then identified, where the best match is the audio fingerprint that is most similar to the first audio fingerprint among multiple audio fingerprints. The identification data associated with the best match is transmitted to the user device, and if a request is received from the user device, content information associated with the identification data is subsequently submitted to the user device.

[0016] According to another aspect of the invention, a method for extracting themes from media content on a workstation device is provided. The method includes acquiring known media content, wherein the known media content includes known audio tracks. A timeline is generated for the known media content, wherein the timeline includes multiple time points.

[0017] For a first time point in the timeline, one or more audio features are extracted from a known audio track at that time point, and a first audio fingerprint containing those features is generated. Then, identification data associated with the known media content at the first time point is linked to the first audio fingerprint. A first audio tag containing that identification data is embedded into the first time point in the known audio track to generate an embedded audio track.

[0018] The method extracts one or more themes from known media content and organizes information associated with those themes to generate content information associated with the identification data. The method also includes storing a first audio fingerprint, an embedded audio track, and the content information. The identification data can be retrieved based on the first audio fingerprint or the embedded audio track, and the content information can be retrieved based on the identification data.

[0019] According to another aspect of the invention, a method for preparing content interaction on a workstation platform is provided. The method includes importing media content into the workstation platform, wherein the media content is associated with a plurality of topics and media content identifiers. A timeline associated with the media content is generated, wherein the timeline includes a plurality of time points. The workstation platform acquires one or more topic files, wherein each topic file is associated with one topic (among the plurality of topics associated with the media content). Each topic file includes a visual description of the topic and / or topic text and is assigned to at least one time point on the timeline. A rich media file containing the timeline, the one or more topic files associated with the timeline, and the media content identifiers is then output from the workstation platform.

[0020] According to another aspect of the present invention, a method for content interaction on a user platform is provided. The method includes listening to an audio signal, wherein the audio signal includes one or more audio tags embedded therein. Selected audio tags are obtained from the audio signal for identifying media content. Identification data is then extracted from the selected audio tags, wherein the identification data is associated with a media content identifier and a first time point.

[0021] If no lock indicator is received on the user platform, the method further includes obtaining a request indicator for a first theme file. This first theme file is associated with a media content identifier (in the identification data) at a first point in time. The identification data is transmitted to a server device, and subsequently, the first theme file is received from the server device. The first theme file includes a first visual depiction and first theme text. The first visual depiction or the first theme text is then displayed, and audio signals are continuously monitored.

[0022] If a lock indicator is received on the user platform, the method also includes reducing listening to audio signals. Identification data is transmitted to a server device, and multiple theme files associated with the media content identifier of the identification data are requested. These multiple theme files include a first theme file. The multiple theme files are received from the server, and the display of a first visual depiction or first theme text is initiated. Upon reaching a second time point, the display of a second visual depiction or second theme text from a second theme file associated with the second time point is initiated.

[0023] According to another aspect of the invention, a computer-readable storage medium, which is transient or non-transient, is provided, comprising one or more program instructions that, when executed by one or more processors, cause the one or more processors to perform any of the methods disclosed herein.

[0024] Benefically, this invention presents aspects of an innovative audio analysis system designed to revolutionize content discovery and interaction in media broadcasting. Unlike traditional methods, this system allows users to identify and engage with media content without prior knowledge of its title or details. This groundbreaking technology not only identifies the content being played but also synchronizes it with it in real time, providing a rich and interactive experience.

[0025] The disclosed system and method employ a dual-approach to content recognition. Direct recognition using inaudible sound codes (i.e., audio tags) combined with complex audio attribute matching using audio fingerprints as a backup sets a new standard in the field of audio analysis. The integration of real-time error correction and environmental noise processing inherent in either audio tags or audio fingerprints demonstrates significant advancements over traditional audio recognition systems. Advantageously, this improves accuracy. By utilizing two complementary techniques, the system and method described herein achieve higher content recognition accuracy even in challenging environments. Further advantageously, this enables real-time interactivity.

[0026] Users can receive instant information about the content, ranging from basic details to in-depth data such as specific product or scenario descriptions. This further enhances the user experience. Seamless operation between the two audio-based methods ensures a consistent and user-friendly experience without human intervention. The technologies encompassed have broad applications, from improving media recognition and knowledge retrieval to providing real-time interactive experiences for the topics seen in these media. Their built-in adaptability makes them suitable for various environments, whether quiet at home or noisy public spaces.

[0027] According to one aspect of the present invention, a method for secure communication and output synchronization between a user application and a transmission application is provided. The method includes scanning content by broadcasting a first request using a bidirectional communication protocol. A first voice code is detected from a microphone controlled by the user application, wherein the first voice code is output from a speaker controlled by the transmission application. The first voice code is configured to be detectable by the microphone and imperceptible to the user of the user application. The method further includes converting the first voice code into an identification code and transmitting the identification code using the bidirectional communication protocol for authentication at the transmission application. A bidirectional connection for real-time data exchange is established with the transmission application, and the bidirectional connection is used to synchronize a first output of the user application with a second output of the transmission application. Advantageously, this method provides a secure and efficient device pairing and synchronization method, enabling real-time data exchange and output synchronization, thereby enhancing user experience and system security.

[0028] According to another aspect of the invention, the bidirectional communication protocol includes WebSocket, and the bidirectional connection includes a full-duplex connection. Advantageously, WebSocket and full-duplex connections support continuous real-time data exchange, reducing latency and improving synchronization accuracy between devices.

[0029] According to another aspect of the invention, the method further includes maintaining the bidirectional connection in response to detecting a second voice code within a predetermined interval after detecting the first voice code. Alternatively, the method further includes terminating or closing the bidirectional connection in response to either: (a) user interaction with a user application, or (b) a location associated with the user application exceeding a predetermined threshold. Advantageously, the invention ensures that the connection remains active only when necessary, enhancing security and resource management by preventing unauthorized access and reducing unnecessary data transmission.

[0030] According to another aspect of the invention, the first sound code is an ultrasonic or near-ultrasonic code. Advantageously, ultrasonic codes are not easily noticeable to the user, providing a discreet and non-intrusive data transmission method, further enhancing the user experience, detection accuracy, decoding accuracy, and system security.

[0031] According to another aspect of the invention, the identification code comprises alphanumeric data, and converting the first voice code into the identification code includes decoding the first voice code. Advantageously, this enhances the accuracy and security of device identification and reduces the risk of errors and unauthorized access.

[0032] According to another aspect of the invention, detecting the first voice code is in response to a transmission application: receiving an identifier associated with a first request from a server; converting the identifier into a first voice code; and initiating the output of the first voice code from a speaker. Advantageously, this results in the secure and accurate transmission of pairing information, enhancing the reliability and security of the pairing process.

[0033] According to another aspect of the invention, the method further includes identifying a current timestamp associated with the first output and a message timestamp associated with the second output sent from the transport application. Reducing the difference between the current timestamp and the message timestamp improves the synchronization between the first and second outputs. Advantageously, this results in precise synchronization of the outputs, enhancing the user experience by providing seamless and synchronized content across devices, and improving security by maintaining time alignment between devices.

[0034] According to one aspect of the present invention, a method for secure communication and output synchronization with a user application at a transport application is provided. The method includes receiving a first request associated with the user application using a bidirectional communication protocol. A pairing request is transmitted to a server, wherein the pairing request includes data associated with the first request. If the server recognizes the pairing request, an identifier is received from the server. The transport application converts the identifier into a first sound code for emission from a speaker, wherein the speaker is controllable from the transport application. The method further includes receiving an identification code from the user application using the bidirectional communication protocol and authenticating the identification code based on a matching identifier and the identification code. If the identification code is authenticated, a handshake is performed to establish a bidirectional connection with the user application for real-time data exchange. A first output of the user application is synchronized with a second output of the transport application using the bidirectional connection. Advantageously, the present invention provides a secure and efficient device pairing and synchronization method, enabling real-time data exchange and output synchronization, thereby enhancing user experience and security.

[0035] According to another aspect of the invention, the bidirectional communication protocol includes WebSocket, and the bidirectional connection includes a full-duplex connection. Advantageously, WebSocket and full-duplex communication support continuous real-time data exchange, reducing latency and improving synchronization accuracy between devices.

[0036] According to another aspect of the invention, synchronization includes receiving a message timestamp associated with a first output using a bidirectional connection. A current timestamp associated with a second output is determined, and synchronization is based on reducing the difference between the message timestamp and the current timestamp, thereby improving the synchronization between the first and second outputs. Advantageously, this achieves precise synchronization of the outputs, enhancing security and user experience by providing seamless and synchronized content across devices.

[0037] According to one aspect of the invention, a transient or non-transient computer-readable medium is provided for storing instructions for performing the methods described above. Advantageously, this medium provides a convenient and efficient way to implement the methods, ensuring consistency and reliability in the execution of pairing and synchronization processes.

[0038] According to one aspect of the present invention, a system for secure communication and output synchronization between user equipment and transmission equipment is provided. The system includes user equipment and transmission equipment configured to perform the methods described above. Advantageously, this system provides a comprehensive solution for secure communication and synchronization between devices, enhancing user experience and security.

[0039] According to one aspect of the present invention, a software development kit (SDK) for secure communication and output synchronization is provided. The SDK includes one or more communication modules, one or more authentication modules, and one or more interaction modules. The one or more communication modules are configured to: send data from a transmission device to a user device using a bidirectional protocol; receive data from the user device to the transmission device using a bidirectional protocol; establish a bidirectional connection between the transmission device and the user device; transmit data to a server; receive data from the server; and transmit instructions to a speaker. The one or more authentication modules are configured to: generate a pairing request associated with data received from the user device; generate a first sound code by encoding an identifier received from the server; and compare the identification code with the identifier received from the server. The one or more interaction modules are configured to: identify a current timestamp associated with the output of the transmission device; calculate the difference between the message timestamp received from the user device and the current timestamp; And it enables adjustments to the output of the transmission device to minimize this difference. Benefically, this SDK provides a robust and flexible framework for achieving secure communication and synchronization, enhancing the development and deployment of such systems.

[0040] According to another aspect of the invention, the one or more communication modules are configured to establish a secure WebSocket connection (WSS) using the TLS / SSL protocol. Advantageously, this supports secure data transmission, prevents unauthorized access and data leakage, thereby enhancing the security of the communication process.

[0041] According to another aspect of the invention, generating the first sound code further includes encoding the current timestamp using added redundancy to facilitate error detection and correction during transmission, broadcasting, playback, or decoding. Advantageously, this improves the reliability and accuracy of data transmission, reduces the risk of errors, and enhances the overall performance of the system.

[0042] In summary, the above methods utilize inaudible voice codes for device pairing and synchronization. These voice codes are not easily detected by the user but are easily captured by the device's microphone. This approach enhances security and user convenience by automating the pairing process and reducing the need for manual input.

[0043] Compared to existing systems, the above aspects offer significant technical advantages in the pairing and synchronization of multimedia content. The disclosed method and system automatically complete the pairing process through voice code detection, eliminating the need for manual input (such as entering a PIN code or scanning a QR code). This automation simplifies the user experience, making it more intuitive and less error-prone, allowing users to easily connect their devices without going through complex setup procedures. Furthermore, the system ensures real-time synchronization of output between devices by establishing a bidirectional connection for real-time data exchange. Using timestamps to synchronize output addresses issues related to latency, network jitter, and varying processing times, resulting in a more cohesive and immersive user experience, particularly beneficial in applications where audio and video synchronization is advantageous.

[0044] The ability to maintain or terminate bidirectional connections based on voice code detection, user interaction, or location thresholds provides dynamic connection management. This adaptability ensures continuous secure communication and synchronization even under changing conditions, enhancing system reliability and robustness. Furthermore, computer-readable media containing a software development kit (SDK) and stored method instructions provide developers with the tools needed to implement secure communication and synchronization in their applications. This promotes the widespread adoption and integration of the technology, driving a more secure and synchronized ecosystem of connected devices. The use of bidirectional communication protocols such as WebSocket further ensures secure data transmission, enhancing the overall security and efficiency of the system.

[0045] According to another aspect of the present invention, a method for maintaining media content synchronization on a user device is provided, comprising receiving a root hash of a hierarchical hash structure associated with the media content. Higher levels of the hash structure include combined hash values ​​from lower levels, ultimately forming the root hash. The system retrieves data from a database associated with the root hash, including content information for media content interaction and context tags for distinguishing known media content (e.g., distinguishing individual songs from the use of the same song in a movie soundtrack). Media content is monitored for a predetermined duration at a first timestamp to obtain a first media segment. Synchronization is verified by calculating the hash of the first media segment and comparing it with the hierarchical hash structure. If synchronization is verified, the hash is used to retrieve and display the content information associated with the first timestamp. If not, the system adjusts the timestamp to resynchronize, or uses a machine learning model to generate a predictive hash to retrieve the content information. Media content is then monitored for a second timestamp within a predetermined duration from the first timestamp to obtain a second media segment overlapping with the first media segment. The hash of the second media segment is calculated to maintain synchronization.

[0046] Benefically, by utilizing a hierarchical hash structure, this method ensures robust and accurate synchronization of media content. It allows for real-time synchronization verification by calculating and comparing the hashes of media segments, improving the accuracy and reliability of content interaction. The ability to adjust timestamps and generate predictive hashes using machine learning models ensures synchronization is maintained even in dynamic or interrupted environments. Furthermore, the method's ability to switch between hash-based synchronization and audio tag detection provides flexibility and resilience, adapting to various media content types and playback conditions. This dual-approach not only improves the user experience by providing seamless and uninterrupted content interaction but also ensures accurate retrieval and display of content information, guaranteeing secure and accurate data synchronization across devices.

[0047] It should be understood that the described apparatus, processes, systems, and methods are not limited to the pairing and synchronization of multimedia content and can be applied to other contexts and use cases. For example, the described technology can be used in secure access control systems where ultrasonic codes are used to authenticate users and grant access to restricted areas. Furthermore, it can be applied to industrial automation to synchronize the operation of different machines or equipment, ensuring precise coordination and timing. The technology can also be applied to smart home environments where various smart devices need to communicate seamlessly and synchronize their actions. Attached Figure Description

[0048] Embodiments of the invention will now be described by way of example only and with reference to the accompanying drawings, wherein: Figure 1 A system architecture diagram for a media interaction system is shown; Figure 2 A system architecture diagram for processing media content on a workstation device is shown; Figure 3A and Figure 3B A system architecture diagram is shown for identifying and synchronizing media content on a user device; Figure 4 A system architecture diagram for identifying media content and retrieving content information is shown; Figure 5 The system architecture diagram for pairing initiation is shown; Figure 6 A system architecture diagram for synchronous output is shown; Figure 7 A system architecture diagram for paired maintenance is shown; Figure 8 A system architecture diagram for identification retrieval is shown; Figure 9 A sequence of user interfaces for initiating content interaction is shown; Figure 10 A user interface sequence for terminating content interaction is shown; Figure 11 A system architecture diagram for the SDK is shown; Figure 12 A flowchart illustrating the method for performing media interaction is shown; Figure 13 A flowchart illustrating a method for preparing content interaction on a workstation platform is shown. Figure 14A and Figure 14B A flowchart illustrating a method for extracting themes from media content on a workstation device; Figure 15 A flowchart illustrating a method for media content identification and content information retrieval on a user device is shown; Figure 16A and Figure 16B A flowchart illustrating a method for content interaction on a user platform is shown; Figure 17 A flowchart illustrating a method for audio track recognition and content information transmission on a server device is shown; Figure 18A , Figure 18B and Figure 18C A flowchart illustrating the method for media content recognition and interaction is shown; Figure 19 A flowchart illustrating a method for establishing secure communication on a user equipment is shown. Figure 20 A flowchart illustrating a method for establishing secure communication in a transport application is shown; Figure 21A and Figure 21B The system architecture diagram of the video interaction system is shown; Figure 22 The system architecture diagram of the dual synchronization system is shown; Figure 23 A flowchart illustrating a method for processing video content is shown; and Figure 24 A flowchart illustrating a method for synchronization based on video hashes is shown; and Figure 25 An example computing environment is shown for performing any of the methods described herein.

[0049] All accompanying drawings are for the purpose of illustrating selected versions of the invention and are not intended to limit the scope of the invention.

[0050] As a preliminary consideration, those skilled in the art will readily understand that this invention has broad applicability and utility. It should be understood that any embodiment may include only one or more of the aspects disclosed above, and may further include only one or more of the features disclosed above. Furthermore, any embodiment discussed and marked as "preferred" is considered part of the best mode contemplated for carrying out embodiments of the invention.

[0051] Other embodiments may also be discussed for the purpose of providing a full and achievable disclosure. Furthermore, many embodiments, such as adaptations, variations, modifications, and equivalent arrangements, will be implicitly disclosed by the embodiments described herein and fall within the scope of the invention. Therefore, although embodiments have been described in detail herein in conjunction with one or more examples, it should be understood that the invention is illustrative and exemplary and has been made solely for the purpose of providing a full and achievable disclosure.

[0052] The detailed disclosure of one or more embodiments herein is not intended, nor should it be construed as, limiting the scope of patent protection conferred by any patent claims granted thereby, which should be defined by the claims and their equivalents. No limitation not explicitly stated in the claims themselves is intended to define the scope of patent protection in any claim. Furthermore, it should be noted that each term used herein refers to its meaning as understood by one of ordinary skill in the art based on its use in the context herein. If the meaning of a term used herein—as understood by one of ordinary skill in the art based on its use in the context—differs from any particular dictionary definition of the term, the meaning as understood by one of ordinary skill in the art shall prevail. Additionally, it should be noted that, as used herein, "a" or "an" generally means "at least one," but does not exclude plural unless the context requires otherwise. When used herein to connect a list of items, "or" means "at least one item," but does not exclude multiple items in the list. Finally, when used herein to connect a list of items, "and" means "all items in the list."

[0053] The following detailed description refers to the accompanying drawings. Where possible, the same reference numerals are used in the drawings and the following description to refer to the same or similar elements. While many embodiments of the invention can be described, modifications, adaptations, and other implementations are possible. For example, elements shown in the drawings may be substituted, added, or modified, and the methods described herein may be modified by replacing, reordering, or adding stages to the disclosed methods. Therefore, the following detailed description does not limit the invention. Rather, the appropriate scope of the invention is defined by the appended claims.

[0054] This invention includes headings. It should be understood that these headings are for reference only and should not be construed as limiting the subject matter disclosed under them. Other technical advantages may become apparent to those skilled in the art upon review of the following drawings and description. It should be understood from the outset that, although exemplary embodiments are shown in the drawings and described below, the principles of the invention can be implemented using any number of techniques, whether currently known or unknown. The invention should in no way be limited to the exemplary implementations and techniques shown in the drawings and described below.

[0055] Unless otherwise stated, the accompanying drawings are intended to be read in conjunction with the specification and are considered an integral part of the entire written description of the invention. As used in the following description, the terms "horizontal," "vertical," "left," "right," "up," "down," etc., and their adjective and adverbial derivatives (e.g., "horizontally," "to the right," "upward," "radially," etc.), refer only to the orientation of the illustrated structure when the particular drawing is facing the reader. Similarly, the terms "inward," "outward," and "radially" generally refer to the orientation of a surface relative to its axis of elongation or axis of rotation (where applicable).

[0056] This invention includes many aspects and features. Furthermore, while many aspects and features are relating to and described in the context of media content interaction or synchronization systems, embodiments of the invention are not limited to use only in that context. In the context of this invention, any system, method, or process disclosed herein includes at least one processing unit, wherein said at least one processing unit performs the method of the invention. Detailed Implementation

[0057] Embodiments of the invention will now be described with reference to the accompanying drawings. It should be noted that the following description is intended only to enable those skilled in the art to understand the invention, and is not intended to limit the applicability of the invention to other embodiments that the reader can readily understand and / or conceive of. In particular, while the invention primarily relates to audio recognition and content recognition systems for enhancing media interaction, those skilled in the art will understand that the methods and systems described herein are more generally applicable to audio signal-based recognition and secure synchronization, as well as related data retrieval.

[0058] One of the advantages of the systems, methods, and techniques described herein is bridging the gap between broadcast media and user interactivity. The embodiments covered, such as those with… Figures 1 to 4Related embodiments improve media content identification and data retrieval, thereby enhancing efficiency, security, and data retrieval reliability compared to known methods. For example, viewers can identify the exact scene of a movie playing on television and access related information, including actors, locations, and products appearing in the scene, such as clothing or accessories. Multiple interactions between the viewer / user and the user device are no longer required, such as the previous need to search for the correct movie, the correct scene, and then further search for related information. This reduction not only improves human-computer interaction, enabling more accurate, time-saving, and secure data retrieval, but also improves the energy efficiency and battery life of user devices.

[0059] Another advantage of the disclosed embodiments is the dual-method approach to content recognition and synchronization. The first method involves embedding unique, preferably inaudible, sound codes as audio tags within the soundtrack of movies and broadcasts. These audio tags, detectable by the user device's microphone, contain information related to the content and its current playback position. If factors hinder the first method, such as environmental factors including background noise, a reliable backup mechanism based on audio attribute analysis exists. This second method ensures the continuity of content recognition and synchronization. The combination of these technologies, the dual-method approach, provides a seamless and rich content interaction experience, offering users instant access to detailed information about what they are watching (e.g., specific scenes and items within them).

[0060] There is a growing demand for more intuitive, efficient, and interactive methods of content discovery and engagement. Today's audiences seek instant recognition and deeper connections to content, craving information beyond basic details, including real-time data, contextual insights, and interactive elements such as product details and scenario-specific facts. Furthermore, as user devices become more efficient, there is a need to reduce human-computer interaction and improve the overall efficiency of data retrieval systems. Moreover, reducing human-computer interaction and improving accuracy and reliability leads to safer and more reliable data retrieval systems, as not only is the retrieved information relevant and, for example, cybersecurity-safe, but the risk of malware or malicious content retrieval is effectively eliminated.

[0061] Furthermore, while this invention primarily addresses the pairing and synchronization of multimedia content, those skilled in the art will understand that the apparatus, processes, methods, and systems described herein are applicable to a wide range of other fields. For example, this technology can be used in secure access control systems where ultrasonic codes are used to authenticate users and grant access to restricted areas. Additionally, it can be applied to industrial automation to synchronize the operation of different machines or devices, ensuring precise coordination and timing. Moreover, this technology can be applied to smart home environments, enabling seamless communication and synchronization between various smart devices.

[0062] The advantages of the systems and methods covered include enhanced security, seamless user experience, real-time synchronization, dynamic connection management, and efficient implementation for developers. Overall, the proposed methods offer significant technical advantages over existing systems in terms of security, user experience, synchronization, and ease of implementation.

[0063] The pairing process is automated through voice code detection, eliminating the need for manual input (such as entering a PIN code or scanning a QR code). This automation simplifies the user experience, making it more intuitive and less error-prone. Users can easily connect their devices without going through complicated setup procedures.

[0064] This method ensures real-time synchronization of output between devices by establishing a bidirectional connection for real-time data exchange. Using timestamps to synchronize output addresses issues related to latency, network jitter, and varying processing times. This results in a more cohesive and immersive user experience, particularly beneficial in applications where audio and video synchronization is crucial, such as mobile applications.

[0065] Furthermore, the ability to maintain or terminate bidirectional connections based on voice code detection, user interaction, or location thresholds provides dynamic connection management. This adaptability ensures continuous secure communication and synchronization even under changing conditions, enhancing the system's reliability and robustness.

[0066] Furthermore, computer-readable media containing software development kits (SDKs) and stored method instructions provide developers with the tools they need to implement secure communication and synchronization in their applications. This facilitates the widespread adoption and integration of the technology, driving a more secure and synchronized ecosystem of connected devices.

[0067] The method presented in this paper significantly improves security by utilizing voice codes, particularly ultrasonic or near-ultrasonic codes, for device pairing and authentication. These voice codes are encoded and decoded on authorized devices and / or using authorized applications or SDKs, reducing the risk of unauthorized access and man-in-the-middle attacks common in traditional methods such as Bluetooth pairing or QR code scanning. The use of bidirectional communication protocols such as WebSocket further ensures secure data transmission.

[0068] Figure 1 A system architecture diagram for a media interaction system is shown. Specifically, Figure 1 A block diagram of a media interaction system 100 according to an exemplary embodiment of the present invention is shown.

[0069] The media interaction system 100 includes: known media content 102, a workstation 104, a user device 106, content information 108, a server 110, a media content source 112, a database 114, an embedded audio track 116, an audio track portion 118, a speaker 120, identification data 122, and a request 124. The media content source 112 includes media content storage 112-A and a media content player 112-B. The audio track portion 118 includes audio signals 118-A and is associated with unidentified media content 118-B. Optionally, the server 110 includes the database 114.

[0070] In one example, the media interaction system 100 is configured to: process known media content 102 at workstation 104; facilitate interaction with the known media content 102 at user device 106; and transmit content information 108 from server 110 to user device 106. Workstation 104 retrieves the known media content 102 from media content source 112 or database 114 and stores the embedded audio track 116 and content information 108 at database 114 or server 110. User device 106 listens to the audio track portion 118 emitted from speaker 120 and uses the audio track portion 118 or retrieves identification data 122 from server 110. Based on the interaction between the user and user device 106, a request 124 is sent to server 110. Server 110 receives request 124 and retrieves the stored content information 108 using identification data 122. Content information 108 is then output to user device 106.

[0071] Workstation 104 is a computing system, device, or software application configured to process known media content 102. Known media content 102 includes any media content associated with audio, such as television programs, movies or films, audiobooks, podcasts, music, video games, social media content, web series, streaming media, or broadcast media. Therefore, known media content 102 includes known audio tracks. Workstation 104 obtains known media content 102 from media content source 112 or database 114.

[0072] Media content source 112 is either media content storage 112-A or media content player 112-B. Media content storage 112-A may optionally be database 114, server 110, user device 106, or another source of known media content 102, such as a device associated with a device or platform external to media interaction system 100, like a device associated with the owner, creator, distributor, or licensor of known media content 102. Media content player 112-B is an optional source of known media content 102 and also a source of unidentified media content 118-B. For example, media content player 112-B is a media broadcaster, streaming provider, or device capable of playing, streaming, or storing media content, such as a device including speaker 120. Optionally, media content source 112 stores known media content 102 in database 114, where known media content 102 can subsequently be retrieved at workstation 104.

[0073] Workstation 104 is configured to generate a timeline, an embedded audio track 116, and content information 108 from known media content 102, as shown in the following text. Figure 2 As shown. Embedded audio track 116 includes one or more audio tags, which are preferably embedded at frequencies detectable at user device 106 but advantageously inaudible to the user of user device 106. The audio tags contain identification data 122, and therefore embedded audio track 116 contains identification data 122. Content information 108 is linked to identification data 122 because content information 108 includes organized information related to one or more topics extracted from known media content 102. Topics are objects, places, people, services, or other data points related to known media content 102. Optionally, workstation 104 includes a machine learning model for identifying and / or classifying objects in known media content 102.

[0074] User equipment 106 is configured to listen to audio track portion 118 emanating from speaker 120, for example by detecting, measuring, and / or collecting audio signals 118-A using a microphone, or by otherwise observing unidentified media content 118-B. For example, user equipment 106 is a mobile phone, software application, personal tablet, or computer, or any other platform or device that communicates or is electronically coupled to a microphone. Unidentified media content 118-B associated with audio track portion 118 can be identified based on identification data 122. Identification data 122 includes a media identifier (or other indicator indicating specific media content, such as known media content 102) and an indicator of a specific time or time interval within that specific media content, such as a point in time on a timeline generated by workstation 104. If user equipment 106 detects an audio tag, identification data 122 is extracted from the audio tag. If user equipment 106 does not detect an audio tag, for example if audio track portion 118 does not embed an audio tag or if there are factors affecting audio tag detection (such as ambient noise), identification data 122 is obtained from server 110.

[0075] For example, server 110 acquires one or more audio features extracted from audio track portion 118, either directly from user device 106 or by acquiring audio track portion 118 from user device 106 and extracting the one or more audio features. An unknown audio fingerprint containing the one or more audio features is then generated and compared with multiple audio fingerprints, for example, stored in database 114. The audio fingerprint most similar to the unknown audio fingerprint among the multiple audio fingerprints is the best match, and the information data associated with the best match is then transmitted to user device 106. Optionally, multiple audio fingerprints are generated by workstation 104, wherein workstation 104 is also configured to generate audio fingerprints for each time point in the timeline and link identification data to each audio fingerprint.

[0076] Server 110 is configured to receive request 124, wherein the request 124 is initiated based on an interaction between the user and user device 106. For example, the user indicates that they wish to receive content information 108 associated with audio track portion 118 at user device 106. Request 124 includes identification data 122, causing server 110 to retrieve the content information 108 associated with identification data 122 from database 114. Optionally, audio signal 118-A is associated with multiple known media contents, such that the interaction between the user and user device 106 includes selecting identification data 122 from the multiple identification data sets.

[0077] Advantageously, user equipment 106 is independent of audio track portion 118, meaning that user equipment 106 does not need to communicate with media content source 112 to identify audio track portion 118 or receive content information associated with audio track portion 118. Furthermore, speaker 120 does not need to emit embedded audio track 116 for user equipment 106 to identify audio track portion 118, because user equipment 106 can identify and obtain content information 108 associated with playback media content that does not contain audio tags generated by workstation 104. Audio track portion 118 can be identified by media interaction system 100 as long as an audio fingerprint associated with it has been generated.

[0078] Another advantage of the media interaction system 100 is the secure and efficient retrieval of content information 108. A single user interaction on user device 106 is sufficient to receive content information 108 related to the audio track portion 118 and originating from a secure and trusted source (i.e., server 110 and / or database 114). This reduces the risk of receiving malware or malicious data on user device 106 and improves efficiency and battery life associated with user device 106. Unlike alternative technologies, user device (a) does not need to link to or communicate with media content player 112-B to receive relevant content information, or (b) does not need to perform numerous search operations to automate such searches, whether through multiple interactions between the user and the user device, or by applying excessive computational resources, such as using complex artificial intelligence algorithms, to automate such searches.

[0079] Figure 2 A system architecture diagram for processing media content on workstation devices is shown. Specifically, Figure 2 A block diagram of a media processing system 200 according to an exemplary embodiment of the present invention is shown.

[0080] The media processing system 200 includes: a workstation 202, known media content 204, audio tracks 206, a server 208, a timeline 210, a first time point 212, audio features 214, audio fingerprints 216, content identifiers 218, embedded audio tracks 220, audio tags 222, themes 224, content files 226, and optionally enriched media files 228.

[0081] In one example, media processing system 200 includes workstation 202 configured to acquire known media content 204, wherein the known media content includes audio track 206. The known media content 204 is acquired from server 208, a database, or other undescribed media content source. A timeline 210 is generated for the known media content 204, the timeline including multiple time points, including a first time point 212. Audio features 214 are extracted from audio track 206 to generate an audio fingerprint 216 linked to identification data containing a media content identifier 218 and the first time point 212. The audio fingerprint 216 is generated at workstation 202, server 208, or other undescribed device. The media processing system is also configured to generate an embedded audio track 220 by embedding an audio tag 222 within audio track 206, wherein the audio tag 222 contains identification data. The embedded audio track 220 is generated at workstation 202 or server 208. The audio fingerprint 216 and the embedded audio track 220 are stored at server 208, workstation 202, or a database (e.g., [database name missing]). Figure 1 The database is located at 114.

[0082] Furthermore, a theme 224 is extracted from the known media content 204 to obtain a content file 226, wherein the content file 226 includes a visual depiction of theme 224-A and thematic text associated with theme 224-A. The rich media file 228 is output to a server 208, a database, or a workstation 202, wherein the rich media file 228 is stored, for example, for access by the server 208 or for retrieval by a user device for the content file 226, or is modified / edited, for example, at the workstation 202. The rich media file 228 includes a timeline 210, the content file 226, and a media content identifier 218.

[0083] Figure 3A and Figure 3B A system architecture diagram is shown for identifying and synchronizing media content on user devices. Specifically, Figure 3A and Figure 3B A block diagram of a content synchronization system 300 according to an exemplary embodiment of the present invention is shown.

[0084] The content synchronization system 300 includes: an audio signal 302, a speaker 304, a first audio tag 306, a second audio tag 308, a user device 310, a user interface 310-A, identification data 312, a server 314, a first theme file 316, a database 318, a first visual description or theme text 320, a second theme file 322, a second visual description or theme text 324, a first audio feature 326, and a second audio feature 328.

[0085] In one example, the content synchronization system 300 is configured to listen to an audio signal 302 emitted from a speaker 304 using a microphone or other audio sensor of a user device 310, wherein a first audio tag 306 and a second audio tag 308 are embedded in the audio signal 302. The user device 310 extracts identification data 312 from the first audio tag 306. When the user interface 310-A ​​receives a request indicator (indicating a request for topic information associated with the audio signal 302), the identification data 312 from the first audio tag 306 is transmitted to a server 314. The server 314 retrieves a first topic file 316, for example from a database 318, and receives it at the user device 310. The first topic file 316 includes a first visual description or topic text 320, which can be displayed at the user device 310. To synchronize the displayed topic information, the audio signal 302 continues to be listened to, and, for example, the user device 310 extracts identification data 312 from the second audio tag 308. When the user interface 310-A ​​receives a request indicator, it transmits the identification data 312 from the second audio tag 308 to the server 314. The server 314 retrieves the second theme file 322 and receives it at the user device 310. The second theme file 322 includes a second visual description or theme text 324, which can be displayed at the user device 310.

[0086] In another example, the content synchronization system 300 is configured to listen to an audio signal 302 emitted from a speaker 304, wherein a first audio tag 306 and a second audio tag 308 are not embedded in the audio signal 302, or cannot be detected using a microphone or other audio sensor of the user device 310. The user device 310 extracts a first audio feature 326 from the audio signal 302 at a first time point, or collects a first portion of the audio signal 302 at a first time point. To identify the audio signal 302, the user device 310 transmits the first audio feature 326 or the first portion to a server 314. The server 314 uses the first audio feature 326 to retrieve identification data 312 and transmits it to the user device 310. When the user interface 310-A ​​receives a request indicator (indicating a request for topic information associated with the audio signal 302), it transmits the identification data 312 to the server 314. The server 314 retrieves a first topic file 316, for example, from a database 318, and receives it at the user device 310. The first theme file 316 includes a first visual description or theme text 320, which can be displayed at the user device 310. Optionally, the user device 310 extracts a second audio feature 328 from the audio signal 302 at a second time point, or collects a second portion of the audio signal 302 at a second time point. To synchronize the display of content information with the audio signal 302, the user device 310 transmits the second audio feature 328 or the second portion to the server 314. The server 314 uses the second audio feature 328 to retrieve identification data 312 and transmits it to the user device 310. When the user interface 310-A ​​receives a request indicator, it transmits the identification data 312 to the server 314. The server 314 retrieves the second theme file 322 and receives it at the user device 310. The second theme file 322 includes a second visual description or theme text 324, which can be displayed at the user device 310.

[0087] Optionally, a lock indicator may be received, as described in method 800 below.

[0088] Figure 4 A system architecture diagram for identifying media content and retrieving content information is shown. Specifically, Figure 4 A block diagram of an audio recognition system 400 according to an exemplary embodiment of the present invention is shown.

[0089] The audio recognition system 400 includes: a server 402, a user device 404, a first part 406, audio features 408, a first audio fingerprint 410, multiple audio fingerprints 412, a database 414, recognition data 416, a request 418, content information 420, and related information 422. The multiple audio fingerprints 412 include a second audio fingerprint 412-A, a third audio fingerprint 412-B, and a best match 412-C. The related information 422 includes related information 422-A associated with the second audio fingerprint 412-A, related information 422-B associated with the third audio fingerprint 412-B, and related information 422-C associated with the best match 412-C.

[0090] In one example, the audio recognition system 400 includes a server 402 configured to acquire a first portion 406 of an audio track or audio features 408 extracted from the first portion 406 from a user device 404. If the first portion 406 is not received, the server 402 extracts the audio features 408 from the first portion 406. A first audio fingerprint 410 is generated from the audio features 408 and compared with a plurality of audio fingerprints 412 stored in a database 414. The best match 412-C is the audio fingerprint among the plurality of audio fingerprints 412 that is most similar to the first audio fingerprint 410. The server 402 retrieves recognition data 416 associated with the best match 412-C and transmits it to the user device 404. If a request 418 for content information 420 is received from the user device 404, the content information 420 is transmitted to the user device 404, wherein the content information 420 is related information 422-C associated with the best match 412-C and the recognition data 416.

[0091] Related information 422 includes information associated with one or more topics related to the first part 406, or specifically, information associated with one or more topics related to the media content associated with the audio track at the time of the first part 406, such as... Figure 3A First subject file 316 or Figure 2 Content file 226.

[0092] Optionally, server 402 or user equipment 404 is configured to adaptively optimize audio processing parameters based on detected media format and transmission quality, thereby enhancing the technology's resilience to signal degradation. Thus, while media format and transmission medium may cause audio signal degradation, audio recognition system 400 is configured to reduce any impact from different media formats and transmission media. Further optionally, audio recognition system 400 is configured to ensure real-time processing of audio data and synchronization between server 402, database 414 (optionally a server database contained within server 402), and user equipment 404, without inefficiently consuming large amounts of computing resources. This is achieved by optimizing data processing workflows and implementing efficient algorithms that allow real-time operation without compromising accuracy or user experience.

[0093] For example, machine learning algorithms can be combined to improve and optimize system performance based on user interaction and feedback. In one example, a machine learning algorithm is used to determine the best match 412-C. For instance, a machine learning algorithm can be used to generate a first audio fingerprint 410 or identify audio tags in the first section 406. In another example, a data caching strategy is employed to reduce server response time and improve the overall speed of content recognition and information retrieval.

[0094] Optionally, the audio recognition system 400 includes a transcription database (not shown in the drawings). The transcription database contains multiple transcribed media contents, allowing word-based content recognition and synchronization. For example, user equipment 404 is configured to transmit spoken or typed queries, such as quoting lines from a movie or describing a scene, to server 402, thereby using the transcription database to identify specific media content.

[0095] In one example, an AI algorithm is integrated at user device 404 and configured to interpret and analyze spoken and / or typed user queries by processing words, phrases, or descriptions provided by the user. The AI ​​algorithm then cross-references these queries with a transcription database via server 402, which also stores metadata, including titles, actor names, subject elements, or any other data relevant to identifying the best match. These trained AI algorithms are updated based on user interactions with user device 404, making them personalized algorithms for the user or user device 404. For example, media content identification or retrieval of relevant information associated with media content is improved based on, for example, individual user preferences or viewing history.

[0096] In another example, the artificial intelligence algorithm learns from user interactions to improve the accuracy and efficiency of recognition and information retrieval. For instance, user device 404 receives information such as whether the best match 412-C is a successful match for media content from the speaker, which is used to improve the artificial intelligence algorithm used to identify the best match from database 414.

[0097] More specifically, one example of an artificial intelligence algorithm is the Support Vector Machine (SVM), and another is the Artificial Neural Network (ANN). An SVM is a supervised learning model used for classification and regression tasks. To identify the best match 412-C within a database 414, the SVM is trained on multiple audio fingerprints that store information related to audio features, such as frequency distribution, rhythmic patterns, or spectral characteristics. The training data consists of input audio fingerprints and corresponding label pairs indicating the most similar audio fingerprint in the database. Once training is complete, the SVM receives the audio fingerprint input and classifies it based on the patterns learned during training to identify the most similar audio fingerprint. An ANN is a deep neural network trained on a dataset similar to an SVM. An ANN typically consists of multiple layers of neurons and is trained using techniques such as backpropagation. Once training is complete, the audio fingerprint is fed into the ANN, passed through its layers, and outputs a prediction based on the learned representation, indicating the most similar audio fingerprint in the database 414. In both cases, the choice to use, for example, an SVM or an ANN, depends on factors such as data complexity, dataset size, and available computational resources. ANNs, especially deep learning models, may offer better performance for tasks involving large datasets and complex patterns, but they also require more computational resources and data for effective training. On the other hand, SVMs are simpler models that work well on smaller datasets and may require less computational overhead. Preferably, the choice of artificial intelligence algorithm is tailored to the available computational resources and training data of the audio recognition system 400.

[0098] Preferably, server 402 and database 414 (optionally integrated with server 402) are cloud-based. Advantageously, such as Figure 1-4 The hybrid use of cloud-based infrastructure and distributed computing resources illustrated improves reliability and scalability. In the server-based cloud example, the choice of AI model involves using a large pre-trained generative algorithm, which is then customized using, for example, the training data described above, to perform any of the methodological steps described in methods 500-1100 below.

[0099] The following text describes in detail Figures 5 to 11This relates to secure communication and output synchronization between user equipment and transmission equipment. According to one embodiment, the user equipment scans content, such as content displayed on the transmission equipment. The transmission equipment communicates with a server to identify a request, and if the server identifies the request, it generates a sound code. This sound code contains encoded text code used to authenticate communication between the user equipment and the transmission equipment. The user equipment detects the sound code, generates decoded text code, and then transmits it to the transmission equipment. If the pre-encoded text code and the decoded text code match, communication is authenticated, and a handshake is performed. The handshake establishes a bidirectional connection between the transmission equipment and the user equipment for real-time data exchange, thereby enabling efficient and secure inter-device communication for content interaction. This bidirectional connection is then used for output synchronization, thereby improving user experience, increasing accuracy, reducing latency, improving reliability, improving system security, improving resource efficiency, and enhancing versatility.

[0100] Synchronizing output between user and transport applications ensures seamless content display across multiple devices, eliminating latency and discrepancies for a more cohesive viewing experience. Synchronization using timestamps enables high precision, particularly beneficial for live streaming, gaming, and interactive multimedia. Real-time data exchange via WebSocket minimizes latency, ensuring instant feedback of actions across devices. Regular heartbeat messages and error correction mechanisms enhance stability and reliability, even in noisy environments. Pairing using ultrasonic or near-ultrasonic codes adds a layer of security, while secure WebSocket connections using TLS / SSL protocols protect data transmission. Ultrasonic or near-ultrasonic codes for device pairing and authentication significantly improve security by reducing the risk of unauthorized access and man-in-the-middle attacks, risks common in traditional methods such as Bluetooth pairing or QR code scanning. Furthermore, ultrasonic or near-ultrasonic frequencies, or other frequencies used for implementation, represent frequencies with less interference, enabling more accurate detection and decoding. Efficient resource management is achieved by maintaining bidirectional connections only when necessary, reducing unnecessary data transmission and power consumption. The methods and systems described are versatile and applicable to various fields such as secure access control, industrial automation, and smart home environments, and their applicability is not limited to the systems described herein.

[0101] Figure 5 A system architecture diagram for pairing initiation is shown. Specifically, Figure 5 An implementation of system 500 is shown, wherein system 500 is configured to pair one or more user equipments with one or more transmission devices. These devices are securely paired using bidirectional connections that allow for real-time data exchange and synchronization.

[0102] System 500 is designed for secure, fast, and preferably simultaneous communication for content synchronization (e.g., output on a content delivery device) and content interaction (e.g., output on a user interaction device). Content is preferably media content, such as visual media content (e.g., movies, television programs, images, or live streams) and / or audio media content (e.g., music and podcasts). However, System 500 is not limited to media content and can be used as an alternative to content and content interaction synchronization and communication, such as games, simulations, social and communication interfaces, forms, surveys, exams, machine instructions, autonomous vehicle data, or other content not disclosed herein.

[0103] To ensure connection security, authentication and pairing are performed between devices. Advantageously, using bidirectional connectivity for device pairing also allows for rapid two-way communication. Furthermore, pairing is only completed after device communication has been authenticated, i.e., a handshake is performed. Once paired, bidirectional connectivity allows for improved synchronization of content and content interaction between devices.

[0104] System 500 includes: user equipment 502, transmission device 504, user interface 506, speaker 508, content interface 510, first request 512, pairing request 514, server 516, identifier 518, first voice code 520, and identification code 522. Optionally, user equipment 502 includes or runs a user application, and those skilled in the art will understand that all references to the methods, processes, and steps performed by user equipment 502 are executed or initiated via the user application. Alternatively, or additionally, transmission device 504 includes or runs a transmission application, and those skilled in the art will understand that all references to the methods, processes, and steps performed by transmission device 504 are executed or initiated via the transmission application.

[0105] System 500 pairs user equipment 502 with transmission device 504 for output synchronization between user interface 506 and content interface 510. User equipment 502 scans for content via broadcasting a first request 512 to enable user-content interaction. Transmission device 504 then receives the first request 512 and initiates the transmission of a first voice code 520 from a speaker 508 associated with transmission device 504. Using a microphone, user equipment 502 detects the first voice code 520 and converts it into an identification code 522. The identification code 522 is transmitted to transmission device 504 for authentication; if authentication is successful, a handshake is performed. A bidirectional connection (not shown) is established between user equipment 502 and transmission device 504, allowing real-time data exchange between the devices.

[0106] Alternatively or additionally, transmission device 504 receives first request 512 and transmits a pairing request 514 associated with first request 512 to server 516. In response to receiving pairing request 514, server transmits identifier 518 to transmission device 504. Identifier 518 is converted (e.g., encoded) into first voice code 520 and then emitted from speaker 508. Transmission device 504 receives identification code 522 from user equipment 502 (e.g., decoded from first voice code 520 at user equipment 502). Identification code 522 is authenticated using identifier 518, and if authentication is successful, a handshake is initiated between user equipment 502 and transmission device 504.

[0107] The scanned content includes broadcasting a first request from user equipment 502 using a two-way communication protocol. User equipment 502 is any device configured for user interaction with content, such as a mobile device, tablet device, portable computing device, wearable technology, or other device that includes a microphone, user interface, and two-way communication protocol capabilities. Optionally, as described above, user equipment 502 includes or runs a user application, such as an application or application that performs or initiates any functions described regarding user equipment 502. For example, a user application running on user equipment 502 initiates a content scan, causing user equipment 502 to broadcast the first request 512.

[0108] User interface 506 is configured for user interaction with user device 502, including navigation, user input, user selection, and / or execution of user commands. In an example where user device 502 is a mobile device, user interface 506 is a screen, such as a touchscreen or a display configured for user interaction, and output is displayed on user interface 506. Examples of user interaction with user interface 506 include any of the following: touching or viewing the output displayed on the touchscreen; clicking (e.g., using a pointer) a button, link, or other interactive element; and / or instructing user device 502 to perform actions related to the output on user interface 506 via voice commands or actions (e.g., using a microphone or motion detection sensors, such as a camera or infrared sensor).

[0109] The bidirectional communication protocol includes technologies and standards that enable user device 502 to communicate with transmission device 504, such as WebSocket. Advantageously, WebSocket-based communication protocols provide a full-duplex channel over a single, long-lived TCP connection. Unlike traditional HTTP (where user device 502 continuously polls transmission device 504 for updates and vice versa), WebSocket allows real-time and near real-time communication between clients and servers. Therefore, WebSocket is particularly useful in scenarios involving real-time media control or interactive content where low-latency communication is beneficial. For example, a user application running on user device 502 sends commands (e.g., play, pause, volume up / down) to the paired transmission device 504, and transmission device 504 sends updates (e.g., current playback position, timestamp, buffer status) to the paired user device 502 using WebSocket. Optionally, communication from user device 502 and communication from transmission device 504 are simultaneous. Using WebSocket not only minimizes latency or lag, but also provides a more stable and secure connection than traditional communication protocols.

[0110] Alternatives to WebSocket include: HTTP / 2; Extensible Message Processing Context Protocol (XMPP); Advanced Message Queuing Protocol (AMQP); Message Queuing Telemetry Transport (MQTT); for example, in bandwidth-constrained situations; Real-Time Streaming Protocol (RTSP) or Real-Time Transport Protocol (RTP); Remote Procedure Call (gRPC); Web Real-Time Communication (WebRTC); CoAP (Co-Restricted Application Protocol); Session Initiation Protocol (SIP); Bluetooth Low Energy (BLE) and General Attribute Protocol (GATT); various protocols or technologies that facilitate bidirectional communication between user equipment and media / transmission equipment, preferably enabling real-time control, feedback, and interaction; or combinations thereof.

[0111] Transmission device 504 is a media device or a device capable of displaying or otherwise playing, transmitting, or emitting content, such as media content. For example, a transmission device is any of the following: a television device running a transmission application (such as Netflix™); a computing device; a media broadcasting device; a micro-console, such as Chromecast™ or Apple TV™; or a device not listed. The transmission device is associated with a content interface 510 (such as a display screen) and a speaker 508, which is preferably configured to emit sound frequencies detectable by a microphone of user device 502 but less perceptible to the user, such as ultrasonic or near-ultrasonic frequencies, for example, audio tags generated by workstation 104. Optionally, transmission device 504 includes and / or runs a transmission application, such as an application configured to initiate or perform functions described regarding transmission device 504.

[0112] Transmission device 504 is configured to transmit content (e.g., via a transmission application) and respond to user equipment 502 scanning content. When user equipment 502 scans content, transmission device 504 receives or detects a first request 512 broadcast from user equipment 502. In an example where the first request 512 is broadcast using WebSocket, transmission device 504 has WebSocket access or functionality, such as transmission device 504 running a transmission application that supports WebSocket access. Preferably, the bidirectional communication protocol used by user equipment 502 and the bidirectional communication protocol used by transmission device 504 are configured for secure communication, such as through the authentication and pairing steps described herein.

[0113] Advantageously, communication security is further ensured by using voice codes (e.g., first voice code 520). Transmission device 504 is configured to convert identifier 518 into first voice code 520, for example, by converting the alphanumeric data of identifier 518 into sound using a predetermined encoding method. The frequency range of the first voice code 520 and any additional voice codes is based on any one or a combination of: detectability at a microphone, e.g., detectability by an industry-standard microphone associated with the user equipment; detectability of the user, e.g., the user is less detectable; the operating range of a loudspeaker, e.g., the accuracy and capability of an industry-standard loudspeaker to emit the first voice code; or limiting interference, e.g., avoiding sound frequencies associated with content, user conversations, traffic, or any other standard noise sources that may affect the detectability of the first voice code 520 and subsequent decoding. Ultrasonic or near-ultrasonic frequencies are preferred.

[0114] User equipment 502 is configured to convert a first audio code 520 into an identification code 522, for example, by using a predetermined decoding method to convert the first audio code 520 into the identification code 522 or by otherwise extracting the identification code 522 from the first audio code 520. Optional predetermined methods include dual-tone multi-frequency (DTMF), Morse code, and / or preferably frequency shift keying (FSK).

[0115] FSK is a modulation scheme in which digital data (such as alphanumeric characters) is encoded by changing the frequency of a carrier wave. Different frequencies represent different binary values ​​(such as 0 and 1). These binary values ​​can then be mapped to alphanumeric characters using a predefined encoding scheme (such as American Standard Code for Information Interchange (ASCII)). An FSK demodulator analyzes the frequency variations in the received audio signal, converting the frequencies back into binary data. This binary data is then converted back into alphanumeric characters.

[0116] User equipment 502 is further configured to interact directly with server 516. For example, output from the user equipment is based on information or data retrieved from or otherwise received from server 516. Preferably, this output is associated with content transmitted from transmission device 504, received from server 516, and synchronized to the content using the methods and systems disclosed herein.

[0117] Both users and content providers are seeking content identification and complex interactions, requiring information beyond basic details, including real-time data, contextual insights, and interactive elements such as product details and scenario-specific facts. User equipment 502 becomes more efficient with reduced human-computer interaction, while the efficiency of communication and content delivery systems is also improved. Furthermore, reduced human-computer interaction and improved accuracy and reliability lead to a safer and more reliable data retrieval system, as not only is the retrieved information relevant and cybersecurity-secure, but the risk of malware or malicious content retrieval is also effectively eliminated. Optionally, user equipment 502 uses a first voice code 520 to identify the output of transmission device 504.

[0118] For example, user equipment 502 executes a method for media content identification and content information retrieval. Initially, the user equipment's microphone listens to a portion of an audio track emitted from speaker 508, where the audio track is associated with specific media content. Then, user equipment 502 detects whether audio tags, such as a first sound code 520, are embedded in that portion of the audio track. These audio tags are embedded at frequencies detectable by the microphone but less noticeable to the user, for example, using the same configuration as the first sound 520. Upon detection of an audio tag, user equipment 502 extracts identification data associated with known media content at a specific point in time. This identification data is linked to the known media content at that point in time; for example, the first sound code 520 includes an identification code 522 and a message timestamp. Alternatively, if no audio tag is detected, the system employs an alternative method: it receives identification data from multiple audio fingerprint databases via server 516 based on the best match of audio fingerprints. These fingerprints consist of one or more audio features extracted from the listened-to audio track portion. Once the identification data is obtained, user equipment 502 identifies the media content based on the known media content associated with the identification data. Subsequently, user equipment 502 retrieves content information linked to the identification data.

[0119] Retrieving content information involves transmitting an information request, including identification data, to a server device (e.g., server 516). Server 516 then responds with information on one or more topics related to known media content at a specified point in time. Each topic is linked to a known point in time, ensuring that the retrieved information is contextually accurate and relevant to the media content being identified. This method provides user devices with a seamless and efficient way to identify media content and access related information in real time.

[0120] The above-described alternative embodiments improve media content identification and data acquisition, thereby enhancing efficiency, security, and data retrieval reliability compared to known methods. For example, a user can identify the exact frame of a movie playing on a television and access related information, including actors, locations, and products appearing in the scene, such as clothing or accessories. Multiple interactions between the viewer / user and user device 502 are no longer required, such as the multiple interactions previously needed to search for the correct movie, the correct scene, and then further search for related information. This reduction not only improves human-computer interaction, enabling more accurate, time-saving, and secure data retrieval, but also improves the energy efficiency and battery life of user device 502. When combined with synchronization performed using the disclosed bidirectional connection, efficiency, speed, system security, and user experience are further enhanced.

[0121] Only devices and / or applications configured with conversion (e.g., related encoding / decoding) capabilities can be paired, preventing unauthorized or malicious devices from compromising communications, devices, or applications. For example, transmission device 504 transmits a pairing request 514 to server 516 for device / user authorization. Device / user authorization may optionally include any of the following: identification of user device 502; identification of the user of user device 502; identification of transmission device 504; identification of the user of transmission device 504; or a combination thereof. Pairing request 514 includes data associated with the first request 502, such as transmission ID, user ID, geographic location, or other identifying information. By identifying pairing request 514, server 516 determines whether pairing request 514 is associated with one or more trusted devices and / or users. Server 516 generates an identifier 518 in response to identifying (and thus authenticating) pairing request 514, or retrieves the identifier from a database, such as... Figure 8 As shown.

[0122] In another example, a software development kit (SDK) is used to provide a framework that allows transport devices and user devices to pair, such as SDK 1100 described below. The SDK for secure bidirectional communication (e.g., WebSocket communication) includes components and / or functions necessary for the devices to perform the processes and methods described herein, enhancing the security of interactions between authorized devices. Transport device 504 sends a first request 512 to the SDK and / or a transport application that utilizes, embeds, or otherwise associates with the SDK. The SDK generates a pairing request 514 and communicates the pairing request 514 to server 516. Server 516 recognizes the pairing request and sends an identifier 518 to the SDK. This allows only secure / protected interaction with server 516, helping to isolate the server from malicious devices or attacks.

[0123] Further, optionally, secure communication can be further enhanced using any of the following: encrypted communication, such as using Transport Layer Security (TLS) to encrypt identifier 518, identification code 522, and / or bidirectional communication; firewall integration, such as authorizing only pairing requests 518 to be received by server 516; or IP filtering, allowing only requests from known IP addresses to establish bidirectional connections. Continuing with the SDK example, SDK components and mechanisms ensure that only authorized devices can communicate via bidirectional connections.

[0124] Advantageously, the use of the SDK or associated transport application allows a variety of smart and non-smart transport devices to operate the methods and processes described herein. For example, even if such functionality may not be securely implemented in other ways, an internet-connected television, monitor, or screen can securely communicate and establish a bidirectional connection with the user device using a bidirectional communication protocol via the SDK or transport application.

[0125] In another example, system 500 includes multiple devices. For example, there are multiple user devices 502, including a first user device and a second user device, and a transmission device 504. Optionally, both the first and second user devices are independently paired and synchronized with the output of transmission device 504 (e.g., displayed media content). The output of the first user device (e.g., first information associated with the displayed media content) is synchronized to be updated and adapted based on the output of transmission device 504. Similarly, the output of the second user device (e.g., second information associated with the displayed media content) is also synchronized to be updated and adapted based on the output of transmission device 504. The first and second information contain the same or different information, such as data based on user selection, data associated with a user of either user device, or another factor affecting the first information relative to the second information. Preferably, both the first and second information are associated with the output of transmission device 504.

[0126] In another scenario, there are multiple transmission devices 504 and one user equipment 502. User equipment 502 scans content and responds by receiving multiple sound codes, for example, one for each transmission device 504. User equipment 502 pairs with a selected transmission device 504 based on, for example, user selection or a selection protocol. A preferred selection protocol depends on any of the following: a proximity threshold, such as using geographic location or sound code volume measurement; data associated with historical user selections, user preferences, or user controls; or the priority of the transmission device output. For example, if a favorable alarm associated with a danger to the user is output from transmission device 504, user equipment 502 will pair with that transmission device 504 even if music is playing from a different device, and updated danger-related information will be output at user equipment 502.

[0127] Optionally, System 500 is a system for pairing and synchronizing a user device (e.g., a mobile device) with a transmission device using a bidirectional communication protocol (e.g., WebSocket) and sound (e.g., ultrasound). System 500 includes a transmission device and a mobile device. The transmission device is configured to: receive a pairing request from the mobile device via WebSocket; send a pairing request to a server and receive an identification code from the server; convert the identification code from text to an ultrasound or near-ultrasound code; send the ultrasound or near-ultrasound code to the mobile device; receive a text code from the mobile device via WebSocket; perform a handshake with the mobile device if the text code matches; and periodically send ultrasound or near-ultrasound signals to ensure the mobile device remains near the transmission device. The mobile device is configured to: initiate a content scan using WebSocket; receive ultrasound or near-ultrasound codes via a microphone; convert the received ultrasound or near-ultrasound codes to text codes; send the text code to the transmission device via WebSocket; perform a handshake with the transmission device if the text code matches; synchronize with the transmission device via WebSocket; and periodically receive ultrasound or near-ultrasound signals to ensure the mobile device remains near the transmission device and prompt the user for confirmation of presence if a signal is missed.

[0128] Figure 6 A system architecture diagram for synchronous output is shown. Specifically, Figure 6 System 600 is illustrated, wherein system 600 is configured to synchronize content between paired devices. For example, system 600 includes system 500, which allows paired devices to synchronize content.

[0129] System 600 includes: user equipment 502, transmission device 504, user interface 506, speaker 508, content interface 510, first sound code 520, first timestamp 602, second timestamp 604, first output 606 and second output 608.

[0130] According to one embodiment, user equipment 502 synchronizes a first output 606 at user interface 506 with a second output 608 at content interface 510 of transmission device 504. For example, user equipment 502 identifies a first timestamp 602 associated with the first output 606, for example, by determining the current timestamp of the first output 606. User equipment 502 then identifies a second timestamp associated with the second output, for example, by extracting the timestamp encoded in the first sound code 520. The second timestamp 604 corresponds to the timestamp of the second output when the first sound code 520 was generated and / or transmitted. Alternatively, the second timestamp 604 is received from transmission device 504 using a bidirectional connection. The difference between the first timestamp 602 and the second timestamp 604 is then reduced, thereby improving the synchronization between the first output 606 and the second output 608. If the first timestamp 602 is earlier than the second timestamp 604, the first output 606 is delayed (e.g., slowed down) to correspond temporally to the second output 608. If the first timestamp 602 is later than the second timestamp 604, then the first output 606 is advanced (e.g., sped up) to correspond in time to the second output 608.

[0131] According to another embodiment, the transmission device 504 synchronizes the second output 608 with the first output 606. The transmission device 504 receives the first timestamp 606 from the user equipment 502 using a bidirectional connection and identifies the second timestamp 608. It then reduces the difference between the first timestamp 602 and the second timestamp 604, thereby improving the synchronization between the first output 604 and the second output 608. If the first timestamp 602 is earlier than the second timestamp 604, the second output 608 is advanced (e.g., sped up) to correspond in time to the first output 606. If the first timestamp 602 is later than the second timestamp 604, the first output 606 is delayed (e.g., slowed down) to correspond in time to the second output 608.

[0132] Optionally, user equipment 502 is further configured for sound-based content synchronization. Using a microphone, user equipment 502 listens to a portion of an audio track emitted from a speaker to detect embedded audio tags. The audio tags are encoded or otherwise associated with a second timestamp 603. These tags are embedded at frequencies detectable by the microphone but less noticeable to the user, for example, using the same configuration as the first sound 520. Upon detection of an audio tag, user equipment 502 extracts identification data associated with known media content at the second timestamp 604. If no audio tag is detected, user equipment 502 identifies the second timestamp 604 via an audio fingerprinting technique using a server (e.g., server 516) by matching audio features against a known fingerprint database. This dual-method approach ensures accurate timestamp identification even in the absence of embedded tags. Once the second timestamp 604 is identified, user equipment 502 modifies a first output 606 such that the first timestamp 602 matches the second timestamp 604.

[0133] Figure 7 A system architecture diagram for paired maintenance is shown. Specifically, Figure 7 System 700 is shown, wherein system 700 is configured to maintain or terminate device pairing. For example, system 700 includes system 500, wherein paired devices maintain or deactivate the connection as needed.

[0134] System 700 includes: user equipment 502, transmission device 504, user interface 506, speaker 508, bidirectional connection 702, second sound code 704, maintenance request 706, first maintenance user selection 708, maintenance instruction 710 and termination instruction 712.

[0135] After user equipment 502 pairs with transmission equipment 504, it sends a second voice code 704. The transmission interval of the second voice code 704 can be a random interval, a consistent interval, or other predetermined interval, and the transmission interval should be after the first voice code 520 is sent. The second voice code 704 ensures that bidirectional connection and / or synchronous output between user equipment 502 and transmission equipment 504 are still required. If user equipment 502 detects the second voice code 704 emitted by transmission equipment 504, then user equipment 502 is still in the vicinity of transmission equipment 504. If user equipment 502 does not detect the second voice code 704, then user equipment 502 may no longer be in the vicinity.

[0136] If a second voice code 704 is detected at user equipment 502 within a predetermined interval, the bidirectional connection is maintained. Maintaining a bidirectional connection, such as a full-duplex connection established via WebSocket, may optionally include one or more processes and mechanisms to ensure that the connection remains stable, secure, and functional over time. Examples include, but are not limited to: connection maintenance, i.e., keep-alive mechanisms, such as ping-pong frames or heartbeat messages; data framing and segmentation; message processing and routing; error handling and recovery, such as detecting invalid frames or unexpected disconnections; flow control and congestion management; session state management and state update mechanisms; or other processes for maintaining a stable and secure bidirectional connection.

[0137] Bidirectional communication is terminated or closed if the following conditions are met: no second voice code 704 is detected at user equipment 502 within a predetermined interval; the geographical location associated with user equipment 502 exceeds a predetermined threshold, such as 10 meters; the interaction indication between the user and user equipment 502 no longer requires output synchronization; a termination threshold is exceeded, such as one hour or less of no user interaction at user equipment 502 or transmission device 504; or a combination of the above. Termination of a bidirectional connection, such as a full-duplex connection established via WebSocket, may optionally include normal closure, such as user equipment 502 sending a close frame to transmission device 504, or forced closure, such as user equipment 502 or transmission device 504 forcibly closing the connection. A timeout mechanism helps prevent stale connections from unnecessarily consuming resources.

[0138] According to one embodiment, a second sound code 704 is emitted from speaker 508. If user equipment 502 detects the code, the bidirectional connection 702 is maintained. If user equipment 502 does not detect the code, a maintenance request 706, such as the question “Are you still watching?”, is displayed on user interface 506. A first maintenance user selection 708 is associated with a user interaction in response to the maintenance request 706 instructing the maintenance of the bidirectional connection to be maintained or terminated. If the first maintenance user selection 708 instructs the maintenance of the bidirectional connection, a maintenance instruction 710 initiates maintenance of the bidirectional connection 702, and optionally, any time-related thresholds used to determine whether to maintain or terminate the bidirectional connection 702 are reset. If the first maintenance user selection 708 instructs the termination of the bidirectional connection, a termination instruction 712 initiates the termination of the bidirectional connection 702. Optionally, a second request 714 is broadcast from user equipment 502 to scan for updated content.

[0139] The second sound code 704 may optionally be the first sound code 520 repeated at predetermined intervals. Alternatively, the second sound code 520 may be a separate sound code associated with updated data. For example, the first sound code 520 may optionally include a first timestamp associated with the output of the transmission device 504 when the first sound code 520 is generated. The second sound code 704 includes a second timestamp associated with the output of the transmission device 504 when the second sound code 704 is generated. If the output has stopped, paused, or otherwise remained unchanged in time, the second timestamp remains unchanged relative to the first timestamp, and the second sound code 704 may optionally be the first sound code 520. If the output has changed, for example, advanced or regressed in time, the second timestamp differs from the first timestamp, and the second sound code 704 differs from the first sound code 704. In another example, the second sound code 704 includes data associated with a pairing between the user equipment 502 and the transmission device 704.

[0140] Figure 8 A system architecture diagram for identification retrieval is shown. Specifically, Figure 8 System 800 is shown, wherein system 800 is configured to recognize pairing authentication requests. For example, system 800 includes system 500, which allows associated devices to pair securely.

[0141] System 800 includes: transmission device 504, first request 512, first request data 512-A, pairing request 514, server 516, identifier 518, user equipment identifier 518-A, database 802, trusted file 804, and transmission application 806.

[0142] As discussed above regarding systems 500, 600, and 700, server 516 receives a pairing request 514 that includes data associated with the first request 512 (e.g., first request data 512-A, such as geographic location, timing data, device identifier, application identifier, or user identifier associated with a device and / or application). The first request data 512-A may optionally be used by the server to search database 802. Database 802 stores trusted files 804 associated with trusted and / or authorized devices and / or users, such as whitelists. The identifier associated with user device 502, i.e., user device identifier 518-A, is then retrieved from the associated trusted file 804. Alternatively, the server generates user device identifier 518-A. Optionally, server 516 includes database 802.

[0143] Further optionally, the server receives pairing request 514 and generates identifier 518 after determining that pairing request 514 meets a predetermined security protocol. Database 802 is optionally not required. Example security protocols include: comparing first request data 512-A with a whitelist to determine whether the first request data 512-A exceeds a time- or location-related threshold, or counting the number of pairing requests 514 received within a predetermined time range.

[0144] Preferably, server 516 communicates with a transmission application 806 running on transmission device 504. Transmission application 806 is a content-embedded or service-embedded SDK that prevents insecure communication between the device and server 516. For example, SDK / transmission application 806 is configured to receive a first request 512, generate and transmit a pairing request 514 to server 516, receive or retrieve an identifier 518 from server 516, convert the identifier 518 into a first sound code 520, and optionally instruct or otherwise activate speaker 508 to emit the first sound code 520, such as... Figure 11 As stated above.

[0145] Those skilled in the art will understand that transmission application 806 is applicable to systems 500, 600, and 700, wherein any or all functions of transmission device 504 are performed by transmission application 806. Similarly, any or all functions of user equipment may optionally be performed by user application (not shown here).

[0146] Figure 9 A sequence of user interfaces for initiating content interaction is shown. Specifically, Figure 9 A sequence of possible interactions initiated at user equipment 504 for searching and / or interacting with the content of transmission device 504 is shown. Any user interface disclosed is an option for user interface 502 of system 500, system 600, or system 700.

[0147] The first user interface sequence 900 includes: a display 902; a search user interface 506-A, including initiating user selection 904 and applying user selection 906; a confirmation user interface 506-B, including confirming user selection 908 and re-initiating user selection 910; and a pairing user interface 506-C, including first pairing content 912 and content interaction user selection 914.

[0148] Optionally, the user interface 506 includes a display 902. Alternatively, the display 902 is optional, and the functions of the display 902 described herein are performed by the user interface 506.

[0149] The first user interface sequence 900 is associated with the user device 502 using the search user interface 506-A to search for content (e.g., initiating step 802 of method 800 below), using the confirmation user interface 506-B to select desired content from one or more detected contents (e.g., generated by steps 804 and 806, and / or initiating step 808 of method 800), and using the pairing user interface 506-C to interact with output synchronized to the selected content (e.g., generated by steps 810 and 812).

[0150] Search user interface 506-A includes initiating user selection 904. Interaction between the user of user interface 506-A and initiating user selection 904 instructs the user to confirm the content scanned by user device 502, for example, initiating method 800. Alternatively, the user of user interface 506-A can choose not to initiate content scanning by not interacting with initiating user selection 904, for example, by interacting with application user selection 906 to initiate alternative functions of user device 502 or a user application running on user device 502.

[0151] After initiating a content scan, the user interface 506-B provides one or more confirmation user selections 908. Each confirmation user selection 908 is associated with an identification code 522 extracted from the first voice code 520, and each confirmation user selection 908 is associated with a different identification code and a different voice code. Interaction between the user of user interface 506-B and the confirmation user selections 908 instructs the user to confirm pairing with the associated transmission device 504 and / or synchronize the output of user interface 506 with associated content (e.g., output 608). Optionally, the user of user interface 506-B interacts with a re-initiating user selection 910 to initiate a content scan. For example, a re-initiating user selection 910 is an initiating user selection 904.

[0152] After confirming the associated content and / or transmission device 504, user equipment 502 and transmission device 504 pair, for example, by performing a handshake and establishing a bidirectional connection. The pairing user interface 506-C includes first pairing content 912 and optionally content interaction user selection 914. The first pairing content 912 and / or pairing user interface 506-C may optionally be an output 606, such as synchronized to output 608. The first pairing content 912 is associated with an output on the paired transmission device. One or more content interaction user selections 914 are associated with the first pairing content 912, wherein user interaction with the content interaction user selection 914 instructs the user to interact with the first pairing content 912 and / or content outputs (e.g., output 608) from the paired transmission device.

[0153] Example methods for a user to interact with any user interface described herein include touching a portion of a touchscreen, using voice commands, using a pointer to click a button or portion of the displayed output, or utilizing other user-machine interaction methods.

[0154] Figure 10 A user interface sequence for terminating content interaction is shown. Specifically, Figure 10 A sequence of possible interactive terminations performed at user equipment 504 is shown for maintaining or terminating / closing the bidirectional connection between user equipment 504 and transmission equipment 506. Terminating user interface 506-D is an option for user interface 506 of system 500, system 600, system 700, or the first user interface sequence 900.

[0155] The second user interface sequence 1000 includes: an optional display 902; a pairing user interface 506-C, including first pairing content 912 and content interaction user selection 914; a termination user interface 506-D, including a termination message 1002, maintaining user selection 1004, and terminating user selection 1006; and a search user interface 506-A, including initiating user selection 904 and applying user selection 906.

[0156] The second user interface sequence 1000 is associated with the user of user device 502 using the pairing user interface 506-C to interact with content, using the termination user interface 506-D to terminate the synchronized content interaction session, and optionally using the search user interface 506-A to search for updated content for interaction.

[0157] The termination user interface 506-D includes a termination message 1002, notifying the user that the conditions for terminating the pairing session have been met. For example, the user device 502 may not have detected the second voice code 704 and / or a timing or location threshold has been exceeded. In another example, the bidirectional connection may be compromised, insecure, or require termination and / or active maintenance, and the termination message 1002 will notify the user.

[0158] Optionally, the termination message 1002 is associated with a termination time threshold. After a predetermined time interval measured from the generation and / or display of the termination message 1002, the pairing session is closed, and the bidirectional connection terminates. For example, no user interaction with the user interface 506-D is required to terminate the bidirectional connection. This timeout mechanism improves the security of the bidirectional connection and enhances efficiency by preserving computing and energy resources.

[0159] Optionally, or additionally, termination message 1002 is associated with a user confirmation request. The user interacts with the maintain user selection 1004, indicating that the pairing session continues and the bidirectional connection is maintained. The user interacts with the terminate user selection 1006, indicating that the pairing session is closed and the bidirectional connection is terminated.

[0160] Figure 11 A system architecture diagram for the SDK is shown. Specifically, Figure 11 SDK 1100 is shown, an SDK for secure communication and output synchronization. For example, SDK 1100 is used by transport application 806 and / or operates on transport device 504, allowing transport device 504 to securely pair with one or more user devices 502.

[0161] SDK 1100 includes: communication module 1102, authentication module 1104, and interaction module 1106.

[0162] The communication module is configured to provide communication between the transmission device 504 and the user equipment 502 using a bidirectional communication protocol (e.g., WebSocket), and to establish a bidirectional connection 702, such as a full-duplex connection, between the transmission device 504 and the user equipment 502. The communication module is further configured to communicate with the server 516. For example, the communication module provides functions for: receiving a first request 512 and an identification code 522 from the user equipment 502; transmitting a pairing request 514 to the server and receiving an identifier 518 from the server; and transmitting a first sound code 520 from the speaker 508. Optionally, the communication module is configured to establish a secure WebSocket connection (WSS) using the TLS / SSL protocol.

[0163] The authentication module 1104 is configured to authenticate user equipment 502, thus improving security when establishing bidirectional communication. For example, the authentication module 1104 provides functions for generating a pairing request 514, generating a first voice code 520, and comparing an identification code 522 with an identifier 518 to authenticate user equipment 502 before initiating a pairing session by establishing a bidirectional connection through the communication module 1102.

[0164] Optionally, generating the first sound code also includes using error correction, forward error correction, or Reed-Solomon encoding of the current timestamp. For example, the identifier is encoded in an acoustic signal (such as an ultrasonic or near-ultrasonic signal) using Reed-Solomon error correction encoding, where the data is converted into a sequence of sound symbols with built-in redundancy to facilitate error detection and correction during transmission, broadcasting, playback, or decoding.

[0165] Interaction module 1106 is configured to help synchronize the output of transmission device 504 (e.g., output 608, such as content provided by a content service provider) with the output of user device (e.g., output 606, such as pairing user interface 506-C). For example, interaction module 1106 provides functions for: determining and / or identifying the current timestamp associated with the output from transmission device 504; identifying the message timestamp received from user device 502 via communication module 1102, optionally extracted at authentication module 1104; calculating the difference between the message timestamp and the current timestamp; and initiating adjustments to the output of transmission device 504 (e.g., speeding up or slowing down content) to minimize the difference; or initiating adjustments to user device 502 (e.g., instructing user device 502 to speed up or slow down content) to minimize the difference.

[0166] According to one option, SDK 1100 also includes a service module (not depicted) configured to embed services (e.g., content service provisioning) into SDK 1100. Additionally, or alternatively, SDK 1100 may optionally include an error correction module (not shown) configured to detect and correct errors associated with communication module 1102, authentication module 1104, or interaction module 1106, such as insecure bidirectional connections or errors in generating the first sound code, leading to pairing failure.

[0167] According to another option, SDK 1100 also includes a user interface module (not shown in the figure) configured to display data associated with any module disclosed herein to the user of transmission device 504 or paired user device 502.

[0168] According to another option, SDK 1100 also includes a storage module (not shown) that is further configured to store data received via communication module 1102 and / or processed by authentication module 1104 or interaction module 1106.

[0169] Any of the optional modules described above can be combined or otherwise integrated into SDK 1100. Furthermore, any functionality provided by SDK 1100 can form dedicated or composite modules. For example, SDK 1100 is used to pair and synchronize a transmitting device with a user device (e.g., a mobile device) using a bidirectional communication protocol (e.g., WebSocket) and sound (e.g., ultrasound). SDK 1100 includes: a module for receiving a pairing request from the mobile device via WebSocket; a module for sending a pairing request to a server and receiving an identification code from the server; a module for converting the identification code from text to ultrasound or near-ultrasound codes; a module for sending ultrasound or near-ultrasound codes to the mobile device; a module for receiving a text code from the mobile device via WebSocket; a module for performing a handshake with the mobile device if the text code matches; and a module for periodically sending ultrasound or near-ultrasound signals to ensure the mobile device remains near the transmitting device.

[0170] In another example, authentication module 1104 includes a generation module configured to generate a pairing request 514 and a first voice code 520, and a comparison module configured to compare an identification code 522 with an identifier 518. Optionally, any combination of the above modules includes one or more additional modules, or combinations forming a composite module. For example, authentication module 1104 and interaction module 1106 form a single module.

[0171] Those skilled in the art will understand that SDK 1100 can be used for transport application 806, and further applicable to systems 500, 600 and 700, wherein any or all functions of transport device 504 are performed using SDK 1100.

[0172] In the context of this invention, Figure 1-4 and Figure 5-11Several features described herein may be identical or have similar functions. Specifically, user equipment 502 may be identical to user equipment 106, providing a platform for user interaction with content. Server 516 may be identical to server 110, facilitating the storage and retrieval of content information and managing pairing requests. Speaker 508 may be identical to speaker 120, emitting audio tracks or sound codes for content identification and pairing. Database 802 may be identical to database 114, storing media content, related information, and trusted device data. Transmission device 504 may be identical to media content source 112, serving as the source of media content and facilitating content transmission. Identification code 522 may be identical to identification data 122, used to identify media content and user equipment. First request 512 may be identical to request 124, initiating a content identification or pairing process. First output 606 may be identical to content information 108, providing information related to media content. User interface 506 may be identical to user interface 310-A, enabling user interaction with devices and content.

[0173] It should be understood that the reference numerals used in this document are not intended to be limiting. Figure 1-4 and Figure 5-11 The elements of the features described herein may optionally be based on shared functionalities. This means that those skilled in the art will readily understand that combining elements from different figures is within the scope of this invention. For example, from Figure 5 User equipment 502 can optionally connect with from Figure 1 The media interaction system 100 is combined. Specifically, user equipment 502 can be used to interact with known media content 102, wherein user equipment 502 detects audio track portion 118 emitted from speaker 508 (which may optionally be speaker 120) and retrieves content information 108 from server 516 (which may optionally be server 110). After, simultaneously with or before, content retrieval based on, for example, system 100, user equipment 502 synchronizes with transmission devices 504 / 112, thereby improving the efficiency and security of communication, data retrieval, and synchronization between devices. In another example, from Figure 5 The transmission device 504 can communicate with devices from... Figure 3A The content synchronization system 300 is combined. Specifically, the transmission device 504 can emit an audio signal 302 from a speaker 508, wherein the user equipment 310 (which may optionally be user equipment 502) detects a first audio tag 306 and a second audio tag 308, and synchronizes the first output 606 with the second output 608 based on the timestamps associated with the audio tags. These examples illustrate... Figure 1-4 and Figure 5-11 The features and elements described herein are interchangeable and can be combined in various ways to achieve the functions described herein.

[0174] Figure 12A flowchart illustrating a method for performing media interaction is shown. Specifically, method 1200 is... Figure 1 The method of executing components of the media interaction system 100.

[0175] Method 1200 includes: step 1202, processing known media content on a workstation device; step 1204, interacting with the media content on a user device; and step 1206, transmitting content information to the user device.

[0176] Step 1202 includes processing known media content on the workstation device. Known media content is media content that is identified or known to the workstation device or media content source, such that the known media content is associated with a unique content identifier. A content identifier is a name, code, or other form of ID that allows the known media content to be identified or otherwise distinguished from other media content. Examples of known media content are movies, television programs, radio broadcasts, podcasts, or other broadcast, streaming, or recorded media.

[0177] A workstation device is a computing system or platform capable of importing or otherwise acquiring known media content and outputting or otherwise storing modified media content. For example, a workstation device is a computer communicatively coupled to a server or database via an internet connection. In another example, a workstation device is... Figure 1 Workstation 104 is configured to process known media content 102. (As mentioned above...) Figure 1 The term "known media content" includes any media content associated with audio, such as television programs, movies or films, audiobooks, podcasts, music, video games, social media content, web series, streaming media, or broadcast media. This allows for audio-based processing on workstation devices, audio-based recognition on server or user devices, and audio-based synchronization on user devices.

[0178] Step 1204 includes interacting with media content emitted from a media playback device (e.g., a speaker) at a user device (e.g., a personal electronic device). The interaction includes identifying the media content and receiving information associated with the media content, such as from a server, database, or workstation device. If the media content being interacted with is modified media content, such as modified media content that has been processed, stored, and then accessed and played on the media playback device, the media content is identifiable on the user device. If the media content being interacted with is unmodified media content, such as media content that has been processed in step 1202 but the original unmodified media content is being played on the media playback device, the media content is identifiable based on comparing it with stored known media content (e.g., the known media content from step 1202). In cases where multiple known media content pieces are similar, such as the same audio track used for multiple movies, the user device may optionally be configured to determine the best-matching known media content, such as based on user selection.

[0179] Step 1206 involves transmitting content information to the user device. After the user device initiates a request for content information, the server or workstation device transmits information related to the media content identified in step 1204. The content information includes topics extracted from the media content, such as locations, people, services, items, objects, quotations, food and beverages, or other identifiable content related to the media content being played, streamed, broadcast, or otherwise emitted. The content information includes a visual depiction of the extracted topics, topic text (such as the name or description of the topics), links (such as hyperlinks for further user-interactive association with the topics), or some combination thereof.

[0180] For example, the topic is a voting option related to a television competition being broadcast to a media playback device. The content includes a link to a website hosting the voting, allowing users on the user's device to interact with and vote. Users do not need to input data or repeatedly interact with their device to access the voting, and only relevant votes are presented to them. This not only improves efficiency in terms of time and human-computer interaction but also requires less energy and computing resources, thus extending the user's device battery life. There is also no opportunity to introduce inaccuracies during this information retrieval process because the user and / or the user's device do not perform extensive searches. Furthermore, security is enhanced because the user's device does not need to access multiple resources (e.g., multiple web pages) or repeatedly interact with search engines or external databases. This reduces the likelihood of the user's device accessing malware. Additionally, fake versions of the vote or other malicious content (which may appear related to the television competition) will not be accessed. Therefore, the user's device is less likely to be subjected to phishing attacks or other cybersecurity threats.

[0181] In another example, the media content is a streaming movie, and the subject is the filming location of the current scene in the streaming movie. Content information includes information about the location, information on getting to the location, nearby hotels, or other location-based data. Retrieval of this type of content information can optionally be monitored, providing accessible metrics that can be analyzed on servers, workstations, or output to external systems. Benefits include market research, trend analysis, demand forecasting, user experience optimization and improvement, traffic management, and environmental protection. For example, conservation organizations can use location lookup data to identify ecologically important areas or areas with high levels of human activity. This information can inform conservation efforts and help mitigate environmental impacts. Overall, monitoring location lookups can provide valuable insights across various fields, facilitating better decision-making and resource allocation.

[0182] In another example, the media content is an emergency weather broadcast, with the subject being the name of a storm. The content information transmitted to the user's device includes safety information related to the storm's impact on the user's device location. In another example, the media content is an emergency cybersecurity broadcast, with content information including patches to prevent malicious viruses from damaging or otherwise negatively impacting the user's device. In yet another example, the media content is a surgical record, with the subject being the patient, and content information including diagnostic and technical information related to the surgical procedure, such as the patient's heart rate and respiratory rate as the surgery progressed. Optionally, using a machine learning model to perform subject extraction and / or content information generation allows for near real-time content information transmission, such as synchronized with a live broadcast. These non-limiting examples lead to targeted and efficient data dissemination, as well as accurate and efficient data retrieval.

[0183] Figure 13 A flowchart illustrating a method for preparing content interaction on a workstation platform is shown. Specifically, method 1300 is... Figure 2 A method performed by a component of the media processing system 200.

[0184] Method 1300 includes: step 1302, importing media content into a workstation platform; step 1304, generating a timeline associated with the media content; step 1306, obtaining one or more topic files; step 1308, assigning each topic file to at least one time point; and step 1310, outputting rich media files.

[0185] Step 1302 includes importing media content into the workstation platform. The media content is associated with multiple topics and a media content identifier. Importing media content includes obtaining media content, such as from a server, database, or other media content source, downloading media content, or creating / generating media content. Multiple topics include those related to the media content, as described above regarding method 1200, and the media content identifier is used to distinguish the media content from other known media content. Preferably, the media content identifier is unique for each imported media content.

[0186] For example, the workstation platform is media processing system 200, or more specifically, workstation 202, the media content is known media content 204, and the media content identifier is... Figure 2 The content identifier is 218. Once imported, the media content can be processed, for example, as described in step 1202 of method 1200. If the media content is to be processed, method 1300 proceeds to step 1304. If the media content is not processed or is not fully processed, for example, if previous media content associated with the same media content identifier has been at least partially processed previously, method 1300 proceeds to any of steps 1304-610 to prevent unnecessary duplication and improve efficiency. In one example, the imported media content is already associated with a timeline (step 1304) and multiple topics (step 1306). However, new topics are associated with the media content, causing steps 1306-610 to be repeated as needed.

[0187] Step 1304 includes generating a timeline associated with the media content. The timeline comprises multiple points in time, each representing a predetermined time interval. Therefore, the timeline is generated by segmenting the imported media content into predetermined time intervals. The points in the timeline can then be used to synchronize content information associated with the media content to its relevant time; for example, technical information related to a vehicle is linked to the time when the vehicle appears in the media content.

[0188] For example, workstation 202 generates and Figure 2The timeline 210 is associated with the media content 204. If the timeline 210 has already been generated, for example, if the media content 204 has been at least partially processed previously, step 1304 includes acquiring the timeline 210 and evaluating it to determine if additional processing is required. Optionally, generating the timeline includes determining predetermined time intervals for multiple time points. If the media content contains relatively homogeneous content, longer time points result in fewer time points, requiring less processing at workstation 202, and the resulting rich media file 228 is smaller, requiring fewer computational resources. However, if the media content contains relatively heterogeneous content over time, shorter time points allow for more accurate synchronization of topic-related content information to the time when that topic is associated with the media content. The predetermined length of the time points may also affect the speed and / or efficiency of media content identification.

[0189] Step 1306 includes acquiring one or more topic files at a workstation platform. Each acquired topic file is associated with one of a plurality of topics and includes a visual description of the topic and topic text. The topic files are acquired from: a database storing the visual descriptions and / or topic text of the topics, such as stored topic files associated with topics of previously processed media content; a workstation, for example, generating topic files using a machine learning model or from a knowledge base associated with the media content source; or an external source, such as searching and retrieving information associated with the topics. Thus, acquiring one or more topic files may optionally include generating a topic file associated with one of a plurality of topics.

[0190] For example, a visual description is an image associated with a subject, such as a picture or video of the subject extracted from media content or from different sources. Visual descriptions can be displayed on a user device, such as on a user interface and / or the user device's display screen. Subject text is a name, description, or information associated with a subject and can also be displayed as text data on a user device, such as on a user interface and / or the user device's display screen.

[0191] Optionally, acquiring one or more topic files includes using artificial intelligence algorithms to: combine visual descriptions and topic text; retrieve or generate visual descriptions or topic text; or provide topic-related data. For example, a generative algorithm system containing one or more large generative models is used to generate topic text or organize information about the topic based on visual description input, such as for manual verification and / or editing.

[0192] Optionally, step 1306 includes editing, removing, adding, or correcting relevant data, visual descriptions, or topic text. For example, if one or more topic files retrieved from the database are associated with previous media content, the information may be outdated or less relevant to the media content imported in step 1302. In another example, if one or more topic files contain data at least partially generated by an artificial intelligence algorithm, the information may be inaccurate or otherwise inconsistent with the media content. Editing or other forms of correction or verification increase the relevance of the information associated with the topic files, thereby improving content interaction.

[0193] Step 1308 includes assigning each topic file to at least one point in time on the timeline. The topic file is linked to a point in time where the topic is relevant to the media content, such as when the topic is displayed, heard, or otherwise referenced in the media content. Optionally, the topic file may be linked to a point in time where the topic is not directly displayed, heard, or otherwise referenced but is indirectly relevant to the media content at that point in time.

[0194] For example, if a watch is visible at a first time point but not at a second time point yet remains relevant—for example, being discussed in media content or becoming visible again at a third time point—then the relevant topic file is assigned to the first time point and optionally to the second. However, if the watch is not displayed or is otherwise relevant at a fourth time point, the relevant topic file is not assigned to the fourth time point. This ensures that content information associated with the watch is displayed on the user's device when the watch is relevant, and prevents irrelevant information from being displayed when the watch is no longer relevant. This targeted data retrieval reduces the amount of data transferred from the server and retrieved on the user's device, thus reducing the memory and storage requirements of the user's device to display relevant information.

[0195] Step 1310 includes outputting a rich media file, such as to a storage medium on a database, server, user device, or workstation device. The rich media file includes a timeline generated in step 1304, one or more topic files associated with the timeline obtained in step 1306, and a media content identifier for linking the rich media file to media content obtained in step 1302. Optionally, outputting the rich media file includes storing it in a database or otherwise downloading it from a workstation platform. For example, importing media content into the workstation platform in step 1302 may optionally include uploading the previously downloaded rich media file output in step 1310. Therefore, one or more steps of method 1300 may be repeated as needed.

[0196] For example, rich media files are Figure 2The rich media file 228 includes a timeline 210, a theme file 226, and a content identifier 218. The rich media file can be accessed on a workstation device or a separate workstation platform, for example, for further editing or modification, such as for editing the timeline or assigning attached or edited theme files to the timeline, or for optimizing the processing of attached media content that shares at least one of multiple themes.

[0197] Optionally, the media content imported in step 1302 includes an audio track. Method 1300 also includes embedding an audio tag within a first time point of the audio track, wherein the audio tag includes identification data associated with a media content identifier and the first time point. Embedding the audio tag generates an embedded audio track, such as step 1410 of method 1400. Further optionally, the enriched media file also includes the embedded audio track, and any topic files assigned to the first time point in step 1308 are associated with the audio tag. This allows for the synchronization of audio-based identification and topic information retrieval. For example, based on the audio tag detected on the user device, at least one topic file assigned to the first time point is provided to the user device.

[0198] Further optionally, the media content imported in step 1302 is associated with a set of media content. This set of media content is associated with a unique group identifier for each set. For example, the set of media content might be individual episodes of a TV series, generated by the same media content creator, acquired or owned by the same media content distributor, or otherwise linked. Enriched media files may also optionally include group identifiers to make acquiring one or more subject files in step 1306 more efficient.

[0199] Figure 14A and Figure 14B A flowchart illustrating a method for extracting themes from media content on a workstation device is shown. Specifically, method 1400 is... Figure 2 Media processing system 200 and / or Figure 4 Another method is performed by the components of the audio recognition system 400.

[0200] Method 1400 includes: step 1402, acquiring known media content including a known audio track; step 1404, generating a timeline for the known media content; step 1406, generating a first audio fingerprint; step 1408, linking identification data to the first audio fingerprint; step 1410, generating an embedded audio track associated with the identification data encoded within the known audio track; step 1412, extracting one or more themes from the known media content; step 1414, generating content information for each theme; and step 1416, storing the first audio fingerprint, the embedded audio track, and the content information.

[0201] Step 1402 includes acquiring known media content at a server or workstation platform. The known media content includes known audio tracks and is associated with one or more themes related to the known media content. The known audio tracks include audio signals and are audio in any form, such as recordings, songs, speech, or other audio data detectable from a standard microphone, such as the microphone of a personal computing device or mobile phone.

[0202] For example, step 1402 includes step 1302 of method 1300 or is included within step 1202 of method 1200. Optionally, the known media content is Figure 2 Known media content 204 or Figure 1 Known media content 102. For example, known media content includes video content, interactive content, streaming media content, broadcast content, user-generated content, or other forms of media. In another example, obtaining known media content includes obtaining it from a server device (e.g., Figure 4 The server 402) or media source (e.g. Figure 1 (112) Import known media content from media content source 1.

[0203] Step 1404 includes generating a timeline for the known media content. The timeline includes multiple points in time, and each point in time is a predetermined time interval, such that the timeline is divided into predetermined time intervals. These segments can be used to synchronize the topic and associated content information to a specific time within the known media content.

[0204] For example, step 1404 includes step 1304 of method 1300 or is included within step 1202 of method 1200. Optionally, the timeline is... Figure 2 Timeline 210. In another example, the total duration of each point in time (e.g., each predetermined time interval) together is at least the duration of the known media content. Alternatively, the known media content comprises multiple timelines, such as the timeline for each episode of a TV series, or the timeline for each known audio track when the known media content comprises multiple known audio tracks, such as when a movie is dubbed into different languages.

[0205] Step 1406 includes extracting one or more audio features from a known audio track and generating an audio fingerprint using the one or more audio features. The audio fingerprint includes the one or more audio features, wherein the audio features are characteristics of the audio signal, such as amplitude, frequency, timbre, duration, phase, envelope, harmonics, noise, dynamics, spatialization, resonance, distortion, pitch, phase coherence, transients, or other features that define and / or distinguish different audio signals.

[0206] For example, extracting one or more audio features involves analyzing the audio signal of a known audio track to identify or quantify selected audio features. Analyzing the audio signal includes performing any of the following: feature extraction techniques including Fourier transform, Mel-frequency cepstral coefficients, chromaticity features, spectral centroid, zero-crossing rate, root-mean-square energy, or other standard techniques; statistical analysis, temporal analysis, feature scaling, or visualization analysis; or one or more machine learning techniques, such as using the Mel-frequency cepstral coefficients (MFCC) method.

[0207] MFCC is a widely used feature extraction method in speech and audio processing tasks such as speech recognition, speaker recognition, and music genre classification. For example, a typical MFCC method includes preprocessing, Mel frequency transformation, cepstral analysis, feature selection, normalization, and feature vector output. Preprocessing: Audio signals are typically preprocessed by applying techniques such as windowing (e.g., using a Hamming window) and framing to divide them into small overlapping segments. Mel Frequency Transformation: The power spectrum for each frame is calculated using techniques such as Fourier transform. The power spectrum is then transformed to a Mel frequency scale, a perceptually meaningful scale that better aligns with human auditory perception. Cepstral Analysis: After transformation to a Mel frequency scale, the logarithm of the power spectrum is taken. This logarithmic spectrum is then transformed using a discrete cosine transform (DCT) to obtain cepstral coefficients. These coefficients capture information about the spectral envelope of the audio signal. Feature Selection: Typically, not all cepstral coefficients are used for further analysis. Usually, a subset of coefficients is selected based on their importance or relevance to the task at hand. Normalization: Optionally, selected cepstral coefficients can be normalized to remain invariant to changes in overall volume or intensity. Feature Vector: The final output is a feature vector representing the audio characteristics of the input signal. This feature vector can then be fed into machine learning algorithms for tasks such as classification, clustering, or regression. MFCCs effectively capture valuable features of audio signals, such as timbre information and spectral characteristics, in a compact and discriminative representation. Due to their robustness and efficiency in representing audio data, they have been successfully applied to a variety of audio processing tasks.

[0208] Generating an audio fingerprint by extracting one or more audio features from a known audio track includes generating a first audio fingerprint, which comprises one or more audio features extracted from the known audio track at a first time point. For example, one or more features are extracted from a known audio track at a first time point on a timeline, and these features are used to generate the first audio fingerprint. Since the known audio track in this example is non-homogeneous and changes from the first time point to a second time point, the one or more audio features extracted from the known audio track at the second time point will differ from the one or more audio features extracted at the first time point. The second audio fingerprint generated using the one or more audio features extracted at the second time point is therefore distinguishable from the first audio fingerprint. The known audio track and the first time point can be identified based on the first audio fingerprint, and the known audio track and the second time point can be identified based on the second audio fingerprint. In other words, the first audio fingerprint is used to identify a known audio track and associated known media content at a specific time point, as described above.

[0209] For example, one or more audio features extracted from a known audio track are hashed or otherwise converted into an encoded identifier that forms an audio fingerprint. This encoded identifier will vary depending on the one or more audio features extracted from the known audio track and can be compared with other encoded identifiers to determine audio fingerprints containing similar or identical extracted audio features. Encoding the one or more audio features (e.g., using hashing techniques or other encryption / encoding techniques to store information about the one or more audio features) into an audio fingerprint allows for the identification of similar or identical audio tracks based on similar or identical audio fingerprints, even if the original audio tracks are at different volumes, affected by different levels of noise, played using different audio codec technologies or with different audio quality, or affected by other different influencing factors.

[0210] Hashing the one or more audio features may optionally involve combining various features extracted from the audio signal into a single hash value. Example methods for hashing multiple audio features include feature extraction, normalization, concatenation, hashing, optional salting, and output. Feature Extraction: Extract one or more audio features from the audio signal, such as in step 1412. Optional features include MFCC (Melbourne Frequency Cepstral Coefficients), spectral features (e.g., spectral centroid, spectral bandwidth), rhythmic features (e.g., tempo, beat), or other relevant features such as zero-crossing rate, energy, or pitch. Normalization: Normalize the extracted features to ensure consistency in scale and range. Normalization techniques (such as minimum-maximum scaling or Z-score normalization) may be applied independently or collectively to each feature. Concatenation: Concatenate the normalized feature vectors into a single feature vector. This creates a uniform representation of all extracted features. Hashing: Apply a hash function to the concatenated feature vectors to generate a fixed-length hash value. Common hash algorithms (such as SHA-1, SHA-256, or MD5) may be used for this purpose, although those skilled in the art will understand that other algorithms may also be used. Salting: Optionally, a salt value is incorporated into the hashing process to enhance security and prevent pre-computed hash attacks. The salt value is a random string added to the input before hashing. This increases the security for future identification and synchronization. Output: The output of the hashing process is a hash value that uniquely represents a combination of one or more audio features extracted from the audio signal. The generated hash values ​​can be efficiently stored or compared to identify similar audio content without storing the original audio data. Furthermore, the choice of hashing algorithm and parameters can optionally vary depending on the specific requirements of the application, including considerations for security, efficiency, and collision resistance. For example, the choice of hashing algorithm and parameters is customized according to the media interaction system 100.

[0211] Step 1408 includes linking the identification data to an audio fingerprint. The identification data is associated with known media content at a relevant time point. For example, the identification data linked to the first audio fingerprint in step 1408 includes a content identifier of the known media content and a first time point. Therefore, the identification data identifies both the known media content associated with the audio track and the relevant time point corresponding to the linked audio fingerprint on the timeline, thereby allowing audio-based content identification and synchronization.

[0212] For example, a song being broadcast to television includes one or more audio features. The song is monitored by a mobile phone's microphone, where the mobile phone and television are independent devices. The one or more features are extracted to form a television audio fingerprint. The television audio fingerprint can then be compared with a first audio fingerprint generated in step 1406. If the first audio fingerprint and the television audio fingerprint are the same, identification data linked to the first audio fingerprint is used to identify the television audio fingerprint, thereby identifying the song. The identification data may optionally be received by the mobile phone, allowing the song to be identified. Further optionally, one or more topic files associated with the song at a first time point (see method 1300 above) are received on the user device for synchronously displaying relevant content information.

[0213] Step 1410 includes embedding an audio tag within the known audio track obtained in step 1402, thereby generating an embedded audio track. The audio tag includes the identification data from step 1408, such that the embedded audio tag encodes the identification data within the known audio track. At a first time point in the timeline generated in step 1404, a first audio tag is embedded within the known audio track, and the first audio tag includes a content identifier of the known media content and a first time point, thereby allowing audio-based content identification and synchronization.

[0214] Preferably, the audio tag can be detected by the microphone of the user equipment, but is less detectable to the user of the user equipment or workstation equipment. For example, the audio tag is embedded at frequencies inaudible to the human ear, such as using frequency shift keying modulation over a defined time interval, wherein the defined time interval is designed to be beyond human detection range under standard environmental conditions.

[0215] Frequency Shift Keying (FSK) modulation is a digital modulation technique used in telecommunications and digital communication systems, involving modulating the frequency of a carrier signal based on a digital input signal. The digital input signal consists of binary data, where each binary symbol (bit) represents a specific digital value, typically 0 or 1. A carrier signal with a fixed frequency is generated. This carrier signal is then modulated at different frequencies according to the digital input signal. In FSK modulation, different frequencies are assigned to represent different digital values. Typically, one frequency is used to represent a binary "1," and another frequency is used to represent a binary "0." Based on the digital input signal, the frequency of the carrier signal is modulated to switch between the assigned frequencies. When the input signal is logic high (e.g., binary 1), the carrier frequency shifts to the first frequency. Conversely, when the input signal is logic low (e.g., binary 0), the carrier frequency shifts to the second frequency. The modulated signal (now carrying information encoded as a frequency shift) is transmitted through the communication channel. At the receiving end, the modulated signal is received and demodulated to recover the original digital input signal. This process involves detecting the carrier frequency shift and mapping it back to the corresponding digital value. FSK modulation has several advantages, including simplicity, noise immunity, and bandwidth utilization efficiency. To overcome potential limitations such as frequency offset and nonlinear distortion in the transmission channel, variants of FSK, such as Continuous Phase Frequency Shift Keying (CPFSK) and Minimum Frequency Shift Keying (MSK), can be used. Alternative techniques such as CPFSK and MSK apply the basic principles of frequency modulation based on digital input signals.

[0216] The first audio tag in step 1410 and the first audio fingerprints in steps 1406 and 1408 can both be used to identify known audio tracks. The first audio tag in step 1410 allows for the direct extraction of identification data associated with known media content at a first time point, resulting in faster identification of known media content. Furthermore, it is not important whether the known audio track is associated with multiple known media contents, multiple time points, or multiple known media contents at multiple time points. The encoded identification data includes a content identifier for the relevant known media, which differs from other content identifiers, for example, from common or mutual audio tracks.

[0217] For audio tags to be detected, the embedded audio track must be played and detectable, which is not always feasible for older media content, such as those without redistributed soundtracks or with low-bitrate audio file formats. For example, streaming content does not always have a stable or high-bandwidth internet connection, which may result in the emitted sound not containing detectable embedded audio tags. Therefore, the audio fingerprints in steps 1406 and 1408 provide a backup mechanism when audio tags are undetectable or not detected. While generating an audio fingerprint and comparing it with other audio fingerprints linked to the identification data takes longer than extracting identification data directly from embedded audio tags, and may require storing multiple audio fingerprints for comparison, it allows older media content and lower-quality audio tracks to be identified, regardless of what media content is being monitored (content identifier) ​​or where in the media content is being monitored (time point).

[0218] Step 1412 involves extracting one or more themes from the known media content. For example, themes are extracted and labeled or otherwise categorized at one or more points in time in the timeline generated in step 1404. The one or more themes are items, places, people, services, or other objects that are associated with the known media content at any one or more points in time. For example, the theme extracted at the first point in time is a singer of a known audio track associated with the known media content, an object worn by a character within the known media content at the first point in time, a location where the known media content is created at the first point in time, or an actor heard at the first point in time.

[0219] Optionally, extraction is performed using one or more artificial intelligence algorithms, such as machine learning models like artificial neural networks or computer vision models. For example, a computer vision model is trained on image data to identify and classify objects in an image. One or more object tracking algorithms are used to analyze objects to track how many consecutive frames or time points a given object is part of, and then linked to a timeline. Furthermore, identified objects may optionally be linked to other objects identified at different time points, resulting in the merging of the same objects at different time points.

[0220] Step 1414 includes organizing information associated with each of the one or more topics extracted in step 1412. This organized information is used to generate content information associated with the identification data, which can be provided to the user device based on the identification data. For example, steps 1412 and 1414 include step 1308 of method 1300, and the information associated with the extracted topics is... Figure 2 The topic file 226. The combination of topic files related to, for example, a first point in time is organized information associated with one or more topics extracted from known media content at that first point in time. Therefore, the generated content information represents information related to known media content at the first point in time.

[0221] Optionally, step 1414 is performed using one or more artificial intelligence algorithms. For example, the object identified in step 1412 is enriched with metadata through, for example, a multimodal generative artificial intelligence algorithm, where the extracted object image (a visual depiction of the subject) is used as input, for example, for more detailed identification or retrieval of information associated with the object. The source of information associated with the object can be controlled or otherwise monitored to achieve, for example, secure and accurate information retrieval from known authoritative sources, or it can be generated, for example, using a large language model or other generative algorithms. The object, along with its associated image and text, may optionally be stored to allow for future identification of existing objects and for training or updating machine learning models.

[0222] Further optionally, verification of object recognition and linking time points on a server or workstation device leads to improved accuracy and relevance in extracting the one or more topics. Verification may optionally include linking objects, adding additional objects, modifying associated metadata, selecting an image of the object, or adding precision to the recognized object. Advantageously, combining multiple artificial intelligence algorithms, such as the classifier in step 1412 and the generative model in step 1414, results in an automated system for extracting topics and / or generating content information. Optionally, the artificial intelligence algorithms are hosted externally, on a server, and / or communicate with the server to reduce the computational resources required by the workstation and improve efficiency.

[0223] Step 1416 includes storing audio fingerprints, identification data, embedding audio tracks, and content information. For example, step 1416 includes... Figure 2 The rich media files 228 are stored on a server or in a database. Identification data can be retrieved based on audio fingerprints or embedded audio tracks (e.g., from a server, workstation, or user device). This allows for audio-based media content identification and synchronization. Content information can be retrieved based on the identification data, enabling the retrieval of relevant information, both in terms of contextual relevance and temporal relevance.

[0224] Optionally, content information associated with additional time points can also be retrieved based on identification data. For example, content information associated with a first time point and a second time point can be retrieved based on identification data associated with the first time point, wherein the second time point is within an adjustable buffer period. The buffer period can be adjusted on a workstation, server, or user device and is based on user selection, customization according to known media content, or on physical or computational limitations of the device used to perform, for example, method 1400, such as the speed of data communication between devices. Alternatively or additionally, the adjustable buffer period can be extended or adjusted based on whether a lock indicator is received, as detailed in method 1600 below.

[0225] In another example, the server will look up content information associated with a point in time within an adjustable buffer period and return all objects associated with the object and point in time within the adjustable buffer period. Optionally, the user device will receive sequentially collected content information, such as a visual depiction of the topic, for example, based on a user's request or selection of content information. Further audio-based monitoring of known media content may optionally be used to keep the sequential collection synchronized with known media content. Further optionally, based on user interaction with the user device, the user device receives and / or displays additional information, such as topic text or additional content information generated in step 1414, on the user device's user interface, such as a display screen. In one example, interaction with the visual depiction of the topic (e.g., clicking or tapping the visual depiction of the topic) results in a detailed view of the topic with different information, as well as possible future interactions based on the topic type.

[0226] Optionally, step 1416 includes storing the first audio fingerprint, embedded audio track, and content information in a database. This database is configured to retrieve identification data based on the audio fingerprint or embedded audio track and / or to retrieve content information based on the identification data. Alternatively, the database is configured for a server to retrieve identification data or content information, wherein the server is connected to the database. Preferably, the workstation device performing steps 1400 communicates with a server configured to access the database. In another example, the workstation device communicates directly with the database, or the server performs steps 1400.

[0227] Optionally, the steps of method 1400 are repeated. For example, step 1406 is repeated for a second audio tag. The second audio tag includes identification data associated with known media content at a second time point, wherein the second time point is one of a plurality of time points within the timeline generated in step 1404. Step 1410 includes embedding the second audio tag within a known audio track such that the embedded audio track includes both the first and second audio tags.

[0228] Figure 15 A flowchart illustrating a method for media content identification and content information retrieval on a user device is shown. Specifically, method 1500 is... Figure 3A and Figure 3B This is one method executed by the components of the content synchronization system 300.

[0229] Method 1500 includes: step 1502, monitoring a portion of an audio track associated with media content; step 1504, determining whether an audio tag is embedded in the portion; step 1506, extracting identification data from the audio tag; step 1508, retrieving the identification data based on one or more extracted audio features; step 1510, identifying media content based on the identification data; and step 1512, retrieving content information associated with the identification data.

[0230] Step 1502 includes monitoring a portion of an audio track emitted from a speaker using a microphone from the user equipment, wherein the audio track is associated with media content. The microphone is attached to or communicatively coupled to the user equipment and is configured to collect audio signals. Monitoring a portion of the audio track includes collecting audio signals emitted from a speaker, wherein the speaker is attached to or communicatively coupled to a media playback device. Preferably, the media playback device and the user equipment are independent devices.

[0231] For example, the media content is a movie streaming on a television device, and the movie soundtrack, including an audio track, is emitted from the television device's speakers. The mobile device's microphone monitors a portion of the audio track, such as a section or interval. This section can be any duration between microseconds and the total duration of the audio track, and is preferably between 1 and 10 seconds, such as 2 seconds.

[0232] Step 1504 includes detecting whether an audio tag is embedded within the portion of the audio track. The audio tag is embedded at a frequency detectable by a microphone, such as a frequency detectable by a standard built-in microphone of a standard mobile or personal computing device. Preferably, the audio tag is embedded at a frequency less detectable to the user of the user device, such as a frequency at which humans cannot distinguish the audio tag from the audio track and / or ambient noise, or otherwise at frequencies inaudible to standard humans.

[0233] Returning to the example of step 1502 above, within the audio signal monitored in step 1502, an audio tag is detected by the microphone of the mobile phone device. Optionally, the detected audio tag is embedded using frequency shift keying modulation at defined time intervals. These time intervals may optionally be customized based on the audio track, associated media content, or meet physical or computational limitations associated with standard microphone, speaker, or human audio detection. Since an audio tag has been detected, method 1500 proceeds to step 1506. Otherwise, method 1500 proceeds to step 1508.

[0234] Optionally, step 1504 includes applying one or more correction algorithms to compensate for noise, interference, or other audio degradation factors. This results in more accurate and efficient audio tag detection. Optionally, or additionally, the detection of the presence of an audio tag is limited to a predefined time range, such as between 1 and 10 seconds, for example, 3 seconds, or depends on the estimated time between audio tags embedded within the audio track, or on the estimated time between time points. For example, if two audio tags should have been detected within the predefined time range but were not, method 1500 proceeds to step 1506.

[0235] Step 1506 includes extracting identification data from the audio tag if an audio tag is detected in step 1504. The identification data is associated with known media content at a known point in time. For example, the identification data includes a content identifier (e.g., ...). Figure 2 The content identifier (218) and time point (e.g., time point 212). Therefore, the audio tag is the encoded identification data within the audio tag. An audio track with an embedded audio tag, i.e., an embedded audio track, is an audio track enriched with machine-readable data (including the content identifier and the time in the audio track) within a given interval, such as step 1410 of method 1400.

[0236] Preferably, recognition data is extracted from audio tags on the user device. Advantageously, this results in fast and accurate recognition on the user device and optionally, synchronization of media content. Alternatively, the audio tags are transmitted to, for example, a server for recognition data extraction. If recognition data is successfully extracted, method 1500 proceeds to step 1510. Otherwise, step 1506 is repeated or method 1500 proceeds to step 1508.

[0237] Step 1508 includes retrieving identification data based on the best match of audio fingerprints from multiple audio fingerprints. The audio fingerprint includes one or more audio features extracted from the portion of the audio track. For example, step 1508 includes extracting one or more audio features to generate an audio fingerprint, such as step 1406 of method 1400, for example by hashing or otherwise encoding the one or more audio features.

[0238] The one or more audio features are extracted at the user device or the server. If the one or more audio features are extracted at the user device, the audio features or an audio fingerprint containing the one or more audio features are transmitted to the server. Extracting the one or more audio features at the user device improves efficiency. If the one or more audio features are not extracted at the user device, the portion of the audio track is transmitted to the server, and an audio fingerprint is generated at the server. Extracting the one or more audio features at the server reduces the computational resource requirements of the user device.

[0239] The generated audio fingerprint is compared with other audio fingerprints, such as multiple audio fingerprints stored in a database or on a server. Multiple audio fingerprints are associated with multiple portions of an audio track, such as an audio track at different times, and / or portions of multiple audio tracks, such as audio tracks associated with one or more media contents that have been previously processed to generate audio fingerprints associated with media content. Since each of the multiple audio fingerprints is associated with identification data, the identification data can be retrieved based on each audio fingerprint. The best match is the stored audio fingerprint that is most similar to the generated audio fingerprint, for example, it is the same audio fingerprint, and the identification data associated with the best match is retrieved by the server and subsequently received on the user's device.

[0240] Step 1510 includes identifying media content based on known media content from the identification data. For example, the best-matching identification data is associated with known media content, so the media content in step 1402 is identified as known media content. If the audio track is optionally associated with multiple known media content items, then the best match in step 1508 may optionally be multiple best matches, and step 1510 then includes selecting media content from the multiple known media content items. Returning to the example above, the audio track is a song played during a movie played on a television device. The song is also included in the soundtracks of different movies not played on the television device. Therefore, in step 1508, identification data associated with the movie on the television and the different movies is received on the user device. Step 1510 includes requesting a selection from the user on the user device and receiving the user's selection, such as displaying the titles of two movies based on content identifiers, and the user clicking or tapping the correct title.

[0241] Step 1510 may optionally also include identifying time points within the media content based on the identification data. This involves identifying not only the media content itself, but also its current time position, such as the minutes and seconds corresponding to the portion of media content being played, for example, on a television device. For example, a movie is identified as Movie A, and the time point is identified as 01:17:33 (using hh:mm:ss format). In this example, the time point includes a predetermined interval of one second, such that subsequent time points will be identified as 01:17:34. If the time point includes a predetermined interval of five seconds, then the time point will be identified as 01:17:30 – 01:17:35.

[0242] Optionally, the microphone is turned on and off to match the defined intervals where audio tags are embedded within the audio track. Advantageously, the audio tag is present and detectable when the microphone is on, and the microphone is off when the audio tag is absent and therefore undetectable, reducing the microphone's power requirements and extending the user device's battery life without affecting the synchronization process based on audio track monitoring. Further optionally, if no match is found, for example due to ambient noise for a period longer than the defined interval, the microphone remains on until an audio tag is detected or step 1508 is performed.

[0243] Step 1512 includes retrieving content information associated with the identified data. For example, retrieving content information generated in step 1414 of method 1400 and / or Figure 2 Subject file 226. Receives requests for content information, such as based on user-device interactions, and transmits identification data from the user device to the server. Optionally, the content information is information associated with one or more topics related to known media at a known point in time. Each of the one or more topics is linked to a known point in time, or an adjacent point in time within an adjustable buffer period.

[0244] If identification data is received from the server, for example, step 1508 is performed to identify media content, then step 1512 may optionally be performed together with step 1508, such that identification data and content information are received from the server without requiring additional requests for content information. Alternatively, additional requests for content information may include user selection, for example, if multiple best matches exist.

[0245] Because the retrieved content information is associated with identification data, which includes the time point of the currently playing media content, the retrieved content information is synchronized with the playback of the identified media content. Returning to the example above, video A is playing on a television device at 01:17:33. Several topics associated with video A at 01:17:33, including a narrating actor (invisible but audible at 01:17:33), a shark species (visible at 01:17:33), the location where the shark was filmed, and a service that donates to a shark conservation charity, are mentioned at different points in time in video A. The content information includes: information associated with the actors, such as their names and other known media content in which they have appeared; information associated with the shark, such as the species name and conservation status; information associated with the location, such as the name and travel costs between the user device location and the filming location; and information associated with the charity, such as its name and a link to its website. Optionally, step 1512 includes displaying the content information on the user device based on the time point. The displayed content information is then associated with the media content at the known time.

[0246] Optionally, step 1512 further includes transmitting an information request containing identification data, for example, from the user device to a server or database. Then, based on the transmitted identification data, information associated with one or more topics related to known media at a known point in time is retrieved on the user device from the server or database. Each of the one or more topics is linked to a known point in time, such as step 1308 of method 1300. Further optionally, the information request is transmitted based on user interaction with the user device. Alternatively, the information request is transmitted automatically after the media content is identified in step 1510, or based on other criteria associated with whether the user is interacting with the identified media content (e.g., being near the identified media content).

[0247] Figure 16A and Figure 16B A flowchart illustrating a method for content interaction on a user platform is shown. Specifically, method 1600 is... Figure 3A and Figure 3B Another method performed by the components of the content synchronization system 300. Any step of method 1600 may be combined with any step of method 1500 for identifying media content and / or synchronizing retrieved content information.

[0248] Method 1600 includes: step 1602, monitoring an audio signal; step 1604, obtaining selected audio tags from the audio signal; step 1606, extracting identification data from the selected audio tags; step 1608, determining whether a lock indicator has been received; step 1610, obtaining a request for a first topic file; step 1612, receiving the first topic file; step 1614, displaying data from the first topic file; step 1616, reducing monitoring of the audio signal; step 1618, requesting multiple topic files; step 1620, receiving multiple topic files; step 1622, displaying data associated with the first topic file; and step 1624, displaying data associated with a second topic file.

[0249] Step 1602 includes monitoring an audio signal, wherein the audio signal includes one or more audio tags embedded therein. For example, the audio signal comes from an embedded audio track generated in step 1410 of method 1400. In another example, step 1602 includes steps 1502, 1504, and 1506 of method 1500. The audio signal is associated with an audio track of the media content and is emitted from a speaker of the media playback device. Preferably, the media playback device is independent of the user device associated with monitoring the audio signal. For example, the audio signal is collected using a microphone connected to the user device and monitored at a user platform on the user device. Optionally, the user platform is the user device. Alternatively, the user platform is a software interface or application running on the user device.

[0250] Step 1604 includes retrieving a selected audio tag from the audio signal. The selected audio tag is used to identify media content and may optionally be the first audio tag detected during monitoring of the audio signal. Alternatively, it may be a subsequent audio tag, for example, when monitoring of the audio signal is repeated in step 1602. The audio tag includes coded identification data associated with specific media content, such as media content that includes the audio signal of step 1602.

[0251] Optionally, obtaining the selected audio tag includes extracting multiple audio tags from one or more audio signals. For example, multiple media contents playing simultaneously result in the detection of multiple audio tags. The selected audio tag is determined from the multiple audio tags based on user interaction with the user platform (e.g., selecting audio tags associated with media content that the user wishes to interact with).

[0252] Step 1606 includes extracting identification data from the selected audio tags. The identification data includes a media content identifier (e.g., a content identifier unique to a specific media content) and a time point, such as a first time point, indicating the current location within the media content where the audio signal was emitted and monitored. Therefore, step 1606 includes identifying the media content associated with the audio signal monitored in step 1602.

[0253] Alternatively, steps 1604 and 1606 may include step 1508 of method 1500. For example, when an audio tag is undetectable, identification data is received at the user platform based on one or more audio features extracted from the audio signal.

[0254] Step 1608 includes determining whether a lock indicator is received at the user platform. The lock indicator indicates whether the user platform should synchronize with the media content identified in step 1606. For example, the lock indicator is a user interaction with the user device running the user platform, such as tapping a touchscreen display, or a criterion determined at the user platform, such as whether the user device is moving away from the media content, for example based on location data or monitoring of audio signals in step 1602, or whether audio tags from different audio signals can be detected.

[0255] If a lock indicator is received, method 1600 proceeds to step 1616. If no lock indicator is received, method 1600 proceeds to step 1610.

[0256] Step 1610 includes obtaining a request indicator for a first subject file, wherein the request indicator is another indication of user interaction with the user platform or of retrieving information associated with the identification data. The first subject file is associated with the media content of the identification data at a point in time (e.g., a first time point) of the identification data. One example indication that continues is an estimated proximity of the user to the media content, such as the volume of an audio signal; another example is whether the user platform is active on the user device, such as the user platform being displayed on the user device.

[0257] Step 1612 includes transmitting the identification data extracted or received in step 1606 to a server device (e.g., a server) and receiving a first subject file from the server device. The first subject file includes a first visual depiction, such as an illustration of the first subject, and first subject text, such as a name and / or text description or information associated with the first subject. For example, the first subject file is... Figure 2 Theme file 226. Optionally, step 1612 includes receiving a plurality of first theme files, wherein each theme file is associated with a theme related to media content at, for example, a first point in time.

[0258] Step 1614 includes initiating the display of a first visual depiction or a first topic text. For example, step 1614 includes displaying both the first visual depiction and the first topic text. In another example, step 1614 includes displaying the first visual depiction and, upon receiving interaction between the user and the user platform, displaying the first topic text. In yet another example, initiating the display includes receiving interaction between the user and the user platform indicating whether to display, for example, the first visual depiction, the first topic text, or neither. Optionally, method 1600 proceeds to step 1602.

[0259] For example, repeating step 1602 includes monitoring an audio signal, wherein the audio signal includes a new audio tag. Step 1604 repeats and retrieves a new audio tag from the audio signal. Step 1606 extracts new identification data from the new audio tag, wherein the new identification data is associated with a media content identifier and a second time point or a new media content identifier.

[0260] Optionally, each of the multiple topics is associated with a category. Each of the multiple topics is associated with a visual description and topic text, such that, for example, a first topic file includes multiple visual descriptions and multiple topic texts. Further optionally, initiating the display in step 1614 includes selecting, for example, a visual description of a topic from the multiple topics based on the category for display.

[0261] Step 1616 includes reducing the monitoring of the audio signal. For example, the microphone is turned on and off to match defined intervals where subsequent audio tags are embedded within the audio track, or the microphone is turned off for longer intervals, with only periodic or sporadic audio tag detection to verify synchronization. Optionally, the microphone remains off until an indicator (such as user interaction with the user platform or a change in the user device's location or activity) triggers the microphone to turn on to monitor the audio signal.

[0262] Step 1618 includes transmitting the identification data to a server device, thereby requesting multiple subject files from the server device. The multiple subject files are associated with the media content of the identification data at multiple points in time, such as a first time point and a predetermined number of additional time points, such as the adjustable buffer period based on method 1500. The predetermined number of additional time points may optionally be subsequent time points, such as a second and third time point, for example, when the media content is being emitted normally. However, if monitoring of the audio signal indicates that the media content is being emitted at an accelerated rate, such as fast forward, the predetermined number of additional time points may optionally be selected subsequent time points, such as a fourth and eighth time point, to maintain synchronization between the user platform and the media content. Alternatively, if monitoring of the audio signal indicates that the media content has been paused, the additional time points may optionally be time points around the first time point. Alternatively, if monitoring of the audio signal indicates that the media content is being rewound or played backward, the additional time points may optionally be time points prior to the first time point.

[0263] Step 1620 includes receiving multiple topic files from a server device. For example, the server uses the identification data received in step 1618 to retrieve multiple topic files from a database, where each topic file is associated with one or more topics extracted from the media content, such as steps 1412 and 1414 of method 1400.

[0264] Step 1622 includes initiating the display of a first visual depiction or first subject text, as described with respect to step 1614.

[0265] Step 1624 includes initiating the display of a second visual depiction or a second topic text, as described with respect to step 1614. The second visual depiction and the second topic text are contained within a second topic file among the plurality of topic files received in step 1620. Preferably, the display of the second visual depiction is performed when a second time point is reached. Optionally, the arrival of the second time point is verified by monitoring a second audio tag in the audio signal.

[0266] Optionally, if a lock indicator is received, method 1600 further includes receiving a monitoring indicator and resuming monitoring of the audio signal. The monitoring indicator indicates that the time difference between arriving at the second time point and displaying the second visual depiction or second subject text exceeds a threshold. For example, the monitoring indicator is received when a user of the user platform leaves a transmission area, which includes a geographical location associated with the transmission, transmission, broadcasting, or streaming of the audio signal or media content. Alternatively, the monitoring indicator is received when the audio signal or media content is paused, sped up, slowed down, or otherwise interrupted or interfered with compared to the temporal continuity of the audio signal or media content.

[0267] Optionally, method 1600 includes receiving user interactions with a user platform associated with displayed data. For example, the user interaction may indicate a request for further information associated with the displayed topic data, or a request to share, comment on, or otherwise interact with components of the topic file.

[0268] Figure 17 A flowchart illustrating a method for audio track recognition and content information transmission on a server device is shown. Specifically, method 1700 is... Figure 4 One method performed by the components of the audio recognition system 400.

[0269] Method 1700 includes: step 1702, acquiring one or more audio features; step 1704, acquiring a first portion of an audio track; step 1706, extracting one or more audio features from the first portion of the audio track; step 1708, generating a first audio fingerprint containing the one or more audio features; step 1710, comparing the first audio fingerprint with a plurality of audio fingerprints; step 1712, identifying the best match most similar to the first audio fingerprint; step 1714, transmitting identification data based on the best match; and step 1716, transmitting content information associated with the identification data. Steps 1704 and 1706 are alternatives to step 1702, and step 1716 is optional, for example, depending on whether a request for content information is received at the server performing method 1700.

[0270] Step 1702 includes acquiring one or more audio features extracted from a first portion of the audio track from the user equipment. For example, the server receives one or more audio features extracted by the user equipment, wherein the user equipment includes a microphone that monitors or otherwise collects audio signals from the audio track. After receiving one or more audio features, method 1700 proceeds to step 1708. In another example, the server receives a first audio fingerprint generated by the user equipment, and method 1700 proceeds to step 1710. If the server does not acquire one or more audio features, method 1700 proceeds to step 1704.

[0271] Step 1704 includes acquiring a first portion of the audio track from the user equipment, such as an audio signal associated with the audio track over a period of time.

[0272] Step 1706 includes extracting one or more audio features from the first portion. For example, if the user equipment is unable to extract one or more audio features from the first portion, the server receives the first portion in step 1704 and extracts one or more audio features in step 1706. For example, extracting one or more audio features includes analyzing the audio signal of the audio track to identify or quantize selected audio features.

[0273] Step 1708 includes generating a first audio fingerprint containing the one or more audio features. The first audio fingerprint is associated with information about the audio signal in the first portion, such that the first audio fingerprint is a hashed or otherwise encoded version of the one or more audio features. The first audio fingerprint not only requires less storage space but also allows for faster transfer between devices. Furthermore, hashing or otherwise encoding the one or more audio features allows for filtering out or otherwise compensating for noise, environmental conditions, volume differences, or other effects on the audio signal. The same first audio fingerprint can still be generated even when the first portion of a soundtrack is played at different volumes on different quality speakers, subjected to different noise levels, and recorded by different microphones.

[0274] Step 1710 includes comparing the first audio fingerprint with a plurality of audio fingerprints. Each of the plurality of audio fingerprints is associated with identification data containing a content identifier and a time point, such that each audio fingerprint is associated with a known audio track at a known time point. For example, the server compares the first audio fingerprint with a plurality of audio fingerprints stored in a database.

[0275] Step 1712 includes identifying the best match, where the best match is the audio fingerprint most similar to the first audio fingerprint. For example, steps 1710 and 1712 use a ball tree algorithm to identify which of the multiple audio fingerprints is most similar to or identical to the first audio fingerprint. Optionally, step 1710 is repeated for each of the multiple audio fingerprints until the best match is identified in step 1712. Alternatively, a hierarchical structure is used to store or retrieve audio fingerprints from a database, such that each repetition of step 1710 identifies a more similar audio fingerprint. Alternatively, steps 1710 and 1712 are combined, and artificial intelligence or computational algorithms are applied to extract the audio fingerprint most similar to the first audio fingerprint from the multiple audio fingerprints to identify the best match. For example, comparing the first audio fingerprint with multiple audio fingerprints includes performing a nearest neighbor search to identify the best match.

[0276] Step 1714 includes transmitting the identification data to the user equipment based on the best match. For example, the identification data associated with the best match identified in step 1712 is transmitted to the user equipment for identifying the audio track in steps 1702–1706.

[0277] Step 1716 is optional and depends on whether a request is received from the user equipment. If a request is received, step 1716 includes transmitting content information associated with the identification to the user equipment. For example, a request for content information, including identification data, is received from the user equipment at the server, such as step 1610 or step 1618 of method 1600. In response to the request, the server retrieves content information associated with the received identification data from storage (e.g., a database), such as content information related to the audio track at the time of identification. If no request is received, step 1716 may optionally include transmitting a query to the user equipment to determine whether the audio track was successfully identified based on the best match. Alternatively, step 1716 may not be performed.

[0278] Optionally, method 1700 further includes obtaining a second portion of the audio track from the user device, or one or more audio features extracted from the second portion. Steps 1708-1714 are then repeated at least for the second portion. A second audio fingerprint is generated, comprising one or more audio features extracted from the second portion. The second audio fingerprint is then compared with multiple audio fingerprints to identify an updated best match, which is the audio fingerprint most similar to the second audio fingerprint. Identification data associated with the updated best match and / or content information associated with the updated best match is then transmitted to the user device. For example, playback of media content associated with the audio track is paused, skipped, or otherwise interrupted, causing it to no longer play in typical chronological order. The second audio fingerprint is then used to maintain synchronization with the audio track. Alternatively, the second portion of the audio track is a part of a second audio track, where the second audio track is associated with different media content, such as different media content currently playing on a media playback device. The second audio fingerprint is then used to identify the second audio track.

[0279] As described above, multiple audio fingerprints may optionally be stored in a database. Preferably, the database is configured such that the database or a server connected to the database can retrieve identification data based on the audio fingerprints and retrieve content information based on the identification data.

[0280] Figure 18A , Figure 18B and Figure 18C A flowchart illustrating a method for media content recognition and interaction is shown. Specifically, method 1800 is... Figure 1-4 Another method performed by components of the media interaction system 100, media processing system 200, content synchronization system 300, or audio recognition system 400. Specifically, method 1800 is an example embodiment of method 1200, although those skilled in the art will understand that alternative embodiments of method 1200 are possible, at least based on the information provided herein.

[0281] Method 1800 includes: step 1802, acquiring known media content; step 1804, acquiring known audio tracks; step 1806, generating a timeline from the known media content; step 1808, generating embedded audio tracks; step 1810, generating content information; step 1812, storing the embedded audio tracks and content information; step 1814, monitoring a portion of an unknown audio track; step 1816, identifying media content by acquiring recognition data; step 1818, determining whether an audio tag is detected; step 1820, extracting recognition data from the audio tag; step 1822, receiving recognition data from a server; step 1824, initiating a request for content information; step 1826, receiving a request for content information; step 1828, retrieving stored content information using the recognition data; and step 1830, outputting the content information to a user device.

[0282] Step 1802 includes acquiring the known media content associated with the known audio track. If the known audio track is not acquired in step 1802, or if the predetermined quality threshold that allows for optimal audio feature extraction or audio tag embedding is not met, then method 1800 proceeds to step 1804. If the known audio track is acquired in step 1802, then method 1800 proceeds to step 1806.

[0283] Step 1804 is optional and includes acquiring a known audio track. For example, if the audio track acquired in step 1802 is below a predetermined quality threshold, then step 1804 includes retrieving the known audio track and determining that the known audio track exceeds the predetermined quality threshold. For example, steps 1802 and 1804 are step 1402 of method 1400.

[0284] Step 1806 includes generating a timeline based on the known media content. The timeline includes multiple points in time such that the timeline is a sequence of segmented time intervals representing the duration of the known media content. For example, step 1806 is step 1404 of method 1400.

[0285] Step 1808 includes generating an embedded audio track by embedding audio tags within a known audio track. The audio tags include identification data associated with known media content at a selected time point (e.g., a first time point). The audio tags are embedded at frequencies detectable by a microphone and imperceptible to the user of the user's device. For example, step 1808 is step 1410 of method 1400. Optionally, the embedded audio track is generated by embedding multiple audio tags within a known audio track. For example, FSK modulation is used to embed the audio tags at regular time intervals, such as every 12 seconds.

[0286] Optionally, step 1808 further includes generating an audio fingerprint for each time point of the timeline, thereby generating multiple audio fingerprints. An audio fingerprint includes one or more audio features extracted from a known audio track at each time point. Based on the multiple time points of the timeline, identification data is linked to each of the multiple audio fingerprints.

[0287] Step 1810 includes generating content information associated with the identified data. Information related to one or more topics extracted from known media content at a selected time point (e.g., a first time point) is organized to form the content information. For example, step 1810 is steps 1412 and 1414 of method 1400.

[0288] Step 1812 includes storing the embedded audio track and content information associated with each point in time on the timeline. For example, step 1812 is step 1416 of method 1400. Optionally, steps 1802–1812 are performed on a workstation or server.

[0289] Step 1814 includes monitoring a portion of the unknown audio track using the microphone of the user device. The unknown audio track is associated with media content emitted from a speaker of a media playback device, which is preferably independent of the user device. Optionally, monitoring this portion includes monitoring an audio signal that includes or is otherwise associated with that portion of the unknown audio track and processing the audio signal to improve the detection of audio tags, such as filtering out noise.

[0290] In one example, the audio signal also includes a portion of another unknown audio track. Therefore, monitoring the audio signal in step 1814 includes detecting multiple audio tags. Step 1820 then includes extracting identification data for each of the multiple audio tags, which is associated with multiple known media content, and receiving user interaction to select known media content from the multiple known media content.

[0291] Step 1816 includes identifying media content by acquiring identification data. If an audio tag is detected in step 1818, then step 1816 includes step 1820. If no audio tag is detected in step 1818, then step 1816 includes step 1822. For example, steps 1814-1822 are steps 1502-1508 of method 1500. For example, if an audio tag is detected within the portion of the unknown audio track, identification data is extracted from that audio tag. The extracted identification data is associated with known media content at a known point in time, thus the unknown audio track is identified as a known audio track of known media content.

[0292] Step 1818 includes determining whether an audio tag has been detected, such as step 1504 of method 1500. If ambient noise affects or otherwise affects the detection of the audio tag required for extracting identification data in step 1820, or affects or otherwise affects the extraction of one or more audio features required for receiving identification data in step 1822, then step 1818 also includes implementing one or more noise reduction algorithms or signal enhancement techniques. These algorithms / techniques minimize the impact of ambient noise, ensuring reliable audio code detection and feature extraction in noisy environments.

[0293] Alternatively, if no audio tag is detected, step 1822 includes acquiring one or more audio features extracted from said portion of the unknown audio track monitored by the user device, wherein the audio features are extracted at the user device or server. An unknown audio fingerprint containing the one or more audio features is generated and compared with multiple audio fingerprints generated in step 1808. Once the best match is identified, i.e., the audio fingerprint most similar to the unknown audio fingerprint, the identification data associated with the best match is transmitted to the user device. Optionally, the server or workstation device communication is coupled to a database, such as a database configured to store the embedded audio track, optional multiple audio fingerprints, and content information associated with each point in time on the timeline.

[0294] Step 1824 includes initiating a request for content information to the server device based on the interaction between the user and the user device. For example, step 1824 includes step 1512 of method 1500.

[0295] Step 1826 includes receiving the request for content information from step 1824, wherein the request includes the identification data from step 1816. For example, step 1826 is step 1610 or step 1618 of method 1600.

[0296] Step 1828 includes retrieving stored content information associated with the identified data. For example, step 1828 is step 1716 of method 1700.

[0297] Step 1830 includes outputting content information to a user device, such as step 1512 of method 1500.

[0298] Optionally, the steps of method 1800 are repeated, for example, steps 1802-1812 are repeated when processing second known media content or a second known audio track. For example, repeating step 1808 includes embedding a second audio tag within a second known audio track on a workstation device. The second audio tag includes identification data associated with the second known audio track, thereby generating a second embedded audio track. Repeating step 1812 includes storing the second embedded audio track. Optionally, additional steps are repeated for additional known media content or known audio tracks. For example, repeating step 1814 includes monitoring a portion of a second unknown audio track using a microphone on a user device. Repeating step 1820 includes extracting identification data from the second audio tag detected within said portion of the second unknown audio track, thereby identifying the second known audio track as the second known audio track. Repeating step 1824 includes requesting content information associated with the second known audio track by transmitting the identification data to a server device, and repeating step 1830 includes receiving the content information associated with the second known audio track on a user device. Therefore, those skilled in the art will understand that not all steps of the overall process outlined in methods 1200 and 1800 need to be repeated.

[0299] Figure 19 A flowchart illustrating a method for establishing secure communication on a user equipment is shown. Specifically, Figure 19 A method 1900 for secure communication and output synchronization with transmission device 504, transmission application and / or SDK as described above (hereinafter referred to as transmission device 504) is shown, executed at user equipment 502 and / or user application (hereinafter referred to as user equipment 502).

[0300] Method 1900 includes: step 1902, scanning content; step 1904, detecting a first sound code 520; step 1906, converting the first sound code 520 into an identification code 522; step 1908, transmitting the identification code 522; and step 1910, establishing a bidirectional connection. Method 1900 may optionally include step 1912, synchronous output.

[0301] Step 1902 includes scanning content by broadcasting a first request 512 using a bidirectional communication protocol. For example, user equipment 502 broadcasts the first request 512, such as using WebSocket, to be detected at a compatible transport device 504. The compatible transport device 504 includes the capability to receive the first request 512 using a bidirectional communication protocol, such as a device with enabled bidirectional communication protocol access, directly or via a transport application and / or SDK. This allows encrypted bidirectional communication or other configuration for secure communication between user equipment 502 and transport device 504. Optionally, user equipment 502 is a user application running on user equipment 502 to provide bidirectional communication functionality or another functionality described with respect to method 1900.

[0302] Step 1904 includes detecting a first voice code 520 using a microphone. The microphone may be controlled by a user application running on user equipment 502. Optionally, the microphone is communicatively coupled to or otherwise associated with user equipment 502. The first voice code 520 includes an identifier 518 and an optional timestamp, wherein the identifier 518 is an alphanumeric code encoded into sound, such as using FSK modulation. The first voice code 520 is emitted from a speaker 508 associated with transmission device 504, such as a speaker controllable by the transmission application / SDK.

[0303] Preferably, the first voice code 520 is detectable by a microphone but not easily detectable by the user of user equipment 502. For example, the first voice code 520 uses ultrasonic or near-ultrasonic frequencies. This allows voice-based communication between transmission device 504 and user equipment 502 without negatively impacting the user's interaction with audio output from transmission device 504, speaker 508, or user equipment 502.

[0304] Step 1906 includes converting the first sound code 520 into an identification code 522. The conversion includes decoding the first sound code 520 into alphanumeric code, for example, by reversing the previous encoding process, and / or extracting the identification code 522 from the first sound code 520. Optionally, the conversion also includes decrypting the first sound code 520 and / or the alphanumeric code.

[0305] Step 1908 includes transmitting an identification code 522 for authentication using a bidirectional communication protocol. For example, the identification code 522 is transmitted to a transmission device 504 for authentication. Authentication may optionally include comparing the identification code 522 with an identifier 518 to determine whether the identification code 522 matches the identifier 518. If the identification code 522 matches the identifier 518, then communication with the user equipment 502 is secure and / or unaffected by errors. Otherwise, communication with the user equipment 502 is insecure or affected by errors.

[0306] Step 1910 includes establishing a bidirectional connection, for example, by performing a handshake via WebSocket, and the bidirectional connection is a full-duplex channel. If the transmission device 504 is authenticated with identification code 522, the bidirectional connection is used to pair the transmission device 504 and the user equipment 502, allowing secure and fast real-time communication between the devices. This achieves efficient and accurate output synchronization.

[0307] Step 1911 is an optional step, which includes synchronizing a first output (e.g., output 606) of user equipment 502 with a second output (e.g., output 608) of transmission device 504. For example, a bidirectional connection is used to transfer time information between user equipment 502 and transmission device 504.

[0308] In one example, time information includes a current timestamp associated with the output at user equipment 502, which is identified by user equipment 502. A message timestamp associated with the output at transmission device 504 is communicated from transmission device 504. If the current timestamp is earlier than the message timestamp, the output at user equipment 502 is modified to slow down or otherwise match the timing associated with the message timestamp. Alternatively, if the current timestamp is later than the message timestamp, the output is sped up or otherwise matched. If the current timestamp matches the message timestamp, no action is required because the output is already synchronized.

[0309] In another example, the time information includes a current timestamp associated with the output at transmission device 504, which is identified by transmission device 504. Message timestamps associated with the output at user equipment 502 are communicated from user equipment 502. If the current timestamp is earlier than the message timestamp, the output at transmission device 504 is modified to slow down or otherwise match the timing associated with the message timestamp. Alternatively, if the current timestamp is later than the message timestamp, the output is sped up or otherwise matched.

[0310] Optionally, one or more instructions to initiate output alternation may also be communicated. For example, user equipment 502 transmits instructions for accelerating the output of transmission device 504 based on time information received from transmission device 504. In another example, transmission device 504 transmits instructions for accelerating transmission device 504 based on time information received from user equipment 502.

[0311] Preferably, method 1900 is used to pair and synchronize user equipment 502 (e.g., mobile device) with transmission device 504 using a bidirectional communication protocol (e.g., WebSocket) and sound (e.g., ultrasound). For example, method 1900 includes: initiating a content scan using WebSocket; receiving ultrasound or near-ultrasound codes via a microphone on the mobile device; converting the received ultrasound or near-ultrasound codes into text codes; sending the text codes to transmission device 504 via WebSocket; performing a handshake with transmission device 504 if the text codes match; synchronizing the mobile device with transmission device 504 via WebSocket; and periodically receiving ultrasound or near-ultrasound signals to ensure the mobile device remains near transmission device 504, and prompting the user to confirm its presence if a signal is missed.

[0312] Figure 20 A flowchart illustrating a method for establishing secure communication in a transport application is shown. Specifically, Figure 20A method 2000 for secure communication and output synchronization with user equipment 502 and / or user application (hereinafter referred to as user equipment 502) is shown, executed at transmission device 504, transmission application and / or the SDK as described above (hereinafter referred to as transmission device 504).

[0313] Method 2000 includes: step 2002, receiving a first request; step 2004, transmitting a pairing request; step 2006, receiving an identifier 518; step 2008, converting the identifier 518 into a first voice code 520; step 2010, receiving an identification code 522; step 2012, authentication; step 2014, performing a handshake; and optional step 2016, error detection.

[0314] Step 2002 includes receiving a first request associated with user equipment 502 using a bidirectional communication protocol (e.g., WebSocket). For example, transmission device 504 receives or otherwise detects the first request broadcast by user equipment 502 in step 1902 of method 1900.

[0315] Step 2004 includes transmitting a pairing request to server 516, such as sending pairing request 514 to server 516. The pairing request includes data associated with the first request, such as identification information associated with user equipment 502 and / or transmission equipment 504. Server 516 uses the pairing request to identify authorized devices, trusted devices, and / or devices previously paired with transmission equipment 504.

[0316] Step 2006 includes receiving an identifier 518 from server 516. For example, the identifier 518 is received from server 516 only if server 516 has recognized the pairing request, such as by identifying or otherwise authorizing user equipment 502, the user associated with user equipment 502, the location of user equipment 502, or other relevant information contained in / associated with the pairing request. Optionally, identifying the pairing request includes identifying or otherwise authorizing transport device 504, and the transport application running or an SDK (e.g., SDK 700).

[0317] Step 2008 includes converting identifier 518 into a first sound code 520. The first sound code 520 includes identifier 518 and optionally a message timestamp, wherein identifier 518 is an alphanumeric code encoded into sound, such as using FSK modulation. Optionally, the first sound is emitted from a speaker 508 associated with transmission device 504, such as a speaker controllable from a transmission application / SDK.

[0318] Step 2010 includes receiving identification code 522 from user equipment 502 using a bidirectional communication protocol. For example, identification code 522 is the identification code 522 transmitted in step 1908 of method 1900.

[0319] Step 2012 includes authenticating the identifier 522 based on a match between the identifier 518 and the identifier 522. For example, the identifier 522 is compared with the identifier 518 to determine whether the user equipment 502 is authenticated and / or whether communication with the user equipment 502 is secure and unaffected by errors. If the identifier 522 matches the identifier 518, then the user equipment 502 and the communication are secure, authorized / trusted, and / or unaffected by errors, and method 2000 proceeds to step 2014. Otherwise, the user equipment 502 or the communication is insecure, unauthorized / untrusted, or affected by errors, and method 2000 proceeds to step 2016.

[0320] Step 2014 includes performing a handshake to establish a bidirectional connection, such as bidirectional connection 702 and / or the bidirectional connection established in step 1910 of method 1900. For example, establishing a bidirectional connection includes performing a handshake using WebSocket to establish a full-duplex connection. Once the WebSocket connection is established, the paired devices (user device 502 and transport device 504) can begin exchanging data frames using the WebSocket protocol. The WebSocket protocol provides low-overhead communication, allowing real-time bidirectional data exchange, and is suitable for applications requiring real-time updates and improved synchronization between complex and / or data-rich outputs.

[0321] Optionally, user equipment 502 sends a handover request to transmission equipment 504 to upgrade the bidirectional communication protocol to a bidirectional connection. For example, the handover request is associated with identification code 522 to initiate a handshake. Then, transmission equipment 504 responds to the handover request in step 2014, agrees to the upgrade, and sends a verification response from user equipment 502 to transmission equipment 504 to confirm the handshake is complete. Further optionally, the first request 512 includes a handover request, and identification code 522 includes a verification response.

[0322] Alternatively, step 2014 includes: after authenticating the identification code 522 in step 2012, the transmission device 504 sends a handover request to the user equipment 502. Then, the user equipment 502 responds to the handover request and establishes a bidirectional connection. Further optionally, the transmission device 504 sends an authentication response to the user equipment 502 to confirm the bidirectional connection.

[0323] Alternatively, server 516 may participate in establishing a bidirectional connection, for example, by sending a handover request as a pairing request 514 to server 516, which processes the request to determine whether user equipment 502 and transmission equipment 504 are authorized / trusted. Optionally, server 516 may determine whether both user equipment 502 and transmission equipment 504 have the functionality required to perform a handshake and maintain a bidirectional connection.

[0324] Those skilled in the art will understand that there are many variations in performing a handshake, and it is not limited to the examples detailed above.

[0325] Step 2016 includes performing error detection, such as if identification code 522 is not authenticated or if the handshake fails.

[0326] For example, if a WebSocket handshake fails, the following error detection and correction processes can be employed: retry mechanisms, such as implementing a retry strategy where the client attempts to handshake multiple times with delays between attempts, optionally including exponential backoff to avoid overloading the server; protocol-level error handling, such as checking specific status codes or error messages indicating specific problems, which can be logged or trigger alternative actions; TLS / SSL certificate verification, such as verifying the relevant certificate if the handshake fails due to SSL / TLS issues, where correction may involve updating the certificate, adjusting security settings, or notifying the user of the problem; alternative connection methods, such as trying alternative protocols or fallback methods (e.g., using HTTP / 2 or long polling) to maintain communication if the handshake continues to fail; network diagnostics, such as performing network diagnostics to check for problems that may be preventing WebSocket connections, such as DNS resolution problems, firewall restrictions, or proxy settings; and user notification, such as notifying the user of the handshake failure and providing specific error details, providing troubleshooting steps, or suggesting alternative actions, such as switching networks or contacting support. These methods help to handle and correct problems when a WebSocket handshake fails, ensuring effective management of potential connection issues.

[0327] Preferably, method 2000 is used to pair and synchronize transmission device 504 with user device 502 (e.g., mobile device) using a bidirectional communication protocol (e.g., WebSocket) and sound (e.g., ultrasound). For example, method 2000 includes: receiving a pairing request from the mobile device via WebSocket; sending the pairing request to server 516 to identify the request using a transmission ID, user ID, and geographic location; receiving an identifier 518 code from server 516; converting the identifier 518 code from text to an ultrasound or near-ultrasound code; sending the ultrasound or near-ultrasound code to the mobile device; receiving a text code from the mobile device via WebSocket; performing a handshake with the mobile device if the text codes match; and periodically sending ultrasound or near-ultrasound signals to ensure that the mobile device remains in the vicinity of transmission device 504.

[0328] In an optimized example, methods 1900 and 2000 are used in combination, for example, using system 500. Optionally, to achieve pairing and synchronization using WebSocket and ultrasound: the transmitting device obtains secure / authorized WebSocket access via the SDK or transmitting application as described above; the transmitting device is configured to send sound or information, data, or instructions to another associated device to send sound upon request; and / or the user equipment runs a user application capable of picking up sound through a microphone.

[0329] For example, according to one embodiment, the user application initiates a content scan using WebSocket. A transport device (TDwWS) with WebSocket access receives this request and sends a pairing request to the SDK embedded in the service. The SDK forwards the request to the server, which identifies the request using the transport ID, user ID, and / or geolocation. The server then sends an identification code to the SDK. This code is converted from text to ultrasound or near-ultrasound code by the SDK, which then sends the ultrasound or near-ultrasound code. The user application picks up this code via the device's microphone and converts it back to text. The user application sends the text code to the TDwWS via WebSocket. If the text code matches, a handshake is performed via WebSocket, and the devices are now paired. Synchronization is then performed via WebSocket. To ensure the consumer remains near the TDwWS, ultrasound or near-ultrasound signals are sent at time intervals. If a signal is missed, the application prompts the user with a message such as "Are you still watching?" The user can confirm by interacting with a confirmation option in the user application.

[0330] Further implementation options include error correction or redundancy in the ultrasonic signal to improve reliability in noisy environments. Optional robust coding schemes such as Reed-Solomon help ensure accurate reception and decoding of the ultrasonic code. To maintain session management and security, ultrasonic or near-ultrasonic codes can be optionally sent at time intervals to revoke or extend access to content synchronization. If a signal is missed, a grace period or retry mechanism can be optionally implemented before revoking access. Additionally, geolocation tagging ensures disconnection when the user leaves the vicinity of TDwWS.

[0331] The managed service's SDK handles receiving pairing requests, converting identifiers to ultrasound signals, and maintaining WebSocket connections. The SDK is configured to receive pairing requests, convert identifiers to ultrasound signals, send ultrasound signals, handle WebSocket communication, and timestamps. Functions provided by the SDK include initiating WebSocket connections with timestamp synchronization, processing pairing requests, sending timestamped ultrasound signals, sending requests to the server, retrieving and / or identifying the current timestamp, and processing timestamped WebSocket messages.

[0332] The user application is configured to receive ultrasonic signals and WebSocket messages, extract timestamps, and perform corresponding synchronization. It is also configured to receive pairing requests, convert ultrasonic waves into identification codes, send the identification codes, extract timestamps, synchronize accordingly, and perform WebSocket communication. The user application provides functionality including initiating WebSocket connections and handling timestamp synchronization, sending scan requests with the current timestamp, initiating ultrasonic signal monitoring, processing recorded ultrasonic data and extracting timestamps, synchronizing based on timestamps, sending decoded codes via WebSocket, processing WebSocket messages with timestamps, and / or retrieving files associated with identified audio and the current timestamp.

[0333] For synchronization, both the SDK and the user application embed timestamps in WebSocket messages and ultrasonic signals. This allows the two devices to synchronize based on the timing of the content being played. The user application adjusts its synchronization based on the difference between the received timestamp and its current timestamp, ensuring that content-related operations on the user device are aligned with the output associated with the SDK. Heartbeat messages are periodically exchanged to continuously synchronize the two devices, including timestamps, to correct any drift that occurs over time.

[0334] Security and reliability are enhanced by ensuring that the SDK and user applications maintain timestamp consistency using the same time source (e.g., Network Time Protocol (NTP)). Error correction is optionally implemented for ultrasonic signals, and the WebSocket connection is optionally secure, for example, using WSS. A connection recovery mechanism is also preferably implemented to handle interruptions.

[0335] Figure 21A and Figure 21B The system architecture diagram of the video interaction system is shown. Specifically, Figure 21A A system 2100 for processing media content, such as video content including audio tracks and accompanying visual elements, is shown. Figure 21B A system 2100 for synchronizing with media content is shown.

[0336] System 2100 includes media content 2102, time interval 2104, media segment 2106, workstation device 2108, hierarchical hash structure 2110, database 2112, transmission device 2114, user device 2116, and data 2118.

[0337] Media content 2102, such as video content, is divided into media segments 2106 based on time intervals 2104, each segment having a predefined duration, such as 1 second. Optionally, the predefined duration is selected based on balancing efficient content synchronization with accurate identification of media segments. Media segments 2106 are hashed, preferably using cryptographic hashing techniques or algorithms such as SHA-256. A hierarchical hash structure 2110, such as a Merkle tree, is constructed using the hashes, where each leaf node represents a segment hash, and the parent node is a hash that combines (e.g., joins) the child hashes, ultimately forming the root hash. The root hash is distributed to user devices 2116, preferably at the start of synchronization with media content 2102, and optionally stored using blockchain technology. Database 2112 includes data associated with media content 2102, such as metadata. For example, descriptions, topics, or other embedded elements described above. The metadata is pre-stored in the database, making content associated with a specific media segment 2106 retrievable based on the segment hash.

[0338] User equipment 2116 (e.g., a user application installed or running on a mobile device, tablet, or electronic device, optionally identical to workstation equipment 2108) also hashes media segments 2106 obtained by monitoring media content 2102 that is being played, transmitted, or otherwise available from transmission equipment 2114. Optionally, the media content 2102 is being played or otherwise displayed on user equipment 2116, for example, transmission equipment 2114 is user equipment 2116. User equipment 2116 uses segment hashes to retrieve data 2118 associated with media segment 2106, such as the content information of media content 2102 and, in particular, the content of media segment 2106, such as the location displayed in media content 2102 during time interval 2104.

[0339] System 2100 is configured to handle various situations where portions of media content 2102 may be lost, corrupted, or unrecognizable. In such cases, user equipment 2116 employs algorithms such as Dynamic Time Warping (DTW) and / or AI-based predictive models to adjust timestamps and predict hashes of lost segments. This ensures continuous and seamless synchronization of media content 2102 even in the event of an interruption.

[0340] One significant advantage of system 2100 is its ability to provide secure and efficient synchronization between devices and media content. The use of cryptographic hashing and hierarchical hashing structures ensures that the media content 2102 is tamper-proof and that its integrity is maintained throughout transmission and playback. Furthermore, the integration of metadata from database 2112 allows for a rich and interactive user experience, as user device 2116 can dynamically retrieve and display relevant information based on the media segment 2106 being played.

[0341] The cryptographic hashing applied to media segment 2106 involves generating a unique, fixed-size hash value for each segment of media content 2102 using a secure hashing algorithm. This process first divides media content 2102 into discrete segments of fixed duration, such as 1-second intervals. Each segment is then processed by a cryptographic hash function (e.g., SHA-256), which generates a 256-bit hash value. This hash value serves as a unique identifier for media segment 2106, ensuring that even slight changes to the segment content result in significantly different hash values.

[0342] The use of cryptographic hashing offers several beneficial advantages. First, it ensures data integrity because any modification to media segment 2106 will be immediately detected through hash value mismatches. Second, it enhances security by making it computationally infeasible to reverse engineer the original media content 2102 from the hash value, thus protecting the content from unauthorized access or tampering. Furthermore, cryptographic hashing facilitates efficient data comparison and retrieval because hash values ​​can be quickly compared to verify the authenticity and integrity of the media segment.

[0343] Alternative cryptographic hash algorithms, such as SHA-3 or Blake2, can also be used depending on specific security requirements and computational efficiency considerations. These algorithms offer different levels of security and performance, allowing for flexibility in implementation based on System 2100 requirements. For example, SHA-3, as part of the NIST standard, provides a higher level of security guarantees, while Blake2 is optimized for faster performance and lower computational overhead.

[0344] The hierarchical hash structure 2110, preferably implemented as a Merkle tree, is constructed using the cryptographic hashes of media segments 2106. In this structure, each leaf node represents the hash of a single media segment 2106. Parent nodes generate hashes by concatenating the hashes of their child nodes and then hashing the concatenation results. This process continues until a single root hash is obtained, representing the entire set of media segments 2106.

[0345] The Merkle tree structure offers several significant advantages. It enables efficient and secure verification of the integrity of the entire media content 2102. By distributing the root hash to the user device 2116 at the start of content consumption, the system ensures that the user device can verify the authenticity of any media segment 2106 by traversing the Merkle tree and comparing the calculated segment hash with the corresponding leaf node. This hierarchical approach also allows for partial verification, where only a subset of media segments needs to be verified, significantly reducing the computational load.

[0346] Furthermore, the hierarchical hash structure 2110 supports efficient data synchronization and resynchronization. In scenarios where parts of media content 2102 are lost, corrupted, or unrecognizable, the system can quickly identify and isolate affected segments by comparing their hashes with a Merkle tree. This facilitates targeted resynchronization efforts, such as retransmitting only the affected segments or using AI-based predictive models to estimate the lost hashes.

[0347] Alternative hierarchical structures, such as hash lists or hash chains, can also be considered. While these alternatives may offer simpler implementations, they generally lack the scalability and efficiency of Merkle trees in handling large datasets and supporting partial validation. However, in scenarios with lower security requirements or smaller datasets, these alternatives may provide viable solutions.

[0348] The combination of cryptographic hashing and hierarchical hashing structures in System 2100 ensures the robustness, security, and efficient synchronization of media content 2102. These technologies provide a comprehensive framework for verifying data integrity, facilitating dynamic metadata integration, and enabling seamless user experiences across various content consumption scenarios.

[0349] AI predictive hashing is an optional technique used to maintain synchronization and data integrity in the event that parts of media content 2102 are lost, corrupted, or unidentifiable. This method utilizes artificial intelligence and machine learning algorithms to predict the hash value of problematic segments based on the context and content of adjacent identified segments. The primary goal of AI predictive hashing is to ensure continuous and seamless synchronization, even in the presence of data anomalies or transmission errors.

[0350] The process of generating AI-predicted hashes begins with identifying unidentifiable or missing media segments. 2106 Once these segments are detected, the system analyzes surrounding identified segments to gather contextual information. Machine learning models, such as recurrent neural networks (RNNs) or convolutional neural networks (CNNs), are then used to predict the hash values ​​of the unidentifiable segments. These models are trained on large datasets containing a wide variety of media content, enabling them to accurately infer possible hash values ​​based on patterns and relevances observed in the data.

[0351] The use of AI predictive hashing offers several beneficial advantages. First, it enhances the robustness of the synchronization process by providing a mechanism to handle data anomalies without disrupting the user experience. Second, it improves synchronization accuracy by leveraging AI's predictive capabilities to generate hash values ​​that closely match the original content. Furthermore, AI predictive hashing reduces the need for retransmissions of lost segments, thereby optimizing bandwidth usage and minimizing latency.

[0352] Alternatives to AI-based hash prediction include heuristic-based methods and statistical models. Heuristic-based methods rely on predefined rules and patterns to estimate the hash value of unidentifiable segments. While these methods are simpler to implement, they may lack the accuracy and adaptability of AI-based methods. Statistical models, such as Markov chains or Bayesian networks, can also be used to predict hash values ​​based on probabilistic relationships between segments. However, these models may require significant computational resources and may perform poorly in complex or highly variable content scenarios.

[0353] By leveraging the predictive power of machine learning algorithms, this technology ensures seamless and accurate synchronization even in the presence of data anomalies. The flexibility and adaptability of AI predictive hashing make it a valuable complement to the overall synchronization framework, supplementing other methods such as cryptographic hashing and hierarchical hash structures.

[0354] In one example, the process of implementing AI-driven predictive hashing begins with collecting a comprehensive and diverse dataset. This dataset covers a wide range of audio and video content, including movies, TV shows, music videos, commercials, and standalone audio tracks. The goal is to capture diverse contexts and scenarios, appropriately reflecting the breadth of the end use case where possible. The dataset should be annotated with metadata, including segment hashes, timestamps, and contextual tags (e.g., whether a song is part of a movie or a standalone track). Data should be collected in formats compatible with machine learning frameworks, such as CSV files for metadata and binary files for media content. The dataset should be large enough to ensure that the AI ​​model can learn complex patterns and relevances within the media content. Typically, a dataset containing several terabytes of annotated media content is recommended for achieving high predictive accuracy.

[0355] Choosing the right machine learning model for AI-based hash prediction is crucial for the success of this technology. Given the sequential nature of media content, Recurrent Neural Networks (RNNs) and their variants, such as Long Short-Term Memory (LSTM) networks and Gated Recurrent Units (GRUs), are well-suited for this task. These models excel at capturing temporal dependencies and patterns in sequential data. Furthermore, Convolutional Neural Networks (CNNs) can be used to extract spatial features from video frames, which can be combined with the temporal features extracted by RNNs. Hybrid models integrating RNNs and CNNs can provide a comprehensive understanding of media content, enabling accurate hash prediction. The model architecture should be designed to handle the high dimensionality of video data and the sequential nature of audio data.

[0356] Training an AI model involves feeding a collected dataset into a chosen model architecture. The training process is iterative and requires significant computational resources, typically involving the use of GPUs or TPUs to accelerate training. The model is trained to minimize a loss function that measures the difference between the predicted hash and the actual hash of a media clip. Techniques such as backpropagation and gradient descent are used to update the model parameters. The training process also involves data augmentation techniques, such as adding noise or changing playback speed, to improve the model's robustness and generalization ability. The training phase can take anywhere from several days to several weeks, depending on the size of the dataset and the complexity of the model.

[0357] Once the model is trained, it is beneficial to validate its accuracy and reliability. This involves evaluating the model on a separate validation dataset that was not used during training. The validation dataset should represent the real-world scenarios where the model will be applied. Metrics such as mean squared error (MSE), mean absolute error (MAE), and accuracy are used to evaluate the model's performance. Furthermore, the model's predictions are compared to the true hashes to ensure that the predicted hashes are within an acceptable range of the actual hashes. Cross-validation techniques, such as k-fold cross-validation, can be employed to further validate the model's performance and ensure that it does not overfit the training data.

[0358] After successful validation, the AI ​​model is deployed within system 2100 to generate predictive hashes in real time. During content playback, user device 2116 monitors media content 2102 and calculates segment hashes. If a segment is unidentifiable or missing, the AI ​​model is invoked to predict the hash based on the context provided by adjacent identified segments. The predicted hash is then used to retrieve the corresponding metadata from database 2112, ensuring continuous synchronization and a seamless user experience. The system uses the predicted hash to dynamically adjust synchronization, utilizing techniques such as dynamic time warping to accurately align timestamps. The AI ​​model continuously learns and improves by incorporating feedback from real-world usage, further enhancing its predictive capabilities.

[0359] The implementation of AI predictive hashing involves a comprehensive process of data collection, model selection, training, validation, and application. By leveraging advanced machine learning techniques and large datasets, the system ensures robust and accurate synchronization of media content, even in the presence of data anomalies. This approach enhances the overall user experience by providing seamless and secure interaction with media content.

[0360] Advantageously, system 2100 is configured to distribute the root hash of hierarchical hash structure 2110 to user device 2116. The distribution of the root hash ensures the integrity and synchronization of media content within the proposed system. The root hash, derived from a hierarchical hash structure (e.g., a Merkle tree), serves as a unique and secure identifier for the entire set of media segments. The primary purpose of distributing the root hash is to provide a reference point for verifying the authenticity and integrity of the media content during playback. This process begins at the start of content consumption, when the root hash is transmitted to user device 2116. Distribution can be achieved through various secure channels, such as encrypted communication protocols, the aforementioned bidirectional communication (e.g., WebSocket), or a security token-based system. By providing the root hash at the beginning, system 2100 ensures that user device 2116 can continuously verify the integrity of each media segment 2106 by comparing the calculated segment hash with the corresponding node in the Merkle tree. This mechanism not only enhances security but also facilitates efficient synchronization, as any tampering with media content 2102 can be detected and addressed immediately.

[0361] Dynamic metadata matching is an optional process that enhances the interactivity and contextual relevance of media content during playback. System 2100 utilizes a pre-built database 2112 that stores metadata 2118 associated with various media segments 2106, such as descriptions, products, locations, and other related elements. During playback, user device 2116 calculates a hash for each media segment 2106 and uses this hash to query database 2112 to retrieve the corresponding metadata 2118. This dynamic retrieval ensures that the metadata 2118 is precisely synchronized with the specific video segment 2106 being played. System 2100 optionally employs an algorithm to match the calculated segment hash with the pre-stored metadata, enabling real-time updates and interactivity. For example, if a specific scene in the video showcases a product, the system can dynamically display relevant information or purchase links.

[0362] Contextualization is a beneficial aspect of the systems described in this specification, ensuring that metadata and synchronization processes accurately reflect the specific use cases of media content. Therefore, the contextualization mechanism described for system 2100 applies to all systems described herein. The metadata database 2118 includes context tags that differentiate the various contexts in which content is used. For example, a song may be used as part of a movie, an advertisement, or as standalone audio on a broadcast or streaming service. By tagging content with these context tags, the system can provide precise differentiation and customized interactions based on specific contexts. This differentiation is achieved through a combination of metadata tags and hierarchical hash relationships. The system can identify whether a song is embedded in larger content (such as a movie) or played independently. This ensures that users receive context-relevant information and interactions, thereby enhancing the overall experience. Furthermore, contextualization allows for accurate tracking and reporting of content usage across different platforms, providing valuable insights for content creators and advertisers.

[0363] Hierarchical context association is a method for maintaining relationships between different levels of content within a proposed system. This method involves linking the hashes of independent media segments (e.g., songs) to the hashes of larger content (e.g., movies or advertisements) that embed them. By creating these hierarchical associations, the system can accurately reflect the context of content usage. For example, a song appearing in a movie will have its hash linked to the overall movie hash, allowing the system to distinguish between the song's independent use and its embedded use within the movie. This hierarchical structure ensures that metadata and synchronization processes are context-accurate and relevant. The system can dynamically adjust the interactions and information displayed to users based on specific contexts, providing a seamless and engaging experience. Furthermore, hierarchical context association facilitates efficient content management and tracking, enabling content creators and distributors to monitor and analyze the use of their content across different platforms and contexts. This approach enhances the robustness and scalability of the system, ensuring its ability to handle a wide range of media types and use cases.

[0364] Optionally, user device 2116 uses a sliding hash window technique to enhance the robustness and efficiency of the synchronization process. This method involves continuously rehashing overlapping segments of the media content to enable rapid resynchronization in the event of data loss or corruption. The sliding window operates by moving incrementally across the media content, hashing each segment and its overlapping neighbors. For example, if a segment duration is predefined as 1 second, the sliding window might hash segments such as 0 to 1 second, 0.5 to 1.5 seconds, etc. This overlap ensures that even if a segment is lost or corrupted, adjacent overlapping segments can still be used to verify and maintain synchronization. The sliding hash window technique significantly reduces computational overhead by avoiding rehashing the entire content and focusing only on the affected segments. This method is particularly useful in dynamic streaming environments where network conditions may vary, ensuring a seamless and uninterrupted user experience.

[0365] Recognition algorithms can be optionally integrated into system 2100 to enable real-time recognition of objects, faces, and other subjects within video content 2102. These algorithms utilize advanced machine learning and computer vision techniques to analyze and interpret the visual and auditory elements of the media. The primary goal of the recognition algorithms is to enhance user interaction and synchronization by dynamically associating the recognized subjects with relevant metadata. For example, a facial recognition algorithm can identify actors in a movie, while an object recognition algorithm can detect products or locations. These algorithms typically involve training deep learning models, such as CNNs and RNNs, on large annotated datasets. Examples include real-time object detection using YOLO (You Only Look Once) and facial recognition using OpenFace. Implementing these algorithms requires significant computational resources and careful tuning to achieve high accuracy and low latency. By integrating recognition algorithms, the system can provide rich metadata, enabling features such as interactive product displays, contextual advertising, and enhanced content recommendations.

[0366] Metadata enrichment is an optional process that enhances the basic metadata 2118 associated with media content 2102 by incorporating additional contextual and dynamic information. This process leverages the capabilities of recognition algorithms and pre-built databases to provide a richer and more interactive user experience. For example, when a recognition algorithm identifies a specific product in a video, system 2100 can automatically retrieve and display detailed information about that product, such as its name, price, and purchase link. Similarly, facial recognition can be used to provide background information about actors or characters in a scene. Metadata enrichment also involves dynamically updating metadata based on user interaction and real-time content analysis. This ensures that the metadata remains relevant and engaging, adapting to the user's specific context and preferences. Rich metadata can be stored in structured formats such as JSON or XML to facilitate efficient retrieval and integration with media content. By providing rich metadata, the system enhances the overall user experience, making the content more informative, interactive, and engaging.

[0367] End-to-end encryption is a beneficial security measure, optionally implemented in the system or any other system, to ensure the confidentiality and integrity of hash lists and metadata. This encryption technique involves encrypting data at the source and decrypting it only at the intended destination, preventing unauthorized access or tampering during transmission. The system employs asymmetric encryption algorithms, such as RSA (Rivest-Shamir-Adleman), to securely encrypt the hash lists and metadata. In this process, data is encrypted using a public key, and decrypted using the corresponding private key. This ensures that only authorized devices with the correct private key can access and decrypt the data. End-to-end encryption provides robust protection against a variety of security threats, including man-in-the-middle attacks and data breaches. Furthermore, it ensures the integrity of the hash lists and metadata is maintained, as any tampering with the encrypted data will result in decryption failure. By implementing end-to-end encryption, the system guarantees that the synchronization process remains secure and trustworthy, protecting the privacy and integrity of media content and associated metadata.

[0368] System 2100 offers significant advantages to streaming platforms by ensuring real-time synchronization between interactive features and video playback. The system's ability to divide video content into fixed-duration segments and hash each segment allows for precise synchronization of metadata and interactive elements. For example, during live sporting events, the system can dynamically display player statistics, match highlights, and advertisements synchronized with the video stream. This dual synchronization mechanism, combining ultrasonic metadata and cryptographic hashing, ensures synchronization is maintained even in the event of network fluctuations or data loss. This robust synchronization capability enhances the user experience by providing seamless and interactive content consumption, making streaming platforms more engaging and informative.

[0369] System 2100 revolutionizes e-commerce integration by seamlessly linking product showcases within video content with actionable consumer interactions. Leveraging dynamic metadata matching, the system identifies products displayed in video clips and retrieves relevant information from a pre-built database. For example, during a live fashion show, the system can display product details, prices, and purchase links for clothing worn by models. Hierarchical contextualization ensures that products are accurately tagged, whether appearing as standalone items or as part of a larger context such as a film or advertisement. This capability not only enhances the viewer experience but also provides a direct and efficient pathway to e-commerce transactions, driving sales and increasing revenue for content creators and advertisers.

[0370] In the gaming industry, System 2100 ensures synchronized in-game actions and live video content, providing players with a more immersive and interactive experience. The system's ability to hash video segments and dynamically match metadata allows for real-time updates and interactions based on game progress. For example, during a live esports tournament, the system can display player profiles, game statistics, and sponsor advertisements synchronized with gameplay. Sliding hash window technology facilitates rapid resynchronization in case of data loss or network issues, ensuring consistent synchronization throughout the live stream. This feature enhances viewer engagement and provides valuable insights and information, making the gaming experience more enjoyable and interactive.

[0371] System 2100 significantly enhances educational content by providing interactive video tutorials with synchronized Q&A sessions. The system's ability to segment and hash video content ensures precise synchronization between educational videos and supplementary materials such as quizzes, annotations, and interactive exercises. For example, during online lectures, the system can dynamically display relevant questions and prompts based on the video content being played. Context-sensitive features ensure that educational content is accurately labeled and differentiated based on its usage context (whether it's part of a larger course or a standalone tutorial). This feature enhances the learning experience by providing interactive and context-sensitive content, making education more engaging and effective.

[0372] System 2100 optionally includes real-time topic recognition, which enhances interactivity through real-time identification and interaction of topics within video content. By implementing AI-based recognition algorithms, System 2100 can identify objects, faces, and other topics in the video and dynamically associate them with rich metadata. For example, during a live concert stream, the system can identify performing artists and display their biographies, album catalogs, and social media links in real time. This feature not only enhances the viewer experience but also provides valuable information and interaction, making the live content more engaging and informative. The system's ability to handle unidentifiable portions using AI prediction and dynamic time warping algorithms ensures seamless synchronization, even in the presence of data anomalies.

[0373] System 2100 ensures accurate context differentiation and seamless identification in scenarios where songs transition between embedded use and standalone playback. The system uses temporal hash clustering to detect whether a song hash appears as part of a broader context or as standalone playback. For example, if a song hash aligns with a movie content hash, it is marked as embedded. If it has no parent context hash, it is marked as standalone. This feature ensures the system accurately differentiates between various use cases, providing precise synchronization and metadata matching. A dual synchronization mechanism, including ultrasonic metadata as a backup system, ensures continuous synchronization even if hashing may be temporarily interrupted. This robust and flexible synchronization capability enhances the user experience by providing seamless and context-accurate content consumption.

[0374] Figure 22 A system architecture diagram of a dual-synchronization system is shown. Specifically, Figure 22 A dual synchronization system 2200 is illustrated, including components of system 2100, a speaker 2202, and an audio tag 2204. The speaker 2202 is configured to emit an audio signal associated with media content 2102, for example, the speaker 2202 is communicatively connected to or integrated into a transmission device 2214, and the audio tag 2204 is embedded in the audio signal, as described in detail above. Advantageously, the dual synchronization system is configured to dynamically switch between video hashing of system 2100 and, for example, ultrasonic synchronization of content synchronization system 300.

[0375] The switching within the dual synchronization system 2200 is a beneficial mechanism that ensures continuous and seamless synchronization of media content, even in the presence of interruptions or differing network conditions. The system dynamically switches between video hashing (as implemented in system 2100) and ultrasonic synchronization (as utilized in content synchronization system 300). This switching process is initiated based on predefined criteria, such as the integrity of the hash verification process or the presence of ultrasonic metadata. When the video hashing mechanism detects a discrepancy or is unable to verify the integrity of a media segment due to data loss or corruption, the system seamlessly switches to ultrasonic synchronization. This is achieved by utilizing an audio tag 2204 embedded in the audio signal emitted by speaker 2202. The audio tag preferably provides a coarse-grained synchronization reference, allowing the system to maintain alignment with the media content. A key advantage of this dynamic switching capability is the fault-tolerant mechanism it provides, ensuring synchronization is maintained even under adverse conditions. This enhances the robustness and reliability of the system, providing an uninterrupted user experience.

[0376] Error detection is a fundamental aspect of the Dual Synchronization System 2200, ensuring the integrity and accuracy of the synchronization process. The system optionally employs a multi-layered error detection mechanism to identify and handle discrepancies in media content. The primary method involves using cryptographic hashing, where each media segment is hashed and the resulting hash value is compared to a hierarchical hash structure (e.g., a Merkle tree). Any mismatch between the calculated hash and the expected hash indicates a potential error, prompting the system to initiate corrective actions. Furthermore, the system incorporates error correction algorithms (e.g., Reed-Solomon) to detect and correct small data losses during video transmission. These algorithms add redundancy to the data, allowing the system to recover lost or corrupted segments. The dual synchronization mechanism further enhances error detection by providing an alternative synchronization reference (via ultrasonic metadata). If the video hashing process fails to verify the integrity of a media segment, the system switches to ultrasonic synchronization, ensuring that errors are detected and handled promptly. This multi-layered error detection approach enhances the reliability and accuracy of the synchronization process, ensuring a seamless user experience.

[0377] System 2200 optionally employs one or more advanced technologies to prevent and detect tampering. In addition to using cryptographic hashing (where each media segment is hashed using a secure algorithm such as SHA-256), the system uses blockchain technology or a distributed ledger to store the root hash. This ensures the root hash is immutable and tamper-proof, providing a trusted reference for verification. Furthermore, the system employs end-to-end encryption to protect the hash list and metadata securely during transmission, preventing unauthorized access or tampering. By combining these advanced security measures, the dual synchronization system 2200 ensures media content remains secure and trustworthy, providing a robust and reliable synchronization framework.

[0378] Scalability is a key design consideration for the Systems 2100 and 2200, ensuring the system can effectively manage large amounts of data and support a growing user base. The system architecture is configured to handle various media types and use cases, leveraging distributed computing resources to process and store hash lists and metadata. Hierarchical hash structures (such as Merkle trees) and sliding hash windows minimize computational overhead, allowing the system to scale efficiently. Furthermore, the use of cloud-based storage and processing solutions enables the system to dynamically allocate resources based on demand, ensuring it can handle peak loads and large-scale deployments. This scalability ensures the system can provide robust and reliable synchronization across different content platforms and user groups.

[0379] Figure 23 A flowchart illustrating a method for processing video content is shown. Specifically, the flowchart illustrates method 2300 for processing known media content.

[0380] Method 2300 begins with step 2302, which involves segmenting known media content into multiple media segments. This is achieved by dividing the media content into segments of fixed duration (e.g., 1-second intervals). A predefined duration is chosen to balance the need for efficient synchronization with the requirement for accurate identification of media segments. This segmentation process ensures that each segment is manageable in size and can be processed, hashed, and verified independently. The segmentation is performed on a workstation device, which can be a server or a dedicated processing unit, ensuring that the media content is systematically divided into discrete, time-constrained segments.

[0381] In step 2304, each of the multiple media segments is hashed using a cryptographic algorithm (e.g., SHA-256). This hashing process generates a unique hash value for each segment, called a segment hash. The segment hash acts as a digital fingerprint of the media segment, ensuring that even slight changes to the segment content will result in significantly different hash values. Each segment hash is associated with known media content at a specific time interval, providing a precise and secure method for verifying the integrity and authenticity of the media content. The use of cryptographic hashes ensures that the segment hashes are tamper-proof and can be reliably used for synchronization and verification purposes.

[0382] Step 2306 involves arranging multiple segment hashes at one or more lower levels of a hierarchical hash structure (such as a Merkle tree). In this structure, each leaf node represents a segment hash. The hierarchical hash structure is configured by combining hash values ​​from lower levels to generate hash values ​​for higher levels. This process involves concatenating the hash values ​​of child nodes and hashing the concatenation results to generate the hash value of the parent node. This hierarchical arrangement continues until a single root hash is generated, representing the entire set of media segments. The root hash acts as a unique identifier for the known media content across all time intervals, providing a secure and efficient method for verifying the integrity of the entire media content.

[0383] In step 2308, multiple segment hashes are associated with data containing one or more context tags. These context tags are used to distinguish known media content and provide additional metadata for each segment. Context tags may include information such as content type (e.g., movie, advertisement, standalone audio), specific use case (e.g., part of larger content or standalone), and other relevant metadata (e.g., description, product, location). This association ensures that each segment hash is enriched in context, allowing for accurate differentiation and dynamic metadata matching during playback. Context tags are stored in a pre-built database, enabling the system to retrieve and display relevant information in real time, enhancing the user experience and providing valuable insights into the media content.

[0384] Method 2300 uses hashing on workstation devices to ensure robust, secure, and efficient processing of media content. The segmentation, hashing, hierarchical arrangement, and contextual association of media content provide a comprehensive framework for verifying the integrity, authenticity, and contextual relevance of media content, ensuring seamless synchronization and enhanced user interaction. Optionally, Method 2300 includes additional steps from any previous method, particularly those related to processing media content for integration or synchronization.

[0385] Figure 24A flowchart of a method for synchronization based on video hashes is shown. Specifically, the flowchart of method 2400 is used to synchronously display a video hash generated by a user application installed on a user device on the user device.

[0386] Method 2400 begins with step 2402, which involves receiving the root hash of a hierarchical hash structure associated with the media content. This root hash is a unique identifier representing the entire set of media segments, ensuring the integrity and authenticity of the media content. The hierarchical hash structure (e.g., a Merkle tree) is constructed by combining hash values ​​from lower levels to generate hash values ​​from higher levels, ultimately forming the root hash. The root hash is transmitted to the user device at the start of content consumption, providing a reference point for verifying the integrity of the media content throughout playback.

[0387] In step 2404, the user device retrieves data from the database associated with the root hash. This data includes contextual tags used to distinguish known media content and content information for media content interaction. Contextual tags provide information about the specific use case of the media content, such as whether it is part of a movie, an advertisement, or standalone audio. Content information includes metadata such as descriptions, products, locations, and other embedded elements that enhance the user experience. By associating the root hash with this data, the system ensures that the media content is context-rich and ready for dynamic interactions during playback.

[0388] Step 2406 involves monitoring the media content at a first timestamp for a predetermined duration to obtain a first media segment. The predetermined duration is typically a fixed interval, such as 1 second, ensuring that each segment is manageable in size and can be processed individually. The user device continuously monitors the media content, capturing segments at specific timestamps to facilitate synchronization and verification. This step ensures that the media content is systematically divided into discrete, time-bound segments that can be hashed and compared with a hierarchical hash structure.

[0389] In step 2408, the user equipment verifies synchronization by calculating a hash of the first media segment and comparing that hash with a hierarchical hash structure. The calculated hash is matched against the corresponding hash in the Merkle tree to verify the integrity and authenticity of the media segment. If the calculated hash matches the expected hash, synchronization is verified, indicating that the media content has not been tampered with or corrupted. This verification process ensures that the media content remains secure and reliable throughout playback.

[0390] If synchronization is verified, the method proceeds to step 2410, where the user device uses the hash to retrieve and display content information associated with the first timestamp. This content information is dynamically matched and retrieved from a pre-built database, providing the user with relevant metadata and interactive elements. For example, if a media clip corresponds to a scene in a movie, the system can display information about the actors, locations, or products in that scene. This dynamic retrieval and display of content information enhances the user experience by providing context-relevant and interactive content.

[0391] If synchronization is not verified, the method proceeds to perform corrective measures to maintain synchronization. In step 2412, the user equipment adjusts the first timestamp to resynchronize and display content information associated with the adjusted first timestamp. This adjustment involves using an algorithm such as DTW to accurately align the timestamps, ensuring that media content remains synchronized despite any discrepancies. By adjusting the timestamps, the system can recover from minor synchronization errors and continue to provide a seamless user experience.

[0392] As an alternative to or supplement to step 2412, in step 2414, the system can use a trained machine learning model to generate a predicted hash based on one or more adjacent media segments. This method leverages AI-based prediction to estimate the hash of unidentifiable or corrupted segments, enabling the system to retrieve and display content information associated with the predicted timestamp. The machine learning model is trained on simulated or historical datasets, enabling it to accurately predict hashes based on the context provided by adjacent segments. This predictive capability ensures continuity and synchronization even in the presence of data anomalies.

[0393] Step 2416 involves monitoring the media content for a predetermined duration at a second timestamp, the second timestamp being within the predetermined duration from the first timestamp. This step ensures that the system captures overlapping media segments, facilitating continuous monitoring and synchronization. By acquiring a second media segment that overlaps with the first media segment, the system can more effectively maintain synchronization and verify the integrity of the media content.

[0394] In step 2418, the user equipment calculates the hash of the second media segment, thereby rehashing the overlapping media segments to maintain synchronization. This sliding hash window method ensures that the system continuously verifies the integrity of the media content, even if some segments are lost or corrupted. By rehashing overlapping segments, the system can quickly resynchronize and maintain a seamless user experience. This method ensures robust and reliable synchronization, providing users with secure and interactive media content consumption.

[0395] Systems 2100 and 2200, along with methods 2300 and 2400, are preferably interconnected to form a comprehensive framework for media content synchronization. System 2100 focuses on hashing the media content, dividing it into segments of fixed duration, hashing each segment, and constructing a hierarchical hash structure (e.g., a Merkle tree) to ensure secure and efficient synchronization. Method 2300 details the steps involved in this process, from segmenting the media content to associating segment hashes with contextual metadata. Building upon System 2100, System 2200 incorporates a dynamically switching video hash and an ultrasonic-based dual synchronization mechanism. This system includes additional components, such as speakers and audio tags, for transmitting and decoding ultrasonic signals for synchronization. Method 2400 outlines the steps for maintaining synchronization on the user device, including receiving the root hash, monitoring the media content, verifying synchronization, and handling unrecognized segments through AI prediction and timestamp adjustment. In summary, these systems and methods provide a robust, secure, and interactive framework for media content synchronization, ensuring a seamless user experience across various platforms and scenarios.

[0396] Systems 2100 and 2200, as well as methods 2300 and 2400, are closely related to systems and methods for synchronizing ultrasonic signals (e.g., systems and methods including audio tag-based processes). The dual synchronization mechanism in system 2200 utilizes ultrasonic metadata as a backup system to maintain coarse-grained synchronization when video hashing encounters problems. This integration ensures continuous synchronization by dynamically switching between ultrasonic-based and hash-based synchronization. The ultrasonic signal emitted by the speaker and decoded by the user equipment provides a reliable synchronization reference, complementing the cryptographic hashing method. This combination enhances the robustness and reliability of the synchronization process, ensuring that media content remains synchronized even in the event of data anomalies or network fluctuations.

[0397] The disclosed systems and methods also involve audio hashing and audio fingerprint-based synchronization. The hashing techniques used in systems 2100 and 2200 and detailed in methods 2300 and 2400 can be applied to audio content as well as video content, or any other form of media content, such as paintings in a gallery, where synchronization is based on location rather than time. By segmenting audio content, hashing each segment, and constructing a hierarchical hash structure, the system can achieve accurate synchronization of audio content across various platforms. Similar to audio fingerprinting (where unique identifiers are generated for audio segments to facilitate synchronization and identification), all described processes can be combined into a comprehensive content integration and synchronization system. For example, communication between devices uses a bidirectional communication protocol as described above to securely retrieve content information based on hashes, audio fingerprints, or audio tags 2118. The integration of AI-based prediction and dynamic metadata matching further enhances the system's ability to handle unrecognized audio segments and provide context-sensitive information. This comprehensive framework leverages the advantages of cryptographic hashing and audio fingerprinting techniques to ensure accurate and reliable synchronization of audio and video content.

[0398] Optionally, the system disclosed herein includes one or more AI models integrated into a structured pipeline to enhance the extraction, analysis, and synthesis of content-related information. Unlike traditional systems that typically operate AI models in isolation, each AI model in this system operates as a modular component within a larger network, allowing features extracted from one model to inform subsequent models. This architecture not only improves accuracy and efficiency but also generates new, derived insights that are difficult to obtain through independent AI models.

[0399] For example, in the context of media content recognition, a primary AI model may first extract audio features from broadcast media clips using techniques such as frequency analysis, spectral decomposition, or waveform analysis. These extracted features, which might include spectral peaks, temporal energy distributions, or Mel-frequency cepstral coefficients (MFCCs), are then fed into a secondary AI model trained for feature matching against a database of known media clips. In scenarios where no exact match is found, an additional AI model can be applied, leveraging machine learning-based inference (such as convolutional neural networks (CNNs) or Transformer-based architectures) to identify the closest match by considering temporal patterns, speech similarity, or semantic relevance.

[0400] This cascaded approach allows for more robust media content recognition in challenging environments, such as those with background noise or partially blurred audio. Furthermore, the system benefits from self-learning capabilities, where feedback from mismatches or user interactions refines the AI ​​model's performance over time. This can be achieved through reinforcement learning techniques or adaptive retraining mechanisms that prioritize frequently encountered misidentifications for further fine-tuning.

[0401] In addition to media recognition, the system may optionally include content synchronization and information retrieval. A three-tiered AI model can be tasked with extracting relevant contextual elements from the recognized media, such as objects, people, or locations appearing in the scene. This is particularly relevant for enhancing interactive experiences in applications such as smart TV systems, second-screen experiences, or augmented reality (AR) overlays. For example, once a media clip is identified, an AI-driven entity recognition model can analyze captions, closed captions, or visual metadata to identify key themes or thematic elements. If a film scene contains a recognizable landmark, such as the Eiffel Tower, the AI ​​model can link this visual entity to a structured knowledge graph to retrieve supplementary data about its history, significance, or relevant cultural references.

[0402] The system's scalability and modularity are further enhanced by its loosely coupled architecture, allowing for seamless expansion and adaptability across different domains. Each AI component operates independently while still contributing to a unified workflow, ensuring that individual models can be updated or replaced without disrupting the entire system. This is particularly advantageous when integrating new AI capabilities, such as sentiment analysis, speech-to-text conversion, or real-time translation, enabling the system to continuously evolve and remain at the forefront without requiring a complete system overhaul.

[0403] Another implementation of this AI integration pipeline is in security and authentication applications, where AI models work together to verify user identity based on biometric or behavioral data. For example, a first AI model might process voice biometrics from a media stream, while a secondary AI model cross-references this data with stored voice signatures. If a discrepancy is detected, an additional AI model can be employed to apply anomaly detection techniques to analyze contextual factors such as background noise or user intent, improving verification accuracy. This approach ensures multi-layered security, reduces false positives, and improves overall system resilience.

[0404] Furthermore, the system utilizes bidirectional communication and data exchange mechanisms to optimize performance. This is particularly evident in real-time content synchronization, where alignment between media playback timestamps and supplementary content must be maintained. In this context, the AI ​​model monitors real-time changes in playback speed, buffering events, or network latency, dynamically adjusting the retrieval and presentation of relevant content (such as player statistics or event highlights) to maintain contextual relevance. For example, if a live sporting event experiences delays due to streaming congestion, the AI ​​system can recalibrate secondary content to maintain contextual relevance.

[0405] Furthermore, the integration of hierarchical hash structures and encryption technologies enhances the system's reliability in verifying content authenticity and mitigating potential tampering. The AI ​​model can generate a unique audio / video fingerprint for each media segment, embedding it in an encrypted chain to ensure that the retrieved content aligns with the original index reference. This has key applications in media copyright management, digital watermarking, and forensic content tracking.

[0406] From a deployment perspective, the system is designed to operate efficiently on a distributed architecture, supporting cloud-based processing and edge computing scenarios. This flexibility allows computationally intensive AI tasks (such as deep learning inference or large-scale data matching) to be offloaded to cloud servers, while lightweight AI models handling real-time interactions can be executed directly on user devices. This hybrid deployment strategy optimizes latency and computational efficiency, ensuring high-priority tasks are processed with minimal latency, while resource-intensive operations benefit from the scalability of the cloud.

[0407] In one exemplary optional use case, interactive media content discovery allows users to use a smartphone app to identify television programs playing in the background. The system captures audio clips, extracts key features, and queries a remote database to retrieve possible matches. If an exact match is found, the app presents the user with contextual information such as actor details, behind-the-scenes footage, and product placement. If no direct match is available, the AI ​​system suggests alternative results in probability rankings, providing the user with informed options rather than outright failure. Furthermore, user interactions (such as confirming or rejecting suggested matches) are recorded to refine the accuracy of the AI ​​model over time.

[0408] Another alternative implementation extends to smart home environments, where AI models orchestrate cross-device synchronization between multimedia devices. For example, if a user starts watching content and then moves to another room, an AI-driven system can facilitate seamless playback continuation on secondary devices, adjusting audio and video parameters to suit different speaker configurations and display settings, such as using any of the methods, systems, or processes described above.

[0409] Benefically, the disclosed system represents a significant advancement in AI-driven content interaction by seamlessly integrating multiple AI tools within a scalable, modular, and loosely coupled framework. This integration enhances content recognition, synchronization, and contextual information retrieval, while ensuring adaptability across diverse applications. By enabling AI models to interact dynamically and refine their outputs based on multimodal inputs, the system achieves greater accuracy, improved user engagement, and enhanced operational efficiency. These optional applications are broadly applicable across industries ranging from entertainment and media to security, e-commerce, and smart home automation.

[0410] This invention extends previously disclosed media interaction systems by combining hybrid radio metadata enrichment, bidirectional device synchronization, and AI-driven topic extraction for broadcasting. The combination of real-time and pre-programmed metadata retrieval methods ensures accurate, dynamic content engagement while maintaining secure, efficient, and low-latency processing. The disclosed system further leverages the previously described bidirectional synchronization techniques to optimize metadata alignment and audience interaction, covering both live and archived broadcast content. These advancements support a unified, scalable approach to interactive broadcast content consumption, reinforcing the principles of seamless media interaction, secure communication, and synchronized metadata enrichment.

[0411] For example, a two-way connection is established between workstations, AI components, and RadioDNS to enhance live and pre-programmed broadcasts. This integration relies on several key protocols and technologies. RadioDNS Hybrid Radio provides metadata links for live broadcasts, while Service and Program Information (SPI - ETSI TS 102 818) enables broadcasters to structure metadata about stations, programs, and content. RadioEPG (Electronic Program Guide - ETSI TS 102 818) synchronizes pre-programmed content with metadata, and RadioTAG (ETSI TS 103 270) allows listeners to tag audio content for future interaction, whether live or on demand. Additional technologies include DAB Slideshow and RadioVIS (ETSI TS 101 499), which support visual enhancements such as advertisements, images, and contextual metadata, as well as RTSP and HLS protocols for streaming broadcasts, ensuring synchronized metadata overlay.

[0412] Workstations and AI preferably interact with live and pre-programmed broadcasts in multiple ways. Metadata enrichment for live broadcasts leverages RadioDNS SPI and AI-enhanced tagging to introduce real-time metadata, while pre-programmed broadcasts rely on RadioEPG to preload metadata. Content recognition for live broadcasts involves AI listening to audio fingerprints and hashes of the live stream, while pre-recorded content is tagged and enriched by AI scanning files. Interactive features, such as search, purchase, booking, sharing, commenting, betting, and rating, are available in real-time via an SDK linked to RadioTAG, and for pre-programmed content, planned and stored interactions, similar to podcast engagement, are supported. Optionally, the system includes live commercial and / or real-time ad insertions using RadioVIS and RadioEPG integrations, while pre-tagged content facilitates e-commerce and booking integrations. Furthermore, dashboards for content creators and media owners provide a real-time interface for real-time enrichment and content synchronization, and offer AI-driven suggestions for pre-programmed tagging.

[0413] To effectively manage live and pre-programmed broadcasts, the workstation dashboard preferably includes the following features. Real-time content enrichment for live broadcasts involves an AI listener that monitors live DAB+ or streaming broadcast sources and matches audio fingerprints with metadata. RadioDNS API integration retrieves station metadata, program details, and service information, while automatically synchronized tags and contextual data enriches content with AI-generated topics, keywords, and business links. RadioTAG event handling allows users to tag moments in broadcasts for later engagement, and automated call action (CTA) triggers introduce interactive options such as "Buy Now," "Book Now," or "Learn More." Furthermore, real-time topic insertion allows for enhanced content control.

[0414] For pre-programmed content management, the dashboard integrates with RadioEPG to load planned program metadata. Batch processing of pre-recorded content uses AI to scan the database to find relevant topics, and marketplace links pre-assign topics to pre-recorded material. A hybrid interactive management system ensures seamless switching between live and pre-programmed content, maintaining a consistent interactive experience regardless of format. API support for third-party application integration connects with broadcast apps, in-vehicle dashboards, and smart home devices. Furthermore, user behavior analytics and reporting features track listener engagement, conversion rates, and the success rate of various interactions.

[0415] The technical workflow is preferably divided into two main processes: live broadcasting and pre-programmed broadcasting. In the live broadcasting process, the RadioDNS API retrieves station and program metadata, while AI listens for and matches content by identifying relevant topics, people, products, and locations. Workstations then synchronize commerce, ratings, and user interactions, allowing listeners to buy, book, comment, share, and bet in real time through consumer-facing applications. In the pre-programmed broadcasting process, RadioEPG preloads planned content, and AI scans audio segments to tag context-relevant information. Interactive elements are preloaded to enable engagement after broadcast, allowing users to re-interact with archived content, purchase products, book experiences, and participate in other post-broadcast activities.

[0416] Other aspects are described in detail below.

[0417] One aspect of the invention provides a modular and scalable approach for integrating metadata enrichment, real-time content interaction, and secure device synchronization for live and pre-programmed broadcasts. This integration leverages hybrid radio standards and streaming protocols to facilitate rich audience engagement through bidirectional communication between workstation platforms, AI-based richness modules, and interactive media retrieval systems.

[0418] According to one aspect of the invention, a system is provided for combining live and pre-programmed broadcast content with real-time metadata extraction, content enrichment, and interactive engagement. The system is configured to establish a bidirectional connection between workstation equipment and a hybrid radio metadata provider. The system utilizes secure communication protocols to ensure the synchronous retrieval and transmission of radio metadata, program information, and enriched content topic data between the workstation equipment and user equipment. This facilitates real-time synchronization, interactivity, and business enablement of broadcasts, whether the content is consumed live or accessed as archived recordings.

[0419] For example, a workstation device configured for hybrid radio integration may optionally receive a live metadata stream from a hybrid radio metadata provider via the RadioDNS API. This metadata includes station identification data, programming schedules, content descriptions, and live audio tagging information. The workstation device processes this data in conjunction with an AI-driven content enrichment module that dynamically links the metadata to contextual topics, such as artists, topics, products, locations, or events mentioned in the broadcast. This enriched metadata is then synchronized with the live broadcast stream and transmitted to the user device via a bidirectional communication protocol.

[0420] In an optional implementation, the DAB+ digital radio receiver within the workstation continuously retrieves metadata from the ETSITS 102 818-compliant Electronic Program Guide (RadioEPG) to provide a structured timeline of upcoming programming. This timeline is pre-loaded into the workstation, allowing the AI ​​enrichment module to pre-tag topics associated with the planned content before programming airs. This pre-tagning process enhances interactive engagement, enabling users to access additional contextual information, such as biographies, product links, or historical references, without post-broadcast processing.

[0421] Alternatively, in environments where DAB+ signals are unavailable or unreliable, the system can optionally utilize HLS or RTSP streaming protocols to receive broadcast streams from internet-based sources. In this case, audio fingerprinting technology is applied to the incoming stream, matching the content against a pre-indexed broadcast segment database. If a match is detected, the system retrieves the corresponding metadata from the workstation database and aligns it with the current broadcast timeline. This backup mechanism ensures that rich metadata remains available even in scenarios where direct integration with hybrid radio providers is not possible.

[0422] Another example of optional system operations is in the context of live audience interaction via the RadioTAG standard (ETSI TS 103 270). When a listener engages with a broadcast by tagging a moment of interest, the system transmits this tag to a central workstation via a two-way communication protocol. The workstation then processes the tag and retrieves the associated metadata, allowing listeners to access additional details, save the content for later use, or trigger interactive actions such as purchasing products mentioned in the broadcast, booking related events, or subscribing to the broadcaster's updates. In a collaborative environment, multiple listeners may tag the same segment, enabling real-time audience interaction aggregation and enhancing broadcaster analytics.

[0423] In an alternative implementation, the system can extend its functionality beyond live broadcasting to include on-demand and archived broadcast content. For pre-programmed content, an AI enrichment module scans audio files before publication, extracting speech-to-text transcriptions, keyword associations, and contextual topic mappings. These enriched segments are then stored in a workstation database, ensuring that users participating in archived content receive the same level of interactivity as those listening to live broadcasts.

[0424] Benefically, by implementing rich, hybrid radio metadata integration driven by real-time AI, the system significantly enhances listener engagement, content accessibility, and broadcaster business opportunities. Seamless synchronization between the metadata provider, AI processing module, and user-facing applications ensures relevant content is presented to users in real time, minimizing latency and maximizing interactivity. By supporting live and pre-programmed broadcasts, the system provides a consistent and interactive experience, regardless of when content is consumed.

[0425] The use of a two-way communication protocol ensures that metadata retrieval, content synchronization, and audience interaction are maintained in real time. Unlike traditional broadcast metadata systems (which are typically one-way and limited to static program guides), the system disclosed in this paper allows for continuous two-way engagement between broadcasters and listeners. This is particularly advantageous in applications requiring business integration, audience analytics, and dynamic content adaptation.

[0426] Furthermore, the system's modularity allows for seamless integration with third-party platforms, in-vehicle infotainment systems, and smart home devices. By supporting standardized hybrid radio protocols such as SPI, RadioEPG, and RadioTAG, as well as AI-driven metadata enrichment, the system ensures that broadcasters and content creators can maximize the value of their content across multiple distribution channels.

[0427] Furthermore, the system's backup mechanisms, including audio fingerprinting and index metadata retrieval, provide resilience against connectivity issues, regional broadcast limitations, and metadata inconsistencies. By combining embedded metadata retrieval with audio-based content recognition, the system ensures high-precision synchronization, even in challenging broadcast environments.

[0428] In summary, the integration of hybrid radio metadata and rich interactive media provides a robust, scalable, and future-proof solution for enhancing live and on-demand broadcast experiences. Whether implemented for traditional broadcast environments, digital streaming platforms, or hybrid radio ecosystems, this system ensures seamless and intuitive delivery of metadata-rich, interactive, and business-enabled content engagement to users.

[0429] Optionally, the system implements a bidirectional communication protocol that enables seamless synchronization between live and pre-programmed broadcasts and rich metadata for interactive engagement. Hybrid radio metadata is retrieved through a service-oriented interface that integrates with external metadata providers, such as those conforming to the ETSI RadioDNS hybrid radio specification. Specifically, the system supports one or more of the following: Service and Program Information (SPI - ETSI TS 102 818): Provides metadata about radio stations, programs and content arrangements.

[0430] RadioEPG (Electronic Program Guide - ETSI TS 102 818): Facilitates the preloading of metadata for pre-programmed content, allowing for context-rich and interactive scheduling.

[0431] RadioTAG (ETSI TS 103 270): Enables listeners to tag audio segments in real time for future participation, supporting interactive elements across live and archived broadcasts.

[0432] DAB Slides with RadioVIS (ETSI TS 101 499): Enhance content with synchronized visual metadata such as images, ads, and interactive theme overlays.

[0433] RTSP and HLS streaming protocols: support synchronous audio streaming processing, ensuring that metadata overlay is always aligned with the timing of real-time broadcast.

[0434] The bidirectional communication protocol further ensures that the system operates independently of proprietary str...

Claims

1. A method for media content identification and content information retrieval on a user device, the method comprising: Use the user's device microphone to listen to a portion of the audio track emanating from the speaker, where the audio track is associated with the media content; Detect whether an audio tag is embedded in the audio track in that section, wherein the audio tag is embedded at a frequency that is detectable by the microphone and imperceptible to the user of the user device; If an audio tag is detected: extract identification data from the audio tag, wherein the identification data is associated with known media content at a known point in time; If no audio tag is detected: Identification data is received based on the best match from multiple audio fingerprints, wherein the audio fingerprints include one or more audio features extracted from that portion of the audio track; Identify media content based on known media content in the identification data; as well as Retrieve and identify content information associated with the data.

2. The method according to claim 1, characterized in that, The search results include: Send a request for information containing identification data to the server device; and Retrieve information from a server device that is associated with one or more topics related to known media content at a known point in time, wherein each of the one or more topics is linked to that known point in time.

3. The method according to claim 2, characterized in that, The information request is sent based on the interaction between the user and the user device.

4. The method according to claim 1, characterized in that, The speaker is associated with a media device that is independent of the user device.

5. The method according to claim 1, characterized in that, The method further includes: Identifying the time point of media content based on known time points in the identification data; and The content information is displayed on the user's device based on this point in time, so that the displayed content information is related to the media content at that known point in time.

6. The method according to claim 1, characterized in that, When listening to this section of the audio track, the microphone is turned on or off at periodic intervals.

7. The method according to claim 1, characterized in that, The audio tags are embedded using frequency shift keying modulation at set time intervals.

8. The method according to claim 7, characterized in that, Adjust the set time interval after identifying the media content.

9. The method according to claim 1, characterized in that, The known time points are predetermined time intervals.

10. The method according to claim 1, characterized in that, The plurality of audio fingerprints include: An audio fingerprint associated with multiple portions of an audio track at multiple points in time; or An audio fingerprint associated with one or more portions of multiple audio tracks.

11. The method according to claim 1, characterized in that, The data received based on the best match includes: Transmit the audio fingerprint or a portion of the audio track to a server device, wherein the audio fingerprint is generated by extracting one or more audio features from the portion of the audio track; and The system retrieves identification data associated with the best match from a server device, which is configured to compare the audio fingerprint with multiple audio fingerprints stored in a database accessible to the server device to determine the best match.

12. The method according to claim 1, characterized in that, Detecting whether audio tags are embedded in this section of the audio track involves applying one or more correction algorithms to compensate for noise, interference, or other audio degradation factors.

13. The method according to claim 1, characterized in that, The detection of the presence of audio tags is limited to a predefined time range.

14. A computer-readable medium storing instructions for performing the method of any one of claims 1-13.

15. A user equipment, comprising: microphone; user interface; as well as A user application installed on a user device provides instructions that, when executed, cause the user device to perform the method described in any one of claims 1-13.

Citation Information

Patent Citations

  • Method and apparatus for the insertion of audio cues in media files by post-production audio and video editing systems

    US9837127B2