Independent routing transmission method and device for heterogeneous audio and video, and storage medium
By decoding and independently transmitting video and audio data in an HDMI matrix, the problem of traditional HDMI matrices being unable to independently select video and audio sources is solved, enabling synchronous output of heterogeneous audio and video sources and reducing the requirements and costs of peripheral hardware.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HDCVT TECH CO LTD
- Filing Date
- 2026-03-31
- Publication Date
- 2026-04-28
AI Technical Summary
Traditional HDMI matrices cannot independently select and maintain synchronized output of video and audio sources internally, requiring an external FPGA or dedicated HDMI audio/video de-embedding chip for separation.
By receiving HDMI input signals from different source devices, decoding them into video and audio data, and transmitting them to the matrix switching module through independent first and second channel groups respectively, the channel group is determined according to the target output terminal and an audio-video reconstruction signal is generated. The output module performs signal restoration and embedding, thereby realizing independent routing of heterogeneous audio and video.
Independent selection and synchronous output of video and audio are achieved within the matrix, reducing the need for external de-embedding devices, lowering transmission costs and complexity, and ensuring audio and video synchronization.
Smart Images

Figure CN121940499A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio and video signal processing technology, and in particular to a method, device and storage medium for independent routing transmission of heterogeneous audio and video. Background Technology
[0002] Traditional high-speed serial high-definition multimedia interface (HDMI) matrix transmission schemes follow the official native encapsulation protocol of the high-definition multimedia interface (hereinafter referred to as HDMI), encapsulating video data, audio data and control / auxiliary information together to form a high-speed serial data stream. This high-speed serial data stream is routed into the matrix through the same set of differential channels.
[0003] In the above scheme, if video and audio are to be obtained separately from different source devices, an additional field-programmable gate array (FPGA) or HDMI dedicated audio and video de-embedding chip needs to be connected after each output end of the matrix to perform de-embedding processing to separate the mixed-package audio and video data. It is not possible to achieve independent selection of video source and audio source and maintain synchronous output within the matrix.
[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a method, device and storage medium for independent routing of heterogeneous audio and video, which aims to solve the technical problem that traditional HDMI matrices bind audio and video routing within the matrix, thus making it impossible to independently select video and audio sources and maintain synchronous output within the matrix.
[0006] To achieve the above objectives, this application proposes an independent routing transmission method for heterogeneous audio and video, the method comprising:
[0007] Receive HDMI input signals from different source devices and decode the HDMI input signals into video data, audio data, and control / auxiliary information; The video data and control / auxiliary information corresponding to each of the source devices are encapsulated into a first signal by the first encapsulation module and transmitted to the matrix switching module through each first channel group; The audio data corresponding to each of the source devices is encapsulated into a second signal by the second encapsulation module and transmitted to the matrix switching module through each second channel. The first channel group and the second channel are independent physical transmission links. Based on the video source device and audio source device corresponding to the target output terminal, determine the target first channel group and the target second channel corresponding to the target output terminal; The target first signal in the target first channel group and the target second signal in the target second channel are transmitted to the output module through the matrix switching module, so that the output module generates an audio and video reconstruction signal based on the target first signal and the target second signal and transmits it to the target output terminal. The first signal does not contain an audio sampling payload for the output module to restore it to an HDMI output audio source.
[0008] In one embodiment, the output module includes an output restoration submodule, an audio restoration submodule, and an audio embedding submodule; After the step of transmitting the target first signal in the target first channel group and the target second signal in the target second channel to the output module through the matrix switching module, the method further includes: The target first signal is restored to an HDMI video signal frame through the output restoration submodule; The target second signal is restored to an audio interface signal through the audio recovery submodule; The audio interface signal is embedded into the HDMI video signal frame through the audio embedding submodule to obtain the audio-video reconstructed signal.
[0009] In one embodiment, the step of embedding the audio interface signal into the HDMI video signal frame through the audio embedding submodule to obtain the audio-video reconstructed signal includes: The audio interface signal is embedded into the blanking period of the HDMI video signal frame to obtain a new HDMI video signal frame. The audio clock regeneration parameters corresponding to the audio interface signal are determined based on the video pixel clock frequency in the HDMI video signal frame. The audio clock regeneration parameters are encapsulated into a data packet and embedded into the control / auxiliary data area of the HDMI video signal frame to obtain the audio-video reconstructed signal.
[0010] In one embodiment, before the step of encapsulating the audio clock regeneration parameters into a data packet and embedding it into the control / auxiliary data area of the HDMI video signal frame to obtain the audio-video reconstructed signal, the method further includes: Audio stream parameters are obtained from the audio interface signal, and the audio stream parameters include one or more of the following: audio sampling rate, number of channels, sampling bit width, and encoding format identifier; The audio stream parameters are encapsulated into an audio information packet and embedded into the control / auxiliary data area of the HDMI video signal frame.
[0011] In one embodiment, before the step of encapsulating the audio stream parameters into an audio information packet and embedding it into the control / auxiliary data area of the HDMI video signal frame, the method further includes: The available transmission bandwidth during the blanking period is determined based on the video pixel clock frequency and the number of channels in the target first channel group; The audio bitrate corresponding to the audio information packet is determined based on the audio stream parameters; The encapsulation size of the audio information packet is determined based on the audio bitrate and the available transmission bandwidth.
[0012] In one embodiment, the step of determining the target first channel group and the target second channel corresponding to the target output terminal based on the video source device and audio source device corresponding to the target output terminal includes: Load the video routing table, which is used to define the mapping relationship between the input and output terminals of each first channel group. The first channel group includes more than 3 differential data channels. Load the audio routing table, which is used to define the mapping relationship between the input and output terminals corresponding to each second channel; Based on the video routing table and the audio routing table, the target first channel group and the target second channel corresponding to the video source device and the audio source device are determined.
[0013] In one embodiment, before the step of determining the target first channel group and the target second channel corresponding to the target output terminal based on the video source device and audio source device corresponding to the target output terminal, the method further includes: Read the EDID data of the receiving device connected to the target output terminal to determine the video format parameters and audio format parameters supported by the receiving device; Based on the video format parameters and the audio format parameters, the video data and audio data corresponding to each of the source devices are traversed, and the video source device and the audio source device that are compatible with the target output are determined among the source devices.
[0014] In addition, to achieve the above objectives, this application also proposes an independent routing transmission device for heterogeneous audio and video, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the independent routing transmission method for heterogeneous audio and video as described above.
[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the independent routing transmission method for heterogeneous audio and video as described above.
[0016] This application provides a method for independent routing transmission of heterogeneous audio and video. It receives HDMI input signals from different source devices and decodes them into video data, audio data, and control / auxiliary information. A first encapsulation module encapsulates the video data and control / auxiliary information corresponding to each source device into a first signal, which is then transmitted to a matrix switching module via each first channel group. A second encapsulation module encapsulates the audio data corresponding to each source device into a second signal, which is then transmitted to the matrix switching module via each second channel. Based on the video and audio source devices corresponding to the target output terminal, the target first channel group and target second channel corresponding to the target output terminal are determined. The matrix switching module transmits the target first signal in the target first channel group and the target second signal in the target second channel to the output module, so that the output module generates an audio-video reconstructed signal based on the target first signal and the target second signal, and transmits it to the target output terminal.
[0017] This method directly decodes the HDMI signal at the matrix input side, splitting it into two independent data categories: "video data plus control / auxiliary information" and "audio data." These are then processed by a first encapsulation module and a second encapsulation module to form a first signal and a second signal, respectively. These signals are then sent to the matrix switching module via two independent physical links: a first channel group and a second channel. The matrix switching module directly switches and allocates the two independent physical links independently, thereby achieving the effect of independent source selection for audio and video output within the matrix, where "video comes from one source and audio comes from another." Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating an embodiment of the independent routing transmission method for heterogeneous audio and video in this application; Figure 2 A flowchart illustrating the second embodiment of the independent routing transmission method for heterogeneous audio and video in this application; Figure 3 A flowchart illustrating the third embodiment of the independent routing transmission method for heterogeneous audio and video in this application; Figure 4 A simplified flowchart is provided for Embodiment 3 of this application; Figure 5 This is a schematic diagram of channel grouping and independent switching provided in Embodiment 3 of this application; Figure 6 This is a flowchart illustrating Embodiment 4 of the independent routing transmission method for heterogeneous audio and video in this application. Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the independent routing transmission method for heterogeneous audio and video in the embodiments of this application.
[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not intended to limit this application.
[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific embodiments. It should be noted that all actions involving the acquisition of signals, information, or data in this application are performed in accordance with the relevant data protection laws and regulations of the country where the application is located, and with authorization from the owner of the corresponding device.
[0024] Traditional HDMI matrix transmission schemes follow the official native encapsulation protocol of HDMI, which encapsulates video data, audio data and control / auxiliary information together to form a high-speed serial data stream. This high-speed serial data stream is routed into the matrix through the same set of differential channels.
[0025] To obtain video and audio separately from different source devices, an external field-programmable gate array or HDMI dedicated audio / video de-embedding chip needs to be connected after each output of the matrix to perform de-embedding processing, so as to separate the mixed-package audio and video data. It is not possible to independently select video and audio sources and keep them synchronously output within the matrix.
[0026] In view of the above problems, this application proposes an independent routing transmission method for heterogeneous audio and video. The method involves receiving HDMI input signals from different source devices and decoding them into video data, audio data, and control / auxiliary information. A first encapsulation module encapsulates the video data and control / auxiliary information corresponding to each source device into a first signal, which is then transmitted to a matrix switching module via each first channel group. A second encapsulation module encapsulates the audio data corresponding to each source device into a second signal, which is then transmitted to the matrix switching module via each second channel. Based on the video and audio source devices corresponding to the target output terminal, the target first channel group and target second channel corresponding to the target output terminal are determined. The matrix switching module transmits the target first signal in the target first channel group and the target second signal in the target second channel to the output module, so that the output module generates an audio-video reconstructed signal based on the target first signal and the target second signal, and transmits it to the target output terminal.
[0027] This method directly decodes the HDMI signal at the matrix input side, splitting it into two independent data categories: "video data plus control / auxiliary information" and "audio data." These are then processed by a first encapsulation module and a second encapsulation module to form a first signal and a second signal, respectively. These signals are then sent to the matrix switching module via two independent physical links: a first channel group and a second channel. The matrix switching module directly switches and allocates the two independent physical links independently, thereby achieving the effect of independent source selection for audio and video output within the matrix, where "video comes from one source and audio comes from another."
[0028] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of implementing the above functions, or an independent routing transmission system for heterogeneous audio and video, etc. The following description uses an independent routing transmission system for heterogeneous audio and video as an example to illustrate this embodiment and the subsequent embodiments.
[0029] Based on this, the first embodiment of this application provides an independent routing transmission method for heterogeneous audio and video, referring to... Figure 1 In this embodiment, the independent routing transmission method for heterogeneous audio and video includes steps S10 to S50: Step S10: Receive HDMI input signals from different source devices and decode the HDMI input signals into video data, audio data, and control / auxiliary information.
[0030] It's important to note that the source device is the device that outputs HDMI audio and video signals; it's the transmitter of the HDMI input signal and can provide a complete HDMI signal including video, audio, and associated control information. The HDMI input signal conforms to the HDMI protocol standard and is a high-speed multimedia signal that integrates video data, audio data, and control / auxiliary information. It can use TMDS (Transition-Minimized Differential Signaling) or FRL (Fixed Rate Link) transmission methods. Video data is the raw data in the HDMI input signal used to represent image content and is the core data for display output. Audio data is the raw data in the HDMI input signal used to represent sound content and is the core data for audio playback. Control / auxiliary information is the supporting control and auxiliary data required to complete the normal transmission and display output of HDMI signals. It includes at least video timing parameters, color format parameters, InfoFrame (InformationFrame), HDCP (High-bandwidth Digital Content Protection) related control information, EDID (Extended Display Identification Data), and capability negotiation related information, and may also include configuration information for link negotiation and output configuration.
[0031] Step S10 described above can be executed by the system's input processing module. The input processing module receives HDMI input signals from different source devices via the HDMI physical interface, completing the reception and initial shaping of the physical layer signal to obtain an HDMI baseband signal suitable for protocol decoding. The input processing module performs protocol parsing on the shaped HDMI baseband signal, separating the fused video and audio data from the HDMI baseband signal according to the TMDS or FRL transmission specifications in the HDMI protocol, resulting in independent video and audio data streams. The input processing module continues to parse the HDMI baseband signal, extracting all control / auxiliary information carried in the signal according to the encapsulation format of control / auxiliary information in the HDMI protocol, forming an independent control / auxiliary information set. After decoding, the input processing module temporarily stores the separated video and audio data, as well as the extracted control / auxiliary information, providing a data foundation for subsequent signal encapsulation steps.
[0032] Step S20: The video data and control / auxiliary information corresponding to each of the source devices are encapsulated into a first signal by the first encapsulation module and transmitted to the matrix switching module through each first channel group.
[0033] It should be noted that the first encapsulation module is a functional module in the system used to encapsulate video data and control / auxiliary information to generate a first signal that conforms to the internal transmission specifications. The first signal is an internal high-speed serial signal carrying video data and control / auxiliary information, transmitted through the first channel group, and does not contain audio sampling payloads that can be used to recover audio interface signals at the output end and serve as the audio source for HDMI output. The first channel group is a differential channel group for transmitting the first signal, which is a group of bound differential data channels and is the physical transmission carrier of the first signal.
[0034] It should also be noted that, in this embodiment, the matrix switching module is a cross-switch matrix-like functional module that independently switches and allocates the first signals of multiple first channel groups and the second signals of multiple second channels, and can realize independent source selection and routing of audio and video signals.
[0035] For example, the first encapsulation module retrieves the decoded and temporarily stored video data and control / auxiliary information from the input processing module, completes data reading and buffering, and performs format encapsulation and data encoding on the retrieved video data and control / auxiliary information according to the system's internal high-speed serial signal transmission protocol to generate a first signal conforming to the internal transmission specifications. During the encapsulation process, it ensures that the first signal does not contain audio sampling payloads that can be used to recover audio interface signals at the output end and serve as HDMI output audio sources. Then, the first encapsulation module allocates the generated first signals corresponding to each source device to pre-matched first channel groups, controls the first signals to perform high-speed serial transmission through the corresponding first channel groups, and accurately sends each first signal to the matrix switching module to complete the output and transmission of the first signal.
[0036] Step S30: The audio data corresponding to each of the source devices is encapsulated into a second signal by the second encapsulation module and transmitted to the matrix switching module through each second channel.
[0037] It should be noted that the second encapsulation module is a functional module in the system used to encapsulate audio data and generate a second signal that conforms to the internal transmission specifications. This second signal is a dedicated high-speed serial signal that carries audio data and is transmitted via the second channel; it is the core signal for audio data transmission within the system. The second channel is a differential channel for transmitting the second signal; it is an independent differential data channel and serves as the physical transmission carrier for the second signal.
[0038] For example, the second encapsulation module retrieves the temporarily stored audio data after decoding from each source device from the input processing module, completes data reading and buffering, and performs format encapsulation and data encoding on the retrieved audio data according to the system's internal high-speed serial signal transmission protocol to generate a second signal conforming to the internal transmission specifications. Then, the second encapsulation module allocates the generated second signals corresponding to each source device to the pre-matched second channels, controls the second signals to perform high-speed serial transmission through the corresponding second channels, and accurately sends each second signal to the matrix switching module to complete the output and transmission of the second signals.
[0039] Step S40: Determine the target first channel group and target second channel corresponding to the target output terminal based on the video source device and audio source device corresponding to the target output terminal.
[0040] It should be noted that the target output terminal is the terminal port used to output the audio and video reconstructed signal. The video source device is the source device that provides video data to the target output terminal, i.e., the original signal output source corresponding to the target first signal. The audio source device is the source device that provides audio data to the target output terminal, i.e., the original signal output source corresponding to the target second signal. The target first channel group is the first channel group in the matrix switching module that is matched with the video source device and can transmit the corresponding first signal to the target output terminal. The target second channel is the second channel in the matrix switching module that is matched with the audio source device and can transmit the corresponding second signal to the target output terminal.
[0041] For example, the system's control module obtains the configuration command from the target output terminal and parses the identification information of the video source device and audio source device corresponding to the target output terminal from the configuration command. Then, the control module retrieves its pre-stored channel mapping table, which contains the binding relationships between all source devices and their corresponding first channel groups and second channels, as well as the transmittable association information between each channel and each output terminal. Based on the parsed video source device identifier, the control module retrieves the first channel group corresponding to the video source device from the channel mapping table and identifies it as the target first channel group corresponding to the target output terminal; and based on the parsed audio source device identifier, it retrieves the second channel corresponding to the audio source device from the channel mapping table and identifies it as the target second channel corresponding to the target output terminal. Then, the determined correspondence between the target first channel group, target second channel, and target output terminal is temporarily stored to provide a configuration basis for subsequent channel switching and signal transmission by the matrix switching module.
[0042] Understandably, in this embodiment, video and audio data undergo physical layer link splitting within the matrix switching module to achieve independent routing. That is, the traditional integrated audio and video internal transmission link is split into two completely independent high-speed serial links: the first channel group and the second channel. The matrix switching module performs switching and allocation operations on the two split links separately. The routing control of the two links does not interfere with each other, breaking the limitation of audio and video bound routing in traditional solutions. This allows video and audio sources at the same output port to be arbitrarily selected and combined from different input ports. Simultaneously, a hard configuration is implemented at the system output end, controlling the target output end not to extract and recover audio data from the video link (i.e., the first signal) as the audio source for HDMI output, but instead to use the audio data recovered from the audio link (i.e., the second signal) as the audio source, achieving audio and video decoupling and independent source selection.
[0043] Video data and control / auxiliary information are characterized by large data volumes and high transmission rate requirements. Differential channel group configurations can provide sufficient bandwidth to meet the transmission needs of video data. Furthermore, binding multiple differential channels together for switching ensures that video data and control / auxiliary information do not separate during transmission and switching, guaranteeing the integrity of the video link and the stability of the display output. Audio data, on the other hand, has a much smaller data volume than video data, and its transmission requirements can be met using separate differential channels. This avoids wasting channel resources and allows the matrix switching module to perform independent switching of audio links on a single channel, without binding it to video channels, thus achieving independence of audio link routing at the physical layer. The independent switching logic of a single channel significantly reduces the control complexity of the matrix module. There is no need for intra-group synchronous control of multi-channel audio links; arbitrary selection of audio sources can be achieved simply through routing mapping, avoiding the complex control logic brought about by FPGA routing in traditional solutions.
[0044] Optionally, step S40 above includes steps S41 to S43: Step S41: Load the video routing table, which is used to define the mapping relationship between the input and output terminals of each first channel group. The first channel group includes three or more channels.
[0045] The video routing table is a predefined configuration table in the system used to record the corresponding mapping relationship between each first channel group and the input end, that is, between each source device and the output end. It is the core basis for the matrix switching module to control the video link routing.
[0046] For example, based on the audio and video source selection requirements of the target output terminal, the control module triggers a routing table loading instruction for the video link, initiating a video routing table read request to the storage module. The control module retrieves the pre-configured video routing table from the system's storage module, completing the loading and caching of the video routing table within the control module. This ensures that the video routing table data can be retrieved in real time, and the loaded video routing table is validated to confirm that the mapping relationship between each first channel group and the input and output terminals is completely defined. The number of channels in each first channel group must be at least three.
[0047] Step S42: Load the audio routing table, which is used to define the mapping relationship between the input and output terminals corresponding to each second channel.
[0048] The audio routing table is a predefined configuration table in the system used to record the corresponding mapping relationship between each second channel and the input and output terminals. It is the core basis for the matrix switching module to control the audio link routing.
[0049] For example, while loading the video routing table, the control module simultaneously triggers an audio routing table loading command for the audio link, initiating a read request for the audio routing table to the storage module. The control module retrieves the pre-configured audio routing table from the system's storage module, completes the loading and caching of the audio routing table within the control module, and performs validity verification on the loaded audio routing table to confirm that the mapping relationships between each second channel and the input and output terminals are completely defined, ensuring the accuracy of the routing table data.
[0050] Step S43: Based on the video routing table and the audio routing table, determine the target first channel group and the target second channel corresponding to the video source device and the audio source device.
[0051] For example, the control module extracts the video source device identifier and audio source device identifier corresponding to the parsed target output terminal, as well as the identifier information of the target output terminal itself, as an index for routing table retrieval. The control module uses the video source device identifier and target output terminal identifier as search conditions to search the loaded video routing table, matching the first channel group corresponding to the search conditions and identifying it as the target first channel group. Simultaneously, the control module uses the audio source device identifier and target output terminal identifier as search conditions to search the loaded audio routing table, matching the second channel corresponding to the search conditions and identifying it as the target second channel. Then, the retrieved correspondence between the target first channel group, target second channel, and target output terminal is associated and stored. At the same time, a channel configuration command is sent to the matrix switching module to synchronize the matching results, providing a routing execution basis for subsequent signal switching and transmission.
[0052] Preferably, in this embodiment, the number of channels in the first channel group is 3, and the second channel is a single channel. Optionally, in other embodiments, the number of channels in the first channel group may be greater than 3, or more than one channel may be selected and bound together as the second channel group for transmitting audio data.
[0053] Step S50: The target first signal in the target first channel group and the target second signal in the target second channel are transmitted to the output module through the matrix switching module, so that the output module generates an audio and video reconstruction signal based on the target first signal and the target second signal and transmits it to the target output terminal.
[0054] The first target signal is the first signal carrying the video data and control / auxiliary information of the corresponding video source device in the first target channel group. The second target signal is the second signal carrying the audio data of the corresponding audio source device in the second target channel. The output module is a functional module in the system that receives the first target signal and the second target signal transmitted by the matrix switching module, performs signal restoration, audio and video reconstruction, and generates an HDMI audio and video reconstruction signal, integrating audio restoration, video restoration, and audio embedding sub-functions. The audio and video reconstruction signal is an HDMI signal that conforms to the HDMI protocol specification and can be output externally, generated by the output module after re-merging the target video data and target audio data from different sources and updating the corresponding audio-related information.
[0055] For example, the matrix switching module receives the channel configuration command issued by the control module and parses out the identification information of the target first channel group and the target second channel corresponding to the target output terminal. The matrix switching module configures its own routing switching logic according to the channel configuration command, establishes the path for the target first signal in the target first channel group, and transmits the target first signal to the corresponding output module for video restoration according to the configured routing rules. While transmitting the target first signal, the matrix switching module independently establishes the signal path for the target second channel, transmitting the target second signal in the target second channel to the same output module for audio restoration and audio embedding, thus achieving synchronous routing of heterogeneous audio and video signals to the output module. The output module receives the target first signal and the target second signal transmitted by the matrix switching module. First, it performs protocol decapsulation and physical layer restoration on the target first signal, parsing out the video data and control / auxiliary information, generating video link physical layer data conforming to the HDMI protocol, such as TMDS or FRL, to obtain an HDMI video signal frame. An HDMI video signal frame is a video signal frame structure conforming to the HDMI protocol specification, containing complete video data and control / auxiliary information. It is the basic unit of HDMI video transmission and can be directly used as the basic signal for video display. Simultaneously, the target second signal undergoes protocol decapsulation and audio recovery, restoring it to standard audio interface signals such as I2S (Inter-IC Sound), TDM (Time-Division Multiplexing), or SPDIF (Sony / Philips Digital Interface Format), thus completing independent parsing of the audio data. Based on the restored video link physical layer data, the output module embeds the recovered audio interface signal into the blanking period of the HDMI video signal frame, completing the physical layer reconstruction of the audio and video data. Finally, the output module performs final encoding of the reconstructed audio and video signal according to the HDMI protocol specification, generating a standard reconstructed audio and video signal, and transmits this reconstructed signal to the corresponding target output terminal, completing the entire independent routing transmission process for heterogeneous audio and video.
[0056] This embodiment directly decodes the HDMI signal at the matrix input side, splitting it into two independent data categories: "video data plus control / auxiliary information" and "audio data." These are then processed by a first encapsulation module and a second encapsulation module to form a first signal and a second signal, respectively. These signals are then sent to the matrix switching module via two independent physical links: a first channel group and a second channel. This achieves audio-video physical decoupling at the input layer. The matrix switching module directly switches and allocates the two links independently, completing the heterogeneous source combination of "video from one source and audio from another" within the matrix. Furthermore, the solution provided in this embodiment does not rely on external de-embedding devices. The output module can directly restore, embed, and output standard HDMI signals from both signals, eliminating the need for external de-embedding chips and FPGAs at each output end. This reduces peripheral components, wiring, and logic design, lowering transmission costs from both hardware and structural perspectives.
[0057] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. On this basis, the output module includes an output restoration submodule, an audio restoration submodule, and an audio embedding submodule.
[0058] The output restoration submodule is used to decapsulate and restore the physical layer of the target first signal, generating a standard HDMI video signal frame. The audio restoration submodule is used to decapsulate and parse the audio data of the target second signal, restoring it to a standard audio interface signal. The audio embedding submodule is used to fuse the standard audio interface signal into the HDMI video signal frame, completing the audio-video reconstruction and generating an audio-video reconstructed signal.
[0059] Please refer to Figure 2 After step S50, the independent routing transmission method for heterogeneous audio and video further includes steps S60 to S80: Step S60: The target first signal is restored to an HDMI video signal frame through the output restoration submodule.
[0060] For example, the output restoration submodule receives the target first signal transmitted by the matrix switching module, performs data buffering and format verification on the target first signal to confirm that no data is lost or corrupted during signal transmission. Then, according to the system's internal high-speed serial signal encapsulation protocol, the target first signal is decapsulated, stripping away the internal transmission encapsulation format to extract the original video data and control / auxiliary information carried within. The output restoration submodule then performs physical layer restoration and frame structure reconstruction of the extracted video data and control / auxiliary information according to the HDMI protocol specification, and encodes it using TMDS or FRL HDMI physical layer transmission methods to generate HDMI video signal frames that conform to the HDMI protocol standard, providing a basic video carrier for subsequent audio embedding.
[0061] Step S70: The target second signal is restored to an audio interface signal through the audio recovery submodule.
[0062] For example, the audio restoration submodule and the output restoration submodule work in parallel. The submodule receives the target second signal transmitted by the matrix switching module, performs data buffering and format verification on the target second signal to confirm the integrity of the audio data. Then, according to the system's internal high-speed serial signal encapsulation protocol, the target second signal is decapsulated, stripping away the internal transmission encapsulation format to extract the original audio data it carries. The audio restoration submodule then parses and converts the extracted original audio data according to a standard audio interface protocol, restoring it to any standard audio interface signal such as I2S, TDM, or SPDIF, ensuring that the audio interface signal can be recognized and processed by the audio embedding submodule.
[0063] Step S80: The audio interface signal is embedded into the HDMI video signal frame through the audio embedding submodule to obtain the audio-video reconstructed signal.
[0064] Optionally, step S80 above includes steps S81 to S83: Step S81: Embed the audio interface signal into the blanking period of the HDMI video signal frame to obtain a new HDMI video signal frame.
[0065] It's important to note that the blanking period is the non-display area within an HDMI video signal frame that contains no image data. It includes the horizontal blanking period and the vertical blanking period, and is a dedicated area for embedding audio data and auxiliary / control signals. The blanking period is further divided into the control period and the data island period. The control period transmits the most basic synchronization control signals, such as horizontal and vertical sync; the data island period is used to transmit audio data and auxiliary information such as audio packets.
[0066] HDMI video signal frames have a three-level nested structure of pixels-scan lines-fields / frames, with all timing parameters based on the pixel clock cycle as the smallest unit. A pixel is the smallest data unit in an HDMI video signal frame, composed of RGB or YCbCr color components, and is the core data of the effective display area. A scan line consists of several consecutive pixels containing effective pixels and line blanking period pixels; it is the basic unit of progressive scanning in HDMI, and the timing structure of each line is consistent. A field / frame consists of several consecutive scan lines containing effective display lines and field blanking period lines. One frame is a complete image, and the frame rate represents the number of frames transmitted per second. For example, assuming a 640x480@60Hz HDMI video signal frame, its total pixel clock cycle is 800 (lines) × 525 (frames), where the effective display area is 640 pixels × 480 lines, and the remainder is the blanking period area.
[0067] Optionally, the aforementioned audio interface signal can be embedded into the data island period of the blanking period in the HDMI video signal frame to obtain a new HDMI video signal frame.
[0068] Step S82: Determine the audio clock regeneration parameters corresponding to the audio interface signal based on the video pixel clock frequency in the HDMI video signal frame.
[0069] The video pixel clock frequency is the clock frequency that drives the transmission and refresh of pixel data in the video data of an HDMI video signal frame. It is the core clock reference for video signal transmission and determines the video resolution and refresh rate. Audio clock reproduction parameters are parameters generated to ensure accurate audio clock reproduction at the receiving end, guaranteeing audio-video synchronization. These parameters include the N (Numerator) value and the CTS (Clock Time Stamp) value.
[0070] For example, the audio embedding submodule extracts the video pixel clock frequency from the control / auxiliary information of the new HDMI video signal frame, and determines the value of N based on the ratio of the video pixel clock frequency to the reference audio defined by the HDMI protocol, for example: N=Fp / Fr×n, where Fp is the video pixel clock frequency, Fr is the reference audio, n is the scaling factor, and the value of N needs to be rounded.
[0071] Next, determine the audio sampling rate and the number of audio samples. First, calculate the product of the video pixel clock frequency and the number of audio samples. Then, determine the CTS value based on the ratio of this product to the audio sampling rate. For example: CTS=(n×Fp×k) / Fa, where Fa is the audio sampling rate and k is the number of audio samples. The CTS value needs to be rounded down.
[0072] After obtaining the N and CTS values mentioned above, the receiving end can reconstruct the audio clock using the following formula to verify the synchronization effect: F'a = Fr × CTS / N. F'a is the audio sampling rate reconstructed by the receiving end. The receiving end does not need to transmit the audio clock separately; it can accurately reconstruct the audio clock synchronized with the video using only the video pixel clock frequency and the embedded N / CTS value, thus solving the problem of clock asynchrony between heterogeneous audio and video sources.
[0073] Step S83: Encapsulate the audio clock regeneration parameters into a data packet and embed it into the control / auxiliary data area of the HDMI video signal frame to obtain the audio-video reconstructed signal.
[0074] For example, the audio embedding submodule calculates the audio clock playback parameters and encapsulates them according to the HDMI protocol's audio clock playback data packet encapsulation specifications to generate a data packet that conforms to the protocol standard, ensuring that the receiving end can parse it correctly. Then, it locates the control / auxiliary data area in the new HDMI video signal frame, determines the data packet encapsulation embedding position as described above, and completely embeds the data packet into the control / auxiliary data area of the HDMI video signal frame, completing the fusion of audio synchronization parameters and video frame to obtain the audio-video reconstructed signal.
[0075] In this embodiment, the decapsulation and restoration of the target first signal and the target second signal are completed within the matrix through the output restoration submodule and the audio recovery submodule, eliminating the need for additional external components, reducing external hardware investment and wiring complexity, and lowering system costs. Furthermore, this embodiment uses the video pixel clock frequency as a reference to calculate audio clock regeneration parameters such as the N value and CTS value and embeds them into the control / auxiliary data area. The receiving end can accurately restore the audio clock synchronized with the video using these parameters, solving the problem of heterogeneous audio and video clock synchronization at the protocol level, ensuring audio-visual synchronization and smooth playback.
[0076] Based on the above embodiments of this application, in the third embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 Before step S80, the independent routing transmission method for heterogeneous audio and video also includes steps S90-S100: Step S90: Obtain audio stream parameters from the audio interface signal. The audio stream parameters include one or more of the following: audio sampling rate, number of channels, sampling bit width, and encoding format identifier.
[0077] Audio stream parameters are a set of parameters that characterize the core attributes of audio interface signals. They are the basis for the receiver to identify and parse audio signals and may include key information such as audio sampling rate, number of channels, sampling bit width, encoding format identifier, audio data block length, and the second channel identifier corresponding to the audio data.
[0078] For example, the audio embedding submodule receives the standard audio interface signal restored by the audio restoration submodule, and extracts preset audio stream parameters from the audio interface signal according to the signal structure of the corresponding audio interface protocol. It can obtain one or more of the following parameters individually or simultaneously: audio sampling rate, number of channels, sampling bit width, and encoding format identifier. Optionally, the audio embedding submodule can verify the validity of the extracted audio stream parameters, confirming that the parameters conform to the audio specifications supported by the HDMI protocol, such as verifying that the audio sampling rate is within the range of 44.1kHz to 192kHz, and cache the verified audio stream parameters in preparation for subsequent encapsulation.
[0079] Step S100: Encapsulate the audio stream parameters into an audio information packet and embed it into the control / auxiliary data area in the HDMI video signal frame.
[0080] The audio embedding submodule retrieves the cached audio stream parameters and, according to the HDMI protocol's audio packet encapsulation specifications, formats and encapsulates the parameters to generate an audio packet conforming to the HDMI protocol standard. After receiving the HDMI video signal frame generated by the output restoration submodule, the audio embedding submodule locates the control / auxiliary data area within the HDMI video signal frame used to carry auxiliary information, determining the embedding position of the audio packet as described above. It should be noted that the embedding position of this audio packet is independent of the embedding position of the audio clock regeneration parameter data packet to avoid data conflicts.
[0081] Optionally, prior to step S100, the independent routing transmission method for heterogeneous audio and video further includes steps S110 to S130: Step S110: Determine the available transmission bandwidth during the blanking period based on the video pixel clock frequency and the number of channels in the target first channel group.
[0082] It should be noted that the available transmission bandwidth is the effective bandwidth available for transmitting auxiliary data such as audio information packets and audio clock regeneration data packets during the blanking period of the HDMI video signal frame, after deducting the bandwidth occupied by the transmission of basic control signals.
[0083] For example, the audio embedding submodule extracts the video pixel clock frequency from the control / auxiliary information of the HDMI video signal frame, and retrieves the number of channels of the target first channel group from the control module. First, it calculates the total transmission bandwidth of the target first channel group based on the video pixel clock frequency: total transmission bandwidth = video pixel clock frequency × single channel bit width × number of channels of the first channel group. The single channel bit width is determined according to the HDMI physical layer specification, such as TMDS, which is usually 10 bits, and FRL, which is the corresponding encoding bit width.
[0084] Next, the blanking period timing parameters of the HDMI video signal frame are analyzed, and the blanking period duty cycle is calculated.
[0085] Optionally, in the HDMI frame format, audio-related data packets are mainly embedded within the blanking period of a single line. Therefore, the blanking period duty cycle can be obtained by calculating the single-line blanking period duty cycle: Single-line blanking period duty cycle = Total number of pixel clock cycles in a single-line blanking period / Total number of pixel clock cycles in a single scan line. Here, the total number of pixel clock cycles in a single-line blanking period is a fixed value specified by the HDMI protocol according to resolution or refresh rate, including the sum of cycles in all sub-regions of the line blanking period, such as the line synchronization period and the data island preparation area. The total number of pixel clock cycles in a single scan line = Number of effective display periods in a single line + Number of blanking period cycles in a single line, which is the complete timing length of a single scan line, calibrated by the HDMI protocol for each resolution or refresh rate. For example, the total number of pixel clock cycles in a single scan line for 1080p@60Hz is 2200, and the single-line blanking period is 448.
[0086] Optionally, the blanking period duty cycle can also be obtained by calculating the blanking period duty cycle of a single frame, adapted to the blanking period bandwidth of the entire frame, for scenarios involving batch transmission of audio data packets. Specifically, the single-frame blanking period duty cycle = the total number of scan lines in the single-frame field blanking period / the total number of scan lines in a single frame.
[0087] Finally, the available transmission bandwidth for transmitting audio-related data packets during the blanking period is determined by multiplying the total transmission bandwidth by the blanking period duty cycle.
[0088] Step S120: Determine the audio bitrate corresponding to the audio information packet based on the audio stream parameters.
[0089] For example, the audio embedding submodule retrieves audio stream parameters and calculates the audio bitrate of the audio signal: Audio bitrate = Audio sampling rate × Sampling bit width × Number of channels in the first channel. If the audio data is in a compressed audio format, the actual audio bitrate is determined according to the bitrate standard of the corresponding compression algorithm.
[0090] Alternatively, the audio bitrate corresponding to the audio information packet can be determined by combining the encapsulation requirements of the HDMI protocol with the additional data such as the protocol header and check bits required for parameter encapsulation on the basis of the original audio bitrate.
[0091] Step S130: Determine the encapsulation size of the audio information packet based on the audio bitrate and the available transmission bandwidth.
[0092] It is understandable that the available transmission bandwidth is the maximum transmission capacity that can be used to transmit audio information packets during the blanking period, which directly determines the maximum encapsulation size of audio information packets that can be transmitted within a single timing period, i.e., a single line or a single frame blanking period; the audio bitrate is derived from the audio stream parameters and reflects the amount of data required to transmit audio information packets per unit time. Therefore, the encapsulation size of the aforementioned audio information packets is proportional to the audio bitrate and the available transmission bandwidth.
[0093] The larger the available transmission bandwidth, the more data can be transmitted within a single time period, and the larger the encapsulation size of the audio information packet. Conversely, the smaller the available transmission bandwidth, the less data can be transmitted within a single time period, and the smaller the encapsulation size of the audio information packet. A higher audio bitrate indicates a larger total amount of data in the audio stream parameters, and thus a larger actual encapsulation size of the audio information packet.
[0094] Optionally, if the minimum encapsulation size corresponding to the audio bitrate is less than or equal to the maximum encapsulation size corresponding to the available transmission bandwidth, then all audio stream parameters are directly encapsulated into a single complete data packet. In this case, the encapsulation size changes synchronously with the audio bitrate, and there is remaining available transmission bandwidth to reserve bandwidth space for subsequent audio clock regeneration data packets, etc. If the minimum encapsulation size corresponding to the audio bitrate is greater than the maximum encapsulation size corresponding to the available transmission bandwidth, the upper limit of the available transmission bandwidth is used as the maximum encapsulation size. The audio stream parameters are divided into blocks and encapsulated into multiple sub-data packets, and a block identifier is added to each sub-data packet. The single-packet encapsulation size of the sub-data packet is determined by the available transmission bandwidth. The total encapsulation size is still positively correlated with the audio bitrate, that is, the higher the audio bitrate, the more blocks there are, and the larger the total data volume. The receiving end reassembles all sub-data packets using the block identifier to restore the complete audio stream parameters.
[0095] For example, to help understand the implementation process of the independent routing transmission method for heterogeneous audio and video obtained by combining this embodiment with the above embodiments, please refer to... Figure 4 , Figure 4 A simplified flowchart illustrating an independent routing transmission method for heterogeneous audio and video is provided, specifically: The independent routing transmission system for heterogeneous audio and video described in this application includes a control module, an input port, an output port, an input processing module, a first encapsulation module, a second encapsulation module, a matrix switching module, and an output module. The output module includes an output restoration submodule, an audio restoration submodule, and an audio embedding submodule.
[0096] The input port receives multimedia interface input signals from multiple different source devices, such as high-definition multimedia interfaces (HDMI), providing multiple audio and video input sources for the system. The input processing module decodes the multiple input signals, separating video data, audio data, and control / auxiliary information to provide basic data for subsequent packaging.
[0097] The first encapsulation module encapsulates the video data and control / auxiliary information corresponding to each source device into a first signal, and transmits it to the matrix switching module through the first channel group to realize an independent transmission link for video and control information.
[0098] The second encapsulation module encapsulates the audio data corresponding to each source device into a second signal and transmits it to the matrix switching module through the second channel, thereby realizing an independent transmission link for the audio data.
[0099] Under the control of the control module, the matrix switching module independently switches and allocates the first signals of multiple first channel groups and the second signals of multiple second channels. It can select video and audio signals from different source devices for combination routing according to the needs of the target output end.
[0100] The output restoration submodule receives the target first signal transmitted by the matrix switching module, restores it to a high-definition multimedia interface video signal frame, and generates a video carrier that conforms to the high-definition multimedia interface protocol.
[0101] The audio recovery submodule receives the target second signal transmitted by the matrix switching module and restores it to a standard audio interface signal, such as I2S / TDM / SPDIF.
[0102] The audio embedding submodule embeds the audio interface signal output by the audio restoration submodule into the high-definition multimedia interface video signal frame generated by the output restoration submodule, thus generating an audio-video reconstructed signal.
[0103] The output port outputs the audio and video reconstructed signal to the target output terminal, completing the combined output of heterogeneous audio and video.
[0104] As the core of the system, the control module is responsible for reading the EDID data of the receiving device, determining the video and audio source devices, loading the routing table and configuring the channel switching rules of the matrix switching module, as well as coordinating the working sequence of each module to achieve automated and intelligent control of the entire system.
[0105] To better understand the solution presented in this example, further explanation is provided in conjunction with specific application scenarios. (Refer to...) Figure 5 , Figure 5 A schematic diagram of channel grouping and independent switching is provided. Assume that the audio and video sources corresponding to the receiving device connected to output port 1 are the video data from input port 1 and the audio data from input port 2.
[0106] Specifically, the input side includes multiple input ports, each corresponding to two independent internal transmission signals. Among them, input ports 1, 2 and 3 transmit the video data and control / auxiliary information of the corresponding source device through the first channel group of 3 channels, and transmit the audio data of the corresponding source device through the second channel of a single channel.
[0107] The matrix switching module executes two independent switching and allocation logics: it switches / allocates the three channels of all input ports by group, which can route the first signal of any input port to the video link of the target output port; and it switches / allocates the single channel of all input ports individually, which can route the second signal of any input port to the audio link of the target output port, so as to realize the heterogeneous combination of "video from A and audio from B" of the same output port.
[0108] Taking output port 1 as an example, the video link on the output side receives the target first signal from the selected input port, which is routed by the receiving matrix switching module, and restores it to a High Definition Multimedia Interface (HDMI) video signal frame. The audio link receives the target second signal from the selected input port, which is routed by the receiving matrix switching module and restored to a standard audio interface signal by the audio restoration submodule. The audio embedding submodule embeds the audio interface signal into the HDMI video signal frame generated by the video link, obtaining a reconstructed audio-video signal, and outputs the final reconstructed audio-video signal to the receiving device.
[0109] Based on the above embodiments of this application, in the fourth embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 6 Before step S40, the independent routing transmission method for heterogeneous audio and video also includes steps S140-S150: Step S140: Read the EDID data of the receiving device connected to the target output terminal to determine the video format parameters and audio format parameters supported by the receiving device.
[0110] EDID data is a standard set of data stored by HDMI receiving devices to describe their hardware capabilities. It includes video format parameters supported by the receiving device, such as resolution, refresh rate, color format, and pixel clock range, as well as audio format parameters, such as audio sampling rate, number of channels, sampling bit width, audio encoding format, and audio interface type. A receiving device is an HDMI terminal device connected to the target output, such as a monitor, television, or amplifier; it is the device that receives and plays back the reconstructed audio and video signals.
[0111] For example, the control module sends an EDID data read command to the target output terminal, triggering the target output terminal to perform EDID data interaction with the connected receiving device; the target output terminal reads the complete EDID data from the EDID storage chip of the receiving device through the EDID read channel specified by the HDMI protocol, and sends the read EDID data back to the control module; the control module parses the EDID data according to the field definition of EDID data in the HDMI protocol, and extracts the video format parameters and audio format parameters that represent the capabilities of the receiving device.
[0112] Step S150: Based on the video format parameters and the audio format parameters, traverse the video data and audio data corresponding to each of the source devices, and determine the video source device and the audio source device that are compatible with the target output terminal among the source devices.
[0113] For example, the control module retrieves the attribute parameters of the video and audio data corresponding to all source devices in the system. These attribute parameters are described in the same specifications as the video and audio format parameters of the receiving device, including the resolution, refresh rate, and other core information of the video data of each source device, as well as the sampling rate and number of channels of the audio data. Using the cached video and audio format parameters of the receiving device as filtering conditions, the control module iterates through the video and audio data attribute parameters of all source devices, respectively filtering out video source devices whose video data matches the video capabilities of the receiving device, and audio source devices whose audio data matches the audio capabilities of the receiving device.
[0114] Optionally, step S150 above includes steps S151 to S154: Step S151: Traverse the video data and audio data corresponding to each of the source devices, and extract the video content features and audio content features from the video data and audio data respectively.
[0115] Video content features are extracted from the video data of the source device and represent the core content of the video, such as frame rate characteristics, dynamic frame ratio, and color space characteristics, used to distinguish video content types. Audio content features are extracted from the audio data of the source device and represent the core content of the audio, such as channel mode, audio type, and frequency range, used to distinguish audio content types.
[0116] For example, the control module retrieves video and audio data corresponding to all source devices, preprocesses the audio and video data such as video frame extraction or audio frame segmentation, and extracts video content features from the video data using a pre-trained video feature extraction model. These features include, but are not limited to, screen resolution distribution, dynamic pixel ratio, color gamut range, brightness distribution, frame rate, and screen switching frequency. Simultaneously, it extracts audio content features from the audio data using a pre-trained audio feature extraction model. These features include, but are not limited to, frequency distribution, harmonic components, sound pressure level, number of channels, and feature values corresponding to the encoding format.
[0117] Step S152: Using a pre-trained semantic classification model, determine the scene semantic labels corresponding to the video content features and the audio content features, respectively.
[0118] The control module inputs the cached video content features into the video classification branch of the pre-trained semantic classification model. The semantic classification model, through the feature mapping rules learned during training, performs scene matching on the video content features and outputs corresponding video scene semantic labels, such as "4K movie playback," "1080p game footage," and "conference recording." Simultaneously, the control module inputs the cached audio content features into the audio classification branch of the semantic classification model. The semantic classification model similarly performs scene matching on the audio content features and outputs corresponding audio scene semantic labels, such as "5.1 surround sound movie audio," "two-channel game sound effects," and "mono conference voice."
[0119] Step S153: Determine the corresponding adaptive receiving device type for the video data and the audio data based on the scene semantic tags.
[0120] For example, the control module retrieves the mapping table of scene semantic tags and adapted receiving device types pre-stored in the system. Using the video scene semantic tag as an index, it retrieves the corresponding video adapted receiving device type in the mapping table. Using the audio scene semantic tag as an index, it retrieves the corresponding audio adapted receiving device type in the mapping table. The video adapted receiving device type and the audio adapted receiving device type are then associated with and stored with the identifier of the corresponding source device.
[0121] Step S154: Based on the adapted receiving device type, the video format parameters, and the audio format parameters, determine the video source device and the audio source device corresponding to the target output terminal in the source devices.
[0122] The control module parses the receiving device EDID data read in step S140. Besides extracting video and audio format parameters, it also extracts the device type identifier, such as the device category field in the EDID data, which indicates the device type, such as a television, monitor, amplifier, or conference screen. The control module first filters out source devices whose "adapted receiving device type" matches the receiving device type connected to the target output, forming a content-adapted source device subset. Then, within this subset, it further filters out source devices whose video data conforms to the receiving device's video format parameters and whose audio data conforms to the receiving device's audio format parameters, forming a source device set with both format and content adaptation. Finally, based on the system's preset priority rules, it determines the video and audio source devices corresponding to the target output from this set of source devices.
[0123] In this embodiment, the format compatibility requirements are automatically obtained by reading the EDID data of the receiving device, enabling source device screening without manual intervention and reducing manual operation costs. Furthermore, this embodiment also adds content and scene adaptation on top of format adaptation by extracting audio and video content features, generating scene semantic tags, and matching device types. This ensures that the selected source devices not only meet the hardware capabilities of the receiving device but also match its usage scenario, further improving adaptation accuracy and practicality.
[0124] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the independent routing transmission method of heterogeneous audio and video in this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0125] This application provides an independent routing transmission device for heterogeneous audio and video, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the independent routing transmission method for heterogeneous audio and video in the above embodiment 1.
[0126] The following is for reference. Figure 7 This document illustrates a structural schematic diagram of an independent routing transmission device suitable for implementing heterogeneous audio and video in the embodiments of this application. The independent routing transmission device for heterogeneous audio and video in the embodiments of this application may include, but is not limited to, mobile terminals such as laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), and fixed terminals such as digital TVs and desktop computers. Figure 7 The illustrated independent routing transmission device for heterogeneous audio and video is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0127] like Figure 7As shown, the independent routing transmission device for heterogeneous audio and video can include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the independent routing transmission device for heterogeneous audio and video. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the independent routing transmission device for heterogeneous audio and video to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows independent routing transmission devices for heterogeneous audio and video with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented or possessed alternatively.
[0128] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0129] This application physically decouples the video and audio links within the matrix and switches them independently, allowing the video and audio sources at the target output to originate from different inputs. Simultaneously, the output does not recover audio from the first signal; instead, it uses the audio recovered from the second signal as the HDMI output audio source and re-embeds it into the video signal frame, achieving independent combination and synchronous output of heterogeneous audio and video. Cost reduction is a byproduct. Compared to existing technologies, the beneficial effects of the heterogeneous audio and video independent routing transmission device provided in this application are the same as those of the heterogeneous audio and video independent routing transmission method provided in the above embodiments. Furthermore, other technical features of this heterogeneous audio and video independent routing transmission device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0130] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0131] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0132] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the heterogeneous audio and video independent routing transmission method in the above embodiments.
[0133] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.
[0134] The aforementioned computer-readable storage medium may be included in a stand-alone routing transmission device for heterogeneous audio and video; or it may exist independently and not be assembled into a stand-alone routing transmission device for heterogeneous audio and video.
[0135] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an independent routing transmission device for heterogeneous audio and video, enable the independent routing transmission device to write computer program code for performing the operations of this application in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, or as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0137] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0138] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described independent routing transmission method for heterogeneous audio and video, thereby solving the technical problem of how to reduce the cost of independent routing transmission of heterogeneous audio and video. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the independent routing transmission method for heterogeneous audio and video provided in the above embodiments, and will not be repeated here.
[0139] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the independent routing transmission method for heterogeneous audio and video as described above.
[0140] The computer program product provided in this application solves the technical problem of how to reduce the cost of independent routing transmission of heterogeneous audio and video. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the independent routing transmission method for heterogeneous audio and video provided in the above embodiments, and will not be repeated here.
[0141] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for independent routing transmission of heterogeneous audio and video, characterized in that, The independent routing transmission method for heterogeneous audio and video includes: Receive HDMI input signals from different source devices and decode the HDMI input signals into video data, audio data, and control / auxiliary information; The video data and control / auxiliary information corresponding to each of the source devices are encapsulated into a first signal by the first encapsulation module and transmitted to the matrix switching module through each first channel group; The audio data corresponding to each of the source devices is encapsulated into a second signal by the second encapsulation module and transmitted to the matrix switching module through each second channel. The first channel group and the second channel are independent physical transmission links. Based on the video source device and audio source device corresponding to the target output terminal, determine the target first channel group and the target second channel corresponding to the target output terminal; The target first signal in the target first channel group and the target second signal in the target second channel are transmitted to the output module through the matrix switching module, so that the output module generates an audio and video reconstruction signal based on the target first signal and the target second signal and transmits it to the target output terminal. The first signal does not contain an audio sampling payload for the output module to restore it to an HDMI output audio source.
2. The independent routing transmission method for heterogeneous audio and video as described in claim 1, characterized in that, The output module includes an output restoration submodule, an audio restoration submodule, and an audio embedding submodule; After the step of transmitting the target first signal in the target first channel group and the target second signal in the target second channel to the output module through the matrix switching module, the method further includes: The target first signal is restored to an HDMI video signal frame through the output restoration submodule; The target second signal is restored to an audio interface signal through the audio recovery submodule; The audio interface signal is embedded into the HDMI video signal frame through the audio embedding submodule to obtain the audio-video reconstructed signal.
3. The independent routing transmission method for heterogeneous audio and video as described in claim 2, characterized in that, The step of embedding the audio interface signal into the HDMI video signal frame through the audio embedding submodule to obtain the audio-video reconstructed signal includes: The audio interface signal is embedded into the blanking period of the HDMI video signal frame to obtain a new HDMI video signal frame. The audio clock regeneration parameters corresponding to the audio interface signal are determined based on the video pixel clock frequency in the HDMI video signal frame. The audio clock regeneration parameters are encapsulated into a data packet and embedded into the control / auxiliary data area of the HDMI video signal frame to obtain the audio-video reconstructed signal.
4. The independent routing transmission method for heterogeneous audio and video as described in claim 3, characterized in that, Before the step of encapsulating the audio clock regeneration parameters into a data packet and embedding it into the control / auxiliary data area of the HDMI video signal frame to obtain the audio-video reconstructed signal, the method further includes: Audio stream parameters are obtained from the audio interface signal, and the audio stream parameters include one or more of the following: audio sampling rate, number of channels, sampling bit width, and encoding format identifier; The audio stream parameters are encapsulated into an audio information packet and embedded into the control / auxiliary data area of the HDMI video signal frame.
5. The independent routing transmission method for heterogeneous audio and video as described in claim 4, characterized in that, Before the step of encapsulating the audio stream parameters into an audio information packet and embedding it into the control / auxiliary data area of the HDMI video signal frame, the method further includes: The available transmission bandwidth during the blanking period is determined based on the video pixel clock frequency and the number of channels in the target first channel group. The audio bitrate corresponding to the audio information packet is determined based on the audio stream parameters; The encapsulation size of the audio information packet is determined based on the audio bitrate and the available transmission bandwidth.
6. The independent routing transmission method for heterogeneous audio and video as described in claim 1, characterized in that, The step of determining the target first channel group and the target second channel corresponding to the target output terminal based on the video source device and audio source device corresponding to the target output terminal includes: Load the video routing table, which is used to define the mapping relationship between the input and output terminals of each first channel group. The first channel group includes more than 3 differential data channels. Load the audio routing table, which is used to define the mapping relationship between the input and output terminals corresponding to each second channel; Based on the video routing table and the audio routing table, the target first channel group and the target second channel corresponding to the video source device and the audio source device are determined.
7. The independent routing transmission method for heterogeneous audio and video as described in claim 1, characterized in that, Before the step of determining the target first channel group and target second channel corresponding to the target output terminal based on the video source device and audio source device corresponding to the target output terminal, the method further includes: Read the EDID data of the receiving device connected to the target output terminal to determine the video format parameters and audio format parameters supported by the receiving device; Based on the video format parameters and the audio format parameters, the video data and audio data corresponding to each of the source devices are traversed, and the video source device and the audio source device that are compatible with the target output are determined among the source devices.
8. A heterogeneous audio and video independent routing transmission device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the independent routing transmission method for heterogeneous audio and video as claimed in any one of claims 1 to 7.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the independent routing transmission method for heterogeneous audio and video as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Embedded audio routing switcher
CN101563915A
Method and system for wireless communication of audio in wireless networks
CN102668547A
Audio and video processing method and system
CN104349205A
Method and device for transmitting HDMI (High-Definition Multimedia Interface) audio-video signal
CN104954722A
Audio-video signal processing device and audio-video signal processing method
CN108737759A