Signal transmission method and device, equipment and storage medium
By generating a first time-frequency map and a second time-frequency map, and combining them with a binarized convolutional neural network, the signal protocol category of multi-source video signals is identified and converted into intermediate format data. This solves the problem of insufficient recognition reliability of signal switching devices in scenarios with low signal-to-noise ratio or ambiguous protocol features in the existing technology, and improves the accuracy and reliability of signal switching.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
- Filing Date
- 2026-01-05
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies are insufficient in their ability to identify protocols when dealing with scenarios with low signal-to-noise ratios or ambiguous protocol features, leading to operational malfunctions and reduced system reliability.
By generating a first time-frequency map and a second time-frequency map, and combining them with a binarized convolutional neural network, the signal protocol category of multi-source video signals is identified, and they are converted into intermediate format data and target signals, thereby achieving accurate signal conversion and output.
It improves the recognition accuracy and robustness of signal protocol categories in complex scenarios, ensures the accuracy and reliability of signal switching, and avoids operational failures caused by misjudgment.
Smart Images

Figure CN121967784A_ABST
Abstract
Description
Signal transmission methods, devices, equipment and storage media Technical Field
[0001] This application belongs to the field of signal switching, specifically relating to a signal transmission method and apparatus, an electronic device, and a storage medium. Background Technology
[0002] With the rapid development of automotive electronics, smart display devices and industrial testing systems, video signal conversion equipment needs to process signals of various protocols.
[0003] In related technologies, traditional adapters use FPGAs (Field Programmable Gate Arrays) for protocol parsing and signal enhancement, and introduce binarized convolutional neural networks (CNNs) for automatic protocol recognition to improve the system's intelligence level. However, such systems are prone to misjudgment when dealing with low signal-to-noise ratio or ambiguous protocol features, resulting in insufficient reliability of protocol recognition results and potentially leading to serious operational failures.
[0004] Therefore, improving the robustness and accuracy of signal switching in complex application scenarios has become an urgent problem to be solved. Summary of the Invention
[0005] The purpose of this application is to provide a signal transmission method that can solve the problems of insufficient robustness and accuracy of signal switching in complex application scenarios.
[0006] Accordingly, embodiments of this application also provide a signal transmission device, an electronic device, and a storage medium to ensure the implementation and application of the above methods.
[0007] To solve the above-mentioned technical problems, this application provides the following: Firstly, embodiments of this application provide a signal transmission method, the method comprising: receiving a multi-source video signal; generating a first time-frequency diagram based on the multi-source video signal; determining a first protocol category corresponding to the multi-source video signal based on the first time-frequency diagram; generating a second time-frequency diagram based on the first protocol category; determining a signal protocol category corresponding to the multi-source video signal based on the first time-frequency diagram and the second time-frequency diagram; converting the multi-source video signal into intermediate format data based on the signal protocol category; generating an adjustment signal based on the intermediate format data; and converting the adjustment signal into a target signal and outputting it to a display device.
[0008] Optionally, generating the first time-frequency diagram based on the multi-source video signal includes: performing signal conditioning on the multi-source video signal to obtain a preprocessed signal; converting the preprocessed signal into a digital signal; determining the frequency energy distribution and time variation characteristics of the digital signal; and generating the first time-frequency diagram based on the frequency energy distribution and the time variation characteristics.
[0009] Optionally, generating the second time-frequency map based on the first protocol category includes: obtaining a protocol feature template corresponding to the first protocol category; and generating the second time-frequency map based on the first time-frequency map and the protocol feature template corresponding to the first protocol category.
[0010] Optionally, determining the signal protocol category corresponding to the multi-source video signal based on the first time-frequency map and the second time-frequency map includes: determining a first confidence level of the first protocol category based on the first time-frequency map; determining a second protocol category corresponding to the multi-source video signal and a second confidence level of the second protocol category based on the second time-frequency map; and determining the signal protocol category corresponding to the multi-source video signal from the first protocol category and the second protocol category based on the first confidence level and the second confidence level.
[0011] Optionally, determining the second protocol category corresponding to the multi-source video signal and the second confidence level of the second protocol category based on the second time-frequency map includes: determining the protocol category probability distribution corresponding to the multi-source video signal based on the second time-frequency map; determining the second protocol category based on the protocol category probability distribution; obtaining the activation value of each pixel in the second time-frequency map; determining the activity and sparsity index of the second time-frequency map based on the activation value; and determining the second confidence level based on the activity and the sparsity index.
[0012] Optionally, determining the signal protocol category corresponding to the multi-source video signal from the first protocol category and the second protocol category based on the first confidence level and the second confidence level includes: determining the signal protocol category corresponding to the multi-source video signal as the second protocol category when the second confidence level is greater than the sum of the first confidence level and the preset confidence level tolerance; and determining the signal protocol category corresponding to the multi-source video signal as the first protocol category when the second confidence level is less than or equal to the sum of the first confidence level and the preset confidence level tolerance.
[0013] Optionally, generating an adjustment signal based on the intermediate format data includes: enhancing the intermediate format data to generate an enhanced signal; determining the pixel coordinates of each pixel in the image of a preset output size in the enhanced signal; determining the neighborhood reference point corresponding to the pixel coordinates; and generating the adjustment signal based on the pixel coordinates, the pixel value of the pixel, and the neighborhood reference point.
[0014] Secondly, embodiments of this application provide a signal transmission apparatus, the apparatus comprising: a receiving module for receiving multi-source video signals; a first time-frequency diagram generation module for generating a first time-frequency diagram based on the multi-source video signals; a first protocol category determination module for determining a first protocol category corresponding to the multi-source video signals based on the first time-frequency diagram; a second time-frequency diagram generation module for generating a second time-frequency diagram based on the first protocol category; a signal protocol category determination module for determining a signal protocol category corresponding to the multi-source video signals based on the first time-frequency diagram and the second time-frequency diagram; a signal conversion module for converting the multi-source video signals into intermediate format data based on the signal protocol category; an adjustment signal generation module for generating an adjustment signal based on the intermediate format data; and a signal output module for converting the adjustment signal into a target signal and outputting it to a display device.
[0015] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0016] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0017] Compared with existing technologies, the embodiments of this application have the following advantages: In this embodiment, multi-source video signals are received; a first time-frequency map is generated based on the multi-source video signals; a first protocol category corresponding to the multi-source video signals is determined based on the first time-frequency map; a second time-frequency map is generated based on the first protocol category; a signal protocol category corresponding to the multi-source video signals is determined based on the first and second time-frequency maps; the multi-source video signals are converted into intermediate format data based on the signal protocol category; an adjustment signal is generated based on the intermediate format data; and the adjustment signal is converted into a target signal and output to a display device. By generating the first and second time-frequency maps, the embodiments of this application can accurately identify the signal protocol category corresponding to the multi-source video signals, improving the identification accuracy of the signal protocol category in complex scenarios. Furthermore, by converting the multi-source video signals into unified intermediate format data based on the identified signal protocol category, generating an adjustment signal from the intermediate format data, and then converting the adjustment signal into a target signal and outputting it to a display device, the robustness and accuracy of signal switching in complex scenarios can be guaranteed. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 is a flowchart of one embodiment of the signal transmission method of this application; Figure 2 is a hardware architecture diagram of one embodiment of the signal transmission method of this application; Figure 3 is a hardware architecture diagram of another embodiment of the signal transmission method of this application; Figure 4 is a structural block diagram of one embodiment of the signal transmission device of this application. Detailed Implementation
[0020] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0022] The signal transmission method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0023] With the rapid development of automotive electronics, intelligent display devices, and industrial testing systems, video signal conversion equipment needs to handle various protocol signals, such as High Definition Multimedia Interface (HDMI), Low Voltage Differential Signaling (LVDS), DisplayPort (DP), Video Graphic Array (VGA), Mobile Industry Processor Interface (MIPI), and automotive bus protocols, such as Controller Area Network (CAN), Local Interconnect Network (LIN), and FlexRay bus. Current technological trends favor highly integrated modular designs, combining FPGAs and dedicated conversion chips to achieve multi-protocol compatibility, high efficiency, and flexible configuration. For example, traditional converters use FPGAs for protocol parsing and signal enhancement, and introduce convolutional neural networks (CNNs) for automatic protocol recognition to improve system intelligence. However, these systems suffer from insufficient reliability in handling low signal-to-noise ratio or ambiguous protocol features, limiting their application in complex industrial environments.
[0024] In existing technologies, binarized CNN protocol recognition systems suffer from "deterministic collapse" in scenarios with low signal-to-noise ratios (SNR) or ambiguous protocol features. The binarization process compresses continuous activation values into discrete ±1 values, leading to the loss of subtle feature information. When the input signal's SNR decreases due to transmission loss or electromagnetic interference, or when different protocol features overlap in the time-frequency domain, the decision-making process of the binarized CNN becomes fragile, and the model outputs incorrect protocol categories with high confidence. Incorrect identification of signal protocol categories can lead to serious operational malfunctions. In automotive electronics testing environments, low-quality signals originate from long-distance wiring or harsh electromagnetic environments; incorrect protocol identification can cause display devices to experience black screens, distorted images, or data parsing failures, thus delaying fault diagnosis and testing processes. In industrial display systems, incorrect protocol identification can cause multi-screen collaboration interruptions or touch function malfunctions, reducing system reliability and user experience. Furthermore, high-confidence erroneous outputs mask the root cause of the problem, complicating debugging and maintenance, and increasing system downtime and costs.
[0025] In some embodiments of this application, a signal transmission method is proposed that can effectively improve the identification accuracy of signal protocol categories in complex scenarios, thereby ensuring the robustness and accuracy of signal switching in complex scenarios.
[0026] Referring to Figure 1, a flowchart of a signal transmission method embodiment of this application is shown, including the following steps: Step 101, receiving multi-source video signals; wherein, multi-source video signals refer to video signals from multiple protocol sources, and the protocol sources corresponding to the multi-source video signals may include at least two of HDMI, LVDS, DP, VGA, MIPI CSI-2 (Mobile Industry Processor Interface Camera Serial Interface-2, a camera serial interface specification defined by the Mobile Industry Processor Interface Alliance), CAN, LIN, FlexRay, Ethernet, and USB (Universal Serial Bus).
[0027] Step 102: Generate a first time-frequency diagram based on the multi-source video signal; wherein, the first time-frequency diagram is used to reflect the characteristic distribution of the multi-source video signal in the time-frequency domain, and each pixel contained in the diagram can reflect the frequency domain energy distribution and time variation characteristics of the multi-source video signal.
[0028] Step 103: Determine the first protocol category corresponding to the multi-source video signal based on the first time-frequency diagram; wherein, the first protocol category is the protocol category corresponding to the multi-source video signal predicted by inputting the first time-frequency diagram into a binary convolutional neural network (CNN).
[0029] Step 104: Generate a second time-frequency map according to the first protocol category; wherein the second time-frequency map can reflect the time-frequency structure and non-dominant feature information of the multi-source video signal, and can maintain the continuity of the input dimension and probability distribution of the multi-source video signal.
[0030] Step 105: Determine the signal protocol category corresponding to the multi-source video signal based on the first time-frequency diagram and the second time-frequency diagram; wherein, the signal protocol category is the signal protocol category with the highest probability predicted from the multi-source video signal.
[0031] Step 106: Convert the multi-source video signal into intermediate format data according to the signal protocol category; wherein, the intermediate format data may include RGB (red, green, and blue primary colors) format data or YUV (a color encoding method) format data. RGB format data can maintain high color accuracy, which is beneficial for high dynamic range data display, while YUV format data allows one signal to serve both black-and-white and color devices simultaneously, and can significantly reduce the amount of data by reducing the chroma sampling rate, thereby improving signal transmission efficiency.
[0032] Step 107: Generate an adjustment signal based on the intermediate format data; wherein, the adjustment signal refers to the signal data generated after signal adjustment of the intermediate format data.
[0033] Step 108: Convert the adjustment signal into a target signal and output it to the display device.
[0034] The target signal is the signal whose signal protocol category matches that of the display device.
[0035] In this embodiment, multi-source video signals from multiple protocol sources can be received. A first time-frequency map is generated based on the multi-source video signals. The first time-frequency map is input into a binarized CNN to predict the first protocol category corresponding to the multi-source video signals. A second time-frequency map is then generated based on the first protocol category. Finally, the signal protocol category corresponding to the multi-source video signals is determined based on the first and second time-frequency maps. After determining the signal protocol category of the multi-source video signals, the multi-source video signals can be converted into intermediate format data according to the signal protocol category. An adjustment signal is generated based on the intermediate format data, and then the adjustment signal is converted into a target signal and output to a display device. Through the above implementation process, using the first and second time-frequency maps to determine the signal protocol category corresponding to the multi-source video signals can effectively improve the recognition accuracy of signal protocol categories in complex scenes. Furthermore, converting the multi-source video signals into unified intermediate format data based on the identified signal protocol category facilitates data transmission compatibility across multiple devices, makes it easier to transmit signals between different devices, optimizes data transmission efficiency, or ensures the color accuracy of the output signal. After generating an adjustment signal from the intermediate format data, the adjustment signal can be converted into a target signal and output to the display device. This enables the transfer of multi-source video signals to the selected display device, ensuring the robustness and accuracy of signal transfer in complex scenarios.
[0036] For example, a multi-source video signal can be represented as a collection of video signals from multiple protocol sources. Where n can represent the number of input channels, i.e., the number of protocol sources. It can represent the i-th video signal in the set, each This corresponds to video signals from different protocol sources. After receiving multi-source video signals, a control signal can be generated by the upper-level configuration module managed by the MCU (Microcontroller Unit) controller. This signal is used to control the multiplexer (MUX) to activate a specified channel and output the video signal corresponding to that specified channel. The upper-layer configuration module manages user-defined protocol priorities and input detection rules. Protocol priorities refer to the user-defined order of priority for different types of signal protocols, typically set based on actual business needs or signal importance. Input detection rules detect the video signals corresponding to each input channel and generate corresponding input detection results based on the detected video signal status and quality. Control signals are generated based on the protocol priorities in the upper-layer configuration module and the input detection results obtained from the input detection rules. Through this implementation process, unified scheduling of input port resources can be achieved, allowing the selection of one or more video signals for transmission from multiple sources. Prioritizing high-priority signal channels based on protocol priorities avoids low-priority traffic consuming resources and causing critical business delays or interruptions. Combining input detection results to determine the output video signal optimizes resource utilization. If some signals are temporarily unavailable, their corresponding channels can be skipped and resources allocated to active signals, reducing interference from low-quality signals and ensuring the reliability of signal transmission.
[0037] In some embodiments of this application, generating a first time-frequency diagram based on the multi-source video signal includes: performing signal conditioning on the multi-source video signal to obtain a preprocessed signal; converting the preprocessed signal into a digital signal; determining the frequency energy distribution and time variation characteristics of the digital signal; and generating a first time-frequency diagram based on the frequency energy distribution and the time variation characteristics.
[0038] After receiving multi-source video signals, signal conditioning can be performed on the multi-source video signals to obtain a pre-processed signal. Signal conditioning may include at least one of impedance matching, level shifting, and noise removal.
[0039] For example, impedance matching for multi-source video signals may include the following steps: obtaining a preset target impedance; and using a matching network to adjust the original impedance of the multi-source video signals to the preset target impedance by adjusting the microstrip line width, dielectric thickness, and compensation capacitor structure. Impedance matching can be expressed by the following formula:
[0040] in, It can represent the original impedance of multi-source video signals; It can represent a preset target impedance, which can be set to 50Ω (ohms) to match the requirements of high-frequency signal transmission, or it can be set according to actual needs. This application does not impose specific restrictions on this. The characteristic impedance of the transmission line. The relative permittivity of the medium, Relative permeability 、, and This information can be obtained from the specifications of the transmission line and medium. By implementing the above process, adjusting the original impedance of the multi-source video signal to match the target impedance can effectively reduce the reflection coefficient. This ensures that the energy transfer efficiency of the signal is maximized during transmission.
[0041] For example, after impedance matching of the multi-source video signal, level conversion can be performed on the impedance-matched multi-source video signal. Level conversion of the multi-source video signal can include the following steps: obtaining a preset reference voltage reference; adjusting the original voltage of the multi-source video signal according to the reference voltage reference using a voltage divider resistor network to adjust the original voltage to an output voltage conforming to the target logic standard. Level conversion can be expressed by the following formula:
[0042] in, The output voltage obtained after level conversion of multi-source video signals; The preset reference voltage reference can be determined according to the input standard of the compatible FPGA, MCU or video decoder; and These represent the resistance values of the upper and lower voltage divider resistors, respectively. By selecting precision resistors as voltage divider resistors, and combining them with a temperature compensation circuit, the influence of temperature on the resistance values can be offset, ensuring that the output voltage after level conversion of multi-source video signals remains stable within the input standard range compatible with FPGAs, MCUs, or video decoders. This implementation process avoids subsequent logic misjudgments caused by excessively high or low levels of multi-source video signals.
[0043] For example, after level conversion of the multi-source video signal, noise removal can be performed on the converted signal. Noise removal of the multi-source video signal can include the following steps: obtaining a preset filtering bandwidth; determining the corresponding filtering resistor and capacitor based on the filtering bandwidth; determining the corresponding cutoff frequency based on the filtering resistor and capacitor; and filtering the multi-source video signal using a passive low-pass filter network based on the cutoff frequency. The cutoff frequency can be expressed by the following formula:
[0044] in, R can represent the cutoff frequency; R can represent the filter resistor; C can represent the filter capacitor. The filter bandwidth can be controlled by adjusting the product of R and C. Through the above implementation process, a passive low-pass filter network can be used to suppress high-frequency interference in multi-source video signals, so that the in-band video signal fidelity is maintained at its highest level, and out-of-band high-frequency noise is effectively attenuated to obtain a high-quality pre-processed signal.
[0045] The preprocessed signal obtained after signal conditioning is an analog signal. This preprocessed signal can be converted from an analog signal to a digital signal. Then, the frequency energy distribution and time variation characteristics corresponding to the digital signal are extracted from the converted digital signal. Based on the extracted frequency energy distribution and time variation characteristics, the one-dimensional digital signal is mapped to a multi-dimensional time-frequency diagram, generating a first time-frequency diagram. For example, the first time-frequency diagram can be a 64×64×1 time-frequency diagram. The dimensions of the first time-frequency diagram can also be set according to actual needs; this application does not impose specific limitations on this.
[0046] Through the above implementation process, signal conditioning of multi-source video signals can significantly improve the signal-to-noise ratio and significantly reduce waveform jitter and amplitude drift. This provides highly stable input features for subsequent signal protocol category identification, making the identified signal protocol category results more accurate and reliable. By converting the preprocessed signal from analog to digital, and then based on the extracted frequency energy distribution and time variation characteristics, the preprocessed signal can be converted into a first time-frequency map where each pixel reflects the frequency domain energy distribution and time variation characteristics.
[0047] For example, the first time-frequency graph can be... After amplitude normalization, the normalization formula for the first time-frequency plot is as follows:
[0048] in, This can represent the first time-frequency graph after normalization. This can represent the original first time-frequency diagram; It can represent the mean of the preprocessed signal, and is used to reflect the overall energy center of the preprocessed signal; This can represent a preset standard deviation, used to control the normalization amplitude range. Through the above implementation process, updating the first time-frequency image to the normalized first time-frequency image can eliminate the influence of different sampling intensities and sensor gain differences on the feature distribution, making the first time-frequency image input to the binarized CNN satisfy the zero mean and unit variance characteristics, thereby improving the stability of protocol category recognition.
[0049] In some embodiments of this application, determining the first protocol category corresponding to the multi-source video signal based on the first time-frequency map may include: inputting the first time-frequency map into a binarized CNN deployed in an FPGA for protocol recognition, and obtaining the output first protocol category.
[0050] For example, the first convolutional layer of a binarized CNN can use 16 convolutional kernels of size 3×3, and the corresponding network weights can be represented as follows after binarization: ,in, This represents the network weights obtained after binarization. It can represent the original real weight matrix of the first convolutional layer; The function can reset positive weights to +1 and negative weights to 1. 1. This allows the binarized result to... The value can only be +1 or 1. Transforming convolution calculations into addition and subtraction operations reduces multiplication operations and lowers hardware implementation complexity, thereby reducing FPGA logic resource consumption. The activation features of the first convolutional layer have also undergone binarization and can be represented as follows: ,in, It can represent the activation features obtained after binarization; This can represent the intermediate activation values of the convolution output, which, after binarization, yield a binary activation matrix. This allows the network to rely solely on symbolic computation during the forward inference phase. The feature map output from the first convolutional layer can be 62×62×16 in size. It can then enter the first pooling layer, where a 2×2 max pooling operation is used to downsample the feature map to a size of 31×31×16. By using reasonable downsampling, the computational complexity can be reduced by compressing the spatial size while preserving the main spatial structure information.
[0051] The second convolutional layer can use 32 3×3 convolutional kernels, and the corresponding network weights, after binarization, can be represented as follows: The corresponding activation features, after binarization, can be represented as: The first convolutional layer outputs a feature map with a size of 29×29×32, which is then compressed using 2×2 max pooling to obtain a feature map with a size of 14×14×32. After two convolution and pooling operations, the frequency domain structure and protocol features of the preprocessed signal (i.e., the first time-frequency map) can be extracted into a high-dimensional sparse representation to obtain the feature map to be identified, which is used to capture multi-scale protocol signal texture features. The feature map to be identified can be stored in hardware as a bitstream to improve throughput. The flattened feature map to be identified can be input into a fully connected layer, which can contain 128 neurons for nonlinear mapping. The output is processed by the Softmax (normalization exponent) function to form a probability distribution of the signal protocol category. , where k is the number of protocol categories that can be supported for identification. The output of the Softmax function can be expressed by the following formula:
[0052] Where z can represent the linear output of the fully connected layer. This can represent the probability that a signal belongs to the i-th type of signal protocol; This can represent the linear combination output corresponding to the identified i-th signal protocol category, used to characterize the signal feature strength at the corresponding location; k is the number of protocol categories supported for identification. The signal protocol category with the highest probability distribution among the signal protocol categories can be determined as the first protocol category. .
[0053] In some embodiments of this application, generating a second time-frequency map based on the first protocol category includes: obtaining a protocol feature template corresponding to the first protocol category; and generating a second time-frequency map based on the first time-frequency map and the protocol feature template corresponding to the first protocol category.
[0054] After determining the first protocol category corresponding to the multi-source video signals, the protocol feature template corresponding to the first protocol category can be obtained. The protocol feature template reflects the average feature shape of the corresponding protocol category and can be generated based on the training dataset of the binarized CNN system used for protocol feature recognition during the offline training phase. The protocol feature template corresponding to the first protocol category can be expressed by the following formula:
[0055] in, This can represent the protocol feature template corresponding to the first protocol category C1; It can represent the number of samples in the training dataset. This can represent the time-frequency plot of the j-th sample in the training dataset. The protocol feature template corresponding to the first protocol category can be obtained by averaging the time-frequency plots of samples corresponding to the same protocol category in the training dataset, and is used to represent the steady-state mode and feature energy distribution of the corresponding protocol.
[0056] A dynamic stripping strength coefficient can be obtained. Based on the stripping strength coefficient and the protocol feature template corresponding to the first protocol category, feature stripping (feature suppression) is performed on the first time-frequency map to generate a second time-frequency map. The second time-frequency map can be represented by the following formula:
[0057] in, can represent the second time-frequency diagram; X can represent the first time-frequency diagram; ⊙ represents element-wise multiplication; α can represent the peel strength coefficient, with a value range of [0,1]. This can represent the protocol feature template corresponding to the first protocol category C1. As the value of α increases, The energy in the high-response region is weakened by a larger proportion, suppressing the dominant features of protocol C1 in X, thereby revealing other masked protocol features. This is equivalent to weighted inverse cancellation of the dominant channel in the feature space to reduce the interference of strong signal patterns on weak signal patterns. The resulting second time-frequency map can retain the time-frequency structure and non-dominant feature information of the preprocessed signal, and maintain the continuity of the input dimension and distribution of the preprocessed signal, ensuring that the second time-frequency signal can also be input into the binarized CNN for protocol category identification.
[0058] Through the above implementation process, a second time-frequency diagram can be generated, which can effectively reduce the masking effect of the main protocol features on the secondary protocol features, thereby ensuring the accuracy and robustness of signal protocol category identification in complex application scenarios such as mixed input of multiple source protocols.
[0059] In some embodiments of this application, determining the signal protocol category corresponding to the multi-source video signal based on the first time-frequency map and the second time-frequency map includes: determining a first confidence level of the first protocol category based on the first time-frequency map; determining a second protocol category corresponding to the multi-source video signal and a second confidence level of the second protocol category based on the second time-frequency map; and determining the signal protocol category corresponding to the multi-source video signal from the first protocol category and the second protocol category based on the first confidence level and the second confidence level.
[0060] The first confidence level can be used to evaluate the credibility of the predicted first protocol category. Determining the first confidence level of the first protocol category based on the first time-frequency map may include: determining the first activation value corresponding to the first time-frequency map, the number of first activation values being determined based on the size of the first time-frequency map and the hyperparameters of the convolutional layer of the binarized CNN, the hyperparameters of the convolutional layer may include parameters such as the size of the convolutional kernel and stride; obtaining the total number of pixels in the first time-frequency map; determining the activity and sparsity indices corresponding to the first time-frequency map based on the first activation value and the total number of pixels in the first time-frequency map; and determining the first confidence level based on the activity and sparsity indices of the first time-frequency map. For example, the stripping strength coefficient can be dynamically associated with the first confidence level. If the first confidence level is lower than a preset first confidence level threshold, the stripping strength coefficient is automatically increased to enhance the weight of weak features and perform adaptive recognition compensation.
[0061] The second confidence level can be used to evaluate the reliability of the predicted second protocol category. The second protocol category and its second confidence level can be determined based on the second time-frequency map. Then, based on the first and second confidence levels, the signal protocol category corresponding to the multi-source video signal is determined from the first and second protocol categories. Through this process, multi-level identification of multi-source video signals can be performed. The first protocol category and first confidence level are obtained from the first time-frequency map, and the second protocol category and second confidence level are obtained from the second time-frequency map. Then, based on the first and second confidence levels, the signal protocol category corresponding to the multi-source video signal is determined from the first and second protocol categories. This effectively avoids the "deterministic collapse" misjudgment caused by feature overload when binarized CNNs identify signal protocol categories, effectively ensuring the accuracy of protocol identification and the reliability of decision-making.
[0062] For example, while outputting the predicted signal protocol category, the preprocessed signal parsing data can also be output. This parsing data is obtained by restoring the feature mapping relationship of the first time-frequency diagram to the protocol frame structure information and payload. It can provide semantic-level input support for the conversion of intermediate format data, and realize high-speed and accurate adaptive signal recognition and parsing in the multi-source protocol signal environment.
[0063] In some embodiments of this application, determining the second protocol category corresponding to the multi-source video signal and the second confidence level of the second protocol category based on the second time-frequency map includes: determining the protocol category probability distribution corresponding to the multi-source video signal based on the second time-frequency map; determining the second protocol category based on the protocol category probability distribution; obtaining the activation value of each pixel in the second time-frequency map; determining the activity and sparsity index of the second time-frequency map based on the activation value; and determining the second confidence level based on the activity and the sparsity index.
[0064] The second protocol category refers to the protocol category corresponding to the multi-source video signal predicted by inputting the second time-frequency image into a binary convolutional neural network (CNN). The second time-frequency image can be input into a binary CNN deployed in an FPGA, and the output is a protocol category probability distribution corresponding to the multi-source video signal formed by the Softmax function. This protocol category probability distribution represents the probability that each video signal in the multi-source video signal conforms to its respective identifiable signal protocol category. The specific generation process of the protocol category probability distribution can be found in the aforementioned content on determining the first protocol category corresponding to the multi-source video signal based on the first time-frequency image, and will not be repeated here. The signal protocol category with the highest probability distribution can be selected as the second protocol category.
[0065] The activation values of each pixel in the second time-frequency image can be obtained, and the activity and sparsity indices of the second time-frequency image can be determined based on these activation values. Activity reflects the degree of signal response to the current network's (binarized CNN) convolutional kernels; a higher value indicates that the signal features of the protocol category have a more significant representation in the feature space. Activity can be expressed by the following formula:
[0066] Where A can represent activity level; N can represent the total number of pixels in the time-frequency graph; It can represent the activation value of the i-th pixel position in the time-frequency graph.
[0067] The sparsity index reflects the stability of the feature distribution in a time-frequency plot (feature plot). The sparsity index can be expressed by the following formula:
[0068] Where σ can represent the sparsity index, which is used to characterize the degree of dispersion of the activation distribution; N can represent the total number of pixels in the time-frequency plot; can represent the activation value of the i-th pixel in the time-frequency graph; μ can represent the mean of all activation values in the time-frequency graph, used to characterize the center position of the feature response. Lower sparsity, i.e., a smaller sparsity index, means that the feature map response is concentrated and stable, while higher sparsity, i.e., a larger sparsity index, indicates that the feature map response has large fluctuations.
[0069] After obtaining the activity and sparsity indices corresponding to the second time-frequency plot, the second confidence level can be determined based on these indices. This includes: assigning dynamic weighting coefficients to the activity and sparsity indices corresponding to the second time-frequency graph, and performing a weighted summation of the activity and sparsity indices corresponding to the second time-frequency graph based on the assigned dynamic weighting coefficients to obtain a second confidence level; or, determining the feature matching scores of the first and second time-frequency graphs, normalizing the feature matching score, activity, and sparsity indices corresponding to the second time-frequency graph, and using the average value of the normalized feature matching score, activity, and sparsity indices of the second time-frequency graph as the second confidence level.
[0070] Through the above implementation process, the activity and sparsity indices of the second time-frequency graph can be obtained. The activity and sparsity indices can be used to judge the effectiveness of the features. The second confidence score generated based on the activity and sparsity indices of the second time-frequency graph can effectively evaluate the credibility of the predicted second protocol category.
[0071] In some embodiments of this application, determining the signal protocol category corresponding to the multi-source video signal from the first protocol category and the second protocol category based on the first confidence level and the second confidence level includes: determining the signal protocol category corresponding to the multi-source video signal as the second protocol category when the second confidence level is greater than the sum of the first confidence level and the preset confidence level tolerance; and determining the signal protocol category corresponding to the multi-source video signal as the first protocol category when the second confidence level is less than or equal to the sum of the first confidence level and the preset confidence level tolerance.
[0072] An arbitration logic comparison mechanism can be adopted to determine the signal protocol category corresponding to the multi-source video signal by comparing the first confidence level and the second confidence level. The confidence tolerance refers to the acceptable range of confidence fluctuations, which is used to offset slight confidence fluctuations caused by network weight binarization or input signal disturbances. The numerical range of the confidence tolerance can be set to [0.1, 0.3], or it can be set separately according to the actual situation. This application does not impose specific restrictions on it.
[0073] If the second confidence level is greater than the sum of the first confidence level and a preset confidence level tolerance, the signal protocol category corresponding to the multi-source video signal is determined as the second protocol category. If the second confidence level is less than or equal to the sum of the first confidence level and a preset confidence level tolerance, the signal protocol category corresponding to the multi-source video signal is determined as the first protocol category. Through the above implementation process, it can be ensured that if the second confidence level corresponding to the second protocol category identified by the signal protocol category identification of the second time-frequency image is significantly higher than the first confidence level corresponding to the first protocol category identified by the first time-frequency image, the final identification result is switched from the first protocol category to the second protocol category. Otherwise, the first protocol category initially identified based on the first time-frequency image is retained, and the first identification category is used as the final identification result. Protocol identification for the first time-frequency image focuses on the extraction of the dominant protocol, while protocol identification for the second time-frequency image focuses on the supplementary identification of weak features. The combination of the two can form a confidence-driven adaptive correction mechanism for protocol categories, thereby improving the stability and accuracy of signal protocol category identification.
[0074] In some embodiments of this application, generating an adjustment signal based on the intermediate format data includes: enhancing the intermediate format data to generate an enhanced signal; determining the pixel coordinates of each pixel in the image of a preset output size in the enhanced signal; determining the neighborhood reference point corresponding to the pixel coordinates; and generating the adjustment signal based on the pixel coordinates, the pixel value of the pixel, and the neighborhood reference point.
[0075] Signal enhancement can be performed on intermediate format data to generate an enhanced signal. Signal enhancement can include noise suppression, color correction, and edge enhancement.
[0076] For example, noise suppression of intermediate format data includes: employing a median filtering method to suppress noise in the time domain of the intermediate format data, thereby eliminating high-frequency isolated noise points and smoothing local brightness variations. The filter window size can be set to w=3, or it can be set differently according to actual needs; this application does not impose specific limitations on this. Using the filter, the median value of each pixel (x, y) in the intermediate format data can be taken as the output value within its local neighborhood. Noise suppression can be expressed by the following formula:
[0077] in, It can represent the output value of the filter, that is, the intermediate format data after noise suppression; As a median function, the median value can be obtained by sorting and taking the middle value, thereby suppressing isolated noise peaks without destroying the image edge structure; This can represent intermediate format data; x and y represent the horizontal and vertical coordinates of each pixel in the intermediate format data, respectively; i and j represent the row and column indices of the feature map, respectively, used to locate the specific position of each pixel in the feature map. Through the above implementation process, median filtering of the intermediate format data can eliminate high-frequency isolated noise points in the intermediate format data, and maintain the temporal smoothness of the noise-suppressed image while preserving high-frequency structural information, providing a stable input for subsequent color enhancement. For example, after noise suppression of the intermediate format data, color correction can be performed on the noise-suppressed intermediate format data, including: performing Gamma correction on the intermediate format data according to a preset correction coefficient; determining the channel average brightness of the gamma-corrected intermediate format data; adjusting the gamma-corrected intermediate format data according to the channel average brightness; and generating color-corrected intermediate format data.
[0078] Gamma correction can be expressed by the following formula: .in, It can represent intermediate format data after gamma correction; It can represent the intermediate format data after noise suppression; x and y represent the horizontal and vertical coordinates of each pixel in the intermediate format data, respectively; γ can represent the correction coefficient, which can be used to map the signal strength from the linear space to the human eye perception space, so that the image brightness distribution is more in line with visual characteristics and improves the overall contrast. γ can be set to 2.2, or it can be set separately according to the actual situation. This application does not impose specific restrictions on it.
[0079] The intermediate format data after gamma correction can then proceed to the white balance adjustment stage, where the grayscale world algorithm is used to correct brightness unevenness between channels. The average channel brightness of the gamma-corrected intermediate format data can be expressed by the following formula:
[0080] in, It can represent the average brightness of a channel in intermediate format data after gamma correction; , and It can represent the average brightness of the RGB three channels.
[0081] The intermediate format data after color correction can be represented by the following formula:
[0082] in, It can represent intermediate format data after color correction; It can represent intermediate format data after gamma correction; x and y represent the horizontal and vertical coordinates of each pixel in the intermediate format data, respectively; It can represent the channel average brightness of intermediate format data after gamma correction; This can represent the average brightness of the channel to which the current pixel belongs. Through the above implementation process, three-channel brightness equalization can be achieved, making the image tone tend towards a natural grayscale distribution and eliminating color shifts caused by signal conversion or illuminance differences.
[0083] For example, after color correction, edge enhancement can be performed on the intermediate format data after color correction. This includes using the Laplacian operator to extract high-frequency gradient information from the intermediate format data after color correction, thereby enhancing the structural boundaries of the intermediate format data. The Laplacian operator K can be expressed by the following formula:
[0084] The enhancement signal is generated by performing a second derivative operation on the color-corrected intermediate format data using this Laplacian operator, which can be expressed by the following formula:
[0085] in, It can represent an enhanced signal; It can represent intermediate format data after color correction; λ is a preset enhancement intensity coefficient, which can be used to control the edge enhancement amplitude. When the λ value is appropriate, the detail edges are clearly enhanced while the noise is suppressed. 2This can represent the Laplacian operator operation. The above implementation process can enhance the structural boundaries of the intermediate format data after color correction, resulting in an enhanced signal.
[0086] After obtaining the enhanced signal, the pixel coordinates of each pixel in the image of the preset output size can be determined. The preset output size can be a 16:9 aspect ratio such as 1920×1080 or 1280×720 to adapt to mainstream display devices. For example, it can be based on the image size of the enhanced signal. With preset output size The proportional relationship between them, where, It can represent the signal width of the enhanced signal. This can represent the signal height of the enhanced signal. It can represent the signal width of the output signal. This can represent the signal height of the output signal. Each output pixel is determined based on this proportional relationship. The corresponding coordinates (x, y) in the image of the corresponding multi-source video signal can be represented by the following formula:
[0087] Then, we can determine the four neighboring reference points corresponding to the pixel coordinates (x, y). The neighboring reference points can be represented as: , , , ,in, , , , , This indicates rounding down. This indicates rounding up to the nearest integer.
[0088] An adjustment signal can be generated by interpolation calculation based on pixel coordinates, pixel values, and neighboring reference points. The interpolation weights can be defined by linear attenuation based on distance. The adjustment signal can be expressed by the following formula:
[0089] in, It can represent an adjustment signal. Indicates the input signal at coordinate point The pixel value at that location; x and y represent the enhanced signal mapping position in continuous coordinates of the preset output size; These are the discrete neighborhood boundaries of the pixel coordinates in the enhanced signal, and the discrete neighborhood boundaries are determined based on the neighborhood reference points.
[0090] Through the above implementation process, signal enhancement of intermediate format data can significantly improve the signal-to-noise ratio, color reproduction, and texture clarity of the generated enhanced signal compared to the original intermediate format data. This provides a high-reliability input for the generation of the adjustment signal, ensuring visual consistency and image quality stability of multi-source video signals during the transition process. Resolution scaling of the enhanced signal can be achieved spatially using a bilinear interpolation algorithm, realizing distance-weighted intensity allocation in the spatial domain. This ensures that the pixel values of the output adjustment signal maintain consistency with the local gradient continuity of the input enhanced signal, thereby avoiding edge blurring and feature distortion during resolution scaling.
[0091] For example, the frame rate of the enhanced signal can also be adjusted in the time dimension to obtain an adjusted signal, including: determining the signal frame rate of the enhanced signal; and determining a preset target frame rate. Signal frame rate with enhanced signal The current ratio, the current ratio The target frame rate can be 25fps (Frames Per Second), 30fps, or 60fps to adapt to the needs of video transmission and display scenarios. It can also be set differently based on actual transmission and real-world scenario requirements; this application does not impose specific restrictions on this. When the current ratio is greater than 1, intermediate frames are inserted between adjacent frames in the enhanced signal. These intermediate frames are generated using an interpolation algorithm by extracting pixel features and motion information from adjacent frames. When the current ratio is less than 1, low-information-content frames in the enhanced signal are identified and discarded. Through the above implementation process, the insertion and discarding strategy of the frame buffer can be controlled by the current ratio obtained from the target frame rate. When ρ > 1, meaning the current enhanced signal frame rate is relatively low, intermediate frames are inserted to improve the frame rate while ensuring the continuity of inter-frame transitions, achieving a smooth transition. When ρ < 1, meaning the current enhanced signal frame rate is relatively high, low-information-content frames are discarded to reduce the frame rate while maintaining temporal balance.
[0092] The resolution and frame rate of the enhanced signal obtained after noise suppression and edge enhancement may not match the requirements of subsequent display output. After adjusting the resolution in the spatial dimension and the frame rate in the temporal dimension, the final adjusted signal can achieve a consistent distribution in both the spatial and temporal domains. The adjusted spatiotemporal resolution can match the feature recognition requirements in the subsequent display output process. The features of the obtained adjusted signal remain stable and the sampling density is uniform, which can provide a standardized feature flow for subsequent recognition and analysis.
[0093] In some embodiments of this application, the adjustment signal is a digital image stream. Converting the adjustment signal into a target signal and outputting it to a display device includes: converting the adjustment signal of the digital image stream into a standard data packet format; encoding the adjustment signal of the standard data packet format to obtain a digital video stream; acquiring an audio signal; generating a target signal based on the digital video stream and the audio signal; and outputting the target signal to the display device.
[0094] For example, converting the adjustment signal of a digital image stream into a standard data packet format may include: using a MIPI converter LT7911UXE chip to perform interface protocol conversion of the input signal, converting the adjustment signal of the digital image stream into a standard data packet format conforming to the MIPI DSI or CSI standard. During the conversion process, a timing-consistent transport stream can be constructed based on a pixel clock synchronization mechanism. The resulting standard data packet format can be expressed by the following formula: ,in, This can represent the signal flow after MIPI encapsulation, which is the standard data packet format. It can represent an adjustment signal. This can represent a set of synchronization control signals, which can be parsed from the timing features extracted from the preprocessed signal. These signals include horizontal synchronization signals, frame synchronization signals, and pixel clock signals, and can be used to synchronize the horizontal and vertical timing of the video signal, ensuring stable image resolution by the display device. This conversion process, through clock domain alignment and bit width matching mechanisms, ensures the timing consistency of the pixel stream and control signals during bus transmission, avoiding phase shifts and data loss between high-frequency channels.
[0095] For example, after converting the adjustment signal of the digital image stream into a standard data packet format, the Gscoolink GSV6172 chip can be invoked to perform second-stage encoding, encoding the adjustment signal of the standard data packet format to obtain a digital video stream. Second-stage encoding can convert the MIPI data stream... (That is, the adjusted signal converted to the standard data packet format) is converted to an HDMI format signal, and encoded using the Transition Minimized Differential Signaling (TMDS) algorithm. The encoded digital video stream can be represented by the following formula:
[0096] in, This refers to the digital video stream encoded with TMDS. This indicates that TMDS encoding is being performed. The TMDS encoding process maps every 8 bits of video data to 10 bits of transmission symbols to reduce DC bias and improve noise immunity.
[0097] The encoding module can also embed audio signal channels via the I2S (Inter-IC Sound, integrated circuit built-in audio) bus. The system acquires audio signals, generates target signals based on the digital video and audio streams, and outputs these target signals to the display device, achieving synchronous output of multimedia signals. The target signal may contain video streams, audio streams, and control channel information, and is output at the HDMI, DP, or Type-C physical layer according to the target interface protocol. The system's physical layer interface features hot-plug detection and EDID reading capabilities. By reading the extended display identification data of the display terminal, it automatically adjusts the output resolution, color depth, and refresh rate to ensure that the signal format is consistent with the terminal's display parameters. This enables plug-and-play and intelligent matching at the output end, ensuring strict synchronization of video and audio in both the time and frequency domains while maintaining signal transmission integrity and electrical compatibility between devices.
[0098] For example, a binary CNN can be pre-trained using a training dataset containing real labels. During the training of the binary CNN, the cross-entropy loss function can be used to measure the difference between the predicted distribution and the true distribution. The cross-entropy loss function can be expressed by the following formula:
[0099] Where L can represent the cross-entropy loss function; The one-hot encoded value of the real label can be obtained from the annotation information (label) of the training dataset by converting the real protocol category into a one-hot vector form; It can represent the predicted probability of the corresponding protocol category.
[0100] For example, the network weights of a binary CNN can also be iteratively updated using the gradient descent algorithm. The update of the network weights can be expressed by the following formula:
[0101] in, It can represent the updated network weights; η can represent the current network weights; η can represent the preset learning rate, used to control the parameter update step size; This can represent the gradient of the loss function with respect to the weights. The process of updating and optimizing the network weights can be performed iteratively on a training dataset containing waveforms with various protocol information. By adjusting the network weights of the binarized CNN layer by layer through gradient propagation, the network can still maintain high recognition accuracy under limited computing resources.
[0102] Referring to Figure 2, which is a hardware architecture diagram of an embodiment of a signal transmission method of this application, the hardware architecture is an HDMI modular adapter, which includes: a USB interface, an ADS signal input module, an FPGA main control chip-MIPI controller, a MIPI DSI / CSI converter (chip LT7911UXE), a MIPI DSI / CSI converter (chip Gscoolink GSV6172), DDR SDRAM (Double Data Rate Synchronous Dynamic Random Access Memory), and a connector.
[0103] The USB interface supports HID or I²C protocols and can be used to connect 4×4 (4 groups × 4 channels) of input signals.
[0104] ADS signal input module: This can be an ADS series analog-to-digital converter (ADC) chip, which can convert the input signal from an analog signal to a digital signal and then send it to the FPGA main control chip - MIPI controller.
[0105] FPGA Main Control Chip - MIPI Controller: This refers to the FPGA main control chip with a MIPI controller, such as the Xilinx Zynq UltraScale+ chip. This chip integrates an FPGA and an ARM-based processing system. The FPGA main control chip - MIPI controller includes a deserializer / transceiver, registers, an aggregation chip, a signal enhancement module, a protocol conversion module, and a resolution / frame rate adjustment module. The deserializer / transceiver extracts timing parameters and pixel format information from the input signal; registers are used to set timing parameters; the aggregation chip combines multi-source input signals, including at least two of the following communication protocols: RGB, YUV, RAW (a computer network protocol), and TTL (Transistor-Transistor Logic), into a single output signal to save resources and improve system efficiency; the signal enhancement module enhances the combined input signal; the resolution / frame rate adjustment module adjusts the resolution and frame rate of the combined input signal; and the protocol conversion module converts the enhanced and adjusted input signal into a unified intermediate format data.
[0106] DDRSDRAM: A dual data rate synchronous dynamic random access memory. Input signals are processed by format recognition and scaling inside the FPGA main control chip - MIPI controller, and can be stored in DDRSDRAM.
[0107] MIPI DSI / CSI Converter (Chip LT7911UXE): This converter can receive intermediate format data sent by the FPGA main control chip - MIPI controller and convert it into the standard data packet format for generating the MIPI DSI / CSI standard; MIPI DSI / CSI Converter (Chip Gscoolink GSV6172): This converter can receive DSI / CSI data packets sent by chip LT7911UXE, perform TMDS encoding, generate digital video streams, and can also combine audio signals from the audio channel to generate target signals containing video streams, audio streams, and control channel information.
[0108] Connector: The generated target signal can be used as the final output, and output at the HDMI, DP or Type-C physical layer according to the target interface protocol.
[0109] In Figure 2, "Touch + Splicing" represents the input signal processing in the HDMI modular adapter. It refers to the process where a 4×4 input signal undergoes analog-to-digital conversion by the ADS signal input module, then enters the FPGA main control chip with a MIPI controller. Within the FPGA main control chip-MIPI controller, the signal is processed and stored using a DDR SDRAM module. Finally, it is converted by two MIPI DSI / CSI converters and output as HDMI, DP, or TYPE-C signals through the connector. The sliding operation on the right can adjust the display area resolution and perform resolution compression, facilitating testing of this hardware architecture.
[0110] Using the hardware architecture shown in Figure 2, the dedicated signal enhancement algorithms integrated on the FPGA (Laplacian edge enhancement and Gamma color correction) effectively improve image details and color accuracy. Meanwhile, the resolution adjustment mechanism based on bilinear interpolation and frame rate conversion ensures perfect compatibility of the output signal with various display devices. Ultimately, this achieves stable output of high-resolution, high-refresh-rate lossless video signals while maintaining low power consumption and high real-time performance, overcoming the technical bottlenecks of signal quality loss, excessive latency, and limited compatibility commonly found in existing conversion equipment.
[0111] Referring to Figure 3, it is a hardware architecture diagram of another embodiment of the signal transmission method of this application. This hardware architecture integrates multi-protocol compatibility, binarized CNN protocol recognition and iterative inference functions, including an input layer, a processing layer, a control layer, an output layer and a power module. The layers are interconnected through a high-speed data bus to ensure low-latency processing.
[0112] Input layer (multi-source video signal access module): Used to receive video signals from various protocol sources, such as HDMI, LVDS, DP, VGA, MIPICSI-2, CAN, LIN, FlexRay, Ethernet, and USB interfaces. This may include modular interfaces, a multiplexer (MUX), and signal conditioning circuitry.
[0113] The modular interface can include centralized input modules such as a display signal input module, a camera signal input module, an ECU signal input module, and an ADS signal input module, which are used to receive display signals, camera signals, ECU signals, and ADS signals, respectively. The modular interface can use pluggable physical connectors, such as MIPI connectors, RJ45, and Type-C connectors. The display signal input module can receive signals such as HDMI, LVDS, DP, and VGA via a serializer / deserializer (SerDes); the camera signal input module can receive signals such as HDMI, LVDS, and MIPI via SerDes; the ECU signal input module can receive signals such as CAN, LIN, and FlexRay via SerDes; and the ADS signal input module can receive signals such as RJ45, Ethernet, and USB via SerDes.
[0114] Among them, the multiplexer MUX can use a switching chip (such as ADG1612) to dynamically select the active input channel and output the selected signal according to the control signal configured by the user.
[0115] The signal conditioning circuit can include at least an impedance matching network, a level shifting chip, and a noise removal circuit. The impedance matching network, designed for high-frequency signals (such as MIPICSI-2) in multi-source video signals, adjusts the input impedance of the multi-source video signal to a target impedance of 50Ω. The level shifting chip (such as TXB0108) converts the input voltage of the multi-source video signal to a preset output voltage. The noise removal circuit uses a low-pass filter to remove high-frequency noise components from the multi-source video signal. After processing the input multi-source video signal, the signal conditioning circuit generates a corresponding pre-processed signal and outputs it to the processing layer.
[0116] Processing layer (signal processing module): This layer integrates binarized CNN protocol recognition and iterative inference logic, with the FPGA as its core. It may include: the FPGA main control chip, high-speed memory, and a MIPI converter.
[0117] The FPGA main control chip (such as Xilinx Zynq UltraScale+MPSoC) can include a binarization CNN module, a confidence evaluation module, a feature stripping module, an arbitration logic module, a protocol conversion submodule, a signal enhancement submodule, and a resolution adjustment submodule. The binarization CNN module, confidence evaluation module, feature stripping module, arbitration logic module, and protocol conversion submodule work together to perform protocol parsing, data extraction, and format conversion on preprocessed signals of various protocol types transmitted by the signal conditioning circuit, obtaining intermediate format data such as RGB or YUV. This data is then transmitted to the integrated ISP chip. The signal enhancement submodule of the ISP chip can perform noise suppression, color correction, and edge enhancement on the intermediate format data, and transmit the generated enhanced signal to the integrated Scaler chip. The resolution adjustment submodule then adjusts the current resolution to the target resolution and the current frame rate to the target frame rate, resulting in parallel RGB data encoding.
[0118] The binarization CNN module can be a lightweight convolutional neural network, with a structure including an input layer, convolutional layers (16 3×3 convolutional kernels, weight binarization, activation binarization), pooling layers (2×2 max pooling), and fully connected layers (128 neurons) used to output the probability distribution of the signal protocol category. The confidence evaluation module can analyze the activity and sparsity of the feature maps of the intermediate layers of the binarized CNN to obtain the confidence of the predicted protocol category. The feature stripping module can load protocol feature templates from high-speed memory to modify the original feature map (first time-frequency map) to obtain the second time-frequency map. The arbitration logic module can compare the first and second confidence scores to output the final signal protocol category prediction result. The protocol conversion submodule can parse the preprocessed signal based on the output signal protocol category and convert it into a unified intermediate format data. The signal enhancement submodule can perform signal enhancement operations such as noise suppression, color correction, and edge enhancement. The resolution adjustment submodule can scale the resolution using bilinear interpolation and adjust the refresh rate using frame rate conversion technology.
[0119] The FPGA main control chip can encode and transmit the generated parallel RGB data to the MIPI converter.
[0120] Among them, MIPI converters (such as the LT7911UXE chip) can convert the signals processed by the FPGA main control chip into MIPI DSI / CSI or LVDS formats, supporting multiple resolutions and frame rates.
[0121] The high-speed memory (such as DDR4 RAM and Flash) can include DDR4 cache (such as the MT40A512M16LY-075E chip) and Flash memory (W25Q256JV chip). DDR4 cache can be used to store multi-frame image data, protocol feature templates, and intermediate feature maps (such as the first time-frequency map, the second time-frequency map, etc.). Flash memory can be used to store binarized CNN model parameters, firmware, and configuration parameters.
[0122] Control layer (system management and configuration module): Used to coordinate the operation of various modules. This may include the MCU controller, user interface, and communication interface.
[0123] The MCU controller is used for scheduling and control layer monitoring of the entire process to monitor system status. It receives configuration commands and manages user configurations through the user interface, and supports remote debugging and firmware upgrades through the communication interface. When input signal loss or format errors are detected, it can also trigger LED alarms and output UART debug logs.
[0124] The user interface is used to provide remote command interaction for I²C / SPI (Serial Peripheral Interface) / physical buttons / LED (light-emitting diode) indicators / displays, and is used for parameter settings (such as resolution and protocol type).
[0125] The communication interface supports multiple communication connections such as USB / UART / Ethernet / Wi-Fi, enabling remote control and OTA (Over-the-Air Technology) upgrades, such as USB firmware upgrades, and can also output UART debug logs.
[0126] Output layer (HDMI signal generation module): Used to convert the processed signal into a standard output format. This may include a MIPI converter, physical interface, and auxiliary circuitry.
[0127] Among them, the MIPI converter (such as the Gscoolink GSV6172 chip) can be used to receive MIPI DSI / CSI signals and convert them into HDMI, DP or Type-C formats, and integrates a TMDS encoder (such as ADIADV7511).
[0128] The physical interface can provide HDMI / DP / Type-C connectors and supports hot plug detection (HPD) and EDID (Extended Display Identification Data) reading.
[0129] The auxiliary circuit may include ESD (Electro-Static discharge) protection (such as TVS (Transient voltage suppression) diodes), filter circuits, and power management modules (such as PMICTPS650861 chips).
[0130] Power module: Can be a DC-DC or LDO power module for power supply. It can accept input voltages of 12V / 24V / USB 5V / battery 3.7V and output voltages of 5V / 3.3V / 1.8V, etc.
[0131] For example, during the overall system integration phase, the overall performance of the system can be evaluated. The evaluation process can use latency, power consumption, and error rate as core indicators to quantify and optimize the overall performance.
[0132] The latency metric consists of input time, processing time, and output time, with the total latency expressed as: .in This indicates the time required for input data to enter the processing unit, reflecting the efficiency of interface response and cache scheduling; This indicates the computation execution time of the core algorithm on an FPGA or embedded platform, and is related to the data block size, parallelism, and pipeline depth. The transmission time of the result signal output to the display or control unit is affected by the output bus bandwidth and communication protocol. , and It can be measured using timing analysis tools or platform timing modules.
[0133] The power consumption model is ,in The dynamic and static power consumption generated by the main computing unit during logic gate switching and on-chip memory access depends on the clock frequency and logic utilization. The energy consumed for signal driving and protocol conversion in high-speed data links reflects the matching between signal transmission rate and interface level. The power consumption of the external power supply regulator module, clock management circuit, and sensor power supply section reflects the energy consumption characteristics of the system's auxiliary circuits. , , Data can be collected using a power analyzer or chip monitoring interface.
[0134] Error rate assessment can be performed using protocol category identification accuracy. ,in The number of signal frames correctly identified by the system represents the number of times the identification module's output matches the reference tag. This represents the total number of signal frames for all recognition tasks. , The accuracy is obtained by statistically analyzing the number of correct identifications and the total number of identifications; this accuracy calculation directly reflects the reliability of the algorithm in time-series signal analysis and protocol classification.
[0135] The core of system integration is ensuring the consistency of data flow transmission between each step. Signals maintain timing continuity between different modules through a bus mapping table and clock synchronization strategy. Finally, the output signal, after passing through a display adapter interface and protocol encapsulation, matches the format and rate with the display device, thereby ensuring the real-time performance and visualization accuracy of the output content and achieving closed-loop stable operation of the entire system from signal generation to result presentation. In the hardware architecture shown in Figure 3, a multiplexer can dynamically select input channels according to user instructions, and in conjunction with an FPGA-implemented protocol conversion, signal enhancement, and resolution adjustment pipeline, a single device can flexibly adapt to various signal sources, from displays and cameras to vehicle buses. This highly integrated architecture replaces the existing solution that requires multiple independent conversion devices for different protocols, significantly reducing system size and cost, and greatly improving testing and deployment efficiency through a unified configuration interface, meeting the stringent requirements of vehicle systems for device multifunctionality and space utilization.
[0136] The hardware architecture shown in Figure 2 or Figure 3 can execute the signal transmission methods provided in some embodiments of this application and be applied to various practical scenarios.
[0137] For example, signal conversion can be performed between display devices, which may include displays, smart TVs, projectors, etc. It supports multiple display interfaces with different protocols such as HDMI / LVDS / DP / VGA input and converts the input signals into HDMI / DP / TYPE-C output high-definition video signals, which can be spliced and used for touch screens.
[0138] For example, it can convert signals from multi-functional camera devices, support multiple camera interface MIPI-CSI-2 / LVDS input, and convert the input signals into HDMI / DP / TYPE-C output high-definition video signals, converting multiple channels into a single channel.
[0139] For example, it can also be applied to automotive electronics testing. It can support ECU input via CAN, LIN, FlexRay, and other automotive electronic interfaces, converting them to HDMI / DP / TYPE-C signals for connection to the display screen of diagnostic equipment, facilitating fault diagnosis by technicians. Multiple DUTs can also be simultaneously input via modular adapters and output to a single display screen for splicing, timing synchronization, resolution compression or enhancement, improving testing efficiency and reducing unnecessary data processing workload.
[0140] For example, it can also be applied to in-vehicle embedded systems, supporting high-speed interface inputs such as multiple cameras, sensors, and Ethernet from Advanced Driving Assistance Systems (ADAS). It supports multiple signal inputs from navigation systems, reversing cameras, and multimedia devices, overlaying them onto the video stream in OSD (on-screen display) form, and converting them into HDMI / DP / TYPE-C output high-resolution video signals to the in-vehicle display screen or monitoring system. It can realize real-time video processing functions such as picture-in-picture or multi-screen touch display.
[0141] It should be noted that the signal transmission method provided in this application embodiment can be executed by a signal transmission device, or a control module in the signal transmission device for executing the loading signal transmission method. This application embodiment uses the execution of the loading signal transmission method by a signal transmission device as an example to illustrate the signal transmission method provided in this application embodiment.
[0142] Referring to Figure 4, a structural block diagram of an embodiment of a signal transmission device according to this application is shown, including the following modules: a receiving module 401, used to receive multi-source video signals; a first time-frequency diagram generation module 402, used to generate a first time-frequency diagram based on the multi-source video signals; a first protocol category determination module 403, used to determine a first protocol category corresponding to the multi-source video signals based on the first time-frequency diagram; a second time-frequency diagram generation module 404, used to generate a second time-frequency diagram based on the first protocol category; a signal protocol category determination module 405, used to determine a signal protocol category corresponding to the multi-source video signals based on the first time-frequency diagram and the second time-frequency diagram; a signal conversion module 406, used to convert the multi-source video signals into intermediate format data according to the signal protocol category; an adjustment signal generation module 407, used to generate an adjustment signal based on the intermediate format data; and a signal output module 408, used to convert the adjustment signal into a target signal and output it to a display device.
[0143] The first time-frequency diagram generation module 402 includes: a signal conditioning submodule for conditioning the multi-source video signal to obtain a preprocessed signal; a conversion submodule for converting the preprocessed signal into a digital signal; a feature determination submodule for determining the frequency energy distribution and time variation characteristics of the digital signal; and a first time-frequency diagram generation submodule for generating a first time-frequency diagram based on the frequency energy distribution and the time variation characteristics.
[0144] The second time-frequency diagram generation module 404 includes: a protocol feature template acquisition submodule, used to acquire the protocol feature template corresponding to the first protocol category; and a second time-frequency diagram generation submodule, used to generate a second time-frequency diagram based on the first time-frequency diagram and the protocol feature template corresponding to the first protocol category.
[0145] The signal protocol category determination module 405 includes: a first confidence level determination submodule, configured to determine a first confidence level of the first protocol category based on the first time-frequency diagram; a second protocol category and second confidence level determination submodule, configured to determine a second protocol category corresponding to the multi-source video signal and a second confidence level of the second protocol category based on the second time-frequency diagram; and a signal protocol category determination submodule, configured to determine the signal protocol category corresponding to the multi-source video signal from the first protocol category and the second protocol category based on the first confidence level and the second confidence level.
[0146] The second protocol category and second confidence level determination submodule includes: a protocol category probability distribution determination unit, configured to determine the protocol category probability distribution corresponding to the multi-source video signal based on the second time-frequency map; a second protocol category determination unit, configured to determine the second protocol category based on the protocol category probability distribution; an activation value acquisition unit, configured to acquire the activation value of each pixel in the second time-frequency map; an activity and sparsity index determination unit, configured to determine the activity and sparsity index of the second time-frequency map based on the activation value; and a second confidence level determination unit, configured to determine the second confidence level based on the activity and sparsity index.
[0147] The signal protocol category determination submodule is further configured to: determine the signal protocol category corresponding to the multi-source video signal as a second protocol category when the second confidence level is greater than the sum of the first confidence level and a preset confidence level tolerance; and determine the signal protocol category corresponding to the multi-source video signal as a first protocol category when the second confidence level is less than or equal to the sum of the first confidence level and a preset confidence level tolerance.
[0148] The adjustment signal generation module 407 includes: an enhancement signal generation submodule, used to enhance the intermediate format data to generate an enhancement signal; a pixel coordinate determination submodule, used to determine the pixel coordinates of each pixel in the image of a preset output size in the enhancement signal; a neighborhood reference point determination submodule, used to determine the neighborhood reference point corresponding to the pixel coordinates; and an adjustment signal generation submodule, used to generate the adjustment signal based on the pixel coordinates, the pixel value of the pixel, and the neighborhood reference point. In this embodiment, the signal transmission device can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. The embodiments of this application do not impose specific limitations.
[0149] The signal transmission device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.
[0150] The signal transmission device provided in this application embodiment can implement the various processes implemented by the signal transmission device in the method embodiments of Figures 1 to 3. To avoid repetition, these processes will not be described again here.
[0151] This application provides a signal transmission device that can receive multi-source video signals from multiple protocol sources, generate a first time-frequency map based on the multi-source video signals, input the first time-frequency map into a binarized CNN to predict the first protocol category corresponding to the multi-source video signals, generate a second time-frequency map based on the first protocol category, and then determine the signal protocol category corresponding to the multi-source video signals based on the first and second time-frequency maps. After determining the signal protocol category of the multi-source video signals, the multi-source video signals can be converted into intermediate format data according to the signal protocol category, an adjustment signal can be generated based on the intermediate format data, and then the adjustment signal can be converted into a target signal and output to a display device. Through the above implementation process, using the first and second time-frequency maps to determine the signal protocol category corresponding to the multi-source video signals can effectively improve the recognition accuracy of signal protocol categories in complex scenes. Furthermore, converting the multi-source video signals into unified intermediate format data based on the identified signal protocol category facilitates data transmission compatibility across multiple devices, makes it easier to transmit signals between different devices, optimizes data transmission efficiency, or ensures the color accuracy of the output signal. After generating an adjustment signal from the intermediate format data, the adjustment signal can be converted into a target signal and output to the display device. This enables the transfer of multi-source video signals to the selected display device, ensuring the robustness and accuracy of signal transfer in complex scenarios.
[0152] Optionally, embodiments of this application also provide an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the various processes of the above-described signal transmission method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0153] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0154] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described signal transmission method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0155] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0156] This application also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described signal transmission method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0157] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0158] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0159] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0160] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A signal transmission method, characterized in that, The method includes: receiving multi-source video signals; generating a first time-frequency diagram based on the multi-source video signals; determining a first protocol category corresponding to the multi-source video signals based on the first time-frequency diagram; generating a second time-frequency diagram based on the first protocol category; determining a signal protocol category corresponding to the multi-source video signals based on the first time-frequency diagram and the second time-frequency diagram; converting the multi-source video signals into intermediate format data based on the signal protocol category; generating an adjustment signal based on the intermediate format data; and converting the adjustment signal into a target signal and outputting it to a display device.
2. The method according to claim 1, characterized in that, The step of generating a first time-frequency diagram based on the multi-source video signal includes: performing signal conditioning on the multi-source video signal to obtain a preprocessed signal; converting the preprocessed signal into a digital signal; determining the frequency energy distribution and time variation characteristics of the digital signal; and generating a first time-frequency diagram based on the frequency energy distribution and the time variation characteristics.
3. The method according to claim 1, characterized in that, The step of generating a second time-frequency map based on the first protocol category includes: obtaining a protocol feature template corresponding to the first protocol category; and generating a second time-frequency map based on the first time-frequency map and the protocol feature template corresponding to the first protocol category.
4. The method according to claim 1, characterized in that, The step of determining the signal protocol category corresponding to the multi-source video signal based on the first time-frequency diagram and the second time-frequency diagram includes: determining a first confidence level of the first protocol category based on the first time-frequency diagram; determining a second protocol category corresponding to the multi-source video signal and a second confidence level of the second protocol category based on the second time-frequency diagram; and determining the signal protocol category corresponding to the multi-source video signal from the first protocol category and the second protocol category based on the first confidence level and the second confidence level.
5. The method according to claim 4, characterized in that, The step of determining the second protocol category corresponding to the multi-source video signal and the second confidence level of the second protocol category based on the second time-frequency map includes: determining the protocol category probability distribution corresponding to the multi-source video signal based on the second time-frequency map; determining the second protocol category based on the protocol category probability distribution; obtaining the activation value of each pixel in the second time-frequency map; determining the activity and sparsity index of the second time-frequency map based on the activation value; and determining the second confidence level based on the activity and the sparsity index.
6. The method according to claim 4, characterized in that, The step of determining the signal protocol category corresponding to the multi-source video signal from the first protocol category and the second protocol category based on the first confidence level and the second confidence level includes: determining the signal protocol category corresponding to the multi-source video signal as the second protocol category when the second confidence level is greater than the sum of the first confidence level and the preset confidence level tolerance; and determining the signal protocol category corresponding to the multi-source video signal as the first protocol category when the second confidence level is less than or equal to the sum of the first confidence level and the preset confidence level tolerance.
7. The method according to claim 1, characterized in that, The step of generating an adjustment signal based on the intermediate format data includes: enhancing the intermediate format data to generate an enhanced signal; determining the pixel coordinates of each pixel in the image of a preset output size in the enhanced signal; determining the neighborhood reference point corresponding to the pixel coordinates; and generating the adjustment signal based on the pixel coordinates, the pixel value of the pixel, and the neighborhood reference point.
8. A signal transmission device, characterized in that, The apparatus includes: a receiving module for receiving multi-source video signals; a first time-frequency diagram generation module for generating a first time-frequency diagram based on the multi-source video signals; a first protocol category determination module for determining a first protocol category corresponding to the multi-source video signals based on the first time-frequency diagram; a second time-frequency diagram generation module for generating a second time-frequency diagram based on the first protocol category; a signal protocol category determination module for determining a signal protocol category corresponding to the multi-source video signals based on the first and second time-frequency diagrams; a signal conversion module for converting the multi-source video signals into intermediate format data based on the signal protocol category; an adjustment signal generation module for generating an adjustment signal based on the intermediate format data; and a signal output module for converting the adjustment signal into a target signal and outputting it to a display device.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the signal transmission method as described in claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the signal transmission method as described in claims 1-7.