Image and video transmission device based on narrowband audio

By using extreme compression and error correction coding techniques based on deep learning networks, images and videos are converted into feature bitstreams for transmission in narrowband audio channels. This solves the problem of low transmission efficiency of images and videos in narrowband audio channels, achieving efficient and reliable image and video transmission, and is suitable for various communication devices and environments.

CN121814962APending Publication Date: 2026-04-07XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing narrowband audio communication methods are difficult to achieve real-time transmission of images and videos, especially in narrowband audio channels with limited bandwidth. The data transmission efficiency of images and videos is low and cannot meet the needs of modern communication.

Method used

An extreme compression method based on deep learning networks is used to convert images and videos into feature bitstreams. Combined with error correction coding and audio modulation techniques, image and video information is transmitted in narrowband audio channels. Generative adversarial networks are used for efficient compression and decoding, supporting real-time transmission of images and videos in narrowband audio channels.

Benefits of technology

It enables efficient transmission of images and videos in narrowband audio channels, with compression ratios of tens to thousands of times, ensuring the recognizability and integrity of images and videos. It is suitable for various communication environments, compatible with existing voice communication systems, supports various topologies and devices, and improves information density and transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121814962A_ABST
    Figure CN121814962A_ABST
Patent Text Reader

Abstract

An image and video transmission device based on narrowband audio comprises a coding sending end, a decoding receiving end and a data transmission module, and the coding sending end comprises an extreme image and video coding compression module, an error correction coding module, an audio modulator and an audio player. The decoding receiving end comprises an audio recorder, an audio demodulator, an error code correction module and an extreme image video decoding and decompression module, and the coding sending end converts image videos into narrowband audios after extreme coding compression and modulation and outputs the narrowband audios; the data transmission module transmits information in a wired or wireless mode; the decoding receiving end demodulates, corrects and decodes the input narrowband audio and then restores the narrowband audio into an image video to be output; the device has the advantages of being high in compression ratio, high in interconnection and intercommunication applicability and narrow in bandwidth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the fields of image and video processing and wireless communication technology, and particularly relates to an image and video transmission device based on narrowband audio. Background Technology

[0002] In fields such as shortwave / ultra-shortwave voice communication, satellite voice communication, WeChat voice calls, wired telephones, video conferencing, and industrial and defense voice communication, narrowband audio is basically used as the voice communication method. Narrowband audio typically transmits speech in the frequency range of 300Hz to 3,400Hz, which basically covers most of the consonants and vowels in speech, ensuring clarity and intelligibility.

[0003] The advantages of using narrowband audio are: 1) Saves bandwidth (the bandwidth of a single voice transmission is generally less than 64kbps), requires less encoding and transmission data, and is more efficient with limited transmission bandwidth resources; 2) Good compatibility, almost all traditional telephone equipment and services are compatible with narrowband audio; 3) Strong anti-interference ability, narrowband encoding is more stable in some environments with weak signals or large interference, such as voice calls in emergency, industrial and defense fields.

[0004] Because narrowband audio (such as the 300Hz~3400Hz bandwidth of telephone voice) has an extremely low data rate, typically only tens to thousands of bits per second (bps), transmitting a simple small image (a few KB) can take several seconds to tens of seconds. For example, in the early days of the internet, internet access was often achieved using a modem connected to a telephone line. With the rapid development of internet technology and the increase in network transmission bandwidth, the method of using a modem to connect to a telephone line for data transmission and reception has become obsolete. Currently, voice communication methods based on narrowband audio are still widely used, such as online video / voice conferencing, walkie-talkies, shortwave / ultra-shortwave communication, satellite communication, emergency communication, and underwater communication. Research on image and video transmission using extremely low-bandwidth narrowband audio channels has not seen major breakthroughs for many years. However, with the rapid development of artificial intelligence and ultra-high image and video compression technology, real-time image and video transmission based on narrowband audio channels will become possible, thus fully leveraging the convenient interconnectivity between narrowband audio devices for efficient image and video transmission. Even in the current era of 5G / 6G broadband communication, this technology still has significant application value in transmitting more multimedia data using limited bandwidth, especially in smartphone satellite communication and shortwave / ultra-shortwave emergency communication. Summary of the Invention

[0005] To address the aforementioned problems, this disclosure provides an image and video transmission device based on narrowband audio, comprising: an encoding transmitter, a decoding receiver, and a data transmission module. The encoding transmitter includes an extreme image and video encoding and compression module, an error correction encoding module, an audio modulator, and an audio player. The decoding receiver includes an audio recorder, an audio demodulator, an error correction module, and an extreme image and video decoding and decompression module. The encoding and transmitting end converts the image and video into narrowband audio output after extreme encoding compression and modulation. The data transmission module transmits information via wired or wireless means; The decoding receiver demodulates, corrects, and decodes the input narrowband audio, then restores it to an image or video output.

[0006] This disclosure also provides an electronic device, including: a memory, a processor, and module programs stored in the memory and capable of running on the processor for an encoding transmitter and a decoding receiver, wherein the processor executes the module programs.

[0007] Through the above technical solution, the device utilizes narrowband audio to transmit images and videos, featuring high compression ratio, strong interoperability and applicability, and narrow bandwidth, thus solving the problem of real-time transmission of images and videos from narrowband audio. This device can be widely used in shortwave / ultra-shortwave communication, BeiDou communication, smartphone satellite communication, video intercoms, online video / voice conferencing, underwater communication, and other fields. Attached Figure Description

[0008] Figure 1 This is a schematic diagram of the structure of an extreme image and video transmission device based on narrowband audio provided in one embodiment of this disclosure; Figure 2 This is a schematic diagram of an application of a device combined with a shortwave / ultra-shortwave radio provided in one embodiment of the present disclosure; Figure 3 This is a schematic diagram illustrating another application of the device combined with a satellite communication system provided in one embodiment of the present disclosure; Figure 4 This is a schematic diagram of the coding network structure of an extreme compression method provided in one embodiment of this disclosure; Figure 5 This is a schematic diagram of the probabilistic model structure of an extreme compression method provided in one embodiment of this disclosure; Figure 6 This is a schematic diagram of the decoding network structure of an extreme compression method provided in one embodiment of this disclosure; Figure 7 This is a schematic diagram of the discriminant network structure of an extreme compression method provided in one embodiment of this disclosure; Figures 8(a) to 8(c) are diagrams showing the extreme compression effect of an image by an extreme image and video transmission device based on narrowband audio provided in one embodiment of this disclosure. Detailed Implementation

[0009] In one embodiment, such as Figure 1 As shown, a narrowband audio-based image and video transmission device is disclosed, comprising: an encoding transmitter, a decoding receiver, and a data transmission module. The encoding transmitter includes an extreme image and video encoding and compression module, an error correction encoding module, an audio modulator, and an audio player. The decoding receiver includes an audio recorder, an audio demodulator, an error correction module, and an extreme image and video decoding and decompression module. The encoding and transmitting end converts the image and video into narrowband audio output after extreme encoding compression and modulation. The data transmission module transmits information via wired or wireless means; The decoding receiver demodulates, corrects, and decodes the input narrowband audio, then restores it to an image or video output.

[0010] In this embodiment, the extreme image and video encoding compression module employs an extreme compression method based on deep learning networks to convert image and video information into a feature bitstream (a bitstream represented by 0s and 1s), with a compression ratio of over 100 times. That is, extreme encoding compression refers to a compression ratio of at least 100 times, meaning the compressed size is 1 / 100th or less of the original size. Generally, the higher the compression ratio, the lower the quality of the decompressed and recovered image and video.

[0011] In this embodiment, the feature stream corresponds to the encoded and compressed image and video content, which is generated by the encoding and compression deep learning network. It is also a bit stream represented by 0 and 1, but the size of the encoded and compressed feature stream is only one-thousandth to one-several-tenths of the original image and video size, that is, it achieves compression of tens to thousands of times. Conversely, the feature stream can be decoded by the original compressed deep learning network to generate the original image and video information.

[0012] In another embodiment, the extreme image and video encoding compression module uses an extreme compression method based on deep learning networks to convert image and video information into a feature bitstream.

[0013] In this embodiment, the extreme compression method based on deep learning networks employs a generative adversarial network (GAN) architecture, which includes an encoding network, a probabilistic model, a decoding network, and a discriminant network. Wherein, as... Figure 4 As shown, the encoding network first uses convolutional layers with a kernel size of 7. Channel normalization ReLU activation extracts shallow features from the input image. Then, a basic module consisting of three convolutional layers with a kernel size of 3, a channel normalization layer, and ReLU activation is used to compress the shallow features into compact features. Finally, a convolutional layer with a kernel size of 3 is used to output 220 latent features. Probabilistic models, see Figure 5 First, use convolutional layers. ReLU activation Convolutional layer ReLU activation The processing mode of convolutional layers transforms latent features into super-prior latent variables. The super-prior latent variable is quantized and entropy encoded using a prior bottleneck entropy model to obtain quantized features. Then use convolutional layers ReLU activation Convolutional layer ReLU activation Convolutional layers process features layer by layer Estimate its mean and scale The Gaussian conditional entropy model uses these parameters to accurately estimate the probability distribution of each element in the latent features. The decoding network, acting as the generator, see... Figure 6 First, channel normalization is adopted. Convolutional layer Channel normalization enhances the decoded features, and then the features are refined using nine residual blocks, each consisting of a convolutional layer with a kernel size of 3. Channel normalization ReLU activated convolutional layer Channel normalization is then applied. Following this, four deconvolutional layers with a kernel size of 3 are used. Channel normalization The basic module composed of ReLU activations performs feature upsampling, ensuring that the features have the same spatial resolution as the original input image. Finally, a convolutional layer with a kernel size of 7 maps the features to the pixel space, yielding the final decoded image. Figure 7 As shown, the discriminant network uses convolutional layers on one hand. Spatial normalization Leak ReLU Activation Nearest neighbor upsampling will input latent variables The data is converted into conditional features, then concatenated with the input image along the channel dimension, and then processed using four convolutional layers with downsampling capabilities. Spatial normalization The process is performed using basic modules composed of Leak ReLU activations, and finally passed through convolutional layers. Spatial normalization Leak ReLU activated convolutional layer Spatial normalization Adjustment operation Sigmoid activation outputs the probability that the original image is a real image. Adversarial training improves the perceptual quality of reconstructed images under extremely low bitrate conditions. Training the above image compression architecture typically employs a two-stage training method. The first stage trains only the encoder network, probabilistic model, and decoder network, using the following loss function:

[0014] in, , and These represent the bit rate, mean squared error loss, and perceptual loss, respectively. Hyperparameters and Normal settings And 5.0. When the total loss tends towards a certain stable value, the current model reaches the optimal trade-off point between compression and distortion. In the second stage, the discriminator network and all models from the first stage are jointly trained, and the above loss function is adjusted as follows:

[0015] in, Set it to 0.15. To generate adversarial loss, it is represented as follows:

[0016] in, This indicates the discriminant network. When the total loss... Once the value drops to a stable level, the current model can be considered to have achieved the optimal trade-off between compression, perception, and distortion.

[0017] In another embodiment, the error correction coding module converts the feature code stream into an encoded code stream through error correction coding, which is used to detect and correct bit errors generated during data transmission.

[0018] In this embodiment, the error correction coding module converts the feature bitstream into an encoded bitstream through error correction coding. This allows for the detection and correction of bit errors caused by noise, interference, or channel distortion during data transmission, ensuring the integrity of the bit data during communication. This module typically employs a low-rate error correction coding method, increasing redundancy to enhance error correction capabilities. One example uses Reed-Solomon codes (RS), Repetition Codes (RC), Long Constrained Length Convolutional Codes (LCLCC), Turbo codes, Low-Density Parity-Check Codes (LDPC), or Polar codes for error correction coding. Another example uses a concatenated error correction coding method, such as concatenating RS codes and convolutional codes. This involves RS coding, interleaving, and convolutional coding of the feature bitstream to output an encoded bitstream; the receiving end then performs convolutional decoding, deinterleaving, and RS decoding on the encoded bitstream to recover the feature bitstream. Concatenated error correction coding methods also include Turbo code + RS code concatenation, LDPC code + RS code concatenation, or adaptive concatenated coding.

[0019] In another embodiment, the audio modulator modulates the encoded bitstream into a digital audio signal in a digital signal manner.

[0020] In this embodiment, the audio modulator "loads" the coded bitstream onto an audio carrier as a digital signal, that is, modulates the digital signals (0s and 1s) into a digital audio signal (PCM format) suitable for transmission in narrow-band acoustic channels. Different audio frequencies, phases, or amplitudes are typically used to represent different digital bits. This audio signal can be captured by an audio recorder (microphone) at the receiving end and then demodulated to recover the digital signal (coded bitstream). One example is the use of Frequency Shift Keying (FSK), Phase Shift Keying (PSK), Quadrature QPSK, or Quadrature Amplitude Modulation (QAM) techniques.

[0021] The audio modulator can skip the error correction coding module and directly modulate the feature bitstream. That is, when the audio transmission channel quality is high, it does not need to go through the error correction coding module and the bit error correction module, which helps to improve the efficiency and real-time performance of image and video transmission.

[0022] In another embodiment, the audio player is used to play and output the digital audio signal generated by the audio modulator.

[0023] In this embodiment, the audio player plays and outputs the digital audio signal generated by the audio modulator. Typically, the sound card hardware of the electronic device outputs an analog narrowband audio signal, and its output hardware interface can be a 3.5mm standard audio interface, a dedicated audio interface, or a wireless Bluetooth audio interface.

[0024] In another embodiment, the audio recorder acquires and captures the analog narrowband audio signal transmitted by the data transmission module through an audio interface at the decoding receiver, and converts the narrowband audio signal into a digital audio signal through pulse code modulation (PCM).

[0025] In this embodiment, the audio recorder, i.e. the microphone, acquires and captures the analog narrowband audio signal transmitted by the data transmission module through an audio interface (wired or wireless) at the decoding receiver, and converts it into a digital audio signal (PCM format) through PCM (Pulse Code Modulation).

[0026] In another embodiment, the audio demodulator demodulates the digital audio signal at the decoding receiver, accurately identifies the original frequency, phase, or amplitude changes, and converts them back into the encoded bitstream.

[0027] In this embodiment, the audio demodulator demodulates the recorded digital audio (which may contain noise and distortion) at the decoding receiver, accurately identifies the original frequency, phase, or amplitude changes, and converts them back into the encoded bitstream.

[0028] The aforementioned audio demodulator can directly demodulate the feature bitstream without going through the error correction coding module and the bit error correction module when the audio transmission channel quality is high.

[0029] In another embodiment, the error correction module restores the characteristic bitstream from the encoded bitstream at the decoding receiver by performing error correction.

[0030] In this embodiment, the error correction module is the decoding process of the error correction coding module, and the method used corresponds to that of the error correction coding module. At the decoding receiving end, the encoded bitstream is corrected for errors to restore the feature bitstream. For example, if the error correction coding module at the transmitting end uses a concatenated error correction coding method, such as concatenating RS codes and convolutional codes, that is, performing RS encoding → interleaving → convolutional encoding on the feature bitstream to output the encoded bitstream; then the error correction module at the receiving end performs convolutional decoding → deinterleaving → RS decoding on the encoded bitstream to recover the feature bitstream.

[0031] In another embodiment, the extreme image and video decoding and decompression module decodes the feature bitstream of the decoding receiver at the decoding receiver based on the same deep learning network extreme compression method to generate the original image and video information.

[0032] In this embodiment, the extreme image and video decoding and decompression module is the reverse process of the extreme image and video encoding and compression. At the decoding receiver, the feature code stream (bit stream represented by 0 and 1) of the decoding receiver is decoded based on the same deep learning network extreme compression method to generate the original image and video information.

[0033] In another embodiment, the audio modulator and demodulator can be implemented purely in software, i.e., in a Software Modem manner, such as by calling a minimodem.

[0034] In another embodiment, the wired or wireless transmission of information includes transmission in the form of narrowband audio, feature streams, coded streams, and digital audio.

[0035] In this embodiment, the data transmission module transmits the information (narrowband audio / feature stream / encoded stream / digital audio) output by the encoding transmitter to the decoding receiver over short or long distances. This can be done via wired or wireless means. Image and video information is transmitted in narrowband audio mode, meaning the narrowband audio output by the audio player at the encoding transmitter is transmitted to the audio recorder input at the decoding receiver via the data transmission module; image and video information is transmitted in feature stream mode, meaning the feature stream output by the extreme image and video encoding compression module at the encoding transmitter is transmitted to the extreme image and video decoding and decompression module input at the decoding receiver via the data transmission module; image and video information is transmitted in encoded stream mode, meaning the encoded stream output by the error correction encoding module at the encoding transmitter is transmitted to the error correction module input at the decoding receiver via the data transmission module; image and video information is transmitted in digital audio mode, meaning the digital audio output by the audio modulator at the encoding transmitter is transmitted to the audio demodulator input at the decoding receiver via the data transmission module.

[0036] In another embodiment, the data transmission module is a shortwave or VHF radio, see [link to relevant documentation]. Figure 2 In one method, the encoding transmitter connects to the radio's audio interface via an audio cable and wirelessly transmits the audio in narrowband mode. The decoding receiver, after receiving the audio wirelessly, outputs it to the audio recorder (microphone) at the decoding receiver via the radio's audio interface for capture. Alternatively, digital information (feature stream / encoded stream / digital audio) can also be transmitted via radio.

[0037] In another embodiment, the data transmission module is a satellite communication system, such as BeiDou satellite, TianTong satellite, or low-orbit satellite, see [link to documentation]. Figure 3In one method, the encoding transmitter wirelessly transmits digital information (feature stream / encoded stream / digital audio) via a smart terminal (BeiDou terminal or satellite phone); after being relayed by satellite, the smart terminal at the decoding receiver wirelessly receives the digital information (feature stream / encoded stream / digital audio) and sends it to the corresponding module at the decoding receiver for image and video reconstruction. Alternatively, analog narrowband audio transmission can also be performed via satellite voice.

[0038] In another embodiment, the data transmission module is WeChat, Lark, or video / voice chat software. That is, the encoding sending end sends digital files (feature stream / encoded stream / digital audio) through WeChat, Lark, or video / voice chat software; the digital files (feature stream / encoded stream / digital audio) received by the decoding receiving end through WeChat, Lark, or video / voice chat software are sent to the corresponding module of the decoding receiving end for image and video reconstruction.

[0039] In another embodiment, Figures 8(a) to 8(c) show the effect of extreme compression of an image by an extreme image and video transmission device based on narrowband audio provided in one embodiment of this disclosure. Figure 8(a) is the input image before extreme compression (resolution 256*256, bmp format, size 192KB), Figure 8(b) is the feature bitstream (hexadecimal bitstream, size 1020 bytes) of the input image after extreme encoding compression, and Figure 8(c) is the output image after decoding and decompression (resolution 256*256, bmp format, size 192KB); the image is converted into a feature bitstream after extreme compression, and its compression ratio is approximately 188 times.

[0040] In another embodiment, an electronic device includes: a memory, a processor, and module programs stored in the memory and executable on the processor for encoding a transmitter and decoding a receiver, wherein the processor executes the module programs.

[0041] Furthermore, the technical effects produced by the key technical means of the present invention are summarized as follows: Achieving compression ratios of tens to thousands of times, the compressed data size is only a few thousandths to a few tens of percent of the original image and video. This compresses image and video data that was originally impossible to transmit over narrowband audio (300Hz~3400Hz, with a rate of only tens to thousands of bps) to a transmittable level, fundamentally solving the bandwidth shortage problem. Although extremely high compression ratios may lead to a decrease in image quality, the deep learning-based encoding method can preserve key visual features with minimal data volume, ensuring that the receiving end can decode and reconstruct recognizable image or video content.

[0042] In narrowband audio channels (such as shortwave, satellite, and telephone lines) susceptible to noise, attenuation, and distortion, adding redundant information effectively detects and corrects bit errors during transmission. Especially in harsh communication environments such as emergency situations, military applications, and underwater missions, it ensures the integrity of the feature or coded bitstream, improving the success rate of image and video reconstruction. It supports skipping the error correction coding / correction module when channel quality is good, improving transmission efficiency; and enabling a strong error correction mechanism when channel quality is poor, ensuring reliability.

[0043] Compatible with existing voice communication systems: It can transmit images and videos using any device that supports audio transmission, such as shortwave / ultra-shortwave radios, satellite phones (e.g., Beidou, Tiantong), WeChat voice messages, video conferencing, walkie-talkies, and traditional telephones, without requiring a dedicated high-speed data link. It transmits via speaker playback and microphone recording, suitable for all devices with audio interfaces. It can directly send feature streams or digital audio files via software such as WeChat and Lark, improving efficiency. With satellite or radio support, it directly transmits encoded streams, avoiding audio conversion loss. It can be implemented in software (e.g., Software Modem) on smartphones, tablets, PCs, and dedicated communication terminals, requiring no additional hardware.

[0044] In the 5G / 6G era, it can still efficiently utilize limited bandwidth to transmit more multimedia information and improve spectrum utilization. The compressed data size is small, the transmission time is short, reducing the working time of communication modules and extending the battery life of mobile devices (such as satellite phones and emergency terminals). Although the encoding is complex, the decoding end only needs to run a trained deep learning model for feature reconstruction, making it suitable for deployment on terminal devices with limited computing resources.

[0045] It supports various topologies such as broadcast image distribution (e.g., emergency command), point-to-point communication, and network communication. It can run as a software module on general-purpose devices such as smartphones, tablets, and PCs, eliminating the need for customized hardware and reducing deployment costs. It provides third-party developers with image transmission APIs or SDKs based on audio channels, promoting innovative applications in military, emergency response, industrial, and IoT fields.

[0046] When public networks are down, satellite phones or shortwave radios can be used to transmit on-site images, improving command efficiency. In secure or low-probability-of-interception communications, narrowband channels can be used to covertly transmit critical image intelligence. Acoustic channels have extremely low bandwidth, enabling image information transmission between underwater devices. Low-cost visual communication can be achieved in bandwidth-constrained areas. Ordinary mobile phones can send images via satellite voice channels, achieving a "voice-to-image" function.

[0047] Furthermore, we can further analyze and infer some technical effects: Even if the error correction module fails to completely correct all errors, the deep learning decoder at the receiving end can still "infer" reasonable image content based on the incomplete feature bitstream, achieving a visual restoration similar to "filling in the blanks" and significantly improving the user experience. The system performance no longer depends entirely on the upper limit of the error correction coding's error correction capability, but is guaranteed by a dual mechanism of "error correction + semantic recovery," forming a hybrid error-tolerant system, which is impossible to achieve with traditional image coding.

[0048] In sensitive scenarios (such as emergency rescue, military reconnaissance, and communications in restricted areas), image information can be transmitted using seemingly normal "call noise" or "background noise" to avoid revealing the intent of data transmission. This circumvents communication protocol reviews: On platforms like WeChat voice and video conferencing, the system typically allows voice streams but may restrict file or data streams. Encoding images as "playable audio" can circumvent these restrictions, enabling efficient transmission outside of compliance regulations.

[0049] The receiving end does not require pre-installation of specific decoding software; it only needs general AI inference capabilities to reconstruct images, greatly improving system deployment flexibility. In the future, image quality can be improved or new types of images (such as infrared and X-rays) can be adapted by updating the decoding model without changing the encoding logic of the transmitting end, forming an ecosystem of "encode once, adapt to multiple ends".

[0050] Audio emitted by shortwave radio can be recorded and decoded by smartphones, enabling seamless image interaction between military equipment and civilian terminals. During a single voice call, both human voice and image / audio packets can be transmitted, achieving hybrid "voice + image" broadcasting and increasing information density.

[0051] In scenarios such as disaster warnings and public safety announcements, "image and audio streams" can be broadcast via radio stations and emergency broadcasting systems. The public can record these streams using their mobile phones to obtain key information such as site maps and evacuation route maps. The images can be stored on USB flash drives or SD cards, and the audio files can be manually transmitted and decoded in environments without a network connection, making it suitable for extremely isolated environments.

[0052] Although embodiments of the present invention have been described above in conjunction with the accompanying drawings, the present invention is not limited to the specific embodiments and application fields described above. The specific embodiments described above are merely illustrative and instructive, and not restrictive. Those skilled in the art can make many other forms based on the guidance of this specification and without departing from the scope of protection of the claims of the present invention, and all of these are within the scope of protection of the present invention.

Claims

1. An image and video transmission device based on narrowband audio, comprising: Encoding transmitter, decoding receiver, and data transmission module. The encoding and transmitting end includes an extreme image and video encoding and compression module, an error correction encoding module, an audio modulator, and an audio player. The decoding receiver includes an audio recorder, an audio demodulator, an error correction module, and an extreme image and video decoding and decompression module. The encoding and transmitting end converts the image and video into narrowband audio output after extreme encoding compression and modulation. The data transmission module transmits information via wired or wireless means; The decoding receiver demodulates, corrects, and decodes the input narrowband audio, then restores it to an image or video output.

2. The apparatus according to claim 1, preferably, wherein the extreme image and video encoding compression module uses an extreme compression method based on deep learning networks to convert image and video information into a feature bitstream.

3. The apparatus according to claim 1, wherein the error correction coding module converts the feature code stream into an encoded code stream through error correction coding, for detecting and correcting bit errors generated during data transmission.

4. The apparatus according to claim 1, wherein the audio modulator modulates the encoded bitstream into a digital audio signal in a digital signal manner.

5. The apparatus according to claim 1, wherein the audio recorder acquires and captures the analog narrowband audio signal transmitted by the data transmission module through an audio interface at the decoding receiver, and converts the narrowband audio signal into a digital audio signal through pulse code modulation (PCM).

6. The apparatus according to claim 1, wherein the audio demodulator demodulates the digital audio signal at the decoding receiver, accurately identifies the original frequency, phase or amplitude changes, and converts them back into the encoded bitstream.

7. The apparatus according to claim 1, wherein the error correction module restores the characteristic bitstream from the encoded bitstream at the decoding receiver by performing error correction.

8. The apparatus according to claim 1, wherein the extreme image and video decoding and decompression module decodes the feature bitstream of the decoding receiver at the decoding receiver based on the same deep learning network extreme compression method to generate the original image and video information.

9. The apparatus according to claim 1, wherein the wired or wireless information transmission includes transmission in the form of narrowband audio, feature stream, coded stream, and digital audio.

10. An electronic device, comprising: The memory, processor, and the various module programs stored in the memory and capable of running on the processor for the encoding transmitter and decoding receiver, wherein, The processor executes the module programs as described in any one of claims 1 to 9.