Method, system, and medium for data transmission

The data transmission system facilitates bidirectional audio and video communication between monitoring centers and locations by modulating and coupling signals, addressing the limitations of existing devices to enhance intercom capabilities without additional wiring.

WO2025179935A1PCT designated stage Publication Date: 2025-09-04ZHEJIANG DAHUA TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/129007
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-27
Filing Date
2024-10-31
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing audio and video intercom devices, particularly analog cameras, lack the capability for bidirectional communication, as they can only transmit video data from the camera to a digital video recorder (DVR) but not audio data from the DVR to the camera, limiting their functionality in scenarios requiring two-way intercom functionality.

Method used

A data transmission system that includes a first and second terminal device, a transmission line, and computing devices with processing and storage components to modulate and couple audio and video signals for bidirectional communication, enabling audio transmission from the DVR to the camera without additional wiring.

Benefits of technology

Enables secure, two-way audio intercom functionality between monitoring centers and locations without the need for extra cabling, enhancing system functionality and user experience in scenarios like community and bank monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129007_04092025_PF_FP_ABST
    Figure CN2024129007_04092025_PF_FP_ABST
Patent Text Reader

Abstract

A method, a system, and medium for data transmission may be provided. The method may be performed by a data transmission system including a first computing device. The first computing device may include a first processing device and a first storage device being integrated in a first terminal device. The method may include obtaining, by the first processing device, an audio signal acquired by the first terminal device; obtaining, by the first processing device, a video signal from by a second terminal device; determining, by the first processing device, a modulated audio signal by modulating the audio signal; determining, by the first processing device, audio and video mixing data by coupling the modulated audio signal and the video signal; and sending, by the first processing device, the audio and video mixing data to the second terminal device.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD, SYSTEM, AND MEDIUM FOR DATA TRANSMISSION

[0001] CROSS-REFERENCE RELATED TO APPLICATIONS

[0002] CROSS-REFERENCE

[0003] The present disclosure claims priority to Chinese Patent Application No. 202410217497.6, filed on February 27, 2024, the entire contents of which are incorporated herein by reference.TECHNICAL FIELD

[0004] The present disclosure relates to the field of communications, and in particular, to a method, system, and medium for data transmission.BACKGROUND

[0005] With the continuous development of communication technology, people's function requirement of audio and video devices is also improving. In more and more scenarios (for example, community monitoring, bank monitoring, school monitoring, etc. ) , audio and video devices need to have intercom function. An existing audio and video intercom device may usually realize the audio transmission from the camera to the digital video recorder (DVR) , but may not transmit audio data from the DVR to the camera. The two-way intercom function cannot be realized between the DVR to the camera. Thus the existing audio and video intercom device may be difficult to meet the needs of certain scenarios.

[0006] Therefore, it is desired to provide a method, a system, and medium for data transmission to realize the function of bidirectional audio intercom, which is applicable to more scenarios and enhances user experience.SUMMARY

[0007] One or more embodiments of the present disclosure provide a data transmission method. The method for data transmission may be implemented on a system for data transmission. The system may include a first computing device including a first processing device and a first storage device. The first computing device may be integrated in a first terminal device. The method may include obtaining, by the first processing device, an audio signal captured by the first terminal device and obtaining, by the first processing device, a video signal from by a second terminal device. The method may also include determining, by the first processing device, a modulated audio signal by modulating the audio signal, and determining, by the first processing device, audio and video mixing data by coupling the modulated audio signal and the video signal. The method may further include sending, by the first processing device, the audio and video mixing data to the second terminal device.

[0008] One embodiment of the present disclosure provides a data transmission system. The system may include a first terminal device, a second terminal device, and a first computing device. The first computing device may include a first processing device and a first storage device. The first computing device may be integrated in the first terminal device. The first processing device is configured to: obtain an audio signal  captured by the first terminal device, and obtain a video signal from by the second terminal device, and determine a modulated audio signal by modulating the audio signal, and determine audio and video mixing data by coupling the modulated audio signal and the video signal, and send the audio and video mixing data to the second terminal device.

[0009] One or more embodiments of the present disclosure provide a non-transitory computer readable medium. The medium may include executable instructions. When the executable instructions are executed by at least one processor, the at least one processor may be directed to perform the method for data transmission. The method may include obtaining, by the first processing device, an audio signal captured by the first terminal device, and obtaining, by the first processing device, a video signal from by a second terminal device. The method may also include determining, by the first processing device, a modulated audio signal by modulating the audio signal and determining, by the first processing device, audio and video mixing data by coupling the modulated audio signal and the video signal. The method may further include sending, by the first processing device, the audio and video mixing data to the second terminal device.

[0010] One or more embodiments of the present disclosure provide a data transmission device. The device may include a first acquisition module configured to acquire an audio signal captured by a first terminal device. The device may also include a second acquisition module configured to acquire a second terminal device capturing a video signal. The device may further include a modulation module configured to modulate the audio signal to determine the modulated audio signal. The device may further include coupling module configured to couple the modulated audio signal. The device may further include a video signal to determine the audio and video mixing data. The device may further include a data transmission module configured to transmit the audio and video mixing data to the second terminal device.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The present disclosure will be further illustrated by way of exemplary embodiments, which will be described in detail by means of the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbering denotes the same structure, wherein:

[0012] FIG. 1 is a schematic diagram illustrating a data transmission system according to some embodiments of the present disclosure;

[0013] FIG. 2 is a schematic diagram illustrating exemplary hardware and / or software components of an exemplary computing device according to some embodiments of the present disclosure;

[0014] FIG. 3 is a block diagram illustrating an exemplary processing device according to some embodiments of the present disclosure;

[0015] FIG. 4 is an exemplary schematic diagram of a method of transmitting data shown according to some embodiments of the present disclosure;

[0016] FIG. 5 is an exemplary schematic diagram for determining the audio and video mixing data as shown according to some embodiments of the present disclosure;

[0017] FIG. 6 is an exemplary flowchart for determining a modulated audio signal as shown according to some embodiments of the present disclosure;

[0018] FIG. 7 is an exemplary schematic diagram for determining the frequency of an audio modulation carrier as shown according to some embodiments of the present disclosure.DETAILED DESCRIPTION

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings to be used in the description of the embodiments will be briefly described below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present disclosure, and that the present disclosure may be applied to other similar scenarios in accordance with these drawings without creative labor for those of ordinary skill in the art. Unless obviously acquired from the context or the context illustrates otherwise, the same numeral in the drawings refers to the same structure or operation.

[0020] It should be understood that “system, ” “device, ” “unit, ” and / or “module” as used herein is a way to distinguish between different components, elements, parts, sections, or assemblies at different levels. However, these words may be replaced by other expressions if they accomplish the same purpose.

[0021] As indicated in the present disclosure and in the claims, the singular forms “a, ” “an, ” and “the” may be intended to include the plural forms as well, unless the context clearly indicates otherwise. In general, the terms “comprise, ” “comprises, ” and / or “comprising, ” “include, ” “includes, ” and / or “including, ” when used in this disclosure, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0022] Flowcharts are used in the present disclosure to illustrate the operations performed by the system according to some embodiments of the present disclosure. It should be understood that the operations described herein are not necessarily executed in a specific order. Instead, they may be executed in reverse order or simultaneously. Additionally, one or more other operations may be added to these processes, or one or more operations may be removed from these processes.

[0023] A camera may include an analog camera, a network camera, etc. Compared to a network camera, an analog camera, because they do not rely on a network connection, may be more secure in scenarios where a physical location needs to be protected or cyberattacks avoided. Existing analog cameras may be usually only capable of transmitting data (e.g., a video) in a forward direction from the camera to a DVR, but not capable of transmitting data (e.g., an audio) in a reverse direction from the DVR to the camera which means that the existing analog cameras may not realize the two-way intercom function, making it difficult to meet certain scenarios. For example, as shown in FIG. 4, there is only a forward video transmission function from the second terminal device to the first terminal device, but not a reverse audio transmission function from the first terminal device to the second terminal device. In order to realize the bidirectional audio transmission, the bidirectional audio intercom function may be realized by means of adding an independent intercom transmission line. But the additional independent intercom transmission lines may cause a problem such as short signal transmission distance and susceptibility to interference.

[0024] Some embodiments of the present disclosure therefore propose a method, system, and medium for data transmission to securely realize the function of two-way audio intercom without the need for additional wiring.

[0025] FIG. 1 is a schematic diagram illustrating a data transmission system 100 according to some embodiments of the present disclosure. In some embodiments, the data transmission system 100 may include  a first terminal device 110, a second terminal device 120, a transmission line 130, a first computing device 140, and a second computing device 150.

[0026] The first terminal device 110 is a device that captures an audio signal. For example, the first terminal device 110 may include a DVR. In some embodiments, an audio input device, e.g., a microphone, a recorder, etc., may be integrated on the first terminal device 110 for receiving the audio signal. In some embodiments, a user may input an audio signal via a user terminal and transmit the audio signal to the first terminal device 110 via a network. In some embodiments, the first terminal device 110 may be located in a monitoring center.

[0027] The second terminal device 120 refers to a device that captures a video signal. For example, the second terminal device 120 may include an analog video camera. In some embodiments, an audio output device, e.g., speakers, headphones, a sound system, etc., may be integrated on the second terminal device 120 to play an audio obtained from or transmitted by the first terminal device 120. The second terminal device 120 may be not connected to a network, but rather is connected to the first terminal device 110 via a transmission line 130. In some embodiments, the second terminal device 120 may be located at a monitoring point location.

[0028] The transmission line 130 refers to a line for transmitting audio data and / or video data. For example, the transmission line 130 may include a coaxial cable, a fiber optic cable, a copper wire, etc. In some embodiments, the first terminal device 110 and the second terminal device 120 may be disposed on sides of the transmission line 130, respectively. In other words, an end of the transmission lines 130 may be connected with the first terminal device 110 and another end of the transmission line 130 may be connected with the second terminal device 120. The first terminal device 110 may be configured to transmit, via the transmission line 130, an audio signal to the second terminal device 120, and the second terminal device 120 may be configured to transmit, via the transmission line 130, a video signal to the first terminal device 110.

[0029] The first computing device 140 refers to an electronic device that is capable of performing mathematical calculations or logical operations. In some embodiments, the first computing device 140 may include a first processing device 141 and a first storage device 142. In some embodiments, the first computing device 140 may be integrated with the first terminal device 110. In some embodiments, the first computing device 140 may be separately from the first terminal device 110. The first computing device 140 may be in communication with the first terminal device 110 via a wireless connection or a wired connection.

[0030] The first processing device 141 may be used to manage the data resources of the first terminal device 110, as well as processing data and / or information from at least one component of the first terminal device 110 or an external data source. For example, the first processing device 141 may obtain an audio signal via the first terminal device 110 and a video signal via the second terminal device 120. The first processing device 141 may modulate the audio signal to determine the modulated audio signal and couple the modulated audio signal and the video signal to determine the audio and video mixing data. The first processing device 141 may send the audio and video mixing data to the second terminal device 120.

[0031] In some embodiments, the first processing device 141 may include a first acquisition module 310, a second acquisition module 320, a modulation module 330, a coupling module 340, and a data transmission module 350. In some embodiments, the first processing device 141 may include the first acquisition module 310, the second acquisition module 320, the modulation module 330, the coupling module 340, the data  transmission module 350, and at least one of a selection module, a speech recognition module, or an anomaly identification module.

[0032] In some embodiments, the first processing device 141 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the first processing device 141 may be local or remote. In some embodiments, the first processing device 141 may be implemented on a cloud platform. Merely by way of example, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an inter-cloud, a multi-cloud, or the like, or any combination thereof. In some embodiments, the first processing device 141 may be implemented by a computing device 200 having one or more components illustrated in FIG. 2.

[0033] The first storage device 142 may be used to store data, instructions, and / or any other information on the first terminal device 110. For example, the first storage device 142 may be used to store audio signal and video signal. As another example, the first storage device 142 may be used to store instructions for processing the audio signal and / or the video signal for data transmission. As a further example, the first storage device 142 may be used to store a set of instructions for performing a method for data transmission, e.g., from the first terminal device to the second terminal device as described elsewhere in the present disclose.

[0034] In some embodiments, the first storage device 142 may include a mass storage, removable storage, a volatile read-and-write memory, a read-only memory (ROM) , or the like, or any combination thereof. In some embodiments, the first storage device 142 may be implemented on a cloud platform. In some embodiments, the first storage device 142 may integrated into the processing device 140 and / or the terminal device.

[0035] In some embodiments, the first storage device 142 may be connected to a network to communicate with one or more other components (e.g., the first processing device 141, etc. ) of the first terminal device 110.

[0036] The second computing device 150 refers to an electronic device that is capable of performing mathematical calculations or logical operations. In some embodiments, the second computing device 150 may include a second processing device and a second storage device 152. In some embodiments, the second computing device 150 may be integrated with the second terminal device 120. In some embodiments, the second computing device 150 may be separately from the second terminal device 120. The second computing device 15 may be in communication with the second terminal device 120 via a wireless connection or a wired connection.

[0037] The second processing device may be used to manage the data resources of the second terminal device 120, as well as processing data and / or information from at least one component involved in the second terminal device 120 and / or external data source. For example, the second processing device may obtain audio and video mixing data from the first terminal device 110. The second processing device may decouple the audio and video mixing data to determine an intermediate audio signal, and demodulate the intermediate audio signal to determine a target audio signal.

[0038] In some embodiments, the second processing device may include a third acquisition module 360, a decoupling module 370, and a demodulation module 380.

[0039] The second storage device 152 may be used to store data, instructions, and / or any other information on the second terminal device 120. For example, the second storage device 152 may be used to store audio and video mixing data. As another example, the second storage device 152 may also be used to store instructions for processing the audio and video mixing data.

[0040] The type of the second processing device is similar to the type of the first processing device 141, and the type of the second storage device 152 is similar to the type of the first storage device 142, and will not be repeated here.

[0041] In some embodiments, the data transmission system 100 may include some or more other devices, e.g., a network, a user terminal, etc.

[0042] The network may facilitate the exchange of information and / or data. The network enables communication between components, and with other parts outside of the system, thereby facilitating the exchange of data and / or information.

[0043] In some embodiments, the network may be or include a public network (e.g., the Internet) , a private network (e.g., a local area network (LAN) ) , a wired network, a wireless network (e.g., an 802.11 network, a Wi-Fi network) , a frame relay network, a virtual private network (VPN) , a satellite network, a telephone network, routers, hubs, switches, server computers, and / or any combination thereof. For example, the network may include a cable network, a wireline network, a fiber-optic network, a telecommunications network, an intranet, a wireless local area network (WLAN) , a metropolitan area network (MAN) , a public telephone switched network (PSTN) , a BluetoothTM network, a ZigBeeTM network, a near field communication (NFC) network, or the like, or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include wired and / or wireless network access points such as base stations and / or internet exchange points through which one or more components of data transmission system100 may be connected to the network to exchange data and / or information.

[0044] In some embodiments, a user may interact with one or more components of the data transmission system 100 (e.g., the first terminal device 110, the second terminal device 120, the first computing device 140 and the second computing device 150, etc. ) via the user terminal. For example, the user may transmit the control signal to the first terminal device 110 via the user terminal.

[0045] In some embodiments, the user terminal may include an embedded device with a relatively small storage capacity. In some embodiments, the user terminal may include a smart phone, a smart camera, a smart audio, a smart TV, a smart fridge, a robot, a tablet, a laptop, a wearable, a payment device, a cashier device, or the like, or any combination thereof.

[0046] In some embodiments, the first terminal device 110 may be connected to the network, and the second terminal device 120 may be connected to the first terminal device 110 via the transmission line 130. The user may log in to the first terminal device 110 via an application (e.g., an app, applet, software, etc. ) on the user terminal. In the software operation, the user may turn on the intercom function. After completing the above operations, the establishment of the intercom link between the application program of the user terminal and the second terminal device 120 is completed, and two-way intercom may be conducted. After the intercom link has been established, the data transmission system 100 may implement the corresponding data transmission method in the embodiment of the present disclosure.

[0047] It should be noted that the above description regarding data transmission system 100 is merely provided for the purposes of illustration, and not intended to limit the scope of the present disclosure. For people having ordinary skills in the art, multiple variations and modifications may be made under the teachings of the present disclosure. However, those variations and modifications do not depart from the scope of the present disclosure. In some embodiments, data transmission system 100 may include one or more additional  components and / or one or more components of data transmission system 100 described above may be omitted. Additionally or alternatively, two or more components of data transmission system 100 may be integrated into a single component. component of data transmission system 100 may be implemented on two or more sub-components.

[0048] FIG. 2 is a schematic diagram illustrating exemplary hardware and / or software components of an exemplary computing device according to some embodiments of the present disclosure. In some embodiments, the first computing device 140, the second computing device 150, and / or the terminal device may be implemented on a computing device 200. As illustrated in FIG. 2, the computing device 200 may include a processor 210, a storage device 220, an input / output (I / O) 230, and a communication port 240.

[0049] The processor 210 may execute computer instructions (e.g., program code) and perform functions in accordance with techniques describable herein. The computer instructions may include, for example, routines, programs, objects, components, data structures, procedures, modules, and functions, which perform particular functions describable herein.

[0050] In some embodiments, the processor 210 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC) , an application specific integrated circuits (ASICs) , an application-specific instruction-set processor (ASIP) , a central processing unit (CPU) , a graphics processing unit (GPU) , a physics processing unit (PPU) , a microcontroller unit, a digital signal processor (DSP) , a field programmable gate array (FPGA) , an advanced RISC machine (ARM) , a programmable logic device (PLD) , any circuit or processor capable of executing one or more functions, or the like, or any combinations thereof.

[0051] Merely for illustration, only one processor is describable in the computing device 200. However, it should be noted that the computing device 200 in the present disclosure may also include multiple processors, thus operations and / or method operations that are performed by one processor as describable in the present disclosure may also be jointly or separately performed by the multiple processors. For example, if in the present disclosure the processor of the computing device 200 executes both operation A and operation B, it should be understood that operation A and operation B may also be performed by two or more different processors jointly or separately in the first computing device 140 (e.g., a first processor executes operation A and a second processor executes operation B, or the first and second processors jointly execute operations A and B) .

[0052] The storage device 220 may store data obtained from one or more components of the data transmission system 100. In some embodiments, the storage device 220 may include a mass storage device, a removable storage device, a volatile read-and-write memory, a read-only memory (ROM) , or the like, or any combination thereof. In some embodiments, the storage device 220 may store one or more programs and / or instructions to perform exemplary methods describable in the present disclosure. For example, the storage device 220 may store a program for the processing device 210 to execute to compress the machine learning model.

[0053] The I / O 230 may input and / or output signal, data, information, etc. In some embodiments, the I / O 230 may enable a user interaction with the processor 210. In some embodiments, the I / O 230 may include an input device and an output device. The input device may include a keyboard, a touch screen, a speech input, an eye tracking input, a brain monitoring system, or any other comparable input mechanism. The input  information received through the input device may be transmitted to another component (e.g., the processor 210) via, for example, a bus, for further processing. Other types of the input devices may include a cursor control device, such as a mouse, a trackball, or cursor direction keys, etc. The output device may include a display (e.g., a liquid crystal display (LCD) , a light-emitting diode (LED) -based display, a flat panel display, a curved screen, a television device, a cathode ray tube (CRT) , a touch screen, a speaker, a printer, or the like, or a combination thereof.

[0054] The communication port 240may be connected to a network to facilitate data communications. The communication port 240 may establish connections between the processor 210 and the terminal device. The connection may be a wired connection, a wireless connection, any other communication connection that may enable data transmission and / or reception, and / or any combination of these connections. The wired connection may include, for example, an electrical cable, an optical cable, a telephone wire, or the like, or any combination thereof. The wireless connection may include, for example, a BluetoothTM link, a Wi-FiTM link, a WiMaxTM link, a WLAN link, a ZigBeeTM link, a mobile network link (e.g., 3G, 4G, 5G) , or the like, or a combination thereof. In some embodiments, the communication port 240 may be and / or include a standardized communication port, such as RS232, RS485, etc. In some embodiments, the communication port 240 may be a specially designed communication port.

[0055] FIG. 3 is a block diagram illustrating an exemplary processing device according to some embodiments of the present disclosure. In some embodiments, the first processing device 140 may include a first acquisition module 310, a second acquisition module 320, a modulation module 330, a coupling module 340 and a data transmission module 350.

[0056] In some embodiments, the first acquisition module 310 may be configured to acquire an audio signal acquired by the first terminal device.

[0057] In some embodiments, the second acquisition module 320 may be configured to obtain a video signal acquired by the second terminal device.

[0058] In some embodiments, the first terminal device and the second terminal device may be provided on both sides of a transmission line 130 respectively. The first terminal device may be configured to transmit an audio signal to the second terminal device and the second terminal device may be configured to transmit a video signal to the first terminal device.

[0059] In some embodiments, the first terminal device may be located at a monitoring center and the second terminal device may be located at a monitoring point location.

[0060] In some embodiments, the modulation module 330 may be configured to modulate the audio signal to determine the modulated audio signal.

[0061] In some embodiments, the modulation module 330 may be further configured to determine, based on the audio signal and the video signal, a bandwidth of the audio signal and a bandwidth of the video signal; determine, based on the bandwidth of the audio signal and the bandwidth of the video signal, an audio modulation carrier; and determining a modulated audio signal based on the audio signal and the audio modulation carrier.

[0062] In some embodiments, the modulation module 330 may be further configured to determine an activation energy threshold based on the ambient noise and the equipment noise; and, in response to the signal energy of the audio signal not being less than the activation energy threshold, modulate the audio signal to  determine the modulated audio signal.

[0063] In some embodiments, the modulation module 330 may be further configured to determine an activation energy threshold based on the signal energy of the ambient noise and the signal energy of the equipment noise.

[0064] In some embodiments, the coupling module 340 may be configured to couple the modulated audio signal and the video signal to determine the audio and video mixing data.

[0065] In some embodiments, the coupling module 340 may be further configured to filter the modulated audio signal to determine a first modulated audio signal; and to couple the first modulated audio signal and the video signal to determine a mix of audio-video data.

[0066] In some embodiments, the coupling module 340 may be further configured to couple the control signal, the first modulated audio signal, and the video signal to determine the audio and video mixing data.

[0067] In some embodiments, the data transmitting module 350 may be configured to send the audio and video mixing data to a second terminal device.

[0068] In some embodiments, the first processing device 140 may also include a selection module, a speech recognition module, and an anomaly identification module (not shown in the figures) .

[0069] In some embodiments, the number of second terminal devices is a plurality of second terminal devices, and the selection module may be configured to control whether the first terminal device transmits an audio signal to any of the plurality of second terminal devices.

[0070] In some embodiments, the speech recognition module may be configured to perform speech recognition on the audio signal and the video signal, via the speech recognition module, to determine a speech recognition result; and in response to the speech recognition result satisfying a first condition, to trigger an emergency treatment program.

[0071] In some embodiments, the anomaly identification module may be configured to, by the anomaly identification module, perform an anomaly identification of the video signal to determine an anomaly identification result, the anomaly identification result comprising at least one of a voice anomaly result and an image anomaly result; and in response to the anomaly identification result satisfying the second condition, the first terminal device transmits an emergency audio signal to the second terminal device.

[0072] In some embodiments, the second processing device 140 may include a third acquisition module 360, a decoupling module 370, and a demodulation module 380.

[0073] In some embodiments, the third acquisition module 360 may be configured to acquire the audio and video mixing data from the first terminal device.

[0074] In some embodiments, the decoupling module 370 may be configured to decouple the audio and video mixing data to determine an intermediate audio signal, the intermediate audio signal being a decoupled modulated audio signal.

[0075] In some embodiments, the demodulation module 380 may be configured to demodulate the intermediate audio signal to determine a target audio signal, the target audio signal being the demodulated audio signal.

[0076] With respect to the first acquisition module 310, the second acquisition module 320, the modulation module 330, the coupling module 340, the data transmission module 350, the third acquisition module 360, the decoupling module 370, the demodulation module 380, the selection module, speech recognition module, and  the anomaly identification module are described in more detail in FIG. 4 -FIG. 6 and their related descriptions.

[0077] It should be noted that the above descriptions of processing device 140 are provided for the purposes of illustration, and not intended to limit the scope of the present disclosure. For people having ordinary skills in the art, various modifications and changes in the forms and details of the application of the above method and system may occur without departing from the principles of the present disclosure. In some embodiments, processing device 140 may include one or more other modules and / or one or more modules described above may be omitted. Additionally or alternatively, two or more modules may be integrated into a single module and / or a module may be divided into two or more units. However, those variations and modifications also fall within the scope of the present disclosure.

[0078] FIG. 4 is an exemplary schematic diagram of a process of data transmission shown according to some embodiments of the present disclosure. In some embodiments, process 400 may be executed by the data transmission system 100. For example, process 400 may be implemented as a set of instructions stored in a storage device. In some embodiments, the first processing device 140 (e.g., the processor 210 of the first computing device 140 and / or one or more modules illustrated in FIG. 3) and the second processing device 150 (e.g., the second processor of the second computing device 150 and / or one or more modules illustrated in FIG. 3) may execute the set of instructions and may accordingly be directed to perform the process 400. The operations of the illustrated process presented below are intended to be illustrative. In some embodiments, the process 400 may be accomplished with one or more additional operations not described and / or without one or more of the operations discussed. Additionally, the order of the operations of process 400 illustrated in FIG. 4 and described below is not intended to be limiting.

[0079] As shown in FIG. 4, the thin solid line and the thin dashed line represent a forward video transmission direction from the second terminal device to the first terminal device, and the thick solid line and the thick dashed line represent the reverse audio transmission direction from the first terminal device to the second terminal device.

[0080] In 410, a first processing device may obtain the audio signal acquired by a first terminal device. In some embodiments, operation 410 may be performed by the first acquisition module 310.

[0081] The first terminal device is a device for obtaining an audio signal. In some embodiments, the first terminal device may include a device such as a DVR. In some embodiments, the first terminal device may include other devices, such as, a video cassette recorder (VCR) , a hard disk recorder (HDR) , or the like, or a combination thereof. More descriptions about the first terminal device may be found in FIG. 1 and its related description.

[0082] The audio signal is sound data acquired by the first terminal device. In some embodiments, the first processing device may obtain the audio signal in a variety of ways through the first terminal device.

[0083] For example, the first terminal device may have an internally integrated audio input device (e.g., a microphone, etc. ) to acquire the audio signal. The first processing device may obtain the audio signal from the audio input device or the first storage device.

[0084] In other embodiments, the first terminal device may be communicatively connected to an external audio input device for acquiring the audio signal. The first processing device may obtain the audio signal acquired by the external audio input device (e.g., a walkie-talkie, a cell phone, etc. ) via a network (e.g., WLAN, Bluetooth, etc. ) or wired transmission, etc.

[0085] In some embodiments, the first processing device may obtain the audio signal through the first terminal device based on any other feasible means. For example, a user may log in to the first terminal device via an application of a user terminal, the user may record audio to generate the audio signal via the application and transmit the audio signal to the first terminal device via a network.

[0086] In 420, the first processing device may obtain a video signal acquired by a second terminal device. In some embodiments, operation 420 may be performed by the second acquisition module 320.

[0087] The second terminal device is a device for acquiring a video signal. In some embodiments, the second terminal device may include an analog camera.

[0088] In some embodiments, the second terminal device may be connected, via a transmission line 130, to the first terminal device. More about the second terminal device may be found in FIG. 1 and its related description.

[0089] In some embodiments, the first processing device may obtain the video signal through the second terminal device. For example, the first processing device may directly acquire the video signal uploaded by the second terminal device.

[0090] In some embodiments, the first terminal device 110 and the second terminal device 120 may be disposed on sides of the transmission line 130, respectively. The first terminal device 110 may be configured to transmit, via the transmission line 130, an audio signal to the second terminal device 120, and the second terminal device 120 may be configured to transmit, via the transmission line 130, a video signal to the first terminal device 110. More about transmission line 130 may be found in FIG. 1 and its related descriptions.

[0091] In some embodiments, the first terminal device may be located in a monitoring center. The second terminal device may be located at a monitoring point location.

[0092] The monitoring center is a center for centralized analysis and management of data collected from one or more monitoring point locations. The monitoring center may collect and analyze the video signal from each monitoring point location, and / or carry out the operations of video monitoring, data processing, and / or emergency handling.

[0093] A monitoring point location refers to an actual location where site information is collected. In other words, the monitoring point location may be a region or a position where needs monitoring. Monitoring data (e.g., images, videos, and / or audios) of the monitoring point location may be acquired may be transmitted, in the form of video signal, to the monitoring center.

[0094] For example, in the application scenario of housing estate monitoring, the monitoring point locations may be distributed in a passageway, an entrance and exit, a location inside an elevator, a garage, and other locations in the housing estate that need to be monitored. The monitoring center may be located in the property center of the housing estate. When a monitoring personnel located in the monitoring center monitors the dangerous factors (e.g., the elevator may not open the door, etc. ) in the elevator through a monitoring system the monitoring personnel may transmit the audio signal to the second terminal device through the first terminal device and receive a video signal transmitted by the second terminal device through the first terminal device, so as to realize the bidirectional intercommunication with the trapped personnel in the elevator to calm the emotions and obtain relevant information about the trapped personnel in time to facilitate the subsequent rescue operations. The data transmission system may be a portion of the monitoring system.

[0095] In some embodiments of the present disclosure, by placing the first terminal device and the second  terminal device at the monitoring center and the monitoring point location, respectively, a bidirectional intercom function between the monitoring center and the monitoring point location may be realized.

[0096] In 430, the first processing device may modulate the audio signal to determine a modulated audio signal. In some embodiments, operation 430 may be performed by the modulation module 330.

[0097] The modulated audio signal is an audio signal that is obtained by a modulation operation. In some embodiments, the modulated audio signal is narrower than the original audio signal frequency band, the modulated power utilization rate is high, and the video signal and the modulated audio signal are separate from each other in the frequency domain, without mutual interference. The video signal and modulated audio signal are separated from each other in the frequency domain, as shown in FIG. 7, the video signal and regulated audio signal are located in different frequency ranges, and there is no overlap between the two.

[0098] The modulation operation may be configured to convert an audio signal into a modulated audio signal suitable for transmission over a channel. In some embodiments, the first processing device may embed the audio signal into an audio modulation carrier to obtain a modulated audio signal. The audio modulation carrier may be configured to carry information related to the audio signal. In some embodiments, the modulation operation may include using a frequency modulation (FM) manner, or the like.

[0099] In some embodiments, since the manifestation of external interference is mostly parasitic amplitude modulation, one or more embodiments of the present disclosure may determine a frequency modulation manner that may eliminate parasitic amplitude modulation based on limiting amplitude, and the modulated audio signal, compared to the original audio signal, has a narrow frequency band, high modulation power utilization, and more suitable for long-distance transmission. On the other hand, after the audio signal is modulated, the video signal and the modulated audio signal are separated from each other in the frequency domain and do not interfere with each other, so that in the subsequent transmission of the audio and video mixing data, it may be ensured that the signal quality is not compromised.

[0100] In some embodiments, the first processing device may determine the modulated audio signal based on the audio signal and the audio modulation carrier. More descriptions for determining the modulated audio signal may be found elsewhere in the present disclosure (e.g., FIG. 5 and the descriptions thereof) .

[0101] In 440, the first processing device may couple the modulated audio signal and the video signal to determine the audio and video mixing data. In some embodiments, operation 440 may be performed by the coupling module 340.

[0102] The audio and video mixing data is data obtained by coupling a modulated audio signal with the video signal. In some embodiments, the audio and video mixing data may be obtained by a coupling operation.

[0103] The coupling operation refers to combining the modulated audio signal and the video signal, into the audio and video mixing data through multiplexing. The multiplexing refers to the technology of merging multiple signals into single transmission channel. In some embodiments, the multiplexing may include frequency division multiplexing (FDM) , time division multiplexing (TDM) , or the like.

[0104] The frequency division multiplexing is the transmission by assigning different signals to different frequency bands. The time division multiplexing is the transmission of different signals in different time periods, and each signal occupies the transmission channel in the allocated time period. In some embodiments, the first processing device may combine the modulated audio signal and the video signal and  send the audio and video mixing data to the second terminal device over the transmission line 130. It is understandable that if the audio signal is not modulated, the audio signal and the video signal may overlap in the frequency domain, the audio signal and the video signal interfere with each other, so the direct transmission of the audio signal and the video signal will cause damage to the signal quality. However, after the audio signal is modulated, the video signal and the modulated audio signal may be separated from each other in the frequency domain and may be independent of each other. The video signal and the modulated audio signal may do not interfere with each other, which ensures that the quality of the signal is not impaired.

[0105] In some embodiments, the first processing device may filter the modulated audio signal to determine the first modulated audio signal, and the first modulated audio signal may be coupled with the video signal to obtain the audio and video mixing data. More descriptions for determining the first modulated audio signal and the audio and video mixing data may be found elsewhere in the present disclosure (e.g., FIG. 5 and the descriptions thereof) .

[0106] In 450, the first processing device may send the audio and video mixing data to the second terminal device. In some embodiments, operation 450 may be performed by data transmission module 350.

[0107] In some embodiments, the first processing device may send the audio and video mixing data to the second terminal device via the transmission line 130. More descriptions for the transmission line 130 may be found elsewhere in the present disclosure (e.g., the FIG. 3 and the descriptions thereof) .

[0108] In some embodiments of the present disclosure, by using the data transmission method as described above, a bidirectional audio transmission function between a monitoring center and a monitoring point location may be realized without arranging additional lines. Thereby, the functional upgrading of the original monitoring system may be realized without adding additional expenses for wiring arrangement.

[0109] In some embodiments, the process 400 may further include operations 460-480.

[0110] In 460, the second processing device may obtain the audio and video mixing data from the first terminal device. In some embodiments, operation 460 may be performed by a third acquisition module 360.

[0111] In 470, the second processing device may decouple the audio and video mixing data to determine an intermediate audio signal. In some embodiments, operation 470 may be performed by the decoupling module 370.

[0112] The intermediate audio signal is the modulated audio signal obtained by decoupling the audio and video mixing data. It is understood that since the video signal and the modulated audio signal are separate from each other in the frequency domain, there is no interference, the process of coupling and uncoupling does not affect the quality of the video signal and the modulated audio signal, and the intermediate audio signal is the same as the modulated audio signal determined in step 430.

[0113] Decoupling may be used to separate signals in the audio and video mixing data. Decoupling is the inverse process of coupling. In some embodiments, the second processing device may separate the audio and video mixing data (i.e., the modulated audio signal and video signal coupled together) by using a filter to extract the modulated audio signal and the video signal so that they may be processed or transmitted independently. It is understood that since the filter separation refers to the process of processing signals by blocking or allowing signals of a specific frequency range. The video signal and the modulated audio signal are separated from each other in the frequency domain, so that the modulated audio signal can pass through the filter, thus achieving the purpose of separating the modulated audio signal and the video signal.

[0114] In 480, the second processing device may demodulate the intermediate audio signal to determine a target audio signal. The target audio signal may be a demodulated audio signal. In some embodiments, operation 480 may be performed by the demodulation module 380.

[0115] The demodulation is the operational step of restoring a modulated audio signal to the original audio signal. The purpose of demodulation is to convert the modulated signal to the original audio signal at the second terminal device so that the audio signal may be played back. The demodulation is the inverse process of modulation. In some embodiments, the second processing device may recover the audio signal from the modulated audio signal and send the audio signal to the second terminal device to play the obtained audio signal via the second terminal device.

[0116] In some embodiments, when there is a pre-existing transmission line 130 between the monitoring center and the monitoring point location, the data transmission system may transmit the audio and video mixing data based on the pre-existing transmission line 130.

[0117] In some embodiments, via a transmission line 130, the first processing device may send audio and video mixing data using frequency division multiplexing, thereby synchronizing and efficiently transmitting audio signal and video signal based on a single transmission line 130. Wherein, frequency division multiplexing refers to a method of transmitting data in which the available transmission frequency of a channel is split into a plurality of non-overlapping sub-transmission frequencies, each of which is used to transmit a portion of the signal.

[0118] In some embodiments, over a transmission line 130, the first processing device may also send the audio and video mixing data using time division multiplexing. For example, the first processing device may divide the transmission time to obtain a number of time segments of very short duration, and by alternating forward transmissions as well as reverse transmissions within the different time segments, the overall audio signal and video signal transmission of the two-way transmission effect. The forward transmission may refer to the transmission of the video signal from the second terminal device to the first terminal device, and the reverse transmission may refer to the transmission of the audio signal from the first terminal device to the second terminal device.

[0119] In some embodiments of the present disclosure, by connecting the first terminal device and the second terminal device using an existing transmission line 130, fast transmission of audio signal as well as video signal may be realized; and when the first terminal device and the second terminal device may be connected directly to each other using an already laid transmission line 130 without having to lay an additional transmission line 130 specifically for audio bidirectional transmission. And when the first terminal device and the second terminal device are connected to each other, the existing transmission line 130 may be used directly, and there is no need to additionally lay a transmission line 130 specialized for audio bi-directional transmission. This allows for the reuse of the original lines and reduces costs.

[0120] In some embodiments of the present disclosure, by using a combination of coupling as well as decoupling, modulation as well as demodulation, it is possible to perform a data extraction of the audio and video mixing data after the audio and video mixing data transmission is completed, and to obtain the original audio signal and video signal therefrom. High-fidelity transmission of audio signal and video signal in an actual scene may be achieved while ensuring that the original audio signal and video signal remain unchanged.

[0121] It should be noted that the above description of the process 400 is merely provided for the purposes  of illustration, and not intended to limit the scope of the present disclosure. For persons having ordinary skills in the art, multiple variations or modifications may be made under the teachings of the present disclosure. However, those variations and modifications do not depart from the scope of the present disclosure.

[0122] In some embodiments, operations 430 and 440 may be replaced by modulating the video signal by the first processing device to determine the modulated video signal, and coupling the audio signal and the modulated video signal by the first processing device to determine the audio and video mixing data.

[0123] In other embodiments, operations 430 and 440 may be replaced by modulating the audio signal and the video signal by the first processing device to determine the modulated audio signal and the modulated video signal, and coupling the modulated audio signal and the modulated video signal to determine the audio and video mixing data.

[0124] The first processing device may determine the modulated audio signal by embedding the video signal into a video modulation carrier. The process of coupling the audio signal to the modulated video signal, the process of coupling a modulated audio signal to a modulated video signal is similar to that of coupling a modulated audio signal to a video signal, and will not be described herein.

[0125] In some embodiments, the number of second terminal devices may exceed 1, and the first processing device may control whether to transmit the audio signal to any one of the second terminal devices. In response to determining transmitting the audio signal to one of the second terminal devices, the first processing device may transmit the audio signal to the one of the second terminal devices according to a portion of process 400 (e.g., operations 410-450) .

[0126] In some embodiments, the first terminal device may include at least one switch, or the second terminal device may include at least one switch, e.g., a physical switch, a virtual switch in an application, etc. The user may control whether or not to transmit the audio signal to any one or more of the plurality of second terminal devices by switching an operating state of the switch (e.g., connected or disconnected) .

[0127] In some embodiments, separate switches may be provided on a plurality of second terminal devices, or a plurality of switches may correspond to the second terminal may be provided on the first terminal device, to enable transmission of an audio signal from the first terminal device to any one or more of the second terminal devices. As another example, at least one switch may be set on the first terminal device, the first processing device may group the second terminal devices in advance and control at least one second terminal device within a group by a switch, thereby realizing functions such as area broadcasting (e.g., emergency broadcasting in case of an emergency event, range broadcasting for information notification, etc. ) . For example, the first processing device may flexibly switch between the above two switching configuration methods according to actual needs.

[0128] In some embodiments of the present disclosure, by setting up a plurality of second terminal devices, the second terminal devices can simultaneously collect video signals at a plurality of monitoring point locations, thereby realizing simultaneous monitoring of a plurality of locations. By setting a switch for controlling the transmission of audio signal or not, the second terminal device which one the audio signal needs to be transmitted to may be conveniently and quickly selected, thereby realizing the monitoring center's simultaneous voice intercom with one or more monitoring point locations.

[0129] In some embodiments, the first processing device may perform speech recognition on each of the audio signal and the video signal to determine speech recognition results corresponding to the audio signal and  the video signal; and in response to determining that one of the speech recognition results satisfies a first condition, the first processing device may trigger an emergency treatment program.

[0130] The speech recognition refers to a recognition operation for obtaining textual information from data. In some embodiments, the first processing device may perform speech recognition in multiple ways. For example, the first processing device may extract an audio portion from the video signal, perform speech recognition on each of the audio signal and the audio portion from the video signal, and thereby obtain a corresponding speech recognition result. The method of speech recognition may include, but is not limited to, a sound template matching technique, a deep learning model, or the like.

[0131] The speech recognition result is the recognized text from the video signal as well as the audio signal.

[0132] In some embodiments, when the speech recognition is completed, the first processing device may analyze the speech recognition results corresponding to the video signal and the audio signal to determine whether the speech recognition results satisfy the first condition using a text analysis algorithm. The text analysis algorithm may include a word frequency statistics (e.g., TF-IDF, etc. ) algorithm, a semantic analysis algorithm (e.g., Word2Vec, BERT, etc. ) , or the like, or a combination thereof.

[0133] The first condition is a judgment condition for determining whether or not the emergency treatment program needs to be triggered. In some embodiments, the first condition may be determined based on actual application scenarios and needs. For example, the first condition may be that a preset sensitive text is present in the speech recognition result. For example, the text “Danger” , “Fire” , “Help” .

[0134] The emergency treatment program is a treatment program responding to an emergency contingency. In some embodiments, in response to the speech recognition result satisfying the first condition, the first processing device may trigger the emergency treatment program.

[0135] In some embodiments, the emergency treatment program may be determined based on actual application scenarios and needs. For example, the first processing device may group the sensitive text according to the type, e.g., grouping the sensitive text such as “Fire” , “Fire-fighting” , “Burn” , etc., into group 1, and grouping the sensitive text such as “Fainting” , “Morbidity” , “First aid” , etc., into group 2. After the grouping is completed, for a group of sensitive text, the first processing device may set a corresponding emergency treatment program. For example, the emergency treatment program for group 1 may be to call a fire alarm, the emergency treatment program for group 2 may be to call a hospital emergency number, etc. Thereby, the emergency treatment programs applicable to different emergency contingencies are obtained.

[0136] In some embodiments of the present disclosure, by performing speech recognition on a video signal as well as an audio signal, and analyzing and processing the speech recognition results, and activating a corresponding emergency treatment program when a first condition is met, the data transmission system may be enhanced in its scope of application, and enhance the emergency handling capability of the data transmission system when facing abnormal situations.

[0137] In some embodiments, the first processing device may perform an anomaly identification of the video signal to determine an anomaly identification result; and in response to determining that the anomaly identification result satisfies a second condition, the first terminal device may transmit an emergency audio signal to the second terminal device.

[0138] The anomaly identification may include identifying one or more anomalous conditions in a video signal. In some embodiments, the anomaly identification may include at least one of voice anomaly  identification or image anomaly identification.

[0139] In some embodiments, by the anomaly identification, the first processing device may determine the anomaly identification result. The anomaly identification result may include at least one of a voice anomaly identification result or an image anomaly identification result. The voice anomaly identification result may indicate whether the sound data in the video signal involves an anomaly; and the image anomaly identification result may indicate whether the image data in the video signal involves g an anomaly. When the anomaly identification does not determine that the anomaly involves in the video signal, the voice anomaly identification result as well as the image anomaly identification result may indicate that the sound data and the image data in the video signal do not involve an anomaly.

[0140] The voice anomaly identification may include performing anomaly identification on sound data contained in a video signal. In some embodiments, the voice anomaly identification may include emotion recognition. The first processing device may sense the type of emotion of a speaker (e.g., excited, calm, panicked, etc. ) through an emotion recognition method. The emotion recognition method may include, but is not limited to, a rhyme characterization algorithm, a spectral characterization algorithm, or the like, or a combination thereof.

[0141] In some embodiments, the first processing device may identify the type of negative emotion present in the speaker as a voice anomaly result. The negative emotion type may include emotions such as panic, fear, annoyance, or the like, that occur in the speaker. The specific negative emotion types may be determined based on the actual application scenarios and requirements.

[0142] The image anomaly identification may include performing the anomaly identification on images contained in a video signal.

[0143] In some embodiments, the first processing device may perform image anomaly identification in multiple ways. For example, the first processing device may perform background recognition, character recognition, or the like, on the image. As a further example, the background recognition may include using a frame difference algorithm, a deep learning algorithm, or the like; and the character recognition may include using a human body key point detection algorithm, a face recognition algorithm, or the like.

[0144] In some embodiments, the first processing device may determine the image anomaly result by determining an anomaly in the image anomaly identification. The anomaly may include, but are not limited to, a significant change in the background, the presence of objects moving too fast in the image, the presence of a person having an abnormal demeanor, or the like. Specific anomalies may be determined based on actual application scenarios and requirements.

[0145] The second condition refers to a judgment condition for determining whether to transmit an emergency audio signal to the second terminal device. The second condition is more similar to the first condition, and both are used for judging whether or not the emergency situation at the monitoring point location needs to be handled. The difference is that the first condition is used to determine whether the speech recognition result is abnormal, and the second condition is used to determine whether the voice and / or image is abnormal.

[0146] In some embodiments, the second condition may be determined based on actual application scenarios and needs. For example, the second condition may be that the voice anomaly result indicates an anomaly involving and the image anomaly result indicates no anomaly involving, or the voice anomaly result  indicates no anomaly involving and the image anomaly result indicates an anomaly involving. As another example, the second condition may also be that each of the voice anomaly result and the image anomaly result indicates an anomaly involving.

[0147] In some embodiments, in response to determining that the anomaly identification result satisfies the second condition, the first terminal device may transmit the emergency audio signal to the second terminal device.

[0148] In some embodiments, the first processing device may construct the emergency audio signal corresponding to the different anomaly identification results according to the manner of constructing the emergency treatment program corresponding to the sensitive text as described above. When the anomaly identification result satisfies the second condition, the first processing device may transmit the emergency audio signal corresponding to the current anomaly identification result to the second terminal device for playback via the transmission line 130. For example, when the image anomaly identification detects the presence of a suspicious person in the image, the first processing device may send a voice inquiry message to communicate with the suspicious person, so as to be informed of information such as the identity and purpose of the suspicious person.

[0149] In some embodiments, when only one of the voice anomaly identification result and the image anomaly identification result indicates an anomaly involving, or when the anomaly types of the voice anomaly identification result and the image anomaly identification result are the same, one of the voice anomaly identification result and the image anomaly identification result, or the shared anomaly identification result, may be identified as a target anomaly identification result.

[0150] In some embodiments, when each of the voice anomaly identification result and the image anomaly identification result indicates that an anomaly involving, and the types of anomalies of the voice anomaly identification result and the image anomaly identification result are different, the target anomaly identification result may be confirmed by the monitoring personnel of the monitoring center. As another example, the monitoring center may carry out an intercom communication with the relevant personnel located at the monitoring point location, and based on the response of the relevant personnel, determine the target anomaly identification result for subsequent transmission and playback of the emergency audio signal.

[0151] In some embodiments, when the current operation of the suspicious person is found to satisfy ta third condition, the first processing device may execute a corresponding emergency treatment program. The third condition refers to the fact that the relevant operation of the suspicious person involves a significant risk or hidden danger, which needs to be immediately interrupted or terminated, or the like. For example, when the suspicious person is found to have committed an unlawful act, the first processing device may trigger an operation such as an automatic alarm.

[0152] In some embodiments, the third condition may be determined based on actual application scenarios and needs. For example, in the application scenario of hospital monitoring, the third condition may be that a patient is found to have an abnormal situation such as fainting and convulsions, the corresponding emergency treatment program may include notifying a nearby medical personnel, etc.

[0153] In some embodiments, the first processing device may send a voice query via a monitor located at the monitoring center. In some embodiments, the first processing device may issue the voice query information via an artificial intelligence customer service.

[0154] In some embodiments, when the language used by the monitoring personnel of the monitoring center and the speaker of the monitoring point location are not consistent, the first processing device may perform a translation function (e.g., voice real-time translation, text translation, etc. ) to ensure that the monitoring personnel of the monitoring center and the speaker of the monitoring point location may communicate smoothly so as enhancing the communication efficiency.

[0155] In some embodiments of the present disclosure, by performing the abnormality identification of the video signal, the monitoring system can promptly sense whether an abnormality occurs in the video signal. When the anomaly identification result satisfies the second condition, the data transmission system may also transmit an emergency audio signal to the second terminal device, which helps to resolve the abnormality in a timely manner, and effectively avoids the potential abnormality not being resolved in a timely manner due to the potential abnormality effectively avoiding serious consequences due to untimely resolution of potential anomalies.

[0156] FIG. 5 is an exemplary schematic diagram for determining the audio and video mixing data as shown according to some embodiments of the present disclosure.

[0157] In some embodiments, the first processing device may filter a modulated audio signal to determine the first modulated audio signal and couple the first modulated audio signal and a video signal to determine an audio and video mixing data.

[0158] The filtering of the modulated audio signal may include adjusting the frequency components of the modulated audio signal. In some embodiments, the first processing device may allow a signal of a particular frequency range to pass through the filtering operation while attenuating or eliminating a signal of other frequencies.

[0159] In some embodiments, the first processing device may filter the modulated audio signal. In some embodiments, the first processing device may filter both the modulated audio signal and the video signal. For example, the first processing device may filter the modulated audio signal via a first filtering module, and the first processing device may filter the video signal via a second filtering module. The first filtering module may block the video signal and allow only the modulated audio signal to pass through, and the second filtering module may allow the video signal to pass through and block the modulated audio signal. The modulated audio signal obtained after filtering by the first filter module is the first modulated audio signal.

[0160] In some embodiments, the first processing device may superimpose the first modulated audio signal and the video signal, thereby accomplishing coupling of the filtered modulated audio signal and the video signal. The first processing device may accomplish the above coupling operation by means of a coupling module 340. Further description of the coupling module 340 may be found in the FIG. 2 related descriptions.

[0161] In some embodiments of the present disclosure, by using the first filtering module and the second filtering module to filter the modulated audio signal and the video signal, respectively, it is possible to ensure that the audio signal and video signal do not interfere with each other in this way, which helps to enhance the signal quality of the audio signal and the video signal.

[0162] In some embodiments, the first processing device may couple the control signal, the first modulated audio signal, and the video signal to determine the audio and video mixing data.

[0163] The control signal is a control instruction for regulating an operation state of the second terminal device. In some embodiments, the control signal may be used to determine a capability level of different  second terminal devices.

[0164] The capability level of a second terminal device measures how many functions the second terminal device has as well as its performance. The more functions the second terminal device supports and the better the performance, the higher the capability level of the second terminal device. In some embodiments, the functions of the second terminal device may include whether the second terminal device supports intercom, whether the second terminal device has a night vision function (e.g., turning on a fill light or switching to an infrared camera) , or the like. The performance of the second terminal device may be indicated by the high or low picture quality of the video signal provided by the second terminal device, the quality of the audio playback when the second terminal device plays the audio signal, or the like, or a combination thereof.

[0165] In other embodiments, the control signal may be used to regulate an operation state of the second terminal device. For example, the control signal may include an instruction for turning on or turn off the intercom function, an instruction for video signal parameter adjustment, an instruction for volume playback parameter adjustment, or the like, or a combination thereof. For example, the first processing device, by sending the instruction for video signal parameter adjustment, may adjust the brightness, contrast, or the like, or a combination thereof of the second terminal device when the second terminal device collects the video signal. For example, by sending the instruction for volume playback parameter adjustment, the first processing device may adjust the volume, sound effect, or the like, or a combination thereof of the second terminal device when playing the audio signal. The sound effect refers to a parameter used to regulate the effect of audio playback. By setting different sound effect parameters, different sound effects may be realized, for example, bass enhancement, treble enhancement, noise reduction and other sound effects.

[0166] In some embodiments, in order to protect the private information of the monitoring personnel in the monitoring center and to reduce the leakage of unnecessary information, the instruction for volume playback parameter adjustment may erase the voice characteristics of the monitoring personnel in the monitoring center by adjusting the sound effects. For example, the first processing device may avoid leaking information such as the gender characteristic of the monitoring personnel at the monitoring center by changing the sound effects (e.g., robotic sound effects) .

[0167] In some embodiments, the first processing device may determine the control signal by a variety of manners. For example, the control signal may be determined based on a first preset table. The first processing device may construct the first preset table based on the actual scene and the control signal. The first preset table contains correspondences between the actual scene and different control signals. In some embodiments, the first processing device may determine the correspondence between the actual scene and the different control signals based on historical data or a priori experience. Exemplarily, the correspondence may be that, when a user located at the monitoring point location reflects that the volume of the second terminal device is too low when talking to the second terminal device, the corresponding control signal in the first preset table may be to increase the playback volume of the second terminal device. For example, when the relevant user at the monitoring center finds that the picture of the video signal is too dark, the control signal may be that the second terminal device equipped with the night vision function may turn on the night vision function, and the second terminal device that does not have the night vision function may increase the brightness when capturing the video signal, and so on.

[0168] In some embodiments, the first processing device may determine the control signal by querying the  first preset table based on the current actual scene.

[0169] In some embodiments, the first processing device may determine the audio and video mixing data via the coupling module 340 based on the control signal, the first modulated audio signal, and the video signal. Further description of the coupling module 340 may be found in the FIG. 2 related description.

[0170] In some embodiments, the control signal may be transmitted via a transmission line 130. The control signal flows in the opposite direction of the forward video transmission and is superimposed directly on the video signal by the first processing device. Because the control signal typically have a low frequency and may not pass through a second filtering module that filters video signal of a higher frequency, the first processing apparatus may further include a control signal extraction module and a control signal superimposition module, the control signal extraction module configured to obtain a control signal from a mixed signal of the video signal and the control signal, and the control signal superimposition module configured to efficiently superimpose the control signal in the audio and video mixing data. Thereby, the control signal, the first modulated audio signal, and the video signal, are jointly determined to be audio and video mixing data by the coupling module 340.

[0171] Further, before acquiring the control signal, the first processing device may also compare the control signal with a preset control signal energy threshold, and extract the control signal only when the signal energy threshold of the control signal is greater than the preset control signal energy threshold.

[0172] Through the above steps, the embodiment of the present disclosure may ensure that the control signal is effectively superimposed on the audio and video mixing data, so that the control signal and the audio and video mixing data may be transmitted at the same time, to avoid the absence of the control signal, and to improve the signal quality and integrity; The first processing device may extract the control signal only when the signal energy threshold of the control signal is greater than a preset control signal energy threshold, which may avoid the first processing device from identifying irrelevant signal or noise or the like in the audio and video mixing data as the control signal, and may effectively avoid the occurrence of situations such as false alarms of the control signal.

[0173] FIG. 6 is an exemplary flowchart illustrating a process for determining a modulated audio signal as shown according to some embodiments of the present disclosure.

[0174] In some embodiments, process 600 may be executed by the data transmission system 100. For example, the process 600 may be implemented as a set of instructions stored in a storage device. In some embodiments, the first processing device 140 (e.g., the processor 210 of the first computing device 140 and / or one or more modules illustrated in FIG. 3) may execute the set of instructions and may accordingly be directed to perform the process 600. The operations of the illustrated process presented below are intended to be illustrative. In some embodiments, the process 600 may be accomplished with one or more additional operations not described and / or without one or more of the operations discussed. Additionally, the order of the operations of process 600 illustrated in FIG. 6 and described below is not intended to be limiting. In some embodiments, operations 610-640 may be performed by the modulation module 330.

[0175] In 610, determining an activation energy threshold based on an ambient noise and an equipment noise.

[0176] The ambient noise may be noise generated in the environment where the first terminal device is located. For example, the ambient noise may include the sound of a crowd clamoring in the environment, the  sound of a vehicle traveling, the sound of wind, or the like.

[0177] The equipment noise may be noise generated by the first terminal device. For example, the equipment noise may include the background noise when the first terminal device is operating.

[0178] The activation energy threshold is a judgment threshold for determining whether the first processing device is required to acquire and send the audio signal.

[0179] In some embodiments, the activation energy threshold may be determined in multiple ways. For example, the first processing device may establish a second preset table based on the historical ambient noise, the historical equipment noise, and the historical activation energy thresholds. The second preset table includes correspondences between the historical ambient noise, the historical equipment noise, and different historical activation energy thresholds. The first processing device may determine the current activation energy threshold based on the current ambient noise and equipment noise by consulting the second preset table. A historical activation energy threshold may be a manual input by a monitoring personnel or a system default activation energy threshold at a historical moment.

[0180] In some embodiments, the first processing device may determine the activation energy threshold based on the signal energy of the ambient noise and the signal energy of the equipment noise.

[0181] The signal energy may reflect the volume level of ambient noise or equipment noise. In some embodiments, the signal energy may be represented in multiple ways. For example, the first processing device may determine the volume statistic (e.g., mean, median, etc. ) of the ambient noise as the signal energy of the ambient noise. The first processing device may determine the volume statistic (e.g., mean, median, etc. ) of the equipment noise as the signal energy of the equipment noise.

[0182] In some embodiments, the activation energy threshold may be positively correlated to the signal energy of the ambient noise and the signal energy of the equipment noise. For example, the first processing device may collect the ambient noise and the equipment noise, statistically obtain the average volume value of the ambient noise and the average volume value of the equipment noise, and designate the sum of the average volume value of the ambient noise and the average volume value of the equipment noise as the activation energy threshold. In some embodiments, the first processing device may determine weights corresponding to the signal energy of the ambient noise and the signal energy of the equipment noise and determine the activation energy threshold based on the weights corresponding to the signal energy of the ambient noise and the signal energy of the equipment noise.

[0183] According to some embodiments of the present disclosure, the activation energy threshold is determined based on the signal energy of ambient noise and the signal energy of equipment noise. More accurate activation energy threshold results may be obtained, which may help to reduce energy consumption when transmitting invalid audio signal.

[0184] According to some embodiments of the present disclosure, power as well as bandwidth is wasted because the data transmission system may misidentify disturbing factors such as ambient noise and equipment noise, as audio signal to be transmitted. Therefore, it is possible to set the activation energy threshold so that the audio signal transmission is performed only when the monitoring personnel at the monitoring center is speaking, which may help to reduce the transmission of ineffective audio signal, and reduce the energy consumption.

[0185] In 620, in response to a signal energy of the audio signal is not less than the activation energy  threshold, modulating the audio signal to determine the modulated audio signal.

[0186] Bandwidth may refer to the frequency range available in the transmission of a video signal or an audio signal.

[0187] In some embodiments, the first processing device may determine the bandwidth of the video signal in various ways. For example, the first processing device may determine the bandwidth of the video signal based on the resolution, the frame rate, and / or the color depth of the video signal.

[0188] The resolution may reflect how many pixels are included in the video signal, which may be expressed as the product of the length and width. The larger the resolution, the richer the image detail included in the video signal. The frame rate of the video signal reflects the number of images included in the video signal per second, and the larger the frame rate is, the more coherent the video signal. The color depth may reflect the color detail of the video signal. The greater the color depth, the more detailed the color of the video signal. But the larger the resolution, the higher the frame rate, and the larger the color depth, the larger the amount of data contained in the video signal per unit of time, and the more bandwidth will be used to transmit the video signal.

[0189] In some embodiments, the bandwidth of the video signal may be positively correlated to the number of pixels, the frame rate, and the color depth of the video signal.

[0190] For example, the bandwidth of the video signal may be obtained based on the following equation (1) :

[0191] T=k1×A×B×C×D                   (1)

[0192] wherein, T is the bandwidth of the video signal, k1 is a coefficient, A is the resolution width, B is the resolution height, C is the frame rate, and D is the color depth. The coefficient k1 may be preset based on a priori experience. In some embodiments, when the video signal is compressed using a video compression algorithm (e.g., H. 264, H. 265, or the like) , the coefficient k1 may be negatively correlated to the compression capability of the video compression algorithm. The stronger the compression ability of video compression algorithm, the smaller the coefficient k1, the compression capability of video compression algorithm may be characterized by compression ratio, compression efficiency, bit rate, etc. The compression capability of the video compression algorithm may be obtained based on actual testing or by querying relevant information such as the description document of the video compression algorithm.

[0193] In some embodiments, the first processing device may determine the bandwidth of the audio signal in various ways. For example, the first processing device may determine the bandwidth of the audio signal based on the sampling frequency of the audio signal as well as the sampling accuracy. The sampling frequency of the audio signal is similar to the frame rate of the video signal, and the higher the sampling frequency is, the more coherent and smooth the audio signal is. The sampling accuracy is similar to the color depth of a video signal, the higher the sampling accuracy, the higher the audio signal reproduction. Similar to a video signal, the higher the sampling frequency and sampling precision, the more bandwidth will be used in the transmission of the audio signal.

[0194] In some embodiments, the bandwidth of the audio signal may be positively correlated to the sampling frequency and sampling accuracy of the audio signal. The first processing device may take a similar approach to determining the bandwidth of the audio signal by referring to equation (1) above.

[0195] In 630, an audio modulation carrier may be determined based on the bandwidth of the audio signal and the bandwidth of the video signal.

[0196] The audio modulation carrier is a carrier used to carry an audio signal. In some embodiments, an audio signal with a lower frequency may be modulated into an audio modulation carrier with a higher frequency in order to reduce attenuation of the audio signal over long distances and ensure that the audio signals may be effectively transmitted. Moreover, in order to improve the separation of the audio and video mixing data during the decoupling and to reduce the performance requirements for the decoupling module 370, the video signal and the modulated audio signal may be separated as much as possible in the frequency domain and be independent of each other by determining a suitable audio modulation carrier frequency.

[0197] In some embodiments, the audio modulation carrier may be represented by an audio modulation carrier frequency.

[0198] In some embodiments, since the modulated audio signal need to be coupled to the video signal in the subsequent operation 440, there is a need to minimize interference between the modulated audio signal and the video signal.

[0199] In some embodiments, as shown in FIG. 7, the abscissa is the frequency, the ordinate is the amplitude, the audio modulation carrier frequency may satisfy the following equation (2) :

[0200] wherein, fAC is the frequency of the audio modulation carrier, fV_BW is the bandwidth of a video signal, and fA_BW is the bandwidth of an audio signal. The determined audio modulation carrier needs to have a frequency higher than the sum of the bandwidth of the video signal and half of the bandwidth of the audio signal. Since the first processing device beds the audio signal into the audio modulation carrier to obtain the modulated audio signal, it may be seen from FIG. 7 that the audio modulation carrier determined by equation (2) may separate the modulated audio signal from the video signal in the frequency domain, and there is no overlap between the modulated audio signal and the video signal.

[0201] In 640, determining the modulated audio signal based on the audio signal and the audio modulation carrier.

[0202] In some embodiments, the first processing device may embed the audio signal into the audio modulation carrier to obtain the modulated audio signal. More on determining the modulated audio signal may be found in FIG. 2 and its related sections.

[0203] According to some embodiments of the present disclosure, by selecting a suitable audio modulation carrier frequency, it may be ensured that the video signal and the modulated audio signal are separated as much as possible from each other in the frequency domain. Thereby, the modulated audio signal are transmitted together with the video signal without affecting the original video signal, thereby realizing clear and high-quality video calls.

[0204] It should be noted that the above description of the process 600 is merely provided for the purposes of illustration, and not intended to limit the scope of the present disclosure. For people having ordinary skills in the art, multiple variations or modifications may be made under the teachings of the present disclosure. However, those variations and modifications do not depart from the scope of the present disclosure.

[0205] For example, operations 610-620 may be deleted.

[0206] Embodiments of the present disclosure also provide a data transmission system including a first terminal device 110, a second terminal device 120, and a first computing device 140. The first computing  device 140 includes a first processing device 141 and a first storage device, and the first computing device 140 is integrated into the first terminal device 110, and the first processing device 141 is configured to: acquire an audio signal acquired by the first terminal device 110, and acquire a video signal acquired by the second terminal device 120, and modulate the audio signal to determine a modulated audio signal, and couple the modulated audio signal and the video signal to determine the audio and video mixing data, and send the audio and video mixing data to the second terminal device 120.

[0207] In some embodiments, the first processing device 141 is further configured to: determine a bandwidth of the audio signal and a bandwidth of the video signal based on the audio signal and the video signal, and determine an audio modulation carrier based on the bandwidth of the audio signal and the bandwidth of the video signal, and based on the audio signal and the audio modulation carrier, determine a modulated audio signal.

[0208] In some embodiments, the first processing device 141 is further configured to: filter the modulated audio signal to determine the first modulated audio signal, and couple the first modulated audio signal and the video signal to determine the audio and video mixing data.

[0209] In some embodiments, the first processing device 141 is further configured to: couple the control signal, the first modulated audio signal, and the video signal to determine the audio and video mixing data.

[0210] In some embodiments, the first processing device 141 is further configured to: determine an activation energy threshold based on the ambient noise and the equipment noise, and in response to the signal energy of the audio signal not being less than the activation energy threshold, modulate the audio signal to determine the modulated audio signal.

[0211] In some embodiments, the first processing device 141 is further configured to: determine an activation energy threshold based on the signal energy of the ambient noise and the signal energy of the equipment noise.

[0212] In some embodiments, the first terminal device 110 and the second terminal device 120 are provided on both sides of the transmission line 130, the first terminal device 110 being configured to transmit the audio signal to the second terminal device 120, and the second terminal device 120 being configured to transmit the video signal to the first terminal device 110.

[0213] In some embodiments, the first terminal device 110 is located at a monitoring center and the second terminal device 120 is located at a monitoring point location.

[0214] In some embodiments, the number of second terminal devices is a plurality of second terminal devices, and the first processing device 141 is further configured to: control whether or not the first terminal device 110 transmits the audio signal to any one of the plurality of second terminal devices.

[0215] In some embodiments, the first processing device 141 is further configured to: perform textual recognition of the audio signal and the video signal to determine a textual recognition result, and in response to the textual recognition result satisfying the first condition, trigger an emergency treatment program.

[0216] In some embodiments, the first processing device 141 is further configured to: perform an anomaly identification of the video signal to determine an anomaly identification result, the anomaly identification result comprising at least one of a voice anomaly result and an image anomaly result, and in response to the anomaly identification result satisfying the second condition, the first terminal device 110 transmits an emergency audio signal to the second terminal device 120.

[0217] In some embodiments, the data transmission system further comprises a second computing device 150; the second computing device 150 comprises a second processing device 151 and a second storage device 152, the second computing device 150 being integrated into the second terminal device 120, the second processing device 151 being configured to: obtain the audio and video mixing data from the first terminal device 110; decouple the audio and video mixing data to determine an intermediate audio signal, the intermediate audio signal being a modulated audio signal after decoupling, and demodulate the intermediate audio signal to determine a target audio signal, the target audio signal being the demodulated audio signal.

[0218] One or more embodiments of the present disclosure provides a computer-readable storage medium that stores computer instructions, and when the computer reads the computer instructions in the storage medium, the computer performs a method of data transmission according to some embodiments of the present disclosure.

[0219] Beneficial effects that may be brought about by the embodiments of the present disclosure include, but are not limited to: (1) without arranging additional wiring, by using the existing transmission line 130 located between the monitoring center and the monitoring point location, a bidirectional audio transmission function between the monitoring center and the monitoring point location may be realized, so as to realize the functional upgrading of the original monitoring system. (2) By determining a suitable audio modulation carrier, it can be ensured that the video signal and the modulated audio signal are as separate and independent of each other as possible in the frequency domain. Thus in the case of not affecting the original video signal, the modulation of the audio signal and the video signal together with the audio and video mixing data transmission form, thus ensuring high signal quality, transmission robustness, transmission distance and not easy to be disturbed. This enables clear, high-quality two-way audio transmission. (3) Through the use of the first filtering module and the second filtering module, respectively, the modulation of audio signal and video signal for filtering, which can ensure that the audio and video signal do not interfere with each other, and help to improve the signal quality of audio signal and video signal. (4) Ensuring that the control signal is effectively superimposed on the audio and video mixing data, so that the control signal and the audio and video mixing data can be transmitted at the same time, avoiding the absence of the control signal, and improving the signal quality and completeness; the processing equipment can only extract the control signal when the signal energy threshold of the control signal is greater than the preset control signal energy threshold, so that the second terminal equipment can avoid determining irrelevant signal or noises, or the like in the audio and video mixing data as control signal, and can effectively avoid the occurrence of false alarms or or the like. (5) Waste of power and bandwidth due to the fact that the data transmission system may misinterpret disturbing factors such as ambient noise and equipment noise as audio signal to be transmitted. By setting the activation energy threshold, the audio signal can be transmitted only when the monitoring personnel in the monitoring center are speaking, which helps to reduce the transmission of invalid audio signal and reduces energy consumption. (6) By setting up a plurality of second terminal devices, video signal from a plurality of monitoring point location can be collected simultaneously. Thereby realizing simultaneous monitoring of a plurality of locations; by setting a switch controlling the transmission of audio signal or not, it is possible to conveniently and quickly select a second terminal device that needs to transmit audio signal, thereby realizing simultaneous communication between the monitoring center and one or more monitoring point location to conduct voice intercom simultaneously.

[0220] Having thus described the basic concepts, it may be rather apparent to those skilled in the art after reading this detailed disclosure that the foregoing detailed disclosure is intended to be presented by way of example only and is not limiting. Various alterations, improvements, and modifications may occur and are intended to those skilled in the art, though not expressly stated herein. These alterations, improvements, and modifications are intended to be suggested by this disclosure, and are within the spirit and scope of the exemplary embodiments of this disclosure.

[0221] Moreover, certain terminology has been used to describe embodiments of the present disclosure; For example, the terms “one embodiment, ” “an embodiment, ” and / or “some embodiments” mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure; Therefore, it is emphasized and should be appreciated that two or more references to “an embodiment” or “one embodiment” or “an alternative embodiment” in various portions of this disclosure are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined as suitable in one or more embodiments of the present disclosure.

[0222] Furthermore, the recited order of processing elements or sequences, or the use of numbers, letters, or other designations therefore, is not intended to limit the claimed processes and methods to any order except as may be specified in the claims. Although the above disclosure discusses through various examples what is currently considered to be a variety of useful embodiments of the disclosure, it is to be understood that such detail is solely for that purpose, and that the appended claims are not limited to the disclosed embodiments, but, on the contrary, are intended to cover modifications and equivalent arrangements that are within the spirit and scope of the disclosed embodiments. For example, although the implementation of various components described above may be embodied in a hardware device, it may also be implemented as a software only solution, e.g., an installation on an existing server or mobile device.

[0223] Similarly, it should be appreciated that in the foregoing description of embodiments of the present disclosure, various features are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure aiding in the understanding of one or more of the various inventive embodiments. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed subject matter requires more features than are expressly recited in each claim. Rather, inventive embodiments lie in less than all features of a single foregoing disclosed embodiment.

[0224] In some embodiments, the numbers expressing quantities or properties used to describe and claim certain embodiments of the present disclosure are to be understood as being modified in some instances by the term “about, ” “approximate, ” or “substantially. ” For example, “about, ” “approximate, ” or “substantially” may indicate ±20%variation of the value it describes, unless otherwise stated. Accordingly, in some embodiments, the numerical parameter set forth in the written description and attached claims are approximations that may vary depending upon the desired properties sought to be obtained by a particular embodiment. In some embodiments, the numerical parameter should be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. Notwithstanding that the numerical ranges and parameter setting forth the broad scope of some embodiments of the present disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as practicable.

[0225] Each of the patents, patent applications, publications of patent applications, and other material, such  as articles, books, specifications, publications, documents, things, and / or the like, referenced herein is hereby incorporated herein by this reference in its entirety for all purposes, excepting any prosecution file history associated with same, any of same that is inconsistent with or in conflict with the present document, or any of same that may have a limiting effect as to the broadest scope of the claims now or later associated with the present document. By way of example, should there be any inconsistency or conflict between the description, definition, and / or the use of a term associated with any of the incorporated material and that associated with the present document, the description, definition, and / or the use of the term in the present document shall prevail.

[0226] In closing, it is to be understood that the embodiments of the present disclosure disclosed herein are illustrative of the principles of the embodiments of the present disclosure. Other modifications that may be employed may be within the scope of the present disclosure. Thus, by way of example, but not of limitation, alternative configurations of the embodiments of the present disclosure may be utilized in accordance with the teachings herein. Accordingly, embodiments of the present disclosure are not limited to that precisely as shown and described.

Claims

1.A method for data transmission implemented on a system for data transmission, the system comprising a first computing device, the first computing device comprising a first processing device and a first storage device, the first computing device being integrated in a first terminal device, the method comprising:obtaining, by the first processing device, an audio signal captured by the first terminal device;obtaining, by the first processing device, a video signal from by a second terminal device;determining, by the first processing device, a modulated audio signal by modulating the audio signal;determining, by the first processing device, audio and video mixing data by coupling the modulated audio signal and the video signal; andsending, by the first processing device, the audio and video mixing data to the second terminal device.2.The method of claim 1, wherein the modulating the audio signal comprises:determining a bandwidth of the audio signal and a bandwidth of the video signal;determining an audio modulation carrier based on the bandwidth of the audio signal and the bandwidth of the video signal; anddetermining the modulated audio signal based on the audio signal and the audio modulation carrier.3.The method of claim 1 or claim 2, wherein the coupling the modulated audio signal and the video signal comprises:filtering the modulated audio signal to determine a first modulated audio signal; andcoupling the first modulated audio signal and the video signal to determine the audio and video mixing data.4.The method of claim 3, wherein the coupling the first modulated audio signal and the video signal to determine the audio and video mixing data comprises:coupling a control signal, the first modulated audio signal, and the video signal to determine the audio and video mixing data.5.The method of any one of claims 1 to 4, wherein the determining a modulated audio signal by modulating the audio signal comprises:determining an activation energy threshold based on an ambient noise and an equipment noise; andin response to a signal energy of the audio signal is not less than the activation energy threshold, modulating the audio signal to determine the modulated audio signal.6.The method of claim 5, wherein the determine an activation energy threshold based on an ambient noise and an equipment noise comprises:determining the activation energy threshold based on a signal energy of the ambient noise and a signal energy of the equipment noise.7.The method of any one of claims 1 to 6, wherein the first terminal device and the second terminal device are connected by a transmission line; the first terminal device being configured to transmit the audio signal to the second terminal device; and the second terminal device being configured to transmit the video signal to the first terminal device.8.The method of claim 7, wherein the first terminal device is located at a monitoring center and the second terminal device is located at a monitoring point location.9.The method of any one of claims 1 to 8, wherein a number of second terminal devices exceeds 1, the method further comprises:determining whether the first terminal device transmits the audio signal to one of a plurality of the second terminal devices.10.The method of any one of claims 1 to 9, wherein the method further comprises:performing a speech recognition on the audio signal and the video signal to determine a speech recognition result; andin response to determining that the speech recognition result satisfies a first condition, triggering an emergency treatment program.11.The method of any one of claims 1 to 10, wherein the method further comprises:performing an anomaly identification of the video signal to determine an anomaly identification result, the anomaly identification result including at least one of a voice anomaly result or an image anomaly result; andin response to determining that the anomaly identification result satisfies a second condition, transmitting, by the first terminal device, an emergency audio signal to the second terminal device.12.The method of any one of claims 1 to 11, wherein the system further comprises a second computing device, the second computing device comprising a second processing device and a second storage device, the second computing device being integrated in the second terminal device, the method further comprising:obtaining, by the second processing device, the audio and video mixing data from the first terminal device;determining, by the second processing device, an intermediate audio signal by decoupling the audio and video mixing data, the intermediate audio signal being a decoupled and modulated audio signal; anddetermining, by the second processing device, a target audio signal by demodulating the intermediate audio signal, the target audio signal being a demodulated audio signal.13.A system for data transmission, wherein the system comprises a first terminal device, a second terminal device, and a first computing device; the first computing device comprising a first processing device and a first storage device, the first computing device being integrated into the first terminal device, the first processing device is configured to:obtain an audio signal captured by the first terminal device;obtain a video signal captured by the second terminal device;modulate the audio signal to determine a modulated audio signal;couple the modulated audio signal and the video signal to determine audio and video mixing data; andsend the audio and video mixing data to the second terminal device by the first processing device.14.The system of claim 13, wherein the first processing device is further configured to:determine a bandwidth of the audio signal and a bandwidth of the video signal based on the audio signal and the video signal;determine an audio modulation carrier based on the bandwidth of the audio signal and the bandwidth of the video signal; anddetermine the modulated audio signal based on the audio signal and the audio modulation carrier.15.The system of claim 13 or claim 14, wherein the first processing device is further configured to:filter the modulated audio signal to determine a first modulated audio signal; andcouple the first modulated audio signal and the video signal to determine the audio and video mixing data.16.The system of claim 15, wherein the first processing device is further configured to:couple a control signal, the first modulated audio signal, and the video signal to determine the audio and video mixing data.17.The system of any one of claims 13 -16, wherein the first processing device is further configured to:determine an activation energy threshold based on an ambient noise and an equipment noise; andin response to a signal energy of the audio signal is not less than the activation energy threshold, modulate the audio signal to determine the modulated audio signal.18.The system of claim 17, wherein the first processing device is further configured to:determine the activation energy threshold based on the signal energy of the ambient noise and the signal energy of the equipment noise.19.The system of any one of claims 13 to 18, wherein the first terminal device and the second  terminal device are disposed on either side of a transmission line; the first terminal device being configured to transmit the audio signal to the second terminal device; and the second terminal device being configured to transmit the video signal to the first terminal device.20.The system of claim 19, wherein the first terminal device is located at a monitoring center and the second terminal device is located at a monitoring point location.21.The system of any one of claims 13 to 20, wherein the number of the second terminal devices is a plurality, and the first processing device is further configured to:control whether or not the first terminal device transmits the audio signal to any one of a plurality of the second terminal devices.22.The system of any one of claims 13 to 21, wherein the first processing device is further configured to:perform a speech recognition on the audio signal and the video signal to determine a speech recognition result; andin response to the speech recognition result meeting a first condition, trigger an emergency treatment program.23.The system of any one of claims 13 to 22, wherein the first processing device is further configured to:perform an anomaly identification of the video signal to determine an anomaly identification result, the anomaly identification result including at least one of a voice anomaly result or an image anomaly result; andin response to the anomaly identification result meeting a second condition, the first terminal device transmits an emergency audio signal to the second terminal device.24.The system of any one of claims 13 to 23, wherein the system further comprises a second computing device, the second computing device comprising a second processing device and a second storage device, the second computing device being integrated into the second terminal device, the second processing device is configured to:obtain the audio and video mixing data from the first terminal device;decoupling the audio and video mixing data to determine an intermediate audio signal, the intermediate audio signal being a decoupled modulated audio signal; anddemodulate the intermediate audio signal to determine a target audio signal, the target audio signal being a demodulated audio signal.25.A non-transitory computer-readable storage medium, comprising a set of instructions, wherein  when a computer reads the computer instructions in the storage medium, the method for data transmission of any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Visual inter-conversation device, visual inter-conversation system and visual inter-conversation signal transmission method

    CN102769805A

  • Data transmission method, device and system, electronic device and storage medium

    CN118283320A

  • Intelligent system for monitoring crimes

    CN203206395U

  • Monitoring system

    US20050212920A1