Remote USB device voice communication method based on edge computing and browser collaboration
This method of remote USB device voice communication through edge computing and browser collaboration solves the problems of high deployment cost and poor real-time performance of existing remote voice intercom systems, and realizes a stable real-time voice channel based on a general-purpose browser, which is suitable for scenarios such as industrial monitoring and security guarding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-29
- Publication Date
- 2026-04-28
AI Technical Summary
Existing remote voice intercom systems, when relying on dedicated clients or cloud forwarding, suffer from high deployment costs, poor scalability, large end-to-end latency, and significant jitter, making it difficult to achieve stable audio communication in scenarios with high real-time requirements, such as industrial monitoring and security duty.
By collaborating with edge computing and a browser, a communication service based on the WebSocket protocol is established to realize a bidirectional, controllable, real-time voice channel between the control terminal and the edge USB audio device. This includes negotiation of audio transmission path capabilities and parameter configuration, and access and control are achieved using a general-purpose browser.
By relying solely on a general-purpose browser, a stable real-time voice channel was established between the control end and the edge USB audio device, reducing deployment costs, improving scalability, and reducing network jitter and latency, thus ensuring the real-time nature and controllability of audio communication.
Smart Images

Figure CN121940439A_ABST
Abstract
Description
Technical Field
[0001] This application relates to voice communication technology, and more particularly to a remote USB device voice communication method based on edge computing and browser collaboration. Background Technology
[0002] Currently, remote voice intercom systems can be broadly classified into two categories:
[0003] One type is the traditional analog and IP intercom system that relies on dedicated handsets, intercom hosts or embedded terminals. This type of system generally achieves audio communication between terminals and devices through dedicated cabling or closed local area networks. Although it can provide a relatively stable voice channel in fixed scenarios, it has high deployment costs, poor scalability, and is usually controlled through dedicated clients or local operation panels, which does not support or has difficulty supporting remote access based on general browsers.
[0004] Another type is the cloud intercom solution that uploads the audio stream to the cloud and then distributes it. This type of solution is convenient for centralized management, but the audio link usually needs to go through multiple public networks and centralized forwarding, resulting in large end-to-end latency and obvious jitter. In scenarios with high requirements for real-time performance and controllability, such as industrial monitoring and security duty, problems such as audio stuttering, echo, and frame loss are likely to occur.
[0005] Therefore, we continue with a communication method that can utilize the browser as a universal access point and directly connect to USB microphones and USB speakers at the edge to achieve remote speaking and remote listening. Summary of the Invention
[0006] This application provides a remote USB device voice communication method based on edge computing and browser collaboration, which can build a bidirectional controllable real-time voice channel between the control end and the edge USB audio device by relying solely on a general browser.
[0007] In a first aspect, this application provides a remote USB device voice communication method based on edge computing and browser collaboration, comprising:
[0008] Obtain the intercom session establishment request from the control terminal, wherein the intercom session establishment request includes the control terminal browser identifier, the edge terminal device identifier, and session parameters;
[0009] A communication service based on the WebSocket protocol is started on the edge processing node, listening on a preset port to receive access requests from the control terminal browser;
[0010] A WebSocket connection is established between the control terminal browser and the edge terminal processing node, and capability negotiation is performed between the control terminal and the edge terminal through the WebSocket connection;
[0011] Based on the capability negotiation results, an uplink audio transmission path from the control end to the edge end and a downlink audio transmission path from the edge end to the control end are established.
[0012] The remote USB device voice communication method based on edge computing and browser collaboration provided in this application obtains the intercom session establishment request from the control end, then starts a communication service based on the WebSocket protocol on the edge processing node, listens on a preset port to receive the access request from the control end browser, then establishes a WebSocket connection between the control end browser and the edge processing node, and performs capability negotiation between the control end and the edge end through the WebSocket connection, and finally establishes an uplink audio transmission path from the control end to the edge end and a downlink audio transmission path from the edge end to the control end based on the capability negotiation result. Thus, under the condition of relying only on a general browser, a bidirectional and controllable real-time voice channel is built between the control end and the edge end USB audio device, so that both uplink remote shouting and downlink remote listening can operate stably in the network environment. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0014] Figure 1 This is a schematic flowchart illustrating a remote USB device voice communication method based on edge computing and browser collaboration, according to an example embodiment of this application.
[0015] Figure 2 This is a flowchart illustrating a specific implementation of S170 according to an example embodiment of this application;
[0016] Figure 3 This is a flowchart illustrating a specific implementation of S180 according to an example embodiment of this application;
[0017] Figure 4 This is a schematic diagram of the structure of a remote USB device voice communication device based on edge computing and browser collaboration, according to an example embodiment of this application;
[0018] Figure 5 This is a schematic diagram of the structure of an electronic device according to an example embodiment of this application.
[0019] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0020] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0021] Figure 1 This is a flowchart illustrating a remote USB device voice communication method based on edge computing and browser collaboration, according to an example embodiment of this application. Figure 1 As shown, the method provided in this embodiment includes:
[0022] S110, Obtain the intercom session establishment request from the control end.
[0023] In this step, you may obtain the intercom session establishment request from the control end. The intercom session establishment request includes the control end browser identifier, the edge device identifier, and session parameters.
[0024] Specifically, a session management service thread can be started in the edge processing node. The session management service thread listens to a preset HTTP interface or WebSocket handshake interface to receive intercom session establishment request messages from the control end browser.
[0025] The web application page is loaded in the browser on the control end. When the user triggers the intercom entry, the web application page sends a session establishment request message to the session management service thread through a preset Uniform Resource Locator. The session establishment request message is encapsulated in JSON format or form format.
[0026] When the control browser generates a session establishment request message, it obtains the control browser identifier from the current browser runtime environment and writes the control browser identifier into the corresponding browser identifier field in the session establishment request message. The control browser identifier includes at least one or more of the following: browser session identifier, browser type information, and / or user account identifier.
[0027] Based on the target device information selected by the user, the browser on the control end writes the corresponding edge device identifier into the device identifier field of the session establishment request message. The edge device identifier is used to uniquely identify the edge processing node participating in the intercom session and / or its connected USB audio device.
[0028] The browser on the control end generates session parameters based on the current business scenario and writes the session parameters into the corresponding session parameter field in the session establishment request message. The session parameters include at least one or more of the following: target audio encoding format, target sampling rate, target number of channels, expected session duration, and / or whether to enable the silence detection flag.
[0029] After receiving a session establishment request message, the session management service thread performs syntax and field integrity checks on the message, parses out the control end browser identifier, edge end device identifier, and session parameters, and creates a corresponding intercom session context based on the parsing results to support the establishment and maintenance of subsequent uplink and downlink audio transmission paths.
[0030] S120. Start the audio backend service on the edge processing node.
[0031] In this step, the audio backend service can be started on the edge processing node, the preset configuration can be loaded, the audio subsystem can be initialized, and the audio subsystem can be bound to the USB microphone and USB speaker to complete the initialization of the USB audio device.
[0032] Specifically, a backend service process based on the Netty framework can be started on an edge ARM architecture board. Then, the backend service process reads a preset audio configuration file. Next, it calls the underlying audio driver interface to initialize the audio input and output channels, binding the audio input channel to a USB microphone and the audio output channel to a USB speaker. Finally, the backend service process creates an audio acquisition thread and an audio playback thread, corresponding to the downlink and uplink audio transmission paths, respectively.
[0033] S130. Start a communication service based on the WebSocket protocol on the edge processing node and listen on a preset port to receive access requests from the control end browser.
[0034] In this step, a WebSocket-based communication service can be started on the edge processing node, listening on a preset port to receive access requests from the control browser.
[0035] Specifically, a network service process based on the Netty framework is started in the edge processing node, and a communication service instance for providing WebSocket protocol access is created in the network service process.
[0036] Configure the preset port number for listening and the corresponding transport layer protocol parameters in the communication service instance. The transport layer protocol parameters include at least one or more of the following: whether to enable TLS encryption, maximum number of connections, connection idle timeout time and / or receive buffer size.
[0037] Build a Netty channel pipeline in a communication service instance. The Netty channel pipeline includes at least one or more of the following: an HTTP request decoding processor, an HTTP response encoding processor, a WebSocket protocol handshake processor, and a WebSocket binary frame processor, to support the control browser to upgrade to a WebSocket long connection via an HTTP handshake.
[0038] Register a connection event listener callback in the communication service instance to accept access requests based on a preset port and trigger the WebSocket handshake process when a TCP connection establishment event from the control browser is detected.
[0039] In the WebSocket handshake process, the control browser performs protocol header verification and path verification on the handshake request message. After confirming that the target URI matches the preset intercom service access path, it returns a handshake response message to complete the upgrade and establishment of the WebSocket connection.
[0040] After completing the upgrade and establishment of the WebSocket connection, the corresponding channel context is associated with the intercom session management module so that uplink audio message frames, downlink audio message frames, and control signaling can be carried through the WebSocket connection in the future.
[0041] S140. Obtain local media streams through the control browser.
[0042] In this step, a local media stream can be obtained through the browser on the control terminal, where the local media stream includes audio data captured by the microphone on the control terminal.
[0043] Specifically, the control browser can request access to the local microphone and speaker devices on the control device via the getUserMedia interface. After user authorization, a local media stream object containing the microphone audio track is obtained. Then, the local media stream object is handed over to the Web Audio API processing node or the front-end audio encoding library for preprocessing and encoding.
[0044] S150. Establish a WebSocket connection between the control end browser and the edge processing node, and negotiate capabilities between the control end and the edge end through the WebSocket connection.
[0045] In this step, a WebSocket connection is established between the control end browser and the edge processing node, and capability negotiation is carried out between the control end and the edge end through the WebSocket connection. The capability negotiation includes at least the target audio encoding format, sampling rate parameters, and silence detection support capability.
[0046] Specifically, a WebSocket client instance is created in the control browser based on the preset intercom service Uniform Resource Locator, and connection establishment callbacks, message reception callbacks, and error callbacks are registered in the WebSocket client instance to handle WebSocket communication events with the edge processing nodes.
[0047] The control browser initiates a WebSocket connection establishment request to the edge processing node through a WebSocket client instance. The WebSocket connection establishment request carries access path information used to identify the intercom service and optional authentication parameters.
[0048] In the edge processing node, the communication service based on the WebSocket protocol receives and accepts the WebSocket connection establishment request, completes the establishment of the WebSocket connection with the control browser through the aforementioned WebSocket handshake process, and assigns a unique connection identifier to each established WebSocket connection to form a corresponding connection context.
[0049] After the WebSocket connection is established, the control browser constructs a capability negotiation request message in the connection establishment callback. The capability negotiation request message is encapsulated in JSON format or a preset binary format, and carries at least one or more of the following: a candidate list of target audio encoding format, a candidate list of sampling rate parameters, a mute detection flag, and a control browser identifier. It is then sent to the edge processing node via the WebSocket connection.
[0050] In the edge processing node, the capability negotiation processing module receives and parses the capability negotiation request message in the WebSocket binary frame processor or text frame processor. It extracts capability parameters such as the target audio encoding format candidate list, the sampling rate parameter candidate list, and the silence detection support capability. It then matches and filters these parameters with the actual support capabilities of the local audio subsystem to determine the negotiation results for the target audio encoding format, target sampling rate parameter, and whether silence detection is enabled for this intercom session.
[0051] After the negotiation result is determined, the edge processing node constructs a capability negotiation response message. The capability negotiation response message carries the target audio encoding format, target sampling rate parameter, and the final determined value of silence detection support capability, and is sent to the control end browser through the established WebSocket connection.
[0052] In the control browser, the capability negotiation response message is received and parsed through message receiving callback. The negotiation results of the target audio encoding format, target sampling rate parameters, and silence detection support capability are obtained from the message. Based on the negotiation results, the local media stream encoding parameters and silence detection strategy are configured to support the establishment and parameter consistency maintenance of the subsequent uplink audio transmission path from the control end to the edge end and the downlink audio transmission path from the edge end to the control end.
[0053] S160. Based on the capability negotiation results, establish the uplink audio transmission path from the control end to the edge end and the downlink audio transmission path from the edge end to the control end.
[0054] Specifically, after completing capability negotiation and obtaining the target audio encoding format, target sampling rate parameters, and silence detection support capabilities, the local media stream acquisition constraints are configured in the control browser based on the target sampling rate parameters, and the audio encoder instance corresponding to the target audio encoding format is initialized.
[0055] In the control browser, an uplink audio transmission channel is created based on the established WebSocket connection. The WebSocket binary frame transmission interface used to send uplink audio message frames is associated with the audio encoder instance to form the control side processing link of the uplink audio transmission path from the control end to the edge end.
[0056] In the edge processing node, an uplink audio decoder is configured according to the capability negotiation result. The audio encoding format of the uplink audio decoder is consistent with the target audio encoding format, and the sampling rate parameter is consistent with the target sampling rate parameter. The uplink audio decoder is bound to the audio playback buffer queue corresponding to the USB speaker to form the edge processing link of the uplink audio transmission path.
[0057] In the control browser, a timed processing task is created to carry uplink audio data. The timed processing task reads audio frame data from the local media stream according to a preset time window, calls an audio encoder instance to encode the audio frame data, and encapsulates it into uplink audio message frames based on a preset frame structure. The uplink audio sending channel is then sent to the edge processing node via a WebSocket connection, thus forming a continuous uplink audio data stream.
[0058] In the edge processing node, the uplink audio message frame is received through the WebSocket binary frame processor. The compressed audio payload in the uplink audio message frame is input into the uplink audio decoder for decoding to obtain uplink PCM audio data that matches the target sampling rate parameter. The uplink PCM audio data is then sent to the audio playback buffer queue and played through the audio output channel bound to the USB speaker to complete the data loop of the uplink audio transmission path.
[0059] In the edge processing node, a downlink audio encoder is configured according to the capability negotiation result. The audio encoding format of the downlink audio encoder is consistent with the target audio encoding format, and the sampling rate parameter is consistent with the target sampling rate parameter. The downlink audio encoder is bound to the audio acquisition thread corresponding to the USB microphone to form the edge processing link of the downlink audio transmission path from the edge end to the control end.
[0060] In the edge processing node, the audio acquisition thread continuously reads PCM data of ambient audio from the USB microphone according to the sampling period matched with the target sampling rate parameter. After the PCM data is processed by framing, it is input into the downlink audio encoder for encoding. Based on the preset frame structure, it is encapsulated into downlink audio message frames and sent to the control end browser in the form of binary frames through the established WebSocket connection, thus forming a continuous downlink audio data stream.
[0061] In the control browser, configure the downlink audio decoder based on the target audio encoding format and target sampling rate parameters, and receive downlink audio message frames from the edge processing node in the WebSocket message receiving callback. Input the compressed audio payload in the downlink audio message frame into the downlink audio decoder for decoding to obtain downlink PCM audio data consistent with the target sampling rate parameters.
[0062] In the control browser, the downlink PCM audio data is sent to the local audio playback unit for playback. The local audio playback unit includes the AudioContext node and its playback channel output to the local speaker, thus completing the data loop of the downlink audio transmission path.
[0063] The silence detection module is selectively enabled in the control browser and edge processing nodes based on silence detection support capabilities. Silence detection and gating control are performed on the corresponding PCM audio data in the uplink and / or downlink audio transmission paths to reduce the transmission load of invalid audio data on the WebSocket connection without affecting voice communication quality, and to ensure that the parameter configuration and capability negotiation results of the uplink and downlink audio transmission paths are consistent.
[0064] S170. In the uplink audio transmission path, the uplink audio data collected and encoded by the microphone at the control end is encapsulated into an uplink audio message frame according to a preset frame structure and sent to the edge processing node via a WebSocket connection.
[0065] In the uplink audio transmission path, the uplink audio data collected and encoded by the microphone at the control end is encapsulated into uplink audio message frames according to a preset frame structure and sent to the edge processing node via a WebSocket connection. The edge processing node parses and decodes the uplink audio message frames to obtain uplink PCM audio data, and outputs the uplink PCM audio data to a USB speaker for playback, thereby enabling remote shouting.
[0066] Figure 2 This is a flowchart illustrating a specific implementation of S170 according to an example embodiment of this application. For example... Figure 2 As shown, the above-mentioned S170 includes:
[0067] S171. On the browser side of the control terminal, the local media stream is segmented with a fixed time window to obtain continuous audio frame data.
[0068] Specifically, the local media stream is obtained through the getUserMedia interface in the control browser, and the local media stream is connected to the AudioContext audio processing graph. The AudioContext audio processing graph includes script processing nodes or audio worker thread nodes for accessing raw audio data.
[0069] In the control browser, the sampling rate of AudioContext is set according to the target sampling rate parameter obtained by capability negotiation, so that the original audio PCM data output from the local media stream is consistent with the target sampling rate parameter, or the original sampling rate PCM data is converted into PCM data corresponding to the target sampling rate parameter through the resampling module.
[0070] Create a timed frame-segmentation task in the control browser. The timed frame-segmentation task is triggered periodically according to a preset time window length (e.g., 10ms, 20ms, or 40ms). Each time it is triggered, it reads the cached PCM sampling data within the current time window from the script processing node or the audio worker thread node.
[0071] The PCM sampled data arranged continuously within each time window is buffered as a single audio frame to form a continuous audio frame data sequence. The adjacent audio frames are continuous and uninterrupted on the time axis, or the preset overlap length can be set as needed to adapt to the frame length requirements of the corresponding audio encoding format.
[0072] S172. After OPUS encoding or G.711 encoding of the audio frame data, the compressed audio payload is obtained.
[0073] Specifically, in the control browser, based on the target audio encoding format obtained through capability negotiation, the audio encoding library corresponding to OPUS encoding or G.711 encoding is loaded locally. The audio encoding library can be provided in the form of JavaScript, WebAssembly, or browser built-in codec modules.
[0074] Initialize the audio encoder instance in the control browser. The encoding parameters of the audio encoder instance include one or more of the following: target audio encoding format, target sampling rate parameter, number of channels, and target bit rate. After initialization, it enters the ready state.
[0075] In the timed framing task, after obtaining each frame of audio data, the corresponding PCM sampling data is input into the audio encoder instance for encoding processing with a preset sampling bit width and channel arrangement.
[0076] When the target audio encoding format is OPUS encoding, the audio encoder instance performs frame encoding on the input PCM sample data according to the OPUS encoding specification, and outputs OPUS compressed bitstream data as the corresponding compressed audio payload.
[0077] When the target audio encoding format is G.711, the audio encoder instance compresses and quantizes the input PCM sample data sample by sample according to the G.711 μ-law or A-law encoding rules, and outputs a G.711 encoded byte sequence as the corresponding compressed audio payload.
[0078] The encoded output obtained from each encoding process is buffered as an independent compressed audio payload unit, which maintains a one-to-one correspondence with the corresponding audio frame data in time, so as to facilitate subsequent frame header splicing and encapsulation.
[0079] S173. Assign a frame sequence number and timestamp to each frame of compressed audio payload, and set the payload type field to form the header of the uplink audio message frame.
[0080] Specifically, a frame sequence counter used to identify the order of uplink audio frames is maintained in the control browser. The frame sequence counter is initialized to zero or to a random starting value when the uplink audio transmission path is established, and it is incremented every time a compressed audio payload is generated to generate the corresponding frame sequence number.
[0081] In the control browser, a timestamp is assigned to each frame of compressed audio payload based on a local high-precision time source, including but not limited to AudioContext time base, performance.now time base, or system clock. The timestamp is used to characterize the position of the moment when the audio acquisition or encoding of that frame is completed in the overall session timeline.
[0082] Set a payload type field for the compressed audio payload according to the target audio encoding format. The payload type field is used to indicate the audio encoding type used in the current frame, including at least a type value to identify OPUS encoding and a type value to identify G.711 encoding.
[0083] In the control browser, the frame sequence number, timestamp, and payload type fields are organized according to the preset header structure. The uplink audio message frame header is constructed using a fixed byte length or a predefined format, and version number fields, session identifier fields, or reserved fields are added when needed to enhance compatibility and scalability.
[0084] S174. The frame header and compressed audio payload are concatenated and encapsulated into a binary uplink audio message frame, and sent to the edge processing node via WebSocket connection.
[0085] Specifically, in the control browser, after generating the frame header for each frame, a binary buffer is allocated to carry the audio data of that frame. The front of the binary buffer is used to store the frame header, and the back is used to store the corresponding compressed audio payload.
[0086] Each field of the frame header is written to the beginning of the binary buffer in the preset byte order, and then the corresponding compressed audio payload is copied to the remaining part of the binary buffer in the original byte order to form a complete binary uplink audio message frame.
[0087] In the control browser, each frame of binary uplink audio message is used as the payload of a WebSocket binary message and sent to the edge processing node through the previously established WebSocket connection. The sending process is executed in a non-blocking or asynchronous callback manner.
[0088] Optionally, in the control browser, the length of the buffer to be sent can be detected before sending binary uplink audio message frames. If the buffer backlog exceeds a preset threshold, some historical frames can be temporarily discarded or the frame interval can be adjusted to reduce the impact of network congestion on the real-time performance of uplink audio transmission, thereby ensuring stable real-time voice communication in the uplink audio transmission path.
[0089] S175. Receive binary uplink audio message frames that arrive via a WebSocket connection in the Netty backend service.
[0090] Specifically, on the edge processing node, a server bootstrap instance is created based on the Netty framework to carry WebSocket connections, the server channel type is configured as NioServerSocketChannel, and an event loop thread group is set in the bootstrap instance to handle network I / O events.
[0091] Configure a channel initializer for the sub-channel in the Netty server bootstrap instance. In the channel initializer, add the HTTP codec processor, HTTP aggregation processor, WebSocket handshake processor, and binary frame processor to the channel pipeline in sequence.
[0092] The WebSocket handshake processor performs handshake verification and upgrade processing on the HTTP upgrade request from the control browser according to the preset WebSocket access path and protocol version. After the handshake is successful, a stable WebSocket connection is established between the control browser and the edge processing node.
[0093] Register a read callback method for BinaryWebSocketFrame type messages in the binary frame processor. When the callback method is triggered, extract the ByteBuf buffer of the payload from the BinaryWebSocketFrame, and copy or reference all the byte data in the ByteBuf buffer as a complete binary uplink audio message frame, and submit it to the subsequent audio parsing and processing logic for processing.
[0094] S176. Parse the frame header of the uplink audio message frame and extract the frame sequence number, timestamp, and payload type.
[0095] Specifically, after receiving a binary uplink audio message frame, the binary frame processor first determines the fixed or variable length information of the frame header according to the preset frame structure, and then reads the corresponding byte from the starting offset position in the ByteBuf buffer.
[0096] The frame header is parsed according to the preset byte order. The first few bytes are parsed into a frame sequence number field to identify the frame order, the next few bytes are parsed into a timestamp field to represent the acquisition or encoding time, and one or more bytes immediately following it are parsed into a payload type field.
[0097] When parsing the frame sequence number field, the read unsigned integer value is converted into a local integer variable for subsequent frame order detection, packet loss detection, or out-of-order reordering.
[0098] When parsing the timestamp field, the read timestamp value is converted into a unified time base format for alignment with the local playback clock or jitter buffer control logic.
[0099] When parsing the payload type field, the read type identifier value is matched with the local preset encoding type mapping table to determine the specific audio encoding format used by the compressed audio payload in the current uplink audio message frame, providing a basis for subsequent decoder selection.
[0100] After parsing the frame header field, the starting offset position and data length of the compressed audio payload in the ByteBuf buffer are calculated based on the frame header length, in preparation for the subsequent extraction of the compressed audio payload.
[0101] S177. Select the corresponding decoder based on the payload type, decode the compressed audio payload, and generate uplink PCM audio data.
[0102] Specifically, an audio decoder management module is pre-built in the edge processing node. The audio decoder management module maintains OPUS decoder instances and G.711 decoder instances respectively, or dynamically creates the corresponding decoder instances when needed.
[0103] After parsing the frame header of the uplink audio message frame and obtaining the payload type, the corresponding decoder instance is selected in the audio decoder management module according to the payload type. When the payload type indicates OPUS encoding, the OPUS decoder instance is selected, and when the payload type indicates G.711 encoding, the G.711 decoder instance is selected.
[0104] In the edge processing node, based on the calculated starting offset position and data length, the corresponding compressed audio payload byte sequence is extracted from the ByteBuf buffer, and the byte sequence is provided as input data to the selected audio decoder instance.
[0105] When the payload type corresponds to OPUS encoding, the OPUS decoder instance performs frame decoding on the input OPUS compressed bitstream according to the OPUS decoding specification, and outputs a PCM audio data buffer that matches the target sampling rate parameter and number of channels obtained through capability negotiation.
[0106] When the payload type corresponds to G.711 encoding, the G.711 decoder instance performs dequantization and restoration on a sample-by-sample basis according to the G.711 μ-law or A-law decoding rules to generate the corresponding PCM audio data buffer.
[0107] In the edge processing node, the decoded PCM audio data may optionally undergo simple audio post-processing, including but not limited to volume gain adjustment, silence detection marking, or format conversion (e.g., from 16-bit signed integer to floating point) to generate uplink PCM audio data that meets the input requirements of the USB speaker driver interface.
[0108] S178. Send the upstream PCM audio data into the audio playback buffer queue and play it through the bound USB speaker driver interface.
[0109] Specifically, the edge processing node maintains an audio playback buffer queue corresponding to the USB speaker. The audio playback buffer queue adopts a circular buffer or blocking queue structure to cache the upstream PCM audio data to be played.
[0110] After decoding and optional post-processing of each frame of uplink PCM audio data, the corresponding PCM data blocks are written to the audio playback buffer queue in frame order, with frame sequence number or timestamp information attached during writing, for use in the playback thread for order verification or jitter control.
[0111] A separate audio playback thread is created and run in the edge processing node. The audio playback thread reads prepared uplink PCM audio data blocks from the audio playback buffer queue according to a reading cycle that matches the target sampling rate parameter.
[0112] In the audio playback thread, the PCM audio data of each data block is written to the corresponding audio output channel through the underlying USB audio driver interface or Alsa audio interface according to the preset sampling bit width, channel arrangement and buffer size.
[0113] When the available data in the audio playback buffer queue is detected to be lower than the preset level, mute PCM data can be optionally written to the USB speaker driver interface to avoid popping or noise caused by playback interruption, thereby continuously and smoothly playing uplink PCM audio data on the USB speaker and achieving the audio output effect of remote shouting.
[0114] S180. In the downlink audio transmission path, the edge processing node continuously collects ambient audio through a USB microphone, performs downlink audio encoding and frame encapsulation to form downlink audio message frames, and sends them to the control end browser through a WebSocket connection.
[0115] In the downlink audio transmission path, the edge processing node continuously collects ambient audio through a USB microphone, performs downlink audio encoding and frame encapsulation to form downlink audio message frames, and sends them to the control end browser via a WebSocket connection. The control end browser parses and decodes the downlink audio message frames to obtain downlink PCM audio data, which is then played through the local audio playback unit to achieve remote monitoring.
[0116] Figure 3 This is a flowchart illustrating a specific implementation of S180 according to an example embodiment of this application. Figure 3 As shown, the above-mentioned S180 includes:
[0117] S181. In the edge processing node, read the PCM data of ambient audio from the USB microphone through the audio acquisition thread.
[0118] Specifically, an independent audio acquisition thread is created for the audio input channel corresponding to the USB microphone in the edge processing node. The audio acquisition thread is started after the audio backend service is started and bound to the USB microphone, and it runs in a loop throughout the entire intercom session.
[0119] In the audio acquisition thread, the underlying audio driver interface or ALSA audio interface is called to read audio sampling data from the USB microphone in batches with preset sampling rate parameters, number of channels and sampling bit width. Each time, a fixed number of sampling points or a fixed byte length is read to form a raw PCM data buffer.
[0120] After each successful reading of the raw PCM data buffer from the USB microphone, timestamp information or acquisition sequence information is appended to the PCM data buffer to identify the time position of the PCM data in the entire downlink audio stream.
[0121] The PCM data buffer with added time information is written to the downlink audio acquisition buffer queue corresponding to the downlink audio transmission path for subsequent downlink audio encoding and frame encapsulation steps.
[0122] During the reading process, if the underlying audio driver interface returns an error code or the length of the read data is less than the expected length, the error retry logic is executed in the audio acquisition thread or a buffer filled with silent PCM data is generated and written to the downlink audio acquisition buffer queue to ensure the continuity of the downlink audio data stream.
[0123] S182. The PCM data is processed into frames according to the same or agreed time window as the uplink, and downlink PCM audio frames are generated.
[0124] Specifically, in the edge processing node, the number of sampling points and byte length corresponding to a single downlink PCM audio frame are calculated in advance based on the sampling rate parameters obtained through capability negotiation and the preset time window (e.g., 10ms, 20ms or 40ms).
[0125] Multiple raw PCM data buffers written by the audio acquisition thread are accumulated in the downlink audio acquisition buffer queue. When the length of the accumulated PCM data bytes is greater than or equal to the length of bytes required for a single downlink PCM audio frame, a continuous segment of PCM data is extracted from the head of the queue according to the frame length as a complete downlink PCM audio frame.
[0126] After capturing a downlink PCM audio frame, the timestamp or acquisition sequence number of the PCM data buffer first written within the corresponding interval of the frame is used as the reference time information of the frame. Optionally, the time information can be converted according to the sampling rate parameter to obtain the timestamp of the downlink PCM audio frame.
[0127] Each generated downlink PCM audio frame is stored in the downlink PCM frame buffer queue, providing ordered frame-level input data for subsequent OPUS encoding or G.711 encoding.
[0128] During frame processing, when the accumulated PCM data is insufficient to form a complete frame and the preset processing timeout period has been reached, the frame length can be supplemented with silent PCM data or zero padding to generate a downlink PCM audio frame with a length that is strictly consistent with the time window, thereby ensuring the uniformity of the downlink audio frames on the time scale.
[0129] S183. Perform OPUS encoding or G.711 encoding on the downlink PCM audio frame to obtain the downlink compressed audio payload.
[0130] Specifically, a downlink audio encoding module is pre-established in the edge processing node. The downlink audio encoding module initializes an OPUS encoder instance and a G.711 encoder instance respectively, and determines the encoding method used by the current session based on the target audio encoding format obtained through capability negotiation.
[0131] When the target audio encoding format is OPUS, the downlink PCM audio frames are sequentially retrieved from the downlink PCM frame buffer queue in the downlink audio encoding module. Each downlink PCM audio frame is input into the OPUS encoder instance as a unit and encoded according to the OPUS encoding specification to obtain the corresponding OPUS downlink compressed audio payload. The bit rate, complexity and delay parameters are configured as needed.
[0132] When the target audio encoding format is G.711, the downlink PCM audio frames are sequentially retrieved from the downlink PCM frame buffer queue in the downlink audio encoding module. Each downlink PCM audio frame is input into the G.711 encoder instance one by one according to the sample. Based on the G.711 μ-law or A-law compression rules, the linear PCM sample value is mapped to an 8-bit compressed encoding value to obtain the corresponding G.711 downlink compressed audio payload.
[0133] After encoding each downlink PCM audio frame, the generated downlink compressed audio payload is associated with the corresponding frame timestamp information and frame sequence number placeholder information, and the downlink compressed audio payload is temporarily stored in the downlink compressed payload buffer queue so that the frame sequence number, timestamp and payload type fields can be added to each downlink compressed audio payload in the future.
[0134] During the encoding process, when a single frame encoding failure is detected or the output payload length is zero, the system can choose to re-encode, discard the frame, or replace it with a silent frame, and record the corresponding error information in the log system to ensure the robustness of the downlink audio stream under abnormal conditions.
[0135] S184. Add a frame sequence number, timestamp, and payload type field to each downlink compressed audio payload and encapsulate it into a downlink audio message frame.
[0136] Specifically, the downlink frame sequence number counter is maintained in the edge processing node and initialized with the establishment of the intercom session. The downlink frame sequence number counter is set to the initial value at the beginning of the session. When a downlink compressed audio payload is generated, the current counter value is used as the frame sequence number corresponding to that frame, and the counter is incremented after use.
[0137] Based on the timestamp information recorded by the downlink compressed audio payload when the downlink PCM audio frame is generated, or based on the time base value calculated by the frame sequence number and time window, a corresponding timestamp field is set for the frame, which is used for timing restoration and jitter control when playing the video in the control browser.
[0138] Based on the target audio encoding format obtained through capability negotiation, a payload type field is set for each frame of downlink compressed audio payload. When the target audio encoding format is OPUS, the payload type field is set to a preset identifier value indicating OPUS encoding. When the target audio encoding format is G.711, the payload type field is set to a preset identifier value indicating G.711 encoding.
[0139] According to the preset downlink audio message frame structure, a buffer is allocated in memory to store the downlink audio message frame header and the downlink compressed audio payload. First, the frame sequence number field, timestamp field, and payload type field are written into the buffer in the agreed byte order and field length to form the downlink audio message frame header.
[0140] After writing the header of the downlink audio message frame, the corresponding downlink compressed audio payload byte sequence is written into the same buffer immediately after the frame header, thereby splicing and encapsulating a complete binary downlink audio message frame. The binary downlink audio message frame is then submitted to the WebSocket send buffer queue for sending to the control browser via the WebSocket connection.
[0141] S185. Send downlink audio message frames to the control browser via WebSocket connection.
[0142] Specifically, in the Netty backend service of the edge processing node, a corresponding WebSocket channel context is maintained for each established control end browser connection, and the downlink audio message source that matches the edge device identifier is associated in the channel context.
[0143] In the Netty backend service, create a downlink audio sending thread or use an event loop thread to read pre-packaged binary downlink audio message frames in frame order from the WebSocket send buffer queue.
[0144] Upon reading each frame of binary downlink audio message, the frame is encapsulated into a BinaryWebSocketFrame object supported by the Netty framework, and the corresponding channel context write and refresh methods are called to send the BinaryWebSocketFrame to the target control browser through the established WebSocket connection.
[0145] During transmission, if network congestion or a backlog in the transmission buffer of the WebSocket connection is detected to exceed a preset threshold, some expired downlink audio message frames can be discarded or a simple frame rate reduction can be performed to avoid continuous accumulation of downlink latency and ensure the real-time performance of remote monitoring.
[0146] When the intercom session ends or the control browser is detected to have closed the connection, the system stops reading new downlink audio message frames from the WebSocket send buffer queue, closes the corresponding WebSocket channel context and releases related resources, thereby completing the orderly closure of the downlink audio transmission path.
[0147] S186. Receive binary downlink audio message frames in the WebSocket callback function in the control browser.
[0148] Specifically, a WebSocket object is created in the browser on the control end using JavaScript, and the message received from the edge processing node is processed in the onmessage event callback function of the WebSocket object.
[0149] In the onmessage event callback function, it is determined that the received message type is binary data. If the binary data is of type ArrayBuffer or Blob, it is converted to type ArrayBuffer for subsequent byte-level parsing.
[0150] In the control browser, a downlink receive buffer queue is maintained for the downlink audio transmission path. After each binary downlink audio message frame is received, the corresponding ArrayBuffer object is inserted as a queue element into the tail of the downlink receive buffer queue to ensure the receiving order of the downlink audio message frames.
[0151] During the receiving process, when a WebSocket connection closure event or error event is detected, the browser on the control side stops writing new binary downlink audio message frames into the downlink receive buffer queue and triggers local cleanup logic to release temporary resources associated with the downlink audio message frames.
[0152] S187. Parse the frame header of the downlink audio message frame and extract the frame sequence number, timestamp, and payload type.
[0153] Specifically, in the control browser, the DataView object is used to access the ArrayBuffer corresponding to the binary downlink audio message frame byte by byte. According to the pre-agreed frame header format, the frame sequence number field, timestamp field, and payload type field are read sequentially from the beginning of the ArrayBuffer.
[0154] When reading the frame sequence number field, consecutive bytes are converted into unsigned integer values according to the agreed byte order and field length to obtain a frame sequence number that corresponds one-to-one with the downlink audio message frame, which is used for subsequent out-of-order detection and reordering processing.
[0155] When reading the timestamp field, based on the agreed time unit, the corresponding field is parsed into a timestamp value, and the timestamp value is compared with the current local playback time base to assist the audio synchronization and jitter buffering strategy on the control end browser side.
[0156] When reading the payload type field, the payload type field is parsed into a preset encoding format identifier value, and the downlink compressed audio payload carried by the downlink audio message frame is determined to be OPUS encoded data or G.711 encoded data based on the parsed encoding format identifier value.
[0157] After parsing the frame header field, the byte offset of the frame header in the binary downlink audio message frame is calculated, and the remaining byte sequence is extracted from the ArrayBuffer based on the offset as the binary data segment of the corresponding downlink compressed audio payload for subsequent decoding processing.
[0158] S188. Select the corresponding audio decoder based on the payload type field, decode the downlink compressed audio payload, and obtain downlink PCM audio data.
[0159] Specifically, the OPUS audio decoder module and the G.711 audio decoder module are preloaded or initialized in the control browser. The OPUS audio decoder module can be implemented through WebAssembly or JavaScript, while the G.711 audio decoder module is implemented through JavaScript table lookup or calculation.
[0160] After parsing the payload type field, when it is determined that the target audio encoding format represented by the payload type field is OPUS, the corresponding downlink compressed audio payload is input to the OPUS audio decoder module in the form of Uint8Array. In the OPUS audio decoder module, the downlink compressed audio payload is decoded according to the OPUS decoding specification to generate downlink PCM audio data that matches the session parameters.
[0161] When the target audio encoding format represented by the payload type field is G.711, the corresponding downlink compressed audio payload is input to the G.711 audio decoder module in the form of Uint8Array. In the G.711 audio decoder module, each 8-bit compressed sample is converted into a linear PCM sample value according to the G.711 μ-law or A-law inverse compression rule, thereby generating downlink PCM audio data.
[0162] During the decoding process, the buffer depth of the decoded downlink PCM audio data is dynamically adjusted based on the deviation between the timestamp field and the local playback clock. When the downlink delay is detected to be too large, some outdated downlink audio frames can be discarded appropriately. When the downlink buffer is detected to be insufficient, it is supplemented by inserting silent PCM data to ensure the smooth continuity of the downlink PCM audio data in the time dimension.
[0163] After successfully decoding each frame of downlink PCM audio data, it is associated with the corresponding frame sequence number and timestamp information and written to the downlink PCM playback buffer queue for orderly reading and playback when the data source is subsequently sent to the AudioContext node.
[0164] S189. Send the downlink PCM audio data to the data source of the AudioContext node, and play the audio through the local speaker on the control end.
[0165] Specifically, an AudioContext object is created in the control browser based on the Web Audio API, and a script processing node, audio work node, or AudioWorkletNode is created in the AudioContext object to play downlink PCM audio data, so as to realize the custom push and playback control of PCM data.
[0166] When playable downlink PCM audio data is detected in the downlink PCM playback buffer queue, several frames of downlink PCM audio data are retrieved sequentially from the downlink PCM playback buffer queue according to the frame sequence number. Based on the sampling rate, number of channels, and sampling bit width agreed in the session parameters, an AudioBuffer object matching the specified data is created in the AudioContext object or a corresponding Float32Array / Int16Array data view is constructed.
[0167] Write the downlink PCM audio data to the AudioBuffer object or to a shared buffer accessible by the AudioWorkletNode in the order of sampling. After writing, connect the AudioBuffer object to the destination output node of the AudioContext object, or push the downlink PCM audio data to the audio rendering pipeline through the processing callback function of the AudioWorkletNode.
[0168] During audio playback, the timestamp field is compared with the current playback time of the local AudioContext. If the playback time of a frame is found to be much earlier than the current time, the playback of that frame can be skipped. If the playback time of a frame is found to be much later than the current time, the playback time can be aligned by triggering playback with a short delay or by inserting mute data, thereby reducing the impact of jitter on the remote monitoring experience.
[0169] When it is necessary to end remote listening or when the WebSocket connection is detected to be closed, stop reading new downlink PCM audio data from the downlink PCM playback buffer queue, and call the close interface of the AudioContext object or disconnect from the destination output node to terminate audio playback through the local speaker on the control end and release related audio resources.
[0170] S190: Control signaling is transmitted between the control browser and the edge processing node based on the WebSocket connection.
[0171] Based on the WebSocket connection, control signaling is transmitted between the control end browser and the edge processing node. The control signaling includes key-button call control commands, device status query commands, and audio parameter configuration commands. The uplink and downlink audio transmission paths are started and stopped and their parameters are adjusted according to the control signaling to complete remote USB device voice intercom.
[0172] In this embodiment, by acquiring the intercom session establishment request from the control terminal, a communication service based on the WebSocket protocol is started on the edge processing node, listening on a preset port to receive access requests from the control terminal browser. Then, a WebSocket connection is established between the control terminal browser and the edge processing node, and capability negotiation is performed between the control terminal and the edge terminal through the WebSocket connection. Finally, based on the capability negotiation results, an uplink audio transmission path from the control terminal to the edge terminal and a downlink audio transmission path from the edge terminal to the control terminal are established. Thus, relying only on a general browser, a bidirectional and controllable real-time voice channel is built between the control terminal and the edge terminal USB audio device, enabling both uplink remote talking and downlink remote listening to operate stably in the network environment.
[0173] Figure 4 This is a schematic diagram illustrating the structure of a remote USB device voice communication device based on edge computing and browser collaboration, according to an example embodiment of this application. Figure 4 As shown, the device 300 provided in this embodiment includes:
[0174] The acquisition module 310 is used to acquire the intercom session establishment request of the control terminal, wherein the intercom session establishment request includes the control terminal browser identifier, the edge terminal device identifier, and session parameters;
[0175] The listening module 320 is used to start a communication service based on the WebSocket protocol on the edge processing node and listen to a preset port to receive access requests from the control end browser.
[0176] The negotiation module 330 is used to establish a WebSocket connection between the control terminal browser and the edge terminal processing node, and to negotiate capabilities between the control terminal and the edge terminal through the WebSocket connection;
[0177] The communication module 340 is used to establish an uplink audio transmission path from the control end to the edge end and a downlink audio transmission path from the edge end to the control end based on the capability negotiation results.
[0178] Figure 5 This is a schematic diagram of the structure of an electronic device according to an example embodiment of this application. For example... Figure 5 As shown, the electronic device 400 provided in this embodiment includes: a processor 401 and a memory 402; wherein:
[0179] Memory 402 is used to store computer programs, and the memory may also be flash memory.
[0180] Processor 401 is used to execute the execution instructions stored in the memory to implement the various steps in the above method. For details, please refer to the relevant descriptions in the preceding method embodiments.
[0181] Alternatively, the memory 402 can be either standalone or integrated with the processor 401.
[0182] When the memory 402 is a device independent of the processor 401, the electronic device 400 may further include:
[0183] Bus 403 is used to connect the memory 402 and the processor 401.
[0184] This embodiment also provides a readable storage medium storing a computer program, which, when executed by at least one processor of an electronic device, enables the electronic device to perform the methods provided in the various embodiments described above.
[0185] This embodiment also provides a program product including a computer program stored in a readable storage medium. At least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to perform the methods provided in the various embodiments described above.
[0186] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0187] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A remote USB device voice communication method based on edge computing and browser collaboration, characterized in that, include: Obtain the intercom session establishment request from the control terminal, wherein the intercom session establishment request includes the control terminal browser identifier, the edge terminal device identifier, and session parameters; A communication service based on the WebSocket protocol is started on the edge processing node, listening on a preset port to receive access requests from the control terminal browser; A WebSocket connection is established between the control terminal browser and the edge terminal processing node, and capability negotiation is performed between the control terminal and the edge terminal through the WebSocket connection; Based on the capability negotiation results, an uplink audio transmission path from the control end to the edge end and a downlink audio transmission path from the edge end to the control end are established.
2. The method according to claim 1, characterized in that, In the uplink audio transmission path, the uplink audio data collected and encoded by the control end microphone is encapsulated into an uplink audio message frame according to a preset frame structure and sent to the edge processing node through the WebSocket connection; The edge processing node parses and decodes the uplink audio message frame to obtain uplink PCM audio data, and outputs the uplink PCM audio data to a USB speaker for playback, thereby enabling remote voice communication.
3. The method according to claim 2, characterized in that, In the uplink audio transmission path, the uplink audio data collected and encoded by the control terminal microphone is encapsulated into an uplink audio message frame according to a preset frame structure, including: On the browser side of the control end, the local media stream is segmented into segments with a fixed time window to obtain continuous audio frame data; After OPUS encoding or G.711 encoding of the audio frame data, a compressed audio payload is obtained. Assign a frame sequence number and timestamp to each frame of compressed audio payload, and set the payload type field to form the header of the uplink audio message frame; The frame header and the compressed audio payload are concatenated and encapsulated into a binary uplink audio message frame, which is then sent to the edge processing node via the WebSocket connection.
4. The method according to claim 3, characterized in that, The step of parsing and decoding the uplink audio message frame at the edge processing node to obtain uplink PCM audio data, and then outputting the uplink PCM audio data to the USB speaker for playback includes: Receive binary uplink audio message frames that arrive via the WebSocket connection in the Netty backend service; The frame header of the uplink audio message frame is parsed to extract the frame sequence number, timestamp, and payload type; Based on the load type, a corresponding decoder is selected to decode the compressed audio load and generate uplink PCM audio data. The upstream PCM audio data is sent to the audio playback buffer queue and played through the bound USB speaker driver interface.
5. The method according to claim 2, characterized in that, In the downlink audio transmission path, the edge processing node continuously collects ambient audio through a USB microphone, performs downlink audio encoding and frame encapsulation to form downlink audio message frames, and sends them to the control end browser through the WebSocket connection; The browser on the control end parses and decodes the downlink audio message frames to obtain downlink PCM audio data, which is then played through the local audio playback unit to achieve remote monitoring.
6. The method according to claim 5, characterized in that, The process at the edge processing node continuously acquires ambient audio via the USB microphone, performs downlink audio encoding and frame encapsulation to form downlink audio message frames, including: In the edge processing node, PCM data of ambient audio is read from the USB microphone via an audio acquisition thread; The PCM data is segmented into frames according to a time window consistent with or agreed upon with the uplink, to generate downlink PCM audio frames. The downlink PCM audio frame is OPUS-coded or G.711-coded to obtain the downlink compressed audio payload. Add a frame sequence number, timestamp, and payload type field to each frame of downlink compressed audio payload, and encapsulate it into a downlink audio message frame; The downlink audio message frame is sent to the control browser via the WebSocket connection.
7. The method according to claim 6, characterized in that, The process of parsing and decoding the downlink audio message frames in the browser on the control end to obtain downlink PCM audio data, and then playing it through the local audio playback unit, includes: Receive binary downlink audio message frames in the WebSocket callback function in the control browser; Parse the frame header of the downlink audio message frame to extract the frame sequence number, timestamp, and payload type; Based on the payload type field, the corresponding audio decoder is selected to decode the downlink compressed audio payload and obtain downlink PCM audio data; The downlink PCM audio data is sent to the data source of the AudioContext node and played through the local speaker on the control terminal.
8. The method according to claim 1, characterized in that, After establishing the uplink audio transmission path from the control terminal to the edge terminal and the downlink audio transmission path from the edge terminal to the control terminal, the following is also included: Based on the WebSocket connection, control signaling is transmitted between the control terminal browser and the edge terminal processing node. The control signaling includes key call control instructions, device status query instructions, and audio parameter configuration instructions. The control signaling is used to start / stop the uplink audio transmission path and adjust the parameters of the downlink audio transmission path to enable remote USB device voice communication.
9. The method according to any one of claims 1-4, characterized in that, Before initiating the WebSocket-based communication service on the edge processing node, the following steps are also included: Start the audio backend service on the edge processing node, load the preset configuration, initialize the audio subsystem, and bind the audio subsystem to the USB microphone and USB speaker to complete the initialization of the USB audio device.
10. The method according to claim 9, characterized in that, The step of starting the audio backend service on the edge processing node, loading the preset configuration, initializing the audio subsystem, and binding the audio subsystem to the USB microphone and USB speaker includes: Start a backend service process based on the Netty framework on an edge ARM architecture board; The preset audio configuration file is read in the backend service process; The underlying audio driver interface is called to initialize the audio input and output channels, bind the audio input channel to the USB microphone, and bind the audio output channel to the USB speaker; An audio acquisition thread and an audio playback thread are created in the backend service process, corresponding to the downlink audio transmission path and the uplink audio transmission path, respectively.