Mining communication system and control method thereof
By designing a mining communication system, a combination of voice input devices, playback devices, and controllers was adopted to enable simultaneous communication between multiple clients, solving the problem of low utilization of communication resources in existing technologies and improving the system's communication efficiency and security.
Patent Information
- Application Number
- CN202511181752.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-28
AI Technical Summary
In existing mining communication systems, communication resource utilization is low, and simultaneous communication between multiple clients is not possible.
Design a mining communication system including multiple clients, each client having a voice input device, a voice playback device and a controller. The controller enables the broadcasting and mixed playback of voice information, supports encoding, decoding, echo suppression, noise suppression and automatic gain control, and uses a ring buffer and key matrix for control.
It improves the utilization rate of system communication resources, enables simultaneous communication between multiple clients, enhances communication security and reliability, and supports dynamic registration of new clients.
Smart Images

Figure CN121036788A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mine communication, in particular to a mine communication system and a control method thereof. BACKGROUND
[0002] In industrial environments such as mines, tunnels, petrochemical and power industries, voice communication is an important communication means, which not only ensures the safety of workers, but also directly relates to the work efficiency of personnel scheduling, safety linkage and emergency handling of unexpected events. At present, there are mainly two kinds of mine voice communication systems, which are half-duplex communication system based on analog wireless intercom and digital intercom system based on IP network.
[0003] The half-duplex communication system and the digital intercom system mentioned above can only realize point-to-point communication or single-point-to-multipoint broadcast information, that is, a single client can only receive voice information of one client at a time. If the clients need to communicate multiple times, the session needs to be established frequently, and thus there is a problem of low utilization of communication resources. SUMMARY
[0004] The present application provides a mine communication system and a control method thereof, which are used to solve the problem of low utilization of communication resources in the prior art mine communication system.
[0005] Specifically, the present application provides a mine communication system, which comprises a plurality of client terminals connected in communication with each other, and each of the client terminals comprises: a voice input device configured to collect voice information of a user; a voice playing device configured to play the voice information; a controller connected to the voice input device and the voice playing device, and configured to: acquire the voice information collected by the voice input device, and broadcast the voice information to other client terminals; receive voice information sent by other client terminals, and store the received voice information in corresponding buffer areas respectively; control the voice playing device to mix and play the voice information in the plurality of buffer areas.
[0006] Further, the controller encodes the voice information before broadcasting the voice information to other client terminals; and decodes the received voice information before storing the received voice information in corresponding buffer areas respectively.
[0007] Further, the voice input device performs echo suppression, noise suppression and automatic gain control on the voice information after collecting the voice information.
[0008] Further, the controller judges whether a cache area has been allocated for the client corresponding to the received voice information before storing the voice information into the corresponding cache area respectively; If not, a cache area is allocated for the client corresponding to the voice information.
[0009] Further, the controller is also connected with a key matrix, and is used to generate a control instruction according to the action of a key in the key matrix, and control the voice input device and / or the voice playing device according to the control instruction.
[0010] Further, the controller performs anti-shake processing on the action signal of the key in the key matrix after detecting the action signal.
[0011] Further, the cache area corresponding to each client is a ring-shaped cache area.
[0012] On the other hand, the application also provides a control method of the mine communication system as any one of the above, comprising: Obtaining the voice information collected by the voice input device, and broadcasting the voice information to other clients; Receiving the voice information sent by other clients, and storing the received voice information into the corresponding cache area respectively; Controlling the voice playing device to mix and play the voice information of the plurality of cache areas.
[0013] Further, before the step of broadcasting the voice information to other clients, the method further comprises: encoding processing the voice information; and Before the step of storing the received voice information into the corresponding cache area respectively, the method further comprises: decoding processing the received voice information.
[0014] Further, before the step of storing the received voice information into the corresponding cache area respectively, the method further comprises: judging whether a cache area has been allocated for the client corresponding to the voice information; If not, a cache area is allocated for the client corresponding to the voice information.
[0015] The technical scheme provided by the present application, after the client collects the voice information of the user, the voice information is broadcasted to other clients to realize the voice information sending of a single client to multiple clients; after the client receives the voice information of other clients, the voice information is respectively stored in the corresponding buffer area, and the voice information of multiple buffer areas is mixed and played by the voice playing device, thereby realizing the receiving and playing of the voice information of a single client to multiple clients. Since the technical scheme of the present application can realize the communication between multiple clients and multiple clients at the same time, compared with the prior art, the communication resource utilization rate of the system can be improved.
[0016] The above and other objects, advantages and features of the present application will become more apparent from the following detailed description of some embodiments thereof, when taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0017] Some specific embodiments of the present application will be described in detail hereinafter with reference to the accompanying drawings, by way of illustration and not by way of limitation. The same reference numerals in the drawings denote the same or similar components or parts. It should be understood by those skilled in the art that the drawings are not necessarily drawn to scale. In the drawings: Figure 1 is a schematic diagram of a client in a mine communication system according to an embodiment of the present application; Figure 2 is a flowchart of a control method of a mine communication system according to an embodiment of the present application; Figure 3 is a schematic diagram of a client in a mine communication system according to another embodiment of the present application. DETAILED DESCRIPTION
[0018] In the description of the present embodiments, the description of the terms "one embodiment", "some embodiments", "exemplary embodiment", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the exemplary description of the above terms does not necessarily mean the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0019] In one embodiment of the present application, a mine communication system is provided, which includes multiple clients, which can be interphones, such as Figure 1As shown, the client comprises a controller 11, and a voice input device 12, a voice playing device 13 and a communication device 14 connected with the controller 11, wherein the voice input device 12 can be a microphone, the voice playing device 13 can be a speaker, and the controller 11 is connected with an I2S interface, and the voice input device 12 and the voice playing device 13 are connected with the controller 11 through the I2S interface.
[0020] Each client in the embodiment is connected with other clients through the communication device 14, and the communication protocol between the communication devices 14 can be a TCP / IP protocol. The controller 11 of each client is provided with a memory, such as a NVS (Non-Volatile Storage) flash memory, and the memory of each controller 11 is provided with a buffer area of other clients.
[0021] The controller 11 is connected with an I2S interface, and the I2S interface is configured as a voice input channel and a voice output channel, wherein the voice input channel is used to connect the voice input device 12, and the voice output channel is used to connect the voice playing device 13, so as to realize plug and play of the voice input device 12 and the voice playing device 13.
[0022] The voice input device 12 is used to collect voice information of a user and send the voice information to the controller 11, and the controller 11 can send the voice information collected by the voice input device 12 to other clients, and control the voice playing device 13 to play voice information received from other clients.
[0023] Specifically, the control method of the mine communication system of the embodiment comprises the following steps as shown in the flowchart. Figure 2 As shown, the steps comprise: In step S101, voice information of a user is collected through a voice input device, and the voice information is broadcasted to other clients; In step S102, voice information sent by other clients is received, the received voice information is respectively stored in the corresponding buffer area, and the voice playing device is controlled to mix and play the voice information in the buffer areas.
[0024] In the embodiment, the controller 11 of each client can start a UDP receiving task thread udp_recv_task and a UDP sending task thread udp_send_task in advance. In step S101, the controller 11 obtains the voice information of a user collected by the voice input device 12 connected therewith through an av_stream module, and then the voice information is broadcasted to other clients through the communication device 14 by the UDP sending task thread udp_send_task according to a UDP protocol.
[0025] In step S102, the controller 11 can receive the voice information received from other clients by the UDP (User Datagram Protocol) task thread udp_recv_task, and store the received voice information in the corresponding buffer area. Then the controller 11 mixes the voice information stored in the buffer area in the memory by the audio mixer audio_mixer, and outputs the mixed voice information to the voice playing device 13 through the I2S interface, and sets the read callback function of the audio mixer using the audio_mixer_set_read_cb function, to realize real-time playing of the voice information.
[0026] The audio mixer audio_mixer described above is a tool for mixing, processing and managing multiple audio signals, which can process multiple audio sources at the same time and support different formats of audio mixing. The audio_mixer_set_read_cb function described above is an interface for setting the callback function of the audio mixer audio_mixer, which is usually used in embedded audio development or audio processing framework.
[0027] According to the above, in the embodiment, after the client collects the voice information of the user, the voice information is broadcasted to other clients to realize the voice information sending from a single client to multiple clients. After the client receives the voice information of other clients, the voice information is stored in the corresponding buffer area, and the voice playing device is controlled to mix and play the voice information in multiple buffer areas, so as to realize the receiving and playing of the voice information of multiple clients by a single client. Since the embodiment can realize the communication between multiple clients and multiple clients at the same time, compared with the prior art, the communication resource utilization rate of the system can be improved.
[0028] In some embodiments of the application, the controller 11 of the client encodes the collected voice information using an encoder before broadcasting the voice information to other clients.
[0029] Correspondingly, the client decodes the received voice information before storing the received voice information in the corresponding buffer area.
[0030] In this embodiment, the controller 11 can first encode the collected voice information through the esp_g711a_decode function, and then send the encoded voice information to other clients. The esp_g711a_decode function is a function for G.711 A-law decoding in ESP-ADF (Espressif Audio Development Framework), which can decode A-law compressed audio data into PCM format, which is usually 16-bit linear PCM (Pulse Code Modulation). G.711 is a commonly used audio codec standard, widely used in telephone communication and low bit rate audio transmission.
[0031] Correspondingly, after receiving the voice information of other clients, the controller 11 decodes the voice information through the av_audio_dec_write function since the voice information has been encoded, and then stores the decoded voice information in the corresponding buffer area. The av_audio_dec_write function is a function in the FFmpeg / Libav audio decoding module, which is used to write compressed audio data such as AAC, MP3, G.711, etc. to the audio decoder, trigger the decoding process and output PCM data. The av_audio_dec_write function is usually used for streaming audio decoding, such as reading compressed data block by block from the network or file and decoding.
[0032] In this embodiment, the client encodes the voice information before sending it to other clients, which can prevent the voice information from being stolen during transmission, thereby improving the security and reliability of communication between clients.
[0033] In some embodiments of the application, after the client collects the voice information of the user through the voice input device 12, it also performs echo suppression, noise suppression and automatic gain control on the collected voice information.
[0034] In this embodiment, double microphone noise or acoustic isolation can be used for echo suppression, where double microphone noise reduction uses two microphones, including a main microphone and a reference microphone, to suppress non-target direction sound (including echo) through beamforming. Acoustic isolation optimizes the physical position of the speaker and microphone (such as distance, angle tilt), reducing direct acoustic coupling.
[0035] Noise suppression can employ spectral subtraction, Wiener filtering or subspace method, wherein spectral subtraction first converts audio to frequency domain through short-time Fourier transform, then subtracts estimated noise spectrum from amplitude spectrum of noisy signal, finally preserves phase information, and reconstructs time-domain signal through inverse Fourier transform. Wiener filtering is to minimize mean square error and estimate spectrum of clean speech: subspace method is to decompose noisy signal into signal subspace (clean speech) and noise subspace (noise), and remove noise through projection.
[0036] Automatic gain control is used to dynamically adjust gain of input signal, ensure that output signal amplitude is stable in target range, and avoid too large (clipping distortion) or too small (low signal-to-noise ratio) volume caused by microphone sensitivity difference, speaker distance change or environmental noise fluctuation.
[0037] In the embodiment, the collected voice information is optimized in sound quality through the triple algorithms of echo suppression, noise suppression and automatic gain control, which can effectively overcome device squeal and environmental noise interference in a small space, and further improve accuracy and reliability of the collected voice information.
[0038] In some embodiments of the present application, the controller of the client further comprises, before storing the received voice information into the corresponding buffer area respectively: determining whether a buffer area has been allocated for the client corresponding to the voice information; if not, allocating a buffer area for the client corresponding to the voice information.
[0039] The technical solution of the present application can allocate a buffer area for a new client in the case that the new client joins the mine-used communication system, thereby realizing dynamic registration of the new client and improving reliability of the mine-used communication system.
[0040] In some embodiments of the present application, the client is further provided with a key matrix 15, as shown in Figure 3 The key matrix 15 has a plurality of keys, and the controller 11 is connected with the key matrix 15 and can detect action signals of the keys in the key matrix 15.
[0041] In the embodiment, the controller 11 initializes the key matrix 15 and establishes a GPIO interrupt response mechanism through the esp_periph_set_init() function, and configures functions of the keys in the key matrix 15 by calling the audio_board_key_init() function The esp_periph_set_init() function described above is a function of the ESP-ADF, which is used to initialize the peripheral control set, which is the core module of the ESP-ADF for managing audio-related peripherals. Through the peripheral set, multiple peripheral controllers can be conveniently registered, started, and stopped, and their events can be uniformly processed.
[0042] The execution process of the esp_periph_set_init() function includes three key stages: first, resource allocation is performed, and memory space is dynamically allocated through audio_calloc to create event groups and mutexes to ensure thread safety; then, the GPIO interrupt service is initialized, and an ESP_INTR_FLAG_LEVEL2 level interrupt is used to establish a basic environment for subsequent key matrix interrupt responses; finally, the event-driven architecture is configured, and the audio_event_iface_init is used to establish a peripheral event processing channel containing task stack size, priority, and running core parameters. The esp_periph_set_init() function implements dynamic peripheral ID allocation through periph_dynamic_id, supports runtime expansion of peripheral devices, and uses a reverse resource release mechanism for error handling. If any initialization step fails, the allocated resources will be automatically rolled back. This initialization process provides a configurable interrupt response infrastructure for the key matrix 15, and subsequent audio_board_key_init() can bind specific key functions to this framework.
[0043] The audio_board_key_init() function described above is a function for initializing the key (button) function on the audio development board, which can configure the GPIO pins connected to the keys (such as play / pause, volume up / down buttons), set the GPIO to input mode, enable the internal pull-up resistor (to avoid floating), and configure the interrupt (such as falling edge trigger) to detect key press events.
[0044] The audio_board_key_init() function implements an ADC-based intelligent key initialization system, and accurately identifies the triggering state of 6 physical keys through a multi-level voltage threshold detection mechanism. The function first initializes the hardware collection parameters of ADC1 channel 4, maps each key to a unique voltage interval using a resistance voltage division principle; then dynamically selects internal or external storage space according to system memory configuration, the storage space is detected by audio_mem_spiram_stack_is_enabled, and finally the initialized ADC key peripheral is mounted in a unified peripheral management set (i.e. esp_periph_set_handle_t), establishing a complete key response channel including hardware abstraction layer, voltage sampling and event reporting. The design realizes the bridging of hardware resources and software events through periph_adc_button_init, supporting low-latency key interrupt response in the audio processing framework.
[0045] The controller 11 can generate a control instruction according to the action of the key in the key matrix 15, and control the voice input device 12 and / or the voice playing device 13 according to the control instruction. For example, the keys in the key matrix 15 can include a volume increase key, a volume decrease key, and a communication switching key. If the controller 11 detects that the user presses the volume increase key, the controller 11 controls the voice playing device 13 to increase the volume. If the controller 11 detects that the user presses the volume decrease key, the controller 11 controls the voice playing device 13 to decrease the volume. The communication switching key includes a voice input key and a voice playing key. If the controller 11 detects that the user is pressing the voice input key, the controller 11 starts to receive the voice information collected by the voice input device 12. If the controller 11 detects that the user is pressing the voice playing key, the controller 11 controls the voice playing device 13 to play the voice information stored in each buffer area.
[0046] In the embodiment, each key in the key matrix 15 is provided with a corresponding indicator light, for example, each key is connected with a light-emitting diode. When a key is pressed, the light-emitting diode corresponding to the key also emits light.
[0047] The client in the embodiment is provided with a key matrix, and the controller can control the client according to the action of the key in the key matrix, thereby improving the controllability of the client.
[0048] In some embodiments of the application, the controller 11 of the client further performs a shake prevention process on the key in the key matrix 15 during the detection of the key in the key matrix 15.
[0049] In the embodiment, the method for preventing the key from shaking includes: The controller 11 detects whether the time length of the key being pressed in the key matrix 15 is greater than a first preset time length after detecting that the key in the key matrix 15 is pressed. If yes, it is determined that the key is pressed; if no, it is determined that the key is not pressed.
[0050] Correspondingly, the controller 11 detects whether the time length of the key being released in the key matrix 15 is greater than a second preset time length after detecting that the key in the key matrix 15 is released. If yes, it is determined that the key is released; if no, it is determined that the key is not released.
[0051] In the embodiment, the first preset time length and the second preset time length are preferably 20 ms.
[0052] The anti-jitter processing is performed on the keys of the client in the embodiment, which can prevent the keys from being mistakenly touched, so as to improve the stability and reliability of the client.
[0053] In some embodiments of the application, each buffer area arranged on the client is a ring buffer area.
[0054] In the embodiment, one of the buffer areas is taken as an example, the buffer area has a plurality of storage addresses, and the plurality of storage addresses are sequentially connected to form a ring buffer area. For example, it is assumed that the buffer area has m storage addresses, the next storage address of the i-th storage address is the i+1-th storage address, and the next storage address of the m-th storage address is the 1-th storage address.
[0055] The buffer area of each client is arranged as a ring buffer area in the embodiment, which can prevent the limited storage space of the buffer area from affecting the storage of voice information.
[0056] At this point, those skilled in the art should recognize that, although the plurality of exemplary embodiments of the application have been shown and described in detail herein, many other variants or modifications conforming to the principles of the application can be directly determined or deduced from the disclosure of the application without departing from the spirit and scope of the application. Therefore, the scope of the application should be understood and recognized as covering all these other variants or modifications.
Claims
1. A mining communication system, characterized in that, It includes multiple clients that are interconnected, and each client includes: A voice input device used to collect the user's voice information; A voice playback device used to play voice information; A controller, which is connected to the voice input device and the voice playback device, and is configured to: Acquire the voice information collected by the voice input device and broadcast the voice information to other clients; It receives voice information sent by other clients and stores the received voice information into their respective buffers. The voice playback device is controlled to play voice information from multiple buffers in a mixed manner.
2. The mining communication system according to claim 1, characterized in that, The controller encodes the voice information before broadcasting it to other clients. as well as Before storing the received voice information into its corresponding buffer, the received voice information is decoded.
3. The mining communication system according to claim 1, characterized in that, After acquiring voice information, the voice input device performs echo suppression, noise suppression, and automatic gain control on the voice information.
4. The mining communication system according to claim 1, characterized in that, Before storing the received voice information into its corresponding buffer, the controller determines whether a buffer has been allocated for the client corresponding to the voice information. If not, a cache area will be allocated for the client corresponding to the voice information.
5. The mining communication system according to claim 1, characterized in that, The controller is also connected to a key matrix and is used to generate control commands based on the actions of the keys in the key matrix, and to control the voice input device and / or the voice playback device based on the control commands.
6. The mining communication system according to claim 5, characterized in that, After detecting the action signal of a key in the key matrix, the controller performs anti-bounce processing on the action signal.
7. The mining communication system according to claim 1, characterized in that, Each client corresponds to a circular cache.
8. A control method for a mining communication system as described in any one of claims 1-7, characterized in that, include: Acquire the voice information collected by the voice input device and broadcast the voice information to other clients; It receives voice information sent by other clients and stores the received voice information into their respective buffers. The voice playback device is controlled to play voice information from multiple buffers in a mixed manner.
9. The control method according to claim 8, characterized in that, Prior to the step of broadcasting the voice information to other clients, the method further includes: encoding the voice information; and Before the step of storing the received voice information into its corresponding buffer, the method further includes: decoding the received voice information.
10. The control method according to claim 8, characterized in that, Before the step of storing the received voice information into its corresponding buffer, the method further includes: determining whether a buffer has been allocated for the client corresponding to the voice information; If not, a cache area will be allocated for the client corresponding to the voice information.