Remote voice communication system suitable for building construction environment
By using multi-microphone arrays and Zigbee technology in building construction environments, combining noise and echo perception modules, dynamically adjusting the sound pickup mode, the noise and echo interference problems at the construction site are solved, and efficient long-distance voice communication is achieved.
Patent Information
- Application Number
- CN202510576978.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In building construction environments, traditional voice communication systems are disturbed by construction noise and echo, resulting in a decline in communication quality, especially when communicating from a long distance, which cannot accurately capture voice signals, affecting the real-time communication effect.
Multi-microphone array and Zigbee technology are used to build an ad hoc network, combining the noise sensing module and echo sensing module on the construction site, monitoring and classification of noise and echo sources in real time, and optimizing the audio link through the dynamic pickup mode adjustment module to reduce noise interference and echo impact.
It improves the voice communication effect at the construction site, reduces communication obstacles caused by echo and noise, and achieves high reliability and high definition voice communication in complex environments.
Smart Images

Figure CN120455596A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice communication technology, and in particular to a long-distance voice communication system suitable for a building construction environment. Background Art
[0002] In traditional building construction environments, long-distance voice communication systems face a series of challenges due to factors such as construction noise and complex building structures, especially in terms of real-time voice transmission and maintaining voice quality at the construction site. As the construction industry continues to increase its requirements for construction site safety, efficiency, and coordination, effective communication between construction workers has become particularly important. Traditional voice communication systems in such environments are often plagued by noise interference and echo, which affects the clarity and accuracy of communication and may even lead to communication failure.
[0003] First, construction sites are often accompanied by numerous noise sources, including the sounds of machinery, building materials, and concrete mixing. These noise sources easily interfere with the signals of traditional voice communication systems, causing received voice signals to be drowned out by the noise, severely impacting the quality of voice communication. This is especially true when construction workers need to communicate over long distances. Microphones are easily affected by this ambient noise, making it difficult to accurately capture voice signals, significantly reducing voice quality.
[0004] Secondly, echo is a common problem in construction sites, especially in large structures. Due to the reflective properties of buildings, voice signals can reflect in open spaces, causing echo. This echo not only significantly reduces voice signal clarity but can also increase voice signal latency, further impacting real-time communication. Echo is particularly severe in construction environments, especially when voice signals are transmitted across multiple microphone arrays. Different echo sources can interfere with signal transmission, causing communication link anomalies. Traditional voice communication systems often lack the ability to effectively identify and classify noise and echo sources, making it impossible to adjust audio link configuration in real time. This results in ineffective solutions to noise interference and echo.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0006] The object of the present invention is to provide a long-distance voice communication system suitable for a building construction environment, so as to solve the problems raised in the above background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A long-distance voice communication system suitable for a building construction environment, comprising:
[0009] The communication deployment module is used to deploy multiple microphone arrays at the building construction site. Each construction worker is equipped with a microphone to form multiple microphone nodes. These microphone nodes are connected via Zigbee technology. The module collects and processes the signals from the multi-microphone array and outputs the direction and distance of the voice source.
[0010] The construction site noise sensing module is used to monitor the background noise data of the building construction site in real time, identify and automatically classify the noise source, construct a background sound pressure index, and determine whether there is noise interference in the environment based on the background sound pressure index. If so, it triggers the first audio link layer adjustment instruction;
[0011] The echo perception module is used to monitor the echo characteristics of the i-th microphone signal in the voice communication system of the building construction site in real time, identify and automatically classify the echo source type, and construct the echo sound pressure index Hs of the i-th microphone. i , to determine whether there is an abnormal echo phenomenon in the current communication link. If so, obtain the abnormal echo group;
[0012] Dynamic pickup pattern adjustment module, used to obtain the arrival angle DoA of the i-th microphone echo signal through cross-correlation analysis based on the echo anomaly group i , and generates corresponding second audio link layer adjustment instructions and third audio link layer adjustment instructions.
[0013] Furthermore, the communication arrangement module includes a building three-dimensional modeling unit, a microphone array arrangement unit and a central processing unit;
[0014] The building 3D modeling unit is used to scan the target building with a laser scanner to obtain 3D point cloud data of the building's exterior and interior, and to construct a 3D building model using 3D modeling software such as Revit, SketchUp, or Navisworks. The unit also marks different construction areas in the 3D building model, including high-altitude work areas, ground work areas, and areas surrounding construction equipment.
[0015] The microphone array arrangement unit is used to arrange multiple microphone arrays according to the spatial layout of the construction site and the needs of voice communication. The multiple microphone nodes are interconnected through Zigbee technology to form an ad hoc network structure, which can transmit the collected voice signals to the central processing unit in real time;
[0016] The central processing unit is used to capture the sound waves of each microphone and convert them into electrical signals, obtain the voice signal of the i-th microphone, and build a microphone array located in the three-dimensional model of the building. Each microphone m1, m2, ..., mM , M represents the total number of microphones; the position vector (x i ,y i ,z i );
[0017] The ambient sound waves captured by each microphone are converted into electrical signals. The central processing unit is responsible for centrally processing the signals collected by all microphones. The signal received by the i-th microphone Expressed as:
[0018]
[0019] Where, is the signal received by the i-th microphone at time j, and A is the signal amplitude;
[0020] f(r i ,θ,φ,j) is the signal function of the sound wave propagating to the i-th microphone; specifically, it can be understood as the signal waveform when the sound wave propagates to the i-th microphone in three-dimensional space, which depends not only on the spatial position of the microphone, but also on the direction, distance, and time of the sound source;
[0021] f(r i ,θ,φ,j) is the waveform signal of the sound signal emitted from the speech source (θ,φ,r) and propagated to the i-th microphone, (x i ,y i , z i ) is the position vector of the i-th microphone, θ and φ are the azimuth and elevation angles of the sound source, respectively, r i is the distance from the sound source to the i-th microphone, j represents time. Considering that the propagation of sound waves is dynamic, the sound signal will change over time; is the interference signal from the noise at time t.
[0022] Furthermore, the construction site noise perception module includes a noise collection unit and a noise data classification unit;
[0023] The noise collection unit is used to collect noise data from the i-th microphone using a noise monitor, establish a noise data set, and process the noise data. The noise data classification unit extracts noise features and then classifies the noise data. The classification results include:
[0024] The frequency of low-frequency mechanical noise of tower cranes is 20Hz to 80Hz;
[0025] The mechanical noise frequency of compressor and air pump equipment is 20Hz to 80Hz;
[0026] The frequency of low-frequency mechanical noise of concrete mixers is 80Hz to 150Hz;
[0027] The frequency range of low-frequency noise from transport vehicles is 80Hz to 150Hz;
[0028] The frequency range of low-frequency mechanical noise of excavators is 80Hz to 500Hz;
[0029] The frequency range of human knocking noise is 150Hz to 500Hz;
[0030] The frequency range of manual handling noise is 500Hz to 5kHz;
[0031] The frequency range of chainsaw noise is 1kHz to 6kHz;
[0032] The frequency range of electric drill noise is 1kHz to 10kHz.
[0033] Furthermore, the construction site noise perception module further includes an on-site sound pressure analysis unit, a sound pressure judgment unit, and a frequency band dynamic gain unit;
[0034] The on-site sound pressure analysis unit is used to construct the background sound pressure index Zs of the i-th microphone based on the noise data set. i , the specific steps are:
[0035] S111. Each construction worker wears a microphone. There is noise in the i-th microphone environment. The sound pressure Px at the i-th microphone environment measurement point is collected. i , calculate the sound pressure level Lp of the i-th microphone using the following formula i :
[0036]
[0037] Where P0 is the reference sound pressure, which is set to 20×10-6Pa;
[0038] S112, and collect the average noise level of the environment in the sampling period T, and calculate the equivalent continuous sound level Leq of the i-th microphone using the following formula: i :
[0039]
[0040] Where T is the sampling time period, P i 2 (j) is the square of the instantaneous sound pressure, P i (j) represents the instantaneous sound pressure of the noise signal of the i-th microphone at time j; represents the integration of the square of the instantaneous sound pressure within the sampling period T, which represents the total energy of the noise; dj represents the time differential, which is integrated over time j;
[0041] S113, and collect the maximum instantaneous sound pressure level of the environment within the sampling period T, which is used to measure those noise events with large instantaneous intensity, reflect the suddenness and instantaneous maximum loudness, and calculate the peak sound level Lpeak of the i-th microphone using the following formula i :
[0042]
[0043] Where, P peak,i Indicates the value of the maximum instantaneous sound pressure of the noise signal;
[0044] S114: Combine the sound pressure level Lp of the i-th microphone obtained in S111-S113 i , equivalent continuous sound level Leq i and peak sound level Lpeak i After dimensionless processing, the background sound pressure index Zs of the i-th microphone is calculated by the following formula: i :
[0045] Zs i =α1Lp i +α2Leq i +α3Lpeak i
[0046] Where α1, α2 and α3 are all constants, and α1+α2+α3=1, α1Lp i +α2Leq i +α3Lpeak i Indicates that the background sound pressure index Zs of the i-th microphone is calculated according to the weights α1, α2 and α3 i ;
[0047] The sound pressure judgment unit is used to preset the background sound pressure threshold X and calculate the background sound pressure index Zs of the i-th microphone. i Comparing with the background sound pressure threshold X to obtain a first determination result includes:
[0048] When the echo sound pressure index Hs of the i-th microphone i > Background sound pressure threshold X, indicating that the construction site background noise of the i-th microphone is abnormal, triggering the first audio link layer adjustment instruction;
[0049] When the echo sound pressure index Hs of the i-th microphone i When ≤ background sound pressure threshold X, it indicates that the background noise of the construction site of the i-th microphone is normal, the noise interference is within the expected range, and communication and monitoring are continuous.
[0050] Furthermore, the frequency band dynamic gain unit, after receiving the first audio link layer adjustment instruction, sets the microphone channel as the "noise interference focus channel" and divides the audio signal of the i-th microphone into 6 sub-bands. According to the classification result obtained by the noise data classification unit, the current signal-to-noise ratio SNR of each sub-band is compared with its corresponding trigger threshold. If the frequency band corresponds to the trigger threshold, the noise reduction gain of the frequency band is automatically increased to the preset value. If SNR> the corresponding trigger threshold but meets Hs i >X, the noise reduction threshold will be temporarily relaxed by 3dB or the noise reduction gain of the frequency band will be increased by 5dB.
[0051] Furthermore, the echo perception module includes an echo feature real-time monitoring unit and an echo source classification and identification unit;
[0052] The real-time echo feature monitoring unit is used to collect audio signals from each microphone position in the voice communication system of the building construction site in real time, and extract echo features based on short-time energy analysis and cross-correlation function analysis to obtain the echo path delay, echo intensity, echo duration, echo frequency characteristics, arrival angle of the echo signal, and echo reverberation time of the i-th microphone, and establish an echo source data set;
[0053] The echo source classification and identification unit is used to extract the six-dimensional vector of echo path delay, echo intensity, echo duration, echo frequency characteristics, arrival angle of the echo signal, and echo reverberation time for each detection signal in the echo source data set. After using the support vector machine as a standard sample, an echo recognition model is established to automatically determine the echo source type.
[0054] Echo source types include: ground reflection echo, equipment surface reflection echo, wall hard reflection echo and structural resonance echo.
[0055] Furthermore, the echo perception module further includes an echo interference analysis unit;
[0056] The echo interference analysis unit is used to extract the echo strength Enregy of the i-th microphone in the echo source data set. i , Echo Duration i and the variance of the echo path delay i After dimensionless processing, the echo sound pressure index Hs of the i-th microphone is calculated using the following formula: i :
[0057] Hs i =α4Enregy i +α5Duration i +α6Variance i
[0058] Where α4, α5 and α6 are all constants, and α4+α5+α6=1, α4Enregy i +α5Duration i +α6Variance i Indicates that the echo sound pressure index Hs of the i-th microphone is calculated according to the weights α4, α5 and α6 i ;
[0059] Among them, the variance of the delay of the i-th echo path is i Calculated using the following formula:
[0060]
[0061] Where μ i is the echo path delay of the ith microphone, represents the average echo path delay of all microphones, and M represents the total number of microphones.
[0062] Furthermore, the echo perception module further includes an echo anomaly determination unit;
[0063] The echo abnormality determination unit is used to preset the echo threshold Y and calculate the echo sound pressure index Hs of the i-th microphone. i Comparing with the echo threshold Y to obtain a second determination result includes:
[0064] When the echo sound pressure index Hs of the i-th microphone i >Echo threshold Y×120%, indicating that the construction site echo of the i-th microphone is abnormal, generating the first echo abnormality level zone;
[0065] When the echo threshold Y≤the echo sound pressure index Hs of the i-th microphone i When the value is ≤ the echo threshold Y×120%, it indicates that the construction site echo of the i-th microphone is abnormal, and a second abnormal echo level zone is generated, which has a lower echo intensity than the first abnormal echo level zone;
[0066] When the echo sound pressure index Hs of the i-th microphone i When the value is less than the echo threshold Y, it means that there is no echo at the construction site of the microphone, and communication and monitoring are continuous;
[0067] The first echo abnormality level area and the second echo abnormality level area are counted to obtain an echo abnormality group.
[0068] Furthermore, the dynamic sound pickup pattern adjustment module includes an echo angle analysis unit, a first sound pickup pattern switching unit, and a second sound pickup pattern switching unit;
[0069] The echo angle analysis unit is used to extract the first echo abnormality level area and the second echo level area in the echo abnormality group, and obtain the arrival angle DoA of the i-th microphone echo signal by cross-correlation analysis. i :
[0070]
[0071] Where, Δ T represents the propagation delay time difference of the echo signal between the two microphones, d represents the distance between the i-th microphone and the adjacent i+1-th microphone, and c represents the speed of sound; 0 degrees means the signal comes from the front of the microphone array, 90 degrees means the signal comes from the right side of the array, and -90 degrees means the signal comes from the left side;
[0072] The first pickup mode switching unit is configured to generate a second audio link layer adjustment instruction based on the first echo abnormality level area, comprising: a first step of switching to a directional pickup mode and adjusting the pickup direction. Instead of continuing to pick up sound in the direction of the noise, the system turns to the "opposite direction" thereof. For example, if the noise comes from 30 degrees to the left, the system adjusts the pickup direction to 30 degrees to the right to avoid the strong noise source;
[0073] Step 2: Dynamically adjust the pickup beam width of the current microphone based on the echo signal strength and the distance between the i-th microphone. The beam width has a nonlinear relationship with the distance:
[0074] When the signal distance is less than one meter, the beam width is expanded to 1.2 times the standard width;
[0075] When the signal distance is greater than or equal to one meter and less than two meters, the standard beam width remains unchanged;
[0076] When the signal distance is greater than two meters and less than or equal to four meters, the beam width is compressed to 75 percent of the standard width;
[0077] When the signal distance is greater than four meters, the beam width is compressed to 50 percent of the standard width;
[0078] The pickup beam direction of the i-th microphone is adjusted to the corresponding target pickup direction, and directional enhancement processing is performed on the current audio frequency band according to the determined beam width.
[0079] Furthermore, the second pickup mode switching unit is configured to generate a third audio link layer adjustment instruction based on the second echo abnormality level area, comprising: a first step of switching the pickup beam mode of the i-th microphone from a directional pickup mode to an omnidirectional pickup mode, i.e., abandoning the pickup focus on a specific direction and no longer limiting the pickup angle range, but receiving audio signals within the full angle range;
[0080] Step 2: Enable echo cancellation mode and generate the corresponding echo suppression ratio for the echo source type, including:
[0081] Eliminate 60%-70% of the ground reflected echo;
[0082] Eliminate 50%-65% of the echo reflected from the device surface;
[0083] For hard-reflected echoes from walls, it can eliminate 71%-80% of the echoes.
[0084] For structural resonance echo, it can eliminate 81%-90% of the echo.
[0085] Compared with the prior art, the present invention has the following beneficial effects:
[0086] The system's construction site noise perception module, equipped with real-time noise source monitoring and automatic classification capabilities, constructs a background sound pressure index that reflects construction site noise characteristics. This index, a key indicator for determining the quality of the communication environment, dynamically determines whether severe noise interference exists. Based on this determination, it issues first audio link layer adjustment instructions, intelligently optimizing audio processing strategies and reducing the risk of voice being drowned out by construction noise.
[0087] The present invention also introduces an echo perception module, which not only supports real-time monitoring of the echo characteristics of each microphone signal, but also establishes an individualized echo index based on the sound pressure characteristics, thereby realizing automatic classification of different types of echo sources. By extracting the "echo anomaly group", the system can quickly lock the abnormal link node to avoid large-scale link performance degradation. And through the coordinated work of the echo angle analysis unit, the first pickup mode switching unit and the second pickup mode switching unit, the system can flexibly adapt to different construction environments. In areas with high echo source intensity, the system will actively avoid the direction of the echo source and adjust the beam width to avoid interference from high-intensity echo sources; in complex construction environments, the system can automatically adjust the pickup mode according to the source and intensity of the echo signal and select the most suitable pickup strategy. This adaptive adjustment greatly improves the voice communication effect at the construction site and reduces communication barriers caused by echo and noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 The figure is a flow chart of a long-distance voice communication system suitable for a building construction environment according to the present invention. DETAILED DESCRIPTION
[0089] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific embodiments.
[0090] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0091] Example 1:
[0092] See also Figure 1 The present invention provides a technical solution: a long-distance voice communication system suitable for a building construction environment, comprising:
[0093] The communication deployment module is used to deploy multiple microphone arrays at the building construction site. Each construction worker is equipped with a microphone to form multiple microphone nodes. These microphone nodes are connected via Zigbee technology. The module collects and processes the signals from the multi-microphone array and outputs the direction and distance of the voice source.
[0094] The construction site noise sensing module is used to monitor the background noise data of the building construction site in real time, identify and automatically classify the noise source, construct a background sound pressure index, and determine whether there is noise interference in the environment based on the background sound pressure index. If so, it triggers the first audio link layer adjustment instruction;
[0095] The echo perception module is used to monitor the echo characteristics of the i-th microphone signal in the voice communication system of the building construction site in real time, identify and automatically classify the echo source type, and construct the echo sound pressure index Hs of the i-th microphone. i , to determine whether there is an abnormal echo phenomenon in the current communication link. If so, obtain the abnormal echo group;
[0096] Dynamic pickup pattern adjustment module, used to obtain the arrival angle DoA of the i-th microphone echo signal through cross-correlation analysis based on the echo anomaly group i , and generates corresponding second audio link layer adjustment instructions and third audio link layer adjustment instructions.
[0097] In this embodiment, multiple microphone arrays are organized and deployed at the construction site through the communication layout module, and each construction worker is equipped with an independent microphone node. This not only improves the stability of voice signal acquisition, but also combines Zigbee network technology to achieve low-latency, high-reliability multi-point voice data transmission, effectively reducing the signal distortion problem caused by noise interference in long-distance communication.
[0098] The system's construction site noise perception module provides real-time noise source monitoring and automatic classification, constructing a background sound pressure index that reflects the characteristics of construction site noise. This index, a key indicator for determining the quality of the communication environment, dynamically determines whether severe noise interference exists. Based on this determination, it issues first-level audio link layer adjustment instructions, intelligently optimizing audio processing strategies and reducing the risk of voice being drowned out by construction noise.
[0099] The system incorporates an echo perception module that not only monitors the echo characteristics of each microphone signal in real time but also establishes a personalized echo index based on sound pressure characteristics, enabling automatic classification of different echo sources. By extracting "echo anomaly clusters," the system can quickly identify abnormal link nodes, preventing widespread link performance degradation.
[0100] Based on the detection results of abnormal echo groups, the system further uses a dynamic pickup pattern adjustment module and a cross-correlation method to calculate the arrival angle of the echo signal. This allows for precise adjustment of parameters such as the pickup beam direction and gain strategy, and generates adjustment instructions for the second and third audio link layers. This source-oriented dynamic optimization mechanism enables adaptive adjustments in complex construction environments, improving the robustness and stability of voice communications.
[0101] Unlike traditional static noise reduction and echo suppression methods, this invention utilizes a modular, multi-index, and adjustable system architecture, resulting in enhanced environmental adaptability and interference resistance. It is particularly well-suited for construction in large, open, or highly reflective building environments, meeting the practical needs of on-site construction for highly reliable, high-definition voice communications.
[0102] Example 2
[0103] This example is explained in Example 1. Figure 1 ,Specifically, the communication arrangement module includes a building 3D modeling unit, a microphone array arrangement unit and a central processing unit;
[0104] The building 3D modeling unit is used to scan the target building with a laser scanner to obtain 3D point cloud data of the building's exterior and interior, and to construct a 3D building model using 3D modeling software such as Revit, SketchUp, or Navisworks. The unit also marks different construction areas in the 3D building model, including high-altitude work areas, ground work areas, and areas surrounding construction equipment.
[0105] The microphone array arrangement unit is used to arrange multiple microphone arrays according to the spatial layout of the construction site and the needs of voice communication. The multiple microphone nodes are interconnected through Zigbee technology to form an ad hoc network structure, which can transmit the collected voice signals to the central processing unit in real time;
[0106] The central processing unit is used to capture the sound waves of each microphone and convert them into electrical signals, obtain the voice signal of the i-th microphone, and build a microphone array located in the three-dimensional model of the building. Each microphone m1, m2, ..., m M , M represents the total number of microphones; the position vector (x i ,y i ,z i );
[0107] The ambient sound waves captured by each microphone are converted into electrical signals. The central processing unit is responsible for centrally processing the signals collected by all microphones. The signal received by the i-th microphone Expressed as:
[0108]
[0109] Where, is the signal received by the i-th microphone at time j, and A is the signal amplitude;
[0110] f(r i ,θ,φ,j) is the signal function of the sound wave propagating to the i-th microphone; specifically, it can be understood as the signal waveform when the sound wave propagates to the i-th microphone in three-dimensional space, which depends not only on the spatial position of the microphone, but also on the direction, distance, and time of the sound source;
[0111] f(r i ,θ,φ,j) is the waveform signal of the sound signal emitted from the speech source (θ,φ,r) and propagated to the i-th microphone, (x i ,y i , z i ) is the position vector of the i-th microphone, θ and φ are the azimuth and elevation angles of the sound source, respectively, r iis the distance from the sound source to the i-th microphone, j represents time. Considering that the propagation of sound waves is dynamic, the sound signal will change over time; is the interference signal from the noise at time t.
[0112] In this embodiment, the module is used for spatial optimization of voice signal acquisition in a construction environment, combining three-dimensional modeling and microphone array technology to achieve efficient communication and signal analysis. Through three-dimensional laser scanning and modeling, the precise positioning and layout of microphones in the construction environment are achieved, thereby improving the spatial coverage and accuracy of sound source acquisition. The self-organizing wireless network built based on Zigbee technology can achieve efficient real-time transmission of multi-point voice signals in complex construction sites and reduce the communication interruption rate. The central processing unit combines the sound wave propagation function with position information to preliminarily filter background noise and improve the clarity and recognition accuracy of voice signals. The system takes into account multi-dimensional factors such as sound source direction, distance, and time changes, has good dynamic adaptability, and is suitable for voice perception needs under different construction stages and spatial structure changes.
[0113] Example 3
[0114] This example is explained in Example 1. Figure 1 ,Specifically, the construction site noise perception module includes a noise ,collection unit and a noise data classification unit;
[0115] The noise collection unit is used to collect noise data from the i-th microphone using a noise monitor, establish a noise data set, and process the noise data. The noise data classification unit extracts noise features and then classifies the noise data. The classification results include:
[0116] The frequency of low-frequency mechanical noise of tower cranes is 20Hz to 80Hz;
[0117] The mechanical noise frequency of compressor and air pump equipment is 20Hz to 80Hz;
[0118] The frequency of low-frequency mechanical noise of concrete mixers is 80Hz to 150Hz;
[0119] The frequency range of low-frequency noise from transport vehicles is 80Hz to 150Hz;
[0120] The frequency range of low-frequency mechanical noise of excavators is 80Hz to 500Hz;
[0121] The frequency range of human knocking noise is 150Hz to 500Hz;
[0122] The frequency range of manual handling noise is 500Hz to 5kHz;
[0123] The frequency range of chainsaw noise is 1kHz to 6kHz;
[0124] The frequency range of electric drill noise is 1kHz to 10kHz.
[0125] In this embodiment, the noise at each microphone node is collected by a noise monitor, and noise characteristics are extracted by frequency domain analysis. This can accurately identify and classify various common noise sources on construction sites, such as tower cranes, compressors, concrete mixers, transport vehicles, electric drills, etc.; and achieve coverage of the frequency range from 20Hz to 10kHz, greatly improving the accuracy of noise data discrimination.
[0126] Example 4
[0127] This example is explained in Example 3. Figure 1 ,Specifically, the construction site noise perception module also includes an on-site sound pressure analysis unit, a sound pressure judgment unit, and a frequency band dynamic gain unit;
[0128] The on-site sound pressure analysis unit is used to construct the background sound pressure index Zs of the i-th microphone based on the noise data set. i , the specific steps are:
[0129] S111. Each construction worker wears a microphone. There is noise in the i-th microphone environment. The sound pressure Px at the i-th microphone environment measurement point is collected. i , calculate the sound pressure level Lp of the i-th microphone using the following formula i :
[0130]
[0131] Here, P0 is the reference sound pressure, set at 20 × 10-6 Pa. This quantifies the actual physical intensity of noise, providing basic assessment data for construction site noise intensity. It truly reflects the noise source intensity of different equipment or areas, facilitating classification and subsequent noise reduction allocation.
[0132] S112, and collect the average noise level of the environment in the sampling period T, and calculate the equivalent continuous sound level Leq of the i-th microphone using the following formula: i :
[0133]
[0134] Where T is the sampling time period, P i 2 (j) is the square of the instantaneous sound pressure, P i (j) represents the instantaneous sound pressure of the noise signal of the i-th microphone at time j; The integral of the square of the instantaneous sound pressure over sampling period T represents the total energy of the noise. dj represents the time differential, integrated over time j. This reflects the long-term noise exposure level and is more suitable for assessing the impact on construction workers' hearing. It smoothes out short-term noise peaks to more closely approximate the "average volume" perceived by the human ear.
[0135] Among them, the sound pressure Px at the i-th microphone environment measurement point i and instantaneous sound pressure P i It is a completely different concept, sound pressure Px i It is the pressure disturbance relative to the atmospheric static pressure generated at a certain point during the propagation of sound waves, and its unit is Pascal; instantaneous sound pressure P i It is the sound pressure value at a certain point in time, which is the original sound wave signal data that changes with time, such as vibration waveform;
[0136] S113, and collect the maximum instantaneous sound pressure level of the environment within the sampling period T, which is used to measure those noise events with large instantaneous intensity, reflect the suddenness and instantaneous maximum loudness, and calculate the peak sound level Lpeak of the i-th microphone using the following formula i :
[0137]
[0138] Where, P peak,i Indicates the value of the maximum instantaneous sound pressure of the noise signal;
[0139] S114: Combine the sound pressure level Lp of the i-th microphone obtained in S111-S113 i , equivalent continuous sound level Leq i and peak sound level Lpeak i After dimensionless processing, the background sound pressure index Zs of the i-th microphone is calculated by the following formula: i :
[0140] Zs i =α1Lp i +α2Leq i +α3Lpeak i
[0141] Where α1, α2 and α3 are all constants, and α1+α2+α3=1, α1Lp i +α2Leq i +α3Lpeak i Indicates that the background sound pressure index Zs of the i-th microphone is calculated according to the weights α1, α2 and α3 i ; Through weighted fusion, the impact of different factors on background noise perception can be dynamically reflected.
[0142] The sound pressure judgment unit is used to preset the background sound pressure threshold X and calculate the background sound pressure index Zs of the i-th microphone. i Comparing with the background sound pressure threshold X to obtain a first determination result includes:
[0143] When the echo sound pressure index Hs of the i-th microphone i > Background sound pressure threshold X, indicating that the construction site background noise of the i-th microphone is abnormal, triggering the first audio link layer adjustment instruction;
[0144] When the echo sound pressure index Hs of the i-th microphone i When ≤ background sound pressure threshold X, it indicates that the background noise of the construction site of the i-th microphone is normal, the noise interference is within the expected range, and communication and monitoring are continuous.
[0145] The frequency band dynamic gain unit, after receiving the first audio link layer adjustment instruction, sets the microphone channel as the "noise interference focus channel" and divides the audio signal of the i-th microphone into 6 sub-bands. According to the classification result obtained by the noise data classification unit, the current signal-to-noise ratio SNR of each sub-band is compared with its corresponding trigger threshold. If the frequency band corresponds to the trigger threshold, the noise reduction gain of the frequency band is automatically increased to the preset value. If SNR> the corresponding trigger threshold but meets Hs i >X, the noise reduction threshold will be temporarily relaxed by 3dB or the noise reduction gain of the frequency band will be increased by 5dB.
[0146] If the sound pressure in a certain frequency band is detected to exceed the background sound pressure threshold X in real time, the system will automatically increase the noise reduction gain by 5dB in that frequency band or relax the noise reduction trigger threshold by 3dB (making it easier to trigger noise reduction).
[0147] Please refer to the following sub-band dynamic gain unit for dividing the audio at the i-th microphone position into 6 sub-bands), and each sub-band processes information independently, as shown in Table 1.
[0148] Table 1 Example information of Band1-Band6 processing
[0149]
[0150]
[0151] In Table 1, when the signal-to-noise ratio (SNR) of Band1 (20Hz - 80Hz) is detected in real time to be lower than or equal to 20dB, strong noise reduction is triggered (such as reducing the gain by -30dB); for Band6, "suppression is applied when SNR ≤ 5dB", which has a stricter requirement because high frequencies are very important and suppression only starts when the SNR is very low. This "threshold" is the dynamic noise reduction trigger threshold based on the signal-to-noise ratio (SNR), and it is set separately for each frequency band. A high SNR indicates a good signal, and there is no rush to reduce noise; a low SNR means heavy noise, and suppression is needed.
[0152] Among them, the characteristics of the s and t sounds are as follows:
[0153] s sound: Pronunciation position: The tip of the tongue is close to the upper teeth, and air flows rapidly through the gap between the tongue and the upper teeth. Frequency range: Usually between 4tHz - 8tHz. Sound characteristics: This is a very sharp high-frequency sound, similar to the sound of "si". Since it is a sound generated by air friction, its sound quality is very bright, sounding "crisp" and "sharp". Examples: s (si), sh (shi), soft (soft).
[0154] t sound: Pronunciation position: The tip of the tongue touches the upper gum, and then suddenly the tongue is lifted from the gum, releasing the air flow. Frequency range: Usually between 3tHz - 6tHz, and sometimes it extends to higher frequencies. Sound characteristics: A "bursting" effect is produced at the moment of pronunciation, making the timbre "crisp" and "clear".
[0155] Examples: t (te), time (time), tight (tight);
[0156] The s sound often appears at the end of words (such as three, book, shi, book, student). If the s sound is not clear, the meaning of the word may become ambiguous. The t sound sometimes appears at the beginning, in the middle or at the end of words. It is an important part of many English words, such as day, he, head. If the t sound is blurred, the recognition of the whole word will be greatly reduced. In the noise reduction scheme you mentioned, the noise reduction gain settings for Band5 (1tHz - 5tHz) and Band6 (5tHz - 10tHz) are to protect these high-frequency consonants (s and t sounds) as much as possible while reducing environmental noise.
[0157] If too much noise reduction is applied to high-frequency noise, high-frequency sounds such as s and t will also be weakened or blurred, causing the speech to sound unclear or even difficult to understand. In a noisy environment, excessive noise reduction will make the speaker's s and t sounds unclear, and the listener may not be able to accurately identify these sounds, thus affecting the overall communication of the message. We must avoid: excessive noise reduction in these frequency bands to ensure that the high-frequency components of the speech (such as s and t sounds) can still be clearly transmitted. When reducing noise (for example, setting Band 5 to -8dB and Band 6 to -5dB) it can ensure that high-frequency noise is removed without reducing the clarity of the human voice. s and t sounds are high-frequency consonants, and they play a vital role in language, especially in long-distance voice communication. Protecting the clarity of these syllables is critical to understanding voice information.
[0158] In this embodiment, in the background art, construction site noise monitoring often relies on a single sound pressure level or a simple environmental noise threshold, which cannot accurately identify different types of noise sources and has difficulty reflecting the impact of transient abnormal noise. However, this embodiment constructs a multi-dimensional background sound pressure index (including sound pressure level, equivalent continuous sound level, and peak sound level) to achieve a comprehensive quantitative assessment of the construction environment noise status, significantly improving the accuracy and sensitivity of on-site background noise anomaly judgment.
[0159] Traditional noise reduction solutions are mostly static strategies that fail to dynamically adjust parameters based on the real-time acoustic environment, often resulting in over-suppression or insufficient noise reduction. This embodiment proposes a sub-band dynamic gain mechanism that divides the audio signal into six frequency bands. Based on the SNR of each frequency band and the characteristics of specific noise sources, trigger thresholds and noise reduction gain values are set separately. The noise reduction intensity can be flexibly increased or decreased based on the real-time noise situation, achieving targeted suppression of each noise type.
[0160] In noisy environments, traditional noise reduction solutions often oversuppress high-frequency noise, resulting in the loss of high-frequency sub-tones such as "s" and "t" sounds in speech, blurring the voice information and seriously affecting the effectiveness of remote communication. This embodiment sets milder noise reduction gain values (such as -8dB and -5dB) in the Band 5 and Band 6 frequency bands, and combines them with an SNR determination strategy to perform flexible trigger noise reduction. This effectively removes sharp environmental noise while retaining the key high-frequency information of the human voice, significantly improving the clarity and intelligibility of voice communication.
[0161] Background technology lacks the ability to dynamically adjust specific microphone channels in real time, and cannot achieve regionalized focused intervention. This embodiment constructs an independent sound pressure model for each microphone through a sound pressure judgment unit, and determines whether it is a "noise interference key channel". It can target areas severely polluted by noise and strengthen the noise reduction strategy of the channel to improve the overall anti-interference ability and reliability of the system. Traditional noise monitoring is often powerless against short-term, high-intensity burst noises, which may cause auditory interference to personnel or communication failure. This embodiment introduces a peak sound level indicator to accurately identify high-intensity instantaneous noises such as blasting and steel impact, and combines real-time SNR comparison with background threshold detection to trigger a dynamic gain boost or threshold relaxation mechanism to quickly respond to sudden interference sources, thereby effectively suppressing the impact of noise and ensuring voice quality.
[0162] Example 5
[0163] This example is explained in Example 1. Figure 1 ,Specifically, the echo perception module includes an echo feature real-time ,monitoring unit and an echo source classification and identification unit;
[0164] The real-time echo feature monitoring unit is used to collect audio signals from each microphone position in the voice communication system of the building construction site in real time, and extract echo features based on short-time energy analysis and cross-correlation function analysis to obtain the echo path delay, echo intensity, echo duration, echo frequency characteristics, arrival angle of the echo signal, and echo reverberation time of the i-th microphone, and establish an echo source data set;
[0165] The echo source classification and identification unit is used to extract the six-dimensional vector of echo path delay, echo intensity, echo duration, echo frequency characteristics, arrival angle of the echo signal, and echo reverberation time for each detection signal in the echo source data set. After using the support vector machine as a standard sample, an echo recognition model is established to automatically determine the echo source type.
[0166] Echo source types include: ground reflection echo, equipment surface reflection echo, wall hard reflection echo and structural resonance echo.
[0167] The following are examples of echo source classification as shown in Table 2:
[0168]
[0169]
[0170] In this embodiment, the real-time echo feature monitoring unit collects and extracts features from audio signals at each microphone location in real time based on short-time energy analysis and cross-correlation function analysis, enabling comprehensive monitoring of echo path delay, echo intensity, echo duration, echo frequency, echo signal arrival angle, and echo reverberation time. This mechanism provides multi-dimensional echo data, making echo detection more precise and reliable, effectively avoiding the limitations of a single echo feature analysis method, and improving the accuracy and applicability of the monitoring system. Support vector machine (SVM) technology is used to classify echo sources. Based on the six-dimensional echo feature vector extracted from each detection signal, the echo source type can be automatically and accurately identified. This effectively improves the automation level of echo source classification, not only reducing the need for manual intervention but also ensuring rapid and accurate echo source identification in the complex environments of various construction sites. This embodiment introduces four common echo source types: ground reflection echo, equipment surface reflection echo, wall hard reflection echo, and structural resonance echo. By extracting multi-dimensional echo features and classifying and identifying them using support vector machines, this solution can effectively distinguish different types of echo sources and provide accurate data support for the formulation of subsequent echo suppression and optimization strategies.
[0171] Example 6
[0172] This example is explained in Example 1. Figure 1 ,Specifically, the echo perception module also includes an echo interference analysis unit;
[0173] The echo interference analysis unit is used to extract the echo strength Enregy of the i-th microphone in the echo source data set. i , Echo Duration i , Variance of echo path delay i After dimensionless processing, the echo sound pressure index Hs of the i-th microphone is calculated using the following formula: i :
[0174] Hs i =α4Enregy i +α5Duration i +α6Variance i
[0175] Where α4, α5 and α6 are all constants, and α4+α5+α6=1, α4Enregy i +α5Duration i +α6Variance i Indicates that the echo sound pressure index Hs of the i-th microphone is calculated according to the weights α4, α5 and α6 i ;
[0176] Among them, the variance of the delay of the i-th echo path is i Calculated using the following formula:
[0177]
[0178] Where μ i is the echo path delay of the ith microphone, represents the average echo path delay of all microphones, and M represents the total number of microphones.
[0179] The echo perception module further includes an echo anomaly determination unit;
[0180] The echo abnormality determination unit is used to preset the echo threshold Y and calculate the echo sound pressure index Hs of the i-th microphone. i Comparing with the echo threshold Y to obtain a second determination result includes:
[0181] When the echo sound pressure index Hs of the i-th microphone i >Echo threshold Y×120%, indicating that the construction site echo of the i-th microphone is abnormal, generating the first echo abnormality level zone;
[0182] When the echo threshold Y≤the echo sound pressure index Hs of the i-th microphone i When the value is ≤ the echo threshold Y×120%, it indicates that the construction site echo of the i-th microphone is abnormal, and a second abnormal echo level zone is generated, which has a lower echo intensity than the first abnormal echo level zone;
[0183] When the echo sound pressure index Hs of the i-th microphone i When the value is less than the echo threshold Y, it means that there is no echo at the construction site of the microphone, and communication and monitoring are continuous;
[0184] The first echo abnormality level area and the second echo abnormality level area are counted to obtain an echo abnormality group.
[0185] In this embodiment, the echo interference analysis unit extracts the variance of the echo intensity, echo duration, and echo path delay from the i-th microphone, performs dimensionless processing, and combines this with a weighting coefficient to calculate an echo sound pressure index. This metric provides a precise method for quantifying echo intensity, helping to more accurately assess the impact of echoes and, in turn, provide data support for subsequent echo anomaly determination. The echo anomaly determination unit implements multi-level echo anomaly classification by comparing the echo sound pressure index with a preset echo threshold value Y. Based on the relationship between the echo sound pressure index and the echo threshold value, the system automatically determines the echo anomaly level, distinguishing between abnormal areas with higher echo intensity (the first echo anomaly level zone) and abnormal areas with lower echo intensity (the second echo anomaly level zone). This detailed classification helps the system more accurately locate areas with greater echo impact within the construction site and effectively identify the source of abnormal echoes. This embodiment utilizes echo anomaly level classification, which can accurately identify the echo intensity at the construction site based on the echo sound pressure index and process echoes of varying degrees accordingly. The first and second echo anomaly levels represent areas of high and low echo intensity, respectively, enabling the selection of appropriate processing solutions for echo sources of varying strengths. Stronger echo cancellation measures are employed in strong echo areas, while appropriate optimization strategies are employed in weaker echo areas.
[0186] Example 6
[0187] This example is explained in Example 5. Figure 1 ,Specifically, the dynamic sound pickup pattern adjustment module includes an echo angle analysis unit, a first sound pickup pattern switching unit, and a second sound pickup pattern switching unit;
[0188] The echo angle analysis unit is used to extract the first echo abnormality level area and the second echo level area in the echo abnormality group, and obtain the arrival angle DoA of the i-th microphone echo signal by cross-correlation analysis. i :
[0189]
[0190] Where, Δ T represents the propagation delay time difference of the echo signal between the two microphones, d represents the distance between the i-th microphone and the adjacent i+1-th microphone, and c represents the speed of sound; 0 degrees means the signal comes from the front of the microphone array, 90 degrees means the signal comes from the right side of the array, and -90 degrees means the signal comes from the left side;
[0191] The first pickup mode switching unit is configured to generate a second audio link layer adjustment instruction based on the first echo abnormality level area, comprising: a first step of switching to a directional pickup mode and adjusting the pickup direction. Instead of continuing to pick up sound in the direction of the noise, the system turns to the "opposite direction" thereof. For example, if the noise comes from 30 degrees to the left, the system adjusts the pickup direction to 30 degrees to the right to avoid the strong noise source;
[0192] Refer to the following table 3 for an example of adjusting the pickup direction:
[0193]
[0194] Step 2: Dynamically adjust the pickup beam width of the current microphone based on the echo signal strength and the distance between the i-th microphone. The beam width has a nonlinear relationship with the distance:
[0195] When the signal distance is less than one meter, the beam width is expanded to 1.2 times the standard width;
[0196] When the signal distance is greater than or equal to one meter and less than two meters, the standard beam width remains unchanged;
[0197] When the signal distance is greater than two meters and less than or equal to four meters, the beam width is compressed to 75 percent of the standard width;
[0198] When the signal distance is greater than four meters, the beam width is compressed to 50 percent of the standard width;
[0199] Refer to the following example table as shown in Figure 4:
[0200]
[0201]
[0202] The pickup beam direction of the i-th microphone is adjusted to the corresponding target pickup direction, and directional enhancement processing is performed on the current audio frequency band according to the determined beam width.
[0203] The second pickup mode switching unit is configured to generate a third audio link layer adjustment instruction based on the second echo abnormality level area, comprising: a first step of switching the pickup beam mode of the i-th microphone from a directional pickup mode to an omnidirectional pickup mode, i.e., abandoning the pickup focus on a specific direction and no longer limiting the pickup angle range, but receiving audio signals within the full angle range;
[0204] Step 2: Enable echo cancellation mode and generate the corresponding echo suppression ratio for the echo source type, including:
[0205] Eliminate 60%-70% of the ground reflected echo;
[0206] Eliminate 50%-65% of the echo reflected from the device surface;
[0207] For hard-reflected echoes from walls, it can eliminate 71%-80% of the echoes.
[0208] For structural resonance echo, it can eliminate 81%-90% of the echo.
[0209] In this embodiment, the first pickup mode switching unit determines the echo's arrival angle based on the echo abnormality level zone and infers the direction of the speech source, achieving intelligent directional sound pickup. By adjusting the microphone's pickup direction in the opposite direction of the echo, it effectively avoids areas with strong echo sources, improving the quality of target speech. This feature avoids background noise interference and optimizes voice communication quality on construction sites. Especially in complex environments, such as when multiple parties are talking simultaneously or when machinery and equipment generate high noise levels, the system automatically adjusts the pickup direction to maintain a clear voice signal.
[0210] When the signal distance is close (such as 0.5-0.99 meters), the system expands the beam width to capture more sound information; when the signal distance is far (such as 2-4 meters), the beam width is reduced to focus on the long-distance voice signal while reducing noise interference far away from the target; for long-distance signals exceeding 4 meters, the beam width is further compressed to improve directionality and suppress interference from non-speech sound sources.
[0211] Through this dynamic adjustment, the system can always maintain the best pickup effect within different distance ranges, ensuring high-quality collection of voice signals.
[0212] The second pickup mode switching unit further generates a third audio link layer adjustment instruction based on the abnormal echo level. Under certain circumstances, it can switch the microphone's pickup mode from directional to omnidirectional, eliminating concentrated pickup in a specific direction and expanding the audio reception range to effectively cope with environments with multiple audio sources. Omnidirectional pickup mode is particularly effective in environments with multiple audio sources, ensuring balanced reception of audio signals from all directions, ensuring that all voice information on the construction site is captured.
[0213] In terms of echo source elimination and suppression, the system automatically adjusts the echo cancellation and suppression strategy according to different echo source types. Ground reflection, device surface reflection, wall hard reflection and structural resonance echo have different suppression capabilities. The system sets a reasonable echo suppression ratio for different echo sources:
[0214] For ground reflected echo, the system can eliminate 60%-70% of the echo;
[0215] Eliminate 50%-65% of the echo reflected from the equipment surface;
[0216] Eliminate 71%-80% of hard-reflected echoes from walls;
[0217] Eliminate 81%-90% of structural resonance echoes.
[0218] This highly targeted echo cancellation capability can precisely adjust the suppression effect according to the echo source type, thereby achieving the optimal echo processing effect in different echo source environments, further improving the clarity of voice communications at the construction site.
[0219] The system's coordinated echo angle analysis unit, first pickup pattern switching unit, and second pickup pattern switching unit enable it to flexibly adapt to diverse construction environments. In areas with strong echo sources, the system proactively avoids the source and adjusts the beam width to mitigate interference from high-intensity echoes. In complex construction environments, the system automatically adjusts the pickup pattern based on the source and intensity of the echo signal, selecting the most appropriate strategy. This adaptive adjustment significantly improves voice communication at the construction site and reduces communication barriers caused by echo and noise.
[0220] It should be noted that all calculation formulas in this application document utilize, including but not limited to, regression analysis within machine learning algorithms to deeply analyze the collected parameters and identify their natural trends and interrelationships. Professional software, such as Python's Scikit-learn library or the R language, is used to automatically generate mathematical models that match the data. Model performance is then objectively evaluated through methods such as cross-validation, combined with continuous feedback and optimization to ensure that the created formulas truly reflect the inherent laws of the data, thereby guaranteeing their validity and accuracy, and ensuring that the calculation process complies with the constraints of natural laws rather than being based on artificially set rules.
[0221] The technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0222] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0223] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
[0224] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A long-distance voice communication system suitable for building construction environment, characterized in that: include: The communication deployment module is used to deploy multiple microphone arrays at the building construction site. Each construction worker is equipped with a microphone to form multiple microphone nodes. These microphone nodes are connected via Zigbee technology. The module collects and processes the signals from the multi-microphone array and outputs the direction and distance of the voice source. The construction site noise sensing module is used to monitor the background noise data of the building construction site in real time, identify and automatically classify the noise source, construct a background sound pressure index, and determine whether there is noise interference in the environment based on the background sound pressure index. If so, it triggers the first audio link layer adjustment instruction; The echo perception module is used to monitor the echo characteristics of the i-th microphone signal in the voice communication system of the building construction site in real time, identify and automatically classify the echo source type, and construct the echo sound pressure index Hs of the i-th microphone. i , to determine whether there is an abnormal echo phenomenon in the current communication link. If so, obtain the abnormal echo group; Dynamic pickup pattern adjustment module, used to obtain the arrival angle DoA of the i-th microphone echo signal through cross-correlation analysis based on the echo anomaly group i , and generates corresponding second audio link layer adjustment instructions and third audio link layer adjustment instructions.
2. A long-distance voice communication system suitable for a building construction environment according to claim 1, characterized in that: The communication arrangement module includes a building three-dimensional modeling unit, a microphone array arrangement unit and a central processing unit; The building 3D modeling unit is used to scan the target building with a laser scanner to obtain 3D point cloud data of the building's exterior and interior, and to construct a 3D building model using 3D modeling software such as Revit, SketchUp, or Navisworks. The unit also marks different construction areas in the 3D building model, including high-altitude work areas, ground work areas, and areas surrounding construction equipment. The microphone array arrangement unit is used to arrange multiple microphone arrays according to the spatial layout of the construction site and the needs of voice communication. The multiple microphone nodes are interconnected through Zigbee technology to form an ad hoc network structure, which can transmit the collected voice signals to the central processing unit in real time; The central processing unit is used to capture the sound waves of each microphone and convert them into electrical signals, obtain the voice signal of the i-th microphone, and build a microphone array located in the three-dimensional model of the building. Each microphone m1, m2, ..., m M , M represents the total number of microphones; the position vector (x i ,y i ,z i ); The ambient sound waves captured by each microphone are converted into electrical signals. The central processing unit is responsible for centrally processing the signals collected by all microphones. The signal received by the i-th microphone Expressed as: Where, is the signal received by the i-th microphone at time j, and A is the signal amplitude; f(r i ,θ,φ,j) is the waveform signal of the sound signal emitted from the speech source (θ,φ,r) and propagated to the i-th microphone, (x i ,y i , z i ) is the position vector of the i-th microphone, θ and φ are the azimuth and elevation angles of the sound source, respectively, r i is the distance from the sound source to the i-th microphone, j represents time. Considering that the propagation of sound waves is dynamic, the sound signal will change over time; is the interference signal from the noise at time t.
3. The long-distance voice communication system suitable for a building construction environment according to claim 2, characterized in that: The construction site noise perception module includes a noise collection unit and a noise data classification unit; The noise collection unit is used to collect noise data from the i-th microphone using a noise monitor, establish a noise data set, and process the noise data. The noise data classification unit extracts noise features and then classifies the noise. The classification results include: The frequency of low-frequency mechanical noise of tower cranes is 20Hz to 80Hz; The mechanical noise frequency of compressor and air pump equipment is 20Hz to 80Hz; The frequency of low-frequency mechanical noise of concrete mixers is 80Hz to 150Hz; The frequency range of low-frequency noise from transport vehicles is 80Hz to 150Hz; The frequency range of low-frequency mechanical noise of excavators is 80Hz to 500Hz; The frequency range of human knocking noise is 150Hz to 500Hz; The frequency range of manual handling noise is 500Hz to 5kHz; The frequency range of chainsaw noise is 1kHz to 6kHz; The frequency range of electric drill noise is 1kHz to 10kHz.
4. The long-distance voice communication system suitable for a building construction environment according to claim 3, characterized in that: The construction site noise perception module also includes an on-site sound pressure analysis unit, a sound pressure judgment unit and a frequency band dynamic gain unit; The on-site sound pressure analysis unit is used to construct the background sound pressure index Zs of the i-th microphone based on the noise data set. i , the specific steps are: S111. Each construction worker wears a microphone. There is noise in the i-th microphone environment. The sound pressure Px at the i-th microphone environment measurement point is collected. i , calculate the sound pressure level Lp of the i-th microphone using the following formula i : Where P0 is the reference sound pressure, which is set to 20×10-6Pa; S112, and collect the average noise level of the environment in the sampling period T, and calculate the equivalent continuous sound level Leq of the i-th microphone using the following formula: i : Where T is the sampling time period, P i 2 (j) is the square of the instantaneous sound pressure, P i (j) represents the instantaneous sound pressure of the noise signal of the i-th microphone at time j; represents the integration of the square of the instantaneous sound pressure within the sampling period T, which represents the total energy of the noise; dj represents the time differential, which is integrated over time j; S113, and collect the maximum instantaneous sound pressure level of the environment within the sampling period T, which is used to measure those noise events with large instantaneous intensity, reflect the suddenness and instantaneous maximum loudness, and calculate the first The peak sound level Lpeak of each microphone i : Where, P peak,i Indicates the value of the maximum instantaneous sound pressure of the noise signal; S114: Combine the sound pressure level Lp of the i-th microphone obtained in S111-S113 i , equivalent continuous sound level Leq i and peak sound level Lpeak i After dimensionless processing, the background sound pressure index Zs of the i-th microphone is calculated by the following formula: i : Zs i =α1Lp i +α2Leq i +α3Lpeak i Where α1, α2 and α3 are all constants, and α1+α2+α3=1, α1Lp i +α2Leq i +α3Lpeak i Indicates that the background sound pressure index Zs of the i-th microphone is calculated according to the weights α1, α2 and α3 i ; The sound pressure judgment unit is used to preset the background sound pressure threshold X and calculate the background sound pressure index Zs of the i-th microphone. i Comparing with the background sound pressure threshold X to obtain a first determination result includes: When the echo sound pressure index Hs of the i-th microphone i > Background sound pressure threshold X, indicating that the construction site background noise of the i-th microphone is abnormal, triggering the first audio link layer adjustment instruction; When the echo sound pressure index Hs of the i-th microphone i When ≤ background sound pressure threshold X, it indicates that the background noise of the construction site of the i-th microphone is normal, the noise interference is within the expected range, and communication and monitoring are continuous.
5. The long-distance voice communication system suitable for a building construction environment according to claim 4, characterized in that: The frequency band dynamic gain unit, after receiving the first audio link layer adjustment instruction, sets the microphone channel as the "noise interference focus channel" and divides the audio signal of the i-th microphone into 6 sub-bands. According to the classification result obtained by the noise data classification unit, the current signal-to-noise ratio SNR of each sub-band is compared with its corresponding trigger threshold. If the frequency band corresponds to the trigger threshold, the noise reduction gain of the frequency band is automatically increased to the preset value. If SNR> the corresponding trigger threshold but meets Hs i >X, the noise reduction threshold will be temporarily relaxed by 3dB or the noise reduction gain of the frequency band will be increased by 5dB.
6. The long-distance voice communication system suitable for a building construction environment according to claim 1, characterized in that: The echo perception module includes an echo feature real-time monitoring unit and an echo source classification and identification unit; The real-time echo feature monitoring unit is used to collect audio signals from each microphone position in the voice communication system of the building construction site in real time, and extract echo features based on short-time energy analysis and cross-correlation function analysis to obtain the echo path delay, echo intensity, echo duration, echo frequency characteristics, arrival angle of the echo signal, and echo reverberation time of the i-th microphone, and establish an echo source data set; The echo source classification and identification unit is used to extract the six-dimensional vector of echo path delay, echo intensity, echo duration, echo frequency characteristics, arrival angle of the echo signal, and echo reverberation time for each detection signal in the echo source data set. After using the support vector machine as a standard sample, an echo recognition model is established to automatically determine the echo source type. Echo source types include: ground reflection echo, equipment surface reflection echo, wall hard reflection echo and structural resonance echo.
7. The long-distance voice communication system suitable for a building construction environment according to claim 6, characterized in that: The echo perception module also includes an echo interference analysis unit; The echo interference analysis unit is used to extract the echo strength Enregy of the i-th microphone in the echo source data set. i , Echo Duration i and the variance of the echo path delay i After dimensionless processing, the echo sound pressure index Hs of the i-th microphone is calculated using the following formula: i : Hs i =α4Enregy i +α5Duration i +α6Variance i Where α4, α5 and α6 are all constants, and α4+α5+α6=1, α4Enregy i +α5Duration i +α6Variance i Indicates that the echo sound pressure index Hs of the i-th microphone is calculated according to the weights α4, α5 and α6 i ; Among them, the variance of the delay of the i-th echo path is i Calculated using the following formula: Where μ i is the echo path delay of the ith microphone, represents the average echo path delay of all microphones, and M represents the total number of microphones.
8. The long-distance voice communication system suitable for a building construction environment according to claim 6, characterized in that: The echo perception module further includes an echo anomaly determination unit; The echo abnormality determination unit is used to preset the echo threshold Y and calculate the echo sound pressure index Hs of the i-th microphone. i Comparing with the echo threshold Y to obtain a second determination result includes: When the echo sound pressure index Hs of the i-th microphone i >Echo threshold Y×120%, indicating that the construction site echo of the i-th microphone is abnormal, generating the first echo abnormality level zone; When the echo threshold Y≤the echo sound pressure index Hs of the i-th microphone i When the value is ≤ the echo threshold Y×120%, it indicates that the construction site echo of the i-th microphone is abnormal, and a second abnormal echo level zone is generated, which has a lower echo intensity than the first abnormal echo level zone; When the echo sound pressure index Hs of the i-th microphone i When the value is less than the echo threshold Y, it means that there is no echo at the construction site of the microphone, and communication and monitoring are continuous; The first echo abnormality level area and the second echo abnormality level area are counted to obtain an echo abnormality group.
9. The long-distance voice communication system suitable for a building construction environment according to claim 8, characterized in that: The dynamic sound pickup pattern adjustment module includes an echo angle analysis unit, a first sound pickup pattern switching unit, and a second sound pickup pattern switching unit; The echo angle analysis unit is used to extract the first echo abnormality level area and the second echo level area in the echo abnormality group, and obtain the arrival angle DoA of the i-th microphone echo signal by cross-correlation analysis. i : Where, Δ T represents the propagation delay time difference of the echo signal between the two microphones, d represents the distance between the i-th microphone and the adjacent i+1-th microphone, and c represents the speed of sound; 0 degrees means the signal comes from the front of the microphone array, 90 degrees means the signal comes from the right side of the array, and -90 degrees means the signal comes from the left side; The first pickup mode switching unit is configured to generate a second audio link layer adjustment instruction according to the first echo abnormality level area, The system switches to directional pickup mode and adjusts the pickup direction. Instead of continuing to pick up the sound in the direction of the noise, it turns to the "opposite direction" of the noise. For example, if the noise comes from 30 degrees to the left, the system adjusts the pickup direction to 30 degrees to the right to avoid the strong noise source. Step 2: Dynamically adjust the pickup beam width of the current microphone based on the echo signal strength and the distance between the i-th microphone. The beam width has a nonlinear relationship with the distance: When the signal distance is less than one meter, the beam width is expanded to 1.2 times the standard width; When the signal distance is greater than or equal to one meter and less than two meters, the standard beam width remains unchanged; When the signal distance is greater than two meters and less than or equal to four meters, the beam width is compressed to 75 percent of the standard width; When the signal distance is greater than four meters, the beam width is compressed to 50 percent of the standard width; The pickup beam direction of the i-th microphone is adjusted to the corresponding target pickup direction, and directional enhancement processing is performed on the current audio frequency band according to the determined beam width.
10. The long-distance voice communication system suitable for a building construction environment according to claim 9, characterized in that: The second pickup mode switching unit is configured to generate a third audio link layer adjustment instruction based on the second echo abnormality level area, comprising: a first step of switching the pickup beam mode of the i-th microphone from a directional pickup mode to an omnidirectional pickup mode, i.e., abandoning the pickup focus on a specific direction and no longer limiting the pickup angle range, but receiving audio signals within the full angle range; Step 2: Enable echo cancellation mode and generate the corresponding echo suppression ratio for the echo source type, including: Eliminate 60%-70% of the ground reflected echo; Eliminate 50%-65% of the echo reflected from the device surface; For hard-reflected echoes from walls, it can eliminate 71%-80% of the echoes. For structural resonance echo, it can eliminate 81%-90% of the echo.