Dual master audiovisual collaboration endpoint and method for hearing protection earmuffs

CN122845984APending Publication Date: 2026-09-29GUANGZHOU OPSMEN TECH CO LTD
View PDF 16 Cites 0 Cited by

Patent Information

Application Number
CN202611147549.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0007]上述路线的不足在于:影像采集件与听力防护件在头盔平台上分体挂载、各自成链,供电、计时与存储彼此独立,佩戴负荷与接口占用成倍增加,声学上发生的关键事件与影像中的对应片段只能事后人工比对;影像采集件以粘合或滑轨方式固定于头盔,其取景方向不随头部转动而需另设旋转装置调节,无法真实反映作业人员的第一视角;两件文献中的拾音与降噪模块均仅作用于音频通道本身,未公开由声音事件触发影像采集,视频通路与音频通路虽分别经无线局域网与蓝牙接入电台,但两者的处理器之间亦未公开任何控制与状态接口及任务分工

Benefits of technology

[0037]1.真正的一体化。在同一符合听力防护要求的耳罩上同时完成听力防护、战术通信与第一视角影像采集,取景方向随头部转动,无需外挂独立记录设备,佩戴负荷、供电与时基统一,声学事件与影像片段由同一事件标识天然对应;同一对镜像对称的扩展接口按所读取的模组标识分别路由至音频链路与显示链路,使送话麦克风模组可依用手习惯或接口占用情况左右互换接入、显示模组可按需单目或双目装配,且全部扩展均不贯穿隔声腔体。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845984A_ABST
    Figure CN122845984A_ABST
Patent Text Reader

Abstract

This invention relates to the field of hearing protection and wearable sensing device technology, specifically to a dual-master control audio-visual collaborative terminal and method for hearing protection earmuffs. The method involves an audio master controller and a visual master controller with independent power domains, connected by an internal cross-master control link, on an earmuff with a soundproof cavity. The audio master controller operates the hearing protection audio link and connects to a first wireless communication circuit. The visual master controller encodes and stores a first-view image stream and connects to a second wireless communication circuit with a higher speed. The audio master controller uses the amplitude-limited sound pressure level detection result to determine acoustic events. Acoustic event information carrying an event identifier is sent via the cross-master control link to trigger the visual master controller to change its acquisition, encoding, or storage state. This identifier is written into the image derived data. The image derived data is transmitted via the cross-master control link to the first wireless communication circuit. The second wireless communication circuit is only activated upon receiving a retrieval command carrying the identifier to transmit the corresponding image segment. Hearing protection communication and first-view perception are integrated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of individual hearing protection and wearable sensing device technology, specifically to a dual-master audiovisual collaborative terminal and method for hearing protection earmuffs. Background Technology

[0002] In high-noise work environments such as shooting training, law enforcement operations, engineering blasting, airport ground support, and metal processing, personnel must wear hearing protection earmuffs that meet standards such as ANSIS 3.19 and EN 352. The core of these earmuffs is a soundproof cavity that isolates the outer ear from ambient noise. An ambient sound pickup unit is located outside the cavity, while a speaker unit is inside. The processing circuitry performs gain control and impact sound limiting on the picked-up ambient sound signal, adapting to sound pressure level changes, before playback through the speaker. This allows the wearer to clearly hear low-level commands and environmental cues while being automatically protected against transient high sound pressure level impacts such as gunshots, hammer blows, and explosions. In recent years, these earmuffs have further integrated walkie-talkie communication functions, forming a head-mounted terminal that combines hearing protection with tactical communication.

[0003] Meanwhile, first-person perspective footage of operational processes is increasingly becoming essential material for training debriefing, law enforcement evidence collection, and accident investigation. Several technical approaches have emerged in the existing field, but all have shortcomings.

[0004] Route 1: Mount the image acquisition device on the helmet platform, separating it from the hearing protection device into a chain.

[0005] CN113303540A discloses a communication helmet with noise reduction function, which specifies that the helmet is equipped with an earphone, a microphone and a communication device. An airbag and an inflation / deflation device are provided between the earcups and the shell of the earphone. By inflating or deflating the airbag, the earcups can move inward or outward to switch the sealing state. Its dependent claims further disclose that a video recorder and a display are provided on the helmet and are respectively connected to the communication device by circuit, and a rotating device is provided on the helmet for adjusting the shooting angle of the video recorder.

[0006] CN212969889U discloses a digital individual soldier intelligent combat communication system, which includes a self-organizing network communication radio, an intelligent display control terminal, a video acquisition device, and a receiving and transmitting device. The video acquisition device is wirelessly connected to the self-organizing network communication radio via a wireless local area network, and the receiving and transmitting device is wirelessly connected to the self-organizing network communication radio via Bluetooth. Its dependent claims further disclose that the video acquisition device is installed on a tactical helmet by adhesive or sliding rail, and the receiving and transmitting device is fixed to both sides of the helmet, covering the outer ears when in use, and is equipped with a sound pickup module to amplify small sounds and a noise reduction module to suppress sharp sounds.

[0007] The shortcomings of the above approach are as follows: the image acquisition device and the hearing protection device are separately mounted on the helmet platform and each is linked independently. The power supply, timing and storage are independent of each other, which increases the wearing load and interface occupation many times over. The key acoustic events and the corresponding segments in the images can only be compared manually afterward. The image acquisition device is fixed to the helmet by adhesive or sliding rail. Its framing direction does not follow the head rotation and requires a separate rotation device for adjustment, which cannot truly reflect the first-person perspective of the operator. The sound pickup and noise reduction modules in the two documents only act on the audio channel itself and do not disclose that the image acquisition is triggered by sound events. Although the video path and audio path are connected to the radio via wireless LAN and Bluetooth respectively, no control and status interface or task division between the processors of the two is disclosed.

[0008] Route 2: Add a camera directly to the earcup itself, but use a single processing unit or two main control boards to operate independently.

[0009] CN206212226U discloses a noise-canceling camera headset, including a headband, left and right earpieces, and a data charging cable. Speakers and microphones are respectively located inside the earpiece covers on the left earpiece. A rechargeable lithium battery and a left earpiece main control board with noise-canceling circuitry are located inside the left earpiece cover. A right earpiece main control board is located inside the right earpiece cover, and a camera is mounted on the cover. The right earpiece main control board is connected to a memory card slot, and a universal serial bus port is located on one side of the camera. The camera is manually started and stopped by a camera power switch and has both photo and video recording modes.

[0010] CN117156334A discloses a head-mounted noise-isolating headset equipped with a visible light communication device. Its claims define it as including five visible light receivers, a control chip, a camera, two noise-isolating earmuffs, two built-in speakers, a microphone, a headband, and a light-emitting diode light signal emitting device; wherein the camera is located on the front of the control chip module box on the top of the head and faces the same direction as the head, and is used to allow managers to view the first-person perspective of the workers.

[0011] US7707035B2 discloses an autonomous integrated head-mounted device and sound processing system for tactical operations, the claims of which define the following: a head-mounted housing that isolates the user's ears from direct exposure to ambient sound; a microphone for picking up ambient sound; a microphone for picking up the user's voice; a processing unit that provides sound filtering and amplification, as well as speech and non-speech sound analysis and recognition, and has language recognition and two-way speech translation functions; an internal speaker; an external speaker for playing translations to others; a light sensor that provides incident light frequency and intensity; a position-variable head-mounted camera (the processing unit issues instructions to the camera in response to the user's voice commands); and interconnection for connection to a target designation system, a communication network, and a radio transmitter; the dependent claims further define the processing unit as providing non-speech sound recognition, including vehicle noise recognition.

[0012] The shortcomings of the above approaches are as follows: CN117156334A and US7707035B2 both use a single control chip or a single processing unit to handle audio and video. However, image signal processing and video encoding are typical bursty high-computing loads, which will consume a large amount of processing resources, bus bandwidth and interrupt response. The impact noise limiting of hearing protection must be completed deterministically within milliseconds. Its processing cycle will be jittered due to the camera business, and the wearer will not receive the necessary protection under impact noise. Although CN206212226U has two main control boards on the left and right and places the noise reduction circuit and camera on the two sides, it does not disclose any control and status channels or task division between the two main control boards. The two operate independently, and the images are only stored locally in the memory card and retrieved via the universal serial bus. No wireless external transmission link is disclosed. Furthermore, although US7707035B2 discloses the analysis and recognition of non-voice sounds, the purpose of its recognition results is only to prompt the wearer when the database is matched. The control source of the camera is the user's voice command, and it does not disclose the automatic start or enable of the imaging subsystem based on the recognition results of ambient sound events.

[0013] Route 3: Set up two processors for division, but the images and their previews are all sent directly to the outside via the high-speed wireless circuit built into the visual side.

[0014] US10129919B2 discloses a video headset whose Bluetooth module includes an audio processor for call audio processing and acoustic echo cancellation. The camera module is a system-on-a-chip containing a video processor and a system processor, and has a built-in wireless LAN interface. The two are interconnected via a four-wire bidirectional serial port that runs through the neckband. This serial port is described in the specification as a control and status interface, through which camera configuration and control messages are sent and camera status messages are returned. The specification also states that the reliable data transmission rate required by this interface is only a few kilobits per second. The specification further states that the wireless LAN interface is only enabled when video streaming is required, and its activation or deactivation is determined by the video start and stop operations on the button or application interface.

[0015] US20180048750A1 and its sibling US20220337693A1 disclose an audio-visual wearable computer system with an integrated projector. In the earphone embodiment disclosed in the specification, the left earcup houses a wireless LAN processor, a wireless LAN chipset, a camera, a microphone for the camera, and a power management integrated circuit card. The right earcup houses a Bluetooth processor, a Bluetooth transceiver, a battery, a voice microphone, and a wind noise-canceling microphone. The two processors are connected to a microcontroller for coordinating the operation of the wireless LAN and Bluetooth via an integrated circuit bus and can communicate directly via a universal asynchronous receiver / transmitter protocol. The earphone has a built-in web server that encodes the camera preview frames and provides a viewfinder preview to the mobile device via the wireless LAN.

[0016] US11115928B2 discloses a method for a wearable device to initiate a handshake. The claims define the wearable device as receiving a new data query via a low-power wireless connection, responding to the query by sending a new data message identifying the data stored in the wearable device to a client device via the low-power wireless connection, receiving a connection communication instruction from the client device via the low-power wireless connection to activate a high-speed wireless circuit, and transmitting the stored data to the client device via the high-speed wireless connection. The device claims define the above division of labor in the form of a low-power circuit and a high-speed circuit, and the specification states that the high-speed circuit is automatically powered off when not in use.

[0017] US12028300B2 discloses a method and system for sending an image after a thumbnail is selected. The claims specify that after a first terminal generates a first image, it first sends a message containing only the thumbnail of the image to a second terminal via a near-field connection. The second terminal displays the thumbnail in a thumbnail notification box. The first terminal only sends the first image to the second terminal after the second terminal detects the user's operation on the thumbnail and sends back a corresponding message. The specification cites the near-field connection as a wireless LAN point-to-point connection or a Bluetooth connection as an example.

[0018] The shortcomings of the above-mentioned routes are as follows: their internal interconnect links do not carry derived data generated from images. The four-wire bidirectional serial port of US10129919B2 only supports a transmission rate of several kilobits per second, which is physically insufficient to carry images. The entire video is transmitted externally through the camera module's own wireless LAN interface. The preview frames of US20180048750A1 are directly output from the headset's built-in web server via wireless LAN, and its cross-earc link is responsible for power supply, analog audio, and integrated circuit bus control. Therefore, as long as the user needs to view the image, the high-speed wireless circuit on the visual side must be online for a long time. The soundproof cavity of hearing protection earmuffs is a sealed cavity designed to isolate noise and fits tightly to the head. Its heat dissipation conditions are far worse than those of open-back headbands. The continuous power consumption and heat generation of high-speed radio frequency are concentrated here, directly leading to a sharp reduction in battery life, an increase in shell temperature, and even forcing designers to make holes in the soundproof cavity for heat dissipation, sacrificing the protection level. Although US11115928B2 discloses that the high-speed circuit is not working by default and is powered on by the low-power side according to the instruction, its low-power side does not undertake any audio functions, let alone the hearing protection audio link. The low-power link transmits signaling that identifies the stored data rather than derived data generated from the image. The thumbnail first and the original image later in US12028300B2 occurs between two independent terminals rather than between two main controllers in the same device. Moreover, its thumbnail and the original image use the same near-field connection, and it does not disclose the distinction between the low-speed and high-speed links.

[0019] Route 4: Trigger image action with sound detection, but set up a separate detection pickup channel; while the existing acoustic analysis results on the hearing protection side are never sent to the image subsystem.

[0020] CN119521064A discloses an earphone system, including a sound acquisition module, a camera module, and a processing module disposed on the earphone; the sound acquisition module is used to acquire ambient sound data of the space where the earphone wearer is located, the camera module is used to take pictures under the control of the processing module, and the processing module is used to determine the current scene where the earphone wearer is located based on the ambient sound data, and to wake up the camera module to take pictures if the current scene is a preset target scene.

[0021] US10939066B2 discloses a wearable camera that can connect to an external server. The controller determines whether the sound captured by the microphone is a predetermined sound. If it is a predetermined sound, video recording is initiated, sound data is generated based on the captured sound, and sent to the external server. When the sound data matches the sound stored in its memory, the external server determines a corresponding emergency and sends a notification to the camera. The controller responds to the notification by streaming video to the external server. The dependent claims specify that the sound stored on the server is at least one of gunshots, the sound of a person falling to the ground, or an explosion. The specification states that the microphone is a built-in microphone housed within the camera housing or a wireless microphone wirelessly connected to it. The camera also states that it cyclically overwrites video footage into a buffer memory of random access memory to retain footage for a predetermined time before receiving the notification.

[0022] US11862189B2 discloses a device for performing sound detection, which defines a first stage of a target sound detector containing a binary target sound classifier and activates a second stage when a target non-speech sound is detected. The second stage receives audio data from a buffer and generates a user interface signal to an output device to output a visual representation. Its dependent claims further define that the binary target sound classifier and the buffer are located in a low-power domain and operate in a normally-on mode, and that the first stage activates a camera when a target non-speech sound is detected. The specification states that the purpose of activating the camera is to detect the environment in which the device is located in order to improve the target sound detection effect.

[0023] US9549273B2 discloses an apparatus for selectively enabling a microphone circuit, wherein a processor of the microphone circuit performs sound detection based on a microphone signal, generates an enable signal based on the sound detection result, and sends the enable signal to enable a codec, and its dependent claims further define that the processor determines whether to generate an application processor enable signal based on the sound detection.

[0024] In contrast to the above is CN117242784A, which discloses a personal protective equipment device including a speaker configured to provide a modified sound to a user, a microphone configured to capture an ambient sound stream, a sound analyzer that receives the ambient sound stream from the microphone and identifies a first sound in the ambient sound, and a sound processor that applies a model to the first sound based on the sound recognition to obtain a modified first sound, the model altering the characteristics of the identified sound and the modified sound being provided to the speaker; its dependent claims define the device as a hearing protection device that provides level-dependent hearing protection, and that it receives ambient sound through a microphone positioned externally to the device.

[0025] The shortcomings of the above approach are that the evidence triggering for sudden acoustic events is neither timely nor reliable, while the high dynamic range sound pressure detection capability inherent in hearing protection itself is not utilized. The cameras in CN206212226U and CN117156334A require manual operation, while impact events such as gunshots and explosions occur on the order of milliseconds, meaning manual operation is inevitably delayed beyond the event itself. US10939066B2 uses a built-in ordinary microphone housed within the camera housing for detection. Its dynamic range is designed for speech scenarios, and under impact sounds at the 140 dB sound pressure level, it will inevitably experience clipping saturation, with the rising edge flattened and the spectrum contaminated by harmonics. Criteria based on this are prone to both missed detections and false alarms, and the initiation of recording requires external server verification before notification is issued. Both US11862189B2 and US9549273B2 have separately set up normally open sound detection circuits. The former is a binary target sound classifier and buffer located in the low power domain, while the latter is a processor in the microphone circuit, which increases the number of devices, static power consumption and the area occupied by the outer wall of the cavity. In addition, the purpose of activating the camera in US11862189B2 is to detect the environment in which the device is located in order to improve the sound detection effect, while the target of US9549273B2 is the codec and application processor. Neither of them discloses that the sound detection results are used for image acquisition, storage and transmission. While CN119521064A discloses a method for determining the scene based on ambient sound data in headphones and activating the camera module accordingly, it only has a single processing module. Furthermore, the headphones lack level-dependent attenuation and impact sound limiting links for hearing protection. Therefore, scene determination relies solely on the output of the sound acquisition module, making it impossible to reuse the sound pressure level detection results obtained during the limiting process. It also fails to disclose how image-derived data is transmitted via an internal link to a low-power wireless circuit controlled by another main controller, how a high-speed link is activated upon retrieval, and how the power supply domains are isolated. While CN117242784A discloses a hearing protection device that simultaneously provides level-dependent hearing protection and ambient sound recognition, the sole purpose of its recognition results is to select and apply a model to alter the acoustic characteristics of the sound before transmitting it to the wearer via an internal speaker. The entire device contains no camera or image acquisition components and does not disclose the external transmission of recognition results to an image subsystem.

[0026] Route 5: The motivation for power-off control of the processor and wireless circuit is to save power, without considering them as fault domains that need to be isolated from each other.

[0027] US11115928B2's high-speed circuit automatically shuts off when not in use; US20180048750A1 selectively enables the high-power wireless LAN side from the Bluetooth side; and US10129919B2's wireless LAN interface is only enabled during video streaming. The motivation behind these power-offs or inactivations is to reduce power consumption. However, their shortcomings lie in the fact that none of them treat the two processors as fault domains requiring mutual isolation, nor do they disclose the availability of hearing protection when the visual side malfunctions. US20180048750A1 only discloses a single power management integrated circuit card, with the battery located on one earcup and powering the other side via a cross-earcup power cable. Therefore, if the video side malfunctions due to a crash, overcurrent, or abnormal reset, the audio side is often affected through the shared power supply and interconnection interface, causing the critical hearing protection function to be interrupted in high-noise environments, posing a safety risk.

[0028] In addition, regarding degradation strategies under high-temperature operating conditions.

[0029] CN110933294B discloses an image processing method for a terminal with a camera module. It specifies acquiring a power mode switching command triggered manually or via voice input, switching the power mode to a lower power mode in response to the command, and acquiring scene types including daytime movement, daytime stillness, nighttime movement, and nighttime stillness. Based on the scene type, it determines the corresponding first power mode and controls the camera module to take pictures accordingly. This first power mode includes image quality reduction processing on the preview image and / or the captured image. Its dependent claims further disclose that the processing strategy includes reducing the image resolution, size, output frame rate, or lowering color, saturation, sharpness, contrast, etc. Its shortcomings are: the degradation action only involves image quality, does not include external transmission link actions in the degradation sequence, and does not disclose completing file storage and integrity processing before stopping image acquisition. Therefore, under high-temperature conditions, it is difficult to balance the continuity of the evidence collection task with the integrity and verifiability of the recorded files.

[0030] In summary, existing technologies lack a terminal architecture and collaborative method that can truly integrate hearing protection communication and first-person perspective intelligent perception on a strictly limited platform such as earmuffs that meet hearing protection requirements, and enable the two to coexist without harming each other. Summary of the Invention

[0031] To address the aforementioned shortcomings, the technical problem this invention aims to solve is: how to integrate hearing protection communication with first-person perspective intelligent sensing on an earmuff platform where the soundproof cavity is indestructible, heat dissipation and power consumption are strictly limited, and hearing protection is a critical safety function. This would ensure that the acquisition, evidence collection, and retrieval of first-person perspective images do not interfere with the real-time performance of hearing protection, do not introduce unacceptable power consumption and temperature rise, can be triggered promptly and reliably by sudden acoustic events, and do not affect hearing protection when anomalies occur on the visual side.

[0032] This technical problem can be broken down into four interrelated sub-problems: how to ensure that the hearing protection audio link and the first-person view image link coexist on the same earmuff platform without competing with each other; how to continuously provide usable image information and access points to external devices without keeping the high-speed wireless circuit online for extended periods; how to reliably and timely trigger the first-person view image evidence collection action without increasing the number of sound pickup channels and without being affected by impact sound clipping; and how to ensure that any abnormality on the visual side does not cause the hearing protection function to be interrupted.

[0033] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is to provide a dual-master audiovisual collaborative terminal for hearing protection earmuffs, including a left earmuff and a right earmuff each having a soundproof cavity, a headband connecting the left earmuff and the right earmuff, an ambient sound pickup unit and a first-view camera module disposed outside the soundproof cavity, and a speaker unit disposed inside the soundproof cavity; it also includes an audio master controller and a visual master controller with independent power domains and an internal cross-master control link connecting the two; the audio master controller operates the hearing protection audio link, limits the ambient sound signal output by the ambient sound pickup unit, plays it back through the speaker unit, and connects to a first wireless communication circuit; the visual master controller encodes and stores the image stream output by the first-view camera module, and connects to a wireless communication circuit with a transmission rate higher than the first wireless communication circuit. The second wireless communication circuit of the communication circuit; the audio master controller reuses the limited sound pressure level detection result to determine acoustic events without setting up a separate microphone for this determination, and sends the acoustic event information carrying the event identifier to the visual master controller via the cross-master control link, triggering the visual master controller to change the acquisition parameters, encoding parameters or storage state of the image stream, and writes the event identifier into image derived data generated by the image stream with a data volume smaller than that of the image stream; the image derived data is sent to the audio master controller via the cross-master control link and transmitted by the first wireless communication circuit, and the second wireless communication circuit is activated after the audio master controller receives the image retrieval instruction carrying the event identifier and forwards it via the cross-master control link, for transmitting the image segment corresponding to the event identifier.

[0034] The present invention also provides a dual-master audiovisual collaboration method for hearing protection earmuffs, wherein steps S1 to S6 correspond to the working process of the aforementioned terminal.

[0035] Further technical solutions involve: independent voltage regulators and independent interrupt controllers for the two main controllers; level isolation and operational status monitoring across the main control link; multi-functional multiplexing of the microphone array and priority preemption of the image-derived data forwarding task for real-time audio tasks; impact sound criteria; encrypted ring pre-recording buffer and non-deletable markers at the acquisition start point; time-division carrying and priority across the main control link; specific forms of the first and second wireless communication circuits; telescopic brackets; mirror-symmetrical expansion interfaces and interface locking seats; arrangement of the first-view camera module; selective plugging and left / right swapping of the microphone module and HUD module using the same pair of mirror-symmetrical expansion interfaces; switching the audio pickup source of the communication audio link to the microphone module or outputting display data to the HUD module and determining the left, right, or binocular configuration according to the read module identifier; encrypted event files triggered by multiple sources such as acoustic events, key inputs, voice commands, or remote commands and sharing the same format of event identifiers; and temperature control graded degradation sequence.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1. True integration. Hearing protection, tactical communication, and first-person view image acquisition are all completed simultaneously on the same earmuff that meets hearing protection requirements. The framing direction changes with head movement. No external independent recording equipment is required. Wearing load, power supply, and time base are unified. Acoustic events and video clips are naturally associated with the same event identifier. The same pair of mirror-symmetrical expansion interfaces are routed to the audio link and display link respectively according to the read module identifier. This allows the microphone module to be swapped left and right according to hand habits or interface occupancy. The display module can be equipped with monocular or binocular lenses as needed. Moreover, all expansions do not penetrate the soundproof cavity.

[0038] 2. The real-time nature of hearing protection is ensured. The burst computing power of image signal processing and video encoding is isolated within the visual master control, and does not enter the interruption and computing power budget of the audio master control; in conjunction with the priority preemption of image-derived data forwarding tasks by the hearing protection task, the processing cycle of impact sound amplitude limiting is not affected by camera operations.

[0039] 3. Reduced power consumption and temperature rise. Daily screen browsing and control interactions reuse the low-power link already connected on the audio side. The second wireless communication circuit changes from being constantly online to being event-driven and momentarily online, thus reducing the duty cycle of the high-speed RF on the visual side. Under the same battery capacity and recording mode, this terminal has a longer battery life under noise reduction conditions than the solution with high-speed RF constantly online, and a lower casing temperature rise. Furthermore, the soundproof cavity does not require ventilation holes for heat dissipation, maintaining the protection level. Specific battery life and casing temperature values ​​depend on the selected main control model, battery capacity, and recording mode; actual measured data for the entire device shall prevail.

[0040] 4. The evidence collection trigger is both timely and reliable. The trigger source is directly taken from the high dynamic range sound pressure detection results necessary for hearing protection to achieve amplitude limiting. The sampling occurs before the amplitude limiting action, and no clipping saturation occurs under impact sound at the level of 140dB. Compared with the solution of using a camera's built-in ordinary microphone and a separate normally open detection circuit, this invention does not require any additional sound pickup channels, analog front-ends, or static power consumption, and the criteria are not affected by clipping distortion. Compared with manual operation, millisecond-level events can be automatically backtracked and retained.

[0041] 5. Dual reduction in transmission volume and storage usage. Because the event identifier connects image-derived data and image retrieval commands, external devices can determine whether the original image is needed with only a very small amount of data received, and accurately retrieve the segment corresponding to the event, avoiding the blind transmission of the entire image; the acoustic alarm and the transmitted image correspond strictly in terms of evidence.

[0042] 6. Safety-critical functions have a minimum availability requirement. The two main control power supply domains are independent of each other and have level isolation across the main control link. Abnormalities, resets, and power failures of the visual main control or the second wireless communication circuit do not affect the audio side. Under no circumstances will the wearer lose hearing protection in a high-noise environment due to camera malfunctions. This is a guarantee that a single main control camera architecture cannot provide in principle.

[0043] 7. The evidence collection task remains uninterrupted and the files are complete and verifiable under high-temperature conditions. The tiered and downgraded sequence places the upload link action before the image quality action and the stop acquisition action last. Before stopping acquisition, the encoding is completed, the file is closed, and the integrity is sealed to ensure that the recorded content is always complete, exportable, and verifiable. Attached Figure Description

[0044] Figure 1 This is a schematic diagram of the overall structure of the dual-master audiovisual collaboration terminal for the hearing protection earmuffs provided in this embodiment of the invention;

[0045] Figure 2 This is an electrical architecture block diagram of the dual-master audiovisual collaboration terminal for hearing protection earmuffs provided in an embodiment of the present invention;

[0046] Figure 3 This is a schematic diagram of cross-master bidirectional collaborative data flow provided in an embodiment of the present invention;

[0047] Figure 4 This is a timing diagram of encrypted pre-accreditation triggered by acoustic events provided in an embodiment of the present invention;

[0048] Figure 5 This is a schematic diagram of the temperature control graded degradation state machine provided in an embodiment of the present invention;

[0049] Figure 6This is a partial schematic diagram of the assembly of the expansion interface and the HUD module provided in an embodiment of the present invention;

[0050] Figure 7 This is a flowchart of the dual-master audiovisual collaboration method for hearing protection earmuffs provided in this embodiment of the invention.

[0051] Explanation of reference numerals in the attached figures:

[0052] 1—Left earcup; 2—Right earcup; 3—Headband; 4—HUD module; 5—Telescopic bracket; 51—Circular rod; 6—Telescopic cantilever; 67—Plug; 68—Extension interface; 690—Interface locking seat; 100—Audio main control; 110—Ambient sound pickup unit; 120—Speaker unit; 130—First wireless communication circuit; 200—Visual main control; 210—First-view camera module; 220—Memory; 221—Circular pre-recording buffer; 230—Second wireless communication circuit; 300—Cross-main control link. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the present invention clearer, the specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments are only some embodiments of the present invention, and not all embodiments.

[0054] Example 1: Overall structure and electrical architecture of the terminal.

[0055] like Figure 1 As shown, the dual-master audiovisual collaboration terminal of the hearing protection earmuffs provided in this embodiment includes a left earmuff 1, a right earmuff 2, and a headband 3 connecting the left earmuff 1 and the right earmuff 2. The left earmuff 1 and the right earmuff 2 each have a sound-insulating cavity for isolating ambient noise. When worn, the sound-insulating cavity covers the wearer's auricle and forms a seal with the head. Its sound insulation performance is designed according to the requirements of ANSIS 3.19 or EN 352.

[0056] like Figure 1 As shown, the ambient sound pickup unit 110 is located on the outer wall of the soundproof cavity and is used to pick up ambient sound signals from outside the cavity; the loudspeaker unit 120 is located inside the soundproof cavity. Figure 1The outline of the soundproof cavity (shown by dashed lines) is used to play back processed ambient sound signals, voice messages, and prompts to the wearer. The first-view camera module 210 is also located on the outer wall of the soundproof cavity. In this embodiment, it is located on the left earcup 1 and below the ambient sound pickup unit 110 on that side. Its viewing direction is towards the wearer's front and rotates with the head, thus obtaining a first-view image stream consistent with the wearer's line of sight. The advantages of this arrangement are: firstly, the camera module and the pickup unit are arranged on the same side and in the same area, making the image and acoustic events spatially co-originating; secondly, placing the camera module below the pickup unit avoids acoustic obstruction of the pickup openings on the outer wall of the cavity by its outer shell; and thirdly, the camera module does not penetrate the cavity wall of the soundproof cavity, so the sound insulation performance and protection level are not affected.

[0057] like Figure 1 As shown, the left earcup 1 and the right earcup 2 are connected to the headband 3 via telescopic brackets 5, and the telescopic brackets 5 have circular rods 51 between the corresponding earcups and the headband 3. The outer walls of the sound insulation cavities of the left earcup 1 and the right earcup 2 are respectively provided with mirror-symmetrically arranged expansion interfaces 68 and interface locking seats 690 located next to the expansion interfaces 68. Figure 1 The diagram also shows a HUD module 4 mounted on a circular rod 51 via a telescopic cantilever 6, and a cable extending from the HUD module 4 with a plug 67 at its end inserted into the left-side expansion interface 68. Details of this part will be discussed in Embodiment 5. Figure 6 Detailed explanation.

[0058] like Figure 2 As shown, in terms of electrical architecture, this terminal has two main controllers, audio main controller 100 and visual main controller 200, set up in the above-mentioned structural platform. The two are divided according to functional domain and interconnected by the internal cross-main controller link 300.

[0059] The audio main controller 100 is connected to the ambient sound pickup unit 110, the speaker unit 120, and the first wireless communication circuit 130, respectively. Figure 2 As shown in the left-middle half, the audio controller 100 operates two audio links: one is a hearing protection audio link, which performs gain control and impact sound limiting on the ambient sound signal output by the ambient sound pickup unit 110, which varies with the sound pressure level, and then plays it back through the speaker unit 120, so that the wearer can hear the environment and commands clearly at low sound levels and is automatically limited for protection when a high-level impact occurs; the sound pressure level detection result obtained before the limiting is performed by this link is also used as the input for acoustic event determination, and the terminal does not have a separate microphone or analog front end for this determination; the second is a communication audio link, which establishes a call with external intercom devices or mobile terminals through the first wireless communication circuit 130 or a wired interface. In this embodiment, the audio controller 100 adopts an audio-specific system-on-a-chip with multi-channel audio encoding and decoding, dynamic range compression and low-latency digital signal processing capabilities, and the first wireless communication circuit 130 adopts a low-power Bluetooth circuit.

[0060] The visual controller 200 is connected to the first-view camera module 210, the memory 220, and the second wireless communication circuit 230, respectively. Figure 2 As shown in the right half of the diagram. The visual controller 200 encodes the image stream output by the first-view camera module 210 and stores it in the memory 220; the memory 220 is divided into a circular pre-recording buffer 221, the function of which will be discussed in Embodiment 3. Figure 4 Note: The transmission rate of the second wireless communication circuit 230 is higher than that of the first wireless communication circuit 130. In this embodiment, a wireless local area network circuit is used. In this embodiment, the vision master controller 200 adopts a vision-specific system-on-a-chip with image signal processing, video hardware encoding and storage control capabilities. The first-view camera module 210 adopts a 12-megapixel image sensor, and the factory default recording mode is 2K resolution, 30 frames per second.

[0061] like Figure 2 As shown in the two dashed boxes, the audio master controller 100 and its associated first wireless communication circuit 130 are located in the first power supply domain, while the visual master controller 200 and its associated second wireless communication circuit 230 are located in the second power supply domain. The two power supply domains are independent of each other. The cross-master control link 300 is... Figure 2 The two arrows pointing in opposite directions indicate that it is a bidirectional link; in this embodiment, the cross-master control link 300 is a wired serial communication link located inside the terminal, such as a universal asynchronous transceiver interface.

[0062] In terms of power supply and interface system, the right earcup 2 has a detachable battery compartment that is compatible with three power modes: two AAA alkaline batteries, two AAA lithium batteries in series, and a detachable lithium polymer battery pack. The overall power management supports wide voltage input and features temperature control, power consumption control, automatic degradation due to heat, and overcharge and over-discharge protection. Figure 1 As shown, the expansion interface 68 on the right earcup 2 side, in addition to being used to connect the expansion module, also serves as the charging port for the aforementioned lithium polymer battery pack, thus eliminating the need to drill a separate hole in the soundproof cavity for charging. The left earcup 1 side has a bottom cable exit, whose connector can be replaced between a six-pin tactical connector, a push-pull self-locking connector, and a universal serial bus connector, while retaining ten-pin and fourteen-pin aviation plug compatible accessories to adapt to different models of walkie-talkies; the replacement of the connector occurs outside the soundproof cavity and does not affect the cavity seal.

[0063] The audio controller 100 is also connected to a wear detection device and a positioning module. When the wear detection device detects that the terminal is being worn, it wakes up the extended module (e.g., the HUD module 4 connected via the extended interface 68) which is in a low-power sleep state. When the terminal is removed, the module automatically returns to sleep mode, preventing continuous power consumption and light leakage from the display module after removal. The location data output by the positioning module is sent to the visual controller 200 via the cross-control link 300. This data is used to write location tags into the metadata of the encrypted event file described in Embodiment 3 below, enabling images to be retrieved by both time and location. This location data can also be used as a geofencing criterion, serving as one of the trigger sources alongside acoustic events. The battery compartment, the lower cable connector, the wear detection device, and the positioning module are all located outside or isolated from the soundproof cavity, and do not penetrate the cavity wall.

[0064] It should be noted that the specific chip type, pixel specifications, wireless standard and interface form described in this embodiment are only given for the purpose of understanding the best implementation method and do not constitute a limitation on the scope of protection. For alternative forms, see Embodiment 8.

[0065] Example 2: Data flow for bidirectional collaboration across master controllers.

[0066] like Figure 3 As shown in the figure, this embodiment illustrates the bidirectional collaboration between the audio master controller 100 and the visual master controller 200 via the cross-master control link 300, and its data stream is as follows: Figure 3 The numbers marked ① to ⑦ are shown in the text.

[0067] ① After determining that an acoustic event has occurred in its hearing protection audio link, the audio master controller 100 sends the acoustic event information carrying the event identifier to the visual master controller 200 via the cross-master control link 300. For example... Figure 3 As shown, the acoustic event information is generated from the multiplexed and limited sound pressure level detection results, that is, no additional microphone or analog front end is set up for this determination.

[0068] ② The visual master controller 200 generates image-derived data from the image stream output by the first-view camera module 210. The amount of image-derived data is smaller than the image stream itself and is written into the aforementioned event identifier. The image-derived data is sent to the audio master controller 100 via the cross-master control link 300. In this embodiment, the image-derived data is a sequence of thumbnail frames extracted and downsampled at preset intervals, and its data volume is approximately one-hundredth of the original encoded bitstream.

[0069] ③ The audio controller 100 pushes the image-derived data to the external device via the first wireless communication circuit 130. Since the first wireless communication circuit 130 is already connected to the external device as a communication audio link, no new wireless connection needs to be established here.

[0070] ④ When an external device determines that it needs the original image based on the received image-derived data, it sends an image retrieval command carrying a corresponding event identifier to the audio master controller 100. The command reaches the audio master controller 100 via the first wireless communication circuit 130.

[0071] ⑤ The audio master controller 100 forwards the image retrieval command to the visual master controller 200 via the cross-master control link 300.

[0072] ⑥ The vision controller 200 activates the second wireless communication circuit 230 only after receiving this instruction. For example... Figure 3 As shown in the dashed box, the second wireless communication circuit 230 remains in a non-operational state until ④ occurs.

[0073] ⑦ The vision master controller 200 locates the event identifier in the memory 220 and transmits only the image segment corresponding to the event identifier through the second wireless communication circuit 230; after the transmission is completed, the second wireless communication circuit 230 returns to the non-working state.

[0074] Depend on Figure 3 As can be seen, the data derived from the image flows from the visual side to the audio side within the camera, and is ultimately emitted by the low-power circuitry on the non-capturing side; while the high-speed wireless circuitry only operates momentarily when explicitly requested. This is exactly the opposite of the existing scheme, which uses the internal link only for control and status, and entrusts the entire image path to the high-speed wireless on the visual side, in terms of path direction.

[0075] In a further embodiment, image-derived data, acoustic event information, and control messages are carried in a time-division multiplexing manner across the main control link 300, with the transmission priority of acoustic event information higher than that of image-derived data. This ensures that the cross-main control transmission of millisecond-level events is not blocked by batch thumbnail transmissions. Furthermore, within the audio main control 100, the execution priority of the processing task of the hearing protection audio link is higher than that of the forwarding task of image-derived data. When the hearing protection audio link occupies processing resources, image-derived data is queued and buffered on the audio main control 100 side, and sent later after the hearing protection audio link releases processing resources. Therefore, even if image services are introduced into the audio main control, the real-time performance of hearing protection is still guaranteed.

[0076] In a further embodiment, the ambient sound pickup unit 110 is a microphone array respectively mounted on the outer walls of the left earcup 1 and the right earcup 2, for example, two microphones on each side, for a total of four microphones. The same ambient sound signal output by this microphone array is used simultaneously for four purposes: amplitude limiting and gain control with varying sound pressure levels in the hearing protection audio link, feedforward pickup for active noise cancellation, call noise reduction in the communication audio link, and acoustic event detection. The four purposes share the same set of transducers and analog front-end, which is of great significance under conditions where the outer wall area of ​​the soundproof cavity and the overall weight of the device are extremely limited.

[0077] Example 3: Encrypted pre-acceptance certificate triggered by acoustic events.

[0078] like Figure 4 As shown, this embodiment illustrates the evidence collection process triggered by an acoustic event. Figure 4 The five lifelines, from left to right, are the ambient sound pickup unit 110, the audio main controller 100, the cross-main controller link 300, the visual main controller 200, and the memory 220.

[0079] like Figure 4 As shown above, before the event occurs, the vision controller 200 has already encrypted the first-view image stream at the point of acquisition and written the encrypted data into the circular pre-recording buffer 221 in the memory 220 without generating a plaintext video file. The circular pre-recording buffer 221 is cyclically overwritten in a first-in-first-out manner, and in this embodiment, its capacity corresponds to 15 to 30 seconds of image stream. This design has two functions: first, the footage before the event occurs is preserved, solving the problem that manual button presses inevitably lag behind the event; second, the buffered content is entirely encrypted and is not written into a formal file, which avoids privacy exposure caused by default recording and also makes the buffered data unreadable.

[0080] like Figure 4 As shown in the middle section, the ambient sound pickup unit 110 sends the ambient sound signal to the audio master control 100. The audio master control 100 first performs gain control and impact sound limiting that vary with the sound pressure level. This is an inherent action of the hearing protection audio link itself. Subsequently, the audio master control 100 reuses the sound pressure level detection results obtained in the above limiting process to determine whether an acoustic event has occurred.

[0081] In this embodiment, the determination is made using a dual-criteria approach: the first is the time-domain rate of change of the sound pressure level detection result, and the second is the spectral characteristics of the ambient sound signal. When the rate of increase of the sound pressure level within a very short time window exceeds a first threshold, and the energy proportion of this signal in the high-frequency band exceeds a second threshold, a transient impact sound event is determined to have occurred. Regarding the calibration of the two thresholds: Both the first and second thresholds are determined statistically from a measured sample set. The sample set includes impact sound samples collected by this terminal in live-fire, blasting, and hammering scenarios, as well as steady-state high-noise samples collected in machining workshops, engine test benches, and wind-noise environments. The values ​​are determined based on the quantiles that minimize the weighted cost of the false alarm rate and the missed detection rate on the above two types of samples. During use, the audio controller 100 adaptively adjusts the first threshold upwards or downwards based on the sliding statistics of the ambient background sound pressure level over a recent period to adapt to different operating environments. When the ambient background sound pressure level is consistently higher than the set upper limit, making it impossible to distinguish between the two types of samples, the judgment function automatically degrades to only retaining manual button triggering, and the wearer is notified via the speaker unit 120. The specific values ​​of the above thresholds are engineering parameters calibrated according to the scenario. Those skilled in the art can complete the calibration on specific products based on the above statistical sources, value criteria, adaptive rules, and degradation strategies.

[0082] It is important to note the difference in mechanism between this invention and a separate detection microphone solution: The analog front-end of the hearing protection audio link is designed with a dynamic range of 140dB sound pressure level to meet impact noise protection requirements, and the sound pressure level detection occurs before the limiting action. Therefore, when impacts such as gunshots occur, the detection result does not experience clipping saturation, and its time-domain rising edge and spectral shape are preserved. In contrast, ordinary detection microphones and front-ends designed for speech scenarios will inevitably experience clipping under the same sound pressure level; their rising edges are flattened, and the spectrum is contaminated by harmonics. Criteria formed based on this are prone to both missed detections and false alarms. This is precisely the reason why this invention chooses to reuse the detection result of the hearing protection link instead of a separate detection channel, and why this reuse cannot be replaced by adding another sound recognition channel.

[0083] Compared with the documents described in the background art, the differences in this embodiment are concentrated in two aspects: the source and destination of the detection results. First, the source is the sound pressure level detection that the hearing protection must complete for its own safety function. Therefore, it does not require a normally open binary target sound classifier and buffer located in the low power domain like US11862189B2, nor does it require an independent detection processor in the microphone circuit like US9549273B2, nor does it use a built-in ordinary microphone housed in the camera housing like US10939066B2. The headphones of CN119521064A do not have a reusable sound pressure level detection result because they do not contain an impact sound limiting link. The scene determination can only be obtained separately from the sound acquisition module. Secondly, the destination is to be sent to another master controller via a cross-master control link and change its image acquisition, encoding or storage state; while in CN117242784A, the destination of the recognition result is only to change the acoustic characteristics of the sound itself and then play it back through the cavity speaker; in US11862189B2, the activated camera is used to detect the environment in which the device is located in order to improve the sound detection effect; in US9549273B2, the enabled object is the codec and application processor; in US7707035B2, the recognition result of non-speech sound is only used to prompt the wearer and its camera is controlled by the user's voice command.

[0084] like Figure 4 As shown in the middle, after the determination is made, the audio master controller 100 sends the acoustic event information to the visual master controller 200 via the cross master controller link 300. This information includes the event identifier, the characteristic parameters of the transient impact sound event, and the timestamp.

[0085] like Figure 4 As shown in the lower part, in response to the acoustic event information, the vision controller 200 merges the image segments in the circular pre-recording buffer 221 located before the timestamp with the image streams acquired after the timestamp into the same encrypted event file; and writes the event identifier and an undeletable marker bound to the terminal's unique identifier into the metadata of the encrypted event file. The undeletable marker is a metadata field that cannot be modified or cleared through the terminal's user interface after being written. Its determination and interception are implemented by the file system access layer of the vision controller 200, thereby ensuring that the critical event segments cannot be destroyed on the device side.

[0086] Subsequently, as Figure 4 As shown at the bottom, the visual master controller 200 writes the event identifier into the image derived data, and passes the image derived data carrying the event identifier to the audio master controller 100 via the cross master controller link 300, and then transmits it to the outside via the first wireless communication circuit 130.

[0087] In addition to being triggered by acoustic events, the visual master controller 200 can also respond to key inputs and voice commands forwarded by the audio master controller 100 and the cross-master control link 300, or to remote commands received by the first wireless communication circuit 130, performing the same merging action: merging the image segments in the circular pre-recording buffer 221 before the arrival of the command with the image streams acquired afterward into the same encrypted event file. The event identifier is generated by the audio master controller 100 in the same format as the acoustic event information. Therefore, events triggered from different sources have a consistent identifier in metadata, indexes, and subsequent retrieval, allowing external devices to browse and retrieve data in a unified manner without distinguishing the trigger source. Acoustic triggering solves the problem of not having enough time to press a key, key and voice triggering solves the problem of needing evidence even without acoustic features, and remote command triggering solves the problem of needing evidence as determined by the command end. These three methods are complementary and not mutually exclusive.

[0088] In a further embodiment, the visual controller 200 also writes the event identifier, timestamp, feature parameters, and keyframes extracted from the image stream as index entries into the index area stored in the memory 220 and the encrypted event file partition. The index entries are included in the image derived data and sent out together, so that the external device can browse and select the events to be retrieved with only a very small amount of data. The event identifier carried in the subsequent image retrieval command becomes the basis for the visual controller 200 to locate the image segment.

[0089] Example 4: Temperature control graded downgrade.

[0090] like Figure 5 As shown in the figure, this embodiment illustrates a graded degradation strategy under high-temperature conditions. The vision master controller 200 collects its own temperature, and when the temperature reaches the first threshold, second threshold, and third threshold respectively, it correspondingly changes from the normal state to degradation level one, degradation level two, and degradation level three.

[0091] like Figure 5 As shown, the first level of degradation is to suspend the transmission of image data through the second wireless communication circuit 230, while keeping the encoding of the image stream and its storage in the memory 220 unchanged; the second level of degradation is to successively reduce the encoding bit rate and frame rate of the image stream; the third level of degradation is to complete the encoding of the current encrypted event file, file closure and integrity sealing before stopping the acquisition of the image stream.

[0092] The sequence of events in this strategy reflects a trade-off for the continuity of the forensic task: first, actions that can be performed retroactively are sacrificed; second, image quality is sacrificed; and only lastly is data collection stopped. Even at the third level, where data collection must be stopped, the recorded content is first compiled into a complete, exportable, and verifiable file, rather than simply interrupting the writing process and causing file corruption. To avoid further hot accumulation caused by the finishing process itself, the third-level finishing phase is forced to run at the lowest bitrate and has a maximum finishing time limit; if the timeout is exceeded, the written portion is directly archived.

[0093] like Figure 5 As shown in the solid box in the middle, during the aforementioned degradation levels, the following three items remain operational and unaffected by the degradation: the hearing protection audio link, the communication audio link, and the transmission of image-derived data via the main control link 300 and the first wireless communication circuit 130. In other words, even under the most severe high-temperature conditions, the wearer's hearing protection and calls will not be interrupted, and external devices will always be able to see the image information.

[0094] like Figure 5 As shown in the dotted-line box at the bottom, before each level of degradation is executed, the terminal outputs a degradation prompt via the speaker unit 120 and the first wireless communication circuit 130, and records the degradation level, degradation reason, and timestamp, making the entire degradation process perceptible and traceable. Figure 5 As shown by the return arrow in the diagram, after the temperature drops below the first threshold hysteresis threshold, the system recovers step by step in the reverse order of the degradation sequence, that is, it first restores the outgoing transmission, and then restores the frame rate and bit rate, to avoid repeated oscillations near the threshold.

[0095] The values ​​of the first, second, and third thresholds and the hysteresis threshold mentioned above are determined by the junction temperature specifications of the selected vision controller and image sensor, the measured thermal resistance curve of the housing material, and the safe upper limit of the temperature of the human contact surface. The mapping relationship between the housing temperature and the junction temperature is determined by whole-machine thermal simulation and actual measurement calibration. The hysteresis threshold is set to a value that is not less than the temperature overshoot within one complete degradation-recovery cycle. When the temperature sensor fails or the reading exceeds the limit, the system directly enters the second-level degradation and provides a prompt to ensure failure safety.

[0096] Example 5: Assembly of the expansion interface and HUD module.

[0097] like Figure 6 As shown in the figure, this embodiment illustrates the assembly relationship between the expansion interface 68 and the HUD module 4.

[0098] like Figure 6As shown, the upper end of the telescopic bracket 5 is connected to the headband 3, and the part located between the earcups and the headband 3 is a circular rod 51. One end of the telescopic arm 6 is connected to the circular rod 51, and the other end is connected to the HUD module 4; the telescopic arm 6 is a multi-stage telescopic structure, which can rotate around the circular rod 51 and move along its axis, thereby adjusting the HUD module 4 to a suitable position in front of the wearer's eyes.

[0099] like Figure 6 As shown, the outer wall of the earcup's sound insulation cavity is provided with an expansion interface 68 and an interface locking seat 690 located beside it. The HUD module 4 is plugged into the expansion interface 68 via a cable and a plug 67 connected to the end of the cable, and the interface locking seat 690 locks the plug 67 to ensure reliable connection during vigorous movement. Figure 6 As shown on the right, the right earcup has an expansion interface and an interface locking seat that are mirror images of the left earcup, so the plug 67 of the HUD module 4 can be plugged into either side according to wearing habits or interface occupancy.

[0100] like Figure 6 As shown, the HUD module 4, telescopic cantilever 6, expansion interface 68 and interface locking seat 690 are all located outside the sound insulation cavity and do not penetrate the cavity wall of the sound insulation cavity. Therefore, the addition of all expansion capabilities does not impair the sound insulation performance and protection level.

[0101] like Figure 6 The expansion interface 68 shown can be used to connect not only the HUD module 4, but also a microphone module (not shown in the diagram). The microphone module has a gooseneck microphone head structure, with one end being the microphone head near the wearer's mouth, and the other end leading out a cable and plugging it into the expansion interface 68 with a plug of the same specification 67. The plug is also locked by the interface locking seat 690. Since the expansion interfaces 68 and the interface locking seat 690 on the outer walls of the soundproof cavity of the left earcup 1 and the right earcup 2 are arranged in a mirror symmetrical manner, the microphone module can be inserted into either side according to the wearer's hand preference or if one side of the expansion interface 68 is occupied by the HUD module 4. This eliminates the need to create different molds for left-handed and right-handed users, and the expansion capability is not lost due to damage or occupation of one side of the interface.

[0102] The audio master controller 100 and the vision master controller 200 read the module identifier of the connected module through the expansion interface 68 on the side where the plug 67 is actually plugged in, and route the module to different functional links according to the identifier. When the module identifier indicates that the connected module is a microphone module, the audio master controller 100 switches the sound pickup source of the communication audio link from the ambient sound pickup unit 110 to the microphone module, so that the voice of the person in the conversation is picked up by the microphone close to the mouth, and the speech signal-to-noise ratio is higher than that of the microphone on the outer wall of the earmuff in a high-noise environment. At the same time, the ambient sound pickup unit 110 continues to supply ambient sound signals for the gain control and limiting, feedforward pickup of active noise reduction, and determination of acoustic events of the hearing protection audio link, without interruption due to the switching of the call sound pickup source. This is an extension of the dual master control functional domain division at the expansion module level: the access of the module only changes the input source of one link inside the audio master controller 100, without affecting the vision master controller 200 and the image link, nor the critical safety function of hearing protection. If the two expansion interfaces 68 are connected to the microphone module and the HUD module 4 respectively, the audio controller 100 and the visual controller 200 will complete the above two types of routing in parallel according to the module identifiers they read, without interfering with each other.

[0103] When the module identifier indicates that the connected device is HUD module 4, the visual controller 200 generates display data and transmits it to HUD module 4 via the expansion interface 68 on the side where the plug 67 is actually plugged in. The display data consists of at least one of two types of content: firstly, image content generated based on the image stream, such as a viewfinder, focus prompts, or identification labels; secondly, prompt content generated based on acoustic event information. This prompt content is generated by the audio controller 100 and sent to the visual controller 200 via the cross-control link 300 before being synthesized into display data, such as location prompts for impact sound events and notifications of automatically recorded evidence. Therefore, the content displayed by HUD module 4 simultaneously originates from both the visual and audio links, constituting an audiovisual collaboration on the output side.

[0104] Furthermore, the vision controller 200 determines the current configuration as left-eye, right-eye, or binocular based on the side of the plug 67 that is actually plugged in and the module identifier of the HUD module 4 read through the expansion interface 68 on that side. It then generates display data according to the determined configuration. For example, in a monocular configuration, it adjusts the image offset and prompt position according to the corresponding eye position; in a binocular configuration, it generates left and right images respectively. It should be noted that the plugged-in side alone cannot distinguish between a single module and two modules. Therefore, the module identifier is read here: each HUD module 4 provides a unique module identifier on its interface pins or through a protocol handshake. The vision controller 200 can unambiguously determine the configuration based on the number and content of the module identifiers read from the expansion interfaces 68 on both sides.

[0105] Example 6: Dual-master audiovisual collaboration method.

[0106] like Figure 7 As shown, this embodiment provides a dual-master audiovisual collaboration method for hearing protection earmuffs, applied to the earmuff terminal described in Embodiment 1, including the following steps:

[0107] like Figure 7 As shown, in step S1, the audio master controller 100 runs the hearing protection audio link, limits the ambient sound signal output by the ambient sound pickup unit 110, and then plays it back through the speaker unit 120.

[0108] like Figure 7 As shown, in step S2, the vision master controller 200 encodes and stores the image stream output by the first-view camera module 210, and generates image derivative data with a data volume smaller than that of the image stream from the image stream.

[0109] like Figure 7 As shown, step S3 is the judgment step: the audio main controller 100 multiplexes the sound pressure level detection result of the amplitude limiting in step S1 to determine whether an acoustic event has occurred. If the determination is no, then... Figure 7 As shown in the branch path on the right, proceed directly to step S5 and push image-derived data at the regular pace.

[0110] like Figure 7 As shown, in step S4, when it is determined in step S3 that it is, the audio master controller 100 sends the acoustic event information carrying the event identifier to the visual master controller 200 via the cross master controller link 300; in response to the acoustic event information, the visual master controller 200 changes the acquisition parameters, encoding parameters or storage state of the image stream, and writes the event identifier into the image derived data.

[0111] like Figure 7 As shown, in step S5, the image-derived data is sent to the audio master control 100 via the cross-master control link 300 and then transmitted by the first wireless communication circuit 130, while the second wireless communication circuit 230 remains in a non-operating state.

[0112] like Figure 7 As shown, step S6 is a judgment step: whether the audio master controller 100 has received an image retrieval command carrying an event identifier. If not, then... Figure 7 As shown in the return path on the right, return to step S5 to continue pushing.

[0113] like Figure 7 As shown, in step S7, when it is determined in step S6 that it is, the audio master controller 100 forwards the image retrieval instruction to the visual master controller 200 via the cross master control link 300; the visual master controller 200 activates the second wireless communication circuit 230 to transmit the image segment corresponding to the event identifier, and after the transmission is completed, the second wireless communication circuit 230 returns to the non-working state.

[0114] like Figure 7As shown, in step S8, when the visual master controller 200 or the second wireless communication circuit 230 is in an abnormal, reset or power-off state, the power supply and interrupt response of the audio master controller 100 are maintained by mutually independent power supply domains, so that the hearing protection audio link and the communication audio link continue to operate.

[0115] In a further embodiment of the method, the encoding and storage in step S2 is written into the circular pre-recording buffer 221 in an encrypted form at the acquisition starting point, as described in Embodiment 3; the determination in step S3 adopts a dual criterion of time-domain change rate and spectral characteristics, as described in Embodiment 3; the change in storage state in step S4 is manifested as the merging of the pre-recorded segment and the subsequent image stream, the writing of the non-deletable mark, and the establishment of the index item, as described in Embodiment 3. This merging action can be triggered not only by acoustic event information but also by key input, voice commands forwarded via the audio master control 100 and the cross-master control link 300, or by remote commands received via the first wireless communication circuit 130. All sources share the same event identifier format. In yet another further embodiment of the method, a temperature control grading and degradation step as described in Embodiment 4 is also included.

[0116] Example 7: Application Scenario Example.

[0117] Shooting training debriefing. The instructor wears this terminal. If the instructor does not press a button during training, the terminal will automatically trigger formal recording and rewind 15 to 30 seconds based on the gunshot detected by the sound pickup link. Throughout the training, the instructor's mobile phone or tablet continuously receives thumbnails and event indexes via low-power Bluetooth, and the whole device does not heat up. After the training, the instructor selects a gunshot event in the event list, and the terminal then uses the wireless LAN to transmit the corresponding segment of that event.

[0118] Law enforcement and security evidence collection. Team members insert the microphone module into either expansion port 68 according to their preferred hand position and lock it in place, while keeping the other expansion port 68 closed or connecting it to the HUD module 4 as needed. The terminal automatically switches the call audio source to that microphone module after reading the module's identifier. Team members do not enable formal recording by default; the terminal only maintains encrypted pre-recorded buffers. In the event of an emergency, triggered by a button press, voice, or acoustic event, a complete encrypted event file is automatically generated with time and location tags. Key segments cannot be deleted. After transmission, the data enters evidence management via an encrypted channel, and is exported with watermarks and audit information.

[0119] Continuous operation under high-temperature conditions. Under summer helmet-covered conditions, the terminal degrades in stages according to Example 4: first, external transmission is turned off, then the frame rate is reduced, and finally, recording stops after completing the closed loop of disk recording; throughout the process, hearing protection and communication are not interrupted, and external devices can always see the image information; after the temperature drops, it recovers in reverse order.

[0120] Example 8: Alternative solutions and variations.

[0121] Main controller and components: The audio main controller 100 and the vision main controller 200 are not limited to the dedicated system-on-a-chip described in this embodiment, but can be any processor with corresponding capabilities; the image sensor can be 5 megapixels or 12 megapixels, etc.; the storage capacity is expandable.

[0122] Cross-master link: The cross-master link 300 is not limited to the general asynchronous transceiver interface, but can be any internal wired communication link such as serial peripheral interface, integrated circuit bus interface or secure digital input / output interface.

[0123] Wireless standard: The first wireless communication circuit 130 is not limited to Bluetooth Low Energy, but can be a proprietary low-power wireless standard; the second wireless communication circuit 230 is not limited to the wireless LAN access point mode, but can be a wireless LAN direct connection, wired Ethernet, wired transmission via an interface, or a cellular module.

[0124] Image-derived data: not limited to thumbnail frame sequences, can be a set of keyframes, or a structured description of the target category and location generated by the vision master 200, or a combination of the above forms, as long as its data volume is less than the image stream.

[0125] Acoustic events and trigger sources: The criterion feature library can be expanded to include types such as explosion sound and glass breaking sound; in addition to acoustic events, trigger sources can also include key input, voice command, remote command, acceleration, geofence or external sensor signal. All types of trigger sources are generated by the audio master controller 100 with the same event identifier and sent to the visual master controller 200 through the cross master control link 300.

[0126] Positioning module: can be at least one of satellite navigation receiving module, cellular base station positioning module and inertial calculation module, and its output is used for location tags and geofencing criteria of event files.

[0127] Degradation state machine: The number of thresholds, action sets and sequence can be trimmed or expanded according to product positioning, but the core feature is that the upload link actions are included in the sequence and the disk closing loop is completed before the final stage of image quality actions stops capturing. Trimming is not recommended.

[0128] Expansion Structure: The dual earcup interfaces can be arranged symmetrically on the left and right or only on one side; the locking structure can be a rotary lock, a buckle, or a magnetic lock; the modules connected through the expansion interface are mainly microphone modules and HUD modules, and can also be supplementary lights, night vision modules, or second cameras. Each module is identified by a module identifier for the main control to identify and route to the corresponding functional link; the HUD module 4 can be configured as a monocular or binocular, and the telescopic cantilever 6 can be multi-stage telescopic, gooseneck, or rail-shaped.

[0129] Power supply and cable routing: The power supply mode of the battery compartment is not limited to two AAA alkaline batteries, two AAA lithium batteries in series, and lithium polymer battery packs, and can be added or removed according to the model; the charging port is not limited to the expansion interface on the right earcup; the bottom cable routing connector is not limited to a six-pin tactical connector, a push-pull self-locking connector, and a universal serial bus connector, but can also be an aviation plug, or be changed to a wireless design to eliminate the bottom cable routing.

[0130] Wear detection: can be at least one of infrared proximity detection, capacitance detection or pressure detection, and its wake-up object is not limited to HUD module.

[0131] Wearing style: The earmuffs can be headband style, helmet rail style or rear-mounted bracket style.

[0132] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A dual-master audiovisual collaborative terminal for hearing protection earmuffs, comprising a left earmuff (1) and a right earmuff (2) each having a soundproof cavity, a headband (3) connecting the left earmuff (1) and the right earmuff (2), an ambient sound pickup unit (110) and a first-view camera module (210) disposed outside the soundproof cavity, and a speaker unit (120) disposed inside the soundproof cavity; characterized in that, It also includes an audio master controller (100) and a visual master controller (200) with independent power supply domains, and an internal cross-master controller link (300) connecting the two; the audio master controller operates a hearing protection audio link, limits the ambient sound signal output by the ambient sound pickup unit, plays it back through the speaker unit, and connects to the first wireless communication circuit (130); the visual master controller encodes and stores the image stream output by the first-view camera module, and connects to a second wireless communication circuit (230) with a transmission rate higher than that of the first wireless communication circuit; the audio master controller reuses the limited sound pressure level detection result to determine acoustic events and does not set up a separate system for this determination. A microphone is used to send acoustic event information carrying an event identifier to the vision master controller via the cross-master control link. This triggers the vision master controller to change the acquisition parameters, encoding parameters, or storage state of the image stream, and writes the event identifier into image derived data generated by the image stream, which has a smaller data volume than the image stream. The image derived data is sent to the audio master controller via the cross-master control link and transmitted externally by the first wireless communication circuit. The second wireless communication circuit is activated after the audio master controller receives an image retrieval command carrying the event identifier and forwards it via the cross-master control link, and is used to transmit the image segment corresponding to the event identifier.

2. The dual-master audiovisual collaborative terminal for hearing protection earmuffs according to claim 1, characterized in that, The audio master controller (100) and the visual master controller (200) are powered by independent voltage regulator circuits and are each equipped with an independent interrupt controller; a level isolation circuit is provided on the cross master controller link (300) so that when the visual master controller (200) is powered off, its interface pins do not generate backfeed current to the audio master controller (100); The audio master controller (100) receives the operating status information of the visual master controller (200) through the cross master controller link (300), and when the operating status information indicates an abnormality or the operating status information is not received for a preset time, the speaker unit (120) outputs an abnormal prompt, while maintaining the processing cycle of the hearing protection audio link unchanged; when the visual master controller (200) or the second wireless communication circuit (230) is in an abnormal, reset or power-off state, the hearing protection audio link and the communication audio link operated by the audio master controller (100) continue to operate.

3. The dual-master audiovisual collaborative terminal for hearing protection earmuffs according to claim 1, characterized in that, The ambient sound pickup unit (110) includes a microphone array disposed on the outer wall of the left earcup (1) and the right earcup (2), respectively. The same ambient sound signal output by the microphone array is used simultaneously for the amplitude limiting and gain control with sound pressure level of the hearing protection audio link, feedforward pickup for active noise reduction, call noise reduction of the communication audio link, and determination of the acoustic event. The processing task of the hearing protection audio link in the audio master controller (100) has a higher execution priority than the forwarding task of the image derived data. During the period when the hearing protection audio link occupies processing resources, the image derived data is queued and cached on the audio master controller (100) side and sent after the hearing protection audio link releases processing resources.

4. The dual-master audiovisual collaborative terminal for hearing protection earmuffs according to claim 3, characterized in that, The audio master controller (100) determines that a transient impact sound event has occurred when the time-domain change rate of the sound pressure level detection result and the spectral characteristics of the ambient sound signal simultaneously meet the preset impact sound criterion. The acoustic event information also includes the characteristic parameters and timestamp of the transient impact sound event. Before receiving the acoustic event information, the visual master controller (200) writes the image stream into the circular pre-recording buffer (221) in the memory (220) in an encrypted form at the start of acquisition and does not generate a plaintext video file. In response to the acoustic event information, or in response to key input, voice command, or remote command received by the first wireless communication circuit (130) forwarded by the audio master controller (100) and the cross-master control link (300), the visual master controller (200) merges the image segments in the circular pre-recording buffer (221) before the timestamp or the arrival time of the command with the image stream acquired afterward into the same encrypted event file, and writes the event identifier and an undeletable mark bound to the unique identity of the terminal into the metadata of the encrypted event file. When triggered by the key input, the voice command, or the remote command, the event identifier is generated by the audio master controller (100) in the same format as the acoustic event information. The undeletable mark is a metadata field that cannot be modified or cleared by the user interface of the terminal after being written.

5. The dual-master audiovisual collaborative terminal for hearing protection earmuffs according to claim 1, characterized in that, The cross-master control link (300) is a wired serial communication link located inside the terminal. The image-derived data, the acoustic event information, and the control message are carried on the wired serial communication link in a time-division multiplexing manner, and the transmission priority of the acoustic event information is higher than that of the image-derived data. The first wireless communication circuit (130) is a low-power Bluetooth circuit, and the second wireless communication circuit (230) is a wireless local area network circuit. The visual master controller (200) runs the built-in hypertext transfer service in access point mode while the second wireless communication circuit (230) is enabled, and returns the second wireless communication circuit (230) to a non-working state after the image segment transmission ends.

6. The dual-master audiovisual collaborative terminal for hearing protection earmuffs according to claim 1, characterized in that, The left earmuff (1) and the right earmuff (2) are respectively connected to the headband (3) via telescopic brackets (5); the outer walls of the sound insulation cavities of the left earmuff (1) and the right earmuff (2) are respectively provided with an expansion interface (68) arranged in a mirror symmetrical manner and an interface locking seat (690) located next to the expansion interface (68). The first perspective camera module (210) is located on the outer wall of the soundproof cavity of one of the left earcups (1) and the right earcups (2), and is located below the ambient sound pickup unit (110) on that side, with its framing direction facing the wearer; the first perspective camera module (210), the expansion interface (68) and the interface locking seat (690) are all located outside the soundproof cavity and do not penetrate the cavity wall of the soundproof cavity.

7. The dual-master audiovisual collaborative terminal for hearing protection earmuffs according to claim 6, characterized in that, The expansion interface (68) is used to selectively connect a microphone module or a HUD module (4). The microphone module and the HUD module (4) are respectively connected to the expansion interface (68) on the left earcup (1) or the right earcup (2) via the plug (67) at the end of their cables. The plug (67) is locked by the interface locking seat (690) on that side. The connection side is selected according to wearing habits or the occupancy of the expansion interface (68) on that side. The audio controller (100) and the visual controller (200) read the module identifier of the connected module through the expansion interface (68) on the side where the plug (67) is actually connected. When the accessed module is the microphone module, the audio master controller (100) switches the sound pickup source of the communication audio link operated by the audio master controller (100) from the ambient sound pickup unit (110) to the microphone module according to the module identifier, and the ambient sound pickup unit (110) continues to supply the ambient sound signal for the hearing protection audio link and the determination of the acoustic event; When the connected module is the HUD module (4), the HUD module (4) is also connected to the circular rod (51) of the telescopic bracket (5) located between the corresponding earcup and the headband (3) via the telescopic cantilever (6). The visual master controller (200) generates display data and transmits it to the HUD module (4) via the expansion interface (68) on that side. The display data consists of at least one of the image content generated based on the image stream and the prompt content generated based on the acoustic event information and sent to the visual master controller (200) via the cross master control link (300). The visual master controller (200) determines the left eye, right eye, or binocular configuration based on the side where the plug (67) is actually plugged in and the module identifier, and generates the display data according to the determined configuration.

8. A dual-master audiovisual coordination method for hearing protection earmuffs, applied to an earmuff terminal, the earmuff terminal comprising a left earmuff (1) and a right earmuff (2) respectively having a soundproof cavity, a headband (3) connecting the left earmuff (1) and the right earmuff (2), an ambient sound pickup unit (110) and a first-view camera module (210) disposed outside the soundproof cavity, a speaker unit (120) disposed inside the soundproof cavity, an audio master controller (100) and a visual master controller (200) with independent power domains, an internal cross-master controller link (300) connecting the two, a first wireless communication circuit (130) connected to the audio master controller, and a second wireless communication circuit (230) connected to the visual master controller and having a transmission rate higher than the first wireless communication circuit (130); characterized in that, The method includes: S1, the audio master controller operates the hearing protection audio link, and after limiting the amplitude of the ambient sound signal output by the ambient sound pickup unit, it is played back by the speaker unit; S2, the vision controller encodes and stores the image stream output by the first perspective camera module, and generates image derivative data with a smaller data volume from the image stream; S3, the audio master controller reuses the limited sound pressure level detection result to determine the acoustic event without setting up a separate microphone for the determination, and sends the acoustic event information carrying the event identifier to the visual master controller through the cross master controller link; S4, the visual master controller responds to the acoustic event information, changes the acquisition parameters, encoding parameters or storage state of the image stream, and writes the event identifier into the image derived data; S5, the image-derived data is sent to the audio master control via the cross-master control link and transmitted by the first wireless communication circuit, while the second wireless communication circuit remains in a non-working state; S6, after receiving the image retrieval instruction carrying the event identifier, the audio master controller forwards it to the visual master controller via the cross-master controller link. The visual master controller enables the second wireless communication circuit to transmit the image segment corresponding to the event identifier. After the transmission is completed, the second wireless communication circuit returns to the non-working state.

9. The dual-master audiovisual coordination method for hearing protection earmuffs according to claim 8, characterized in that, The encoding and storage in step S2 includes: writing the image stream into the circular pre-recording buffer (221) of the memory (220) of the vision master controller in an encrypted form at the acquisition starting point, without generating a plaintext video file; The determination of acoustic events in step S3 includes: determining that a transient impact sound event has occurred when the time-domain change rate of the sound pressure level detection result and the spectral characteristics of the ambient sound signal simultaneously meet the preset impact sound criterion, and generating the acoustic event information carrying the event identifier, the characteristic parameters of the transient impact sound event and the timestamp; The change of the storage state of the image stream in step S4 includes: in response to the acoustic event information, or in response to key input, voice command, or remote command received by the first wireless communication circuit via the audio master control and the cross-master control link, merging the image segments in the circular pre-recording buffer (221) before the timestamp or the arrival time of the command with the image streams acquired afterward into the same encrypted event file, writing an undeletable mark bound to the unique identifier of the earpiece terminal in the metadata of the encrypted event file, and writing the event identifier, the timestamp, the feature parameters, and the keyframes extracted from the image stream as index entries into the index area of ​​the memory (220) and the encrypted event file partition; the index entries are included in the image derived data and sent out in step S5, and serve as the basis for the visual master control to locate the image segments in step S6.

10. The dual-master audiovisual coordination method for hearing protection earmuffs according to claim 8, characterized in that, Also includes: The temperature of the vision controller is collected, and the following degradation sequence is executed when the temperature reaches the first threshold, the second threshold and the third threshold respectively, which increase sequentially: First level, the transmission of image data through the second wireless communication circuit is suspended, while the encoding and storage of the image stream remain unchanged; The second stage involves sequentially reducing the encoding bitrate and frame rate of the image stream; the third stage involves completing the encoding termination, file closure, and integrity sealing of the current encrypted event file before stopping the acquisition of the image stream. Before each level of degradation is executed, a degradation prompt is output through the speaker unit and the first wireless communication circuit, and the degradation level, degradation reason and timestamp are recorded; in each level of the degradation sequence, the push of the image derived data in step S5 and the operation of the hearing protection audio link are maintained; after the temperature drops below the hysteresis threshold below the first threshold, it is restored level by level in the reverse order of the degradation sequence.

Citation Information

Patent Citations

  • An image processing method, a terminal, and a computer storage medium

    CN110933294B

  • Communication type helmet with noise reduction function

    CN113303540A

  • Head-mounted sound insulation earphone additionally provided with visible light communication equipment

    CN117156334A

  • System and method for sound processing in personal protection devices

    CN117242784A

  • Earphone system, earphone and control method of camera module of earphone

    CN119521064A