Emergency broadcasting method, terminal and storage medium
By collecting data through microphone arrays and camera devices, and combining neural network and knowledge graph technologies, the emergency broadcasting system has achieved precise directional projection and dynamic voice generation in emergencies. This solves the problems of insufficient environmental perception and response delay in existing technologies, and improves the effectiveness and safety of emergency broadcasting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-14
AI Technical Summary
Existing emergency broadcasting systems cannot effectively separate human cries for help from environmental noise during sudden emergencies. They lack environmental awareness and semantic understanding, resulting in broadcast content that cannot be dynamically adjusted according to the situation on site. Furthermore, the omnidirectional broadcasting method causes sound wave energy to be dispersed, making it difficult to achieve precise directional projection. The system response delay is high, making it impossible to achieve real-time emergency response in extreme environments.
The system uses a microphone array and camera device to collect mixed audio and video streams from the environment. It separates human voice signals from environmental sounds through a convolutional neural network, combines knowledge graphs and emotion analysis to generate directional guidance commands, and projects directional sound waves through a phased array speaker array to achieve accurate voice projection.
It achieves efficient separation and directional projection of human voices in high-noise environments, generates dynamic and emotional guidance instructions, improves the real-time performance and accuracy of emergency broadcasts, reduces panic, and enhances evacuation efficiency.
Smart Images

Figure CN121865243A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent emergency response and public safety technology, specifically relating to an emergency broadcasting method, terminal, and storage medium. Background Technology
[0002] In sudden emergencies such as fires, earthquakes, and terrorist attacks, existing emergency broadcasting systems mainly rely on pre-recorded audio combined with omnidirectional loudspeakers. This traditional model has significant technical shortcomings in the complex disaster scene: First, it lacks environmental perception capabilities, failing to effectively separate and amplify critical human cries for help amidst strong noise (such as alarms and collapse sounds), and also failing to understand the dangers represented by background sounds (such as the sound of burning). Second, it lacks semantic understanding and logical decision-making, resulting in monotonous broadcast content that cannot provide dynamic and safe targeted instructions based on the specific location of the crowd, their state of panic, and the distribution of hazards. Third, it lacks emotional interaction and psychological intervention mechanisms, and the mechanical broadcast tone may induce secondary disasters among panicked crowds. Finally, the system has high response latency, especially in extreme environments where edge computing power is limited and there are network and power outages, making it difficult to achieve millisecond-level real-time emergency response.
[0003] Therefore, there is an urgent need for a new generation of emergency broadcasting system that can achieve a complete system from multimodal environmental perception, intelligent logical decision-making, psychoacoustic intervention to precise directional projection. Summary of the Invention
[0004] This application aims to address at least one of the technical problems existing in the prior art or related technologies.
[0005] Therefore, the first aspect of this application proposes an emergency broadcasting method.
[0006] The second aspect of this application proposes an emergency broadcast terminal.
[0007] The third aspect of this application proposes a storage medium.
[0008] In view of this, according to the first aspect of this application, an emergency broadcasting method is proposed, comprising: Receives mixed audio streams and synchronized video streams from the environment; The environmental mixed audio stream is processed based on a preset category cue vector, and an enhanced output is generated. Target human voice audio signal and environmental sound category information; The video stream is processed based on the target human voice audio signal to locate the sound source and segment the corresponding image region, and the face image of the person is detected and extracted from the image region; Based on the image region, the personnel target in it is tracked, and the continuous spatial coordinates of the personnel target are output; Based on the target human voice audio signal and the extracted facial image, emotion analysis is performed, and a panic index is output. Based on the continuous spatial coordinates, the panic index, and the environmental sound category information, reasoning is performed based on a preset knowledge graph and rules to output guidance instruction text; A speech signal is generated based on the guidance instruction text, and the acoustic parameters of the speech signal are adjusted based on the panic index to output a modulated speech stream. Based on the continuous spatial coordinates, beamforming parameters for directional sound projection are generated; Based on the beamforming parameters and the modulated speech stream, a driving signal is generated to drive the sound wave projection device to directionally project the modulated speech stream onto the target area corresponding to the continuous spatial coordinates according to the driving signal.
[0009] According to a second aspect of this application, an emergency broadcast terminal is provided, comprising: Microphone array for capturing mixed audio streams from the environment; Camera device used to capture video streams; processor; Memory used to store computer programs that can be executed by the processor; Sound projection device, used to project speech; The microphone array, camera device, memory, and sound wave projection device are all communicatively connected to the processor. The processor controls the terminal to perform the method described in the first aspect by executing the computer program.
[0010] According to a third aspect of this application, a storage medium is provided that implements the method described in the first aspect when a computer program is executed by a processor.
[0011] Additional aspects and advantages of this application will become apparent in the following description or may be learned by practice of this application. Attached Figure Description
[0012] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 A schematic block diagram of an emergency broadcast terminal according to an embodiment of this application is shown; Figure 2 A flowchart illustrating an embodiment of the emergency broadcasting method of this application is shown; Figure 3 This illustration shows a flowchart of the steps for processing the environmental mixed audio stream based on a preset category cue vector according to an embodiment of this application; Figure 4 This illustration shows a flowchart of an embodiment of the present application, illustrating the steps of processing the video stream based on the target human voice audio signal, locating the sound source and segmenting the corresponding image region, and detecting and extracting the face image of a person from the image region. Figure 5 This illustration shows a flowchart of the steps of tracking a person target in the image region and outputting the continuous spatial coordinates of the person target based on an embodiment of this application. Figure 6 This illustration shows a flowchart of the steps in an embodiment of the present application to perform emotion analysis based on the target human voice audio signal and the extracted face image, and output a panic index. Figure 7 This document illustrates a flowchart of an embodiment of the present application showing the steps of reasoning based on the continuous spatial coordinates, the panic index, and the environmental sound category information, and outputting guidance instruction text based on a preset knowledge graph and rules. Figure 8 A flowchart illustrating the steps of generating a speech signal based on the guidance instruction text, adjusting the acoustic parameters of the speech signal based on the panic index, and outputting a modulated speech stream according to an embodiment of this application is shown. Figure 9 A flowchart illustrating the steps of generating beamforming parameters for directional sound projection based on the continuous spatial coordinates, according to one embodiment of this application, is shown. Detailed Implementation
[0013] To better understand the above-mentioned objectives, features, and advantages of this application, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0014] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below.
[0015] In sudden emergencies such as fires, earthquakes, and terrorist attacks, emergency broadcasting systems are core facilities for guiding crowd evacuation and issuing critical instructions. However, in real-world high-dynamic, high-noise disaster scenes (such as shopping mall fires filled with smoke, flames, collapses, and shouts), the traditional emergency broadcasting system's "pre-recorded audio combined with omnidirectional loudspeakers" model faces fundamental failure: the broadcasting system cannot perceive the real-time status and specific location of the audience, resulting in ineffective information transmission. Specifically, existing technologies cannot distinguish between background environmental noise (such as alarms and collapses) and the cries for help from trapped individuals, leading to a lack of information at the command level; at the same time, the fixed broadcast content pattern cannot adjust the tone of voice according to the emotions of the crowd on site, and mechanical instructions often exacerbate panic or are ignored by the audience; in addition, the omnidirectional broadcasting method causes sound wave energy dispersion, resulting in a low signal-to-noise ratio in noisy environments, making it difficult for the audience to hear instructions clearly, and easily transmitting unnecessary tension to non-dangerous areas. The intelligent terminal and method described in this invention can autonomously perceive the on-site situation, locate specific trapped individuals, generate reassuring instructions, and accurately project them through directional sound waves, thereby efficiently guiding evacuation and improving the success rate of rescue.
[0016] Please see Figure 1 This embodiment provides an emergency broadcast terminal 100, which serves as the carrier for executing the method of this application. The terminal 100 is specifically designed for harsh disaster environments and possesses explosion-proof, waterproof, dustproof, and wide-temperature-range operating capabilities. This terminal is an embodiment provided for ease of understanding the technical solution and is not intended to limit the specific structure of the terminal.
[0017] Terminal 100 mainly includes: Microphone array 110: Used for high-fidelity acquisition of ambient mixed audio streams in the field environment.
[0018] Camera device 120: typically a wide-angle or binocular camera, used to capture synchronized video streams.
[0019] Processor 130: This is the core control unit, responsible for processing, analyzing, making decisions, and generating signals for all audio and video data.
[0020] Memory 140: It stores computer programs that can be executed by the processor.
[0021] Sound wave projection device 150: typically composed of multiple speaker units arranged in a specific geometry, controlled by a processor, capable of generating a highly directional sound beam. In a preferred embodiment of the invention, this device is a phased array speaker array.
[0022] The microphone array 110, camera device 120, memory 140, and sound wave projection device 150 are all connected to the processor 130.
[0023] Optionally, the terminal also includes a communication interface electrically connected to the processor. The communication interface is the channel for the terminal to interact with the external command system. It is used to receive remote start-up, manual instructions or rule updates, and to transmit on-site analysis results (such as panic index and personnel location) and equipment status back in real time, thereby realizing the coordination between the overall command at the rear and the autonomous execution at the front end, and enhancing the system's flexibility and overall response capability under network conditions.
[0024] In the event of a disaster or emergency, it is crucial to accurately deliver reassuring instructions to the target area on-site. Based on this, this application proposes an emergency broadcasting method to control the terminal to accurately project reassuring voice messages.
[0025] See Figure 2 This embodiment provides an emergency broadcasting method.
[0026] The method includes: S1: Receives a mixed audio stream and a synchronized video stream from the environment.
[0027] In this step, after the terminal powers on, its built-in microphone array collects sound wave signals from the surrounding environment in real time. These signals may include a mixture of various sound sources such as human voices, ambient noise, and alarm sounds. The mixed environmental audio stream is then converted into a digital format by an analog-to-digital converter. Simultaneously, the terminal's built-in camera (e.g., a wide-angle RGB camera or a binocular camera) captures video footage from the scene, generating a video stream. The system utilizes a hardware clock or software timestamp mechanism to ensure precise time alignment between the collected audio and video frames, forming synchronized multimodal data input.
[0028] S2: Process the ambient mixed audio stream based on a preset category cue vector, separate and output the augmented audio. Strong target human voice audio signal and environmental sound category information.
[0029] In this step, the target human voice is further extracted from the high-noise mixed audio, and the environment is assessed. (See also...) Figure 3 Its specific implementation is a goal-oriented decoupling process based on neural audio coding.
[0030] S21: Encode the environmental mixed audio stream into a feature representation; S22: Modulate the feature representation based on the category cue vector; S23: Quantize the modulated feature representation; S24: Decode the quantized feature representation to generate the target human voice audio signal and environmental sound class. Other information.
[0031] Specifically, the ambient mixed audio stream obtained from S1 is input into a pre-set convolutional neural network encoder. This encoder transforms the time-domain waveform data into a high-dimensional, continuous latent spectral feature representation through multi-layer downsampling convolution operations. This representation is a compressed representation of the features of all sound sources in the mixed audio.
[0032] Subsequently, the processor retrieves pre-trained category cue vectors from memory. When human voice needs to be extracted, the "human voice" cue vector is invoked; when environmental analysis is required, cue vectors such as "burning sound" and "explosion sound" can be invoked in parallel. Taking the "human voice" cue as an example, it is input together with the features from step S21 into a conditional feature extraction module. The core of this module is a feature-level linear modulation mechanism: first, a global association between the cue and audio features is established through an attention layer; then, a set of channel-level scaling and translation parameters are generated based on the cue vector.
[0033] Then, the scaling and translation parameters are applied to the original features to perform an affine transformation. This operation suppresses noise components unrelated to "human voice" at the feature level while enhancing the feature channels related to human voice, achieving preliminary feature-level source separation. Subsequently, the separated human voice tendency features are fed into a residual vector quantizer, where the features are quantized and residually encoded through a multi-level codebook, converting them into a series of discrete codebook indices. This process compresses the data while further filtering out redundant noise beyond quantization errors.
[0034] Finally, the discrete codebook index is input into the corresponding transposed convolutional decoder to reconstruct the time-domain waveform. When the input is a "human voice" prompt, the decoder outputs a high-purity target human voice audio signal; when the input is a "burning sound" prompt, it outputs the corresponding environmental sound classification result, i.e., environmental sound category information (such as "fire - high intensity").
[0035] S3: Process the video stream based on the target human voice audio signal, locate the sound source and segment the corresponding image region, and detect and extract the face image of the person from the image region.
[0036] This step aims to perform cross-modal visual localization of sound sources and extraction of face images. (See [link to relevant documentation]). Figure 4 Specific methods include: S31: Convert the target human voice audio signal into an acoustic query vector; S32: Perform correlation calculation between the acoustic query vector and the visual feature map of the video frame to generate an acoustic... Source location information; S33: Based on the sound source localization information, segment the corresponding image region, and detect from the image region... To measure and extract facial images of personnel.
[0037] Specifically: The clean target human voice audio signal output by S2 is input into a lightweight acoustic feature extraction network. The network is transformed into a high-dimensional acoustic query vector that represents the time-frequency characteristics of the current human voice.
[0038] Then, for the synchronized current video frame, a convolutional neural network is used to extract its visual feature map. The sound... The query vector serves as the query, and the expanded features of the visual feature map serve as the key and value, performing cross-attention calculation. By calculating the association weights between the acoustic features and the visual features at each spatial location in the image, a sound source localization heatmap is generated. The regions with high response values in the heatmap correspond to the image locations where the sound source (speaker) is most likely to appear.
[0039] Finally, based on the sound source localization heatmap, threshold segmentation or region growing algorithms are used to segment the corresponding image regions. Subsequently, a face detector based on a convolutional neural network (such as MTCNN or RetinaFace) is run within this region to detect and crop out the faces of people for subsequent analysis.
[0040] S4: Based on the image region, track the personnel target within it and output the continuous spatial coordinates of the personnel target.
[0041] This step employs long-term tracking based on semantic memory to address occlusion and rapid movement. (See [link / reference]). Figure 5 Specific methods include: S41: Establish and maintain a memory bank containing a short-term feature queue and a long-term feature pool; S42: For the current video frame, match the features in the memory bank with the features of the current frame to determine... The position of the personnel target in the current frame; S43: Based on the confidence level of the matching result, decide whether to store the current frame features into the long-term feature memory. Zhengchi; S44: Calculate and output the spatial coordinates of the personnel target based on the determined location.
[0042] Specifically, when the terminal first locks onto a target (such as a group of people with the highest panic index), it extracts the depth visual features (including color, texture, and semantic features) of its initial image region and stores them as the root node in a dynamic two-level memory bank. This memory bank contains a short-term memory queue (storing features from the most recent N frames) and a long-term diversity memory pool (storing features from historical keyframes).
[0043] Then, for a new video frame, its global visual features are extracted. The current frame features are used as the query, and all features stored in the memory are used as the key / value pair, and a spatiotemporal cross-attention mechanism is used for matching. Even if the target is currently partially occluded, the system can "infer" the target's current position by matching the features from the memory when it was not occluded in the past, and generate a pixel-level binary segmentation mask.
[0044] Then, the quality of the generated mask is evaluated, and its intersection-union ratio (IU) confidence score is calculated. An update threshold is set: If the confidence level is higher than the threshold, the tracking is considered reliable. The features of the current frame are then used to determine whether to store them in the long-term memory pool based on a diversity sampling strategy (e.g., if the cosine similarity with features in the long-term memory pool is low, it is considered a new perspective), and the short-term queue is updated simultaneously. If the confidence level is lower than the threshold (indicating that the target may be severely occluded or have left the frame), the writing of new features to the long-term memory pool is paused to prevent erroneous memory contamination, and inertial prediction is performed based on the position and motion model of the previous frame.
[0045] Finally, for frames with high confidence, the geometric centroid of the segmentation mask is calculated. Combining the depth information provided by the camera device, the image's two-dimensional coordinates (u, v) and depth d are mapped to three-dimensional spatial coordinates (x, y, z) in the terminal coordinate system using a camera calibration model. The coordinate sequence is then smoothed using Kalman filtering or Gaussian regression, ultimately outputting smooth and continuous spatial coordinates of the person target.
[0046] S5: Based on the target human voice audio signal and the extracted facial image, perform emotion analysis and output the panic index.
[0047] This step quantifies sentiment through multimodal time-series analysis. (See [link / reference]). Figure 6 Specific methods include: S51: Extract a first feature from the target human voice audio signal using a first pre-trained network; S52: Extract a second feature from the face image using a second pre-trained network; S53: Fuse the first feature and the second feature; S54: Input the fused feature sequence into the selective state-space model for time series modeling; S55: Calculate the panic index based on the output state of the state-space model.
[0048] Specifically, the target human voice audio signal output by S2 is input into a pre-trained speaker verification model. Feature extraction layer. This model is sensitive to subtle physiological changes in the vocal cords and outputs a high-dimensional acoustic embedding vector that reflects fundamental frequency jitter, speech rate, and spectral tension.
[0049] Then, the face image extracted by S3 is input into the feature extraction layer of a pre-trained face recognition model. The model has high resolution for subtle facial muscle movements (micro-expressions) such as the orbicularis oculi and orbicularis oris muscles, and outputs high-dimensional visual embedding vectors.
[0050] Subsequently, the aforementioned acoustic and visual embedding vectors are input into two independent lightweight multilayer perceptron adaptation layers. The adaptation layers use nonlinear transformations to project high-dimensional features containing individual identity information into a low-dimensional common feature space related to emotion, filtering out static identity features and enhancing dynamic emotion change features.
[0051] Subsequently, the adapted acoustic and visual feature sequences are concatenated along the time dimension to form a unified multimodal temporal feature sequence, which is then input into a selective state space model (e.g., a MAMBA architecture). This selective state space model (e.g., a model based on the MAMBA architecture) is a pre-trained model. Before the terminal leaves the factory or is deployed, it has been trained on a large number of multimodal (audio-video) disaster exercise or simulation datasets labeled with time-series sentiment tags, and is stored in the memory 140. When the system performs a sentiment analysis task, the processor 130 loads the model's parameters from the memory 140 into memory, preparing for inference computation. This model, through a selective scanning mechanism, can dynamically decide whether to remember or forget historical information based on the importance of the current input, and models long sequences with linear computational complexity, capturing the trajectory of sentiment evolution over time. It outputs a final hidden state vector that condenses the sentiment dynamics throughout the entire analysis period.
[0052] Finally, the final hidden state of the output is input into a linear classification head, which calculates and outputs a continuous scalar value between 0 and 100 as the panic index, while also outputting the dominant emotion category label (such as "panic" or "stunted").
[0053] S6: Based on the continuous spatial coordinates, the panic index, and the environmental sound category information, reasoning is performed based on a preset knowledge graph and rules to output guidance instruction text.
[0054] This step aims to ensure the security and rationality of the instructions. (See [link / reference]). Figure 7 Specific methods include: S61: Convert the continuous spatial coordinates, the panic index, and the environmental sound category information into symbols. Chemical representation; S62: Inject the symbolic representation into the knowledge graph and update the state of the knowledge graph; S63: Generate one or more candidate guidance instructions based on the updated knowledge graph state; S64: Using a differentiable inference engine, based on the current state of the knowledge graph and preset rules, process candidate instructions. Conduct an assessment; S65: Select the final instruction from the candidate instructions that have passed the evaluation as the guide instruction text.
[0055] Specifically, firstly, by setting a threshold, continuous spatial coordinates, the panic index, and the environmental sound category information are mapped into discrete logical predicates. For example: The panic index is mapped to either Panic Level (High) or Panic Level (Low).
[0056] Map the spatial coordinates to LocatedIn(Zone_A).
[0057] Map ambient sound category information to event predicates, such as Event(Fire).
[0058] Then, an emergency scenario is pre-configured in the memory, with physical space regions as nodes and connection relationships as edges. Knowledge graph. The symbolic predicates generated by S61 are injected into the corresponding nodes or edges of the graph in real time to update their state attributes (e.g., marking the "Corridor 1" node as Status(Hazardous)).
[0059] Subsequently, the updated knowledge graph state is input into a differentiable logic reasoning module for evaluation. This module pre-defines expert-defined contingency rules (e.g., "IF PanicLevel(High) AND Connected(Current_Zone,Safe_Zone) THEN Action(Guide, To_Safe_Zone)"). The module uses fuzzy logic or the T-norm to transform traditional Boolean logic operations into differentiable continuous function operations, and calculates the logical satisfaction score of all candidate instructions with the current rule set.
[0060] Subsequently, the high-scoring candidate instructions are subjected to hard constraint verification, including checking whether the coordinates pointed to by the instruction are within a walkable area (obstacle avoidance) and whether there are any logical contradictions. From all the instructions that pass the verification, the one with the highest logical satisfaction score is selected as the final decision.
[0061] Finally, the final abstract logical decision (such as Action(Guide, To_Exit_3)) is combined with the panic level and filled into a preset natural language template to generate specific guidance instruction text (such as "Don't panic, please move to Exit 3 in an orderly manner").
[0062] S7: Generate a speech signal based on the guidance instruction text, adjust the acoustic parameters of the speech signal based on the panic index, and output the modulated speech stream.
[0063] This step aims to synthesize speech with a psychologically soothing effect; see [link / reference]. Figure 8 Specific methods include: S71: Convert the guidance instruction text into an initial voice signal; S72: Map the panic index to an audio processing target; S73: In the parameter space of the audio effects unit, search for an algorithm that can make the processed initial speech signal conform to the desired parameters. The optimal parameter set for the audio processing target; S74: Process the initial speech signal using the optimal parameter set to generate the modulated speech signal. Sound flow.
[0064] Specifically, a lightweight streaming text-to-speech model is used to convert the guidance instruction text output by S6 into the corresponding Mel spectrogram, and then synthesize a clear but emotion-neutral initial speech signal (speech substrate) through a vocoder.
[0065] Then, based on the panic index output by S5, it is mapped to a set of semantic cue pairs describing the characteristics of the target sound. For example, when the panic index is high, the positive cue is "deep, calm, clear," and the negative cue is "sharp, piercing, mechanical." These textual cuees are then converted into high-dimensional acoustic target vectors and negative vectors through a predictive contrastive language-audio pre-trained model.
[0066] Subsequently, a programmable audio effects chain including modules such as a parametric equalizer, dynamic compressor, and reverb unit is constructed. Since the relationship between the effects parameters and the final auditory perception is complex and non-differentiable, a Bayesian optimization strategy is employed for the search: Randomly initialize or generate a set of effects parameters based on history (e.g., EQ boosted by 3dB at 200Hz, reverb time 0.3 seconds).
[0067] The initial speech signal is processed using this parameter combination to obtain candidate audio.
[0068] The candidate audio is input into the CLAP model to extract features. The cosine similarity between the feature and the target vector (positive score) is calculated, and the similarity with the negative vector (negative score) is subtracted to obtain a comprehensive psychoacoustic score.
[0069] Using a Gaussian process as a surrogate model, a black-box function relating "parameter combination" and "score" is fitted. Then, based on the expected improvement of the acquisition function, balancing "exploration" and "exploitation," the next set of parameters most likely to improve the score is selected for trial.
[0070] Repeat the cd step approximately 10-20 times to quickly converge to the optimal parameter set.
[0071] Finally, the determined optimal parameter set is loaded into the audio effects chain to process the continuous initial speech signal in real time, outputting the final modulated speech stream. This speech stream, while semantically clear, has undergone targeted optimization in terms of spectrum, dynamic range, and spatial awareness to achieve a specific acoustic effect that soothes the target audience.
[0072] S8: Based on the continuous spatial coordinates, generate beamforming parameters for directional sound projection.
[0073] This step calculates how to precisely "project" the sound to the target location; see [link / reference]. Figure 9 Specific methods include: S81: Based on the continuous spatial coordinates, construct an optimization model that includes digital precoding weights, analog phase shifting parameters, and physical aperture configuration; S82: Solve the optimization problem using fractional programming and alternating optimization methods to obtain the beamforming parameters.
[0074] Specifically, a beamforming optimization model is first constructed based on the target continuous spatial coordinates output by S4. This model aims to maximize the received signal-to-noise ratio at the target point while minimizing acoustic energy leakage in non-target areas. Optimization variables include: the digital precoding matrix for baseband processing, the analog beamforming matrix controlling the phase of each loudspeaker unit, and the physical aperture configuration parameters that determine which loudspeaker units are activated.
[0075] Then, fractional programming techniques are used to transform the original fractional objective function into an easily manageable subtraction form. Subsequently, an alternating optimization strategy is employed. Step A (Fixed Aperture, Optimized Beam Weight): Temporarily fix the physical aperture configuration, and use the Riemannian manifold optimization method to iteratively solve for the optimal simulated phase shift parameters under the unit mode constraint (adjusting only the phase, not the amplitude) so that the main lobe of the beam is accurately pointed to the target.
[0076] Step B (Fixed Beam Weights, Optimized Aperture): With the current simulated beamforming weights fixed, an adaptive differential evolution algorithm based on successful history is used to search among possible speaker activation combinations for discrete physical aperture configuration parameters, in order to further improve the sound field focusing effect in the target area.
[0077] Alternately execute steps A and B until the terminal performance converges.
[0078] Finally, the final determined digital precoding weights, analog phase-shifting parameters, and physical aperture activation mask are output, which together constitute the beamforming parameters required to drive the acoustic wave projection device.
[0079] S9: Based on the beamforming parameters and the modulated speech stream, a driving signal is generated to drive the sound wave projection device to directionally project the modulated speech stream to the target area corresponding to the continuous spatial coordinates according to the driving signal.
[0080] This step aims to precisely project the intelligently generated reassuring voice to the designated target personnel using sound wave focusing technology, ensuring that they can clearly hear and understand the voice in the noisy and chaotic disaster scene.
[0081] Specifically, based on the beamforming parameters generated in S8 (including the digital precoding matrix, analog beamforming matrix, and physical aperture configuration) and the modulated speech stream generated in S7, the processor synthesizes the final drive signal: First, in the digital domain, the speech stream undergoes multi-channel weighting processing according to the digital precoding matrix. Then, after digital-to-analog conversion, a programmable phase shifter network controlled by the analog beamforming matrix applies precise phase delay to each channel. Simultaneously, the corresponding units in the phased array speaker array are activated according to the physical aperture configuration parameters. All activated units emit sound synchronously according to the drive signal. The emitted sound waves interfere and superimpose in space, forming constructive interference at the continuous spatial coordinates of the target tracked and locked in S4, producing a high-intensity focusing effect. This allows clear guidance instructions to be precisely directed to the target personnel, while in non-target areas, destructive interference is significantly suppressed.
[0082] Understandably, this method is not limited to any specific execution hardware.
[0083] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the emergency broadcasting method of any embodiment of this application.
[0084] It will naturally have all the beneficial effects of the emergency broadcasting method as described in any embodiment of this application, which will not be repeated here.
[0085] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance, unless otherwise expressly specified and limited. The terms "connection," "installation," and "fixing," etc., should be interpreted broadly. For example, "connection" can mean a fixed connection, a detachable connection, or an integral connection; it can mean a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0086] In the description of this specification, the terms "one embodiment," "some embodiments," "specific embodiment," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0087] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An emergency broadcasting method, characterized in that, The method includes the following steps: Receives mixed audio streams and synchronized video streams from the environment; The environmental mixed audio stream is processed based on a preset category cue vector to separate and output the enhanced target human voice audio signal and environmental sound category information; The video stream is processed based on the target human voice audio signal to locate the sound source and segment the corresponding image region, and the face image of the person is detected and extracted from the image region; Based on the image region, the personnel target in it is tracked, and the continuous spatial coordinates of the personnel target are output; Based on the target human voice audio signal and the extracted facial image, emotion analysis is performed, and a panic index is output. Based on the continuous spatial coordinates, the panic index, and the environmental sound category information, reasoning is performed based on a preset knowledge graph and rules to output guidance instruction text; A speech signal is generated based on the guidance instruction text, and the acoustic parameters of the speech signal are adjusted based on the panic index to output a modulated speech stream. Based on the continuous spatial coordinates, beamforming parameters for directional sound projection are generated; Based on the beamforming parameters and the modulated speech stream, a driving signal is generated to drive the sound wave projection device to directionally project the modulated speech stream onto the target area corresponding to the continuous spatial coordinates according to the driving signal.
2. The method according to claim 1, characterized in that, The processing of the environmental mixed audio stream based on the preset category cue vector includes: The environmental mixed audio stream is encoded into a feature representation; The feature representation is modulated based on the category cue vector; Quantize the modulated feature representation; The quantized feature representation is decoded to generate the target human voice audio signal and environmental sound category information.
3. The method according to claim 1, characterized in that, The process of processing the video stream based on the target human voice audio signal, locating the sound source and segmenting the corresponding image region, and detecting and extracting the face image of the person from the image region includes: Convert the target human voice audio signal into an acoustic query vector; The acoustic query vector is correlated with the visual feature map of the video frame to generate sound source localization information; Based on the sound source localization information, the corresponding image region is segmented, and the face image of the person is detected and extracted from the image region.
4. The method according to claim 1, characterized in that, The step of tracking a person target within the image region and outputting the continuous spatial coordinates of the person target includes: Establish and maintain a memory containing a short-term feature queue and a long-term feature pool; For the current video frame, the features in the memory bank are matched with the features of the current frame to determine the position of the person target in the current frame; Based on the confidence level of the matching results, a decision is made on whether to store the current frame features into the long-term feature pool of the memory bank. Based on the determined location, calculate and output the spatial coordinates of the personnel target.
5. The method according to claim 1, characterized in that, The process of performing emotion analysis based on the target human voice audio signal and the extracted facial image, and outputting a panic index, includes: A first feature is extracted from the target human voice audio signal using a first pre-trained network; A second pre-trained network is used to extract a second feature from the face image; The first feature and the second feature are fused together; The fused feature sequence is input into a selective state-space model for time series modeling. The panic index is calculated based on the output state of the state-space model.
6. The method according to claim 1, characterized in that, The system, based on the continuous spatial coordinates, the panic index, and the environmental sound category information, performs reasoning based on a pre-set knowledge graph and rules, and outputs guiding instruction text, including: The continuous spatial coordinates, the panic index, and the environmental sound category information are converted into symbolic representations. The symbolic representation is injected into the knowledge graph, and the state of the knowledge graph is updated; Based on the updated knowledge graph state, generate one or more candidate guidance instructions; Based on the current state of the knowledge graph and the preset rules, the candidate instructions are evaluated; The final instruction is selected from the candidate instructions that pass the evaluation as the guiding instruction text.
7. The method according to claim 1, characterized in that, The process of generating a speech signal based on the guidance instruction text, adjusting the acoustic parameters of the speech signal based on the panic index, and outputting a modulated speech stream includes: Convert the guidance instruction text into an initial speech signal; Map the panic index to an audio processing target; In the parameter space of the audio effects unit, search for the optimal set of parameters that enables the initial speech signal to meet the audio processing objective after processing; The initial speech signal is processed using the optimal parameter set to generate the modulated speech stream.
8. The method according to claim 1, characterized in that, The generation of beamforming parameters for directional sound projection based on the continuous spatial coordinates includes: Based on the continuous spatial coordinates, an optimization model is constructed that includes digital precoding weights, analog phase shifting parameters, and physical aperture configuration. The optimization problem is solved using fractional programming and alternating optimization methods to obtain the beamforming parameters.
9. An emergency broadcasting terminal, characterized in that, include: Microphone array for capturing mixed audio streams from the environment; Camera device used to capture video streams; processor; Memory used to store computer programs that can be executed by the processor; Sound projection device, used to project speech; The microphone array, camera device, memory, and sound wave projection device are all connected to the processor. The processor controls the terminal to perform the method as described in any one of claims 1 to 8 by executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the intelligent emergency broadcasting method according to any one of claims 1 to 9.