Sound beam forming pointing control method based on crowd recognition and related equipment thereof

Through multimodal sensor data processing and dynamic optimization technology based on crowd recognition, the problem that traditional sound beamforming technology is difficult to dynamically adjust the beam direction in complex environments is solved, and high-precision sound source positioning and sound quality improvement are achieved.

CN120128849AInactive Publication Date: 2025-06-10SHENZHEN JINGJING TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510617099.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional sound beamforming technology is difficult to dynamically adjust the beam direction in complex and changing environments, resulting in reduced accuracy of sound source delivery and poor sound quality.

Method used

Using a crowd recognition method, multi-modal sensor data is synchronously collected and calibrated in space-time, multi-modal sensor data, layered signal enhancement, dense crowd detection and tracking, sound source positioning and beam weight calculation, multi-constraint dynamic optimization is performed, and real-time beam synthesis is finally realized.

Benefits of technology

It realizes accurate capture of the space-time trajectory of the crowd and positioning of sound source, dynamically adjusts the beam direction, improves sound quality and anti-interference ability, and enhances the stability and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128849A_ABST
    Figure CN120128849A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of sound beam control, and provides a sound beam forming pointing control method based on crowd recognition and related equipment thereof. The method comprises the following steps: performing time-space synchronous acquisition and calibration on multi-modal sensor data to obtain a time-space aligned original data set, and performing layered signal enhancement on the original data set to obtain a visual feature tensor and an audio feature tensor; dense crowd detection and tracking are carried out according to the visual feature tensor to obtain a target crowd space-time trajectory matrix, and sound source localization and beam weight calculation are carried out in combination with the audio feature tensor to obtain a beam forming weight vector; performing multi-constraint dynamic optimization in combination with environmental parameters in the multi-modal sensor data to obtain an adaptive control parameter set, and finally performing real-time beam forming in combination with a loudspeaker sequence to obtain a directional sound field output signal. According to the invention, target detection, sound source localization and dynamic adaptive control are combined on the multi-source data, so that the beam direction is dynamically adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of sound beam control, and in particular to a sound beamforming pointing control method based on crowd recognition and related equipment. Background Art

[0002] As an important innovation in the field of audio processing, sound beamforming technology is widely used in conferences, education, entertainment and other fields. Its core lies in accurately controlling the delivery of sound energy to enhance the audio experience. Traditional sound beamforming technology mainly relies on preset fixed directions or simple sound source localization algorithms to deliver sound energy. These methods achieve sound enhancement or attenuation in a specific direction by presetting the weights and delays of each unit in the microphone array, thereby ensuring that the sound in the target area is clear and background noise is suppressed.

[0003] In complex and changing environments, traditional sound beamforming technology that uses fixed directions or simple sound source localization algorithms to deliver sound energy has exposed obvious deficiencies. First, the accuracy of its sound source delivery is limited. Especially when the sound source position changes frequently or there are multiple sound source interferences, the preset fixed direction or simple positioning algorithm often cannot accurately capture the sound source, resulting in reduced sound quality. Secondly, in the face of complex environments, traditional technologies lack the ability to dynamically adjust the beam direction. Whether it is echo interference in a conference room or background noise such as wind and traffic in an outdoor environment, traditional technologies are often unable to adapt to environmental changes in real time, resulting in a significant reduction in the sound beamforming effect. Summary of the invention

[0004] In view of this, the present application provides a sound beamforming pointing control method based on crowd recognition and related equipment to solve the problem of being unable to dynamically adjust the beam direction.

[0005] The first aspect of the present application provides a sound beamforming pointing control method based on crowd recognition, the method comprising: Perform spatiotemporal synchronous acquisition and calibration of multimodal sensor data to obtain spatiotemporal aligned raw data sets; Performing hierarchical signal enhancement processing on the original data set to obtain a visual feature tensor and an audio feature tensor after signal gain optimization; Performing dense crowd detection and tracking processing according to the visual feature tensor to obtain a spatiotemporal trajectory matrix of the target crowd; Performing sound source localization and beam weight calculation processing on the audio feature tensor according to the target population spatiotemporal trajectory matrix to obtain a beamforming weight vector; Performing multi-constraint dynamic optimization processing on the beamforming weight vector according to the environmental parameters in the multimodal sensor data to obtain an adaptive control parameter set; Perform real-time beamforming processing on the adaptive control parameter set according to a preset speaker sequence to obtain a directional sound field output signal.

[0006] In an optional embodiment, the multimodal sensor data includes audio signals, visual signals, and environmental parameters. The spatio-temporal synchronous acquisition and calibration processing of the multimodal sensor data to obtain a spatio-temporally aligned original data set includes: Perform real-time environmental detection processing on the target area according to a preset synchronous acquisition method to obtain the audio signal, the visual signal, and the environmental parameters with a unified initial timestamp; Perform timestamp attachment processing on the audio signal, the visual signal, and the environmental parameters to obtain time-tagged multi-source data; Perform asynchronous sampling interpolation compensation processing on the time-tagged multi-source data to obtain a temporally aligned intermediate data set; Perform sensor calibration processing on the intermediate data set to obtain the original data set.

[0007] In an optional embodiment, the hierarchical signal enhancement processing of the original data set to obtain a visually characteristic tensor and an audio characteristic tensor with optimized signal gain includes: Perform audio-visual signal classification processing on the original data set to obtain an original audio signal and an original visual signal; Perform Wiener filter noise reduction processing on the original audio signal to obtain a noise-reduced optimized audio signal; Perform Mel spectrum modulation and demodulation processing on the noise-reduced optimized audio signal to obtain the audio characteristic tensor; Perform image distortion correction processing on the original visual signal to obtain a corrected image; Perform spatial attention weighting processing on the corrected image to obtain the visually characteristic tensor.

[0008] In an optional embodiment, the dense crowd detection and tracking processing according to the visually characteristic tensor to obtain a target crowd spatio-temporal trajectory matrix includes: Perform Anchor box size adjustment processing on a preset YOLO v5 model according to the visually characteristic tensor to obtain an improved YOLO v5 model; Perform target detection processing on the visually characteristic tensor through the improved YOLO v5 model to obtain a candidate target set; Perform occlusion-aware loss compensation processing on the candidate target set to obtain an optimized target position prediction result; Perform Kalman filtering on the predicted result of the target position to obtain the spatio-temporal trajectory matrix of the target population.

[0009] In an alternative embodiment, the performing sound source localization and beam weight calculation on the audio feature tensor according to the spatio-temporal trajectory matrix of the target population to obtain a beamforming weight vector includes: Perform generalized cross-correlation processing on the audio feature tensor to obtain a candidate set of sound source directions; Perform visual trajectory matching and angle screening on the candidate set of sound source directions according to the spatio-temporal trajectory matrix of the target population to obtain a target direction; Perform minimum variance distortionless response weight calculation according to the target direction to obtain the beamforming weight vector.

[0010] In an alternative embodiment, the performing multi-constraint dynamic optimization on the beamforming weight vector according to the environmental parameters in the multi-modal sensor data to obtain an adaptive control parameter set includes: Perform noise covariance matrix update on the beamforming weight vector according to the environmental parameters to obtain a covariance matrix; Perform temperature attenuation factor correction on the covariance matrix according to the real-time temperature in the environmental parameters to obtain a corrected weight; Perform normalization constraint on the corrected weight to obtain the adaptive control parameter set.

[0011] In an alternative embodiment, the performing real-time beam synthesis on the adaptive control parameter set according to a preset speaker sequence to obtain a directional sound field output signal includes: Step S61: Perform carrier modulation on the adaptive control parameter set according to the speaker sequence to obtain adjustment signals for each speaker channel; Step S62: Perform coherent superposition of the adjustment signals for multiple speaker channels according to the speaker sequence to obtain the directional sound field output signal; Step S63: Perform real-time monitoring of the power spectrum on the directional sound field output signal to obtain a real-time control power spectrum; Step S64: When there is a power spectrum value in the real-time control power spectrum greater than a preset power threshold, perform weight adjustment and optimization according to a preset interruption weight update method to update the adaptive control parameter set; Repeat the execution of step S61 to step S64 until the power spectrum values in the real-time control power spectrum are all less than or equal to the power threshold.

[0012] The second aspect of the present application provides a voice beamforming pointing control device based on crowd recognition, and the device includes: A spatio-temporal alignment module, configured to perform spatio-temporal synchronous acquisition and calibration processing on multi-modal sensor data to obtain a spatio-temporally aligned original data set; A hierarchical gain module, configured to perform hierarchical signal enhancement processing on the original data set to obtain a visually characteristic tensor and an audio characteristic tensor with optimized signal gain; A visual feature module, configured to perform dense crowd detection and tracking processing according to the visually characteristic tensor to obtain a target crowd spatio-temporal trajectory matrix; An audio feature module, configured to perform sound source localization and beam weight calculation processing on the audio characteristic tensor according to the target crowd spatio-temporal trajectory matrix to obtain a beamforming weight vector; A dynamic optimization module, configured to perform multi-constraint dynamic optimization processing on the beamforming weight vector according to the environmental parameters in the multi-modal sensor data to obtain an adaptive control parameter set; A beam synthesis module, configured to perform real-time beam synthesis processing on the adaptive control parameter set according to a preset speaker sequence to obtain a directional sound field output signal.

[0013] The third aspect of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned voice beamforming pointing control method based on crowd recognition are implemented.

[0014] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned voice beamforming pointing control method based on crowd recognition are implemented.

[0015] In summary, the present application at least includes the following beneficial technical effects: 1. By combining the improved YOLO v5 model with Kalman filtering for dense crowd detection and tracking, the spatio-temporal trajectory of the target crowd can be accurately captured to compensate for the occlusion problem, effectively improving the robustness of detection and tracking and ensuring the accuracy of pointing control decisions.

[0016] 2. The noise covariance matrix is updated, the temperature attenuation is corrected, and the normalization constraint is performed on the beamforming weight vector according to the environmental parameters, and the control parameters are dynamically adjusted, so as to be able to adapt to different environmental conditions, realize real-time adaptive control, and further improve the stability and anti-interference ability. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 is a flowchart of a method for controlling the sound beamforming direction based on crowd recognition provided by an embodiment of the present application; Figure 2 is a functional module diagram of a device for controlling the sound beamforming direction based on crowd recognition provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0020] As Figure 1 shown, it is a flowchart of a method for controlling the sound beamforming direction based on crowd recognition provided by an embodiment of the present application. The method for controlling the sound beamforming direction based on crowd recognition provided by an embodiment of the present application includes the following steps.

[0021] The method for controlling the sound beamforming direction based on crowd recognition provided by an embodiment of the present application is executed by an electronic device. Correspondingly, the device for controlling the sound beamforming direction based on crowd recognition runs in the electronic device. The electronic device collects the required objective parameters through various sensors deployed inside the target area. These sensors include a microphone array for collecting the sound field information in the environment, a camera for collecting the scene image and video information, and an environmental parameter sensor for detecting auxiliary information such as temperature and humidity. The following will describe the method for controlling the sound beamforming direction based on crowd recognition provided by an embodiment of the present application from the perspective of the electronic device in combination with the process of adjusting the sound wave direction of the speaker in the target area.

[0022] Step S1: Perform spatio-temporal synchronous acquisition and calibration processing on the multi-modal sensor data to obtain a spatio-temporally aligned original data set.

[0023] It should be understood that if the acquisition start times are inconsistent, the data will be mismatched due to time deviation, affecting the multi-modal data fusion effect. Therefore, in the embodiments of the present application, using a hardware trigger to uniformly start the acquisition not only improves the accuracy of data acquisition but also ensures strict requirements for real-time performance and synchronization of the entire system. First, sampling parameters are set for the sensors within the target area. For example, the microphone sampling rate is set to 48 kHz, the camera frame rate is set to 30 fps, and the sampling period of the temperature and humidity sensor is set to 0.1 second. Secondly, at a preset time point, a unified acquisition start signal is sent to all sensors through the hardware trigger. This signal ensures that all sensors start acquiring data at the same moment, thus providing a unified initial timestamp for subsequent spatio-temporal synchronization. After receiving the trigger signal, the sensors immediately start acquisition. All acquisition devices attach a globally unified initial timestamp to the data packet, and this timestamp usually comes from a high-precision atomic clock or GPS clock. Further, to ensure the continuity and real-time performance of the data, it is also necessary to monitor the acquisition status of the sensors to ensure that there are no missed acquisitions or abnormal time delays during the acquisition process.

[0024] Furthermore, since the sampling moments of the sensors themselves may have slight deviations or small differences due to hardware delays when acquiring data, it is necessary to attach precise time tags to each data sample in the data stream. To eliminate the time delay differences caused by different sampling frequencies and hardware response times of different acquisition devices. Timestamp attachment processing is performed on the acquired multi-modal sensor data to obtain time-tagged multi-source data. Among them, the multi-modal sensor data includes but is not limited to audio signals, visual signals, and environmental parameters. First, after receiving the acquisition instruction with a unified initial timestamp, each sensor records a timestamp at each sampling point (for example, audio sampling, each frame of image, or each environmental parameter value) according to the preset sampling parameters. For example, for audio signals, the sampling frequency is 48 kHz, and the interval between each sampling point is approximately 20.83 microseconds; for visual signals, the frame rate is 30 fps, and the interval between each frame is approximately 33.33 milliseconds; while the environmental parameter sensor acquires data at a period of 0.1 second. After receiving the multi-modal sensor data, the electronic device uses a high-precision timer or clock counter to attach an accurate time scale to each sampled data, combines the data with timestamps into a set of multi-source data (i.e., time-tagged multi-source data), forming a data set with time tags. Each data item in this set can correspond to the data at the corresponding moment acquired by other sensors.

[0025] Due to the large differences in the sampling rates of various sensors, and the sampling moments of some sensors (such as environmental parameter sensors) not being integer multiples of time points, there will be inconsistencies in the data on the time axis. To achieve precise alignment of the data, it is necessary to perform interpolation processing on the multi-source data after time tagging, fill in the missing points of the asynchronous sampling data, and unify all the data to a common time reference to form an intermediate data set with time-domain alignment. Specifically, a common time axis is selected. This common time axis can be chosen with a high-frequency sampling (such as the sampling rate of audio signals) as the reference, or a suitable time step can be set according to the system requirements. For sensor data with a lower sampling rate (such as visual signals or environmental data), the cubic spline interpolation algorithm is used to compensate for the gaps on the common time axis. During the interpolation process, the data value at each missing moment is estimated through a cubic polynomial function based on the values of the surrounding existing data points to ensure a smooth transition in the time continuity of the data. The obtained intermediate data set contains the values from all sensors at each common time point, thus achieving the time-domain alignment of the data. The interpolation process can be represented by the following formula: where, represents the signal value obtained after interpolation at time . represents the time tags of the known data points. are the spline coefficients calculated by the cubic spline interpolation algorithm. is the number of adjacent data points participating in the interpolation calculation. The interpolation processing enables continuous and smooth numerical values to be obtained at each common time point, thus achieving the time-domain alignment of the data.

[0026] Finally, the intermediate data set that has been temporally aligned is further processed for sensor calibration. Sensor calibration includes image undistortion for the visual sensor and acoustic calibration for the microphone array, ensuring that the data is not only aligned in time but also matches the actual physical coordinates in space. This process transforms the intermediate data set into a high-precision raw data set, laying the foundation for subsequent multi-modal data fusion and sound field control. Specifically, for the visual signal data, undistortion processing is performed based on the pre-obtained camera intrinsic matrix. The camera intrinsic matrix includes the focal length, principal point coordinates, and distortion coefficients. By using the pinhole model and distortion correction algorithm, the images captured by the camera are geometrically corrected to generate corrected image data. For the audio signal data, acoustic calibration is performed based on the geometric arrangement information of the microphone array. This calibration includes measuring the relative positions between each microphone and using this information to calculate the steering vector, providing accurate spatial information for subsequent sound source localization and beamforming. For the environmental parameter data, the installation position and response characteristics of the sensors are confirmed, and the measured data such as temperature and humidity are mapped to the actual environmental coordinates to ensure the accuracy of the environmental data. After the calibration process, all data is corrected both temporally and spatially, forming a high-precision raw data set that strictly corresponds to the actual scenario.

[0027] Step S2: Perform hierarchical signal enhancement processing on the raw data set to obtain a visually characteristic tensor and an audio characteristic tensor with optimized signal gain.

[0028] As a mixed data set, the raw data set is prone to errors due to data mixing during subsequent processing. Only by separately extracting the audio signal and the visual signal can targeted signal enhancement algorithms be used in their respective fields. First, the raw data set passes through a preprocessing module. After the internal data format is converted according to a unified standard, the data will be stored in the form of data packets, where each data packet carries a timestamp and a sensor identifier. By parsing the identifier information in the data packet, the system can accurately determine the data source belonging to which modality. For example, the data packets collected by the microphone array contain specific identifiers for audio signals, while the data packets collected by the camera contain identifiers for image or video data. The data classification processing module classifies the data of different modalities according to a pre-set data format protocol and stores them separately in different data buffers. After this process, the audio data forms a set of raw audio signals; while the image data forms a set of raw visual signals.

[0029] For the original audio signal, noise in the original audio signal is suppressed through Wiener filtering to extract a clearer and purer audio signal, thereby laying a foundation for subsequent feature extraction. Wiener filtering is a classic signal processing method, and its core idea lies in using the ratio of the power spectral density of the target signal and noise to adaptively weight the signal in the frequency domain, thereby achieving the noise reduction effect. First, perform a short-time Fourier transform on the original audio signal to convert the time-domain signal to the frequency domain. This process divides the continuous audio signal into short-time frames, and each frame is Fourier-transformed to obtain a spectral representation. In the frequency domain, according to a pre-set algorithm, estimate the power spectral density of the target signal and the power spectral density of the noise for each frequency component. The noise power spectrum is usually obtained through statistical methods in the non-speech or silent segments. Apply the Wiener filtering formula to weight each frequency component, suppress the noise component, and retain the target signal component. Perform an inverse Fourier transform on the processed frequency-domain data to restore it to the time domain to obtain the noise-reduced audio signal (i.e., the noise-reduced and optimized audio signal). The noise reduction process of Wiener filtering can be represented by the following formula: where, represents the representation of the noise-reduced audio signal in the frequency domain. represents the representation of the original audio signal in the frequency domain. represents the power spectral density of the target signal at frequency . represents the power spectral density of the noise at frequency . The noise reduction process of Wiener filtering is used to perform noise reduction on the original audio signal in the frequency domain. The noise component is attenuated through a weighted filter, and a relatively pure target signal is retained, thereby improving the signal-to-noise ratio.

[0030] The denoised audio signal is further converted into an audio feature tensor, providing high-discrimination time-frequency features for subsequent recognition and classification tasks. In the embodiments of the present application, Mel spectrum modulation and demodulation processing is further performed on the denoised and optimized audio signal. Mel spectrum modulation and demodulation converts the spectrum of the audio signal into the Mel scale that conforms to the human auditory characteristic, and extracts feature coefficients therefrom (for example, Mel Frequency Cepstral Coefficients, MFCC). Specifically, the short-time Fourier transform is performed on the denoised audio signal to obtain the frequency-domain representation of the signal. According to a preset Mel filter bank, the spectral data is filtered. The Mel filter bank is a series of band-pass filters, the center frequencies of which are distributed on the Mel scale and can simulate the sensitivity of the human ear to different frequency bands. Taking the logarithm of the filtered energy values forms the Mel spectrum. Using the discrete cosine transform to transform the logarithmic Mel spectrum, low-dimensional feature coefficients are extracted, and these feature coefficients constitute the audio feature tensor. The formed audio feature tensor usually has a fixed dimension. For example, each frame contains 128-dimensional or other preset-dimensional features, which can describe the time-frequency structure of the signal. The Mel spectrum modulation and demodulation processing can be represented by the following formula: wherein, represents the energy output of the th Mel filter. represents the representation of the audio signal in the frequency domain. represents the th frequency response function of the Mel filter. and are respectively the boundary frequencies of the th filter. represents the th Mel Frequency Cepstral Coefficient. represents the total number of Mel filters. is the energy output of the th filter. represents the sequence number of the cepstral coefficient.

[0031] For the original visual signal, the geometric distortion in the original visual signal collected by the camera is eliminated, so as to obtain a corrected image consistent with the geometric structure of the actual scene. There may be lens distortion problems during the manufacturing and installation of the camera. Common ones are radial distortion and tangential distortion. These distortions will cause the straight lines in the image to bend, affecting object detection and subsequent feature extraction. The image distortion correction process is based on the pre-calibrated camera internal parameter matrix and distortion parameters. Through a mathematical model, the position of each pixel in the original image is remapped to the corrected image. Specifically, the camera is calibrated to obtain the camera internal parameter matrix and distortion parameters (for example, radial distortion coefficient k 1 、k2 , k 3 and the tangential distortion coefficient p 1 , p 2 ) According to the pinhole camera model and the corresponding distortion model, calculate the ideal position of each pixel in the original image. Use the interpolation algorithm to resample the original image data to generate the corrected image. During the correction process, the mapping relationship of each pixel is calculated through mathematical formulas to ensure the high precision and continuity of the correction result. Taking radial distortion correction as an example, the distortion correction process can be expressed by the following formula: Among them, represents the normalized coordinates of the pixel in the original image. represents the corrected pixel coordinates. represents the square of the radial distance of the normalized coordinates. is the radial distortion coefficient. Radial distortion correction is used to calculate the radial distortion compensation amount of each pixel in the original image, and add it to the original coordinates to obtain the corrected pixel position, ensuring that the corrected image is geometrically consistent with the actual scene.

[0032] Furthermore, use the spatial attention mechanism to further enhance the features of the corrected image to generate a visual feature tensor with high discrimination. Spatial attention weighting is usually achieved through a convolutional attention module. This module calculates the attention weights in both the channel and spatial dimensions, strengthens the key information in the image, and thus improves the accuracy of target detection and recognition. Specifically, input the corrected image into a pre-trained convolutional neural network to extract the initial image feature map. On the extracted initial image feature map, first perform channel attention calculation. The channel attention module uses global average pooling and global maximum pooling to statistically analyze each channel feature, and then calculates the importance weight of each channel through a shared multi-layer perceptron. After obtaining the channel attention weight, multiply it with the original feature map channel by channel to achieve channel weighting. Subsequently, perform spatial attention calculation on the weighted feature map. The spatial attention module generates a two-dimensional attention map by pooling the feature map in the channel dimension (e.g., max pooling and average pooling), and then obtains the spatial attention weight through a convolutional operation. Multiply the spatial attention weight with the weighted feature map pixel by pixel to obtain the final visually feature tensor after attention weighting. The visually feature tensor contains the importance information of each region in the image, which can highlight the target region and suppress background noise.

[0033] Step S3: Perform dense crowd detection and tracking processing according to the visual feature tensor to obtain the spatio-temporal trajectory matrix of the target crowd.

[0034] It should be understood that the traditional Anchor box size design is usually based on larger targets or conventional scenarios. When directly applied to crowded scenes, targets may be missed due to size mismatch. Adjusting the Anchor box can improve the detection accuracy of small targets in crowded scenes. Specific parameters of the pre-trained YOLO v5 object detection model are adjusted, especially the adjustment of the Anchor box size. The original YOLO v5 model already has high object detection performance in conventional scenarios. However, when dealing with crowded scenes, due to the small target size, high density, and severe occlusion, the preset Anchor sizes in the model often cannot fully cover all targets, resulting in missed detection or misjudgment of small targets. To solve this problem, it is necessary to optimize the size distribution of the Anchor box to make it more suitable for detecting small targets in crowded scenes. The entire adjustment process is carried out based on the visual feature tensor, in which high-level feature information of the image has been extracted. These information include the edges, textures, and local semantic information of the target area, providing rich basis for the subsequent adjustment of the Anchor box. Specifically, first, the size range of pedestrians in the target scene is determined. By analyzing a large amount of annotated data of crowded scenes, it can be determined that the size range of pedestrians usually presents as a relatively small pixel area, and then a new set of Anchor box sizes is selected. For example, {(8, 8), (16, 16), (32, 32)}, etc. The set of Anchor box sizes is more suitable for small target detection than the large-sized Anchor boxes used in traditional models and can cover the local areas of each target in crowded scenes. Next, for the preset YOLO v5 model, the Anchor box parameters in the model are modified, usually by resetting the Anchor box parameters in the model configuration file or network definition. The adjusted model is called the improved YOLO v5 model. Its network structure remains basically unchanged, only the prior box sizes in the output layer are modified, so that the model can generate more candidate regions that fit the crowded small targets during forward propagation.

[0035] After obtaining the improved YOLO v5 model, forward inference is performed on the visual feature tensor to complete the detection of all potential targets in the image. The main task of object detection is to find each potential crowd target in the image and identify it with a bounding box. The specific process includes the model extracting features through multiple levels of the network for the input visual feature tensor, generating region candidates, and subsequent non-maximum suppression processing, and finally outputting a set of candidate targets. This set of candidate targets includes information such as the detected person boxes, corresponding confidence scores, and position coordinates. Specifically, the visual feature tensor obtained through the aforementioned hierarchical signal enhancement processing is used as the input, and the input data undergoes forward propagation through the improved YOLO v5 network. During the forward propagation process, the network continuously abstracts and extracts feature information in the image through hierarchical processing such as convolution, normalization, and activation. The output layer of the network predicts the corresponding object category and position offset at each position based on the adjusted Anchor boxes. The output candidate target information includes bounding box coordinates (usually represented by the center point coordinates, width, and height), confidence, and class scores. All candidate boxes output by the network are initially screened to remove candidate boxes with low confidence. This screening process uses a set confidence threshold to only retain candidate boxes with a confidence higher than this threshold, ensuring that the information in the candidate target set is relatively accurate. Further non-maximum suppression (NMS) processing is performed on the candidate boxes. This step removes redundant boxes by calculating the overlap ratio (IoU) between candidate boxes to ensure that each target corresponds to only one candidate box, thereby generating the final candidate target set.

[0036] Occlusion is the main difficulty in dense crowd detection. Since the target parts overlap or are completely occluded, the detector will output inaccurate candidate bounding boxes. Introducing an occlusion-aware loss term gives a greater weight to the prediction error of occluded targets during the training process, forcing the model to be more sensitive to the information in the occluded area in feature extraction and location regression, thus improving the accuracy of object detection. In a dense crowd scenario, due to the mutual occlusion between people, the detected bounding boxes in the candidate target set may have position deviations, or multiple overlapping bounding boxes may not accurately distinguish each target. To improve the detection accuracy, an occlusion-aware loss term needs to be introduced into the loss function of the candidate targets, so as to impose a greater penalty on the occluded area during the training stage, forcing the model to pay more attention to the detailed information of the occluded targets when outputting the prediction bounding boxes. This step calculates the error between the predicted position and the true position for each target bounding box based on the candidate target set, and at the same time imposes an additional penalty on those occluded targets, making the final predicted result of the target position more accurate. Specifically, for each detected bounding box in the candidate target set, use depth information or an additional occlusion determination module to determine whether the candidate target is occluded. For each candidate bounding box, calculate the standard cross-entropy loss and the additional occlusion compensation loss. The occlusion compensation loss usually uses the mean squared error loss. The two parts of the loss are weighted and summed to obtain the total loss function. During the model training process, the parameters of the improved YOLO v5 model are updated through the backpropagation algorithm, so that the predicted positions of the output candidate targets can also be more accurately predicted in the occluded area. After several training cycles, the model's ability to recognize and locate occluded targets is significantly improved, and the predicted results of the target positions in the candidate target set are optimized. Among them, the total loss function can be expressed by the following formula: where represents the total loss. represents the standard cross-entropy loss, which is used to measure classification errors. is the balance coefficient, usually set to 0.5. represents the index set of occluded targets. represents the center coordinates of the th candidate target prediction bounding box. represents the center coordinates of the

[0037] th candidate target ground truth bounding box. The total loss function is used to impose an additional penalty on the occluded area during the candidate target detection process, ensuring that the predicted results of the positions of the occluded targets are more accurate, thereby improving the overall object detection accuracy.

[0037] Due to the mutual occlusion of targets and detection errors, the single-frame detection results often fluctuate. The Kalman filter can utilize historical information and motion models to smoothly predict and update the target state, thereby effectively reducing the impact of detection noise and obtaining the motion trajectory matrix of the target population in consecutive frames. Dense crowd detection not only requires localizing targets in a single frame but also tracking targets in consecutive frames to obtain their motion states (e.g., position and velocity). As a classic state estimation method, the Kalman filter can smoothly predict and track the motion state of targets in the presence of noise and partial occlusion, thus outputting a high-precision spatio-temporal trajectory matrix. Specifically, first, a state vector is established for each candidate target in the first frame , where is the center coordinate of the target in the image, and are the initial velocities (which can be initially set to 0 or estimated based on the detection results). In each subsequent frame, the state of each target is predicted using the prediction formula of the Kalman filter. The state transition model can be set as a linear model and can be represented by the following formula: where, is the state transition matrix, describing the linear change of position and velocity over time. is the control matrix. is the control input (usually set to 0 in the absence of external control). represents the predicted state vector at time t + 1. represents the state vector at time t. The state transition matrix can be represented by the following formula: Meanwhile, the Kalman filter calculates the covariance matrix of the predicted state to quantify the prediction uncertainty. The covariance update can be represented by the following formula: where, represents the state covariance matrix at time t. represents the transpose of matrix A. represents the process noise covariance matrix, reflecting the random perturbation in the motion process.

[0038] After the target detection module outputs the new target position, the predicted state is combined with the observed data through the Kalman filter to update the state estimation. This update process uses the Kalman gain to balance the predicted value and the observed value. The update formula can be represented by the following formula: where, Represents the updated state vector at time t+1. Represents the Kalman gain, which can be expressed by the following formula: . R Represents the observation noise covariance matrix. Represents the target position observed at time t+1. Is the observation matrix that maps the state vector to the observation space.

[0039] Repeat the above prediction and update steps, record the target states of multiple consecutive frames in chronological order, and form a target population spatio-temporal trajectory matrix. Each row or each component of this matrix corresponds to the position and velocity information of a certain target at consecutive time points, thereby achieving precise tracking of the dynamic movement of the population.

[0040] Step S4: Perform sound source localization and beam weight calculation processing on the audio feature tensor according to the target population spatio-temporal trajectory matrix to obtain a beamforming weight vector.

[0041] In the microphone matrix, there is a time difference in the signals received by each microphone. This time difference directly reflects the order of arrival of sound waves at each microphone and the propagation path. The data in the audio feature tensor is first processed in the frequency domain to extract the phase information between different microphones, and the generalized cross-correlation algorithm is used to process the collected audio signals, thereby calculating the time delay difference between each pair of microphones. This time delay difference is the delay difference in the signals received by different microphones when the sound source propagates in different directions in space. During the processing, the signals collected by all microphones are transformed to the frequency domain through Fourier transform, the complex representation of each frequency component is calculated, then the signals of each pair of microphones are multiplied and normalized, and finally the time domain response of the time delay difference is obtained through inverse Fourier transform. All calculated time delay differences correspond to the possible arrival directions of the sound source, and after analysis, a set of candidate sound source directions is generated. Each candidate direction corresponds to a combination of time delay differences, and its value reflects the position angle of the sound source relative to the microphone array. Specifically, the signals from each microphone in the audio feature tensor are extracted, and short-time Fourier transform (STFT) is performed on each signal to convert the time domain signal to a frequency domain representation. For any pair of microphones i and j, calculate their cross-spectrum as , where Represents the frequency domain signal of the i-th microphone, Represents the conjugate of the j-th microphone signal. For the obtained cross-spectrum Perform normalization processing to form a spectral function that preserves phase information , where Represents The modulus length of. Perform inverse Fourier transform on the normalized spectral function to obtain the time delay difference function , the peak of the time delay difference function at the time delay corresponds to the optimal time delay difference between the received signals of microphone pair i, j. The calculated time delay differences for all microphone pairs are statistically analyzed and combined to form a candidate set of sound source directions. Each direction in the candidate set is corresponding to a set of time delay difference values, and its value is further mapped to the actual angle, usually by using the geometric layout relationship of the microphone array to achieve angle conversion.

[0042] Furthermore, by matching the candidate set of sound source directions with the spatio-temporal trajectory matrix of the target population, further screening and confirmation of the sound source direction are achieved. The spatio-temporal trajectory matrix of the target population provides the position information of each detected target in consecutive frames, and this position information can be collected by a camera and obtained through target detection and tracking processing. Using this information, each direction in the candidate set of sound source directions can be compared with the motion direction of the target obtained by visual detection, and an angle screening strategy is adopted to regard the direction whose angle with the visual target direction in the candidate directions is less than a preset threshold (for example, 10°) as the final target direction. This process first maps the candidate direction to convert the time delay difference into the actual angle; then, extracts the average motion direction of the target from the spatio-temporal trajectory matrix of the target population; finally, calculates the angle between the candidate direction and the target direction, and screens out the candidate directions that meet the conditions. Specifically, convert the peak of the time delay difference function into the sound source direction angle, and through the geometric configuration of the microphone array, use the known microphone spacing and the speed of sound to calculate the conversion formula to map the time delay difference to the angle space. Extract the motion direction of each target from the spatio-temporal trajectory matrix of the target population, calculate the displacement vector of each target in consecutive frames, and use the vector averaging method to obtain one or more representative target motion directions. For each direction in the candidate sound source directions, compare it with the extracted target motion direction and calculate the angle between the two. Screen the angle, set a threshold (for example, 10°), and only when the angle between the candidate direction and the target direction is less than this threshold, the candidate direction is regarded as matching the target. The finally selected candidate direction is the target direction, which is used for subsequent beamforming weight calculation processing. The angle calculation can be expressed by the following formula: where represents the candidate sound source direction angle. represents the visual target direction extracted from the spatio-temporal trajectory matrix of the target population. represents the angle between the candidate direction and the visual direction. represents the speed of sound. represents the time delay difference obtained by GCC-PHAT. represents the microphone spacing.

[0043] Finally, the determined target direction is used as the steering direction, and the beamforming weight vector is calculated based on the Minimum Variance Distortionless Response (MVDR) criterion. The basic goal of the MVDR beamforming method is to minimize the power of the entire beam output while keeping the signal in the target direction undistorted, so as to suppress the noise and interference from other directions. The MVDR beamforming method relies on calculating the covariance matrix of the sound field and the steering vector corresponding to the target direction, and obtains the weighting coefficients of each microphone by solving an optimization problem, thereby forming an optimal beamforming weight vector. Specifically, first, according to the audio signals collected by the microphone array, the covariance matrix \(R\) of the sound field is calculated. The covariance matrix reflects the correlation between the signals received by each microphone and is usually calculated from the audio feature tensor. According to the target direction, the steering vector is calculated. The calculation of the steering vector depends on the geometric structure of the microphone array and the acoustic wave propagation characteristics, and can be expressed by the following formula: where, represents the time delay of the acoustic wave from the target direction to the th microphone, represents the number of microphones.

[0044] Use the MVDR criterion to solve the beamforming weight vector. The optimization goal is to minimize the output power while ensuring that the signal in the target direction is undistorted, that is, to satisfy the constraint (where represents the conjugate transpose of ). The formula for the obtained weight vector is as follows: where, represents the beamforming weight vector. represents the covariance matrix of the sound field, which is obtained by calculating the audio feature tensor. represents the inverse matrix of. represents the steering vector corresponding to the target direction . represents the conjugate transpose of . On the premise of ensuring that the signal in the target direction is undistorted, the weighting coefficients of each microphone are obtained by solving the optimization problem, thereby forming an optimal beamforming weight vector to achieve the purpose of suppressing noise and interference and enhancing the target signal.

[0045] The calculated weight vector \(w\) is a complex vector, and each component corresponds to the gain and phase adjustment of each microphone. This vector, as the core parameter of beamforming, can keep the signal undistorted in the target direction while suppressing the noise in other directions.

[0046] Step S5: Perform multi-constraint dynamic optimization processing on the beamforming weight vector according to the environmental parameters in the multi-modal sensor data to obtain an adaptive control parameter set.

[0047] In practical applications, environmental noise exhibits dynamic changes in time and space. Through a smoothing update algorithm, the impact of instantaneous noise fluctuations can be reduced, and a stable covariance matrix that reflects the current actual situation can be obtained. In this way, when calculating the beamforming weights subsequently, noise can be more precisely suppressed, ensuring that the target-direction signal passes through without distortion, thereby achieving an optimized sound field output effect. The beamforming weight vector calculated from the collected audio signals is optimized using environmental parameters (such as real-time temperature, real-time humidity, etc.). The noise covariance matrix reflects the statistical correlation between noise signals in multiple microphone channels. This matrix is used to describe the distribution of noise at different spatial positions, and its update process requires combining historical noise information with currently collected noise data to more accurately reflect the statistical characteristics of the current environmental noise. Specifically, within a determined time window, the system samples the audio signals of each microphone using real-time environmental parameters, calculates the covariance value of the current sample, and then uses a smoothing factor to fuse and update this value with the previous covariance matrix. This process uses a recursive algorithm to ensure that the noise covariance matrix can quickly track environmental changes and provide an accurate statistical basis for subsequent beam weight calculations. Specifically, for the audio signals from the microphone array, the noise sample covariance matrix at the current moment is determined through the environmental noise parameters in the multi-modal data. The calculation process of this matrix is based on the correlation between channels in the sampled data and is obtained through matrix operations.

[0048] Exemplarily, assume a smoothing factor α (e.g., α = 0.9) for balancing the influence of historical noise information and current noise samples , then the currently updated noise covariance matrix is calculated through the following recursive formula: where, represents the updated noise covariance matrix at time . represents the noise covariance matrix at the previous moment. represents the sample covariance matrix calculated at the current moment. is the smoothing factor, whose value range is usually between 0 and 1, controlling the weight ratio between historical data and new data. The recursive formula is used to perform smooth fusion between real-time environmental noise data and historical noise information, update the noise covariance matrix, and provide a statistical basis for subsequent weight optimization.

[0049] After calculating the updated covariance matrix, it is used as an accurate representation of the noise characteristics in the current environment, providing a basis for the correction of the temperature attenuation factor in subsequent steps. This matrix records the correlation characteristics between the microphone signals under the influence of the current environmental parameters, and its accuracy directly affects the dynamic optimization effect of the beamforming weights.

[0050] The temperature attenuation factor is used to correct the previously updated noise covariance matrix according to the real-time temperature, further adjusting the beamforming weights. Temperature has a direct impact on the propagation of sound waves and the performance of devices. Especially in high-temperature environments, the frequency response and amplification characteristics of speakers will change, affecting the sound field output. Therefore, a temperature attenuation factor is introduced based on the covariance matrix, and the weights are corrected through a formula, enabling the output weights to automatically adapt to environmental changes under different temperature conditions, thus ensuring the accuracy of sound field directional control. Specifically, real-time temperature information is extracted from multi-modal sensor data, which is measured by environmental sensors in real time and features high precision and timely updates. According to the design principle of the temperature attenuation factor, a reference temperature (usually set to 25 °C) and an attenuation coefficient β (e.g., 0.05) obtained through experimental verification are set. The attenuation factor represents the impact on the gain of high-frequency signals for every 1 °C increase in temperature. The weights reflected in the covariance matrix are corrected using the temperature attenuation formula. The correction process can be expressed by the following formula: where, represents the corrected beamforming weights. represents the original beamforming weight vector (or the preliminary weights optimized from the covariance matrix). represents the current real-time temperature. represents the reference temperature, usually 25 °C. represents the temperature attenuation coefficient, determined through experiments, reflecting the proportion of weight adjustment for every 1 °C increase in temperature (e.g., 0.05). The correction process is used to correct the original weights according to the real-time temperature, ensuring that the weights decay proportionally when the temperature is higher than the reference value and relatively increase when the temperature is lower than the reference value, thus adaptively compensating for the impact of environmental temperature changes on sound wave propagation and device performance.

[0051] After being corrected by the temperature attenuation factor, the output corrected weights can better reflect the actual impact of the current environmental temperature on signal processing, thus ensuring the distortion-free transmission of the target signal in the subsequent beam synthesis process.

[0052] Finally, normalize the beamforming weight vector after being corrected by the temperature attenuation factor to ensure that the amplitudes of the output signals of each microphone are within a predetermined range, preventing system instability caused by individual weights being too large or too small. The normalization constraint process calculates the modulus length of each component during the calculation and scales all weights proportionally so that the maximum weight does not exceed a preset threshold (e.g., 1.2). This process aims to keep the entire beamforming control system stable in a dynamic environment and effectively prevent hardware overload or signal distortion caused by weights that are too large or too small. Specifically, obtain the weight vector after being corrected by the temperature attenuation factor. The weight vector is a complex vector, and each element corresponds to the gain and phase of a microphone. Further calculate the modulus length of all elements in the weight vector and determine the maximum value among them. Set a normalization threshold. If the maximum value exceeds the normalization threshold, scale the entire vector, and the scaling factor is the normalization threshold / maximum value. Perform a normalization operation on each element in the weight vector to obtain the finally normalized weight vector. The normalized weight vector is the set of adaptive control parameters for subsequent beam synthesis processing.

[0053] Step S6: Perform real-time beam synthesis processing on the set of adaptive control parameters according to a preset speaker sequence to obtain a directional sound field output signal.

[0054] Convert the beamforming weight vector after normalization and temperature correction in the set of adaptive control parameters into an adjustment signal for each speaker through carrier modulation processing. The preset speaker sequence defines the specific positions and arrangements of each speaker in physical space. The basic principle of carrier modulation processing is to multiply the weight by a sine or complex exponential carrier signal on each speaker channel, thereby introducing necessary phase shifts and amplitude adjustments. This process enables each speaker to adjust the amplitude and phase of its sound output according to the beamforming requirements when outputting a signal, forming an overall directional sound field. Specifically, first, extract the corrected weight vector from the set of adaptive control parameters. The weight vector represents the optimal weighting coefficients of each microphone or speaker channel. According to the preset speaker sequence, determine the number and position in space of each speaker. The preset speaker sequence is determined during the system design phase and fixed in the device configuration. Set the carrier frequency, which is selected according to the frequency band of the desired sound field (e.g., 1 kHz), and at the same time determine the modulation clock t as the current moment of the system. For each speaker channel, calculate its modulation signal. The modulation process uses the carrier modulation formula to multiply the corrected weight by the carrier signal. The carrier modulation processing can be represented by the following formula: where, represents the th speaker channel at time The output modulated signal. Represents the corrected weight of the th channel in the adaptive control parameter set, and this weight has undergone the aforementioned temperature correction and normalization processing. Represents the th input signal of the speaker, usually a preset baseband signal or the pre-modulation signal. Is the carrier signal, where Represents the carrier frequency, Is the time variable, Represents the imaginary unit. After the carrier modulation process modulates the corrected weight with the corresponding input signal and the carrier signal, the modulated signal of each speaker is formed. The modulated signal will carry the corresponding amplitude and phase information, serving as the basic data for the subsequent coherent superposition of multiple channels.

[0055] Repeat the above process to calculate the modulated signal for each preset speaker channel in turn. In physical implementation, the modulated signal will directly drive the output of the speaker to form the sound wave emission of each channel. Finally, all the individually modulated speaker signals are stored in the buffer for use in subsequent steps.

[0056] After obtaining the adjustment signals of each speaker channel, the modulated signals of each speaker channel are coherently superposed to form an overall directional sound field output signal. The coherent superposition process is the core of beamforming, aiming to adjust the phase and amplitude of each channel signal so that at the target direction, the wave peaks of each signal coincide, thus forming a signal enhancement effect; while at non-target directions, each signal cancels each other out due to phase misalignment, achieving the effect of suppressing interference. This process requires that the physical layout of the speakers must be determined in advance, and the signals of each channel need to be strictly superposed after time synchronization and modulation processing according to the preset. Specifically, obtain the modulated signal y k (t) of each speaker channel, where k = 1, 2, ……, M, and M is the total number of speakers. According to the preset speaker sequence information, the output time and phase information of each channel signal have been matched with the carrier modulation process. Perform point-by-point summation on all channel signals and calculate the overall superposed signal using the following formula: Where, Represents the directional sound field output signal obtained after superposition. Represents the th speaker channel at time The output modulated signal. Represents the total number of preset speakers.

[0057] After coherent superposition, the output signal of the directional sound field contains the phase and amplitude information of all speaker signals in the target direction. Since the signals of each channel have been phase-adjusted according to the preset carrier during the modulation process, this summation process ensures that all signals are superimposed and enhanced in the target direction, while interference suppression occurs in other directions due to phase differences.

[0058] In order to promptly detect the situation of power exceeding the standard caused by environmental interference or equipment anomalies, and then take measures to adjust the control parameters to ensure that the output signal will not cause damage to the equipment or lead to distortion. After obtaining the output signal of the directional sound field, drive the speakers in the target area according to the output signal of the directional sound field, and at the same time, monitor the power spectrum of the output signal of the directional sound field in real time to ensure that the power distribution of the output signal in each frequency band meets the design requirements, and promptly detect the situation of power exceeding the standard caused by environmental changes or equipment anomalies. By performing fast Fourier transform processing on the output signal of the directional sound field, the power distribution of the signal in the frequency domain, that is, the power spectrum, can be obtained. Real-time monitoring of the power spectrum is crucial for dynamic control and fault protection. When the power of the output signal in one or more frequency bands exceeds the preset safety threshold, subsequent steps should be triggered to dynamically optimize and update the control parameters to prevent equipment overload or signal distortion. Specifically, input the output signal of the directional sound field into the spectrum analysis module, and use the fast Fourier transform to convert the time-domain signal into a frequency-domain representation. Then calculate the square of the modulus value of the frequency-domain representation of the signal to obtain the energy value of each frequency component. And convert the energy value to decibel units to form the power spectrum. The power spectrum calculation can be expressed by the formula shown below: where, represents the power spectrum value (in decibels dB) at frequency . represents the frequency-domain representation obtained by converting the output signal of the directional sound field through fast Fourier transform. represents 's modulus length. The power spectrum is used to convert the directional sound field signal into a frequency-domain power spectrum representation for real-time monitoring of the power distribution of the signal in each frequency band.

[0059] Furthermore, compare the obtained real-time control power spectrum values of each frequency band with the preset power threshold. Among them, the power threshold is the predetermined safety value in the design (such as 90 dB) to determine whether the output signal is within the safe range.

[0060] When it is monitored in real time that the power spectrum of the directional sound field output signal exceeds the safety threshold, the control parameter update process is immediately triggered. When the output signal power exceeds the standard, by recalculating and adjusting the beamforming weights, the adaptive control parameter set is optimized, so as to reduce the output signal power and achieve the goal of safety and stability. The update process is carried out according to the preset interruption weight update strategy to ensure that it can automatically return to the design range. Specifically, when the real-time control power spectrum exceeds the preset threshold in one or more frequency bands, an interruption signal is immediately generated to trigger the weight adjustment and optimization process. The interruption signal pauses the current beam synthesis output and saves the state of the current electronic device (including the current weight vector, target direction, environmental parameters, etc.) for subsequent resumption. At the same time, the electronic device re-executes the sound source localization and weight calculation processing process (that is, the above steps S4 to step S5) to obtain a new beamforming weight vector. And the updated weight vector is normalized and then replaces the original adaptive control parameter set, so as to further obtain the updated real-time control power spectrum through steps S61 to S63, and then verify the power threshold according to the updated real-time control power spectrum. When the power spectrum values in the real-time control power spectrum are all less than or equal to the power threshold, the closed-loop feedback process terminates and enters the stable working state. When there are power spectrum values greater than the power threshold in the real-time control power spectrum, the above real-time control power spectrum update process is repeated until the safety requirements are met.

[0061] This application is applied to the field of sound beam control technology. By synchronously collecting and calibrating multi-modal sensor data in space and time, a raw data set with space-time alignment is obtained. The raw data set is subjected to hierarchical signal enhancement to obtain a visual feature tensor and an audio feature tensor. Then, based on the visual feature tensor, dense crowd detection and tracking are performed to obtain a target crowd space-time trajectory matrix. Combining with the audio feature tensor, sound source localization and beam weight calculation are carried out to obtain a beamforming weight vector. Furthermore, multi-constraint dynamic optimization is carried out in combination with the environmental parameters in the multi-modal sensor data to obtain an adaptive control parameter set. Finally, real-time beam synthesis is carried out in combination with a speaker sequence to obtain a directional sound field output signal. This application organically combines the space-time alignment, signal enhancement, target detection, sound source localization and dynamic adaptive control of multi-modal sensor data, improves the pointing accuracy of sound beamforming, and enhances the adaptability to environmental changes.

[0062] As Figure 2 shown, it is a functional module diagram of a sound beamforming pointing control device based on crowd recognition provided by an embodiment of this application.

[0063] In some embodiments, the voice beamforming pointing control device 2 based on crowd recognition may include a plurality of functional modules composed of computer program segments. The computer programs of each program segment in the voice beamforming pointing control device 2 based on crowd recognition may be stored in the memory of the server and executed by at least one processor to execute (see details in Figure 1 the description) the functions of the voice beamforming pointing control method based on crowd recognition.

[0064] In this embodiment, according to the functions it performs, the voice beamforming pointing control device 2 based on crowd recognition can be divided into a plurality of functional modules. The functional modules may include: a spatio-temporal alignment module 21, a hierarchical gain module 22, a visual feature module 23, an audio feature module 24, a dynamic optimization module 25, and a beam synthesis module 26. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0065] The spatio-temporal alignment module 21 is used to perform spatio-temporal synchronous acquisition and calibration processing on multi-modal sensor data to obtain a spatio-temporally aligned original data set.

[0066] In an alternative embodiment, the spatio-temporal alignment module 21 is specifically used for: Performing real-time environment detection processing on the target area according to a preset synchronous acquisition method to obtain the audio signal, the visual signal, and the environmental parameters with a unified initial timestamp; Performing timestamp attachment processing on the audio signal, the visual signal, and the environmental parameters to obtain time-tagged multi-source data; Performing asynchronous sampling interpolation compensation processing on the time-tagged multi-source data to obtain a temporally aligned intermediate data set; Performing sensor calibration processing on the intermediate data set to obtain the original data set.

[0067] The hierarchical gain module 22 is used to perform hierarchical signal enhancement processing on the original data set to obtain a visually feature tensor and an audio feature tensor with optimized signal gain.

[0068] In an alternative embodiment, the hierarchical gain module 22 is specifically used for: Performing audio-visual signal classification processing on the original data set to obtain an original audio signal and an original visual signal; Performing Wiener filtering noise reduction processing on the original audio signal to obtain a noise-reduced and optimized audio signal; Perform Mel-spectrum modulation and demodulation processing on the noise-reduced and optimized audio signal to obtain the audio feature tensor; Perform image distortion correction processing on the original visual signal to obtain a corrected image; Perform spatial attention weighting processing on the corrected image to obtain the visual feature tensor.

[0069] The visual feature module 23 is used to perform dense crowd detection and tracking processing according to the visual feature tensor to obtain the target crowd spatio-temporal trajectory matrix.

[0070] In an optional embodiment, the visual feature module 23 is specifically used for: Perform Anchor box size adjustment processing on the preset YOLO v5 model according to the visual feature tensor to obtain an improved YOLO v5 model; Perform target detection processing on the visual feature tensor through the improved YOLO v5 model to obtain a candidate target set; Perform occlusion-aware loss compensation processing on the candidate target set to obtain an optimized target position prediction result; Perform Kalman filtering processing on the target position prediction result to obtain the target crowd spatio-temporal trajectory matrix.

[0071] The audio feature module 24 is used to perform sound source localization and beam weight calculation processing on the audio feature tensor according to the target crowd spatio-temporal trajectory matrix to obtain a beamforming weight vector.

[0072] In an optional embodiment, the audio feature module 24 is specifically used for: Perform generalized cross-correlation processing on the audio feature tensor to obtain a sound source direction candidate set; Perform visual trajectory matching and angle screening processing on the sound source direction candidate set according to the target crowd spatio-temporal trajectory matrix to obtain a target direction; Perform minimum variance distortionless response weight calculation processing according to the target direction to obtain the beamforming weight vector.

[0073] The dynamic optimization module 25 is used to perform multi-constraint dynamic optimization processing on the beamforming weight vector according to the environmental parameters in the multi-modal sensor data to obtain an adaptive control parameter set.

[0074] In an optional embodiment, the dynamic optimization module 25 is specifically used for: Perform noise covariance matrix update processing on the beamforming weight vector according to the environmental parameters to obtain a covariance matrix; Perform temperature attenuation factor correction processing on the covariance matrix according to the real-time temperature in the environmental parameters to obtain a correction weight; Perform normalization constraint processing on the correction weight to obtain the adaptive control parameter set.

[0075] The beamforming module 26 is configured to perform real-time beamforming processing on the adaptive control parameter set according to a preset speaker sequence to obtain a directional sound field output signal.

[0076] In an alternative embodiment, the beamforming module 26 is specifically configured to: Step S61: Perform carrier modulation processing on the adaptive control parameter set according to the speaker sequence to obtain adjustment signals for each speaker channel; Step S62: Perform coherent superposition processing on the adjustment signals for multiple speaker channels according to the speaker sequence to obtain the directional sound field output signal; Step S63: Perform real-time power spectrum monitoring processing on the directional sound field output signal to obtain a real-time control power spectrum; Step S64: When there is a power spectrum value in the real-time control power spectrum that is greater than a preset power threshold, perform weight adjustment and optimization processing according to a preset interruption weight update method to update the adaptive control parameter set; Repeat steps S61 to S64 until the power spectrum values in the real-time control power spectrum are all less than or equal to the power threshold.

[0077] It should be understood that the various change modes and specific embodiments in the methods provided in the above embodiments are equally applicable to the sound beamforming pointing control device based on crowd recognition in this embodiment. Through the foregoing detailed description of the sound beamforming pointing control method based on crowd recognition, those skilled in the art can clearly know the implementation method of the sound beamforming pointing control device based on crowd recognition in this embodiment. For the sake of brevity of the specification, it will not be elaborated here.

[0078] As Figure 3 shown, it is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0079] In a preferred embodiment of the present invention, the electronic device 3 may include, but is not limited to: a memory 31, at least one processor 32, and at least one communication bus 33.

[0080] Those skilled in the art should understand that Figure 3 the structure of the electronic device 3 shown does not constitute a limitation on the embodiments of the present invention. The electronic device 3 may further include more or fewer other hardware or software than shown, or different component arrangements.

[0081] In some embodiments, the electronic device 3 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits, programmable gate arrays, digital signal processors, and embedded devices, etc.

[0082] It should be noted that the electronic device 3 is only an example, and other existing or future electronic products that can be adapted to this application should also be included within the protection scope of this application and are included herein by reference.

[0083] In some embodiments, a computer program is stored in the memory 31, and when the computer program is executed by the at least one processor 32, all or part of the steps in the above-mentioned method for controlling the pointing of a sound beamforming based on crowd recognition are implemented. The memory 31 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disc memories, tape memories, or any other computer-readable medium that can be used to carry or store data. Further, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area can store an operating system, application programs required for at least one function, etc.

[0084] In some embodiments, the at least one processor 32 is the control core (Control Unit) of the electronic device 3, connecting various components of the entire electronic device 3 through various interfaces and lines. By running or executing programs or modules stored in the memory 31, and by invoking data stored in the memory 31, it performs various functions of the electronic device 3 and processes data. For example, when the at least one processor 32 executes the computer program stored in the memory 31, it implements all or part of the steps of the method for controlling the sound beamforming direction based on crowd recognition described in the embodiments of the present application; or implements all or part of the functions of the device for controlling the sound beamforming direction based on crowd recognition. The at least one processor 32 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions packaged together, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc.

[0085] In some embodiments, the at least one communication bus 33 is configured to enable connection communication between the memory 31 and the at least one processor 32, etc. Although not shown, the electronic device 3 may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 32 through a power management device, so as to implement functions such as management of charging, discharging, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 3 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0086] The above-mentioned integrated unit implemented in the form of software function modules can be stored in a computer-readable storage medium. The above-mentioned software function modules are stored in a storage medium, including several instructions for causing an electronic device (which may be a personal computer, an electronic device, or a network device, etc.) or a processor to execute part of the methods described in the various embodiments of the present application.

[0087] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0088] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, and it may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0089] The above are all preferred embodiments of this application. The protection scope of this application is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of this application shall be covered within the protection scope of this application.

Claims

1. A sound beamforming directional control method based on crowd recognition, characterized in that: The method comprises: Perform spatiotemporal synchronous acquisition and calibration of multimodal sensor data to obtain spatiotemporal aligned raw data sets; Performing hierarchical signal enhancement processing on the original data set to obtain a visual feature tensor and an audio feature tensor after signal gain optimization; Performing dense crowd detection and tracking processing according to the visual feature tensor to obtain a spatiotemporal trajectory matrix of the target crowd; Performing sound source localization and beam weight calculation processing on the audio feature tensor according to the target population spatiotemporal trajectory matrix to obtain a beamforming weight vector; Performing multi-constraint dynamic optimization processing on the beamforming weight vector according to the environmental parameters in the multimodal sensor data to obtain an adaptive control parameter set; The adaptive control parameter set is processed by real-time beam synthesis according to a preset loudspeaker sequence to obtain a directional sound field output signal.

2. The sound beamforming directional control method based on crowd recognition according to claim 1 is characterized in that: The multimodal sensor data includes audio signals, visual signals and environmental parameters, and the step of performing spatiotemporal synchronous acquisition and calibration processing on the multimodal sensor data to obtain a spatiotemporal aligned original data set includes: Performing real-time environmental detection processing on the target area according to a preset synchronous acquisition method to obtain the audio signal, the visual signal and the environmental parameters with a unified initial timestamp; Performing time stamp addition processing on the audio signal, the visual signal and the environmental parameter to obtain time-tagged multi-source data; Performing asynchronous sampling interpolation compensation processing on the time-tagged multi-source data to obtain a time-domain aligned intermediate data set; The intermediate data set is subjected to sensor calibration processing to obtain the original data set.

3. The sound beamforming directional control method based on crowd recognition according to claim 1 is characterized in that: The performing hierarchical signal enhancement processing on the original data set to obtain a visual feature tensor and an audio feature tensor after signal gain optimization includes: Performing audio-visual signal classification processing on the original data set to obtain an original audio signal and an original visual signal; Performing Wiener filtering noise reduction processing on the original audio signal to obtain a noise reduction optimized audio signal; Performing Mel spectrum modulation and demodulation processing on the noise reduction optimized audio signal to obtain the audio feature tensor; Performing image distortion correction processing on the original visual signal to obtain a corrected image; The rectified image is subjected to spatial attention weighted processing to obtain the visual feature tensor.

4. The sound beamforming directional control method based on crowd recognition according to claim 1 is characterized in that: The performing dense crowd detection and tracking processing according to the visual feature tensor to obtain the spatiotemporal trajectory matrix of the target crowd includes: Performing anchor box size adjustment processing on a preset YOLO v5 model according to the visual feature tensor to obtain an improved YOLO v5 model; Performing target detection processing on the visual feature tensor through the improved YOLO v5 model to obtain a candidate target set; Performing occlusion perception loss compensation processing on the candidate target set to obtain an optimized target position prediction result; The target position prediction result is subjected to Kalman filtering to obtain the spatiotemporal trajectory matrix of the target population.

5. The sound beamforming directional control method based on crowd recognition according to claim 1 is characterized in that: The performing sound source localization and beam weight calculation processing on the audio feature tensor according to the target group spatiotemporal trajectory matrix to obtain a beamforming weight vector includes: Performing generalized cross-correlation processing on the audio feature tensor to obtain a sound source direction candidate set; Performing visual trajectory matching and angle screening processing on the sound source direction candidate set according to the target population spatiotemporal trajectory matrix to obtain the target direction; Minimum variance distortion-free response weight calculation processing is performed according to the target direction to obtain the beamforming weight vector.

6. The sound beamforming directional control method based on crowd recognition according to claim 1 is characterized in that: The performing multi-constraint dynamic optimization processing on the beamforming weight vector according to the environmental parameters in the multimodal sensor data to obtain an adaptive control parameter set includes: Performing noise covariance matrix update processing on the beamforming weight vector according to the environmental parameters to obtain a covariance matrix; Performing temperature attenuation factor correction processing on the covariance matrix according to the real-time temperature in the environmental parameter to obtain a correction weight; The modified weights are subjected to normalization constraint processing to obtain the adaptive control parameter set.

7. The sound beamforming directional control method based on crowd recognition according to claim 6 is characterized in that: The performing real-time beam synthesis processing on the adaptive control parameter set according to the preset speaker sequence to obtain a directional sound field output signal comprises: Step S61, performing carrier modulation processing on the adaptive control parameter set according to the speaker sequence to obtain adjustment signals of each speaker channel; Step S62, performing coherent superposition processing of multiple speaker channels on the adjustment signal according to the speaker sequence to obtain the directional sound field output signal; Step S63, performing real-time monitoring and processing of the power spectrum of the directional sound field output signal to obtain a real-time control power spectrum; Step S64: when there is a power spectrum value in the real-time control power spectrum that is greater than a preset power threshold, weight adjustment optimization processing is performed according to a preset interruption weight update method to update the adaptive control parameter set; The step S61 to the step S64 are repeatedly executed until the power spectrum values ​​in the real-time control power spectrum are all less than or equal to the power threshold.

8. A sound beamforming pointing control device based on crowd recognition, characterized in that: The device comprises: The spatiotemporal alignment module is used to perform spatiotemporal synchronous acquisition and calibration processing on multimodal sensor data to obtain a spatiotemporal aligned original data set; A hierarchical gain module, used for performing hierarchical signal enhancement processing on the original data set to obtain a visual feature tensor and an audio feature tensor after signal gain optimization; A visual feature module, used for performing dense crowd detection and tracking processing according to the visual feature tensor to obtain a spatiotemporal trajectory matrix of the target crowd; An audio feature module, used for performing sound source localization and beam weight calculation processing on the audio feature tensor according to the target population spatiotemporal trajectory matrix to obtain a beamforming weight vector; A dynamic optimization module, configured to perform multi-constraint dynamic optimization processing on the beamforming weight vector according to the environmental parameters in the multimodal sensor data to obtain an adaptive control parameter set; The beam synthesis module is used to perform real-time beam synthesis processing on the adaptive control parameter set according to a preset speaker sequence to obtain a directional sound field output signal.

9. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the sound beamforming pointing control method based on crowd recognition according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the sound beamforming pointing control method based on crowd recognition according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Audio system fault detection method and device based on dual-mode audio recovery

    CN120980433A

  • Nondestructive testing method for evaluating fatigue characteristics of long-life pavement material

    CN122084881A

  • A non-destructive testing method for evaluating the fatigue properties of long-life pavement materials

    CN122084881B