Directional sound fusion analysis system and method based on voiceprint splicing

By collecting data on visitor status and environmental noise, and dynamically adjusting the beam direction of the ultrasonic array, the problems of decreased voice clarity and sound wave overlap interference in the directional broadcasting system were solved, thus achieving high-quality personalized tour guide services.

CN122372896APending Publication Date: 2026-07-10LANGFANG HONGSHU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LANGFANG HONGSHU TECHNOLOGY CO LTD
Filing Date
2026-04-02
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing directional broadcasting systems cannot comprehensively perceive visitor behavior characteristics and changes in environmental noise, resulting in decreased speech clarity and sound pollution. Furthermore, when multiple exhibition areas are arranged adjacently, sound wave overlap and interference are likely to occur.

Method used

By collecting tourist status data and environmental noise data, and combining the status of adjacent broadcasting units, the reflected sound interference factor is calculated, the beam direction and main lobe width of the ultrasonic array are dynamically adjusted, and the optimal sound status parameters are output to achieve directional projection.

Benefits of technology

Under varying crowd density and ambient noise levels, the system ensures visitors have a clear and comfortable auditory experience, resolves acoustic conflicts across multiple exhibition areas, and improves speech clarity and system adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122372896A_ABST
    Figure CN122372896A_ABST
Patent Text Reader

Abstract

This invention discloses a directional sound fusion analysis system and method based on voiceprint splicing, belonging to the field of directional sound fusion technology. Based on visitor status data within the target area, this invention distinguishes the current scene's browsing mode, determines the interference risk status, acquires the exhibition hall's spatial data and material information, calculates the reflected sound and its loudness distribution at different positions and angles reaching the target point according to the relative orientation of the visitor's position coordinates and the ultrasonic transducer array, obtains the reflection interference factor, outputs the optimal sound state parameters, calculates the beam pointing angle based on the visitor's real-time coordinates, and adjusts the main lobe width of the ultrasonic array under the constraint of the optimal sound state parameters for directional projection. This invention ensures that the optimal sound state of the broadcasting unit can be clearly projected to the target area, while minimizing crosstalk to adjacent areas, providing visitors with a high-quality guided tour service in the optimal listening area with appropriate volume and no interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of directional sound fusion technology, and in particular to a directional sound fusion analysis system and method based on voiceprint splicing. Background Technology

[0002] In exhibition venues such as museums, art galleries, and science and technology museums, providing clear, accurate, and personalized audio guides is a core element in enhancing the visitor experience. Traditional guide methods, such as human explanations, handheld audio guides, or uniform explanations based on public address systems, often result in problems such as sound interference and fragmented experiences. Therefore, directional sound technology is widely used to provide independent audio narration for specific exhibition areas, aiming to achieve a visitor experience where sound stops as visitors move around and does not interfere with each other.

[0003] However, existing directional broadcasting systems typically rely solely on simple infrared or pressure sensors to trigger playback. They fail to comprehensively perceive visitor crowd density, flow speed, and other behavioral characteristics, and struggle to monitor changes in ambient background noise in real time. This prevents the system from automatically adjusting its broadcasting strategy based on actual conditions, resulting in a somewhat rigid user experience. Furthermore, in scenarios where multiple exhibition areas are arranged adjacently, the lack of communication and status interaction between units means that if adjacent units are activated simultaneously while one broadcasting unit is in operation, their sound wave coverage areas can easily overlap, leading to sound field superposition, standing waves, or reverberation. This causes a significant decrease in speech clarity and noise pollution, interfering with the visitor's auditory experience. Moreover, existing technologies often assume an ideal free sound field, neglecting the influence of the specific geometry of the exhibition hall and wall materials on sound wave reflection. In practical applications, ultrasonic directional beams hitting glass, metal, or concrete surfaces produce significant reflected sound. If these reflected sounds are not controlled, they can create interference, resulting in unclear audio.

[0004] To address the aforementioned problems, this invention provides a directional sound fusion analysis system and method based on voiceprint splicing. Summary of the Invention

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] In a first aspect, the present invention provides a directional sound fusion analysis method based on voiceprint splicing, comprising the following specific steps:

[0007] S1. Collect tourist status data within the target area, monitor environmental background noise data, and simultaneously obtain the working status of adjacent broadcasting units in real time.

[0008] S2. Based on the tourist status data within the target area, distinguish the browsing mode of the current scene, combine the spatial overlap between the tourist location and the coverage area of ​​the adjacent broadcasting unit, as well as the working status of the adjacent broadcasting unit, to determine the interference risk status.

[0009] S3. Obtain the spatial data and material information of the exhibition hall. Based on the relative orientation of the visitor's standing coordinates and the ultrasonic transducer array, calculate the reflected sound and its loudness distribution at different positions and angles reaching the target point, and obtain the reflection interference factor.

[0010] S4. Combines ambient background noise, browsing mode, interference risk status, and reflection interference factor to output the best sound status parameters.

[0011] S5. Calculate the beam pointing angle based on the tourist's real-time coordinates, and adjust the main lobe width of the ultrasonic array under the constraint of optimal sound state parameters to perform directional projection.

[0012] Preferably, step S1 includes the following specific steps:

[0013] S11. Collect tourist status data within the target area, including station coordinates, crowd density and flow speed; monitor environmental background noise data, including environmental background noise sound pressure level.

[0014] S12. Extract the target voiceprint from the environmental background noise data, obtain the playback energy of the target voiceprint, construct the target voiceprint feature sequence, obtain the standard voiceprint feature sequence of the preset playback content of the adjacent broadcasting unit, perform feature splicing and alignment of the target voiceprint feature sequence and the standard voiceprint feature sequence, calculate the voiceprint similarity between the target voiceprint and the standard voiceprint using cosine similarity, obtain the voiceprint similarity threshold and the playback energy threshold, determine the broadcasting status of the adjacent broadcasting unit, when the voiceprint similarity is greater than or equal to the voiceprint similarity threshold, determine the broadcasting status as broadcasting, when the playback energy is less than the playback energy threshold, determine the broadcasting status as standby, when the playback energy is greater than or equal to the playback energy threshold and the voiceprint similarity is less than the voiceprint similarity threshold, determine the broadcasting status as fault, and allocate the broadcasting activity coefficient to the broadcasting status.

[0015] Preferably, step S2 includes the following specific steps:

[0016] S21. Obtain the crowd density threshold and flow speed threshold. Based on the crowd density and flow speed of tourists in the target area, determine the browsing mode of the current scene. The browsing mode includes sparse flow mode, dense flow mode, sparse dwell mode and dense dwell mode. Sparse flow mode is when the crowd density is less than or equal to the crowd density threshold and the flow speed is greater than the flow speed threshold. Dense flow mode is when the crowd density is greater than the crowd density threshold and the flow speed is greater than the flow speed threshold. Sparse dwell mode is when the crowd density is less than or equal to the crowd density threshold and the flow speed is less than or equal to the flow speed threshold. Dense dwell mode is when the crowd density is greater than the crowd density threshold and the flow speed is less than or equal to the flow speed threshold. Assign an interference sensitivity coefficient to each browsing mode.

[0017] S22. Obtain the intersection area between the target area and the coverage area of ​​the adjacent broadcasting unit, filter out the number of tourists in the intersection area, and obtain the spatial overlap by dividing the number of tourists in the intersection area by the total number of tourists in the target area.

[0018] S23. Obtain the interference risk index by multiplying the broadcast activity coefficient, interference sensitivity coefficient, and spatial overlap.

[0019] Preferably, step S3 includes the following specific steps:

[0020] S31. Obtain the spatial data of the exhibition hall, construct the three-dimensional geometric model of the exhibition hall, obtain the center coordinates and principal axis pointing vector of the ultrasonic transducer array, obtain the material information of each boundary surface of the exhibition hall, and assign the corresponding material sound absorption coefficient according to the acoustic characteristics of the material information.

[0021] S32. Obtain the average ear height and standing coordinates of tourists, construct the three-dimensional coordinates of tourists, and obtain the distance from the three-dimensional coordinates of tourists to the center coordinates of the ultrasonic transducer array, which is set as the direct sound propagation distance.

[0022] S33. Obtain the mirror sound source of the ultrasonic transducer array for each boundary surface, connect the mirror sound source with the three-dimensional coordinates of the tourist, obtain the reflection point through the intersection of the line connecting the mirror sound source and the boundary surface, and obtain the sum of the distance from the center coordinates of the ultrasonic transducer array to the reflection point and the distance from the reflection point to the three-dimensional coordinates of the tourist, and set it as the reflection path length.

[0023] S34. Obtain the direct sound deviation angle by the angle between the vector pointing from the center coordinates of the ultrasonic transducer array to the three-dimensional coordinates of the tourist and the vector pointing from the principal axis; obtain the reflected sound emission angle by the angle between the vector pointing from the center coordinates of the ultrasonic transducer array to the reflection point and the vector pointing from the principal axis.

[0024] S35. Introduce the transducer directivity function, obtain the directivity gain of the direct sound by taking the value of the transducer directivity function at the direct sound deviation angle, and obtain the directivity gain of the reflection path by taking the value of the transducer directivity function at the reflected sound emission angle.

[0025] S36. Calculate the direct sound pressure level at the tourist's three-dimensional coordinates using the direct sound pressure level calculation formula, wherein the direct sound pressure level calculation formula is:

[0026]

[0027] In the formula, To reach the sound pressure level directly, As the reference sound pressure level, For the directional gain of direct sound, For the direct sound propagation distance, The air attenuation coefficient is used to convert the direct sound pressure level into direct sound energy. The sound pressure level reflected from the boundary surface to the tourist's three-dimensional coordinates is calculated using the reflected sound pressure level calculation formula, which is:

[0028] In the formula, Let i be the sound pressure level reflected from the i-th boundary surface to the tourist's three-dimensional coordinates. For the directional gain of the i-th reflection path, Let be the reflection path length reflected by the i-th boundary surface. Let be the sound absorption coefficient of the material at the i-th boundary surface. Let be the incident angle of the sound wave incident on the i-th boundary surface. The sound pressure level reflected from the boundary surface to the tourist's three-dimensional coordinates is converted into reflected sound energy. The total sound pressure level at each tourist's three-dimensional coordinates is calculated by superimposing all direct sound energy and reflected sound energy. The reflection interference factor is calculated by the difference between the total sound pressure level and the direct sound pressure level.

[0029] Preferably, step S4 includes the following specific steps:

[0030] S41. Obtain the reference playback sound pressure level by summing the ambient background noise sound pressure level and the target signal-to-noise ratio; obtain the browsing mode gain correction value by multiplying the interference sensitivity coefficient corresponding to the browsing mode with the dynamic gain correction reference value; obtain the avoidance attenuation value by multiplying the interference risk index with the avoidance attenuation coefficient; and obtain the reflection clarity compensation gain value by multiplying the reflection interference factor with the compensation coefficient.

[0031] S42. Obtain the optimal sound state parameters, i.e., the optimal playback sound pressure level, by subtracting the avoidance attenuation value from the sum of the reference playback sound pressure level, the browsing mode gain correction value, and the reflection clarity compensation gain value.

[0032] Preferably, step S5 includes the following specific steps:

[0033] S51. Obtain the standing coordinates of tourists within the target area and calculate the center of the tourist's standing coordinates, set it as the projection target point, combine it with the average ear height of tourists to construct the coordinates of the projection target point, and obtain the direction vector from the center of the array to the projection target point by subtracting the center coordinates of the ultrasonic transducer array from the coordinates of the projection target point.

[0034] S52. Calculate the horizontal and vertical deflection angles of the beam relative to the array normal using the deflection angle calculation formula. The deflection angle calculation formula is as follows:

[0035]

[0036] In the formula, It is the horizontal deflection angle. It is the vertical deflection angle. For the distance the sound wave travels, The coordinates of the projection target point, Let the center coordinates of the ultrasonic transducer array be given. The phase offset of the array elements is calculated using the phase offset calculation formula, which is:

[0037]

[0038] In the formula, Let be the phase offset of the j-th array element. The center frequency of the ultrasound. For the speed of sound wave propagation, Given the coordinates of the j-th array element relative to the array center, output the phase array based on the phase offset of all array elements;

[0039] S53. Obtain the distribution radius of tourists in the target area relative to the projection target point, and calculate the minimum beamwidth required to cover the tourists using the minimum beamwidth calculation formula. The minimum beamwidth calculation formula is as follows:

[0040]

[0041] In the formula, The radius of the distribution of tourists in the target area relative to the projection target point;

[0042] S54. Obtain array structure parameters, including the number of array elements and the spacing between array elements. Using the window function method, based on the minimum beamwidth and the preset sidelobe level, the amplitude weighting coefficient of each array element is inversely calculated, and the amplitude array is output based on the amplitude weighting coefficients of all array elements.

[0043] S55. Calculate the total gain value using the total gain value calculation formula, which is:

[0044]

[0045] In the formula, This is the total gain value. For optimal playback sound pressure level, This is a baseline correction term for air absorption;

[0046] S56. Obtain the phase array, amplitude array, optimal playback sound pressure level, and total gain value, and input them into the controller for directional projection.

[0047] Secondly, the present invention provides a directional sound fusion analysis system based on voiceprint splicing, comprising:

[0048] The data monitoring module is used to collect tourist status data within the target area, monitor environmental background noise data, and simultaneously obtain the working status of adjacent broadcasting units in real time.

[0049] The interference risk assessment module is used to distinguish the browsing mode of the current scene based on the tourist status data in the target area, and to determine the interference risk status by combining the spatial overlap between the tourist location and the coverage area of ​​the adjacent broadcasting unit, as well as the working status of the adjacent broadcasting unit.

[0050] The reflection interference factor analysis module is used to obtain spatial data and material information of the exhibition hall. Based on the relative orientation of the visitor's standing coordinates and the ultrasonic transducer array, it calculates the reflected sound and its loudness distribution at different positions and angles reaching the target point, and obtains the reflection interference factor.

[0051] The optimal sound level analysis module is used to combine ambient background noise, browsing mode, interference risk status, and reflection interference factor to output the optimal sound state parameters.

[0052] The directional projection module is used to calculate the beam pointing angle based on the real-time coordinates of the tourists, and adjust the main lobe width of the ultrasonic array under the constraints of optimal sound state parameters to perform directional projection.

[0053] Thirdly, the present invention provides a storage medium comprising stored instructions, wherein, when the instructions are executed, the device in which the storage medium is located executes the directional sound fusion analysis method based on voiceprint splicing as described above.

[0054] Fourthly, the present invention provides an electronic device, including a memory and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described above for directional sound fusion analysis based on voiceprint splicing.

[0055] Compared with the prior art, the beneficial effects of the present invention are:

[0056] This invention accurately distinguishes different browsing modes by collecting tourists' station coordinates, crowd density, and flow speed. Combined with real-time environmental background noise data, it dynamically outputs the best sound state parameters to ensure that tourists can obtain a clear and comfortable auditory experience under different crowd densities and environmental noise levels.

[0057] By introducing real-time acquisition of the working status of adjacent broadcasting units and spatial overlap calculation, the interference risk status is judged. When a high-risk overlap is detected, the system can automatically adjust the broadcasting strategy, effectively solving the acoustic conflict problem when multiple exhibition areas are open at the same time, and improving the independence and coordination of multi-area explanation.

[0058] By combining the geometric structure of the exhibition hall space with the sound absorption coefficient of the materials, the system can predict and compensate for the sound field distortion caused by reflectors such as walls and display cases. This allows the system to predict and compensate for the sound blur caused by wall reflections, effectively reducing the interference of reflected sound on direct sound and significantly improving speech clarity.

[0059] The ultrasonic array's beam pointing angle and main lobe width are dynamically adjusted based on the real-time coordinates of tourists, enabling the directional sound field to adaptively adjust according to environmental changes and tourist status, always providing tourists with high-quality guided tour services that are in the optimal listening area, at a suitable volume, and without interference. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 A schematic diagram of the directional sound fusion analysis method based on voiceprint splicing provided in an embodiment of the present invention;

[0062] Figure 2 This is a schematic diagram of the S3 process of the directional sound fusion analysis method based on voiceprint splicing provided in an embodiment of the present invention;

[0063] Figure 3 This is a schematic diagram of the S5 flow of the directional sound fusion analysis method based on voiceprint splicing provided in an embodiment of the present invention;

[0064] Figure 4 This is a schematic diagram of the structure of a directional sound fusion analysis system based on voiceprint splicing provided in an embodiment of the present invention. Detailed Implementation

[0065] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0066] In this invention, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, or apparatus. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element.

[0067] Please see Figure 1 This invention provides a method for directional sound fusion analysis based on voiceprint splicing, including the following specific steps:

[0068] S1. Collect tourist status data within the target area, monitor environmental background noise data, and simultaneously obtain the working status of adjacent broadcasting units in real time.

[0069] In this embodiment, S1 includes the following specific steps:

[0070] S11. Set the area containing the collections in the entire exhibition hall as the monitoring area. Establish a rectangular coordinate system with any fixed corner of the monitoring area as the origin, such as the southwest corner. Define the x-axis as due east and the y-axis as due north. Obtain the overall spatial parameters of the monitoring area, such as length, width, and area. The target area is a specific spatial subset of a single collection within the monitoring area. Its core function is to accurately focus the projection object, ensuring that the projected content accurately covers the collection while avoiding interference with the surrounding environment. Collect visitor status data within the target area, including standing coordinates, crowd density, and flow speed. Identify all individual visitors within the target area using cameras and obtain their images in the images. Pixel coordinates are mapped to the monitoring area coordinate system using a homography matrix. The total number of tourists in the target area is obtained by counting the number of station coordinate points. The crowd density is obtained by dividing the total number of tourists in the target area by the area of ​​the target area. The instantaneous movement rate of a single tourist is obtained by the change of station coordinates of a single tourist in continuous time frames. The average instantaneous movement rate of all tourists in the target area is used to obtain the group flow speed. Environmental background noise data is monitored, including environmental background noise sound pressure level. Instantaneous time-domain sound pressure signals are collected by acoustic sensors, and the effective value of the sound pressure is converted into a logarithmic scale sound pressure level. The calculation of the environmental background noise sound pressure level can be expressed as:

[0071]

[0072] In the formula, This is the root mean square value of the sound pressure level. The standard sound pressure level is typically taken as 20 μPa (2× );

[0073] S12. Extract the target voiceprint from the ambient background noise data, obtain the playback energy of the target voiceprint, and construct the target voiceprint feature sequence. The target voiceprint refers to the acoustic characteristics of the sound signal emitted by the adjacent broadcasting unit, which is collected from the ambient background noise. The acquisition steps are as follows: After converting the collected ambient background noise signal into a digital signal, perform bandpass filtering to filter out the audible frequency band below the ultrasonic carrier frequency. Perform frame processing on the filtered signal, with each frame having a length of 20 milliseconds. Adjacent frames overlap by half and are multiplied by a Hamming window to reduce spectral leakage. Then, use endpoint detection technology to determine the start and end positions of the sound. Calculate the short-time energy and short-time zero-crossing rate of each frame. Set a threshold based on the pure background noise during silence or historical closure. If multiple consecutive frames exceed the threshold... The value indicates the existence of a valid sound segment. All detected valid segments are spliced ​​together to obtain the target sound signal. Playback energy is used to measure the intensity of the target sound. The acquisition steps are as follows: calculate the root mean square sound pressure (RMS) of the target voiceprint: sum the squares of all sampling points within the target sound segment, divide by the number of sampling points, and then take the square root to obtain the RMS. Then, convert the RMS to sound pressure level, i.e., take the logarithmic decibel value with 20 micropascals as the reference sound pressure, to obtain the playback energy of the target voiceprint. Mel frequency cepstral coefficients (MFCCs) are extracted from the target sound signal as voiceprint features. The acquisition steps are as follows: perform a short-time Fourier transform on each frame of the signal to obtain the spectrum, pass the spectrum through a set of Mel triangular filters (24-40), and calculate the power density of each filter. The output energy is logarithmically calculated and then subjected to a discrete cosine transform. The first 13 dimensions of coefficients (excluding the energy term) are taken as the feature vector of that frame. The feature vectors of all frames are arranged in chronological order, thus forming the target voiceprint feature sequence. For subsequent alignment and comparison, the feature sequence needs to be normalized. The standard voiceprint feature sequence of the preset playback content of the adjacent broadcasting units is obtained. The adjacent broadcasting units will pre-store the audio content they will play. For each preset content, its MFCC feature sequence is extracted offline in advance, normalized, and a standard voiceprint feature sequence is obtained. The index of the content currently being played (or about to be played) by the adjacent broadcasting units is obtained through a real-time communication protocol, thereby obtaining the corresponding standard voiceprint feature sequence. The target voiceprint feature sequence and the target voiceprint feature sequence are then compared. The standard voiceprint feature sequence is spliced ​​and aligned using a Dynamic Time Warping (DTW) algorithm. The target voiceprint feature sequence is aligned with the standard voiceprint feature sequence along the time axis. The steps are as follows: The optimal matching path is found based on the Euclidean distance of the feature vectors; the target sequence is mapped onto the time axis of the standard sequence; if there are missing frames in the target sequence due to noise interference, the feature vectors at the corresponding positions in the standard sequence are used for imputation or interpolation to eliminate the impact of time delay; cosine similarity is used to calculate the voiceprint similarity between the target and standard voiceprints; and a voiceprint similarity threshold and a playback energy threshold are obtained. The playback energy threshold is used to distinguish between background noise and valid sound signals, measuring whether the sound signal has analytical value.The acquisition steps are as follows: Measure the sound pressure level of adjacent units reaching the microphone of this unit under typical operating conditions, take the average value and add a preset signal-to-noise ratio margin (6dB). Determine the broadcasting status of adjacent broadcasting units. When the voiceprint similarity is greater than or equal to the voiceprint similarity threshold, the broadcasting status is determined to be broadcasting. When the playback energy is less than the playback energy threshold, the broadcasting status is determined to be standby. Standby indicates that no valid audio signal is detected in the background sound or that feature segments related to the standard voiceprint feature sequence cannot be extracted. When the playback energy is greater than or equal to the playback energy threshold and the voiceprint similarity is less than the voiceprint similarity threshold, the broadcasting status is determined to be faulty, such as playing an incorrect file, severe distortion, stuttering, or only noise remaining. A broadcasting activity coefficient is assigned to the broadcasting status. The broadcasting activity coefficient reflects the performance of adjacent units. The actual impact of a unit on the current sound field is determined by dividing the sound pressure level of adjacent unit signals separated from the environmental background noise data by the maximum reference sound pressure level of the system design to obtain the loudness ratio. Then, a broadcast activity coefficient is obtained by weighted summation of the loudness ratio and voiceprint similarity. The weights corresponding to the broadcast activity coefficient are set according to the acoustic characteristics of the exhibition hall and the type of exhibit. In voice-dominated scenarios (historical explanations, storytelling), the clarity of the voice content is crucial, and content matching is the core criterion for judging interference. The weights corresponding to the loudness ratio and voiceprint similarity are set to 0.3 and 0.7 (the sum of the weights is 1), respectively. In atmosphere-based scenarios (immersive experiences, background music), the loudness of the sound has a more direct impact on the environment, while content details are relatively secondary. The weights corresponding to the loudness ratio and voiceprint similarity are set to 0.6 and 0.4, respectively.

[0074] S2. Based on the tourist status data within the target area, distinguish the browsing mode of the current scene, combine the spatial overlap between the tourist location and the coverage area of ​​the adjacent broadcasting unit, as well as the working status of the adjacent broadcasting unit, to determine the interference risk status.

[0075] In this embodiment, S2 includes the following specific steps:

[0076] S21. Obtain crowd density thresholds and flow speed thresholds. Based on the crowd density and flow speed of tourists within the target area, determine the current browsing mode. The browsing modes include sparse flow mode, dense flow mode, sparse dwell mode, and dense dwell mode. Sparse flow mode is characterized by crowd density less than or equal to the crowd density threshold and flow speed greater than the flow speed threshold. Tourists in sparse flow mode pass through quickly, requiring only brief announcements, and have relatively high tolerance for audio interference. Dense flow mode is characterized by crowd density greater than the crowd density threshold and flow speed greater than the flow speed threshold. Dense flow mode has higher crowd density but faster flow. A balance needs to be struck between broadcast length and audio clarity. The sparse dwell mode is defined as a population density less than or equal to a population density threshold and a flow speed less than or equal to a flow speed threshold. In this mode, the population density is low and people tend to stay longer, primarily requiring basic information delivery in the audio. The dense dwell mode is defined as a population density greater than a population density threshold and a flow speed less than or equal to a flow speed threshold. Visitors in this mode tend to stay longer and have a higher demand for clear, low-interference audio. An interference sensitivity coefficient is assigned to each browsing mode. This coefficient reflects the weight of audio interference on the visitor experience in the current scenario; higher density and lower speed result in higher interference sensitivity. The higher the density, the greater the crowd density feature. Crowd density is obtained by dividing the crowd density by the maximum crowd density within the exhibition hall, and flow speed is obtained by dividing the average flow speed by the flow speed. The interference sensitivity coefficient is obtained by weighted summation of the crowd density and flow speed features. The weights corresponding to the interference sensitivity coefficients are set according to the browsing mode. Visitors in the densely populated mode need to concentrate on listening to explanations for a longer period, so density has a greater impact on interference sensitivity. For example, in crowded conditions, even if the flow is slow, high density will exacerbate interference. Therefore, the weights corresponding to crowd density and flow speed are set to 0.7 and 0.3 respectively (the sum of the weights is 1). Visitors in the sparsely populated mode pass through quickly, so speed has a greater impact on interference sensitivity. Interference sensitivity has a greater impact. For example, at high speeds, even at low density, the delay or blurring of sound transmission can affect the experience. Therefore, the weights for crowd density and flow speed are set to 0.3 and 0.7, respectively. In dense flow patterns, high density leads to sound overlap, and high speed leads to short sound reception time. Both have an impact, but density has a slightly greater impact. Therefore, the weights for crowd density and flow speed are set to 0.6 and 0.4, respectively. In sparse dwell patterns, low speed leads to long dwell time and high requirements for sound clarity, but low density results in moderate overall interference sensitivity. Therefore, the weights for crowd density and flow speed are set to 0.4 and 0.6, respectively.

[0077] S22. Obtain the intersection area between the target area and the coverage area of ​​the adjacent broadcasting unit, filter out the number of tourists in the intersection area, and obtain the spatial overlap by dividing the number of tourists in the intersection area by the total number of tourists in the target area.

[0078] S23. Obtain the interference risk index by multiplying the broadcast activity coefficient, interference sensitivity coefficient, and spatial overlap.

[0079] S3. Obtain the spatial data and material information of the exhibition hall. Based on the relative orientation of the visitor's standing coordinates and the ultrasonic transducer array, calculate the reflected sound and its loudness distribution at different positions and angles reaching the target point, and obtain the reflection interference factor.

[0080] Please see Figure 2 In this embodiment, S3 includes the following specific steps:

[0081] S31. Acquire spatial data of the exhibition hall and construct a three-dimensional geometric model of the exhibition hall. This involves collecting the size, shape, spatial position, and geometric features of each boundary surface (walls, floor, display panels, etc.) of the exhibition hall through laser scanning. Use CAD software and point cloud processing tools to generate a three-dimensional geometric model of the exhibition hall. Obtain the center coordinates and principal axis pointing vector of the ultrasonic transducer array. The principal axis pointing vector is the main sound wave propagation direction of the array, i.e., the core emission direction of the sound source or the core receiving direction of the sensor. Obtain the material information of each boundary surface of the exhibition hall, such as wood or stone for walls, carpet or tile for floors, and metal or plastic for display panels. Assign the corresponding material absorption coefficient based on the acoustic characteristics of the material information. The steps for obtaining the material absorption coefficient are as follows: search for the absorption coefficient of the material at the target frequency from the acoustic material database or measure the absorption coefficient of the material in the laboratory using the standing wave tube method. The absorption coefficient determines the energy loss (absorption) and reflection ratio when the sound wave encounters the surface, which directly affects the distribution of the sound field. Obtain the sound energy reflection coefficient by the difference between the value 1 and the material absorption coefficient.

[0082] S32. Obtain the average ear height and standing coordinates of tourists. The average ear height of tourists is usually taken as 1.5m-1.7m. Construct the three-dimensional coordinates of tourists and obtain the distance from the three-dimensional coordinates of tourists to the center coordinates of the ultrasonic transducer array. Set as the direct sound propagation distance. The distance is obtained by calculating the Euclidean distance of the corresponding point coordinates.

[0083] S33. Obtain the mirror sound source of the ultrasonic transducer array for each boundary surface, connect the mirror sound source with the three-dimensional coordinates of the visitor, obtain the reflection point through the intersection of the line connecting the mirror sound source and the boundary surface, obtain the sum of the distance from the center coordinates of the ultrasonic transducer array to the reflection point and the distance from the reflection point to the three-dimensional coordinates of the visitor, and set it as the reflection path length. The number of reflection paths corresponds to the number of boundary surfaces in the exhibition hall.

[0084] S34. Obtain the direct sound deviation angle by the angle between the vector pointing from the center coordinates of the ultrasonic transducer array to the three-dimensional coordinates of the tourist and the vector pointing from the principal axis; obtain the reflected sound emission angle by the angle between the vector pointing from the center coordinates of the ultrasonic transducer array to the reflection point and the vector pointing from the principal axis.

[0085] S35. Introduce the transducer directivity function. The transducer directivity function is a dimensionless normalized function that describes the radiation capability of the transducer in different directions. The directivity gain of the direct sound is obtained by taking the value of the transducer directivity function at the direct sound deviation angle, and the directivity gain of the reflection path is obtained by taking the value of the transducer directivity function at the reflected sound emission angle.

[0086] S36. Calculate the direct sound pressure level at the tourist's three-dimensional coordinates using the direct sound pressure level calculation formula, wherein the direct sound pressure level calculation formula is:

[0087]

[0088] In the formula, To reach the sound pressure level directly, To obtain the reference sound pressure level, the steps are as follows: Place the sound level meter 1 meter in front of the sound source, ensure that the ambient noise is at least 10 dB lower than the sound pressure level of the source, and measure the sound pressure level of the source at that location. For the directional gain of direct sound, For the direct sound propagation distance, The air attenuation coefficient is the rate attenuation of sound waves as they propagate through air. It is calculated by measuring the temperature, humidity, and frequency of the current environment, using empirical formulas or by consulting air attenuation coefficient tables according to ISO 9613-1 standard. The direct sound pressure level is converted into direct sound energy. The sound pressure level reflected from the boundary surface to the visitor's three-dimensional coordinates is then calculated using the reflected sound pressure level calculation formula:

[0089] In the formula, Let i be the sound pressure level reflected from the i-th boundary surface to the tourist's three-dimensional coordinates. For the directional gain of the i-th reflection path, Let be the reflection path length reflected by the i-th boundary surface. Let be the sound absorption coefficient of the material at the i-th boundary surface. Let be the incident angle of the sound wave when it strikes the i-th boundary surface. The incident angle is the angle between the sound wave and the interface normal, which affects the effectiveness of the interface sound absorption coefficient. This represents the proportion of sound energy reflected after a sound wave is incident on the boundary surface. Since the material's sound absorption coefficient represents the proportion of absorbed sound energy, the sound energy reflection coefficient is obtained by the difference between the value of 1 and the material's sound absorption coefficient. Furthermore, the sound absorption coefficient varies with the incident angle. Only reflected sound energy contributes to the reflected sound pressure level at the visitor's coordinates; absorbed sound energy is consumed by the surface and does not participate in reflection. The sound pressure level reflected from the boundary surface to the visitor's three-dimensional coordinates is converted into reflected sound energy. Sound pressure level is a logarithmic unit describing sound intensity, defined based on the logarithmic perception of sound by the human ear. Since sound pressure level is a logarithmic unit, it cannot be directly linearly added; therefore, each sound pressure level needs to be converted into a linear energy value. The original definition of sound pressure level is 20log... 10 Sound intensity (energy density) is proportional to the square of sound pressure, and energy is defined as 10log⁻¹. 10 The logarithmic relationship between sound pressure level and sound intensity means that for every 10 dB increase in sound pressure level, the sound intensity increases tenfold. The total sound pressure level at each tourist's three-dimensional coordinates is calculated by superimposing all direct and reflected sound energy. The total sound pressure level can be expressed as:

[0090]

[0091] The reflection interference factor is calculated by the difference between the total sound pressure level and the direct sound pressure level. The calculation of the reflection interference factor can be expressed as:

[0092]

[0093] In the formula, As a reflection interference factor, The total sound pressure level is represented by the reflection interference factor. The larger the reflection interference factor value, the more severe the sound interference from reflected sound from the boundary surface at that location, and the lower the sound clarity.

[0094] S4. Combines ambient background noise, browsing mode, interference risk status, and reflection interference factor to output the best sound status parameters.

[0095] In this embodiment, S4 includes the following specific steps:

[0096] S41. Obtain the reference playback sound pressure level by summing the ambient background noise sound pressure level and the target signal-to-noise ratio. The target signal-to-noise ratio is usually set to 10-15dB to ensure speech clarity. Obtain the browsing mode gain correction value by multiplying the interference sensitivity coefficient corresponding to the browsing mode by the dynamic gain correction reference value. The steps for obtaining the dynamic gain correction reference value are as follows: select an environment with no or low interference as the reference state, record the system gain value with a sound level meter, take the average value through multiple tests, obtain the avoidance attenuation value by multiplying the interference risk index by the avoidance attenuation coefficient, and obtain the reflection clarity compensation gain value by multiplying the reflection interference factor by the compensation coefficient. The compensation coefficient is usually set to 0.2-0.5 to avoid the total sound pressure level being too high due to overcompensation.

[0097] S42. Obtain the optimal sound state parameters, i.e., the optimal playback sound pressure level, by subtracting the avoidance attenuation value from the sum of the reference playback sound pressure level, the browsing mode gain correction value, and the reflection clarity compensation gain value.

[0098] It should be noted that the steps for obtaining the avoidance attenuation coefficient, voiceprint similarity threshold, crowd density threshold, and flow velocity threshold are as follows: A historical dataset containing multiple known interference risk indices and their corresponding best measured avoidance attenuation values ​​is collected as labels. The data is imported into MATLAB using the `readtable` function for preprocessing, including outlier detection and handling, missing value imputation, and feature standardization. A fitting model (such as support vector regression or multinomial regression) is selected, and the model hyperparameters are adjusted through cross-validation to minimize the root mean square error (RMSE) between the predicted avoidance attenuation value and the historical measured value, while ensuring that the avoidance attenuation coefficient meets physical constraints (such as non-negativity). This yields the avoidance attenuation coefficient, which is then used for voiceprint similarity, crowd density threshold, and flow velocity threshold. The goal of this method is to distinguish between high-risk and low-risk states. Historical data is sorted by the interference risk index, and the critical point that best distinguishes high-risk and low-risk samples is selected as the initial threshold. The receiver operating characteristic curve is calculated using the ROC function. The initial threshold is determined by balancing the false alarm rate and the detection rate (finding the point with the largest Youden index). Using a reserved test dataset, the confusion matrix is ​​calculated using the confusionmat function. The accuracy, recall, and other indicators are analyzed to verify the classification accuracy. If the verification results do not meet the requirements, the initial threshold is adjusted and re-verified. Finally, the voiceprint similarity threshold, the crowd density threshold, and the flow speed threshold are determined. At the same time, a dynamic update mechanism can be established to retrain the thresholds with new data periodically to adapt to scene changes.

[0099] S5. Calculate the beam pointing angle based on the tourist's real-time coordinates, and adjust the main lobe width of the ultrasonic array under the constraint of optimal sound state parameters to perform directional projection.

[0100] Please see Figure 3In this embodiment, S5 includes the following specific steps:

[0101] S51. Obtain the standing coordinates of tourists within the target area and calculate the center of the tourist's standing coordinates, set it as the projection target point, combine it with the average ear height of tourists to construct the coordinates of the projection target point, and obtain the direction vector from the center of the array to the projection target point by subtracting the center coordinates of the ultrasonic transducer array from the coordinates of the projection target point.

[0102] S52. Calculate the horizontal and vertical deflection angles of the beam relative to the array normal using the deflection angle calculation formula. The deflection angle calculation formula is as follows:

[0103]

[0104] In the formula, The horizontal deflection angle is the angle by which the sound beam emitted by the ultrasonic transducer array deflects in the horizontal plane relative to the array normal. The vertical deflection angle is the angle at which the sound beam emitted by the ultrasonic transducer array deflects relative to the array normal in the vertical plane. This refers to the sound wave propagation distance, which is the straight-line distance from the center of the ultrasonic transducer array to the target projection point. The coordinates of the projection target point, Let the center coordinates of the ultrasonic transducer array be given. The phase offset of the array elements is calculated using the phase offset calculation formula, which is:

[0105]

[0106] In the formula, Let be the phase offset of the j-th array element. The center frequency of the ultrasound wave is acquired using a spectrum analyzer. This refers to the speed of sound propagation, for example, the speed of sound in air is approximately 340 m / s. The coordinates of the j-th array element relative to the array center are used to locate the position of the array element in the array. The phase array is output based on the phase offset of all array elements. The phase array is directly input to the phase shifter (the core control component of the array). By controlling the phase delay of each array element, the spatial pointing of the beam is adjusted, that is, the sound wave is accurately aligned with the target point.

[0107] S53. Obtain the distribution radius of tourists in the target area relative to the projection target point. The distribution radius can be calculated as follows:

[0108]

[0109] In the formula, The total number of tourists in the target area. Given the coordinates of the nth tourist's position, the minimum beamwidth required to cover the tourist is calculated using the minimum beamwidth calculation formula:

[0110]

[0111] In the formula, The minimum beamwidth is the distribution radius of tourists in the target area relative to the projection target point. It should be noted that the minimum beamwidth is corrected according to the reflection interference factor. That is, the minimum value between the beamwidth limited by reflection interference and the calculated minimum beamwidth is selected. If the reflection interference factor is high, the beam needs to be narrowed manually to reduce wall penetration. The beamwidth limited by reflection interference can be obtained by dividing the environmental constant by the reflection interference factor. The environmental constant needs to be calibrated according to the actual scene, or it can be obtained by creating a corresponding table. For example, when the reflection interference factor is high, the beamwidth limited by reflection interference is 10°; when the reflection interference factor is medium, the beamwidth limited by reflection interference is 15°; and when the reflection interference factor is low, the beamwidth limited by reflection interference is 20°.

[0112] S54. Obtain array structure parameters, including the number of array elements and the element spacing. The element spacing affects beam directivity. A window function method is used, where a Chebyshev window can be selected. By adjusting the amplitude of each array element, the main lobe width of the beam conforms to the minimum beam width, while simultaneously reducing the sidelobe level to a preset value. Based on the minimum beam width and the preset sidelobe level (the preset sidelobe level represents the energy propagated by the beam in non-target directions, and its magnitude needs to be controlled to reduce interference; a value of -25dB can be used), the amplitude weighting coefficient of each array element is derived. The window function method can be expressed as:

[0113]

[0114] In the formula, Number The amplitude coefficient of the array element, For the number of array elements, and Determined by the minimum beamwidth and sidelobe level, and obtained by looking up the window function parameter table. As the center of the array, the calculated amplitude weights are normalized, and an amplitude array is output based on the amplitude weighting coefficients of all array elements. The amplitude array controls the amplification factor of each channel signal and determines the beam shape so that the beam width covers the distribution of tourists.

[0115] S55. Calculate the total gain value using the total gain calculation formula. The total gain is the total amplification factor that the system (including the sound source, array, amplifier, etc.) needs to provide. Its core purpose is to compensate for the sound pressure level attenuation caused by spherical wave diffusion and air absorption during the propagation of sound waves from the sound source to the target point, ultimately enabling the sound pressure level at the target point to reach the optimal playback sound pressure level. The total gain value calculation formula is as follows:

[0116]

[0117] In the formula, This is the total gain value. For optimal playback sound pressure level, This is the baseline correction term for air absorption. When sound waves propagate in the air, high-frequency components are additionally attenuated due to absorption by air molecules. The greater the distance, the more significant the high-frequency attenuation. Under standard atmospheric pressure and 20°C conditions, the baseline correction term for air absorption is approximately 11 dB. This refers to the attenuation of sound pressure level due to the diffusion of spherical waves as sound waves propagate in free space; the attenuation is more pronounced with increasing distance.

[0118] S56. Obtain the phase array, amplitude array, optimal playback sound pressure level, and total gain value, and input them into the controller for directional projection.

[0119] Please see Figure 4 The present invention also provides a directional sound fusion analysis system based on voiceprint splicing, comprising:

[0120] The data monitoring module is used to collect tourist status data within the target area, monitor environmental background noise data, and simultaneously obtain the working status of adjacent broadcasting units in real time.

[0121] The interference risk assessment module is used to distinguish the browsing mode of the current scene based on the tourist status data in the target area, and to determine the interference risk status by combining the spatial overlap between the tourist location and the coverage area of ​​the adjacent broadcasting unit, as well as the working status of the adjacent broadcasting unit.

[0122] The reflection interference factor analysis module is used to obtain spatial data and material information of the exhibition hall. Based on the relative orientation of the visitor's standing coordinates and the ultrasonic transducer array, it calculates the reflected sound and its loudness distribution at different positions and angles reaching the target point, and obtains the reflection interference factor.

[0123] The optimal sound level analysis module is used to combine ambient background noise, browsing mode, interference risk status, and reflection interference factor to output the optimal sound state parameters.

[0124] The directional projection module is used to calculate the beam pointing angle based on the real-time coordinates of the tourists, and adjust the main lobe width of the ultrasonic array under the constraints of optimal sound state parameters to perform directional projection.

[0125] This invention also provides a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to perform the directional sound fusion analysis method based on voiceprint splicing as described above.

[0126] This invention also provides an electronic device, specifically including a memory and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described above for directional sound fusion analysis based on voiceprint splicing.

[0127] The various embodiments in this specification are described in a progressive manner. Similar parts between embodiments can be referred to mutually. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the description of the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0128] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art. The general principles defined in this invention may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A directional sound fusion analysis method based on voiceprint splicing, characterized in that, The specific steps include the following: S1. Collect tourist status data within the target area, monitor environmental background noise data, and simultaneously obtain the working status of adjacent broadcasting units in real time. S2. Based on the tourist status data within the target area, distinguish the browsing mode of the current scene, combine the spatial overlap between the tourist location and the coverage area of ​​the adjacent broadcasting unit, as well as the working status of the adjacent broadcasting unit, to determine the interference risk status. S3. Obtain the spatial data and material information of the exhibition hall. Based on the relative orientation of the visitor's standing coordinates and the ultrasonic transducer array, calculate the reflected sound and its loudness distribution at different positions and angles reaching the target point, and obtain the reflection interference factor. S4. Combines ambient background noise, browsing mode, interference risk status, and reflection interference factor to output the best sound status parameters. S5. Calculate the beam pointing angle based on the tourist's real-time coordinates, and adjust the main lobe width of the ultrasonic array under the constraint of optimal sound state parameters to perform directional projection.

2. The directional sound fusion analysis method based on voiceprint splicing according to claim 1, characterized in that, S1 includes the following specific steps: S11. Collect tourist status data within the target area, including station coordinates, crowd density and flow speed; monitor environmental background noise data, including environmental background noise sound pressure level. S12. Extract the target voiceprint from the environmental background noise data, obtain the playback energy of the target voiceprint, construct the target voiceprint feature sequence, obtain the standard voiceprint feature sequence of the preset playback content of the adjacent broadcasting unit, perform feature splicing and alignment of the target voiceprint feature sequence and the standard voiceprint feature sequence, calculate the voiceprint similarity between the target voiceprint and the standard voiceprint using cosine similarity, obtain the voiceprint similarity threshold and the playback energy threshold, determine the broadcasting status of the adjacent broadcasting unit, when the voiceprint similarity is greater than or equal to the voiceprint similarity threshold, determine the broadcasting status as broadcasting, when the playback energy is less than the playback energy threshold, determine the broadcasting status as standby, when the playback energy is greater than or equal to the playback energy threshold and the voiceprint similarity is less than the voiceprint similarity threshold, determine the broadcasting status as fault, and allocate the broadcasting activity coefficient to the broadcasting status.

3. The directional sound fusion analysis method based on voiceprint splicing according to claim 2, characterized in that, S2 includes the following specific steps: S21. Obtain the crowd density threshold and flow speed threshold. Based on the crowd density and flow speed of tourists in the target area, determine the browsing mode of the current scene. The browsing mode includes sparse flow mode, dense flow mode, sparse dwell mode and dense dwell mode. Sparse flow mode is when the crowd density is less than or equal to the crowd density threshold and the flow speed is greater than the flow speed threshold. Dense flow mode is when the crowd density is greater than the crowd density threshold and the flow speed is greater than the flow speed threshold. Sparse dwell mode is when the crowd density is less than or equal to the crowd density threshold and the flow speed is less than or equal to the flow speed threshold. Dense dwell mode is when the crowd density is greater than the crowd density threshold and the flow speed is less than or equal to the flow speed threshold. Assign an interference sensitivity coefficient to each browsing mode. S22. Obtain the intersection area between the target area and the coverage area of ​​the adjacent broadcasting unit, filter out the number of tourists in the intersection area, and obtain the spatial overlap by dividing the number of tourists in the intersection area by the total number of tourists in the target area. S23. Obtain the interference risk index by multiplying the broadcast activity coefficient, interference sensitivity coefficient, and spatial overlap.

4. The directional sound fusion analysis method based on voiceprint splicing according to claim 3, characterized in that, S3 includes the following specific steps: S31. Obtain the spatial data of the exhibition hall, construct the three-dimensional geometric model of the exhibition hall, obtain the center coordinates and principal axis pointing vector of the ultrasonic transducer array, obtain the material information of each boundary surface of the exhibition hall, and assign the corresponding material sound absorption coefficient according to the acoustic characteristics of the material information. S32. Obtain the average ear height and standing coordinates of tourists, construct the three-dimensional coordinates of tourists, and obtain the distance from the three-dimensional coordinates of tourists to the center coordinates of the ultrasonic transducer array, which is set as the direct sound propagation distance. S33. Obtain the mirror sound source of the ultrasonic transducer array for each boundary surface, connect the mirror sound source with the three-dimensional coordinates of the tourist, obtain the reflection point through the intersection of the line connecting the mirror sound source and the boundary surface, and obtain the sum of the distance from the center coordinates of the ultrasonic transducer array to the reflection point and the distance from the reflection point to the three-dimensional coordinates of the tourist, and set it as the reflection path length. S34. Obtain the direct sound deviation angle by the angle between the vector pointing from the center coordinates of the ultrasonic transducer array to the three-dimensional coordinates of the tourist and the vector pointing from the principal axis; obtain the reflected sound emission angle by the angle between the vector pointing from the center coordinates of the ultrasonic transducer array to the reflection point and the vector pointing from the principal axis. S35. Introduce the transducer directivity function, obtain the directivity gain of the direct sound by taking the value of the transducer directivity function at the direct sound deviation angle, and obtain the directivity gain of the reflection path by taking the value of the transducer directivity function at the reflected sound emission angle. S36. Calculate the direct sound pressure level at the tourist's three-dimensional coordinates using the direct sound pressure level calculation formula, wherein the direct sound pressure level calculation formula is: In the formula, To reach the sound pressure level directly, As the reference sound pressure level, For the directional gain of direct sound, For the direct sound propagation distance, The air attenuation coefficient is used to convert the direct sound pressure level into direct sound energy. The sound pressure level reflected from the boundary surface to the tourist's three-dimensional coordinates is calculated using the reflected sound pressure level calculation formula, which is: In the formula, Let i be the sound pressure level reflected from the i-th boundary surface to the tourist's three-dimensional coordinates. For the directional gain of the i-th reflection path, Let be the reflection path length reflected by the i-th boundary surface. Let be the sound absorption coefficient of the material at the i-th boundary surface. Let be the incident angle of the sound wave incident on the i-th boundary surface. The sound pressure level reflected from the boundary surface to the tourist's three-dimensional coordinates is converted into reflected sound energy. The total sound pressure level at each tourist's three-dimensional coordinates is calculated by superimposing all direct sound energy and reflected sound energy. The reflection interference factor is calculated by the difference between the total sound pressure level and the direct sound pressure level.

5. The directional sound fusion analysis method based on voiceprint splicing according to claim 4, characterized in that, S4 includes the following specific steps: S41. Obtain the reference playback sound pressure level by summing the ambient background noise sound pressure level and the target signal-to-noise ratio; obtain the browsing mode gain correction value by multiplying the interference sensitivity coefficient corresponding to the browsing mode with the dynamic gain correction reference value; obtain the avoidance attenuation value by multiplying the interference risk index with the avoidance attenuation coefficient; and obtain the reflection clarity compensation gain value by multiplying the reflection interference factor with the compensation coefficient. S42. Obtain the optimal sound state parameters, i.e., the optimal playback sound pressure level, by subtracting the avoidance attenuation value from the sum of the reference playback sound pressure level, the browsing mode gain correction value, and the reflection clarity compensation gain value.

6. The directional sound fusion analysis method based on voiceprint splicing according to claim 5, characterized in that, S5 includes the following specific steps: S51. Obtain the standing coordinates of tourists within the target area and calculate the center of the tourist's standing coordinates, set it as the projection target point, combine it with the average ear height of tourists to construct the coordinates of the projection target point, and obtain the direction vector from the center of the array to the projection target point by subtracting the center coordinates of the ultrasonic transducer array from the coordinates of the projection target point. S52. Calculate the horizontal and vertical deflection angles of the beam relative to the array normal using the deflection angle calculation formula. The deflection angle calculation formula is as follows: In the formula, It is the horizontal deflection angle. It is the vertical deflection angle. For the distance the sound wave travels, The coordinates of the projection target point, Let the center coordinates of the ultrasonic transducer array be given. The phase offset of the array elements is calculated using the phase offset calculation formula, which is: In the formula, Let be the phase offset of the j-th array element. The center frequency of the ultrasound. For the speed of sound wave propagation, Given the coordinates of the j-th array element relative to the array center, output the phase array based on the phase offset of all array elements; S53. Obtain the distribution radius of tourists in the target area relative to the projection target point, and calculate the minimum beamwidth required to cover the tourists using the minimum beamwidth calculation formula. The minimum beamwidth calculation formula is as follows: In the formula, The radius of the distribution of tourists in the target area relative to the projection target point; S54. Obtain array structure parameters, including the number of array elements and the spacing between array elements. Using the window function method, based on the minimum beamwidth and the preset sidelobe level, the amplitude weighting coefficient of each array element is inversely calculated, and the amplitude array is output based on the amplitude weighting coefficients of all array elements. S55. Calculate the total gain value using the total gain value calculation formula, which is: In the formula, This is the total gain value. For optimal playback sound pressure level, This is a baseline correction term for air absorption; S56. Obtain the phase array, amplitude array, optimal playback sound pressure level, and total gain value, and input them into the controller for directional projection.

7. A directional sound fusion analysis system based on voiceprint splicing, used to implement the directional sound fusion analysis method based on voiceprint splicing as described in any one of claims 1-6, characterized in that, include: The data monitoring module is used to collect tourist status data within the target area, monitor environmental background noise data, and simultaneously obtain the working status of adjacent broadcasting units in real time. The interference risk assessment module is used to distinguish the browsing mode of the current scene based on the tourist status data in the target area, and to determine the interference risk status by combining the spatial overlap between the tourist location and the coverage area of ​​the adjacent broadcasting unit, as well as the working status of the adjacent broadcasting unit. The reflection interference factor analysis module is used to obtain spatial data and material information of the exhibition hall. Based on the relative orientation of the visitor's standing coordinates and the ultrasonic transducer array, it calculates the reflected sound and its loudness distribution at different positions and angles reaching the target point, and obtains the reflection interference factor. The optimal sound level analysis module is used to combine ambient background noise, browsing mode, interference risk status, and reflection interference factor to output the optimal sound state parameters. The directional projection module is used to calculate the beam pointing angle based on the real-time coordinates of the tourists, and adjust the main lobe width of the ultrasonic array under the constraints of optimal sound state parameters to perform directional projection.

8. A storage medium, characterized in that, The storage medium includes stored instructions, wherein, when the instructions are executed, the device containing the storage medium is controlled to perform the directional sound fusion analysis method based on voiceprint splicing as described in any one of claims 1-6.

9. An electronic device, characterized in that, It includes a memory, and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described in any one of claims 1-6.