Method and device for pushing directional audio for appreciating large-sized painting

By combining an infrared motion capture system and an improved escape optimization algorithm with the Levy flight mechanism, the problems of multi-target acoustic crosstalk and computational complexity were solved, enabling precise directional audio delivery in high-density crowds, thus enhancing the immersive experience and computational efficiency of art appreciation.

CN122349035APending Publication Date: 2026-07-07ZHONGCHUAN YUEZHONG (BEIJING) CULTURE DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHONGCHUAN YUEZHONG (BEIJING) CULTURE DEVELOPMENT CO LTD
Filing Date
2026-04-14
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Existing targeted audio delivery technologies suffer from severe acoustic crosstalk between multiple targets, the algorithm is prone to getting trapped in local optima under high density, and there is a contradiction between computational complexity and real-time performance, making it impossible to achieve accurate targeted audio delivery in high-density crowd environments.

Method used

An infrared motion capture system is used to collect six-degree-of-freedom head motion data of the audience, construct a global state matrix and an array coordinate matrix, and combine an improved escape optimization algorithm and Levy flight mechanism to generate a complex beamforming weighted coefficient matrix. Directional audio push is achieved through dynamic sound field construction.

Benefits of technology

It achieves a zero-interference experience of 'one voice for one thousand people' in extremely close-range environments with high-density crowds, enhancing the immersiveness of art appreciation. Its computational efficiency far exceeds that of traditional matrix inversion algorithms, enabling real-time tracking of the audience's gaze and seamless audio switching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122349035A_ABST
    Figure CN122349035A_ABST
Patent Text Reader

Abstract

This invention discloses a method and device for directional audio delivery for appreciating large-scale paintings, relating to the field of human-computer interaction technology. The method includes: collecting six-degree-of-freedom motion data of the viewer's head and constructing a global state matrix and an array coordinate matrix; based on the global state matrix, mapping and matching the gaze point with audio, calculating the coordinates of each viewer's gaze point on the large-scale painting, and matching the corresponding local audio narration data; constructing an acoustic evacuation topology map for the local audio narration data; generating a complex beamforming weighting coefficient matrix based on the acoustic evacuation topology map using an improved escape optimization algorithm; generating an audio signal vector, combining it with the complex beamforming weighting coefficient matrix to construct a dynamic sound field, and using a directional audio array to execute directional audio stream delivery. This invention solves the problems of severe multi-target acoustic crosstalk, the algorithm's susceptibility to local optima under high density, and the contradiction between computational complexity and real-time performance in existing technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, and in particular to a method and apparatus for directional audio push for appreciating large-scale paintings. Background Technology

[0002] With the rapid development of ultra-high-definition display technology, ultra-high-definition large-format display terminals are increasingly widely used in cultural exhibition venues such as museums and art galleries. In order to provide an immersive experience, existing technologies are beginning to explore the use of directional sound technology to provide "sound-following-the-person" narration services.

[0003] However, existing targeted audio delivery technologies have the following significant drawbacks: 1) Severe acoustic crosstalk in multiple targets: When there are multiple audience members in the viewing area at the same time and they are close to each other, the traditional beamforming algorithm cannot form a sufficiently deep "zero trap" at the location of the audience members when dealing with large-scale multi-target constraints, resulting in "crosstalk" phenomenon. The audience is very susceptible to interference from the audio of other people's explanations.

[0004] 2) Algorithms are prone to getting trapped in local optima under high density: In complex environments with dense crowds, traditional gradient descent or conventional heuristic algorithms are very prone to getting trapped in local extrema when solving for complex weighted coefficients, which can lead to distortion or incorrect pointing of the generated sound beam.

[0005] 3) The contradiction between computational complexity and real-time performance: It is impossible to track the specific landing point of multiple viewers' gazes on the huge screen in real time and accurately. In addition, after increasing the array size in order to pursue sound field accuracy, the matrix inversion calculation of traditional algorithms increases exponentially, which cannot meet the millisecond-level real-time response requirements of "sound follows people".

[0006] Therefore, there is an urgent need for a new type of directional audio push method that can adapt to high-density populations, has strong anti-interference capabilities, and is computationally efficient. Summary of the Invention

[0007] This invention provides a method and apparatus for directional audio delivery for appreciating large-scale paintings. This invention solves the problems of severe multi-target acoustic crosstalk, easy entrapment of algorithms in high-density environments, and the contradiction between computational complexity and real-time performance in existing technologies.

[0008] In a first aspect, embodiments of the present invention provide a method for targeted audio delivery for appreciating large-scale paintings, the method comprising: Using an infrared motion capture system deployed around a high-definition large-format display terminal, six degrees of freedom motion data of each viewer's head are collected, and a global state matrix and an array coordinate matrix of the directional audio array are constructed for all viewers. Based on the global state matrix, gaze point mapping and audio matching are performed. The coordinates of each viewer's gaze point on the giant painting on the high-definition giant display terminal are calculated, and the corresponding local audio narration data is matched. For each local audio narration data, an acoustic evacuation topology map including the particles to be evacuated, the target exit, and dynamic obstacles is constructed based on the global state matrix and the array coordinate matrix. Based on the acoustic evacuation topology, an improved escape optimization algorithm is used to simulate the emergency evacuation behavior of particles to be evacuated. By iteratively calculating the force state of the particles to be evacuated, a complex beamforming weighting coefficient matrix of the directional audio array is generated. The audio signal vector of each local audio narration data is generated, and a dynamic sound field is constructed by combining it with a complex beamforming weighting coefficient matrix. A directional audio array is then used to perform directional audio stream delivery.

[0009] The technical solution provided in this application has at least the following beneficial effects: The beamforming problem is cleverly transformed into a "multi-objective independent evacuation behavior simulation" to improve the escape optimization algorithm. By introducing a dynamic obstacle repulsion mechanism, the algorithm is forced to automatically generate an extremely deep acoustic null in the location of the audience, achieving a zero-interference experience of "a thousand people, a thousand voices" at extremely close range. The improved escape optimization algorithm innovatively introduces the Levy flight mechanism. When the sound beam optimization is blocked by obstacles and falls into a dead zone, the step size mutation capability can force the algorithm to jump out of the local extremum. At the same time, with the nonlinear convergence factor, it can achieve rapid coarse adjustment in the early stage of iteration and micron-level precise fine adjustment in the later stage. The computational efficiency is far superior to the traditional matrix inversion algorithm. By acquiring the audience's six-degree-of-freedom head data through an infrared motion capture system, and combining ray intersection and affine transformation, the audience's gaze point on the giant painting can be accurately calculated. With the cross-fade-in and fade-out mechanism of the grid boundary buffer, even the slightest movement of the audience's gaze on the giant painting is converted into a seamless switching of the corresponding local narration audio in real time. Moreover, the audio accurately locks onto and follows the audience in the form of an invisible sound beam, greatly enhancing the immersive experience of art appreciation.

[0010] In one alternative implementation, an infrared motion capture system deployed around a high-definition large-format display terminal is used to collect six-degree-of-freedom motion data of each viewer's head, and to construct a global state matrix and an array coordinate matrix of the directional audio array for all viewers, including: Using an infrared motion capture system deployed around a high-definition large-format display terminal, the system tracks reflective markers on each viewer's headband and uses a rigid body calculation algorithm to collect six degrees of freedom motion data of each viewer's head in the physical world coordinate system. The six degrees of freedom motion data includes the three-dimensional coordinates of the head's center of mass and the Euler angles of the head posture. By combining the timestamps, the six-degree-of-freedom motion data of all viewers are assembled to obtain the global state matrix; Obtain the fixed spatial coordinates of each sound-emitting unit in the directional audio array deployed around the high-definition large-format display terminal; By combining the timestamps, the fixed spatial coordinates of all the sound-generating units are assembled to obtain the array coordinate matrix of the directional audio array.

[0011] In one optional implementation, based on the global state matrix, gaze point mapping and audio matching are performed. The coordinates of each viewer's gaze point on the large-scale painting on the high-definition giant display terminal are calculated, and the corresponding local audio narration data is matched, including: For each viewer, extract the corresponding six-degree-of-freedom motion data from the global state matrix and generate the corresponding gaze direction unit vector. Starting from the three-dimensional coordinates of the head's centroid, the gaze ray of the viewer is generated along the corresponding gaze direction unit vector, and the three-dimensional coordinates of the intersection point of the gaze ray and the physical plane equation of the screen of the high-definition giant display terminal are obtained. The three-dimensional coordinates of the intersection point are mapped to the two-dimensional pixel coordinate system of the giant painting through a pre-calibrated affine transformation matrix to obtain the coordinates of the point where the line of sight falls on the giant painting. By looking up a pre-defined grid matrix of artwork content, the audio data segment corresponding to the grid to which the gaze point's coordinates belong can be obtained. If the duration of the gaze point's dwell time on its corresponding grid does not reach the preset trigger threshold, then the local audio narration data from the previous moment will continue to be output. If the duration reaches the preset trigger threshold and the audio data segment is in the boundary buffer between two grids, the mixing factor is calculated inversely proportional to the distance, and the current audio data segment and the target audio data segment are cross-faded in and out to obtain the corresponding local audio narration data. Otherwise, the audio data segment is directly used as the corresponding local audio narration data.

[0012] In one alternative implementation, for each local audio narration data, an acoustic evacuation topology map is constructed based on the global state matrix and the array coordinate matrix, including the particles to be evacuated, the target exit, and dynamic obstacles, comprising: Extract all sound-emitting units from the array coordinate matrix and define them as several particles to be evacuated. Each particle to be evacuated has fixed spatial coordinates and initializes the corresponding two-dimensional virtual initial velocity vector. For the current audience, extract the three-dimensional coordinates of the audience's head centroid from the global state matrix, and define the three-dimensional coordinates of the head centroid as the target exit assigned to the audience; For the current audience, the three-dimensional coordinates of the head centroid of all other audiences in the global state matrix, excluding the current audience, are defined as the dynamic obstacles of the current audience. For the current audience, the three-dimensional coordinates of the particles to be evacuated, the target exit, and the dynamic obstacles in the three-dimensional space are orthogonally projected onto the two-dimensional plane where the directional audio array is located, generating the corresponding two-dimensional topological coordinates of the fixed space, the two-dimensional topological coordinates of the target exit, and the two-dimensional topological coordinates of the dynamic obstacles. By integrating the fixed two-dimensional topological coordinates of each particle to be evacuated, the initial velocity vector, the corresponding two-dimensional topological coordinates of the target exit, and the two-dimensional topological coordinates of all dynamic obstacles, the acoustic evacuation topology map of the current audience is obtained.

[0013] In one alternative implementation, based on an acoustic evacuation topology, an improved escape optimization algorithm is used to simulate the emergency evacuation behavior of particles to be evacuated. By iteratively calculating the force state of the particles, a complex beamforming weighting coefficient matrix for the directional audio array is generated, including: Based on the acoustic evacuation topology, the fixed two-dimensional topological coordinates of all particles to be evacuated are taken as the initial positions, and the corresponding initial velocity vectors are taken as the initial velocities of the particles to be evacuated. Initiate an improved escape optimization algorithm to calculate the combined forces experienced by the current evacuated particle as it moves to the corresponding target exit during the current iteration. By introducing a convergence factor and the Levy flight mechanism, and combining the combined forces, the initial velocity and initial position of the particle to be evacuated are updated to obtain the updated position and updated velocity. Based on the updated position, the fitness value of the current particle to be evacuated is calculated using a preset fitness function; Repeatedly update the position of the current particle to be evacuated. When the current iteration count reaches the maximum iteration count or the rate of change of fitness value is less than the rate of change threshold, terminate the position update and output the velocity of the particle to be evacuated with the best fitness value as the optimal velocity. The optimal velocity direction is mapped to the beam steering angle, and acoustic phase compensation is performed by combining the actual three-dimensional physical distance between the sound unit and the audience to generate complex beamforming weighting coefficients. By integrating the complex beamforming weighting coefficients of all particles to be evacuated, a complex beamforming weighting coefficient matrix for the directional audio array is obtained.

[0014] In one alternative implementation, an improved escape optimization algorithm is initiated. During the current iteration, the combined forces experienced by the current particle to be evacuated to the corresponding target exit are calculated, including: Initiate an improved escape optimization algorithm to calculate the driving force received by the current evacuation particle toward the corresponding target exit during the current iteration. Calculate the repulsive force exerted by dynamic obstacles on the current particle to be evacuated, and integrate the total repulsive force of all dynamic obstacles on the corresponding particle to be evacuated; Calculate the combined force on the current particle to be evacuated based on the driving force and the total repulsive force.

[0015] In one alternative implementation, the update formulas for the velocity and position are: The update formulas for the velocity and position are as follows: In the formula, For the first t+ 1, t The iteration of the ... k The velocity of the particles to be evacuated; This is the acceleration coefficient; Let be a random perturbation vector in the interval [0,1]. For the first t The convergence factor of the next iteration; Scaling factor Corresponding Levy flight stride; For the first t +1 iterations of the first iteration k Particles to be dispersed to the first n The combined forces affecting the target exit of the audience; The symbol for element-wise product; In the formula, It is a random vector that follows a normal distribution; is the stability index of the Levy distribution; For the first t The optimal speed for the next iteration; In the formula, For the first t The convergence factor of the next iteration; These are the maximum and minimum values ​​of the convergence factor; This represents the maximum number of iterations. In the formula, For the first t+ 1, t The iteration of the ... k The updated position of the particles to be evacuated.

[0016] In one alternative implementation, the fitness function is formulated as follows: In the formula, For the first t+ The 1st iteration k Fitness value of the particles to be evacuated; For the first n The target exit for the audience is the first k The target attraction value of the particles to be evacuated; For the first m Dynamic obstacles for the first k The barrier repulsion value of the particles to be evacuated; The radius of the acoustically quiet zone (danger radius, such as 0.3 meters); For fitness weighting coefficients; For the first n The target exit of the audience is located in two-dimensional topological coordinates. For the first m Two-dimensional topological coordinates of dynamic obstacles; It is an exponential function; This is a danger penalty value; This is the penalty coefficient.

[0017] In one alternative implementation, an audio signal vector for each local audio narration data is generated, combined with a complex beamforming weighting coefficient matrix to construct a dynamic sound field, and a directional audio stream is delivered using a directional audio array, including: Integrate local audio narration data from all viewers to generate corresponding audio signal vectors; The audio signal vector is multiplied by the complex beamforming weighting coefficient matrix to generate the final time-domain drive signal vector of all sound units in the array coordinate matrix. Each driving signal in the final time-domain driving signal vector is converted by a DAC and amplified by power, and then input to the ultrasonic parametric array of piezoelectric ceramics in the corresponding sound unit of the directional audio array. By executing the drive signal, the ultrasonic parametric array of sound-generating units radiates an ultrasonic beam carrying audio information outward, thereby enabling directional audio delivery for appreciating large-scale paintings.

[0018] Secondly, embodiments of the present invention provide a directional audio push device for appreciating large-scale paintings, used to implement a directional audio push method for appreciating large-scale paintings, the device comprising: The infrared motion capture unit is used to collect six degrees of freedom motion data of each viewer's head using an infrared motion capture system deployed around a high-definition large-format display terminal, and to construct a global state matrix and an array coordinate matrix of the directional audio array for all viewers. The audio matching unit is used to perform gaze point mapping and audio matching based on the global state matrix, calculate the coordinates of each viewer's gaze point on the giant painting on the high-definition giant display terminal, and match the corresponding local audio narration data. The acoustic evacuation topology construction unit is used to construct an acoustic evacuation topology map, including the particles to be evacuated, the target exit, and dynamic obstacles, based on the global state matrix and the array coordinate matrix for each local audio interpretation data. The beamforming weighting solution unit is used to simulate the emergency evacuation behavior of particles to be evacuated based on the acoustic evacuation topology map and using an improved escape optimization algorithm. It generates the complex beamforming weighting coefficient matrix of the directional audio array by iteratively calculating the force state of the particles to be evacuated. The directional audio push unit is used to generate the audio signal vector of each local audio narration data, combine it with the complex beamforming weighting coefficient matrix to construct a dynamic sound field, and use the directional audio array to perform directional audio stream push.

[0019] A third aspect of this invention provides an electronic device, which includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, such that the at least one processor can perform the method proposed in the first aspect of the present invention.

[0020] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in the first aspect of the present invention. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the electronic device structure of the hardware operating environment involved in the embodiments of the present invention; Figure 2 This is a flowchart illustrating the steps of a method for directional audio push for appreciating large-scale paintings, provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the functional units of a directional audio push device for appreciating large-scale paintings, provided in an embodiment of the present invention. Detailed Implementation

[0022] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0023] The present invention will be further described below with reference to the accompanying drawings.

[0024] Reference Figure 1 , Figure 1 This is a schematic diagram of the electronic device structure of the hardware operating environment involved in the embodiments of the present invention.

[0025] like Figure 1 As shown, the electronic device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.

[0026] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0027] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and an electronic program for a directional audio push device for appreciating large-scale paintings.

[0028] exist Figure 1 In the electronic device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the electronic device of the present invention can be set in the electronic device. The electronic device calls the electronic program of the directional audio push device for appreciating large-scale paintings stored in the memory 1005 through the processor 1001, and executes the directional audio push method for appreciating large-scale paintings provided in the embodiment of the present invention.

[0029] Reference Figure 2 The present invention provides a method for targeted audio push for appreciating large-scale paintings, the method comprising: S201: Using an infrared motion capture system deployed around a high-definition large-format display terminal, six degrees of freedom motion data of each viewer's head are collected, and a global state matrix and an array coordinate matrix of the directional audio array are constructed for all viewers. S202: Based on the global state matrix, perform gaze point mapping and audio matching, calculate the coordinates of each viewer's gaze point on the giant painting on the high-definition giant display terminal, and match the corresponding local audio narration data. S203: For each local audio narration data, construct an acoustic evacuation topology map including the particles to be evacuated, the target exit, and dynamic obstacles based on the global state matrix and the array coordinate matrix; S204: Based on the acoustic evacuation topology, an improved escape optimization algorithm is used to simulate the emergency evacuation behavior of particles to be evacuated. By iteratively calculating the force state of the particles to be evacuated, a complex beamforming weighting coefficient matrix of the directional audio array is generated. S205: Generate the audio signal vector for each local audio narration data, combine it with the complex beamforming weighting coefficient matrix to construct a dynamic sound field, and use a directional audio array to perform directional audio stream delivery.

[0030] The technical solution provided in this application has at least the following beneficial effects: The beamforming problem is cleverly transformed into a "multi-objective independent evacuation behavior simulation" to improve the escape optimization algorithm. By introducing a dynamic obstacle repulsion mechanism, the algorithm is forced to automatically generate an extremely deep acoustic null in the location of the audience, achieving a zero-interference experience of "a thousand people, a thousand voices" at extremely close range. The improved escape optimization algorithm innovatively introduces the Levy flight mechanism. When the sound beam optimization is blocked by obstacles and falls into a dead zone, the step size mutation capability can force the algorithm to jump out of the local extremum. At the same time, with the nonlinear convergence factor, it can achieve rapid coarse adjustment in the early stage of iteration and micron-level precise fine adjustment in the later stage. The computational efficiency is far superior to the traditional matrix inversion algorithm. By acquiring the audience's six-degree-of-freedom head data through an infrared motion capture system, and combining ray intersection and affine transformation, the audience's gaze point on the giant painting can be accurately calculated. With the cross-fade-in and fade-out mechanism of the grid boundary buffer, even the slightest movement of the audience's gaze on the giant painting is converted into a seamless switching of the corresponding local narration audio in real time. Moreover, the audio accurately locks onto and follows the audience in the form of an invisible sound beam, greatly enhancing the immersive experience of art appreciation.

[0031] In one alternative implementation, an infrared motion capture system deployed around a high-definition large-format display terminal is used to collect six-degree-of-freedom motion data of each viewer's head, and to construct a global state matrix and an array coordinate matrix of the directional audio array for all viewers, including: S2011: Using an infrared motion capture system deployed around a high-definition large-format display terminal, by tracking reflective markers on each viewer's headband, and using a rigid body calculation algorithm, six degrees of freedom motion data of each viewer's head in the physical world coordinate system are collected. The six degrees of freedom motion data includes the... n Three-dimensional coordinates of the audience's head centroid Euler angles of head posture , For the first n The three-dimensional coordinates of the audience in the physical world coordinate system. For the first n The yaw and pitch angles of the audience n To indicate the quantity to the audience, l For time indication; In this embodiment, an infrared motion capture system consisting of infrared motion capture cameras is deployed above and around a high-definition large-format display terminal (such as an 8K digital mural screen). The audience wears a lightweight headband with at least three non-collinear reflective markers. The infrared motion capture system captures the reflective points in real time at a high frame rate (such as 60Hz-120Hz) and obtains the three-dimensional coordinates and Euler angles of each audience member's head through rigid body calculation. S2012: Combining timestamps, the six-DOF motion data of all viewers are assembled to obtain the global state matrix, with the formula as follows: In the formula, For the first l The global state matrix at time t; For the first l The first moment The three-dimensional coordinates of the centroid of an audience member's head; For the first l The first moment Euler angles of the head posture of each audience member; N Total number of viewers; S2013: Obtain the fixed spatial coordinates of each sound-emitting unit in the directional audio array deployed around the high-definition large-format display terminal. , For the first k The fixed spatial coordinates of the sound-generating unit. k This is the indicator value for the sound-emitting unit; S2014: Combining the timestamps, the fixed spatial coordinates of all sound-emitting units are assembled to obtain the array coordinate matrix of the directional audio array, as shown in the formula: In the formula, For the first l The array coordinate matrix at each moment; For the first The fixed spatial coordinates of each sound-producing unit; K This represents the total number of sound-producing units.

[0032] In one optional implementation, based on the global state matrix, gaze point mapping and audio matching are performed. The coordinates of each viewer's gaze point on the large-scale painting on the high-definition giant display terminal are calculated, and the corresponding local audio narration data is matched, including: S2021: For each viewer, extract the corresponding six-degree-of-freedom motion data from the global state matrix and generate the corresponding gaze direction unit vector, using the following formula: In the formula, For the first n The unit vector representing the direction of the viewer's gaze; T It is the transpose symbol; S2022: Starting from the three-dimensional coordinates of the head's centroid, generate the viewer's gaze ray along the corresponding gaze direction unit vector, and calculate the three-dimensional coordinates of the intersection point of the gaze ray and the physical plane equation of the high-definition giant display terminal screen. The formula is: In the formula, For the first n The three-dimensional coordinates of the audience's intersection point; The specific distance parameter at which the line-of-sight ray hits the screen; The plane equation parameters are the physical plane equations of the screen of a high-definition large-format display terminal in the global physical world coordinate system; These are absolute spatial coordinate variables in the global physical world coordinate system; For the first n The components of the unit vector of the viewer's line of sight in the world coordinate system; For the first n The unit vector representing the direction of the viewer's gaze; For the first l The first moment n The three-dimensional coordinates of the audience's head centroid; l For time indication; S2023: Map the three-dimensional coordinates of the intersection point to the two-dimensional pixel coordinate system of the large-scale painting through a pre-calibrated affine transformation matrix to obtain the coordinates of the point where the gaze falls on the large-scale painting. The formula is: In the formula, For the first n The viewer's gaze rests on the coordinates of the focal point of the giant painting; It is the affine transformation matrix; For the first n The three-dimensional coordinates of the audience's intersection point; S2024: Obtain the coordinates of the viewpoint by looking up the preset grid matrix of the artwork content. The audio data segment corresponding to the grid; S2025: If the duration of the gaze point's dwell time on its grid does not reach the preset trigger threshold, then the local audio narration data from the previous moment will continue to be output. S025: If the retention time reaches the preset trigger threshold and the audio data segment is within the two-grid boundary buffer, the mixing factor is calculated inversely proportional to the distance, and the current audio data segment and the target audio data segment are cross-faded in and out to obtain the corresponding local audio narration data. Otherwise, the audio data segment is directly used as the corresponding local audio narration data, and the formula is: In the formula, For the first l The audio data segment after the crossfade in and fade-out at each moment, as the first n Local audio narration data of the audience; For the first l The current audio data segment and the target audio data segment at any given moment; It is a mixing factor; In this embodiment, since the screen may have physical splicing gaps or non-planar micro-deformations, an affine transformation matrix obtained in advance through a calibration board is used to map the three-dimensional coordinates of the intersection point to the coordinates of the viewpoint in the pixel coordinates of the painting. The backend server has a "painting content grid matrix" that divides the huge painting into several grids (such as a 50×50 grid). Each grid is bound to an independent audio data segment. When the viewer's gaze moves between the grids, cross-fade-in and fade-out are performed through a mixing factor to avoid sudden sound interruptions.

[0033] In one alternative implementation, for each local audio narration data, an acoustic evacuation topology map is constructed based on the global state matrix and the array coordinate matrix, including the particles to be evacuated, the target exit, and dynamic obstacles, comprising: S2031: Extract all sound-emitting units from the array coordinate matrix, defining them as several particles to be evacuated. Each particle to be evacuated has fixed spatial coordinates, and initializes the corresponding two-dimensional virtual initial velocity vector, as shown in the formula: In the formula, For the first k The initial velocity vector of the sound-producing unit will k As a particle indicator; The initial velocity vector is represented by its two-dimensional component in the two-dimensional plane containing the directional audio array. In this embodiment, several sound-emitting units are defined as "particles to be evacuated". The reason for the need for "virtual velocity" is that in the physical world, the sound-emitting units are fixedly welded to the speaker box and cannot be moved to avoid obstacles. Therefore, the algorithm gives the particles the ability to move in a virtual two-dimensional topological space. The optimal "safe avoidance position" that the particles find actually corresponds to the "optimal emission phase and amplitude" that the sound-emitting unit should adopt. S2032: For the current audience, extract the three-dimensional coordinates of the audience's head centroid from the global state matrix, and define the three-dimensional coordinates of the head centroid as the target exit assigned to the audience; S2033: For the current audience, define the three-dimensional coordinates of the head centroid of all other audiences in the global state matrix except the current audience as the dynamic obstacle of the current audience; In this embodiment, in the topology diagram, the target audience is the "exit" that the particles must reach, while other audiences are "obstacles" that they absolutely cannot touch. This setting directly transforms the mathematical constraints of "maximizing the sound pressure at the target" and "minimizing the sound pressure at non-target" in beamforming into intuitive geometric attraction and repulsion constraints. S2034: For the current audience, the three-dimensional coordinates of the particles to be evacuated, the target exit, and the dynamic obstacles in the three-dimensional space are orthogonally projected onto the two-dimensional plane where the directional audio array is located, generating the corresponding fixed space two-dimensional topological coordinates, the target exit two-dimensional topological coordinates, and the dynamic obstacle two-dimensional topological coordinates. In this embodiment, orthogonal projection (usually projecting along the Z-axis to the XY horizontal plane) is a necessary dimensionality reduction method. Since the beamwidth of the ultrasonic parametric array in the vertical direction (Z-axis) is usually wide and relatively fixed, the main acoustic interference and crosstalk occur in the horizontal plane (when the audience stands side by side). Therefore, stripping the Z-axis and constructing the topology map only in the two-dimensional plane can reduce the computational complexity, which is a key prerequisite for meeting millisecond-level real-time performance. S2035: Integrate the fixed two-dimensional topological coordinates of each particle to be evacuated, the initial velocity vector, the target exit two-dimensional topological coordinates of the corresponding target exit, and the dynamic obstacle two-dimensional topological coordinates of all dynamic obstacles to obtain the current acoustic evacuation topology map of the audience. In this embodiment, the problem of solving the beamforming weighting coefficient of the directional audio array is transformed into the particle evacuation problem in the "improved escape optimization algorithm". The sound-emitting unit is regarded as "particles to be evacuated", the target audience is regarded as "target exit", and other audiences are regarded as "dynamic obstacles". By calculating the repulsive force (to avoid the sound beam disturbing others) and the driving force (to accurately focus on the target)", the industry problem of crosstalk between directional audio in dense multi-audience scenarios is fundamentally solved.

[0034] In one alternative implementation, based on an acoustic evacuation topology, an improved escape optimization algorithm is used to simulate the emergency evacuation behavior of particles to be evacuated. By iteratively calculating the force state of the particles, a complex beamforming weighting coefficient matrix for the directional audio array is generated, including: S2041: Based on the acoustic evacuation topology map, the fixed spatial two-dimensional topological coordinates of all particles to be evacuated are taken as the initial positions, and the corresponding initial velocity vectors are taken as the initial velocities of the particles to be evacuated. The formula is as follows: In the formula, For the first k The initial position of the particles to be evacuated; For the first k The fixed-space two-dimensional topological coordinates of the particles to be evacuated; For the first k The initial velocity of the particles to be evacuated; The initial velocity vector; t This is an indicator of the number of iterations. S2042: Initiate the improved escape optimization algorithm and calculate the comprehensive force on the current particle to be evacuated to the corresponding target exit during the current iteration. S2043: Introducing a convergence factor and Levy flight mechanism, combined with comprehensive forces, the initial velocity and initial position of the current particle to be evacuated are updated to obtain the updated position and updated velocity. S2044: Based on the updated position, calculate the fitness value of the current particle to be evacuated using a preset fitness function; S2045: Repeatedly update the position of the current particle to be evacuated. When the current iteration count reaches the maximum iteration count or the rate of change of fitness value is less than the rate of change threshold, terminate the position update and output the velocity of the particle to be evacuated with the best fitness value as the optimal velocity. S2046: The optimal velocity direction is mapped to the beam steering angle, and acoustic phase compensation is performed in conjunction with the actual three-dimensional physical distance between the sound unit and the audience to generate complex beamforming weighting coefficients. The formula is as follows: In the formula, For the first k The sound-emitting unit corresponding to the particles to be evacuated is the first n Complex beamforming weighting coefficients for the three-dimensional coordinates of the head centroid corresponding to the target exit of the audience; For the first k The optimal velocity of the particles to be evacuated Two-dimensional components in the two-dimensional plane containing the directional audio array; e The base of the exponent; The imaginary unit; It is the arctangent function in the four quadrants; For beam steering angle compensation; For the final applied acoustic phase; The carrier frequency of the ultrasonic parametric array; c Speed ​​of sound; For the first k The sound-emitting unit corresponding to the particles to be evacuated and the first n The physical distance of the audience; It is a scaling factor; This is a step from the "virtual mechanical world" back to the "physical acoustic world," achieving the optimal speed. The modulus (velocity magnitude) represents the energy (amplitude-weighted) that the sound-generating unit should contribute when constructing this sound beam, and the direction angle of the optimal virtual velocity (via...) The precise solution for the quadrant represents the time delay (phase weighting) that should be compensated when the sound-emitting unit emits sound waves. S2047: Integrating the complex beamforming weighting coefficients of all particles to be evacuated, the size of the directional audio array is obtained as follows: K × N The complex beamforming weighting coefficient matrix, N For the total number of viewers, K This represents the total number of sound-producing units.

[0035] In one alternative implementation, an improved escape optimization algorithm is initiated. During the current iteration, the combined forces experienced by the current particle to be evacuated to the corresponding target exit are calculated, including: S20421: Initiate the improved escape optimization algorithm. In the current iteration, calculate the driving force received by the current particle to be evacuated towards the corresponding target exit, using the following formula: In the formula, For the first t +1 iterations of the first iteration k The driving force received by the particles to be evacuated toward the corresponding target exit; For the first n The target exit of the audience is located in two-dimensional topological coordinates. For the first k The location of the particles to be evacuated, in the first iteration, is... ; The driving force coefficient controls the "determination" of the main lobe of the sound beam towards the target. This force is essentially the direction vector between the target position and the current particle position. It ensures that no matter how the particle avoids obstacles, its macroscopic movement trend is always towards the audience. S20422: Calculate the repulsive force exerted by dynamic obstacles on the current particle to be evacuated, and integrate the total repulsive force of all dynamic obstacles on the corresponding particle to be evacuated. The formula is: In the formula, For the first t +1 iterations of the first iteration k The particles to be evacuated are subjected to the first m The repulsive force of dynamic obstacles; For the first k Particles to be dispersed to the first n The audience's target exit is subject to the total repulsive force of dynamic obstacles; For the first m Two-dimensional topological coordinates of dynamic obstacles; m For dynamic obstacle indication; N Total number of viewers; To achieve the repulsion effect, when a particle approaches an obstacle in virtual space, the repulsive force increases exponentially. In the final sound field, this manifests as the algorithm actively adjusting the phase of certain sound-emitting units, causing the sound waves they emit to undergo "destructive interference (180-degree phase difference)" at the obstacle location, thus forming an invisible "soundproof wall" in the ears of bystanders. It is the repulsive force coefficient; S20423: Calculate the combined force on the particle to be evacuated based on the driving force and the total repulsive force, using the following formula: In the formula, For the first t +1 iterations of the first iteration k Particles to be dispersed to the first n The combined forces affecting the audience's target exit.

[0036] In one alternative implementation, the update formulas for the velocity and position are: In the formula, For the first t+ 1, t The iteration of the ... k The velocity of the particles to be evacuated; This is the acceleration coefficient; Let be a random perturbation vector in the interval [0,1]. For the first t The convergence factor of the next iteration; Scaling factor The corresponding Levy flight step size is usually very short (local fine-tuning), but there is a very small probability of producing a huge step size (long-distance jump). In acoustic optimization, when the sound beam is trapped in the "acoustic dead angle" (local minimum) formed by two obstacles, the conventional algorithm will completely stop, while the sudden jump of Levy flight can instantly "eject" the particles out of the dead angle and find the correct acoustic path again. For the first t +1 iterations of the first iteration k Particles to be dispersed to the first n The combined forces affecting the target exit of the audience; The symbol for element-wise product; In the formula, It is a random vector that follows a normal distribution; is the stability index of the Levy distribution; For the first t The optimal speed for the next iteration; In the formula, For the first t+ 1, t The iteration of the ... k The updated position of the particles to be evacuated.

[0037] In one alternative implementation, the fitness function is formulated as follows: In the formula, For the first t+ The 1st iteration k Fitness value of the particles to be evacuated; For the first n The target exit for the audience is the first k The target attraction value of the particles to be evacuated; For the first m Dynamic obstacles for the first k The barrier repulsion value of the particles to be evacuated; The radius of the acoustically quiet zone (danger radius, such as 0.3 meters); For fitness weighting coefficients; For the first n The target exit of the audience is located in two-dimensional topological coordinates. For the first m Two-dimensional topological coordinates of dynamic obstacles; It is an exponential function; This is a danger penalty value; This is the penalty coefficient.

[0038] In one alternative implementation, an audio signal vector for each local audio narration data is generated, combined with a complex beamforming weighting coefficient matrix to construct a dynamic sound field, and a directional audio stream is delivered using a directional audio array, including: S2051: Integrate the local audio narration data of all viewers to generate the corresponding audio signal vector; S2052: Perform matrix multiplication on the audio signal vector and the complex beamforming weighting coefficient matrix to generate the final time-domain drive signal vector for all sound-generating units in the array coordinate matrix. The formula is as follows: In the formula, For the first l The final time-domain driving signal vector at time step; This is the complex beamforming weighting coefficient matrix; For the first l The audio signal vector at that moment; For the first l The first moment k The driving signal for the sound-producing unit; These are complex beamforming weighting coefficients; For the first l The first moment n Local audio narration data of the audience; In this embodiment, for the first k The sound-producing unit, which ultimately plays the sound It is not a pure narration of any one person, but a "weighted superposition mixture" of all the narration audio of the audience. When this mixture is played by an ordinary sound unit, it is noisy noise. However, due to the complex weight filtering based on the improved escape optimization algorithm, this mixture signal will magically "separate" in space after passing through the nonlinear self-demodulation effect of the ultrasonic parametric array, and accurately reassemble into the original narration of each audience member at the position of each person's ear. S2053: After each driving signal in the final time-domain driving signal vector is converted by a digital-to-analog converter (DAC) and amplified, it is input to the ultrasonic parametric array of piezoelectric ceramics in the corresponding sound-generating unit in the directional audio array. S2054: Executes the drive signal and uses the ultrasonic parametric array of sound-generating units to radiate an ultrasonic beam carrying audio information outward, thereby realizing directional audio push for the appreciation of large-scale paintings. In this embodiment, since audible sound waves (20Hz-20kHz) diffuse significantly in the air, ultrasonic waves with a carrier frequency of approximately 40kHz are used. The DAC then samples the sound at a rate of at least 44.1kHz. Converted into an analog signal, the signal is amplified and then driven by a piezoelectric ceramic. Utilizing the nonlinear propagation effect of ultrasound in the air (minor deformation of the medium), the low-frequency voice signal it carries demodulates into audible sound along the propagation path like a ghost. This sound wave has extremely strong directivity (like the beam of a flashlight), thus perfectly achieving the ultimate goal of "sound following people without interference".

[0039] This invention also provides a directional audio push device 300 for appreciating large-scale paintings, see reference. Figure 3 The device may include the following units: The infrared motion capture unit 301 is used to collect six degrees of freedom motion data of each viewer's head using an infrared motion capture system deployed around a high-definition large-format display terminal, and to construct a global state matrix and an array coordinate matrix of the directional audio array for all viewers. The audio matching unit 302 is used to perform gaze point mapping and audio matching based on the global state matrix, calculate the coordinates of the gaze point of each viewer on the giant painting on the high-definition giant display terminal, and match the corresponding local audio narration data. The acoustic evacuation topology construction unit 303 is used to construct an acoustic evacuation topology map, including the particles to be evacuated, the target exit, and dynamic obstacles, based on the global state matrix and the array coordinate matrix for each local audio interpretation data. The beamforming weighting solution unit 304 is used to simulate the emergency evacuation behavior of particles to be evacuated based on the acoustic evacuation topology map and using an improved escape optimization algorithm. By iteratively calculating the force state of the particles to be evacuated, it generates a complex beamforming weighting coefficient matrix for the directional audio array. The directional audio push unit 305 is used to generate the audio signal vector of each local audio narration data, combine it with the complex beamforming weighting coefficient matrix to construct a dynamic sound field, and use the directional audio array to perform directional audio stream push.

[0040] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. Memory, used to store computer programs; When a processor executes a program stored in memory, it implements the directional audio push method for appreciating large-scale paintings according to the present invention.

[0041] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EI) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned terminal and other devices. The memory can include Random Access Memory (RAM), or non-volatile memory, such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.

[0042] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0043] Furthermore, to achieve the above objectives, embodiments of the present invention also propose a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the directional audio push method for appreciating large-scale paintings according to embodiments of the present invention.

[0044] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable hardware devices (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0045] The embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (apparatus), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0046] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0047] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0048] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. "And / or" indicates that either one or both can be chosen. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element.

[0049] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for targeted audio delivery for appreciating large-scale paintings, characterized in that, The method includes: Using an infrared motion capture system deployed around a high-definition large-format display terminal, six degrees of freedom motion data of each viewer's head are collected, and a global state matrix and an array coordinate matrix of the directional audio array are constructed for all viewers. Based on the global state matrix, gaze point mapping and audio matching are performed. The coordinates of each viewer's gaze point on the giant painting on the high-definition giant display terminal are calculated, and the corresponding local audio narration data is matched. For each local audio narration data, an acoustic evacuation topology map including the particles to be evacuated, the target exit, and dynamic obstacles is constructed based on the global state matrix and the array coordinate matrix. Based on the acoustic evacuation topology, an improved escape optimization algorithm is used to simulate the emergency evacuation behavior of particles to be evacuated. By iteratively calculating the force state of the particles to be evacuated, a complex beamforming weighting coefficient matrix of the directional audio array is generated. The audio signal vector of each local audio narration data is generated, and a dynamic sound field is constructed by combining it with a complex beamforming weighting coefficient matrix. A directional audio array is then used to perform directional audio stream delivery.

2. The method for directional audio push for appreciating large-scale paintings according to claim 1, characterized in that, Using an infrared motion capture system deployed around a high-definition large-format display terminal, six-degree-of-freedom motion data of each viewer's head is collected, and a global state matrix and an array coordinate matrix of the directional audio array are constructed for all viewers, including: Using an infrared motion capture system deployed around a high-definition large-format display terminal, the system tracks reflective markers on each viewer's headband and uses a rigid body calculation algorithm to collect six degrees of freedom motion data of each viewer's head in the physical world coordinate system. The six degrees of freedom motion data includes the three-dimensional coordinates of the head's center of mass and the Euler angles of the head posture. By combining the timestamps, the six-degree-of-freedom motion data of all viewers are assembled to obtain the global state matrix; Obtain the fixed spatial coordinates of each sound-emitting unit in the directional audio array deployed around the high-definition large-format display terminal; By combining the timestamps, the fixed spatial coordinates of all the sound-generating units are assembled to obtain the array coordinate matrix of the directional audio array.

3. The method for directional audio push for appreciating large-scale paintings according to claim 2, characterized in that, Based on the global state matrix, gaze point mapping and audio matching are performed. The coordinates of each viewer's gaze point on the giant painting on the high-definition giant display terminal are calculated, and the corresponding local audio narration data is matched, including: For each viewer, extract the corresponding six-degree-of-freedom motion data from the global state matrix and generate the corresponding gaze direction unit vector. Starting from the three-dimensional coordinates of the head's centroid, the gaze ray of the viewer is generated along the corresponding gaze direction unit vector, and the three-dimensional coordinates of the intersection point of the gaze ray and the physical plane equation of the screen of the high-definition giant display terminal are obtained. The three-dimensional coordinates of the intersection point are mapped to the two-dimensional pixel coordinate system of the giant painting through a pre-calibrated affine transformation matrix to obtain the coordinates of the point where the line of sight falls on the giant painting. By looking up a pre-defined grid matrix of artwork content, the audio data segment corresponding to the grid to which the gaze point's coordinates belong can be obtained. If the duration of the gaze point's dwell time on its corresponding grid does not reach the preset trigger threshold, then the local audio narration data from the previous moment will continue to be output. If the duration reaches the preset trigger threshold and the audio data segment is in the boundary buffer between two grids, the mixing factor is calculated inversely proportional to the distance, and the current audio data segment and the target audio data segment are cross-faded in and out to obtain the corresponding local audio narration data. Otherwise, the audio data segment is directly used as the corresponding local audio narration data.

4. The method for directional audio push for appreciating large-scale paintings according to claim 3, characterized in that, For each local audio narration data point, an acoustic evacuation topology map is constructed based on the global state matrix and array coordinate matrix, including the particles to be evacuated, the target exit, and dynamic obstacles, comprising: Extract all sound-emitting units from the array coordinate matrix and define them as several particles to be evacuated. Each particle to be evacuated has fixed spatial coordinates and initializes the corresponding two-dimensional virtual initial velocity vector. For the current audience, extract the three-dimensional coordinates of the audience's head centroid from the global state matrix, and define the three-dimensional coordinates of the head centroid as the target exit assigned to the audience; For the current audience, the three-dimensional coordinates of the head centroid of all other audiences in the global state matrix, excluding the current audience, are defined as the dynamic obstacles of the current audience. For the current audience, the three-dimensional coordinates of the particles to be evacuated, the target exit, and the dynamic obstacles in the three-dimensional space are orthogonally projected onto the two-dimensional plane where the directional audio array is located, generating the corresponding two-dimensional topological coordinates of the fixed space, the two-dimensional topological coordinates of the target exit, and the two-dimensional topological coordinates of the dynamic obstacles. By integrating the fixed two-dimensional topological coordinates of each particle to be evacuated, the initial velocity vector, the corresponding two-dimensional topological coordinates of the target exit, and the two-dimensional topological coordinates of all dynamic obstacles, the acoustic evacuation topology map of the current audience is obtained.

5. The method for directional audio push for appreciating large-scale paintings according to claim 4, characterized in that, Based on the acoustic evacuation topology, an improved escape optimization algorithm is used to simulate the emergency evacuation behavior of particles to be evacuated. By iteratively calculating the force state of the particles, a complex beamforming weighting coefficient matrix for the directional audio array is generated, including: Based on the acoustic evacuation topology, the fixed two-dimensional topological coordinates of all particles to be evacuated are taken as the initial positions, and the corresponding initial velocity vectors are taken as the initial velocities of the particles to be evacuated. Initiate an improved escape optimization algorithm to calculate the combined forces experienced by the current evacuated particle as it moves to the corresponding target exit during the current iteration. By introducing a convergence factor and the Levy flight mechanism, and combining the combined forces, the initial velocity and initial position of the particle to be evacuated are updated to obtain the updated position and updated velocity. Based on the updated position, the fitness value of the current particle to be evacuated is calculated using a preset fitness function; Repeatedly update the position of the current particle to be evacuated. When the current iteration count reaches the maximum iteration count or the rate of change of fitness value is less than the rate of change threshold, terminate the position update and output the velocity of the particle to be evacuated with the best fitness value as the optimal velocity. The optimal velocity direction is mapped to the beam steering angle, and acoustic phase compensation is performed by combining the actual three-dimensional physical distance between the sound unit and the audience to generate complex beamforming weighting coefficients. By integrating the complex beamforming weighting coefficients of all particles to be evacuated, a complex beamforming weighting coefficient matrix for the directional audio array is obtained.

6. The method for directional audio push for appreciating large-scale paintings according to claim 5, characterized in that, Initiate the improved escape optimization algorithm. In the current iteration, calculate the comprehensive forces experienced by the current particle to be evacuated to the corresponding target exit, including: Initiate an improved escape optimization algorithm to calculate the driving force received by the current evacuation particle toward the corresponding target exit during the current iteration. Calculate the repulsive force exerted by dynamic obstacles on the current particle to be evacuated, and integrate the total repulsive force of all dynamic obstacles on the corresponding particle to be evacuated; Calculate the combined force on the current particle to be evacuated based on the driving force and the total repulsive force.

7. The method for directional audio push for appreciating large-scale paintings according to claim 6, characterized in that, The update formulas for the velocity and position are as follows: In the formula, For the first t+ 1, t The iteration of the ... k The velocity of the particles to be evacuated; This is the acceleration coefficient; Let be a random perturbation vector in the interval [0,1]. For the first t The convergence factor of the next iteration; Scaling factor Corresponding Levy flight stride; For the first t +1 iterations of the first iteration k Particles to be dispersed to the first n The combined forces affecting the target exit of the audience; The symbol for element-wise product; In the formula, It is a random vector that follows a normal distribution; is the stability index of the Levy distribution; For the first t The optimal speed for the next iteration; In the formula, For the first t+ 1, t The iteration of the ... k The updated position of the particles to be evacuated.

8. The method for directional audio push for appreciating large-scale paintings according to claim 7, characterized in that, The formula for the fitness function is: In the formula, For the first t+ The 1st iteration k Fitness value of the particles to be evacuated; For the first n The target exit for the audience is the first k The target attraction value of the particles to be evacuated; For the first m Dynamic obstacles for the first k The barrier repulsion value of the particles to be evacuated; The radius of the acoustically quiet zone (danger radius, such as 0.3 meters); For fitness weighting coefficients; For the first n The target exit of the audience is located in two-dimensional topological coordinates. For the first m Two-dimensional topological coordinates of dynamic obstacles; It is an exponential function; This is a danger penalty value; This is the penalty coefficient.

9. The method for directional audio push for appreciating large-scale paintings according to claim 8, characterized in that, The audio signal vector for each local audio narration data is generated, and a dynamic sound field is constructed by combining it with a complex beamforming weighting coefficient matrix. A directional audio array is then used to perform directional audio stream delivery, including: Integrate local audio narration data from all viewers to generate corresponding audio signal vectors; The audio signal vector is multiplied by the complex beamforming weighting coefficient matrix to generate the final time-domain drive signal vector of all sound units in the array coordinate matrix. Each driving signal in the final time-domain driving signal vector is converted by a DAC and amplified by power, and then input to the ultrasonic parametric array of piezoelectric ceramics in the corresponding sound unit of the directional audio array. By executing the drive signal, the ultrasonic parametric array of sound-generating units radiates an ultrasonic beam carrying audio information outward, thereby enabling directional audio delivery for appreciating large-scale paintings.

10. A directional audio push device for appreciating large-scale paintings, used to implement the directional audio push method for appreciating large-scale paintings as described in any one of claims 1-9, characterized in that, The device includes: The infrared motion capture unit is used to collect six-degree-of-freedom motion data of each viewer's head using an infrared motion capture system deployed around a high-definition large-format display terminal, and to construct a global state matrix and an array coordinate matrix of the directional audio array for all viewers. The audio matching unit is used to perform gaze point mapping and audio matching based on the global state matrix, calculate the coordinates of each viewer's gaze point on the giant painting on the high-definition giant display terminal, and match the corresponding local audio narration data. The acoustic evacuation topology construction unit is used to construct an acoustic evacuation topology map, including the particles to be evacuated, the target exit, and dynamic obstacles, based on the global state matrix and the array coordinate matrix for each local audio interpretation data. The beamforming weighting solution unit is used to simulate the emergency evacuation behavior of particles to be evacuated based on the acoustic evacuation topology map and using an improved escape optimization algorithm. It generates the complex beamforming weighting coefficient matrix of the directional audio array by iteratively calculating the force state of the particles to be evacuated. The directional audio push unit is used to generate the audio signal vector of each local audio narration data, combine it with the complex beamforming weighting coefficient matrix to construct a dynamic sound field, and use the directional audio array to perform directional audio stream push.