Stage light following positioning method based on acousto-optic synchronization

By using an infrared camera and microphone array for audio-visual synchronization, the problem of inaccurate positioning and response delay in stage follow spot systems under complex environments has been solved, achieving high-precision and stable automated follow spot effects.

CN120871033APending Publication Date: 2025-10-31ZHEJIANG DAFENG IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510929289.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing stage follow spot systems suffer from problems such as response delay, low accuracy, high labor costs, image recognition being easily affected by clothing reflections and changes in lighting, and pure audio positioning being easily interfered with by noise in complex performance environments, making it difficult to achieve stable and accurate automated tracking.

Method used

An audio-visual synchronization method combining an infrared camera and a microphone array is adopted. Through infrared image delighting, dynamic threshold binarization, autocorrelation echo recognition, and sound wave localization, an adaptive audio-visual fusion algorithm is constructed to achieve accurate and real-time tracking of actors.

Benefits of technology

It significantly improves the robustness and accuracy of the stage follow spot system in complex environments, ensuring that the beam stably follows the actors, and enhances the level of automation and the viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120871033A_ABST
    Figure CN120871033A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of light control, and discloses a stage light following positioning method based on acousto-optic synchronization, which comprises the following steps of: uniformly arranging infrared cameras and microphone arrays on the edge and the top of a stage, acquiring a target heat source through an infrared image, and extracting a stable heat target area in combination with a dynamic threshold and a reflection suppression algorithm. Then, the gravity center of a heat source is extracted through connected domain analysis and converted into three-dimensional space coordinates; in parallel, a sound source localization model is constructed by using a time difference between microphones, echo interference is identified and filtered through autocorrelation analysis, and an accurate sound source frequency is extracted. And finally, through constructing a fusion positioning function, carrying out minimum distance matching on image space candidate points and sound source coordinates, and determining a final light following position. The method has high robustness and real-time performance, obviously improves the light following precision in the stage performance environment, is suitable for complex photoacoustic interference scenes, and has high practical value and innovativeness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of lighting control technology, specifically to a stage follow spot positioning method based on sound and light synchronization. Background Technology

[0002] Follow spotting systems in stage performances play a crucial role in shaping character portrayals, creating atmosphere, and enhancing stage effects. Traditional stage follow spotting methods primarily rely on manual operation, with operators visually observing actor positions and manually adjusting the beam direction. However, this method suffers from numerous drawbacks, including significant response delays, low accuracy, and high labor costs. Especially in complex stage environments, tracking errors are prone to occur, impacting performance quality. Therefore, automating and intelligentizing stage follow spotting has become a significant trend in the development of performing arts technology.

[0003] Currently, some automatic tracking technologies on the market are primarily based on image recognition methods. For example, they capture stage images using visible light cameras and then use computer vision algorithms to identify and track targets. However, in practical applications, these methods have significant limitations: on the one hand, actors' costumes on stage often contain highly reflective materials, which can easily create bright areas under strong lighting, interfering with the accuracy of visual recognition; on the other hand, the lighting at performance venues changes frequently, and the image acquisition process is easily affected by fluctuations in light intensity, resulting in unstable image quality. Furthermore, current methods generally neglect the auxiliary value of audio signals in actor positioning, making it difficult to maintain stable tracking in scenarios with multiple target occlusions or visual blur.

[0004] To address these issues, some studies have attempted to incorporate audio information into tracking systems, using microphone arrays to locate actors' voices. However, due to frequent audience noise and complex stage structures that easily create sound echoes, pure audio localization often lacks accuracy and is prone to misjudgment and other issues. Furthermore, most current fusion methods rely on manually setting weighting factors to adjust the degree of sound-light fusion, lacking stability and universality, making it difficult to flexibly adapt to different performance scenarios.

[0005] Therefore, there is an urgent need for a stage tracking and positioning method based on sound and light synchronization that can effectively cope with costume reflections, lighting changes, audience noise, and stage echo interference in complex performance environments. This method should integrate infrared image processing and microphone array sound source localization technology, and implement a highly robust fusion algorithm at the data level to ensure that the tracking system can accurately and in real-time track performers in fully automated operation, thereby significantly improving the level of stage automation control and the audience experience. Summary of the Invention

[0006] This invention provides a stage spotlight positioning method based on sound and light synchronization, which helps to solve the problems mentioned in the background art.

[0007] This invention provides the following technical solution: a stage spotlight positioning method based on sound and light synchronization, comprising:

[0008] Get the stage edge outline and the stage top outline;

[0009] N infrared cameras, denoted as C1, C2, ..., C6, are evenly distributed along the edge and top contours of the stage. N ;

[0010] Among them, C N This represents the Nth infrared camera;

[0011] K microphone arrays, denoted as M1, M2, ..., M, are evenly distributed along the edge of the stage. K ;

[0012] Each microphone array contains n k One microphone;

[0013] Image data is acquired through an infrared camera, and the infrared intensity value in the image is used for light removal processing.

[0014] Based on the image data after light removal processing, background candidate regions and binarized images are extracted;

[0015] Obtain the connected components in the binarized image and extract the barycenter coordinates;

[0016] Collect raw sound source data, and preliminarily determine the sound source coordinates based on the difference in sound wave transmission distance and time between microphones;

[0017] Construct an autocorrelation function for echo recognition;

[0018] Extract the echo frequency and construct a band-stop filter function to filter the echo;

[0019] By fusing acoustic and optical data, the final spatial coordinates of the sound source are calculated.

[0020] Optionally, the step of acquiring image data via an infrared camera and performing light-removing processing based on the infrared intensity values ​​in the image includes:

[0021] The image reflection suppression algorithm is constructed as follows:

[0022] The image captured by the infrared camera is denoted as...

[0023] in:

[0024] t represents the time frame index;

[0025] Indicates infrared camera C iImage data acquired at time frame index t, where i∈[1,N];

[0026] Get Image The infrared intensity value at pixel (x, y) is denoted as...

[0027] The dynamic threshold is constructed as follows:

[0028]

[0029] in:

[0030] Representing an image The dynamic threshold;

[0031] The average intensity of an image is expressed as follows:

[0032]

[0033] Where H and W represent the height and width of the image, respectively, in pixels;

[0034] The standard deviation of an image is expressed as follows:

[0035]

[0036] like Determine the image The pixel at (x, y) is a highly reflective area, and intensity suppression is applied, specifically as follows:

[0037] Image The infrared intensity value at pixel (x,y) is set to

[0038] Optionally, the step of extracting background candidate regions and binarized images based on the image data after light removal processing includes:

[0039] Calculate dynamic thermal threshold Specifically as follows:

[0040] Get Image Enter all pixel values ​​and input the pixel value set.

[0041] Extract background candidate region pixel set Specifically:

[0042]

[0043] Among them, T prelimThis indicates the initial upper limit threshold of thermal intensity, specifically expressed as: Indicates exclusion The top 10% of pixel values;

[0044] Calculate the set of background candidate region pixels mean Specifically as follows:

[0045]

[0046] in, express The number of elements in the array, Ip is Elements within;

[0047] Calculate the set of background candidate region pixels Standard deviation Specifically as follows:

[0048]

[0049] Build The expression is

[0050] Where k is the dynamic thermal coefficient; it is set to 3, i.e., the 3σ criterion, to ensure that the heat source area is fully separated from the background.

[0051] Define and extract binarized images Specifically as follows:

[0052]

[0053] in:

[0054] 1 and 0 represent the binarized image, respectively. The mask at pixel (x,y);

[0055] express The mask at pixel (x,y);

[0056] This indicates the image before intensity suppression. The infrared intensity value at pixel (x,y).

[0057] Optionally, obtaining the connected components in the binarized image and extracting the barycenter coordinates includes:

[0058] extract Connected components in the array, and set the j-th connected component as...

[0059] calculate barycentric coordinates Specifically as follows:

[0060]

[0061] in, express The number of connected components in the array.

[0062] Optionally, the acquisition of raw sound source data, and the preliminary determination of sound source coordinates based on the difference in sound wave transmission distance and time between microphones, includes:

[0063] Construct a Cartesian coordinate system for the stage space as follows:

[0064] Obtain the center of the stage surface and use it as the origin of the coordinate system (0,0,0);

[0065] Draw a ray perpendicular to the plane of the stage curtain from the origin of the coordinate system, and take the direction of this ray as the positive direction;

[0066] Draw a ray through the origin of the coordinate system, with the direction of the ray opposite to the positive direction, and use this ray as the Y-axis;

[0067] On the stage surface, draw a ray perpendicular to the Y-axis through the origin, with the ray located to the left of the Y-axis. Use this ray as the X-axis.

[0068] Draw a vertical ray through the origin of the coordinate system and use this ray as the Z-axis;

[0069] Define the location of the sound source as

[0070] Define M K The position coordinates of each microphone are as follows:

[0071] in, This represents the spatial coordinates of the nth microphone in the Kth microphone array;

[0072] Select M K Using the first microphone as the reference microphone, calculate M. K The reception time difference between the other microphones and the reference microphone is as follows:

[0073]

[0074] in:

[0075] M represents K The reception time of the first microphone;

[0076] M represents K The reception time of the i-th microphone;

[0077] M represents K middle and The time difference in acceptance;

[0078] Based on the relationship between the sound wave transmission distance difference and the time difference, the following equation holds:

[0079]

[0080] in:

[0081] M represents K The spatial coordinates of the first microphone in the middle;

[0082] M represents K The spatial coordinates of the i-th microphone;

[0083] V s This indicates the speed at which sound travels through the air.

[0084] Calculate M K The differences between the remaining microphones and the reference microphone are squared and then summed to construct a least-squares objective function, and the sound source coordinates are solved, as follows:

[0085]

[0086]

[0087] Optionally, the construction of the autocorrelation function for echo recognition includes:

[0088] Acquire microphone array M K The sound signal acquired at time frame index t is denoted as...

[0089] τ represents the sampling point index, N represents the number of sampling points per frame, and the specific calculation formula is N = F s ×ΔTime;

[0090] Among them, F s This indicates the audio acquisition frequency, and ΔTime represents the length of the time window for each frame.

[0091] The expression for the autocorrelation function is constructed as follows:

[0092]

[0093] in:

[0094] s_k represents the number of lag points, s_k∈[1,N-2];

[0095] ψ(s_k) represents the autocorrelation value at the lag number s_k;

[0096] Set the autocorrelation threshold to 0.3;

[0097] If there are two or more instances of ψ(s_k) greater than the autocorrelation threshold, an echo is considered to exist.

[0098] Optionally, the step of extracting the echo frequency and constructing a band-stop filter function for filtering echoes includes:

[0099] right Perform a Fast Fourier Transform:

[0100]

[0101] in:

[0102] Represents the Fast Fourier Transform operator;

[0103] This represents the frequency amplitude at the sampling point index τ;

[0104] Select The resonant frequency in the middle is denoted as f. e ;

[0105] The band-stop filter function is constructed as follows:

[0106]

[0107] in:

[0108] δ represents the frequency band tolerance, with a value of ±10 Hz;

[0109] Perform the filtering operation as follows:

[0110]

[0111] right Perform the inverse Fourier transform, specifically:

[0112]

[0113] in:

[0114] This represents the sound signal at index τ_i of the filtered sample point;

[0115] This represents the inverse Fourier transform.

[0116] Optionally, the fusion of acoustic-optical data to calculate the final spatial coordinates of the sound source includes:

[0117] Obtain the centroid coordinates of the images captured by each camera. Convert this to stage space coordinates, as follows:

[0118]

[0119] in:

[0120] express The corresponding stage space coordinates;

[0121] Indicates camera C i The inverse of the projection matrix;

[0122] Indicates camera C i The depth of the human body below;

[0123] Get each corresponding And input the set of spatial candidate points, denoted as PH_SET;

[0124] The final localization function is constructed, and the final spatial coordinates of the sound source are calculated, as follows:

[0125]

[0126] in:

[0127] Indicates the final spatial coordinates of the sound source;

[0128] P represents an element in PH_SET;

[0129] argmin represents a mathematical operator that minimizes the value of the following expression.

[0130] ||·||2 represents the Euclidean distance symbol, indicating the three-dimensional straight-line distance between two points.

[0131] The present invention has the following beneficial effects:

[0132] 1. This stage lighting positioning method based on audio-visual synchronization analyzes the infrared intensity distribution in each frame of the image, dynamically calculates the average intensity and standard deviation of the current image, and constructs an adaptive threshold function. When the infrared intensity of a pixel exceeds this dynamic threshold, the system classifies it as a highly reflective area and suppresses its intensity, thereby reducing the interference of abnormally bright areas in the image on subsequent target recognition. This processing not only maintains the integrity of the image's thermal source features but also effectively eliminates noise from non-target reflections. This significantly improves the overall contrast of the image and makes the difference between the target area and the background clearer, which is helpful for subsequent target extraction and binarization. Secondly, by dynamically adjusting the threshold, the algorithm can adapt to the infrared intensity fluctuations caused by changes in illumination in different time frames, exhibiting strong environmental adaptability. Finally, this method provides higher-quality image input for subsequent steps to extract human connected components and centroid coordinates, significantly improving the robustness and accuracy of the overall positioning system, and laying the foundation for stable and efficient automatic lighting in complex performance environments.

[0133] 2. This method for stage lighting positioning based on sound and light synchronization first de-lights the infrared image to eliminate strong interference caused by clothing reflections or uneven lighting. Based on this, the system further calculates the pixel intensity set of the entire image and sorts this set, removing the top 10% of high-intensity pixels to initially eliminate heat source interference, thereby extracting candidate pixel regions that may represent the stage background. For the extracted background candidate regions, the system calculates their mean and standard deviation, and constructs a dynamic thermal threshold based on the three-times-standard-deviation principle to dynamically delineate the boundary between the background and the heat source. Subsequently, the image is binarized using this thermal threshold to form a black-and-white mask image, clearly distinguishing the target area from the background area. Further, the system extracts connected regions from the binarized image and considers each connected region as a possible heat source target. By calculating the centroid position of each connected region, the approximate position of the target heat source in the image from the stage viewpoint can be obtained. This method has significant advantages. First, the use of a dynamic threshold avoids the problem of poor adaptability of fixed thresholds in different performance environments, effectively handling complex situations such as changes in lighting and frequent switching of stage lights. Secondly, by eliminating highly reflective pixels and dynamically identifying background areas, the positioning accuracy can be significantly improved under conditions of highly reflective clothing or strong lighting. Furthermore, the centroid extraction method is not only computationally efficient but also possesses good stability, providing a reliable basis for subsequent 3D spatial positioning and acoustic-optical fusion. In summary, this image processing scheme enhances the system's adaptability to complex performance environments while ensuring accuracy, demonstrating significant practical value.

[0134] 3. This method, based on sound and light synchronization for stage tracking and positioning, first constructs a three-dimensional Cartesian coordinate system with the center of the stage surface as the origin. Coordinate axes are set according to the direction of the stage curtain and the ground orientation to ensure accurate description of the spatial positions of each microphone and sound source. Then, multiple microphone arrays are deployed along the stage edge, and the specific spatial position of each microphone in each array is recorded. During the actual performance, the system collects sound wave signals generated by the actors' voices in real time. To accurately calculate the sound source position, a specific microphone is selected as a reference point, and the time it takes to receive the sound wave is recorded. Subsequently, the system calculates the time difference between the sound waves received by other microphones and, based on the fixed relationship between the time difference and the speed of sound, calculates the propagation distance difference between different microphones. Since the microphone positions are known, the system establishes an optimization model that minimizes the error based on these propagation distance differences. By solving for the minimum value, the three-dimensional spatial coordinates of the sound source are calculated, achieving preliminary positioning. By using a multi-microphone array for collaborative computation, the error accumulation problem of single-point positioning is effectively overcome, improving the spatial accuracy of sound source localization. Secondly, by utilizing the physical characteristics of sound wave propagation instead of empirical parameters or weighting factors, the system's localization process is ensured to be stable and reliable, and not easily affected by parameter fluctuations. Thirdly, this method does not rely on image information and can continuously track the approximate location of actors even when they are temporarily out of the camera's field of view, providing a solid foundation for subsequent sound and light data fusion. Finally, this localization method has strong real-time performance and is suitable for dynamic and fast-moving scenes, ensuring that the tracking light response during the performance is not delayed or drifted, thereby improving the overall performance quality and automation level.

[0135] 4. This method for stage tracking and localization based on sound and light synchronization first utilizes sound signals acquired by a microphone array. An autocorrelation function is constructed for each time frame, and the similarity between different lag points of the sound signal is analyzed to determine the presence of echo. When multiple significant peaks appear in the autocorrelation result, sound wave echo interference is confirmed in the current frame. After identifying the echo, the sound signal is further analyzed in the frequency domain. A Fast Fourier Transform (FFT) is used to convert the time-domain signal into a frequency-domain signal, extracting the dominant frequency component corresponding to the echo. Based on this, a band-stop filter function is constructed to effectively block the dominant frequency and its adjacent frequency bands, thereby achieving precise suppression of the echo component. Finally, an Inverse Fourier Transform (IFT) is used to restore the processed signal back to the time domain, obtaining sound signal data with echo interference removed, providing high-quality audio information support for subsequent sound source localization. This method can adaptively identify and suppress sound wave echoes without relying on weight parameter adjustment, effectively enhancing the robustness of the system in complex stage structure environments. Secondly, by combining autocorrelation identification and frequency domain filtering, it achieves fine processing of various interference factors in sound signals, avoiding the problems of misjudging sound sources or location drift in traditional methods.

[0136] 5. This stage lighting positioning method based on sound and light synchronization acquires the target's centroid coordinates from images using multiple cameras and maps them to the stage's three-dimensional spatial coordinate system using depth information. The spatial coordinates obtained from each camera are uniformly recorded as a set of candidate positioning points. Simultaneously, the system accurately analyzes the three-dimensional coordinates of the sound source using a microphone array, and then selects the candidate point closest to the sound source coordinates as the final target positioning point using the minimum geometric distance criterion. This step not only achieves deep fusion of visual and sound data in spatial coordinate matching but also provides robust positioning assurance in noisy environments. Compared to traditional methods relying on a single image or sound source for positioning, this invention effectively avoids visual recognition errors such as high reflectivity of actor costumes and strong stage lighting interference, while significantly reducing sound source offset problems caused by audience noise and stage echoes. This not only significantly improves the positioning accuracy and response speed of the spotlight system in complex performance environments, ensuring that the beam can stably follow the target actor in real time, but also achieves adaptive acoustic-optical coordinated positioning without weighting factors, enhancing the system's versatility and environmental adaptability. Through multi-source redundancy verification, it improves the ability to eliminate abnormal signals, thereby significantly improving the intelligence level of the automatic stage spotlight system and the stability of actual performances. Attached Figure Description

[0137] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation

[0138] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0139] Example 1, see Figure 1 A stage spotlight positioning method based on sound and light synchronization includes:

[0140] Get the stage edge outline and the stage top outline;

[0141] N infrared cameras, denoted as C1, C2, ..., C6, are evenly distributed along the edge and top contours of the stage. N ;

[0142] Among them, C N This represents the Nth infrared camera;

[0143] K microphone arrays, denoted as M1, M2, ..., M, are evenly distributed along the edge of the stage. K ;

[0144] Each microphone array contains n k One microphone;

[0145] Image data is acquired through an infrared camera, and the infrared intensity value in the image is used for light removal processing.

[0146] Based on the image data after light removal processing, background candidate regions and binarized images are extracted;

[0147] Obtain the connected components in the binarized image and extract the barycenter coordinates;

[0148] Collect raw sound source data, and preliminarily determine the sound source coordinates based on the difference in sound wave transmission distance and time between microphones;

[0149] Construct an autocorrelation function for echo recognition;

[0150] Extract the echo frequency and construct a band-stop filter function to filter the echo;

[0151] By fusing acoustic and optical data, the final spatial coordinates of the sound source are calculated.

[0152] The step of acquiring image data via an infrared camera and performing light removal processing based on the infrared intensity values ​​in the image includes:

[0153] The image reflection suppression algorithm is constructed as follows:

[0154] The image captured by the infrared camera is denoted as...

[0155] in:

[0156] t represents the time frame index;

[0157] Indicates infrared camera C i Image data acquired at time frame index t, where i∈[1,N];

[0158] Get Image The infrared intensity value at pixel (x, y) is denoted as...

[0159] The dynamic threshold is constructed as follows:

[0160]

[0161] in:

[0162] Representing an image The dynamic threshold;

[0163] The average intensity of an image is expressed as follows:

[0164]

[0165] Where H and W represent the height and width of the image, respectively, in pixels;

[0166] The standard deviation of an image is expressed as follows:

[0167]

[0168] like Determine the image The pixel at (x, y) is a highly reflective area, and intensity suppression is applied, specifically as follows:

[0169] Image The infrared intensity value at pixel (x,y) is set to

[0170] By analyzing the infrared intensity distribution in each frame of the image, the system dynamically calculates the average intensity and standard deviation of the current image to construct an adaptive threshold function. When the infrared intensity of a pixel exceeds this dynamic threshold, the system classifies it as a highly reflective area and suppresses its intensity, thereby reducing the interference of abnormally bright areas in the image on subsequent target recognition. This processing not only maintains the integrity of the image's thermal source features but also effectively eliminates noise from non-target reflections. This significantly improves the overall contrast of the image and makes the difference between the target area and the background clearer, which is helpful for subsequent target extraction and binarization. Secondly, by dynamically adjusting the threshold, the algorithm can adapt to the fluctuations in infrared intensity caused by changes in illumination in different time frames, demonstrating strong environmental adaptability. Finally, this method provides higher-quality image input for subsequent steps to extract human connected components and centroid coordinates, significantly improving the robustness and accuracy of the overall positioning system and laying the foundation for stable and efficient automatic light tracking in complex performance environments.

[0171] The step of extracting background candidate regions and binarized images based on the image data after light removal processing includes:

[0172] Calculate dynamic thermal threshold Specifically as follows:

[0173] Get Image Enter all pixel values ​​and input the pixel value set.

[0174] Extract background candidate region pixel set Specifically:

[0175]

[0176] Among them, T prelim This indicates the initial upper limit threshold of thermal intensity, specifically expressed as: Indicates exclusion The top 10% of pixel values;

[0177] Calculate the set of background candidate region pixels mean Specifically as follows:

[0178]

[0179] in, express The number of elements in the array, Ip is Elements within;

[0180] Calculate the set of background candidate region pixels Standard deviation Specifically as follows:

[0181]

[0182] Build The expression is

[0183] Where k is the dynamic thermal coefficient; it is set to 3, i.e., the 3σ criterion, to ensure that the heat source area is fully separated from the background.

[0184] Define and extract binarized images Specifically as follows:

[0185]

[0186] in:

[0187] 1 and 0 represent the binarized image, respectively. The mask at pixel (x,y);

[0188] express The mask at pixel (x,y);

[0189] This indicates the image before intensity suppression. The infrared intensity value at pixel (x,y).

[0190] The step of obtaining connected components in the binarized image and extracting the barycenter coordinates includes:

[0191] extract Connected components in the array, and set the j-th connected component as...

[0192] A connected component is a region centered at a specified pixel and consisting of surrounding pixels with a foreground mask of 1.

[0193] calculate barycentric coordinates Specifically as follows:

[0194]

[0195] in, express The number of connected components in the array.

[0196] This method first performs de-glare processing on the infrared image to eliminate strong interference information caused by clothing reflections or uneven lighting. Based on this, the system further calculates the pixel intensity set of the entire image and sorts this set, removing the top 10% of high-intensity pixels to initially eliminate heat source interference, thereby extracting candidate pixel regions that may represent the stage background. For the extracted background candidate regions, the system calculates their mean and standard deviation, and constructs a dynamic thermal threshold based on the principle of three times the standard deviation, thereby dynamically defining the boundary between the background and the heat source. Subsequently, the image is binarized using this thermal threshold to form a black-and-white mask image, clearly distinguishing the target area from the background area. Further, the system extracts connected regions from the binarized image and considers each connected region as a possible heat source target. By calculating the centroid position of each connected region, the approximate location of the target heat source in the image from the stage viewpoint can be obtained. This method has significant advantages. First, the use of a dynamic threshold avoids the problem of poor adaptability of fixed thresholds in different performance environments, effectively handling complex situations such as changes in lighting and frequent switching of stage lights. Secondly, by eliminating highly reflective pixels and dynamically identifying background areas, the positioning accuracy can be significantly improved under conditions of highly reflective clothing or strong lighting. Furthermore, the centroid extraction method is not only computationally efficient but also possesses good stability, providing a reliable basis for subsequent 3D spatial positioning and acoustic-optical fusion. In summary, this image processing scheme enhances the system's adaptability to complex performance environments while ensuring accuracy, demonstrating significant practical value.

[0197] The process of collecting raw sound source data, and preliminarily determining the sound source coordinates based on the difference in sound wave transmission distance and time between microphones, includes:

[0198] Construct a Cartesian coordinate system for the stage space as follows:

[0199] Obtain the center of the stage surface and use it as the origin of the coordinate system (0,0,0);

[0200] Draw a ray perpendicular to the plane of the stage curtain from the origin of the coordinate system, and take the direction of this ray as the positive direction;

[0201] Draw a ray through the origin of the coordinate system, with the direction of the ray opposite to the positive direction, and use this ray as the Y-axis;

[0202] On the stage surface, draw a ray perpendicular to the Y-axis through the origin, with the ray located to the left of the Y-axis. Use this ray as the X-axis.

[0203] Draw a vertical ray through the origin of the coordinate system and use this ray as the Z-axis;

[0204] Define the location of the sound source as

[0205] Define M K The position coordinates of each microphone are as follows:

[0206] in, This represents the spatial coordinates of the nth microphone in the Kth microphone array;

[0207] Select M K Using the first microphone as the reference microphone, calculate M. K The reception time difference between the other microphones and the reference microphone is as follows:

[0208]

[0209] in:

[0210] M represents K The reception time of the first microphone;

[0211] M represents K The reception time of the i-th microphone;

[0212] M represents K middle and The time difference in acceptance;

[0213] Based on the relationship between the sound wave transmission distance difference and the time difference, the following equation holds:

[0214]

[0215] in:

[0216] M represents K The spatial coordinates of the first microphone in the middle;

[0217] M represents K The spatial coordinates of the i-th microphone;

[0218] V s This indicates the speed at which sound travels through the air.

[0219] Calculate M K The differences between the remaining microphones and the reference microphone are squared and then summed to construct a least-squares objective function, and the sound source coordinates are solved, as follows:

[0220]

[0221] This method first constructs a three-dimensional Cartesian coordinate system, with the center of the stage surface as the origin. Coordinate axes are set according to the direction of the stage curtain and the orientation of the ground to ensure accurate description of the spatial positions of each microphone and sound source. Then, multiple microphone arrays are deployed along the stage edge, and the specific spatial position of each microphone in each array is recorded. During the actual performance, the system collects the sound wave signals generated by the actors' voices in real time. To accurately calculate the sound source position, a specific microphone is first selected as a reference point, and the time it takes to receive the sound wave is recorded. Subsequently, the system calculates the time difference between the sound waves received by other microphones and, based on the fixed relationship between the time difference and the speed of sound, calculates the propagation distance difference between different microphones. Since the microphone positions are known, the system establishes an optimization model that minimizes the error based on these propagation distance differences. By solving for the minimum value, the three-dimensional spatial coordinates of the sound source are calculated, achieving preliminary localization. By using a multi-microphone array for collaborative computation, the error accumulation problem of single-point positioning is effectively overcome, improving the spatial accuracy of sound source localization. Secondly, by utilizing the physical characteristics of sound wave propagation instead of empirical parameters or weighting factors, the system's localization process is ensured to be stable and reliable, and not easily affected by parameter fluctuations. Thirdly, this method does not rely on image information and can continuously track the approximate location of actors even when they are temporarily out of the camera's field of view, providing a solid foundation for subsequent sound and light data fusion. Finally, this localization method has strong real-time performance and is suitable for dynamic and fast-moving scenes, ensuring that the tracking light response during the performance is not delayed or drifted, thereby improving the overall performance quality and automation level.

[0222] The construction of the autocorrelation function for echo recognition includes:

[0223] Acquire microphone array M K The sound signal acquired at time frame index t is denoted as...

[0224] τ represents the sampling point index, N represents the number of sampling points per frame, and the specific calculation formula is N = F s ×ΔTime;

[0225] Among them, F s This indicates the audio acquisition frequency, and ΔTime represents the length of the time window for each frame.

[0226] The expression for the autocorrelation function is constructed as follows:

[0227]

[0228] in:

[0229] s_k represents the number of lag points, s_k∈[1,N-2];

[0230] ψ(s_k) represents the autocorrelation value at the lag number s_k;

[0231] Set the autocorrelation threshold to 0.3;

[0232] If there are two or more instances of ψ(s_k) greater than the autocorrelation threshold, an echo is considered to exist.

[0233] The extraction of echo frequency and the construction of a band-stop filter function for echo filtering include:

[0234] right Perform a Fast Fourier Transform:

[0235]

[0236] in:

[0237] Represents the Fast Fourier Transform operator;

[0238] This represents the frequency amplitude at the sampling point index τ;

[0239] Select The resonant frequency in the middle is denoted as f. e ;

[0240] The band-stop filter function is constructed as follows:

[0241]

[0242] in:

[0243] δ represents the frequency band tolerance, with a value of ±10 Hz;

[0244] Perform the filtering operation as follows:

[0245]

[0246] right Perform the inverse Fourier transform, specifically:

[0247]

[0248] in:

[0249] This represents the sound signal at index τ_i of the filtered sample point;

[0250] This represents the inverse Fourier transform.

[0251] This method first utilizes sound signals acquired by a microphone array to construct an autocorrelation function for each time frame. By analyzing the similarity of the sound signal across different lag points, it determines whether an echo phenomenon exists. The presence of multiple significant peaks in the autocorrelation result confirms the existence of acoustic echo interference in the current frame. After identifying the echo, the sound signal undergoes further frequency domain analysis. A Fast Fourier Transform (FFT) is used to convert the time-domain signal to the frequency-domain signal, extracting the dominant frequency component corresponding to the echo. Based on this, a band-stop filter function is constructed to effectively block the dominant frequency and its adjacent frequency bands, thereby achieving precise suppression of the echo component. Finally, an Inverse Fourier Transform (IFT) is used to restore the processed signal back to the time domain, obtaining sound signal data with echo interference removed, providing high-quality audio information support for subsequent sound source localization. This method can adaptively identify and suppress sound wave echoes without relying on weight parameter adjustment, effectively enhancing the robustness of the system in complex stage structure environments. Secondly, by combining autocorrelation identification and frequency domain filtering, it achieves fine processing of various interference factors in sound signals, avoiding the problems of misjudging sound sources or location drift in traditional methods.

[0252] The fused acoustic-optical data is used to calculate the final spatial coordinates of the sound source, including:

[0253] Obtain the centroid coordinates of the images captured by each camera. Convert this to stage space coordinates, as follows:

[0254]

[0255] in:

[0256] express The corresponding stage space coordinates;

[0257] Indicates camera C i The inverse of the projection matrix;

[0258] Indicates camera C i The depth of the human body below;

[0259] Get each corresponding And input the set of spatial candidate points, denoted as PH_SET;

[0260] The final localization function is constructed, and the final spatial coordinates of the sound source are calculated, as follows:

[0261]

[0262] in:

[0263] Indicates the final spatial coordinates of the sound source;

[0264] P represents an element in PH_SET;

[0265] argmin represents a mathematical operator that minimizes the value of the following expression.

[0266] ||·||2 represents the Euclidean distance symbol, indicating the three-dimensional straight-line distance between two points.

[0267] The system acquires the target's centroid coordinates from images using multiple cameras and maps them to the stage's three-dimensional spatial coordinate system using depth information. The spatial coordinates obtained from each camera are uniformly recorded as a set of candidate positioning points. Simultaneously, the system precisely analyzes the three-dimensional coordinates of the sound source using a microphone array, and then selects the candidate point closest to the sound source coordinates as the final target positioning point using the minimum geometric distance criterion. This step not only achieves deep fusion of visual and sound data in spatial coordinate matching but also provides robust positioning assurance in noisy environments. Compared to traditional methods relying on a single image or sound source for positioning, this invention effectively avoids visual recognition errors caused by high reflectivity of actor costumes and strong stage lighting interference, while significantly reducing sound source offset problems caused by audience noise and stage echoes. This not only greatly improves the positioning accuracy and response speed of the tracking system in complex performance environments, ensuring the beam can stably follow the target actor in real time, but also achieves adaptive acoustic-optical collaborative positioning without weighting factors, enhancing the system's versatility and environmental adaptability. Through multi-source redundancy verification, it improves the ability to eliminate abnormal signals, thereby significantly improving the intelligence level and actual performance stability of the automatic stage tracking system.

[0268] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0269] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A stage spotlight positioning method based on sound and light synchronization, characterized in that, include: Get the stage edge outline and the stage top outline; N infrared cameras, denoted as C1, C2, ..., C6, are evenly distributed along the edge and top contours of the stage. N ; Among them, C N This represents the Nth infrared camera; K microphone arrays, denoted as M1, M2, ..., M, are evenly distributed along the edge of the stage. K ; Each microphone array contains n k One microphone; Image data is acquired through an infrared camera, and the infrared intensity value in the image is used for light removal processing. Based on the image data after light removal processing, background candidate regions and binarized images are extracted; Obtain the connected components in the binarized image and extract the barycenter coordinates; Collect raw sound source data, and preliminarily determine the sound source coordinates based on the difference in sound wave transmission distance and time between microphones; Construct an autocorrelation function for echo recognition; Extract the echo frequency and construct a band-stop filter function to filter the echo; By fusing acoustic and optical data, the final spatial coordinates of the sound source are calculated.

2. The stage spotlight positioning method based on sound and light synchronization according to claim 1, characterized in that: The step of acquiring image data via an infrared camera and performing light removal processing based on the infrared intensity values ​​in the image includes: The image reflection suppression algorithm is constructed as follows: The image captured by the infrared camera is denoted as... in: t represents the time frame index; Indicates infrared camera C i Image data acquired at time frame index t, where i∈[1,N]; Get Image The infrared intensity value at pixel (x, y) is denoted as... The dynamic threshold is constructed as follows: in: Representing an image The dynamic threshold; The average intensity of an image is expressed as follows: Where H and W represent the height and width of the image, respectively, in pixels; The standard deviation of an image is expressed as follows: like Determine the image The pixel at (x, y) is a highly reflective area, and intensity suppression is applied, specifically as follows: Image The infrared intensity value at pixel (x,y) is set to 3. The stage spotlight positioning method based on sound and light synchronization according to claim 1, characterized in that: The step of extracting background candidate regions and binarized images based on the image data after light removal processing includes: Calculate dynamic thermal threshold Specifically as follows: Get Image Enter all pixel values ​​and input the pixel value set. Extract background candidate region pixel set Specifically: Among them, T prelim This indicates the initial upper limit threshold of thermal intensity, specifically expressed as: Indicates exclusion The top 10% of pixel values; Calculate the set of background candidate region pixels mean Specifically as follows: in, express The number of elements in the array, Ip is Elements within; Calculate the set of background candidate region pixels Standard deviation Specifically as follows: Build The expression is Where k is the dynamic thermal coefficient; it is set to 3, i.e., the 3σ criterion, to ensure that the heat source area is fully separated from the background. Define and extract binarized images Specifically as follows: in: 1 and 0 represent the binarized image, respectively. The mask at pixel (x,y); express The mask at pixel (x,y); This indicates the image before intensity suppression. The infrared intensity value at pixel (x,y).

4. The stage spotlight positioning method based on sound and light synchronization according to claim 1, characterized in that: The step of obtaining connected components in the binarized image and extracting the barycenter coordinates includes: extract Connected components in the array, and set the j-th connected component as... calculate barycentric coordinates Specifically as follows: in, express The number of connected components in the array.

5. A stage spotlight positioning method based on sound and light synchronization according to claim 1, characterized in that: The process of collecting raw sound source data, and preliminarily determining the sound source coordinates based on the difference in sound wave transmission distance and time between microphones, includes: Construct a Cartesian coordinate system for the stage space as follows: Obtain the center of the stage surface and use it as the origin of the coordinate system (0,0,0); Draw a ray perpendicular to the plane of the stage curtain from the origin of the coordinate system, and take the direction of this ray as the positive direction; Draw a ray through the origin of the coordinate system, with the direction of the ray opposite to the positive direction, and use this ray as the Y-axis; On the stage surface, draw a ray perpendicular to the Y-axis through the origin, with the ray located to the left of the Y-axis. Use this ray as the X-axis. Draw a vertical ray through the origin of the coordinate system and use this ray as the Z-axis; Define the location of the sound source as Define M K The position coordinates of each microphone are as follows: in, This represents the spatial coordinates of the nth microphone in the Kth microphone array; Select M K Using the first microphone as the reference microphone, calculate M. K The reception time difference between the other microphones and the reference microphone is as follows: in: M represents K The reception time of the first microphone; M represents K The reception time of the i-th microphone; M represents K middle and The time difference in acceptance; Based on the relationship between the sound wave transmission distance difference and the time difference, the following equation holds: in: M represents K The spatial coordinates of the first microphone in the middle; M represents K The spatial coordinates of the i-th microphone; V s This indicates the speed at which sound travels through the air. Calculate M K The differences between the remaining microphones and the reference microphone are squared and then summed to construct a least-squares objective function, and the sound source coordinates are solved, as follows:

6. The stage spotlight positioning method based on sound and light synchronization according to claim 1, characterized in that: The construction of the autocorrelation function for echo recognition includes: Acquire microphone array M K The sound signal acquired at time frame index t is denoted as... τ represents the sampling point index, N represents the number of sampling points per frame, and the specific calculation formula is N = F s ×ΔTime; Among them, F s This indicates the audio acquisition frequency, and ΔTime represents the length of the time window for each frame. The expression for the autocorrelation function is constructed as follows: in: s_k represents the number of lag points, s_k∈[1,N-2]; ψ(s_k) represents the autocorrelation value at the lag number s_k; Set the autocorrelation threshold to 0.3; If there are two or more instances of ψ(s_k) greater than the autocorrelation threshold, an echo is considered to exist.

7. A stage spotlight positioning method based on sound and light synchronization according to claim 1, characterized in that: The extraction of echo frequency and the construction of a band-stop filter function for echo filtering include: right Perform a Fast Fourier Transform: in: Represents the Fast Fourier Transform operator; This represents the frequency amplitude at the sampling point index τ; Select The resonant frequency in the middle is denoted as f. e ; The band-stop filter function is constructed as follows: in: δ represents the frequency band tolerance, with a value of ±10 Hz; Perform the filtering operation as follows: right Perform the inverse Fourier transform, specifically: in: This represents the sound signal at index τ_i of the filtered sample point; This represents the inverse Fourier transform.

8. The stage spotlight positioning method based on sound and light synchronization according to claim 1, characterized in that: The fused acoustic-optical data is used to calculate the final spatial coordinates of the sound source, including: Obtain the centroid coordinates of the images captured by each camera. Convert this to stage space coordinates, as follows: in: express The corresponding stage space coordinates; Indicates camera C i The inverse of the projection matrix; Indicates camera C i The depth of the human body below; Get each corresponding And input the set of spatial candidate points, denoted as PH_SET; The final localization function is constructed, and the final spatial coordinates of the sound source are calculated, as follows: in: Indicates the final spatial coordinates of the sound source; P represents an element in PH_SET; argmin represents a mathematical operator that minimizes the value of the following expression. ||·||2 represents the Euclidean distance symbol, indicating the three-dimensional straight-line distance between two points.