Gate trailing detection directional voice intervention method, device and equipment and medium
By using image recognition and spatial sound field control, directional voice intervention is dynamically generated, which solves the problems of insufficient real-time performance and accuracy in gate tailgating detection, realizes real-time directional intervention for tailgating behavior, and avoids interference from traditional methods.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-17
AI Technical Summary
Existing gate tailgating detection technology cannot achieve real-time, targeted voice intervention, resulting in unauthorized tailgating behavior not being handled in a timely and effective manner, and the broadcast-style prompts throughout the area cause interference.
Image coordinates of the person following the camera are obtained through image recognition. Combined with the installation posture parameters of the camera and directional speaker, the three-dimensional spatial coordinates and beam pointing angle of the person following the camera are calculated. The directional voice intervention is dynamically generated. The beam direction of the directional speaker is adaptively controlled to make the main lobe of the sound beam align with the head area of the person following the camera.
It enables real-time, directional voice intervention during gate passage, improves the accuracy of tailgating detection and intelligent control, avoids interference from broadcast prompts throughout the area, and ensures timely handling of unauthorized personnel.
Smart Images

Figure CN121884490A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of access control and security monitoring, and in particular to a method, device, equipment and medium for directional voice intervention for gate tailgating detection. Background Technology
[0002] Currently, turnstile access systems commonly employ image recognition-based identity verification and access control to determine the legitimacy of individuals entering and exiting. However, when multiple people pass through turnstiles consecutively, unauthorized individuals can still tailgate, following legitimate users. Existing tailgating detection technologies mostly analyze the distance between people, the timing of passage, or human posture features in video frames. They can only issue audible and visual alarms after identifying tailgating behavior, but cannot achieve real-time, targeted intervention for specific individuals. Summary of the Invention
[0003] To improve the collaborative performance of detection intervention in the gate tailgating detection process, this application provides a method, device, equipment, and medium for directional voice intervention in gate tailgating detection.
[0004] The above-mentioned objective of this application is achieved through the following technical solution:
[0005] A method for directional voice intervention in gate tailgating detection, comprising:
[0006] Acquire passage image information, perform human detection and identity recognition on the passage image information, distinguish between legally passing target personnel and tailgating personnel, and output the image coordinate information of tailgating personnel in real time;
[0007] The installation posture parameters of the camera and directional speaker are obtained. Based on the installation posture parameters, the image coordinate information of the following person is converted into spatial three-dimensional coordinate information. The azimuth and pitch angles of the following person are calculated, and the target sound beam pointing angle of the following person in the directional sound field is determined based on the azimuth and pitch angles.
[0008] The target sound beam pointing angle is compared with a preset angle threshold. When the angle separation corresponding to the target sound beam pointing angle is greater than the preset angle threshold and the distance between the tailing person and the gate threshold line is less than the preset distance threshold, a directional voice intervention trigger command is generated, and a target stability window is established for the tailing person.
[0009] Based on the spatial coordinate sequence within the target stabilization window, the beam direction adaptive control of the directional loudspeaker is executed, and the current beam pointing angle is obtained by correcting the real-time spatial position of the following personnel.
[0010] The azimuth error is determined by the difference between the beam pointing angle and the preset target beam direction angle. The direction of the main lobe of the sound beam is limited to the preset main lobe width range. The side lobe energy overflow is suppressed by phase smoothing compensation and beam calibration strategy. Under the condition that the azimuth error is less than the preset angle error threshold, the preset voice prompt information is issued to the head area of the following person.
[0011] By adopting the above technical solution, real-time identification and precise positioning of tailgating personnel can be achieved during gate passage by integrating image detection and spatial sound field control. The system establishes the spatial three-dimensional coordinate relationship of the tailgating personnel and calculates their azimuth and pitch angles. Based on the spatial position of the tailgating personnel, the target sound beam pointing angle is dynamically generated. Through adaptive control of the beam direction of the directional speaker, the main lobe energy of the sound beam is always aligned with the head area of the tailgating personnel, achieving real-time, directional voice intervention for tailgating behavior. This method differs from traditional image-based tailgating detection, which only reaches the "identification-alarm" stage. Instead, it further establishes a mapping relationship between spatial coordinates and sound beam pointing, automatically completing sound beam calibration and voice output the instant tailgating occurs. This solves the deficiency of existing technologies in being unable to provide real-time sound field intervention for unauthorized personnel, while avoiding interference from full-area broadcast prompts. It achieves high-precision tailgating detection and intelligent voice intervention control in gate scenarios.
[0012] In a preferred embodiment, this application can be further configured as follows: acquiring passage image information, performing human detection and identity recognition on the passage image information, distinguishing between legally passing target personnel and tailgating personnel, and outputting the image coordinate information of tailgating personnel in real time, including:
[0013] Perform target region extraction on the traffic image information to determine the candidate pedestrian area located within the gate channel;
[0014] Human feature detection is performed on the candidate pedestrian area to obtain the human feature detection results;
[0015] Based on the results of human feature detection and the preset identity recognition model, the identity tag information of the person passing through is determined;
[0016] The identity tag information is matched with the gate access authorization information to obtain the matching result;
[0017] When the matching results are consistent, the corresponding pedestrian candidate area is determined as the target person. When the matching results are inconsistent, the corresponding pedestrian candidate area is determined as the tailing person, and the coordinate information of the tailing person in the image is extracted.
[0018] By adopting the above technical solution, the automatic detection and area division of pedestrians in the gate channel can be achieved by extracting the target area based on the passage image information during the gate passage process. The identity tag information of the passage personnel is judged by combining the human feature detection results with the preset identity recognition model, and the gate passage authorization information is used as the comparison benchmark to realize the synchronous determination of identity verification. When the matching results are inconsistent, the trailing person is identified in real time and the coordinate information of the trailing person in the image is extracted. Thus, the accurate identification and spatial position calibration of the trailing person is achieved at the image level, which establishes an accurate input data foundation for subsequent three-dimensional spatial transformation and directional sound beam calculation, and solves the problems of trailing person detection delay and insufficient coordinate positioning accuracy in the existing passage recognition technology.
[0019] In a preferred embodiment, this application can be further configured as follows: The step of converting the image coordinate information of the following person into three-dimensional spatial coordinate information based on the installation pose parameters, and calculating the azimuth and pitch angles of the following person, includes:
[0020] Based on the installation pose parameters, the installation height, horizontal orientation angle and pitch installation angle of the camera are obtained, and the installation offset parameters of the directional speaker relative to the camera are acquired.
[0021] Establish a spatial coordinate system corresponding to the gate channel, map the image coordinate information of the tailing person to the spatial coordinate system, and calculate the initial position coordinate information of the tailing person in the spatial coordinate system.
[0022] Based on the camera's installation height and pitch angle, and combined with the initial position coordinate information, the depth distance information of the following person in the vertical direction is calculated.
[0023] Based on the horizontal orientation angle and installation offset parameters of the camera, a spatial geometric transformation is performed on the depth distance information to obtain the three-dimensional position coordinates of the following person in the spatial coordinate system.
[0024] Based on the spatial relationship between the three-dimensional position coordinates and the directional loudspeaker, the azimuth and pitch angles of the following person relative to the directional loudspeaker are calculated.
[0025] By adopting the above technical solution, the installation height, horizontal orientation angle and pitch angle of the camera can be obtained based on the installation posture parameters. A spatial coordinate system can be established by combining the installation offset parameters of the directional speaker. The image coordinate information of the following person is mapped into three-dimensional spatial position coordinates. The azimuth and pitch angles are calculated based on the spatial position relationship between the three-dimensional position coordinates and the directional speaker, so as to realize the real-time positioning and angle determination of the following person in three-dimensional space. This solves the problem that traditional tailing detection is based only on image planar information and cannot determine the spatial direction.
[0026] In a preferred embodiment, this application can be further configured as follows: comparing the target sound beam pointing angle with a preset angle threshold, and when the angle separation corresponding to the target sound beam pointing angle is greater than the preset angle threshold and the distance between the tailgating person and the gate threshold line is less than a preset distance threshold, generating a directional voice intervention trigger command, and establishing a target stabilization window for the tailgating person, including:
[0027] The difference between the target beam pointing angle and the preset angle threshold is calculated to obtain the angle separation amount;
[0028] Detect the coordinate changes of the following person in consecutive image frames, and calculate the real-time distance and movement direction angle of the following person;
[0029] When the angle separation is greater than the preset angle threshold and the real-time distance information is less than the preset distance threshold, and the movement direction angle tends to the gate passage direction, a directional voice intervention trigger command is generated.
[0030] After generating the directional voice intervention trigger command, the trajectory change rate of the following person is calculated based on the coordinate sequence of continuous frame images to obtain the spatial position stability parameter, and a target stability window is established based on the spatial position stability parameter.
[0031] By adopting the above technical solution, when a tailing person is detected approaching the gate area, the timing of the tailing behavior can be accurately identified by the joint judgment of angle separation, real-time distance information and movement direction angle. After the threshold condition is met, a directional voice intervention trigger command is generated. The trajectory change rate is calculated by combining the coordinate sequence of continuous frame images to establish a target stability window, thereby realizing dynamic locking of the target area of the tailing person and precise control of the intervention timing, solving the problems of intervention delay or false triggering in traditional methods.
[0032] In a preferred embodiment, this application can be further configured as follows: The step of performing adaptive beam direction control of the directional loudspeaker based on the spatial coordinate sequence within the target stabilization window, and correcting and obtaining the current beam pointing angle based on the real-time spatial position of the following person, includes:
[0033] Time series smoothing is performed on the spatial coordinate sequence within the target stable window to generate continuous spatial coordinate information;
[0034] Based on continuous spatial coordinate information, the changes in orientation and distance between adjacent coordinate points are calculated to obtain spatial displacement trend information.
[0035] The beam pointing angle correction is calculated based on the spatial displacement trend information and the preset target beam pointing angle.
[0036] The beam pointing angle correction is applied to beam direction adaptive control, and phase compensation and amplitude correction are performed to obtain the updated current beam pointing angle.
[0037] By adopting the above technical solution, the movement trajectory of the following person can be smoothed based on the spatial coordinate sequence within the target stable window, ensuring the continuity and stability of spatial position information. On this basis, the azimuth change and distance change of adjacent coordinate points are calculated to dynamically reflect the displacement trend of the following person during passage. Then, the beam pointing angle correction amount is determined according to the spatial displacement trend information and the preset target beam pointing angle, and the sound direction and amplitude distribution of the directional loudspeaker are adjusted in real time to achieve adaptive tracking and continuous locking of the sound beam pointing, thereby improving the spatial accuracy and response consistency of voice intervention.
[0038] In a preferred embodiment, this application can be further configured as follows: the step of performing beam direction adaptive control on the beam pointing angle correction amount, performing phase compensation and amplitude correction, to obtain the updated current beam pointing angle includes:
[0039] The beam pointing angle correction is received as a control input to initiate iterative calculations for beam direction adaptive control.
[0040] During the iterative calculation, the phase difference between adjacent sound units is calculated based on the beam pointing angle correction. The phase of each sound unit is adjusted sequentially according to the preset phase update step size to obtain the adjusted phase difference.
[0041] The amplitude correction is calculated based on the adjusted phase difference, and the output amplitude of the sound unit is proportionally corrected according to the amplitude correction to form an amplitude distribution result corresponding to the adjusted phase difference.
[0042] Based on the corrected amplitude distribution results and the changing trend of the beam pointing angle correction, the control weights of the beam direction adaptive control are adjusted to obtain the updated control weight distribution results.
[0043] The updated current beam pointing angle is calculated based on the adjusted phase difference, the corrected amplitude distribution, and the updated control weight distribution.
[0044] By adopting the above technical solution, the beam pointing angle correction can be used as a dynamic control input to initiate the iterative calculation process of beam direction adaptive control. During the iteration process, the phase difference between adjacent sound units is calculated based on the beam pointing angle correction, and the phase of each sound unit is adjusted sequentially according to the preset phase update step size, so that the beamforming direction gradually approaches the target pointing angle. At the same time, the amplitude correction is calculated based on the adjusted phase difference and the output amplitude of each sound unit is proportionally corrected to form the corrected amplitude distribution result. Then, the control weight distribution is dynamically adjusted in combination with the changing trend of the beam pointing angle correction. Finally, the updated current beam pointing angle is calculated based on the adjusted phase difference, the corrected amplitude distribution result, and the updated control weight distribution result, thereby ensuring the continuity and response accuracy of beam direction control.
[0045] In a preferred embodiment, this application can be further configured as follows: determining the azimuth error based on the difference between the beam pointing angle and the preset target beam direction angle, limiting the main lobe direction of the sound beam within a preset main lobe width range, and suppressing sidelobe energy overflow through phase smoothing compensation and beam calibration strategies, including:
[0046] The azimuth error is compared with a preset angle error threshold. When the azimuth error is less than the preset angle error threshold, the main lobe direction limitation operation is performed.
[0047] In the main lobe direction limiting operation, the sound beam radiation direction is adjusted according to the preset main lobe width range to limit the main lobe direction of the sound beam.
[0048] After the main lobe direction is defined, phase smoothing compensation is performed to correct the phase change of the beam direction. After the phase smoothing compensation is completed, a beam calibration strategy is performed to adjust the amplitude distribution of the beam and suppress sidelobe energy overflow.
[0049] By adopting the above technical solution, after determining the azimuth error, the azimuth error can be compared with a preset angle error threshold. When the azimuth error is lower than the preset angle error threshold, the main lobe direction limiting operation is initiated. In the main lobe direction limiting operation, the beam radiation direction is adjusted according to the preset main lobe width range to keep the main lobe of the beam within the limited angle range. Then, the phase smoothing compensation process is performed. By correcting the phase change of the sound unit point by point, the smooth transition of the beam direction is achieved. Finally, after the phase compensation is completed, the beam calibration strategy is executed to proportionally adjust the amplitude distribution of the sound unit to reduce the leakage of sidelobe energy, thereby achieving stable control and energy concentration of the beam radiation direction.
[0050] The second objective of this invention is achieved through the following technical solution:
[0051] A gate tailgating detection and directional voice intervention device, the gate tailgating detection and directional voice intervention device comprising:
[0052] The image recognition module is used to acquire passage image information, perform human detection and identity recognition on the passage image information, distinguish between legally passing target personnel and tailgating personnel, and output the image coordinate information of tailgating personnel in real time.
[0053] The pose calculation module is used to obtain the installation pose parameters of the camera and the directional speaker. Based on the installation pose parameters, the image coordinate information of the following person is converted into spatial three-dimensional coordinate information. The azimuth and pitch angles of the following person are calculated, and the target sound beam pointing angle of the following person in the directional sound field is determined based on the azimuth and pitch angles.
[0054] The intervention trigger module is used to compare the target sound beam pointing angle with a preset angle threshold. When the angle separation corresponding to the target sound beam pointing angle is greater than the preset angle threshold and the distance between the tailing person and the gate threshold line is less than the preset distance threshold, a directional voice intervention trigger command is generated, and a target stability window is established for the tailing person.
[0055] The beam control module is used to perform adaptive beam direction control of the directional loudspeaker based on the spatial coordinate sequence within the target stabilization window, and to correct and obtain the current beam pointing angle based on the real-time spatial position of the following person.
[0056] The voice output module is used to determine the azimuth error based on the difference between the beam pointing angle and the preset target beam direction angle, limit the main lobe direction of the sound beam within the preset main lobe width range, and suppress side lobe energy overflow through phase smoothing compensation and beam calibration strategies. Under the condition that the azimuth error is less than the preset angle error threshold, the module sends a preset voice prompt message to the head area of the following person.
[0057] By adopting the above technical solution, real-time identification and precise positioning of tailgating personnel can be achieved during gate passage by integrating image detection and spatial sound field control. The system establishes the spatial three-dimensional coordinate relationship of the tailgating personnel and calculates their azimuth and pitch angles. Based on the spatial position of the tailgating personnel, the target sound beam pointing angle is dynamically generated. Through adaptive control of the beam direction of the directional speaker, the main lobe energy of the sound beam is always aligned with the head area of the tailgating personnel, achieving real-time, directional voice intervention for tailgating behavior. This method differs from traditional image-based tailgating detection, which only reaches the "identification-alarm" stage. Instead, it further establishes a mapping relationship between spatial coordinates and sound beam pointing, automatically completing sound beam calibration and voice output the instant tailgating occurs. This solves the deficiency of existing technologies in being unable to provide real-time sound field intervention for unauthorized personnel, while avoiding interference from full-area broadcast prompts. It achieves high-precision tailgating detection and intelligent voice intervention control in gate scenarios.
[0058] The above-mentioned objective three of this application is achieved through the following technical solution:
[0059] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned gate tailgating detection directional voice intervention method.
[0060] The fourth objective of this application is achieved through the following technical solution:
[0061] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned gate tailgating detection directional voice intervention method.
[0062] In summary, this application includes at least one of the following beneficial technical effects:
[0063] 1. This system enables real-time identification and precise positioning of tailgating individuals during gate passage by integrating image detection and spatial sound field control. It establishes the spatial three-dimensional coordinate relationship of the tailgating individual and calculates the azimuth and pitch angles. Based on the spatial position of the tailgating individual, it dynamically generates the target sound beam pointing angle. Through adaptive control of the beam direction of the directional speaker, the main lobe energy of the sound beam is always aligned with the head area of the tailgating individual, achieving real-time, directional voice intervention for tailgating behavior. This method differs from traditional image analysis-based tailgating detection, which only stays at the "identification-alarm" stage. Instead, it further establishes a mapping relationship between spatial coordinates and sound beam pointing, automatically completing sound beam calibration and voice output the instant tailgating occurs. This solves the shortcomings of existing technologies that cannot provide real-time sound field intervention for unauthorized personnel, while avoiding interference from full-area broadcast prompts. It achieves high-precision tailgating detection and intelligent voice intervention control in gate scenarios. Attached Figure Description
[0064] Figure 1 This is a flowchart of a gate tailgating detection directional voice intervention method according to one embodiment of this application.
[0065] Figure 2 This is a flowchart illustrating the implementation of step S10 in a gate tailgating detection directional voice intervention method according to an embodiment of this application;
[0066] Figure 3 This is a flowchart illustrating the implementation of step S20 in a gate tailgating detection directional voice intervention method according to an embodiment of this application;
[0067] Figure 4This is a flowchart illustrating the implementation of step S30 in a gate tailgating detection directional voice intervention method according to an embodiment of this application;
[0068] Figure 5 This is a flowchart illustrating the implementation of step S40 in a gate tailgating detection directional voice intervention method according to an embodiment of this application;
[0069] Figure 6 This is a flowchart illustrating the implementation of step S404 in a gate tailgating detection directional voice intervention method according to an embodiment of this application;
[0070] Figure 7 This is a flowchart illustrating the implementation of step S50 in a gate tailgating detection directional voice intervention method according to an embodiment of this application;
[0071] Figure 8 This is a schematic diagram of a gate tailgating detection directional voice intervention device according to one embodiment of this application. Detailed Implementation
[0072] The present application will be further described in detail below with reference to the accompanying drawings.
[0073] In one embodiment, such as Figure 1 As shown, this application discloses a method for directional voice intervention to detect tailgating at turnstiles, which specifically includes the following steps:
[0074] S10: Acquire passage image information, perform human detection and identity recognition on the passage image information, distinguish between legally passing target personnel and tailgating personnel, and output the image coordinate information of tailgating personnel in real time.
[0075] In this embodiment, the passage image information consists of continuous image frames captured by a camera installed above the turnstile, covering the turnstile channel area. Human detection processing is performed on the passage image information, extracting regions with human structural features from each image frame to obtain pedestrian candidate regions, and acquiring their corresponding image coordinates. Identity recognition processing is performed based on these pedestrian candidate regions, matching the personnel feature information corresponding to the candidate regions with identity feature templates in the turnstile access authorization database, and determining target personnel and unauthorized personnel based on the matching results. When an unauthorized person is detected entering the turnstile channel together with the target person in terms of spatial location or temporal sequence, the unauthorized person is determined to be a tailgating person, and their image coordinates in the passage image information are extracted to characterize their position in the image coordinate system.
[0076] Specifically, a camera is fixedly installed above the turnstile, and the installation angle parameters are calibrated. When the turnstile detects a person approaching, the image acquisition process is initiated. Continuous image frames covering the turnstile channel area are acquired at a fixed frame rate to form passage image information. Human detection processing is performed frame by frame on the passage image information. Pedestrian candidate areas with human shape features are extracted from the images, and the corresponding area position coordinates are output. After obtaining the pedestrian candidate areas, identity recognition processing is performed on the pedestrian candidate areas. Based on the pedestrian candidate areas, the personnel identity feature information is extracted and matched with the identity feature template in the turnstile access authorization database. According to the matching results, the corresponding personnel are distinguished as target personnel or unauthorized personnel. Position association and trajectory analysis are performed on the pedestrian candidate areas in continuous image frames. When an unauthorized person follows the target person into the turnstile channel in terms of spatial position or passage time, the unauthorized person is determined to be a tailgating person, and the center coordinates of the pedestrian candidate area of the tailgating person in the passage image information are extracted as the image coordinate information of the tailgating person.
[0077] S20: Obtain the installation posture parameters of the camera and directional speaker. Based on the installation posture parameters, convert the image coordinate information of the following person into spatial three-dimensional coordinate information, calculate the azimuth and pitch angles of the following person, and determine the target sound beam pointing angle of the following person in the directional sound field based on the azimuth and pitch angles.
[0078] In this embodiment, the installation pose parameters refer to the set of parameters used to describe the installation position and orientation of the camera and the directional speaker in the gate channel space, including the installation height, horizontal orientation angle, pitch installation angle and installation offset distance between them. The target sound beam pointing angle refers to the direction angle of the main lobe of the sound beam emitted by the directional speaker relative to the reference coordinate system in three-dimensional space, which is used to determine the concentrated radiation direction of the sound field energy.
[0079] Specifically, during the equipment deployment phase, a laser rangefinder is used to obtain the camera's installation height hc and the directional speaker's emission center height hs, while an electronic compass is used to record the camera's horizontal orientation angle. With pitch installation angle The horizontal offset distance (dcs) between the camera and the directional speaker is measured, and the measured value is used as the input for the installation pose parameters. After obtaining the installation pose parameters, a spatial coordinate system corresponding to the gate channel is established. Where the X-axis is along the gate's passage direction, the Y-axis is horizontal, and the Z-axis is vertical, with the camera's optical center position as the origin, the image coordinate information (u,v) of the trailing person is input into the back projection function of the pinhole camera model, through the camera's intrinsic parameter matrix:
[0080] Performing a transformation from pixel coordinates to normalized planar coordinates yields: .
[0081] Normalized plane coordinates Combined with camera tilt angle The spatial depth distance Zp of the following person relative to the camera is calculated based on the installation height hc. The calculation formula is as follows: And obtain the horizontal displacement from the perspective relationship:
[0082] This allows us to obtain the three-dimensional spatial coordinates of the person following us in the camera's coordinate system. Then, the camera coordinate system is translated to the directional speaker reference coordinate system using a spatial coordinate transformation matrix:
[0083] , three-dimensional coordinate points Convert to coordinates in the directional speaker coordinate system Based on the geometric relationship between the coordinate point position vector and the speaker center axis vector, calculate the azimuth and elevation angles, where: , Then calculate the azimuth angle With pitch angle Substitute the sound field direction mapping function: in The target beam pointing angle, This represents the angular synthesis relationship based on the array geometry and the distribution of the sound-generating units. It is used to map spatial angles to the main lobe direction angle of the sound beam, thereby obtaining the target sound beam pointing angle of the following person in the directional sound field.
[0084] S30: Compare the target sound beam pointing angle with the preset angle threshold. When the angle separation corresponding to the target sound beam pointing angle is greater than the preset angle threshold and the distance between the tailing person and the gate threshold line is less than the preset distance threshold, generate a directional voice intervention trigger command and establish a target stability window for the tailing person.
[0085] In this embodiment, the angle threshold refers to the angle boundary value used to limit the deviation of the target sound beam pointing angle from the preset reference angle range, the preset distance threshold refers to the distance boundary value used to determine the spatial positional relationship between the tailing person and the gate threshold line, the directional voice intervention trigger command refers to the trigger signal generated to start sound beam control after the combined conditions of angle and distance are met, and the target stabilization window refers to the continuous coordinate data interval used to track the stable spatial position of the tailing person in the time series.
[0086] Specifically, the target beam pointing angle is input into the comparison module, and the difference between the target beam pointing angle and the reference angle is calculated using the angle difference calculation function. The formula for calculating the difference is as follows: ,in The target beam pointing angle, For the preset reference angle, Indicates the angular separation amount, when When the current frame is in an angle deviation state, record that the current frame is in such a state. To preset the angle threshold, and based on the spatial three-dimensional coordinate information of the trailing person obtained in the previous stage... The actual distance Dp between the trailing person and the gate threshold is calculated using the Euclidean distance formula: ,in The coordinates of the reference point of the gate threshold line in the spatial coordinate system are given when the distance calculation result satisfies... And the angular separation amount satisfies When the comparison module outputs a trigger signal logic "1", the control unit generates a directional voice intervention trigger command. After generating the directional voice intervention trigger command, a target stabilization window is established. The target stabilization window is constructed based on the spatial coordinate information of the trailing person in multiple consecutive frames, and the displacement difference between coordinates in every two frames is calculated. When the displacement difference is lower than the stability threshold When the number of frame sequences reaches the set window length Nw, the time interval is determined as the target stable window.
[0087] S40: Based on the spatial coordinate sequence within the target stabilization window, perform adaptive beam direction control of the directional loudspeaker, correct the beam pointing angle according to the real-time spatial position of the following personnel, and obtain the current beam pointing angle.
[0088] In this embodiment, beam direction adaptive control refers to a control method that dynamically adjusts the phase and amplitude distribution of the directional loudspeaker's sound-emitting unit based on the spatial coordinate sequence within the target stabilization window to change the beam radiation direction. The current beam pointing angle refers to the real-time direction angle of the directional loudspeaker's main lobe relative to the reference coordinate system after the spatial position correction is completed.
[0089] Specifically, the spatial coordinate sequence within the target stabilization window The coordinate sequence is input to the spatial variation analysis module, which performs time-series smoothing to remove short-term disturbances and calculates the average coordinates of three consecutive frames using a moving average algorithm.
[0090] After smoothing, a continuous spatial coordinate sequence is generated. Calculate the spatial displacement increment between adjacent coordinate points:
[0091] The instantaneous velocity vector of the trailing person is calculated based on the spatial displacement increment and the time interval Δt.
[0092] And calculate the motion trend angle based on the direction of the velocity vector. The previously calculated target beam pointing angle is corrected using the motion trend angle. The corrected formula is: in This is an adaptive correction coefficient, with a range of values. Used to control the response speed of direction correction, and to correct the pointing angle. As the input for beam direction adaptive control, the spatial phase difference between each sound-emitting unit of the directional loudspeaker and the main lobe direction is calculated:
[0093] Where dn is the distance from the nth sound-emitting unit to the center of the array. The wavelength of the sound wave. The geometric arrangement angle of the sound-generating units is determined by calculating the phase adjustment value of each sound-generating unit using the phase difference formula and performing a phase synchronization operation to complete one beam direction update. The updated pointing angle is then recorded as the current beam pointing angle.
[0094] S50: Determine the azimuth error based on the difference between the beam pointing angle and the preset target beam direction angle, limit the main lobe direction of the sound beam to the preset main lobe width range, and suppress side lobe energy overflow through phase smoothing compensation and beam calibration strategies. Under the condition that the azimuth error is less than the preset angle error threshold, issue a preset voice prompt message to the head area of the following person.
[0095] In this embodiment, azimuth error refers to the angular difference between the beam pointing angle and the target beam direction angle; main lobe width range refers to the allowable expansion angle range within the main lobe region of the directional loudspeaker beam; angle error threshold refers to the maximum limiting angle of the allowable azimuth error, used to control the beam pointing accuracy; and voice prompt information refers to a fixed voice segment stored in the voice signal library corresponding to the trailing event.
[0096] Specifically, calculate the beam pointing angle. relative to the target beam direction angle The difference is used to obtain the azimuth error. Its calculation formula is
[0097] The azimuth error is compared with a preset angle error threshold. When comparing, At that time, enter the main lobe-limited operation.
[0098] In the main lobe limiting operation, the beam radiation direction range is constrained to the main lobe width range with the target beam azimuth angle as the center. Inside, among which To preset the main lobe half-width, the weights of the directional loudspeaker array's sound-emitting units are redistributed according to the defined angular range. The weight calculation formula is as follows: ,in Given the radiation direction angle corresponding to the nth sound-generating unit, after defining the main lobe direction, a phase smoothing compensation operation is performed to calculate the phase difference change between adjacent sound-generating units: , And the phase difference change is processed by a moving average. Through the smoothed phase difference The emission phase of each sound-generating unit is corrected to ensure continuous change in the main lobe direction and concentrated energy. After phase smoothing compensation, a beam calibration strategy is executed to perform amplitude regularization calculation on the output amplitude of the sound-generating unit.
[0099] Where An' is the calibrated output amplitude, An is the original amplitude, and N is the number of sound generating units. Amplitude regularization ensures the stability of the overall sound pressure distribution and suppresses sidelobe energy overflow. When the azimuth error is kept less than the angle error threshold, the voice output module is activated, and the digital voice signal corresponding to the preset voice prompt information is loaded to the speaker driver port. After being modulated by the acoustic filter, it is output. The center direction of the main lobe of the sound beam points to the head area of the following person, realizing directional voice playback.
[0100] In one embodiment, such as Figure 2 As shown, in step S10, the passage image information is acquired, and human detection and identity recognition are performed on the passage image information to distinguish between legally passing target personnel and tailgating personnel, and the image coordinate information of tailgating personnel is output in real time, including:
[0101] S101: Perform target area extraction operation on the passage image information to determine the pedestrian candidate area located in the gate channel.
[0102] In this embodiment, the target region extraction operation refers to separating the image region with human features from the background region in the passing image information using spatial segmentation and motion feature recognition methods. The pedestrian candidate region refers to the set of pixels that have human morphological features and are initially determined through the target region extraction operation.
[0103] Specifically, the image information is input into the image preprocessing flow. First, image grayscale conversion and noise smoothing are performed. High-frequency noise is removed by mean filtering to maintain the overall brightness balance of the image. After preprocessing, a foreground segmentation step is performed. The brightness change of consecutive frames is calculated by the pixel difference method of adjacent frames. A motion region mask image is generated according to the pixel change amplitude. Pixel regions with brightness changes exceeding a preset threshold are marked as potential active regions. Then, morphological dilation and erosion operations are used to correct the boundaries of the motion regions and eliminate isolated small regions to form continuous region edges. Then, the spatial filtering range is set according to the geometric constraints of the gate channel, and only active regions located within the effective detection area of the channel are retained. Each independent active target region is identified by the connected component labeling method. The aspect ratio and area range of each active target region are calculated. Regions with aspect ratios within the human body proportion range and areas within the size range of passing personnel are identified as pedestrian candidate regions.
[0104] S102: Perform human feature detection on the pedestrian candidate region to obtain the human feature detection results.
[0105] In this embodiment, human feature detection refers to determining whether a pedestrian candidate region has human appearance features through image feature extraction and classification recognition algorithms and outputting the corresponding detection results. The human feature detection results refer to the human presence confidence value and corresponding bounding box information generated for each pedestrian candidate region.
[0106] Specifically, the pedestrian candidate region is input into the human feature detection model. First, local texture information and contour gradient distribution information are extracted through a convolutional feature extraction layer. During the feature extraction process, a multi-scale sliding window structure is used to scan the pedestrian candidate region pixel by pixel and generate candidate feature maps. Then, the candidate feature maps are input into the feature aggregation module. Max pooling is used to extract the main morphological features and suppress background interference. After feature aggregation is completed, the model enters the classification and recognition stage. In the classification and recognition stage, feature matching is performed based on the weights of the trained human recognition network. The similarity score between the pedestrian candidate region and the human template features is calculated. Based on the similarity score, the confidence information of human presence is output, and the position and size parameters of the detection box are marked in the pedestrian candidate region. Finally, the corresponding human feature detection result is generated.
[0107] S103: Based on the human feature detection results and the preset identity recognition model, determine the identity tag information of the person passing through.
[0108] In this embodiment, the preset identity recognition model refers to the identity discrimination network trained based on image samples of registered passers-in, and the identity tag information refers to the identity identification category data generated for each passer-in during the recognition process, which is used to indicate whether the passer-in is an authorized passer-in.
[0109] Specifically, the human feature detection results are input into a preset identity recognition model. First, image regions containing facial features and upper body shape are extracted within the detection box. Corresponding high-dimensional feature vectors are generated through a feature extraction network. The feature extraction network uses a deep convolutional structure to refine local detail features and overall contour features layer by layer. The generated feature vectors are then input into a matching module. The matching module performs similarity calculations based on the feature templates registered in the passerby database. It uses a feature distance metric algorithm to compare the similarity between the input features and each template feature, and selects the most similar template label as the identity label information of the current passerby based on the similarity ranking results. During the feature matching process, the confidence value is recorded simultaneously. When the confidence value is lower than a set threshold, the corresponding passerby is marked as unauthorized. When the confidence value is higher than the set threshold, the passerby is marked as authorized, thus outputting complete identity label information data.
[0110] S104: Match the identity tag information with the gate access authorization information to obtain the matching result.
[0111] In this embodiment, the gate access authorization information refers to the authorized personnel identity dataset stored by the gate control terminal, which includes personnel number, image feature identifier and access permission status. The matching result refers to the corresponding relationship status output after comparing the identity tag information with the gate access authorization information, which is used to characterize whether the current personnel have the qualification to pass.
[0112] Specifically, the identity tag information is input into the matching process and compared item by item with the personnel records in the gate access authorization information. First, the personnel number field corresponding to the identity tag information is retrieved in the authorization information table. Candidate records are located by indexing and the associated image feature identifier and permission status parameters are read. After reading, the identity consistency verification step is performed. The identity tag information is compared with the number and feature identifier fields of the authorization record by string matching and hash verification. When all field values are consistent, the identity matching is determined to be successful. When any field value is inconsistent, the identity matching is determined to be unsuccessful. A matching result is generated and a matching status flag is attached. The matching status flag is in Boolean form to indicate whether the match is valid. After the matching result is output, it is passed to the subsequent identity classification and tailgating determination steps.
[0113] S105: When the matching results are consistent, the corresponding pedestrian candidate area is determined as the target person; when the matching results are inconsistent, the corresponding pedestrian candidate area is determined as the tailing person, and the coordinate information of the tailing person in the image is extracted.
[0114] In this embodiment, the target person refers to the person who is identified as an authorized access subject after the identity tag information and the gate access authorization information are successfully matched. The trailing person refers to the unauthorized person who appears in the gate access area but fails to match. The image coordinate information refers to the set of two-dimensional pixel position coordinates corresponding to the trailing person in the access image information.
[0115] Specifically, after receiving the matching result, logical judgment is performed according to the matching status flag. When the matching status flag is true, the pedestrian candidate area corresponding to the target person area is marked, and the identification number and detection box information of the target person are registered in the data buffer. When the matching status flag is false, the corresponding pedestrian candidate area is marked as the tailgating person area, and the coordinate extraction process is entered. In the coordinate extraction process, the center point coordinate position is calculated according to the detection box boundary parameters of the tailgating person area. The center point pixel coordinate is calculated by reading the pixel index values of the upper left and lower right corners of the detection box, and the center point coordinate is recorded as the image coordinate information of the tailgating person. After the marking is completed, all tailgating person areas are numbered and managed so that the image coordinate information of the tailgating person can be called in the subsequent spatial position calculation and sound beam pointing angle determination steps to realize the continuous tracking processing of the tailgating person.
[0116] In one embodiment, such as Figure 3 As shown, in step S20, the image coordinate information of the following person is converted into three-dimensional spatial coordinate information based on the installation pose parameters, and the azimuth and pitch angles of the following person are calculated, including:
[0117] S201: Based on the installation pose parameters, obtain the camera's installation height, horizontal orientation angle, and pitch installation angle, and acquire the installation offset parameters of the directional speaker relative to the camera.
[0118] In this embodiment, the camera's installation height refers to the vertical distance from the camera's optical center to the ground, the horizontal orientation angle refers to the horizontal rotation angle of the camera's optical axis relative to the gate's passage direction, the pitch installation angle refers to the pitch tilt angle of the camera's optical axis relative to the horizontal plane, and the installation offset parameter of the directional speaker relative to the camera refers to the horizontal and vertical distance offset between the two in space.
[0119] Specifically, after fixing the camera above the turnstile channel, the vertical distance from the optical center of the camera to the ground is read by a ranging device and recorded as the installation height of the camera. The horizontal orientation angle of the camera is calibrated using an angle measuring instrument. The optical axis centerline is kept at a predetermined deflection angle with the turnstile passage direction by adjusting the rotation angle of the camera bracket. Then, the tilt angle of the camera optical axis relative to the horizontal plane is measured by an electronic level and recorded as the pitch installation angle. After the camera parameters are determined, the horizontal and vertical distances between the sound center point of the directional speaker and the optical center point of the camera are measured by a laser ranging device. The horizontal and vertical distances are recorded as the installation offset parameters of the directional speaker relative to the camera.
[0120] S202: Establish a spatial coordinate system corresponding to the gate channel, map the image coordinate information of the following person to the spatial coordinate system, and calculate the initial position coordinate information of the following person in the spatial coordinate system.
[0121] In this embodiment, the spatial coordinate system refers to a three-dimensional reference coordinate frame established based on the physical structure of the gate channel, which is used to describe the spatial positional relationship between the camera, the directional speaker and the following person. The initial position coordinate information refers to the three-dimensional position data of the following person when they are first located in the spatial coordinate system.
[0122] Specifically, in the gate channel modeling stage, a spatial coordinate system is established based on the gate passage direction, ground plane, and camera installation position. The main axis direction of the spatial coordinate system is determined with the gate channel centerline as the reference direction. The vertical direction is defined as the height axis, the passage direction is defined as the front-to-back axis, and the channel lateral direction is defined as the left-to-right axis. After the coordinate system is established, the image coordinate information of the following person is input into the coordinate transformation module. The coordinate transformation module performs a mapping operation according to the camera installation height, horizontal orientation angle, and pitch installation angle parameters. The projection position of the following person in the spatial coordinate system is determined by the relative distribution of pixel positions in the image plane. The projection position is geometrically corrected by combining the installation offset parameters of the camera and directional speaker to obtain the initial position coordinate information of the following person in the spatial coordinate system.
[0123] S203: Calculate the depth and distance information of the following person in the vertical direction based on the camera's installation height and pitch angle, combined with the initial position coordinate information.
[0124] In this embodiment, depth distance information refers to the distance data of the following person in the spatial coordinate system relative to the camera mounting reference point in the vertical direction, which is used to characterize the spatial displacement of the following person in the height direction.
[0125] Specifically, after obtaining the camera's installation height and pitch angle, the initial position coordinate information is used as input parameters to establish a vertical correspondence between the camera's optical center position and the spatial projection position of the following person. The vertical depth distance information is determined by calculating the relative displacement of the following person in the height axis direction in the spatial coordinate system. During the calculation process, the angle between the optical axis direction and the vertical line of the ground is adjusted according to the camera's pitch angle to ensure that the vertical projection conforms to the geometric reference range of the gate channel. By using the camera's installation height as the initial height reference and the spatial projection point of the following person as the measurement endpoint, the vertical depth distance of the following person is calculated, and the depth distance information is obtained.
[0126] S204: Based on the horizontal orientation angle and installation offset parameters of the camera, perform spatial geometric transformation on the depth distance information to obtain the three-dimensional position coordinates of the following person in the spatial coordinate system.
[0127] In this embodiment, spatial geometric transformation refers to the process of rotating and translating the spatial position of the following person using the horizontal orientation angle and installation offset parameters of the camera. The three-dimensional position coordinates refer to the three-dimensional coordinate point data of the following person in the spatial coordinate system after geometric transformation, which are used to describe the spatial positioning result of the following person in the gate channel.
[0128] Specifically, after obtaining the horizontal orientation angle and installation offset parameters of the camera, the depth distance information is input into the coordinate transformation process. First, the coordinate rotation direction is determined based on the horizontal orientation angle of the camera. The camera's viewpoint is aligned with the main direction of the gate channel by rotating around the vertical axis of the spatial coordinate system. After the rotation is completed, the coordinate translation operation is performed according to the horizontal and vertical offsets in the installation offset parameters. The spatial difference between the camera and the directional speaker is corrected by adjusting the position of the spatial coordinate origin, so that the depth distance information in the translated coordinate system remains consistent with the actual physical distance. After the rotation and translation operations are completed, the corrected spatial coordinate data is integrated into the three-dimensional position coordinates of the following person in the spatial coordinate system.
[0129] S205: Based on the spatial relationship between the three-dimensional position coordinates and the directional loudspeaker, calculate the azimuth and pitch angles of the following person relative to the directional loudspeaker.
[0130] Specifically, based on the spatial relationship between the three-dimensional position coordinates and the directional loudspeaker, the reference point position of the directional loudspeaker in the spatial coordinate system is first determined. By extracting the installation height, horizontal offset, and pitch angle of the directional loudspeaker, a spatial reference vector of the directional loudspeaker is established. The difference between the three-dimensional position coordinates and the spatial reference vector of the directional loudspeaker is calculated to obtain the spatial position difference vector of the following person relative to the directional loudspeaker. Based on the changes in the components of the spatial position difference vector in the horizontal and vertical planes, the azimuth and pitch angles of the following person relative to the directional loudspeaker are calculated. The azimuth angle is obtained by analyzing the change in the angle between the difference vector and the main sound emission direction of the loudspeaker in the horizontal plane, and the pitch angle is obtained by analyzing the change in the angle between the difference vector and the horizontal plane in the vertical direction.
[0131] In one embodiment, such as Figure 4 As shown, in step S30, the target sound beam pointing angle is compared with a preset angle threshold. When the angle separation corresponding to the target sound beam pointing angle is greater than the preset angle threshold and the distance between the tailgating person and the gate threshold line is less than the preset distance threshold, a directional voice intervention trigger command is generated, and a target stability window is established for the tailgating person, including:
[0132] S301: Calculate the difference between the target beam pointing angle and the preset angle threshold to obtain the angle separation amount.
[0133] Specifically, based on the target beam pointing angle and the preset angle threshold, the angle information of the current target beam pointing angle is first read from the beam control module, and the preset angle threshold parameter is called from the parameter configuration. The angle separation amount is obtained by performing a difference calculation operation on the two. The difference calculation operation includes comparing the target beam pointing angle as an input variable with the preset angle threshold step by step, and determining the direction and magnitude of the deviation based on the comparison result. When the target beam pointing angle is greater than the preset angle threshold, the calculation result is a positive offset, and when the target beam pointing angle is less than the preset angle threshold, the calculation result is a negative offset. After the difference calculation is completed, the angle separation amount is processed into an absolute value to eliminate the directional influence, so that the angle separation amount can accurately reflect the magnitude of the target beam pointing angle deviating from the preset angle threshold.
[0134] S302: Detect the coordinate changes of the following person in consecutive image frames, and calculate the real-time distance information and movement direction angle of the following person.
[0135] Specifically, based on the image coordinate information of the trailing person extracted from consecutive image frames, each frame is first numbered according to time sequence and a coordinate change record table is established. By calculating the horizontal and vertical differences of the trailing person's image coordinates in adjacent frames, the displacement vector of the trailing person on the image plane is obtained. Combining the camera's installation pose parameters with the spatial coordinate system mapping relationship, the displacement vector is mapped to the spatial coordinate difference to obtain the displacement distance of the trailing person in the spatial coordinate system. The real-time distance information of the trailing person is calculated based on the displacement distance and the time interval between image frames, thereby reflecting the instantaneous approach degree of the trailing person relative to the gate threshold line. Furthermore, by analyzing the component changes of the displacement vector in the horizontal axis direction within adjacent time periods, the movement direction angle of the trailing person is determined. When the movement direction angle is close to the gate's passage direction angle, it indicates that the trailing behavior is more obvious. After the calculation is completed, the real-time distance information and the movement direction angle are used as input data for subsequent directional voice intervention trigger condition judgment to support the accurate execution of the trigger logic.
[0136] S303: When the angle separation is greater than the preset angle threshold and the real-time distance information is less than the preset distance threshold, and the movement direction angle tends to the gate passage direction, a directional voice intervention trigger command is generated.
[0137] Specifically, based on the angle separation, real-time distance information, and movement direction angle, three sets of values are input into the condition judgment process. In the condition judgment process, the angle separation is first compared with a preset angle threshold. When the angle separation is greater than the preset angle threshold, it is determined that there is a significant deviation between the spatial position of the following person and the main emission direction of the sound beam. Then, the real-time distance information is compared with a preset distance threshold. When the real-time distance information is less than the preset distance threshold, it is determined that the following person has approached the gate threshold area. Further, the movement trend of the following person is judged based on the angle relationship between the movement direction angle and the gate passage direction. When the movement direction angle tends to the gate passage direction, the following behavior is considered to be true. When all three conditions are determined to be true, the calculation logic generates a directional voice intervention trigger command based on the condition matching result, and stores the trigger command in association with the spatial coordinate information of the following person for subsequent beam direction control and target stabilization window establishment.
[0138] S304: After generating the directional voice intervention trigger command, calculate the trajectory change rate of the following person based on the coordinate sequence of continuous frame images to obtain the spatial position stability parameter, and establish the target stability window based on the spatial position stability parameter.
[0139] In this embodiment, the trajectory change rate refers to the rate of change of the position of the motion trajectory of the following person in the continuous frame images over time. The spatial position stability parameter is a quantitative index calculated based on the trajectory change rate to measure the spatial position stability of the following person. The target stability window is a continuous time interval established under the condition that the spatial position of the following person is stable.
[0140] Specifically, after generating the directional voice intervention trigger command, the coordinate sequence in the continuous frame image is called, and the spatial position data of the following person in the time series is time-synchronized and serialized. The trajectory change rate is obtained by calculating the displacement difference of the coordinate points in adjacent frames and dividing it by the inter-frame time interval. After obtaining the trajectory change rate, the change rate sequence is smoothed and filtered to remove instantaneous jitter. The mean and variance of the change rate are calculated. The spatial position stability parameter is determined based on the degree of fluctuation reflected by the variance of the change rate. When the spatial position stability parameter is lower than the preset stability threshold, it is determined that the spatial position of the following person remains stable for a short period of time. A target stability window is established based on the duration and fluctuation range of the stability parameter. The duration of the target stability window corresponds to the duration of the following person's stable state. After the establishment is completed, the target stability window is used as the input interval for the adaptive control of the directional speaker beam direction for subsequent dynamic correction of the beam pointing angle.
[0141] In one embodiment, such as Figure 5 As shown, in step S40, based on the spatial coordinate sequence within the target stabilization window, adaptive beam direction control of the directional loudspeaker is performed, and the current beam pointing angle is obtained by correcting the real-time spatial position of the following person, including:
[0142] S401: Perform time series smoothing on the spatial coordinate sequence within the target stable window to generate continuous spatial coordinate information.
[0143] In this embodiment, continuous spatial coordinate information refers to smoothed spatial position data formed after time series smoothing, which is used to reflect the continuous displacement change trend of the following person within a stable window.
[0144] Specifically, after acquiring the spatial coordinate sequence within the target stabilization window, the spatial coordinate points in the time series are arranged according to the sampling time. A smoothing operation is then performed on the arranged spatial coordinate sequence. The displacement changes of the spatial coordinate points are weighted and averaged using a moving average method or a weighted sliding filter method to reduce coordinate jumps caused by camera sampling jitter or instantaneous motion. During the smoothing process, the width of the smoothing window is determined based on the time length of the target stabilization window and the sampling frame rate to ensure the continuity of coordinate changes in the time dimension. After the smoothing process is completed, interpolation is performed on the smoothing result to fill in the missing coordinate points within the time interval, forming continuous spatial coordinate information with uniform temporal distribution and continuous spatial changes.
[0145] S402: Calculate the azimuth and distance changes of adjacent coordinate points based on continuous spatial coordinate information to obtain spatial displacement trend information.
[0146] In this embodiment, the orientation change refers to the change in the horizontal angle of the following person relative to the main sound-emitting direction of the directional loudspeaker between adjacent spatial coordinate points, the distance change refers to the difference in distance between the following person and adjacent spatial coordinate points along the axis of the spatial coordinate system, and the spatial displacement trend information refers to the movement direction and displacement trend of the following person within the target stability window, which is characterized by the orientation change and the distance change.
[0147] Specifically, based on continuous spatial coordinate information, spatial coordinate points are paired in chronological order, and two adjacent sets of spatial coordinate points are selected as the calculation objects. First, the spatial distance difference between the two points is calculated. The distance change is determined by comparing the numerical changes of adjacent coordinate points on the horizontal axis and depth axis. After obtaining the distance change, the relative angle change of the two points on the spatial plane is further calculated. The azimuth change is determined by analyzing the offset angle of the following person's position on the horizontal plane relative to the main sound direction of the directional loudspeaker. After the azimuth change and distance change are calculated, they are combined to form spatial displacement trend information, which is used to describe the direction of change of the following person's trajectory and the continuity of its displacement within the target stability window. The spatial displacement trend information is stored in the data buffer after generation, providing input basis for the subsequent calculation of beam pointing angle correction.
[0148] S403: Calculate the beam pointing angle correction based on the spatial displacement trend information and the preset target beam pointing angle.
[0149] In this embodiment, the preset target beam direction angle refers to the ideal beam direction parameter used as a control reference in the directional loudspeaker sound field, and the beam pointing angle correction amount refers to the angle correction value calculated based on the difference between the spatial displacement trend information and the preset target beam direction angle.
[0150] Specifically, based on the spatial displacement trend information and the preset target beam direction angle, the azimuth change in the spatial displacement trend information is first extracted as the angle input variable. This variable is then numerically compared with the preset target beam direction angle. The beam direction offset is determined by calculating the angle difference between the two. After determining the offset, the weighting coefficient of the angle difference is corrected by combining the distance change in the spatial displacement trend information. This ensures that the angle correction process reflects the impact of the distance change of the following personnel on the beam pointing accuracy. Subsequently, the beam pointing angle correction is calculated based on the angle offset and the weighting coefficient. During the calculation, a weighted difference method is used to smooth the angle correction result to avoid abrupt changes that could cause beam instability. After the calculation is completed, the beam pointing angle correction is recorded in the beam control parameter set as the input parameter for the next stage of beam direction adaptive control.
[0151] S404: Perform beam direction adaptive control on the beam pointing angle correction, perform phase compensation and amplitude correction, and obtain the updated current beam pointing angle.
[0152] In this embodiment, the updated current beam pointing angle refers to the new main beam emission direction angle information formed after phase compensation and amplitude correction.
[0153] Specifically, after obtaining the beam pointing angle correction, the beam pointing angle correction is input into the beam direction adaptive control process. In this process, the adjustment of the sound unit parameters is initiated according to the control step size. By calculating the relative phase difference between the sound units and adjusting the phase angle of each sound unit based on the phase difference change trend, dynamic compensation of the sound beam radiation direction is achieved. After the phase adjustment is completed, the amplitude correction coefficient is calculated based on the amplitude information of the beam pointing angle correction, and the output power of the sound unit is proportionally adjusted so that the amplitude distribution of the sound unit output is consistent with the updated beam pointing requirements. After the phase compensation and amplitude correction are completed, the phase and amplitude state parameters of each sound unit are summarized, the new sound beam direction vector is calculated, and the updated current beam pointing angle is obtained. The updated current beam pointing angle is recorded in the beam control parameter set to maintain the continuity and traceability of beam direction adjustment.
[0154] In one embodiment, such as Figure 6 As shown, in step S404, the beam pointing angle correction is performed using beam direction adaptive control, including phase compensation and amplitude correction, to obtain the updated current beam pointing angle, including:
[0155] S4041: Receives the beam pointing angle correction as a control input and initiates iterative calculations for beam direction adaptive control.
[0156] Specifically, after receiving the beam pointing angle correction during the beam direction control phase, the beam pointing angle correction is input into the initial calculation unit of the beam direction adaptive control process. The iterative calculation process is started by setting the number of iterations and step size parameters. In each iteration, the deviation between the current beam direction and the target beam direction is calculated. The phase and amplitude input values of the sound unit are adjusted according to the deviation, and the adjusted results are fed back to the next iteration, forming an adaptive adjustment loop that gradually approaches the preset direction. During the iteration process, the convergence trend of the beam direction error is monitored in real time. When the beam direction error remains within the preset stable range in multiple consecutive iterations, it is confirmed that the beam direction has reached the control convergence state, thus completing the iterative calculation process of beam direction adaptive control.
[0157] S4042: During the iterative calculation, the phase difference between adjacent sound units is calculated based on the beam pointing angle correction, and the phase of each sound unit is adjusted sequentially according to the preset phase update step size to obtain the adjusted phase difference.
[0158] In this embodiment, the adjusted phase difference value refers to the phase difference result between adjacent sound generating units recalculated after performing phase update.
[0159] Specifically, during the iterative calculation, the beam pointing angle correction is first input into the phase calculation process. Based on the spatial distribution order of the sound-emitting units in the array, the spatial spacing information of adjacent sound-emitting units is extracted one by one. By comparing the changing trend of the product of the beam pointing angle correction and the spatial spacing, the initial phase difference between adjacent sound-emitting units is calculated. After obtaining the initial phase difference, the preset phase update step size parameter is called, and the phase adjustment operation is performed sequentially according to the arrangement order of the sound-emitting units. By correcting the current phase difference by increasing or decreasing the step size, the sound beam radiation direction gradually approaches the target beam direction. After all sound-emitting units have completed phase adjustment, the phase difference of adjacent units is calculated again to obtain the adjusted phase difference. The adjusted phase difference is then used as the input parameter for the next stage of amplitude correction and beam direction optimization calculation.
[0160] S4043: Calculate the amplitude correction amount based on the adjusted phase difference value, and proportionally correct the output amplitude of the sound unit according to the amplitude correction amount to form an amplitude distribution result corresponding to the adjusted phase difference value.
[0161] Specifically, after acquiring the adjusted phase difference value, the phase difference value information corresponding to adjacent sound units is read sequentially according to the spatial arrangement order of the sound units in the sound beam array. The phase difference value is used as the basis for amplitude adjustment. The direction of amplitude increase and amplitude decrease are distinguished according to the trend of the phase difference value change in the array direction, and the corresponding amplitude correction amount is determined based on the change amplitude of the phase difference value. After determining the amplitude correction amount, the output amplitude of each sound unit is proportionally scaled according to the amplitude correction amount, so that the output amplitude adjustment amount corresponding to the sound unit with a larger phase difference change is increased accordingly, and the output amplitude adjustment amount corresponding to the sound unit with a smaller phase difference change is decreased accordingly. After completing the proportional correction of the output amplitude of all sound units, the corrected output amplitude of each sound unit is uniformly collected and sequentially reorganized according to the arrangement relationship of the sound beam array to form an amplitude distribution result that corresponds one-to-one with the adjusted phase difference value.
[0162] S4044: Adjust the control weights of the beam direction adaptive control based on the corrected amplitude distribution results and the changing trend of the beam pointing angle correction, and obtain the updated control weight distribution results.
[0163] Specifically, based on the changes in the corrected amplitude distribution and the beam pointing angle correction, two sets of data are input into the control weight adjustment process. By comparing the energy concentration of the amplitude distribution in the sound unit array with the rate of change of the beam pointing angle correction over time, the dynamic adjustment direction of the control weight is determined. During the adjustment process, the weight gain is allocated according to the unevenness of the amplitude distribution. The control weight is reduced for sound units in the energy concentration area and increased for sound units in the energy attenuation area, so that the overall beam direction control process maintains a balance in energy distribution. After the control weight correction of all sound units is completed, the correction results are integrated into a control weight matrix according to the array order to generate the updated control weight distribution result.
[0164] S4045: Calculate the updated current beam pointing angle based on the adjusted phase difference, the corrected amplitude distribution, and the updated control weight distribution.
[0165] Specifically, in the beam direction adaptive control stage, the adjusted phase difference, the corrected amplitude distribution, and the updated control weight distribution are input into the beam pointing angle comprehensive calculation process. The beam direction vector refers to the main radiation direction vector of the sound beam formed by superimposing the phase and amplitude of the sound unit array in space. The angle calculation method refers to the mathematical transformation steps of decomposing the beam direction vector in the spatial coordinate system and calculating its corresponding angle value. First, the phase distribution structure of the sound unit array is determined based on the adjusted phase difference. By calculating the components of the phase difference of each sound unit on the array coordinate axis, a phase distribution matrix is constructed. Then, the energy output ratio of each sound unit is determined based on the corrected amplitude distribution. The energy weight coefficient is formed by performing normalization processing on the amplitude data. Then, combined with the updated control weight distribution, a weighted superposition operation is performed on the phase distribution matrix and the energy weight coefficient to obtain the comprehensive sound field response distribution. After the sound field response distribution is formed, the main radiation direction of the sound beam is calculated using the beam direction vector. The angle value corresponding to this direction is extracted by the angle calculation method to generate the updated current beam pointing angle.
[0166] In one embodiment, such as Figure 7 As shown, in step S50, the azimuth error is determined based on the difference between the beam pointing angle and the preset target beam direction angle. The main lobe direction of the sound beam is limited to a preset main lobe width range, and sidelobe energy overflow is suppressed through phase smoothing compensation and beam calibration strategies. This includes:
[0167] S501: Compare the azimuth error with the preset angle error threshold. When the azimuth error is less than the preset angle error threshold, perform the main lobe direction limitation operation.
[0168] Specifically, in the beam control stage, the azimuth error is compared with a preset angle error threshold. The azimuth error refers to the angle difference between the current beam pointing angle and the target beam direction angle. The preset angle error threshold is an angle limitation parameter used to determine whether the beam direction deviation is within the allowable range. The main lobe direction limitation operation is a control process that limits and fixes the main radiation direction of the beam when the angle error condition is met. When performing the comparison, the angle difference between the current beam pointing angle and the target beam direction angle is read first and the azimuth error is calculated. Then, the azimuth error is compared with the preset angle error threshold. When the azimuth error is less than the preset angle error threshold, the main lobe direction limitation stage is entered. In the main lobe direction limitation stage, the phase synchronization state of the sound-emitting units in the array is corrected according to the current beam pointing angle to make the phase center of the sound-emitting unit consistent with the main radiation direction of the beam and to lock the energy output range of the main lobe region. By limiting the phase offset of adjacent sound-emitting units, the main lobe radiation direction is ensured to be stable, thus completing the main lobe direction limitation operation.
[0169] S502: In the main lobe direction limiting operation, the sound beam radiation direction is adjusted according to the preset main lobe width range to limit the main lobe direction of the sound beam.
[0170] Specifically, during the main lobe direction limiting operation, a preset main lobe width range is input into the beam direction adjustment process. The preset main lobe width range refers to the angular range within which the energy distribution on both sides of the main radiation direction of the sound beam is allowed to deviate. The beam radiation direction refers to the direction of sound energy propagation formed by the directional loudspeaker array in space. The main lobe direction refers to the direction of the position where the sound beam radiation energy is concentrated and the amplitude is the largest. Based on the preset main lobe width range, the center angle of the main lobe is first determined. The upper and lower limit angle values of the main lobe range are calculated by analyzing the distribution position of the current beam pointing angle in the spatial coordinate system. After obtaining the upper and lower limit angle values, the boundary of the beam radiation direction is adjusted so that the center angle of the main lobe remains within the preset main lobe width range. Subsequently, the phase distribution of the sound unit array is fine-tuned. By controlling the change in phase difference between adjacent sound units, the radiation direction of the main lobe energy concentration area is made consistent with the center angle of the main lobe. Finally, the main lobe direction of the sound beam is limited to be stable within the preset main lobe width range, and the limited main lobe direction parameters are written into the beam control parameter set to maintain the consistency of the sound beam direction.
[0171] S503: After completing the main lobe direction limitation, perform phase smoothing compensation to correct the phase change of the beam direction. After the phase smoothing compensation is completed, perform beam calibration strategy to adjust the beam amplitude distribution and suppress sidelobe energy overflow.
[0172] Specifically, after defining the main lobe direction, phase smoothing compensation and beam calibration strategies are executed sequentially. Phase smoothing compensation refers to the process of reducing abrupt changes in beam direction phase by adjusting the phase change rate of the sound unit array. Beam calibration strategy refers to the calculation steps in beam control to optimize the distribution of sound beam radiation energy through amplitude redistribution. Side lobe energy overflow refers to the non-target energy radiation phenomenon generated by the sound beam in the region outside the main lobe. When performing phase smoothing compensation, the phase gradient difference between adjacent sound units is calculated based on the phase distribution state after defining the main lobe direction. The phase change rate is adjusted by interpolation smoothing to keep the phase difference distribution curve continuous. After phase smoothing compensation is completed, the beam calibration strategy stage is entered. Amplitude distribution data is extracted based on the corrected phase distribution, and regions with uneven amplitude in the sound unit array are detected. The amplitude distribution is adjusted by amplitude normalization and energy constraint methods to concentrate the sound beam energy in the main lobe direction and reduce the energy amplitude in the side lobe region, thereby suppressing side lobe energy overflow. After adjustment, the updated phase and amplitude parameters are synchronized to the beam control parameter set to maintain beam direction stability and continuous sound field distribution.
[0173] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0174] In one embodiment, a gate tailgating detection directional voice intervention device is provided, which corresponds one-to-one with the gate tailgating detection directional voice intervention method described in the above embodiments. For example... Figure 8 As shown, the gate tailgating detection directional voice intervention device includes an image recognition module, a pose calculation module, an intervention trigger module, a beam control module, and a voice output module.
[0175] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A gate tail detection directed voice intervention method, characterized by, A method for directional voice intervention to detect tailgating at turnstiles includes: Acquire passage image information, perform human detection and identity recognition on the passage image information, distinguish between legally passing target personnel and tailgating personnel, and output the image coordinate information of tailgating personnel in real time; The installation posture parameters of the camera and directional speaker are obtained. Based on the installation posture parameters, the image coordinate information of the following person is converted into spatial three-dimensional coordinate information. The azimuth and pitch angles of the following person are calculated, and the target sound beam pointing angle of the following person in the directional sound field is determined based on the azimuth and pitch angles. The target sound beam pointing angle is compared with a preset angle threshold. When the angle separation corresponding to the target sound beam pointing angle is greater than the preset angle threshold and the distance between the tailing person and the gate threshold line is less than the preset distance threshold, a directional voice intervention trigger command is generated, and a target stability window is established for the tailing person. Based on the spatial coordinate sequence within the target stabilization window, the beam direction adaptive control of the directional loudspeaker is executed, and the current beam pointing angle is obtained by correcting the real-time spatial position of the following personnel. The azimuth error is determined by the difference between the beam pointing angle and the preset target beam direction angle. The direction of the main lobe of the sound beam is limited to the preset main lobe width range. The side lobe energy overflow is suppressed by phase smoothing compensation and beam calibration strategy. Under the condition that the azimuth error is less than the preset angle error threshold, the preset voice prompt information is issued to the head area of the following person.
2. The method for directional voice intervention in gate tailgating detection according to claim 1, characterized in that, The process of acquiring passage image information, performing human detection and identity recognition on the passage image information, distinguishing between legally passing target personnel and tailgating personnel, and outputting the image coordinate information of tailgating personnel in real time includes: Perform target region extraction on the traffic image information to determine the candidate pedestrian area located in the gate channel; Human feature detection is performed on the candidate pedestrian area to obtain the human feature detection results; Based on the results of human feature detection and the preset identity recognition model, the identity tag information of the person passing through is determined; The identity tag information is matched with the gate access authorization information to obtain the matching result; When the matching results are consistent, the corresponding pedestrian candidate area is determined as the target person. When the matching results are inconsistent, the corresponding pedestrian candidate area is determined as the tailing person, and the coordinate information of the tailing person in the image is extracted.
3. The method for directional voice intervention in gate tailgating detection according to claim 1, characterized in that, The step of converting the image coordinate information of the following personnel into three-dimensional spatial coordinate information based on the installation posture parameters, and calculating the azimuth and pitch angles of the following personnel, includes: Based on the installation pose parameters, the installation height, horizontal orientation angle and pitch installation angle of the camera are obtained, and the installation offset parameters of the directional speaker relative to the camera are acquired. Establish a spatial coordinate system corresponding to the gate channel, map the image coordinate information of the tailing person to the spatial coordinate system, and calculate the initial position coordinate information of the tailing person in the spatial coordinate system. Based on the camera's installation height and pitch angle, and combined with the initial position coordinate information, the vertical depth distance information of the following person is calculated. Based on the horizontal orientation angle and installation offset parameters of the camera, a spatial geometric transformation is performed on the depth distance information to obtain the three-dimensional position coordinates of the following person in the spatial coordinate system. Based on the spatial relationship between the three-dimensional position coordinates and the directional loudspeaker, the azimuth and pitch angles of the following person relative to the directional loudspeaker are calculated.
4. The method for directional voice intervention in gate tailgating detection according to claim 1, characterized in that, The step involves comparing the target sound beam pointing angle with a preset angle threshold. When the angle separation corresponding to the target sound beam pointing angle is greater than the preset angle threshold and the distance between the tailgating person and the gate threshold line is less than a preset distance threshold, a directional voice intervention trigger command is generated, and a target stabilization window is established for the tailgating person, including: The difference between the target beam pointing angle and the preset angle threshold is calculated to obtain the angle separation amount; Detect the coordinate changes of the following person in consecutive image frames, and calculate the real-time distance and movement direction angle of the following person; When the angle separation is greater than the preset angle threshold, the real-time distance information is less than the preset distance threshold, and the movement direction angle tends to the gate passage direction, a directional voice intervention trigger command is generated. After generating the directional voice intervention trigger command, the trajectory change rate of the following person is calculated based on the coordinate sequence of continuous frame images to obtain the spatial position stability parameter, and a target stability window is established based on the spatial position stability parameter.
5. The method for directional voice intervention in gate tailgating detection according to claim 1, characterized in that, The step of performing adaptive beam direction control of the directional loudspeaker based on the spatial coordinate sequence within the target stabilization window, and correcting and obtaining the current beam pointing angle based on the real-time spatial position of the following person, includes: Time series smoothing is performed on the spatial coordinate sequence within the target stable window to generate continuous spatial coordinate information; Based on continuous spatial coordinate information, the changes in orientation and distance between adjacent coordinate points are calculated to obtain spatial displacement trend information. The beam pointing angle correction is calculated based on the spatial displacement trend information and the preset target beam pointing angle. The beam pointing angle correction is applied to beam direction adaptive control, and phase compensation and amplitude correction are performed to obtain the updated current beam pointing angle.
6. The method for directional voice intervention in gate tailgating detection according to claim 5, characterized in that, The step of performing beam pointing angle correction on beam direction adaptive control, performing phase compensation and amplitude correction, to obtain the updated current beam pointing angle includes: The beam pointing angle correction is received as a control input to initiate iterative calculations for beam direction adaptive control. During the iterative calculation, the phase difference between adjacent sound units is calculated based on the beam pointing angle correction. The phase of each sound unit is adjusted sequentially according to the preset phase update step size to obtain the adjusted phase difference. The amplitude correction amount is calculated based on the adjusted phase difference value, and the output amplitude of the sound unit is proportionally corrected according to the amplitude correction amount to form an amplitude distribution result corresponding to the adjusted phase difference value. Based on the corrected amplitude distribution results and the changing trend of the beam pointing angle correction, the control weights of the beam direction adaptive control are adjusted to obtain the updated control weight distribution results. The updated current beam pointing angle is calculated based on the adjusted phase difference, the corrected amplitude distribution, and the updated control weight distribution.
7. The method for directional voice intervention in gate tailgating detection according to claim 1, characterized in that, The step of determining the azimuth error based on the difference between the beam pointing angle and the preset target beam direction angle, limiting the main lobe direction of the sound beam within a preset main lobe width range, and suppressing sidelobe energy overflow through phase smoothing compensation and beam calibration strategies includes: The azimuth error is compared with a preset angle error threshold. When the azimuth error is less than the preset angle error threshold, the main lobe direction limitation operation is performed. In the main lobe direction limiting operation, the sound beam radiation direction is adjusted according to the preset main lobe width range to limit the main lobe direction of the sound beam. After the main lobe direction is defined, phase smoothing compensation is performed to correct the phase change of the beam direction. After the phase smoothing compensation is completed, a beam calibration strategy is performed to adjust the amplitude distribution of the beam and suppress sidelobe energy overflow.
8. A directional voice intervention device for detecting tailgating at turnstiles, characterized in that, The gate tailgating detection and directional voice intervention device includes: The image recognition module is used to acquire passage image information, perform human detection and identity recognition on the passage image information, distinguish between legally passing target personnel and tailgating personnel, and output the image coordinate information of tailgating personnel in real time. The pose calculation module is used to obtain the installation pose parameters of the camera and the directional speaker. Based on the installation pose parameters, the image coordinate information of the following person is converted into spatial three-dimensional coordinate information. The azimuth and pitch angles of the following person are calculated, and the target sound beam pointing angle of the following person in the directional sound field is determined based on the azimuth and pitch angles. The intervention trigger module is used to compare the target sound beam pointing angle with a preset angle threshold. When the angle separation corresponding to the target sound beam pointing angle is greater than the preset angle threshold and the distance between the tailing person and the gate threshold line is less than the preset distance threshold, a directional voice intervention trigger command is generated, and a target stability window is established for the tailing person. The beam control module is used to perform adaptive beam direction control of the directional loudspeaker based on the spatial coordinate sequence within the target stabilization window, and to correct and obtain the current beam pointing angle based on the real-time spatial position of the following person. The voice output module is used to determine the azimuth error based on the difference between the beam pointing angle and the preset target beam direction angle, limit the main lobe direction of the sound beam within the preset main lobe width range, and suppress side lobe energy overflow through phase smoothing compensation and beam calibration strategies. Under the condition that the azimuth error is less than the preset angle error threshold, the module sends a preset voice prompt message to the head area of the following person.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the gate tailgating detection directional voice intervention method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the gate tail-following detection directional voice intervention method as described in any one of claims 1 to 7.