A multimodal registration and target localization method and system

By employing a three-modal collaborative mechanism combining voiceprint signals, image data, and temperature data, the problems of modal incompleteness and rigid computing power allocation in target localization under complex environments are solved, achieving high-precision and real-time target localization and improving localization accuracy and system efficiency.

CN122260235APending Publication Date: 2026-06-23重庆中科汽车软件创新中心
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610456168.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-08
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing technologies face challenges in target localization in complex environments, including incomplete modal dimensions, interference removal, and spatiotemporal registration. Furthermore, rigid computing power allocation leads to a trade-off between efficiency and accuracy, making it difficult to achieve high-precision real-time localization.

Method used

A three-modal collaborative mechanism of voiceprint signal, image data and temperature data is adopted. Multimodal data time alignment is achieved by timestamp marking. Data fusion and registration are performed by combining perspective projection transformation model. Target localization is achieved in stages. Temperature data is used to verify the effectiveness of the target and eliminate interference signals.

Benefits of technology

It achieves high-precision (error ≤ 0.05m) and real-time (response time ≤ 0.5s) target positioning in complex environments, reducing system computation and cost, and improving positioning accuracy (up to 92%).

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122260235A_ABST
    Figure CN122260235A_ABST
Patent Text Reader

Abstract

This invention relates to the field of industrial automation technology, specifically to a multimodal registration and target localization method and system. The system includes a sensing module for acquiring acoustic signature signals, image data, and temperature data; a control module including a pan-tilt controller and a data synchronization unit, wherein the pan-tilt controller receives positioning coordinates and drives the pan-tilt to the target direction, and the data synchronization unit uses timestamps to achieve time alignment of the multimodal data; a fusion processing module for calculating the initial positioning coordinates of the sound source based on the acoustic signature signal, performing visual refinement by combining it with image data, and registering and fusing the acoustic signature coordinates, image coordinates, and temperature data to output the target localization result; and an execution module for displaying the target localization coordinates and temperature information, and triggering an alarm when a valid target is located. This technical solution can achieve high-precision target localization while ensuring real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial automation technology, specifically to a multimodal registration and target localization method and system. Background Technology

[0002] Achieving rapid and accurate location and tracking of targets (such as personnel, equipment malfunction points, and heat source leaks) in complex environments such as industrial workshops, hazardous chemical warehouses, and disaster sites is a key technological requirement for ensuring production safety and improving rescue efficiency.

[0003] In the existing target localization technology system, it is mainly divided into two categories: single-modal localization and multi-modal fusion localization.

[0004] While single-modal localization technology has certain application value in specific scenarios, it also faces significant technical bottlenecks in complex environments. Single voiceprint localization uses a microphone array to calculate the location of a sound source through beamforming or time difference of arrival (TDOA) algorithms. Its advantages include fast response and insensitivity to lighting conditions, but its disadvantages are significant: in high-noise environments such as industrial workshops, non-target interference sources such as wind and mechanical vibrations (i.e., "noise interference") overlap with the effective target sound source. Traditional algorithms lack a mechanism to eliminate interference sources, resulting in a very high misjudgment rate (measured accuracy is only about 65%). Furthermore, the physical characteristics of voiceprint localization determine its low accuracy, with positioning errors typically exceeding 1 meter, failing to meet the needs of precision operations.

[0005] Single-vision localization relies on images captured by a camera, achieving positioning through feature extraction and matching. While it boasts high accuracy under ideal conditions (error approximately 0.1-0.3m), it is highly dependent on lighting conditions and line-of-sight. In industrial scenarios with smoke, dim lighting, target occlusion, or complex backgrounds, the visual algorithm is prone to failure or drift, resulting in poor recognition stability (accuracy approximately 78% in workshop environments). Furthermore, achieving high-precision localization across all scenarios requires continuous full-frame image processing and computationally intensive feature matching, leading to high system latency and energy consumption.

[0006] Single-temperature positioning can only detect areas of abnormal temperature through infrared thermal imaging. It is an "attribute detection" method rather than a "spatial positioning" method and cannot output the specific coordinates of the target. It has no positioning value when used alone.

[0007] To overcome the limitations of single-modality localization, multimodal fusion localization technology has emerged. Existing technologies, such as the prior art document CN120254762A, have proposed basic ideas for multimodal data fusion and spatiotemporal synchronization. However, through research and analysis, it has been found that existing multimodal localization technologies still have the following deep-seated technical shortcomings, especially when dealing with complex industrial scenarios: First, there is an incompleteness in modal dimensions and a lack of verification mechanisms. Existing technologies are mostly limited to binary fusion of "sound + vision" or "light + temperature," which prevents the system from constructing a closed-loop logic of initial location determination, visual refinement, and attribute verification. For example, after locating a sound source or visual feature, it is impossible to perform secondary screening using the physical attribute of temperature. This results in a lack of recognition ability for interference sources without temperature characteristics (such as simple mechanical impact sounds or visual images of cold objects), leading to a large number of false alarms in the localization results.

[0008] Second, there are challenges in interference removal and spatiotemporal registration in complex environments. In highly interference scenarios such as industrial workshops, existing technologies lack specialized algorithms for removing non-target sound signatures such as wind noise, echoes, and mechanical vibrations, resulting in the initial location coordinates of the sound source containing a significant amount of noise. Furthermore, the inherent time differences in data acquisition from different modal sensors, coupled with their independent coordinate systems, lead to spatial misalignment during multimodal data fusion. This not only fails to leverage the complementary advantages of multimodal data but also amplifies the final location deviation due to error aggregation.

[0009] Third, the rigid allocation of computing power leads to a contradiction between efficiency and accuracy. Existing robot or positioning systems often employ fixed configurations for heterogeneous computing power allocation, either by pre-setting computing power schemes based on hardware models or by coarsely scheduling based solely on a single task type. In practical applications, continuously operating high-computing-power visual processing in pursuit of accuracy results in system lag and resource waste; conversely, relying solely on low-computing-power voiceprint triggering fails to meet accuracy requirements. It is difficult to achieve high accuracy while maintaining real-time performance. Summary of the Invention

[0010] The purpose of this invention is to propose a multimodal registration and target localization method and system, which can achieve high-precision target localization while ensuring real-time performance.

[0011] To achieve the above objectives, in a first aspect, the present invention proposes a multimodal registration and target localization method and system, comprising: The sensing module is used to collect voiceprint signals, image data, and temperature data; The control module includes a gimbal controller and a data synchronization unit. The gimbal controller is used to receive positioning coordinates and drive the gimbal to turn in the target direction. The data synchronization unit is used to achieve time alignment of multimodal data through timestamp marking. The fusion processing module is used to calculate the initial location coordinates of the sound source based on the voiceprint signal, perform visual refinement by combining image data, and register and fuse the voiceprint coordinates, image coordinates and temperature data to output the target location result. The execution module is used to display the target location coordinates and temperature information, and to trigger an alarm when a valid target is located.

[0012] The beneficial effects of the basic scheme: In existing technologies, single-vision positioning is easily affected by lighting and occlusion, while single-voiceprint positioning is easily interfered with by irrelevant sound sources such as wind and noise. Dual-modal positioning also suffers from a lack of effective target validity verification criteria, making it difficult to distinguish between interference signals and real target signals, leading to frequent positioning misjudgments in complex environments. This invention innovatively introduces temperature data as a third modality, forming a sound-light-temperature three-modal collaborative mechanism with voiceprint signals and image data. By verifying the validity of the target through temperature data, it can accurately eliminate interference sound sources without temperature anomalies, such as wind and environmental noise, while also eliminating false visual targets without temperature characteristics. This reduces positioning misjudgments from the source, enabling the system to work stably in noisy, variable lighting, and complex interference scenarios (such as workshops and outdoor monitoring).

[0013] Existing positioning systems often suffer from a trade-off between accuracy and efficiency: pursuing high accuracy through full-scene visual computing leads to a surge in computational load, response delays, and high operating costs; while pursuing high efficiency through rapid positioning using a single sensor results in positioning accuracy that fails to meet practical needs. This invention achieves target positioning in three stages: the first stage rapidly locates the approximate target area using acoustic signature signals, with a response time ≤0.5s, ensuring positioning efficiency; the second stage, based on the initial acoustic signature positioning result, drives the gimbal to turn towards the target direction, performing visual refinement only on the target area, avoiding the high computational load of full-scene visual computing; the third stage verifies the target's validity using temperature data, further correcting the positioning result. This design ensures positioning accuracy while significantly reducing system computational costs, achieving a dual optimization of accuracy and efficiency. Compared to existing full-vision positioning systems, it significantly reduces computational load and improves response speed.

[0014] Existing single-modal positioning systems (single vision or single voiceprint) have low positioning accuracy. Single vision positioning achieves only 78% accuracy in noisy workshop environments, while single voiceprint positioning achieves only 65%, failing to meet the high-precision positioning requirements of industrial monitoring, security, and other scenarios. Existing dual-modal positioning systems lack effective multimodal data registration mechanisms, resulting in poor data fusion and limited improvement in positioning accuracy due to the inability to fully leverage the advantages of each modality. This invention utilizes a perspective projection transformation model in the fusion processing module to achieve precise registration of voiceprint coordinates, image coordinates, and temperature data. This allows for the complementary advantages of the three modalities—voiceprint signals enable rapid range locking, image data enables precise coordinate positioning, and temperature data enables target validity verification. With the synergistic effect of these three technologies, target positioning accuracy in noisy workshop environments can reach 92%, fully meeting the high-precision positioning needs of industries such as industrial and security.

[0015] This invention utilizes a data synchronization unit within the control module to achieve time alignment of multimodal data using timestamp marking. Simultaneously, it incorporates a trigger-based data acquisition mechanism—using voiceprint localization results as a trigger signal to initiate visual and temperature data acquisition—avoiding redundant acquisition of irrelevant data and further enhancing data synchronization. Furthermore, the invention, combined with the perspective projection transformation model of the fusion processing module, achieves precise fusion of multimodal coordinates, controlling the registration error to ≤0.05m. As a feasible preferred embodiment, the sensing module includes an acoustic sensor array, which comprises a reference receiver and peripheral receivers; the fusion processing module is used to calculate the time difference of sound waves arriving at each receiver and, in conjunction with the sound wave propagation speed, calculate the distance difference. The formula is as follows:

[0016] Where t0 is the reference receiver reception time, t i Let v be the reception time of a certain peripheral receiver, and v be the speed of sound propagation. The three-dimensional world coordinates of the sound source are calculated using the spherical intersection algorithm and used as the initial localization coordinates. The formula for the spherical intersection algorithm is:

[0017] in,( , , () represents the world coordinates of the sound source to be determined. , , Let be the coordinates of the i-th receiver. Let be the time difference between the i-th receiver and the reference receiver.

[0018] As a feasible preferred embodiment, the perception module includes an image acquisition unit equipped with an ORB feature extraction algorithm. The fusion processing module uses the ORB feature point matching algorithm to extract and match features from the acquired image, and combines a perspective projection transformation model to convert the image pixel coordinates into world coordinates. This achieves visual refinement of the initial voiceprint localization coordinates and filters the random noise in the initial voiceprint localization output signal to reduce jitter in the output signal. The formula is as follows:

[0019] in, These are the filter coefficients.

[0020] As a feasible preferred embodiment, the fusion processing module is used for error correction, specifically including the following: Radial distortion correction of the image is performed using a polynomial model. Combining the deviation between the visual localization results and the voiceprint coordinates, an error compensation function is established to correct the coordinate deviation. The correction formula is as follows:

[0021]

[0022] in, and The distortion coefficient is... The distance from the pixel to the principal point. and These are the coordinates of the camera's principal point.

[0023] Coordinate deviation correction combines the deviation between visual positioning results and voiceprint coordinates to establish an error compensation function:

[0024] in, and For compensation coefficient, As a feasible preferred solution, the perspective projection transformation model registers and fuses voiceprint coordinates, image coordinates, and temperature data. The perspective projection transformation model is used to establish the mapping relationship between world coordinates and pixel coordinates, and its expression is:

[0025] Where (u, v) are the image pixel coordinates, (X, Y, Z) are the world coordinates obtained from voiceprint localization, and K is the camera intrinsic parameter matrix. This is the extrinsic parameter matrix.

[0026] As a feasible and preferred solution, the pan-tilt controller receives the three-dimensional world coordinates (X, Y, F, Z) of the sound source output by the fusion processing module. w , Y w Z w Using the gimbal rotation center as the reference point (X0, Y0, Z0), calculate the target's relative coordinates:

[0027]

[0028]

[0029] The horizontal rotation angle Pan is calculated using the arctangent function in the four quadrants, as shown in the following formula:

[0030] Normalized to 0°–360° range:

[0031] The formula for calculating the pitch angle (Tilt) is as follows:

[0032]

[0033] Zero protection is prevented using the following formula:

[0034]

[0035] And limited to the range of -30° to +90°:

[0036] Secondly, this application also provides a multimodal registration and target localization method, which utilizes the aforementioned multimodal registration and target localization system. This method includes the following steps: Acquire voiceprint signals, image data, and temperature data, and achieve time alignment of multimodal data through timestamp marking; The three-dimensional world coordinates of the sound source are calculated based on the voiceprint signal using a time difference algorithm to obtain the initial location coordinates of the sound source. Based on the initial location coordinates of the sound source, the gimbal is driven to turn towards the target direction. The initial location coordinates of the sound source are then visually refined using image data, a feature point matching algorithm, and a perspective projection transformation model to obtain the visually refined coordinates. The voiceprint coordinates, visual refinement coordinates, and temperature data are registered using a perspective projection transformation model. The validity of the target is verified by judging whether the temperature data matches the target attributes. If the verification is successful, the final target positioning result is output.

[0037] As a feasible preferred solution, the step of visually refining the initial localization coordinates of the sound source using image data through a feature point matching algorithm combined with a perspective projection transformation model specifically includes: The acquired image is converted to grayscale and subjected to Gaussian filtering. The ORB feature point matching algorithm is used to extract feature points from the image and calculate descriptors, which are then matched with the target template. Perspective transformation relationships are established by matching feature point pairs, and the precise coordinates of the target in the world coordinate system are calculated by combining the camera intrinsic parameter matrix.

[0038] As a feasible and preferred solution, the method of verifying the effectiveness of the target by determining whether the temperature data conforms to the target attributes specifically involves: Determine whether the collected temperature data is within the preset abnormal temperature range; If the temperature data is within the abnormal temperature range, then the location result is confirmed as a valid target; If the temperature data is within the normal temperature range, it is determined to be an interfering sound source and the localization result is removed. The initial sound source localization step is then repeated. Attached Figure Description

[0039] Figure 1 This is a schematic diagram of the architecture of a multimodal registration and target localization system. Detailed Implementation

[0040] To make the technical solution and advantages of this application clearer, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only some embodiments of the present invention, and are only used to explain this application, not to limit it. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered isolated; they can be combined with each other to achieve better technical effects. The same reference numerals appearing in the accompanying drawings of the following embodiments represent the same features or components, and can be applied to different embodiments.

[0041] Furthermore, unless otherwise defined, the technical or scientific terms used in this invention description shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains.

[0042] The present invention will now be described in further detail with reference to the accompanying drawings.

[0043] This disclosure provides a multimodal registration and target localization system, referring to... Figure 1 It includes a perception module, a control module, a fusion processing module, and an execution module.

[0044] The sensing module is a system consisting of an acoustic sensor array, an image acquisition unit, and a temperature acquisition unit, which are responsible for acquiring acoustic signature, image, and temperature data, respectively.

[0045] In this embodiment, an infrasound receiver array is used, specifically a hexagonal circular distribution of six receivers. The geometric center (center of the circle) of the infrasound receiver array is set as the origin of the world coordinate system (0,0,0). A reference receiver is placed at the center, with a receiver spacing of 1.2m and an array diameter of 3m. This distribution ensures that sound source signals from any direction within a 360° range can be captured by at least three receivers, avoiding directional blind spots. Compared to a linear array, the time difference calculation error is reduced by 40%.

[0046] The image acquisition unit includes a high-definition camera equipped with the ORB (Oriented Fast and Rotated BRIEF) feature extraction algorithm to acquire image data of the target area. The ORB algorithm has rotation invariance and good matching performance, making it suitable for feature extraction and matching in complex environments.

[0047] The temperature acquisition unit includes an infrared temperature sensor. In this embodiment, the temperature detection range is -20℃ to 150℃ with an accuracy of ±0.5℃. It is used to collect temperature data of the target area to provide a basis for subsequent target validity verification.

[0048] The control module includes a gimbal controller and a data synchronization unit, which are used to drive the gimbal to turn towards the target direction based on the initial acoustic signature, and to achieve time alignment of the three-modal data through timestamp marking.

[0049] The gimbal controller receives the initial acoustic signature coordinates from the fusion processing module, calculates the gimbal turning angle, and drives the gimbal to turn the camera and infrared temperature sensor toward the target direction, providing a foundation for subsequent visual refinement and temperature verification.

[0050] Specifically, the random noise in the initial voiceprint localization output signal is filtered to reduce output signal jitter, as shown in the following formula:

[0051] in, The filter coefficient is 0.1 to 0.3, which is used in this embodiment depending on the application scenario.

[0052] The PTZ controller receives the three-dimensional world coordinates (X, Y) of the sound source output from the fusion processing module. w , Y w Z w Using the gimbal rotation center as the reference point (X0, Y0, Z0), calculate the target's relative coordinates:

[0053]

[0054]

[0055] The horizontal rotation angle Pan is calculated using the arctangent function in the four quadrants, as shown in the following formula:

[0056] Normalized to 0°–360° range:

[0057] The formula for calculating the pitch angle (Tilt) is as follows:

[0058]

[0059] Zero protection is prevented using the following formula:

[0060]

[0061] And limited to the range of -30° to +90°:

[0062] Because the input of the acoustic signal may cause the pan-tilt unit to tilt too far during operation, affecting the angle of each movement. The angle was changed to 0.0001, and angle changes of less than 1 degree were no longer input in real time to avoid the pain of frequent fine-tuning of the motor ESC, resulting in smoother rotation. The gimbal uses closed-loop control, triggering a visual refinement process upon completion.

[0063] The data synchronization unit is used to achieve time alignment of voiceprint, image and temperature data through timestamp marking, ensuring the time synchronization of different modal data and providing a foundation for subsequent data fusion and registration.

[0064] The fusion processing module incorporates a time difference algorithm, an ORB feature point matching algorithm, and a perspective projection transformation model for data processing, coordinate registration, and target validity assessment.

[0065] The time-difference algorithm is used to calculate the three-dimensional world coordinates of a sound source. Let the reception time of the reference receiver be t0, and the reception time of a certain peripheral receiver be t. i If the speed of sound propagation is v (taken as 340 m / s at room temperature), then the distance difference between the receiver and the reference receiver is:

[0066] By combining the distance difference data of each receiver, the three-dimensional coordinates of the sound source are calculated using the spherical intersection algorithm, and the initial positioning error is controlled within 1m.

[0067] The formula for the spherical intersection algorithm is:

[0068] in,( , , () represents the world coordinates of the sound source to be determined. , , Let be the coordinates of the i-th receiver. Let be the time difference between the i-th receiver and the reference receiver.

[0069] The perspective projection transformation model is used to establish the mapping relationship between world coordinates and pixel coordinates, achieving accurate registration of voiceprint coordinates (3D world coordinates), image coordinates (2D pixel coordinates), and temperature data (attribute data). The model expression is as follows:

[0070] Where (u, v) are the image pixel coordinates, (X, Y, Z) are the world coordinates obtained from voiceprint localization, and K is the camera intrinsic parameter matrix.

[0071] The camera intrinsic parameter matrix K (including focal length) is obtained using Zhang Zhengyou's calibration method. , coordinates of the principal point , ) and extrinsic parameter matrix (R is the rotation matrix, and t is the translation vector). Temperature data is bound to pixel coordinates through sensor IDs, establishing a relationship table between pixel coordinates and temperature values.

[0072] The error correction mechanism employs a two-step correction method. Radial distortion correction eliminates the effects of lens distortion through a polynomial model, and the correction formula is as follows:

[0073]

[0074] in, and The distortion coefficients are obtained through calibration. The distance from the pixel to the principal point. and The coordinates of the camera's principal point (optical center).

[0075] Coordinate deviation correction combines the deviation between visual positioning results and voiceprint coordinates to establish an error compensation function:

[0076] in, and The compensation coefficient was obtained by fitting experimental data, and the registration error was ultimately controlled within 0.05m.

[0077] The ORB feature point matching algorithm is used for precise target localization in the visual refinement stage. The algorithm steps include: Image preprocessing: The image captured by the camera is converted to grayscale and Gaussian filtered, with a 5×5 convolution kernel (grayscale conversion and Gaussian filtering) to remove environmental noise; Feature extraction involves detecting feature points in the image using the ORB algorithm, calculating the orientation and descriptor of each feature point, and ensuring the rotation invariance of the feature points. Feature matching involves performing Hamming distance matching between feature points in the current image and feature points in the target template, and filtering feature point pairs with a matching degree ≥ 0.85. Specifically, the ORB feature descriptor is a 256-bit binary string, and the Hamming distance is the number of corresponding bits that differ between two binary strings. Let the current image feature point descriptor be... The target template feature point descriptor is Hamming distance calculation formula:

[0078] Where ⊕ represents the XOR operation, n is the binary bit index, and H is the Hamming distance value.

[0079] The matching degree is the Hamming distance similarity, and the formula is as follows:

[0080] Matching score range: 0~1; the higher the matching score, the better the consistency of feature points. Feature point pairs with a Hamming distance ≤ 38 are retained (matching score ≥ 0.85).

[0081] Coordinate calculation establishes perspective transformation relationships by matching feature point pairs, and calculates the precise coordinates of the target in the world coordinate system by combining the camera intrinsic parameter matrix, achieving a visual refinement accuracy of 0.1m.

[0082] Specifically, based on matching feature point pairs, the perspective homography matrix H (3×3) is solved, satisfying the following formula: =H

[0083] in, This indicates the current pixel coordinates of the image.

[0084] Matrix H:

[0085] Combining the homography matrix H with the camera intrinsic parameter matrix K, we can decompose the rotation matrix R and the translation vector t:

[0086] The rotation matrix R is an orthogonal identity matrix with a determinant of 1.

[0087] The execution module includes a display unit and an alarm unit, which are used to output the target positioning coordinates and temperature information, and trigger an alarm when a valid target is located.

[0088] The display unit uses an LCD screen or an LED screen to display the target positioning coordinates (X, Y, Z) and temperature information (T) in real time, so that operators can intuitively understand the target status.

[0089] The alarm unit uses an audible and visual alarm. When the fusion processing module confirms that a valid target has been located, it triggers an alarm to remind the operator to take timely measures.

[0090] A multimodal registration and target localization method, which utilizes the aforementioned multimodal registration and target localization system, includes the following steps.

[0091] Initial sound source localization: Vibrations from the operating heat-generating equipment in the workshop radiate infrasound waves, which are collected in real time by an acoustic sensor array. The three-dimensional world coordinates of the sound source are calculated using a time-difference algorithm to obtain an approximate location (e.g., X=15.2m, Y=8.6m, Z=1.2m, with an error of 0.8m). The control module drives the pan-tilt unit to turn the camera and infrared sensor to this location, completing the initial localization. Actual measurements show this process takes ≤0.5s.

[0092] Visual refinement is then performed. After the gimbal turns towards the target direction, the camera immediately captures an image of the target area. The image is processed using the ORB feature point matching algorithm: feature points of the device in the image are extracted and matched with a template. Combined with a perspective projection transformation model, the image pixel coordinates are converted into world coordinates to obtain a precise positioning result (e.g., X=15.23m, Y=8.58m, Z=1.21m, error 0.09m). Simultaneously, an infrared sensor collects temperature data for the area, for example, obtaining a target temperature of 68℃.

[0093] Temperature verification and fusion processing modules register the initial acoustic signature coordinates, visual refinement coordinates, and temperature data using a perspective projection transformation model. The module then determines whether the temperature data matches the target attributes (normal workshop equipment temperature ≤ 45℃, 68℃ is considered abnormal). If the temperature is abnormal, the positioning result is confirmed as a valid target, and the final coordinates and temperature information are output. If the temperature is normal (e.g., the temperature in the positioning area corresponding to wind noise is 25℃), it is determined to be an interfering sound source, the positioning result is discarded, and the initial sound source positioning is performed again.

[0094] The above content is merely an embodiment of the present invention. Commonly known structures and characteristics of the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can improve and implement this solution based on the guidance provided in this application and their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A multimodal registration and target localization system, characterized in that, include: The sensing module is used to collect voiceprint signals, image data, and temperature data; The control module includes a gimbal controller and a data synchronization unit. The gimbal controller is used to receive positioning coordinates and drive the gimbal to turn in the target direction. The data synchronization unit is used to achieve time alignment of multimodal data through timestamp marking. The fusion processing module is used to calculate the initial location coordinates of the sound source based on the voiceprint signal, perform visual refinement by combining image data, and register and fuse the voiceprint coordinates, image coordinates and temperature data to output the target location result. The execution module is used to display the target location coordinates and temperature information, and to trigger an alarm when a valid target is located.

2. The multimodal registration and target localization system according to claim 1, characterized in that, The sensing module includes an acoustic sensor array, which includes a reference receiver and a peripheral receiver; The fusion processing module is used to calculate the time difference of sound waves arriving at each receiver, and to calculate the distance difference by combining the sound wave propagation speed. The formula is as follows: Where t0 is the reference receiver reception time, t i Let v be the reception time of a certain peripheral receiver, and v be the speed of sound propagation. The three-dimensional world coordinates of the sound source are calculated using the spherical intersection algorithm and used as the initial localization coordinates. The formula for the spherical intersection algorithm is: in,( , , () represents the world coordinates of the sound source to be determined. , , Let be the coordinates of the i-th receiver. Let be the time difference between the i-th receiver and the reference receiver.

3. The multimodal registration and target localization system according to claim 1, characterized in that, The perception module includes an image acquisition unit equipped with an ORB feature extraction algorithm. The fusion processing module uses the ORB feature point matching algorithm to extract and match features from the acquired image, and combines a perspective projection transformation model to convert the image pixel coordinates into world coordinates to achieve visual refinement of the initial voiceprint localization coordinates and filter random noise in the initial voiceprint localization output signal to reduce output signal jitter. The formula is as follows: in, These are the filter coefficients.

4. The multimodal registration and target localization system according to claim 1, characterized in that, The fusion processing module is used for error correction, specifically including the following: Radial distortion correction of the image is performed using a polynomial model. Combining the deviation between the visual localization results and the voiceprint coordinates, an error compensation function is established to correct the coordinate deviation. The correction formula is as follows: in, and The distortion coefficient is... The coordinates of the camera's principal point; Coordinate deviation correction combines the deviation between visual positioning results and voiceprint coordinates to establish an error compensation function: in, and This is the compensation coefficient.

5. The multimodal registration and target localization system according to claim 1, characterized in that, The perspective projection transformation model is used to register and fuse voiceprint coordinates, image coordinates, and temperature data. This model establishes the mapping relationship between world coordinates and pixel coordinates, expressed as: Where (u, v) are the image pixel coordinates, (X, Y, Z) are the world coordinates obtained from voiceprint localization, and K is the camera intrinsic parameter matrix. This is the extrinsic parameter matrix.

6. The multimodal registration and target localization system according to claim 1, characterized in that, The PTZ controller receives the three-dimensional world coordinates (X, Y) of the sound source output from the fusion processing module. w , Y w Z w Using the gimbal rotation center as the reference point (X0, Y0, Z0), calculate the target's relative coordinates: The horizontal rotation angle Pan is calculated using the arctangent function in the four quadrants, as shown in the following formula: Normalized to 0°–360° range: The formula for calculating the pitch angle (Tilt) is as follows: Zero protection is prevented using the following formula: And limited to the range of -30° to +90°: 。 7. A multimodal registration and target localization method, characterized in that, The method, employing the multimodal registration and target localization system as described in any one of claims 1-6, includes the following steps: Acquire voiceprint signals, image data, and temperature data, and achieve time alignment of multimodal data through timestamp marking; The three-dimensional world coordinates of the sound source are calculated based on the voiceprint signal using a time difference algorithm to obtain the initial location coordinates of the sound source. Based on the initial location coordinates of the sound source, the gimbal is driven to turn towards the target direction. The initial location coordinates of the sound source are then visually refined using image data, a feature point matching algorithm, and a perspective projection transformation model to obtain the visually refined coordinates. The voiceprint coordinates, visual refinement coordinates, and temperature data are registered using a perspective projection transformation model. The validity of the target is verified by judging whether the temperature data matches the target attributes. If the verification is successful, the final target positioning result is output.

8. The multimodal registration and target localization method according to claim 7, characterized in that, The step of visually refining the initial location coordinates of the sound source using image data through a feature point matching algorithm combined with a perspective projection transformation model specifically includes: The acquired image is converted to grayscale and subjected to Gaussian filtering. The ORB feature point matching algorithm is used to extract feature points from the image and calculate descriptors, which are then matched with the target template. Perspective transformation relationships are established by matching feature point pairs, and the precise coordinates of the target in the world coordinate system are calculated by combining the camera intrinsic parameter matrix.

9. The multimodal registration and target localization method according to claim 7, characterized in that, The method of verifying the effectiveness of the target by judging whether the temperature data matches the target attributes specifically involves: Determine whether the collected temperature data is within the preset abnormal temperature range; If the temperature data is within the abnormal temperature range, then the location result is confirmed as a valid target; If the temperature data is within the normal temperature range, it is determined to be an interfering sound source and the localization result is removed. The initial sound source localization step is then repeated.

Citation Information

Patent Citations

  • In-vehicle target positioning method, device, equipment and medium

    CN120254762A