Target tracking method, system, device, electronic device, storage medium and product

CN122815329APending Publication Date: 2026-09-25INSPUR (SHANDONG) COMPUTER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610941623.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,纯视觉系统却存在固有的技术局限性:(1)系统通常需要在摄像头“看到”异常行为后才能触发报警,无法对发生在视觉盲区或具有声音征兆(如争吵、呼救、破坏声)的潜在威胁进行提前预警,存在感知被动性与延迟性;(2)摄像头视野受建筑物、绿化带、车辆等物体遮挡,无法做到全域覆盖,难以保证准确性

Benefits of technology

[0020]本发明提供了一种目标跟踪方法,包括:获取各音频采集设备发送的音频数据,所述音频数据包括目标对象的原始音频数据以及关于所述目标对象的初始识别数据和初始定位数据;分别对各所述初始识别数据和各所述初始定位数据进行融合处理,得到所述目标对象的声纹事件类型和声纹活跃区域;计算各视频采集设备与所述声纹活跃区域之间的关联度,并根据各所述关联度和所述声纹事件类型确定目标视频采集设备;控制所述目标视频采集设备获取所述目标对象的视频数据,并将所述视频数据、所述音频数据、所述声纹活跃区域进行关联处理,得到所述目标对象的目标跟踪结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122815329A_ABST
    Figure CN122815329A_ABST
Patent Text Reader

Abstract

The application discloses a target tracking method, system and device, electronic equipment, storage medium and product, relates to the technical field of Internet, and is used for solving the problems of delay and inaccuracy in target tracking only relying on computer vision in traditional technology, and the method comprises the steps of obtaining audio data sent by each audio acquisition device, wherein the audio data comprises original audio data of a target object and initial identification data and initial positioning data about the target object; performing fusion processing on each initial identification data and each initial positioning data respectively to obtain a voiceprint event type and a voiceprint active area of the target object; calculating the correlation degree between each video acquisition device and the voiceprint active area, determining a target video acquisition device according to each correlation degree and the voiceprint event type; controlling the target video acquisition device to obtain video data of the target object, and performing correlation processing on the video data, the audio data and the voiceprint active area to obtain a target tracking result of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet technology, and in particular to a target tracking method, as well as a target tracking system, device, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] With the deepening of smart city and safe city construction, video surveillance networks have widely covered public areas and become the core of the security system. Existing intelligent security systems rely heavily on computer vision technology, such as face recognition, behavior analysis and target tracking. However, pure vision systems have inherent technical limitations: (1) The system usually needs to "see" abnormal behavior before it can trigger an alarm, and it cannot provide early warning of potential threats that occur in visual blind spots or have sound signs (such as arguments, calls for help, or destructive sounds), resulting in passive perception and delay; (2) The camera's field of view is blocked by objects such as buildings, green belts, and vehicles, making it impossible to achieve full coverage and difficult to guarantee accuracy.

[0003] Therefore, how to achieve more timely and accurate target tracking in order to realize safer and more effective intelligent security is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] The purpose of this invention is to provide a target tracking method that can achieve more timely and accurate target tracking, thereby enabling safer and more effective intelligent security. Another purpose of this invention is to provide a target tracking system, device, electronic device, computer-readable storage medium, and computer program product, all of which have the aforementioned beneficial effects.

[0005] In a first aspect, the present invention provides a target tracking method, comprising: Acquire audio data sent by each audio acquisition device, the audio data including the original audio data of the target object as well as the initial identification data and initial positioning data of the target object; The initial identification data and the initial positioning data are fused separately to obtain the voiceprint event type and voiceprint active area of ​​the target object; Calculate the correlation degree between each video acquisition device and the active area of ​​the voiceprint, and determine the target video acquisition device based on the correlation degree and the voiceprint event type; The target video acquisition device is controlled to acquire video data of the target object, and the video data, audio data, and active voiceprint region are correlated to obtain the target tracking result of the target object.

[0006] The initial identification data includes the initial voiceprint event type and the confidence level corresponding to the initial voiceprint event type; Accordingly, the initial identification data are fused to obtain the voiceprint event type of the target object, including: The weight value of each initial voiceprint event type is determined based on the confidence level corresponding to each initial voiceprint event type. The initial voiceprint event types are weighted and fused according to their respective weight values ​​to obtain the voiceprint event type of the target object.

[0007] The initial positioning data are fused to obtain the active voiceprint region of the target object, including: The weight value of each initial location data is determined based on the confidence level corresponding to each initial voiceprint event type. The weighted fusion calculation is performed on each of the initial positioning data according to the weight value of each initial positioning data to obtain the fused positioning data of the target object; The fused positioning data is corrected using a sound field propagation model to obtain the active voiceprint region of the target object; the active voiceprint region is based on elliptic probability representation.

[0008] The calculation of the correlation between each video acquisition device and the active area of ​​the acoustic signature includes: Calculate the overlap area between the screen area of ​​each video acquisition device and the active area of ​​the audioprint; Calculate the distance between each video acquisition device and the active area of ​​the acoustic signature; For each video acquisition device, the correlation between the video acquisition device and the active area of ​​the audioprint is calculated based on the overlapping area, distance length, resolution coefficient, and pan-tilt availability coefficient corresponding to the video acquisition device.

[0009] The determination of the target video acquisition device based on the correlation degree and the voiceprint event type includes: Search the video acquisition device scheduling rules corresponding to the aforementioned voiceprint event type in the strategy knowledge base; Arrange the aforementioned correlation degrees in a preset order to obtain a video acquisition device correlation list; The target video acquisition device is determined in the video acquisition device association list according to the video acquisition device scheduling rules.

[0010] The initial identification data includes the initial voiceprint event type and the timestamp corresponding to the initial voiceprint event type; Accordingly, the video data, the audio data, and the active voiceprint region are correlated to obtain the target tracking result of the target object, including: Calculate the fusion timestamp based on the timestamp corresponding to each of the initial voiceprint event types; Extract the audio segment corresponding to the fusion timestamp from each of the audio data; Project the active voiceprint region onto each of the audio segments; In the active voiceprint region, the target object and its trajectory information are identified and marked within the projection area of ​​each audio segment to obtain each target audio segment; By binding each target audio segment with the voiceprint event type, the target tracking result of the target object is obtained.

[0011] This includes acquiring the audio data sent by each audio acquisition device, including: A collaborative instruction is sent to each of the audio acquisition devices, so that each audio acquisition device can use an acoustic model to process the raw audio data to determine the initial recognition data, and determine the initial positioning data based on the audio time difference and spatial position relationship between itself and the adjacent audio acquisition devices.

[0012] The process of sending collaborative instructions to each of the aforementioned audio acquisition devices includes: When an abnormal video alert is received from any video capture device, and it is determined that no audio data has been received from any of the audio capture devices within the abnormal area, the coordination instruction is sent to each audio capture device within the abnormal area; wherein, the abnormal area is determined based on the abnormal video in the abnormal video alert.

[0013] The process of sending collaborative instructions to each of the aforementioned audio acquisition devices includes: When an abnormal audio alert is received from any audio acquisition device, and it is determined that the distance between the audio acquisition device that sent the abnormal audio alert and the estimated location of the abnormal sound source exceeds a preset threshold, the coordination instruction is sent to each audio acquisition device within the target area where the estimated location of the abnormal sound source is located; wherein, the estimated location of the abnormal sound source is determined based on the abnormal audio in the abnormal audio alert.

[0014] The process of sending collaborative instructions to each of the aforementioned audio acquisition devices includes: Based on the historical tracking results of the target object, the future movement area of ​​the target object is estimated, and the collaborative command is sent to each audio acquisition device within the future movement area.

[0015] Secondly, the present invention also discloses a target tracking system, including a central server, various audio acquisition devices, and various video acquisition devices, wherein each audio acquisition device is communicatively connected to the central server, and each video acquisition device is communicatively connected to the central server. The central server is used to acquire audio data sent by each of the audio acquisition devices. The audio data includes the original audio data of the target object and initial identification data and initial positioning data of the target object. The server performs fusion processing on each of the initial identification data and initial positioning data to obtain the voiceprint event type and voiceprint active area of ​​the target object. The server calculates the correlation degree between each of the video acquisition devices and the voiceprint active area, and determines the target video acquisition device based on the correlation degree and the voiceprint event type. The server controls the target video acquisition device to acquire the video data of the target object, and performs correlation processing on the video data, the audio data, and the voiceprint active area to obtain the target tracking result of the target object.

[0016] Thirdly, the present invention also discloses a target tracking device, comprising: The acquisition module is used to acquire audio data sent by each audio acquisition device. The audio data includes the original audio data of the target object and the initial identification data and initial positioning data of the target object. The fusion module is used to perform fusion processing on each of the initial identification data and each of the initial positioning data to obtain the voiceprint event type and voiceprint active area of ​​the target object; The calculation module is used to calculate the correlation degree between each video acquisition device and the active area of ​​the voiceprint, and to determine the target video acquisition device based on the correlation degree and the voiceprint event type. The control module is used to control the target video acquisition device to acquire video data of the target object, and to perform correlation processing on the video data, the audio data, and the active area of ​​the voiceprint to obtain the target tracking result of the target object.

[0017] Fourthly, the present invention also discloses an electronic device, comprising: Memory, used to store computer programs; A processor for executing the computer program to implement any of the target tracking methods described above.

[0018] Fifthly, the present invention also discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the target tracking methods described above.

[0019] In a sixth aspect, the present invention also discloses a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of any of the target tracking methods described above.

[0020] This invention provides a target tracking method, comprising: acquiring audio data sent by various audio acquisition devices, the audio data including original audio data of a target object and initial identification data and initial positioning data of the target object; performing fusion processing on each of the initial identification data and each of the initial positioning data to obtain the voiceprint event type and voiceprint active region of the target object; calculating the correlation degree between each video acquisition device and the voiceprint active region, and determining the target video acquisition device based on each correlation degree and the voiceprint event type; controlling the target video acquisition device to acquire video data of the target object, and performing correlation processing on the video data, the audio data, and the voiceprint active region to obtain the target tracking result of the target object.

[0021] By applying the technical solution provided by this invention, audio and video acquisition devices are simultaneously deployed in the target tracking area to combine computer vision and acoustic technologies. In the process, firstly, audio data collected by each audio acquisition device is fused to determine the target object's voiceprint event type and active voiceprint area. Then, the correlation between each video acquisition device and the active voiceprint area is calculated, combined with the target object's voiceprint event type, to select the target video acquisition device. Finally, the target video acquisition device is used to acquire the target object's video data and establish a correlation with the video data and the active voiceprint area to obtain the target tracking result. Therefore, this technical solution combines computer vision and acoustic technologies to achieve target tracking. The audio acquisition device, with its characteristics of being unaffected by light, having no visual obstruction, and being able to sense events outside the line of sight, helps to achieve more timely and accurate target tracking, thereby realizing safer and more effective intelligent security.

[0022] The target tracking system, device, electronic equipment, computer-readable storage medium, and computer program product provided by this invention also have the above-mentioned technical effects, and will not be described in detail here. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the prior art and the embodiments of the present invention, the accompanying drawings used in the description of the prior art and the embodiments of the present invention will be briefly introduced below. Of course, the accompanying drawings described below with respect to the embodiments of the present invention are only a part of the embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort, and such other drawings also fall within the protection scope of the present invention.

[0024] Figure 1 This is a schematic diagram of the structure of a target tracking system provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating a target tracking method provided in an embodiment of the present invention; Figure 3 This is an overall architecture diagram of a target tracking system provided in an embodiment of the present invention; Figure 4 A timing diagram of a target tracking method provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a target tracking device provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0025] The core of this invention is to provide a target tracking method that can achieve more timely and accurate target tracking, thereby enabling safer and more effective intelligent security. Another core aspect of this invention is to provide a target tracking system, device, electronic device, computer-readable storage medium, and computer program product, all of which have the aforementioned beneficial effects.

[0026] To provide a clearer and more complete description of the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be introduced below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0027] This invention provides a target tracking method.

[0028] First, please refer to Figure 1 , Figure 1 This is a schematic diagram of a target tracking system provided in an embodiment of the present invention. The target tracking system includes a central server 100, multiple audio acquisition devices 200, and multiple video acquisition devices 300. Each audio acquisition device 200 is communicatively connected to the central server 100, and each video acquisition device 300 is communicatively connected to the central server 100. The target tracking method provided in this embodiment of the present invention is implemented based on this target tracking system.

[0029] For further information, please refer to [link / reference]. Figure 2 , Figure 2This is a flowchart illustrating a target tracking method provided in an embodiment of the present invention. The target tracking method is applied to... Figure 1 The central server 100 shown may include the following steps S101 to S104.

[0030] S101: Acquire audio data sent by each audio acquisition device. The audio data includes the original audio data of the target object as well as the initial identification data and initial positioning data of the target object.

[0031] This step aims to acquire audio data. Specifically, each audio acquisition device deployed in the intelligent security area (target tracking area) can collect audio data and upload it to the central server. This audio data mainly includes the raw audio data of the target object, as well as initial identification data and initial positioning data about the target object. The target object is the target tracking object. The initial identification data and initial positioning data about the target object can be obtained by analyzing the corresponding raw audio data. The initial identification data is the identification result of the target object, such as event type, timestamp, confidence level, etc.; the initial positioning data is the positioning result of the target object, such as coordinate information, latitude and longitude information, etc.

[0032] It should be noted that the central server may acquire audio data in various ways. It may issue instructions to each audio acquisition device to actively acquire audio data, or it may receive audio data actively uploaded by each audio acquisition device. This application does not limit the method. The following embodiment provides a method for the central server to acquire audio data.

[0033] In one embodiment of the present invention, acquiring audio data sent by each audio acquisition device may include: sending a coordination instruction to each audio acquisition device so that each audio acquisition device can use an acoustic model to process the raw audio data to determine initial recognition data, and determine initial positioning data based on the audio time difference and spatial position relationship between itself and adjacent audio acquisition devices.

[0034] As mentioned above, the audio data mainly includes the raw audio data of the target object, as well as initial identification and initial positioning data of the target object. The initial identification and initial positioning data can be obtained by analyzing the corresponding raw audio data. In specific implementation, acoustic models can be used to process the raw audio data to obtain the initial identification data; the initial positioning data can be obtained based on the audio time difference and spatial relationship between adjacent audio acquisition devices. Furthermore, since the initial positioning data requires collaborative processing between adjacent audio acquisition devices, each audio acquisition device can perform the above operations based on collaborative instructions issued by the central server. Further, when an audio acquisition device fails to receive a collaborative instruction from the central server, it can simply perform real-time / timed acquisition of the raw audio data and / or anomaly analysis and processing of the raw audio data (i.e., without data communication with other audio acquisition devices).

[0035] Based on this, the following embodiments further propose multiple triggering conditions for the central server to issue collaborative instructions to each audio acquisition device.

[0036] For example, issuing coordination instructions to each audio acquisition device can include: when receiving an abnormal video alert from any video acquisition device, and determining that no audio data has been received from any audio acquisition device within the abnormal area, issuing coordination instructions to each audio acquisition device within the abnormal area; wherein, the abnormal area is determined based on the abnormal video in the abnormal video alert. Specifically, multiple video acquisition devices will also be deployed within the intelligent security area to achieve intelligent monitoring of the intelligent security area, and each video acquisition device can perform video anomaly analysis in addition to its video acquisition function. Once an abnormal video is detected, it will upload an abnormal video alert to the central server; at this time, if the central server finds that it has not received audio data uploaded by each audio acquisition device within the abnormal area corresponding to the abnormal video, it can issue coordination instructions to each audio acquisition device within that abnormal area to obtain the audio data within the abnormal area.

[0037] For example, issuing collaborative instructions to each audio acquisition device can include: when receiving an abnormal audio alert from any audio acquisition device, and determining that the distance between the audio acquisition device sending the abnormal audio alert and the estimated location of the abnormal sound source exceeds a preset threshold, issuing collaborative instructions to each audio acquisition device within the target area where the estimated location of the abnormal sound source is located; wherein, the estimated location of the abnormal sound source is determined based on the abnormal audio in the abnormal audio alert. As described above, when each audio acquisition device fails to receive the collaborative instructions issued by the central server, it can simply perform real-time / timed acquisition of raw audio data and / or abnormal analysis processing of raw audio data. Then, when an audio acquisition device detects abnormal audio, it can upload an abnormal audio alert to the central server. At this time, if the central server finds that the distance between the audio acquisition device sending the abnormal audio alert and the estimated location of the abnormal sound source corresponding to the abnormal audio is too long, it can directly issue collaborative instructions to each audio acquisition device within the target area where the estimated location of the abnormal sound source is located, so as to acquire audio data within that target area. The estimated location of the abnormal sound source can be achieved by the audio acquisition device sending the abnormal audio alert using its built-in sound source localization model.

[0038] For example, issuing coordination instructions to each audio acquisition device could include: estimating the target object's future movement area based on its historical tracking results, and issuing coordination instructions to each audio acquisition device within that future movement area. It's understandable that the target object is not always stationary, such as a "rapidly moving distress call." In this case, the central server can aggregate the target object's historical tracking results to determine its future movement area—the area the target object is likely to move to—and issue coordination instructions to each audio acquisition device within that future movement area. This allows for the acquisition of audio data within that area, thereby ensuring the continuity of target tracking and the accuracy of the tracking results.

[0039] S102: Perform fusion processing on each initial identification data and each initial positioning data to obtain the voiceprint event type and voiceprint active area of ​​the target object.

[0040] This step aims to achieve the fusion processing of various initial identification data and various initial positioning data, so as to obtain the voiceprint event type and voiceprint active area of ​​the target object, respectively. The following two embodiments propose methods for fusion processing of multiple initial identification data and methods for fusion processing of multiple initial positioning data, respectively.

[0041] In one embodiment of the present invention, the initial identification data includes initial voiceprint event types and corresponding confidence levels. Accordingly, fusing each initial identification data to obtain the voiceprint event type of the target object can include: determining the weight value of each initial voiceprint event type based on its corresponding confidence level; and performing a weighted fusion calculation on each initial voiceprint event type based on its weight value to obtain the voiceprint event type of the target object. Specifically, each audio acquisition device can analyze the raw audio data it acquires to obtain an initial voiceprint event type and its corresponding confidence level. This process can be implemented using the acoustic model built into the audio acquisition device. Therefore, the central server can set the weight value of each initial voiceprint event type based on its corresponding confidence level (i.e., the higher the confidence level, the larger the weight value), and then perform a weighted fusion calculation on each initial voiceprint event type using its weight value to obtain the voiceprint event type of the target object, i.e., the fusion result of multiple initial identification data.

[0042] In one embodiment of the present invention, fusing the initial positioning data to obtain the voiceprint active region of the target object may include: determining the weight value of each initial positioning data according to the confidence level corresponding to each initial voiceprint event type; performing weighted fusion calculation on the initial positioning data according to the weight value of each initial positioning data to obtain the fused positioning data of the target object; and correcting the fused positioning data using a sound field propagation model to obtain the voiceprint active region of the target object; the voiceprint active region is based on elliptic probability representation. That is, the central server can also set the weight value of each initial positioning data according to the confidence level corresponding to each initial voiceprint event type, i.e., the higher the confidence level, the larger the weight value. Then, by using the weight value of each initial positioning data to perform weighted fusion calculation on the initial positioning data, the fused positioning data of the target object can be obtained, i.e., the preliminary fusion result of multiple initial positioning data. Based on this, the central server can also use a pre-deployed sound field propagation model to correct the fused positioning data to obtain the voiceprint active region of the target object based on elliptic probability representation, rather than a single coordinate point, which better reflects the positioning uncertainty in a real environment. At this point, the final fusion result of multiple initial positioning data can be obtained.

[0043] The process of using a sound field propagation model to correct the fused positioning data and obtain the active area of ​​the target object's voiceprint can include: based on the three-dimensional geometric model of the smart security area (including the position information of reflective surfaces such as building exterior walls and ground), the theoretical path of sound waves reaching each audio acquisition device through each reflective surface is simulated in reverse. Each theoretical path is matched with the measured path, and the error between the theoretical path value and the measured path value is minimized through iterative optimization, thereby compensating for the positioning drift caused by building reflections, etc., and finally outputting a "active area of ​​voiceprint" represented by a probability ellipse.

[0044] S103: Calculate the correlation between each video acquisition device and the active area of ​​the voiceprint, and determine the target video acquisition device based on the correlation and the type of voiceprint event.

[0045] This step aims to select the target video capture device, that is, to choose a suitable target video capture device from all video capture devices within the intelligent security area to acquire the corresponding video data. This video data will then be combined with audio data to determine the target object's tracking results. Specifically, the correlation between each video capture device within the intelligent security area and the area with active voiceprint activity can be calculated, and then the target video capture device can be selected based on the target object's voiceprint event type. It can be understood that the higher the correlation value, the more suitable the corresponding video capture device is as the target video capture device for acquiring the relevant video data.

[0046] In one embodiment of the present invention, calculating the correlation between each video acquisition device and the active area of ​​the audioprint may include: calculating the overlap area between the screen area of ​​each video acquisition device and the active area of ​​the audioprint; calculating the distance between each video acquisition device and the active area of ​​the audioprint; and for each video acquisition device, calculating the correlation between the video acquisition device and the active area of ​​the audioprint based on the overlap area, distance, resolution coefficient, and pan-tilt availability coefficient corresponding to the video acquisition device.

[0047] This embodiment provides a method for calculating the correlation between a video capture device and an active area of ​​soundprints. This can be achieved by referring to the overlapping area between the screen area of ​​the video capture device and the active area of ​​soundprints, the distance between the video capture device and the active area of ​​soundprints, the resolution coefficient of the video capture device (whether the image quality is good), and the pan-tilt availability coefficient of the video capture device (whether the pan-tilt can be rotated and zoomed in and out to adjust the angle and focal length).

[0048] The calculation of the correlation between video capture devices and active audioprint areas, based on the overlapping area, distance, resolution coefficient, and pan-tilt availability coefficient of the corresponding video capture devices, can include: using a correlation model to calculate the correlation between the corresponding video capture devices and active audioprint areas based on the overlapping area, distance, resolution coefficient, and pan-tilt availability coefficient of the video capture devices; the correlation model is: Score_i = α × (OverlapArea_i / Area_A) + β × (1 / Distance_i) + γ × ResolutionCoeff_i + δ ×PTZAvailability_i, Score_i is the correlation degree, OverlapArea_i represents the overlap area between the screen area of ​​the video capture device and the active area of ​​the audioprint, Area_A represents the area of ​​the active area of ​​the audioprint, Distance_i represents the distance between the video capture device and the active area of ​​the audioprint, ResolutionCoeff_i represents the resolution coefficient of the video capture device, PTZAvailability_i represents the pan-tilt availability coefficient of the video capture device, and α, β, γ, δ are adjustable weights. In one possible implementation, α > β > γ = δ.

[0049] It should be noted that since the active area of ​​the audioprint is a probabilistic ellipse rather than a single coordinate point, the distance between the video capture device and the active area of ​​the audioprint can be calculated using the "distance from the audio capture device to the center of the ellipse". Clearly, the closer the audio capture device is to the center of the ellipse, the clearer the image, the less interference, and the higher the score. Furthermore, if the audio capture device is located inside the ellipse, Distance_i can be minimized (e.g., 1 meter) to avoid the distance being zero, which would cause (1 / Distance_i) in the formula to become infinite.

[0050] In one embodiment of the present invention, determining the target video acquisition device based on each correlation degree and voiceprint event type may include: querying the video acquisition device scheduling rules corresponding to the voiceprint event type in the policy knowledge base; arranging each correlation degree in a preset order to obtain a video acquisition device association list; and determining the target video acquisition device in the video acquisition device association list according to the video acquisition device scheduling rules.

[0051] This embodiment provides a method for selecting target video capture devices based on correlation. Specifically, a strategy knowledge base can be pre-created to store video capture device scheduling rules corresponding to different voiceprint event types. Simultaneously, the correlation between each video capture device and the voiceprint active area is arranged in descending order to obtain a video capture device association list. Thus, the video capture device scheduling rules corresponding to the voiceprint event type can be determined from the strategy knowledge base, and the target video capture device can be determined by combining this with the video capture device association list.

[0052] For example, when the voiceprint event type is a moving sound source event, the corresponding video acquisition equipment scheduling rule might be: "Prioritize scheduling the PTZ camera (video acquisition equipment) with the highest correlation to quickly turn to the general area for scanning; after scanning the target, lock and track it; at the same time, notify the fixed bullet camera (video acquisition equipment) in front of the target's movement direction to prepare for image comparison, and guide the next PTZ camera to perform relay preset." As another example, when the voiceprint event type is a high-priority event (such as an explosion), the corresponding video acquisition equipment scheduling rule might be: "Immediately schedule all associated video acquisition equipment (including fixed bullet cameras and PTZ cameras) within the active voiceprint area to simultaneously turn to the event area, with the PTZ camera performing both wide-angle scanning and zoomed-in close-up tasks; simultaneously trigger the highest-level alarm, push the real-time image to the command center's large screen; simultaneously lock the full audio and video data for 30 seconds before and after the event, marking it as high-priority evidence to prevent it from being repeatedly overwritten."

[0053] S104: Control the target video acquisition device to acquire video data of the target object, and perform correlation processing on the video data, audio data, and active area of ​​the voiceprint to obtain the target tracking result of the target object.

[0054] This step aims to correlate the audio and video data of the target object to obtain target tracking results. Specifically, the target video acquisition device can be controlled to acquire the video data of the target object and correlate it with the audio data and active areas of the speaker's voiceprint to obtain the target tracking results.

[0055] In one embodiment of the present invention, the initial identification data includes an initial voiceprint event type and a timestamp corresponding to the initial voiceprint event type. Accordingly, the video data, audio data, and voiceprint active region are correlated to obtain the target tracking result of the target object, which may include: calculating a fusion timestamp based on the timestamp corresponding to each initial voiceprint event type; extracting the audio segment corresponding to the fusion timestamp from each audio data; projecting the voiceprint active region onto each audio segment; identifying and marking the target object and the trajectory information of the target object within the projection area of ​​the voiceprint active region in each audio segment to obtain each target audio segment; and binding each target audio segment with the voiceprint event type to obtain the target tracking result of the target object.

[0056] This embodiment provides a method for associating video data, audio data, and active voiceprint regions. Specifically, the initial identification data may include the initial voiceprint event type and its corresponding timestamp, i.e., the specific time node when the corresponding voiceprint event occurred. Based on this, a fusion timestamp can be calculated first. This process can adopt a method similar to the determination method of voiceprint event type and active voiceprint region based on weighted fusion processing described above. Then, audio segments corresponding to the fusion timestamp are extracted from each audio data, such as audio segments 30 seconds before and after the fusion timestamp. At this time, the active voiceprint region is projected onto each audio segment to identify and mark the target object and its trajectory information within the projection area, obtaining each target audio segment. Finally, each target audio segment is bound to the voiceprint event type as the target tracking result of the target object.

[0057] As can be seen, the target tracking method provided in this embodiment of the invention deploys both audio and video acquisition devices in the target tracking area to combine computer vision technology with acoustic technology. In the implementation process, firstly, the audio data collected by each audio acquisition device is fused to determine the voiceprint event type and voiceprint active area of ​​the target object. Then, the correlation between each video acquisition device and the voiceprint active area is calculated and combined with the voiceprint event type of the target object to select the target video acquisition device. Finally, the target video acquisition device is used to acquire the video data of the target object and establish a correlation between the video data and the voiceprint active area to obtain the target tracking result. Therefore, this technical solution combines computer vision technology with acoustic technology to achieve target tracking. The audio acquisition device possesses characteristics such as being unaffected by light, having no visual obstruction, and being able to sense events outside the line of sight, which helps to achieve more timely and accurate target tracking, thereby achieving safer and more effective intelligent security.

[0058] This invention provides another target tracking method.

[0059] Please refer to Figure 3 and Figure 4, Figure 3 This is an overall architecture diagram of a target tracking system provided in an embodiment of the present invention. Figure 4 This is a timing diagram of a target tracking method provided in an embodiment of the present invention. The implementation process of the target tracking method provided in an embodiment of the present invention may specifically include the following steps.

[0060] 1. Acoustic perception layer and edge computing layer.

[0061] The acoustic perception layer and edge computing layer consist of distributed microphone array nodes (multiple audio acquisition devices) widely deployed in the monitoring area (intelligent security area). Each node has a built-in DSP (Digital Signal Processor) processing unit, which is mainly responsible for the following operations.

[0062] (1) Audio acquisition and processing: Acquire raw audio data and perform pre-processing such as pre-filtering and noise reduction.

[0063] (2) Initial judgment of voiceprint events: Run a lightweight acoustic model to detect abnormal sounds of preset categories in real time (such as explosion, glass breaking, violent arguments, calls for help, etc.) and generate event tags (initial identification data) containing event type, timestamp, and confidence level.

[0064] (3) Cooperative positioning calculation: When a single node detects an event or a cooperative instruction issued by the central layer, adjacent nodes calculate the signal arrival time difference through the GCC-PHAT (Generalized Cross Correlation with Phase Transform) algorithm, and combine it with the known node spatial coordinates to preliminarily estimate the location of the sound source.

[0065] In the specific implementation process, the collected raw audio data can first be denoised to remove environmental noise. Then, any two microphone nodes are selected, and the timing of their reception of the same sound signal is compared to determine the time difference. Finally, the distance difference between the two microphone nodes and the sound source is calculated using "time difference × speed of sound (approximately 340 m / s)". Combining the time difference information between multiple microphone nodes and the installation coordinates of each microphone node, the approximate location and distance of the sound source (initial positioning data) can be determined through geometric calculation. In addition, to overcome reverberation and multipath effects, this layer will upload an independently generated audio data packet from each microphone node participating in the collaborative positioning to the central layer. Each data packet contains three parts: raw audio data, preliminary judgment result (initial recognition data), and preliminary positioning data.

[0066] Among them, the triggering conditions for the central layer to issue instructions are roughly as follows: (1) Multi-node event reporting requires collaborative positioning: When multiple microphone nodes detect the same abnormal sound event simultaneously or sequentially, but because a single microphone node is too far away to complete reliable positioning independently, the central layer, after integrating the information reported from multiple sources, determines that multi-node collaborative positioning needs to be initiated and actively issues collaborative instructions to require multiple related microphone nodes (node ​​groups) to execute the GCC-PATH calculation process. (2) Continuous tracking and positioning during target movement: After the system enters the tracking phase, the position of the moving target needs to be continuously updated. The central layer predicts the area where the target object may be located at the current moment based on the historical (such as the previous moment) tracking results, actively wakes up the microphone node group around the area, and issues instructions to require it to start collaborative positioning calculation in order to achieve continuous position perception of the target object. (3) Cross-modal verification trigger: When the video analysis module detects suspicious behavior (such as people climbing over walls, abnormal gatherings, etc.), but the microphone nodes in the corresponding area do not actively report sound events, the central layer actively instructs the microphone node group in that area to start the above-mentioned audio acquisition and GCC_PATH collaborative positioning calculation in order to obtain audio evidence or assist visual positioning, so as to realize mutual verification of audio and video.

[0067] 2. Central Intelligent Collaboration Layer.

[0068] The central intelligent collaboration layer is the brain of the system, deployed on cloud or edge servers, and mainly includes the following core engines.

[0069] (1) Voiceprint region calculation engine: First, it receives multiple audio data packets from the edge computing layer, and then fuses the preliminary positioning coordinates reported by multiple microphone nodes according to the confidence level to obtain a fused coordinate point. Then, it uses a sound field propagation model based on three-dimensional ray tracing to correct the fused coordinates and obtain a "voiceprint active region" represented by a probability ellipse, rather than a single coordinate point, which is more in line with the positioning uncertainty in the real environment.

[0070] (2) Audio-visual dynamic mapping engine: Maintains the spatial database of all camera nodes in the system (including geographical location, installation height, orientation, focal length, field of view, PTZ capability, etc.). When a "soundprint active area" is received, the engine uses the correlation model to calculate the dynamic correlation between each camera and the area in real time, and generates a dynamic correlation camera list based on each dynamic correlation. Then, the cameras are sorted by correlation, and the one with the highest score is the first camera called by the system. The one with the second highest score is ready to take over at any time.

[0071] (3) Collaborative Strategy Engine: It has a built-in configurable strategy knowledge base that can generate specific collaborative tracking instruction sequences based on the type of voiceprint event (such as "rapidly moving distress call" or "fixed glass breakage") and a dynamically associated list of cameras.

[0072] (4) Device control and data fusion engine: Converts the instructions of the collaborative strategy engine into specific camera control protocols and sends them to the cameras. At the same time, it receives the video stream and analysis results (such as target boxes and trajectories) returned by the cameras, and performs spatiotemporal alignment and correlation with the initial judgment data of the voiceprint event (event type and timestamp) and the calculated probability ellipse "voiceprint active area" to generate a structured cross-modal event log (including fields such as event ID, initial judgment data, voiceprint active area, and video information) to facilitate subsequent retrieval, playback and evidence presentation. In the specific implementation process, time alignment is first performed. Based on the timestamp of the voiceprint event, video segments of N seconds (e.g., 30 seconds) before and after the time point are extracted from the video stream returned by the camera to ensure that the video evidence and the voiceprint event occur within the same time period. Then, spatial alignment is performed. The probability ellipse of the "voiceprint active area" is projected onto the video screen coordinate system and spatially matched with the field of view of the camera to determine the corresponding area in the video screen where the sound occurred. Finally, target association is performed. Moving targets appearing in the video screen at the time point and in the area are identified, and the target trajectory is bound to the voiceprint event to generate a "sound-target" association evidence chain.

[0073] 3. Video perception and control layer.

[0074] It consists of network cameras (including PTZ PTZ cameras and fixed bullet cameras) and a built-in lightweight video analysis module. In addition to executing control commands (rotation, zoom) issued by the central layer, it can also detect and track moving targets and transmit the tracking data back in real time.

[0075] 4. Data fusion application layer.

[0076] It stores all merged cross-modal event logs and can perform the following functions.

[0077] (1) Real-time visual command: Dynamically display the active area of ​​voiceprint, the field of view of the mobilized camera, and the trajectory of the moving target on the map.

[0078] (2) Efficient event retrieval: Supports joint retrieval by “sound event type”, “time” and “location”, and allows one-click access to all related audio and video clips.

[0079] (3) Post-event review and analysis: Provides a timeline tool that can replay the entire process of audio-visual linkage during the event.

[0080] As can be seen, the target tracking method provided in this embodiment of the invention has the following technical advantages: (1) Proactive early warning and proactive response: The system can be triggered before visually visible conflicts or damage occur (such as arguments escalating into fights or preparatory actions before breaking windows) through voiceprint perception, thus achieving proactive early warning; (2) Break through visual limitations and achieve seamless tracking across blind spots: When the target enters the visual blind spot covered by the acoustics from the field of view of camera A, the system can continuously locate the target based on the sound (footsteps, conversations) generated by the target in the blind spot, and intelligently wake up camera B on the other side of the blind spot to wait in advance, so as to achieve “seamless” relay tracking in vision, which greatly improves the ability to continuously control the target. (3) Significantly improve investigation and retrieval efficiency: Investigators can directly search through "specific voiceprint events". The system automatically associates and presents all relevant video clips and target trajectories, reducing post-event video screening from "hours" to "minutes", which can greatly improve efficiency; (4) Intelligent resource scheduling to improve system efficiency: Change the 24 / 7 indiscriminate recording to an active tracking mode that is "triggered by events and focused on a specific purpose", which saves network bandwidth, storage space and manpower in the monitoring center, and can maximize the efficiency of limited security resources; (5) Strong system adaptability and scalability: The collaborative strategy engine can be flexibly configured for different scenarios (such as the area around schools and city squares), and the voiceprint recognition model can also be continuously updated, with good scalability and scenario adaptability.

[0081] This invention provides a target tracking system.

[0082] like Figure 1 As shown, Figure 1 This is a schematic diagram of the structure of a target tracking system provided by the present invention. The target tracking system may include a central server 100, audio acquisition devices 200, and video acquisition devices 300. Each audio acquisition device 200 is communicatively connected to the central server 100, and each video acquisition device 300 is communicatively connected to the central server 100. The central server 100 is used to acquire audio data sent by each audio acquisition device 200. The audio data includes the original audio data of the target object as well as initial identification data and initial positioning data of the target object. The initial identification data and initial positioning data are fused to obtain the voiceprint event type and voiceprint active area of ​​the target object. The correlation degree between each video acquisition device 300 and the voiceprint active area is calculated, and the target video acquisition device is determined according to the correlation degree and voiceprint event type. The target video acquisition device is controlled to acquire the video data of the target object, and the video data, audio data and voiceprint active area are correlated to obtain the target tracking result of the target object.

[0083] As can be seen, the target tracking system provided in this embodiment of the invention deploys both audio and video acquisition devices in the target tracking area to combine computer vision technology with acoustic technology. In the implementation process, firstly, the audio data collected by each audio acquisition device is fused to determine the voiceprint event type and voiceprint active area of ​​the target object. Then, the correlation between each video acquisition device and the voiceprint active area is calculated and combined with the voiceprint event type of the target object to select the target video acquisition device. Finally, the target video acquisition device is used to acquire the video data of the target object and establish a correlation between the video data and the voiceprint active area to obtain the target tracking result. Therefore, this technical solution combines computer vision technology with acoustic technology to achieve target tracking. The audio acquisition device possesses characteristics such as being unaffected by light, having no visual obstruction, and being able to sense events outside the line of sight, which helps to achieve more timely and accurate target tracking, thereby achieving safer and more effective intelligent security.

[0084] For a description of the system provided in the embodiments of the present invention, please refer to the above method embodiments; the present invention will not be described in detail here.

[0085] This invention provides a target tracking device.

[0086] Please refer to Figure 5 , Figure 5 This is a schematic diagram of a target tracking device provided by the present invention. The target tracking device may include: The acquisition module 1 is used to acquire audio data sent by each audio acquisition device. The audio data includes the original audio data of the target object as well as the initial identification data and initial positioning data of the target object. Fusion module 2 is used to fuse each initial identification data and each initial positioning data to obtain the voiceprint event type and voiceprint active area of ​​the target object; Calculation module 3 is used to calculate the correlation between each video acquisition device and the active area of ​​the voiceprint, and to determine the target video acquisition device based on the correlation and the type of voiceprint event. Control module 4 is used to control the target video acquisition device to acquire video data of the target object, and to perform correlation processing on the video data, audio data, and active area of ​​the voiceprint to obtain the target tracking result of the target object.

[0087] As can be seen, the target tracking device provided in this embodiment of the invention deploys both audio and video acquisition devices in the target tracking area to combine computer vision technology with acoustic technology. In the implementation process, firstly, the audio data collected by each audio acquisition device is fused to determine the voiceprint event type and voiceprint active area of ​​the target object. Then, the correlation between each video acquisition device and the voiceprint active area is calculated and combined with the voiceprint event type of the target object to select the target video acquisition device. Finally, the target video acquisition device is used to acquire the video data of the target object and establish a correlation between the video data and the voiceprint active area to obtain the target tracking result. Therefore, this technical solution combines computer vision technology with acoustic technology to achieve target tracking. The audio acquisition device possesses characteristics such as being unaffected by light, having no visual obstruction, and being able to sense events outside the line of sight, which helps to achieve more timely and accurate target tracking, thereby achieving safer and more effective intelligent security.

[0088] In one embodiment of the present invention, the initial identification data includes an initial voiceprint event type and a confidence level corresponding to the initial voiceprint event type; accordingly, the fusion module 2 can be specifically used to determine the weight value of each initial voiceprint event type according to the confidence level corresponding to each initial voiceprint event type; and to perform weighted fusion calculation on each initial voiceprint event type according to the weight value of each initial voiceprint event type to obtain the voiceprint event type of the target object.

[0089] In one embodiment of the present invention, the fusion module 2 can be specifically used to determine the weight value of each initial positioning data according to the confidence level corresponding to each initial voiceprint event type; calculate the weighted fusion of each initial positioning data according to the weight value of each initial positioning data to obtain the fused positioning data of the target object; and correct the fused positioning data using a sound field propagation model to obtain the voiceprint active region of the target object; the voiceprint active region is based on elliptic probability representation.

[0090] In one embodiment of the present invention, the above-mentioned calculation module 3 can be specifically used to calculate the overlapping area between the screen area of ​​each video acquisition device and the active area of ​​the audioprint; calculate the distance between each video acquisition device and the active area of ​​the audioprint; and for each video acquisition device, calculate the correlation between the video acquisition device and the active area of ​​the audioprint based on the overlapping area, distance, resolution coefficient, and pan-tilt availability coefficient corresponding to the video acquisition device.

[0091] In one embodiment of the present invention, the above-mentioned calculation module 3 can be specifically used to query the video acquisition device scheduling rules corresponding to the voiceprint event type in the strategy knowledge base; arrange each correlation degree in a preset order to obtain a video acquisition device association list; and determine the target video acquisition device in the video acquisition device association list according to the video acquisition device scheduling rules.

[0092] In one embodiment of the present invention, the initial identification data includes an initial voiceprint event type and a timestamp corresponding to the initial voiceprint event type; accordingly, the control module 4 can be specifically used to calculate a fusion timestamp based on the timestamp corresponding to each initial voiceprint event type; extract audio segments corresponding to the fusion timestamp from each audio data; project the voiceprint active region onto each audio segment; identify and label the target object and the trajectory information of the target object within the projection area of ​​the voiceprint active region in each audio segment to obtain each target audio segment; bind each target audio segment to the voiceprint event type to obtain the target tracking result of the target object.

[0093] In one embodiment of the present invention, the acquisition module 1 described above may be specifically used to issue collaborative instructions to each audio acquisition device, so that each audio acquisition device can use an acoustic model to process the original audio data to determine the initial recognition data, and determine the initial positioning data based on the audio time difference and spatial position relationship between itself and the adjacent audio acquisition devices.

[0094] In one embodiment of the present invention, the acquisition module 1 is specifically used to send a coordination instruction to each audio acquisition device in the abnormal area when it receives an abnormal video alert from any video acquisition device and determines that it has not received audio data from each audio acquisition device in the abnormal area; wherein, the abnormal area is determined based on the abnormal video in the abnormal video alert.

[0095] In one embodiment of the present invention, the acquisition module 1 is specifically used to send a coordination instruction to each audio acquisition device in the target area where the abnormal sound source is located when it receives an abnormal audio alert sent by any audio acquisition device and determines that the distance between the audio acquisition device that sent the abnormal audio alert and the estimated location of the abnormal sound source exceeds a preset threshold; wherein, the estimated location of the abnormal sound source is determined based on the abnormal audio in the abnormal audio alert.

[0096] In one embodiment of the present invention, the acquisition module 1 can be specifically used to predict the future movement area of ​​the target object based on the historical tracking results of the target object, and issue collaborative instructions to each audio acquisition device in the future movement area.

[0097] For a description of the apparatus provided in the embodiments of the present invention, please refer to the above method embodiments; the present invention will not be described in detail here.

[0098] This invention provides an electronic device.

[0099] Please refer to Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided by the present invention. The electronic device may include: Memory 11 is used to store computer programs; The processor 10 is configured to execute computer programs to implement the steps of any of the target tracking methods described above.

[0100] like Figure 6 The diagram shows the structural composition of an electronic device, which may include a processor 10, a memory 11, a communication interface 12, and a communication bus 13. The processor 10, memory 11, and communication interface 12 all communicate with each other through the communication bus 13.

[0101] In this embodiment of the invention, the processor 10 may be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field-programmable gate array, or other programmable logic devices.

[0102] The processor 10 can call programs stored in the memory 11. Specifically, the processor 10 can execute operations in the embodiments of the target tracking method.

[0103] The memory 11 is used to store one or more programs. The programs may include program code, which includes computer operation instructions. In this embodiment of the invention, the memory 11 stores at least a program for implementing the following functions: The system acquires audio data sent by each audio acquisition device, including the original audio data of the target object as well as initial identification and initial positioning data. It then fuses each initial identification and initial positioning data to obtain the target object's voiceprint event type and active voiceprint region. The system calculates the correlation between each video acquisition device and the active voiceprint region, and determines the target video acquisition device based on the correlation and voiceprint event type. Finally, it controls the target video acquisition device to acquire the target object's video data and correlates the video data, audio data, and active voiceprint region to obtain the target tracking result.

[0104] In one possible implementation, the memory 11 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; and the data storage area may store data created during use.

[0105] In addition, memory 11 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device or other volatile solid-state storage device.

[0106] Communication interface 12 can be an interface for the communication module, used to connect with other devices or systems.

[0107] Of course, it should be noted that, Figure 6The structure shown does not constitute a limitation on the electronic device in the embodiments of the present invention. In practical applications, the electronic device may include more than Figure 6 More or fewer components as shown, or combinations of certain components.

[0108] This invention provides a computer-readable storage medium.

[0109] The computer-readable storage medium provided in this embodiment of the invention stores a computer program, which, when executed by a processor, can implement the steps of any of the target tracking methods described above.

[0110] The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a server or data center that integrates one or more available media. For example, it can be any medium that can store computer program code, such as magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives).

[0111] For a description of the computer-readable storage medium provided in the embodiments of the present invention, please refer to the above method embodiments; the present invention will not be described in detail here.

[0112] This invention provides a computer program product.

[0113] The computer program product provided in this embodiment of the invention includes a computer program / instruction, which, when executed by a processor, can implement the steps of any of the target tracking methods described above.

[0114] Specifically, in the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product.

[0115] The computer program product may include one or more computer programs / instructions, which, when loaded and executed on a computer, can generate all or part of the processes or functions described in the embodiments of the present invention. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line, etc.) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0116] For a description of the computer program product provided in the embodiments of the present invention, please refer to the above method embodiments; the present invention will not be described in detail here.

[0117] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0118] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0119] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0120] The technical solution provided by this invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make several improvements and modifications to this invention without departing from the principles of this invention, and these improvements and modifications also fall within the protection scope of this invention.

Claims

1. A target tracking method, characterized in that, include: Acquire audio data sent by each audio acquisition device, the audio data including the original audio data of the target object as well as the initial identification data and initial positioning data of the target object; The initial identification data and the initial positioning data are fused separately to obtain the voiceprint event type and voiceprint active area of ​​the target object; Calculate the correlation degree between each video acquisition device and the active area of ​​the voiceprint, and determine the target video acquisition device based on the correlation degree and the voiceprint event type; The target video acquisition device is controlled to acquire video data of the target object, and the video data, audio data, and active voiceprint region are correlated to obtain the target tracking result of the target object.

2. The target tracking method according to claim 1, characterized in that, The initial identification data includes the initial voiceprint event type and the confidence level corresponding to the initial voiceprint event type; Accordingly, the initial identification data are fused to obtain the voiceprint event type of the target object, including: The weight value of each initial voiceprint event type is determined based on the confidence level corresponding to each initial voiceprint event type. The initial voiceprint event types are weighted and fused according to their respective weight values ​​to obtain the voiceprint event type of the target object.

3. The target tracking method according to claim 2, characterized in that, The initial positioning data are fused to obtain the active voiceprint region of the target object, including: The weight value of each initial location data is determined based on the confidence level corresponding to each initial voiceprint event type. The weighted fusion calculation is performed on each of the initial positioning data according to the weight value of each initial positioning data to obtain the fused positioning data of the target object; The fused positioning data is corrected using a sound field propagation model to obtain the active voiceprint region of the target object; the active voiceprint region is based on elliptic probability representation.

4. The target tracking method according to claim 1, characterized in that, Calculate the correlation between each video acquisition device and the active area of ​​the acoustic signature, including: Calculate the overlap area between the screen area of ​​each video acquisition device and the active area of ​​the audioprint; Calculate the distance between each video acquisition device and the active area of ​​the acoustic signature; For each video acquisition device, the correlation between the video acquisition device and the active area of ​​the audioprint is calculated based on the overlapping area, distance length, resolution coefficient, and pan-tilt availability coefficient corresponding to the video acquisition device.

5. The target tracking method according to claim 1, characterized in that, Determining the target video acquisition device based on the aforementioned correlation degree and the aforementioned voiceprint event type includes: Search the video acquisition device scheduling rules corresponding to the aforementioned voiceprint event type in the strategy knowledge base; Arrange the aforementioned correlation degrees in a preset order to obtain a video acquisition device correlation list; The target video acquisition device is determined in the video acquisition device association list according to the video acquisition device scheduling rules.

6. The target tracking method according to any one of claims 1 to 5, characterized in that, The initial identification data includes the initial voiceprint event type and the timestamp corresponding to the initial voiceprint event type; Accordingly, the video data, the audio data, and the active voiceprint region are correlated to obtain the target tracking result of the target object, including: Calculate the fusion timestamp based on the timestamp corresponding to each of the initial voiceprint event types; Extract the audio segment corresponding to the fusion timestamp from each of the audio data; Project the active voiceprint region onto each of the audio segments; In the active voiceprint region, the target object and its trajectory information are identified and marked within the projection area of ​​each audio segment to obtain each target audio segment; By binding each target audio segment with the voiceprint event type, the target tracking result of the target object is obtained.

7. The target tracking method according to claim 1, characterized in that, Acquire audio data sent by each audio acquisition device, including: A collaborative instruction is sent to each of the audio acquisition devices, so that each audio acquisition device can use an acoustic model to process the raw audio data to determine the initial recognition data, and determine the initial positioning data based on the audio time difference and spatial position relationship between itself and the adjacent audio acquisition devices.

8. The target tracking method according to claim 7, characterized in that, Sending collaborative instructions to each of the aforementioned audio acquisition devices, including: When an abnormal video alert is received from any video capture device, and it is determined that no audio data has been received from any of the audio capture devices within the abnormal area, the coordination instruction is sent to each audio capture device within the abnormal area; wherein, the abnormal area is determined based on the abnormal video in the abnormal video alert.

9. The target tracking method according to claim 7, characterized in that, Sending collaborative instructions to each of the aforementioned audio acquisition devices, including: When an abnormal audio alert is received from any audio acquisition device, and it is determined that the distance between the audio acquisition device that sent the abnormal audio alert and the estimated location of the abnormal sound source exceeds a preset threshold, the coordination instruction is sent to each audio acquisition device within the target area where the estimated location of the abnormal sound source is located; wherein, the estimated location of the abnormal sound source is determined based on the abnormal audio in the abnormal audio alert.

10. The target tracking method according to claim 7, characterized in that, Sending collaborative instructions to each of the aforementioned audio acquisition devices, including: Based on the historical tracking results of the target object, the future movement area of ​​the target object is estimated, and the collaborative command is sent to each audio acquisition device within the future movement area.

11. A target tracking system, characterized in that, It includes a central server, various audio acquisition devices, and various video acquisition devices. Each audio acquisition device is communicatively connected to the central server, and each video acquisition device is communicatively connected to the central server. The central server is used to acquire audio data sent by each of the audio acquisition devices. The audio data includes the original audio data of the target object and the initial identification data and initial positioning data of the target object. The initial identification data and the initial positioning data are fused to obtain the voiceprint event type and voiceprint active area of ​​the target object; the correlation degree between each video acquisition device and the voiceprint active area is calculated, and the target video acquisition device is determined according to the correlation degree and the voiceprint event type; the target video acquisition device is controlled to acquire the video data of the target object, and the video data, the audio data, and the voiceprint active area are correlated to obtain the target tracking result of the target object.

12. A target tracking device, characterized in that, include: The acquisition module is used to acquire audio data sent by each audio acquisition device. The audio data includes the original audio data of the target object and the initial identification data and initial positioning data of the target object. The fusion module is used to perform fusion processing on each of the initial identification data and each of the initial positioning data to obtain the voiceprint event type and voiceprint active area of ​​the target object; The calculation module is used to calculate the correlation degree between each video acquisition device and the active area of ​​the voiceprint, and to determine the target video acquisition device based on the correlation degree and the voiceprint event type. The control module is used to control the target video acquisition device to acquire video data of the target object, and to perform correlation processing on the video data, the audio data, and the active area of ​​the voiceprint to obtain the target tracking result of the target object.

13. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the target tracking method as described in any one of claims 1 to 10.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the target tracking method as described in any one of claims 1 to 10.

15. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the target tracking method according to any one of claims 1 to 10.