Target tracking method and system based on artificial intelligence, and storage medium

By integrating thermal imaging with visual data to reconstruct occluded target outlines, the method addresses the issue of tracking interruptions and errors in complex scenarios, ensuring accurate and continuous target tracking.

CN120321491APending Publication Date: 2025-07-15SHENZHEN SHARE VISION CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510451243.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Current target tracking methods in video surveillance rely solely on visual data, which fail to accurately determine the outline of obscured targets, leading to tracking interruptions and errors due to incomplete or incorrect data when targets are occluded.

Method used

Integrate thermal imaging data with visual data to reconstruct occluded target outlines by aligning the temporal and spatial relationships between the two data types, overlaying the thermal data on the visual data to provide complete target outlines.

Benefits of technology

Ensures continuous and accurate target tracking by complementing visual data with thermal information, enhancing the visual representation of obscured targets and improving the overall accuracy and continuity of target tracking in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321491A_ABST
    Figure CN120321491A_ABST
Patent Text Reader

Abstract

The invention discloses a target tracking method and system based on artificial intelligence, and a storage medium, relates to the technical field of video monitoring, and discloses a target tracking method based on artificial intelligence, and the method comprises the steps: determining a human body contour of a tracking object in visual data collected by a target camera in response to a target tracking process triggering instruction; if it is detected that the shielding area appears in the human body contour, thermal imaging data collected by a target camera are obtained; according to the time-space relationship between the visual data and the thermal imaging data and the thermal imaging data, generating human body contour supplementary information corresponding to the occlusion area; in the display interface of the visual data, the human body contour supplementary information and the visual data are displayed at the same time, and the human body contour supplementary information is displayed on the upper layer of the visual data. According to the method, the video data and the thermal imaging data are combined, the human body contour information is supplemented when the target is shielded, the problem of tracking interruption or misjudgment caused by shielding is solved, and the monitoring accuracy and continuity are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video surveillance technology, and particularly to an object tracking method, system and storage medium based on artificial intelligence. Background Art

[0002] Object tracking technology is one of the core technologies in the field of video surveillance, aiming to continuously locate and analyze the trajectory of specific objects (such as pedestrians, vehicles, etc.) in a video through algorithms. However, in complex scenarios, the object may be occluded (such as being occluded by other objects or crowds), resulting in partial or complete loss of the visual features of the object, thus affecting the accuracy and continuity of tracking. Current object tracking methods mainly rely on visual data collected by cameras and detect and track objects through deep learning models. However, when the object is occluded, it is difficult to accurately infer the contour information of the occluded part only relying on visual data, resulting in the tracking algorithm may lose the object or make misjudgments.

[0003] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide an object tracking method, system and storage medium based on artificial intelligence, aiming to solve the technical problem of tracking interruption or misjudgment caused by object occlusion.

[0005] To achieve the above object, this application proposes an object tracking method based on artificial intelligence, and the method includes:

[0006] In response to an object tracking process trigger instruction, determine the human contour of the tracking object in the visual data collected by the target camera;

[0007] If an occlusion area is detected in the human contour, obtain the thermal imaging data collected by the target camera;

[0008] Generate human contour supplementary information corresponding to the occlusion area according to the spatio-temporal relationship between the visual data and the thermal imaging data, and the thermal imaging data;

[0009] In the display interface of the visual data, simultaneously display the human contour supplementary information and the visual data, wherein the human contour supplementary information is displayed on the upper layer of the visual data.

[0010] In one embodiment, the step of generating human contour supplementary information corresponding to the occlusion area according to the spatio-temporal relationship between the visual data and the thermal imaging data, and the thermal imaging data includes:

[0011] Determine the human body area corresponding to the tracking object in the thermal imaging images corresponding to each moment according to the spatio-temporal relationship;

[0012] Generate a thermal imaging human body contour according to the human body area in the thermal imaging image;

[0013] Determine the supplementary information of the human body contour according to the thermal imaging human body contour, the human body contour and the occlusion area.

[0014] In one embodiment, after the step of simultaneously displaying the supplementary information of the human body contour and the visual data in the display interface of the visual data, the method further includes:

[0015] Input the visual data into a scene recognition model, and obtain a scene recognition result through the scene recognition model;

[0016] Input the scene recognition result and the visual data into a behavior analysis model to obtain a behavior description text.

[0017] In one embodiment, the step of inputting the scene recognition result and the visual data into a behavior analysis model to obtain a behavior description text includes:

[0018] In response to the input data, the behavior analysis model loads the atomic event recognition rules corresponding to the target scene according to the scene recognition result;

[0019] Generate atomic events according to the atomic event recognition rules and the visual data;

[0020] Generate composite events based on the spatio-temporal combination of at least two of the atomic events;

[0021] Generate the behavior description text based on the composite events and predefined sentence patterns.

[0022] In one embodiment, the step of generating atomic events according to the atomic event recognition rules and the visual data includes:

[0023] Generate the movement trajectory of the tracking object according to the spatio-temporal position relationship between the tracking object and a preset reference object in the visual data, and the pre-stored earth coordinates associated with the preset reference object;

[0024] Determine the recognition rules corresponding to the stay event, the area switching event and / or the sensitive area entry event according to the atomic event recognition rules;

[0025] When it is detected that the movement trajectory matches the recognition rules, generate the corresponding atomic events.

[0026] In one embodiment, after the step of inputting the scene recognition result and the visual data into the behavior analysis model to obtain the behavior description text, the method further includes:

[0027] Input the behavior description text into a natural language processing model to obtain the behavior characteristics corresponding to the tracking object;

[0028] Determine the similarity between the behavior characteristics and the historical abnormal behavior characteristics;

[0029] If the similarity is greater than or equal to a preset threshold, output an abnormal behavior alarm.

[0030] In one embodiment, after the step of simultaneously displaying the human body contour supplementary information and the visual data in the display interface of the visual data, the method further includes:

[0031] If it is detected that the tracking object leaves the monitoring range of the target camera;

[0032] Determine the moving trend of the tracking object according to the moving trajectory of the tracking object within a preset time period before leaving the monitoring range;

[0033] Determine the handover camera according to the moving trend;

[0034] Set the handover camera as the target camera and generate a target tracking process trigger instruction;

[0035] Jump to execute the step of determining the human body contour of the tracking object in the visual data collected by the target camera.

[0036] In one embodiment, before the step of determining the human body contour of the tracking object in the visual data collected by the target camera, the method further includes:

[0037] In response to a control operation received on the monitoring video display interface, determine the tracking object and generate the target tracking process trigger instruction.

[0038] In addition, to achieve the above object, the present application also proposes an artificial intelligence-based target tracking system, the system includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program is configured to implement the steps of the artificial intelligence-based target tracking method as described above.

[0039] In addition, to achieve the above object, the present application also proposes a storage medium, the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the artificial intelligence-based target tracking method as described above.

[0040] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the above-mentioned artificial intelligence-based target tracking method.

[0041] One or more technical solutions proposed by the present application have at least the following technical effects:

[0042] When the present application detects an occlusion area in the human body contour in visual data, it synchronously acquires the thermal imaging data collected by the target camera, and generates human body contour supplementary information corresponding to the occlusion area based on the spatio-temporal relationship between the visual data and the thermal imaging data, solving the problem that traditional target tracking methods may cause tracking interruption or misjudgment due to relying on single visual data in case of occlusion. Thermal imaging technology can penetrate visual occluders and capture the human body's thermal radiation characteristics. Through spatio-temporal alignment and feature fusion with visual data, it ensures that physical occlusion no longer completely blocks the acquisition of the target contour. By supplementing the occlusion area in visual data with thermal imaging data, the present application realizes the restoration of the occluded part of the contour information by superimposing it on the display interface. Compared with traditional technologies, the present application can still maintain continuous and accurate human target tracking in a complex occlusion environment. At the same time, by superimposing the supplementary information on the upper layer of the visual data, it ensures the user's synchronous perception of the original scene and the supplementary information, improving the intuitiveness of the video surveillance system and the accuracy of target tracking. Description of the Drawings

[0043] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0044] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 Schematic flowchart provided for the first embodiment of the artificial intelligence-based target tracking method of the present application;

[0046] Figure 2 Schematic flowchart provided for the second embodiment of the artificial intelligence-based target tracking method of the present application;

[0047] Figure 3 Schematic flowchart provided for the third embodiment of the artificial intelligence-based target tracking method of the present application;

[0048] Figure 4Schematic flowchart provided for the fourth embodiment of the object tracking method based on artificial intelligence in this application;

[0049] Figure 5 Schematic flowchart provided for the fifth embodiment of the object tracking method based on artificial intelligence in this application;

[0050] Figure 6 Schematic diagram of the structure of the hardware operating environment involved in the object tracking system based on artificial intelligence in the embodiments of this application.

[0051] The implementation, functional features, and advantages of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0052] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0053] Object tracking technology is one of the core technologies in the field of video surveillance, aiming to continuously locate and analyze the trajectories of specific objects (such as pedestrians, vehicles, etc.) in videos through algorithms. However, in complex scenarios, the object may be occluded (such as being occluded by other objects or crowds), resulting in partial or complete loss of the visual features of the object, thus affecting the accuracy and continuity of tracking. Current object tracking methods mainly rely on the visual data collected by cameras to detect and track objects through deep learning models. However, when the object is occluded, it is difficult to accurately infer the contour information of the occluded part only relying on visual data, resulting in the tracking algorithm may lose the object or produce misjudgments.

[0054] To better understand the technical solutions of this application, the following will be described in detail in combination with the accompanying drawings of the specification and specific implementation manners.

[0055] The main solution of the embodiments of this application is: in response to an object tracking process trigger instruction, determine the human contour of the tracking object in the visual data collected by the target camera; if an occlusion area appears in the detected human contour, obtain the thermal imaging data collected by the target camera; generate human contour supplementary information corresponding to the occlusion area according to the spatio-temporal relationship between the visual data and the thermal imaging data, and the thermal imaging data; in the display interface of the visual data, simultaneously display the human contour supplementary information and the visual data, where the human contour supplementary information is displayed on the upper layer of the visual data.

[0056] In this embodiment, for the sake of easy description, the following will be described with an object tracking system based on artificial intelligence (hereinafter referred to as the "system") as the execution subject. The system includes an object camera, a data fusion module, and an enhanced display module. The object camera can collect visual data and thermal imaging data simultaneously. Among them, the visual data is obtained by a visible light acquisition component, and the thermal imaging data is collected by a built-in thermal imaging acquisition component.

[0057] Since there are technical limitations such as missing contour information in occlusion scenarios when only relying on visual data for object tracking in the current technology, which leads to tracking interruption or misjudgment, etc., this application provides a solution: fusing visual data and thermal imaging data to achieve the reconstruction of the object contour in the occluded area. In the specific implementation process, a synchronous trigger acquisition device is used to minimize the time alignment error of the data of the object camera as much as possible. At the same time, based on a feature fusion network with an attention mechanism, spatial weighted fusion is performed on the human body contour segments and the thermal imaging area, automatically generating human body contour supplementary information and superimposing it on the upper layer of the visual data, so as to provide a display interface that can display the superimposed effect of the human body contour.

[0058] Specifically, the system first detects the human body bounding box in the visual data, and then makes an occlusion determination on the human body bounding box. For example, if in consecutive video frames, the pixel confidence of a certain area within the human body bounding box drops sharply from a relatively high value to nearly zero and remains low in subsequent video frames, the system determines that occlusion occurs in this area. At this time, the system obtains the thermal imaging data collected by the object camera, and fuses the visual data and the thermal imaging data through the data fusion module. Using the thermal radiation characteristics of the human body in the thermal imaging data and combining the existing human body contour information in the visual data, spatial weighted fusion is performed through the feature fusion network to generate human body contour supplementary information for the occluded area. Finally, the enhanced display module superimposes and displays the generated human body contour supplementary information in the form of a semi-transparent or specific color contour on the display interface of the visual data, providing users with complete and intuitive object contour information and improving the accuracy of object tracking in occlusion scenarios.

[0059] As can be seen from the above embodiments, since this application adopts the technical means of fusing visual data and thermal imaging data and analyzing their spatio-temporal relationships to generate human body contour supplementary information, it overcomes the technical problems of tracking interruption or misjudgment when the object is occluded in the current technology, thus achieving the technical effect of improving the accuracy and integrity of object tracking in complex scenarios.

[0060] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, data fusion, and program running functions, such as an edge computing terminal, an intelligent camera, a personal computer, etc. in a video surveillance system, or an electronic device, a surveillance system, etc. that can implement the above functions. Hereinafter, taking a video surveillance system as an example (hereinafter referred to as "system"), this embodiment and the following embodiments will be described.

[0061] Based on this, an embodiment of the present application provides an object tracking method based on artificial intelligence. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the object tracking method based on artificial intelligence of the present application.

[0062] In this embodiment, the object tracking method based on artificial intelligence includes steps S10 to S40:

[0063] Step S10, in response to a target tracking process trigger instruction, determine the human contour of the tracking object in the visual data collected by the target camera.

[0064] It should be noted that the target tracking process trigger instruction can be manually issued by the user on the system operation interface, or automatically generated by a specific event preset in the system. For example, when the system detects an abnormal behavior or a person breaks into a specific area, the target tracking program is automatically started. After receiving the target tracking process trigger instruction, the system will call the visual data collected by the target camera. The target camera is a device deployed in the monitoring area for collecting video images, and the visual data collected by the target camera is used to provide basic image information for subsequent target tracking.

[0065] It can be understood that in actual application scenarios, it is necessary to quickly and accurately track a specific target. If the system fails to respond to the target tracking process trigger instruction in time and determine the human contour of the tracking object, it may miss the best tracking opportunity due to processing irrelevant picture content. Therefore, step S10 can avoid ineffective data processing, thereby improving the timeliness and accuracy of target tracking.

[0066] In a feasible implementation manner, before step S10, step S101 may be included:

[0067] Step S101, in response to a control operation received on the monitoring video display interface, determine the tracking object and generate the target tracking process trigger instruction.

[0068] It should be noted that the monitoring video display interface is a user interaction interface in the video monitoring system for real-time displaying the video images captured by the target camera. This interface displays the visual data captured by the target camera in real time and provides human-computer interaction controls (such as selection tools, touch gestures, voice commands, etc.) for users to operate. The control operation refers to the operations performed by the user on the monitoring video display interface, such as using a mouse or touch screen to click on the preset tracking button, selecting the target, inputting commands, etc., to specify the target object to be tracked; or the system automatically triggers the target tracking program according to the preset rules to specify the target object to be tracked. When the monitoring video display interface receives the corresponding control operation, the system will determine the specific tracking object according to the control operation and generate a corresponding target tracking process trigger command.

[0069] In addition, it should be noted that the target tracking process trigger command is a signal used to notify the system to start the target tracking program, and it contains relevant information about the tracking task, such as the initial position of the tracking object, the human body contour features, etc.

[0070] In this embodiment, by receiving the control operation on the monitoring video display interface to determine the tracking object and generating the target tracking process trigger command, it can meet the user's need to track specific targets, avoid the system from tracking all targets indiscriminately, and thus improve the pertinence and effectiveness of target tracking.

[0071] Step S20, if it is detected that the human body contour has an occlusion area, obtain the thermal imaging data collected by the target camera;

[0072] In the traditional tracking method based on pure visual data, when the target is occluded, the corresponding target contour information in the visual data will be lost, resulting in the system being unable to accurately identify and continue to track the target, thus causing tracking interruption or misjudgment. To solve this problem, in this embodiment, when the system detects that the human body contour has an occlusion area, it will trigger the target camera to collect thermal imaging data. The thermal imaging data is obtained through thermal imaging technology. The thermal imaging acquisition component in the target camera can sense the infrared rays emitted by the object and convert them into images. Since the human body itself emits a certain amount of heat, even if it is visually occluded, the thermal signal formed by the heat emitted by the human body can still be captured by the thermal imaging acquisition component. Therefore, the thermal imaging data provides an additional tracking basis for the system, effectively making up for the deficiency of visual data.

[0073] It should be noted that in this embodiment, the target camera is only triggered to collect thermal imaging data when it is detected that the human body contour in the visual data has an occlusion area, effectively reducing the processing overhead of the system.

[0074] Step S30: Generate the supplementary human body contour information corresponding to the occluded area according to the spatio-temporal relationship between the visual data and the thermal imaging data, and the thermal imaging data.

[0075] It should be noted that the spatio-temporal relationship refers to the corresponding association between the visual data and the thermal imaging data in terms of time and space. Specifically, the spatial relationship is reflected in the description of the target position in the same scene by the visual data and the thermal imaging data, while the time relationship is reflected in the continuity of the target motion trajectory in consecutive frames.

[0076] In a feasible implementation manner, the system realizes the time alignment of the visual data and the thermal imaging data through timestamps. After the visual data and the thermal imaging data are collected, the system adds timestamps to each frame of the visual data and the thermal imaging data respectively, and the timestamps record the acquisition moment of each frame of data. The system finds the frames with the closest time by comparing the timestamps of the visual data and the thermal imaging data. For example, when the timestamp of a certain frame of visual data is detected as T1, the system will look for the frame with the closest timestamp to T1 in the thermal imaging data. If the timestamp of a frame in the thermal imaging data is T2 and |T1 - T2| is less than the preset time error threshold, it is considered that these two frames of data are time-aligned.

[0077] In another feasible implementation manner, when the system acquires the visual data and the thermal imaging data by the target camera, it uses a synchronous clock device. The synchronous clock device provides a unified clock signal for the visible light acquisition component and the thermal imaging acquisition component to ensure that the two acquisition components are strictly synchronized when acquiring data. For example, at the beginning of each acquisition cycle, the synchronous clock device simultaneously sends trigger signals to the visible light acquisition component and the thermal imaging acquisition component to make them start acquiring data at the same time. In this way, the system ensures the time consistency of the visual data and the thermal imaging data during data acquisition, reduces the time alignment error in post-processing, and improves the processing efficiency of the system.

[0078] In yet another feasible implementation manner, the system realizes the spatial alignment of the visual data and the thermal imaging data through a spatial calibration algorithm. Since there may be differences in resolution and scale between the thermal imaging data and the visual data, the system first uses a calibration tool (such as a checkerboard) to acquire the visual data and the thermal imaging data in the same scene respectively. Through image processing technology, the system calculates the geometric transformation parameters between the visual data coordinate system and the thermal imaging data coordinate system, including scaling, rotation, and translation parameters. Using the calculated geometric transformation parameters, the system maps the thermal imaging data to the same coordinate system as the visual data to ensure the spatial consistency of the visual data and the thermal imaging data.

[0079] Furthermore, the system extracts the edge features of the human body contour in the unoccluded part of the visual data through an edge detection algorithm. Meanwhile, in the thermal imaging data, according to the thermal radiation distribution of the human body, the contour features of the human body's thermal region are identified. Among them, both the edge features and the contour features are represented by a sequence of coordinate points. Then, the system uses a feature point matching algorithm, such as the SIFT (Scale-invariant feature transform) algorithm, to find the corresponding relationship between the edge feature points of the visual data and the contour feature points of the thermal imaging data. By detecting the key points in the visual data and the thermal imaging data, and calculating the feature descriptors of the neighborhoods around the key points, the matching point pairs are determined based on the similarity of the feature descriptors. These matching point pairs will be used for subsequent contour stitching.

[0080] First, the system converts the coordinates of the contour feature points in the occluded area of the thermal imaging to the same coordinate system as the visual data according to the previously established spatio-temporal alignment relationship. Then, the converted thermal imaging contour features are stitched with the contour edge features of the unoccluded part in the visual data. During the stitching process, if there is a discontinuity problem at the junction, the system can use an interpolation algorithm to construct a smooth curve passing through the given data points, and calculate the coordinates of the intermediate transition points based on the contour feature points on both sides of the junction, so that the thermal imaging contour and the visual contour are naturally connected, and finally generate the human body contour supplementary information corresponding to the occluded area.

[0081] It can be understood that in a complex target tracking scenario, the target may cause partial or complete loss of its visual features due to occlusion. When the target is occluded, relying solely on single visual data will result in the lack of target contour information, affecting the accuracy and continuity of tracking. However, the thermal imaging data can capture the target thermal signal information that still exists even when occluded, and there is a correlation between the visual data and the thermal imaging data in terms of time and space. Therefore, through the spatio-temporal relationship between the visual data and the thermal imaging data, combined with the thermal imaging data, the human body contour supplementary information corresponding to the occluded area can be generated, thus avoiding the incomplete human body contour information caused by occlusion and improving the accuracy of target tracking in the case of occlusion.

[0082] Step S40, in the display interface of the visual data, display both the human body contour supplementary information and the visual data, where the human body contour supplementary information is displayed on the upper layer of the visual data.

[0083] It should be noted that the display interface of visual data is the window for users to view the captured images of target cameras, usually composed of a Graphical User Interface (GUI), and has functions such as display and interaction. Through this display interface, users can observe the situation of the monitored area in real time. The human body contour supplementary information is a visual presentation of the shape, position, etc. of the occluded part of the human body obtained after fusing the thermal imaging data and visual data, and is displayed in the form of contour lines, filled areas, etc.

[0084] It can be understood that during the target tracking process, the occlusion of the target makes it difficult to obtain complete information about the tracked object only through visual data, which may lead to tracking interruption or misjudgment. And the human body contour supplementary information can provide the display content of the occluded part, so this embodiment can prevent users from being unable to fully understand the state of the tracked object due to image occlusion, and improves the visualization effect of target tracking.

[0085] In this embodiment, the system superimposes the human body contour supplementary information on the display interface of the visual data in a manner with adjustable transparency. Users can adjust the transparency of the human body contour supplementary information according to actual needs to more intuitively observe the complete contour of the tracked object.

[0086] Based on the first embodiment of this application, in the second embodiment of this application, for the same or similar content as the above first embodiment, reference can be made to the above introduction and will not be elaborated hereinafter. On this basis, please refer to Figure 2 , step S30 includes steps S31 to S33:

[0087] Step S31: Determine the human body area corresponding to the tracked object in the thermal imaging images corresponding to each moment according to the spatio-temporal relationship;

[0088] It should be noted that the thermal imaging image refers to the image generated from the thermal imaging data captured by the target camera, which can reflect the thermal radiation distribution of the object. Since thermal radiation can penetrate some occluders, compared with visual data, the thermal imaging data can still provide human body contour information under occlusion.

[0089] In this embodiment, the system ensures the time alignment of visual data and thermal imaging data through a timestamp or a synchronous clock device. Then, the system obtains the geometric transformation parameters between the coordinate system of the thermal imaging data and the coordinate system of the visual data through a spatial calibration algorithm, so as to map the thermal imaging data into the same coordinate system as the visual data. After completing the spatio-temporal alignment of the visual data and the thermal imaging data, the system matches the tracked object in the visual data with the corresponding area in the thermal imaging image according to the spatio-temporal relationship. Specifically, the system analyzes the thermal signal intensity in the thermal imaging image and uses a threshold segmentation algorithm to identify the area with a temperature higher than the preset temperature threshold in the thermal imaging image as the potential human body area. Then, the system matches the feature points of the tracked object in the visual data with the feature points of the potential human body area in the thermal imaging image through a feature point matching algorithm, so as to determine the corresponding human body area of the tracked object in the thermal imaging image.

[0090] Step S32: Generate a thermal imaging human body contour according to the human body area in the thermal imaging image;

[0091] It should be noted that the thermal imaging human body contour refers to the contour information of the tracked object extracted by analyzing the human body area in the thermal imaging image.

[0092] In addition, it should be noted that the purpose of generating the thermal imaging human body contour is to supplement the occluded part in the visual data. The thermal imaging data can provide information different from the visual data. Especially in the case of occlusion, the thermal imaging data can provide the thermal radiation contour of the target, thus helping the system to more accurately infer the complete contour of the target.

[0093] In this embodiment, the system can use an edge detection algorithm (such as the Canny operator) to process the human body area in the thermal imaging image and extract the contour features of the human body.

[0094] Step S33: Determine the human body contour supplementary information according to the thermal imaging human body contour, the human body contour and the occluded area.

[0095] It should be noted that the human body contour supplementary information refers to the additional contour information generated by combining the thermal imaging human body contour and the human body contour in the visual data to supplement the occluded area. The occluded area refers to the part of the target that cannot be observed in the visual data.

[0096] In this embodiment, the visual data and the thermal imaging data have been spatio-temporally aligned, and the thermal imaging human contour has a corresponding relationship with the human contour in the visual data in the spatial coordinate system. The system calculates the translation, rotation, and scaling parameters of the thermal imaging human contour relative to the human contour in the visual data by comparing the key points (such as the vertex of the head, the endpoints of the shoulders, etc.) on the thermal imaging human contour and the human contour in the visual data. Then, the system performs a geometric transformation on the thermal imaging human contour according to these parameters to make the thermal imaging human contour as aligned as possible with the human contour in the visual data.

[0097] Next, based on the thermal imaging human contour aligned with the human contour in the visual data, the system identifies the area of the thermal imaging human contour corresponding to the occluded part of the human contour in the visual data, so as to determine the human contour supplementary information.

[0098] It can be understood that by combining the thermal imaging human contour and the human contour in the visual data to generate the human contour supplementary information, it can effectively supplement the additional contour information of the occluded area, avoid the target failure caused by occlusion, prevent the target tracking from being interrupted or misjudged, and improve the continuity of target tracking.

[0099] Based on the above embodiments of the present application, in the third embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , after step S40, the artificial intelligence-based target tracking method further includes steps S41 to S42:

[0100] Step S41: Input the visual data into the scene recognition model, and obtain the scene recognition result through the scene recognition model;

[0101] It should be noted that the scene recognition model is used to identify the scene type in the video data, such as a school playground, a shopping mall, etc. The scene recognition model is trained by a large number of labeled scene images to learn the features of different scenes, so as to accurately classify the input visual data. The scene recognition result refers to the scene classification information output by the scene recognition model about the video data, such as a clear scene description like "school playground".

[0102] It can be understood that in the target tracking scenario, the behavior and events of the target may be affected by the scene type, and it is difficult to accurately understand the behavior of the target only relying on the visual data. For example, in the scene of a school playground, running is a normal behavior; but in the scene of a hospital corridor, running may pose a safety hazard. Therefore, by analyzing the visual data through the scene recognition model, it can provide scene context information for subsequent behavior analysis, so as to more accurately understand the behavior of the target and make corresponding judgments.

[0103] In this embodiment, different scenarios correspond to different event points of interest. Event points of interest refer to behaviors or events that the system needs to pay special attention to and analyze in specific scenarios. These event points of interest are pre-set according to the functions, rules and actual usage requirements of the scenarios, with the purpose of focusing on behaviors or events with special significance in specific scenarios, so that the system can analyze and judge the target behaviors more specifically.

[0104] In an optional implementation, the event points of interest can be identified through expert experience to identify key events or behaviors that need to be paid attention to in various common scenarios. Alternatively, by collecting a large amount of historical video data in different scenarios, statistical analysis of various events occurring in each scenario can be performed to mine frequent and representative events in specific scenarios, thereby determining the event points of interest corresponding to each scenario.

[0105] For example, in a shopping mall scenario, the focus event points can be set as "children alone", "physical conflict", "crowds running", and "abnormal stay in the valuables area". Among them, "children alone" means that there is a child missing; "physical conflict" means that there is a dispute between customers; "crowds running" means that there is an emergency, such as fire or violence; "abnormal stay in the valuables area" means that there is theft or scouting.

[0106] Step S42: input the scene recognition result and the visual data into a behavior analysis model to obtain a behavior description text.

[0107] It should be noted that the behavior analysis model is a model used to analyze the target behavior in the video data. The model can identify the target's behavior pattern, such as walking, running, gathering, etc., and generate corresponding behavior description text. The behavior description text refers to the natural language description of the target behavior output by the behavior analysis model, such as "tracking target A stayed in area P1 for 3 minutes and then moved to area P2." This behavior description text can provide users with intuitive behavior information, making it easier to understand whether the target's behavior conforms to the normal behavior pattern in the corresponding scenario.

[0108] In a feasible implementation, step S42 may include steps S421 to S424:

[0109] Step S421: In response to the input data, the behavior analysis model loads the atomic event recognition rules corresponding to the target scene according to the scene recognition result;

[0110] It should be noted that atomic events refer to the basic behavior units that constitute complex behaviors, such as "stay", "walk", "wander", "run", etc. Atomic event recognition rules refer to the rules and algorithms used to recognize these basic behavior units, usually based on the target's motion characteristics, posture characteristics, etc.

[0111] Step S422: Generate atomic events according to the atomic event recognition rule and the visual data;

[0112] It should be noted that generating atomic events refers to the process of extracting basic behavior units from visual data based on the atomic event recognition rule.

[0113] Step S423: Generate compound events based on the spatio-temporal combinations of at least two of the atomic events;

[0114] Step S424: Generate the behavior description text based on the compound event and a predefined sentence pattern.

[0115] It should be noted that spatio-temporal combination means considering the relevance of atomic events in terms of time and space. For example, a certain tracked target continuously has multiple atomic events in a short period of time, or multiple tracked targets have related atomic events at the same location.

[0116] It should be noted that a compound event is a complex behavior event composed of multiple atomic events combined in time and space. The generated atomic events are the basis for generating compound events, and the accuracy and integrity of atomic events will affect the quality of the finally generated behavior description text.

[0117] In addition, it should be noted that the behavior description text refers to the finally generated natural language description of the target behavior, providing intuitive behavior information of the tracked target. The predefined sentence pattern refers to a pre-designed natural language template for describing behavior, such as "A person performs [behavior] at [location] and transfers from [location] to [location] within [time]". The predefined sentence pattern can be adjusted according to different scenarios and behavior types.

[0118] It can be understood that since the defined focus event points may be different in different scenarios, and complex behaviors are usually composed of multiple basic behaviors, it is difficult to accurately describe the behavior pattern of the tracked target only by using the atomic event recognition rule to extract the atomic events of the target. Therefore, generating compound events through spatio-temporal combination and combining with predefined sentence patterns can generate accurate and easy-to-understand behavior description texts, so as to more comprehensively reflect the behavior pattern of the target and improve the accuracy of behavior analysis.

[0119] In another feasible implementation manner, step S422 may include steps S4221 to S4223:

[0120] Step S4221: Generate the movement trajectory of the tracked object according to the spatio-temporal position relationship between the tracked object and a preset reference object in the visual data, and the pre-stored earth coordinates associated with the preset reference object;

[0121] It should be noted that the preset reference object refers to an object or identifier with a fixed position and easy to identify in the target tracking scenario, such as a pillar in a shopping mall, a flagpole on a school playground, a street lamp on a street, etc. These preset reference objects are relatively fixed in position in the corresponding scenario, facilitating the use as a reference for determining the position of the tracking object. The pre-stored geodetic coordinates refer to the precise coordinate information of the preset reference object obtained and stored in advance through professional measurement means in the geodetic coordinate system. The geodetic coordinate system is a standard coordinate system used to determine the position of objects on the earth.

[0122] In addition, it should be noted that the spatio-temporal position relationship between the tracking object and the preset reference object in the visual data is obtained by processing the visual data, including information such as the distance, angle of the tracking object relative to the preset reference object, and the relative position change at different time points.

[0123] It can be understood that since an accurate movement trajectory is crucial for subsequent behavior analysis and event judgment, by using the preset reference object and the pre-stored geodetic coordinates to generate the movement trajectory of the tracking object, it is possible to avoid the error accumulation caused by solely relying on the relative position judgment of visual data, thereby improving the accuracy of movement trajectory generation.

[0124] Step S4222: Determine the recognition rules corresponding to the stay event, region switching event, and / or sensitive area entry event according to the atomic event recognition rules;

[0125] It should be noted that the atomic event recognition rules are pre-set rules and algorithms for recognizing basic behavior units based on the motion characteristics, posture characteristics, etc. of the target, and can be used as the criteria and conditions for judging whether a specific event occurs. The atomic event recognition rules are obtained through the analysis and summary of common behaviors and events in different scenarios. For example, in different scenarios, there are different criteria for the length of stay time and the definition of region boundaries. The stay event refers to the situation where the tracking object stays at a specific position or region for a time exceeding the preset stay threshold. The region switching event refers to the behavior of the tracking object moving from one region to another. The sensitive area entry event refers to the situation where the tracking object enters a pre-set area that needs to be focused on, such as a warehouse in a shopping mall, a laboratory in a school, etc.

[0126] In addition, it should be noted that there are mutual associations and influences among different atomic event recognition rules. For example, when judging the region switching event, factors such as the stay time of the tracking object in the original region may need to be considered simultaneously.

[0127] It can be understood that the criteria for judging the occurrence of events may vary in different scenarios. By clarifying the recognition rules corresponding to different atomic events, misjudgment and missed judgment of events in different scenarios can be avoided, thereby improving the accurate analysis of target behaviors and the ability to recognize events, and promptly discovering potential problems and risks.

[0128] Step S4223: When it is detected that the movement trajectory matches the recognition rule, generate the corresponding atomic event.

[0129] In this embodiment, by comparing the movement trajectory of the tracked object with the recognition rules of each atomic event one by one, when the information such as position and time in the movement trajectory meets the recognition rule of a certain atomic event, it can be determined that the atomic event occurs.

[0130] It can be understood that only by accurately generating atomic events can a reliable basis be provided for subsequent behavior analysis. Generating corresponding atomic events when the movement trajectory matches the recognition rule can avoid ignoring key behavior information in the current scenario, thereby improving the system's understanding of target behaviors and meeting the requirements of target tracking and event management.

[0131] Based on the above embodiments of the present application, in the fourth embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 4 , after step S42, the target tracking method based on artificial intelligence further includes steps S43 to S45:

[0132] Step S43: Input the behavior description text into a natural language processing model to obtain the behavior characteristics corresponding to the tracked object;

[0133] It should be noted that the natural language processing model is a model trained with a large amount of text data and can perform semantic analysis, feature extraction, etc. on the input natural language text. The behavior characteristics are the key information extracted from the behavior description text that can represent the behavior pattern of the tracked object, including but not limited to behavior actions, the time and location of the behavior occurrence.

[0134] It can be understood that since the behavior description text is an intuitive description of the behavior of the tracked object, step S43 can avoid subjective interpretation of behavior characteristics, thereby improving the accuracy and objectivity of behavior analysis.

[0135] Step S44: Determine the similarity between the behavior characteristics and the historical abnormal behavior characteristics;

[0136] It should be noted that the similarity is obtained by calculating the distance or similarity measure between behavior feature vectors, and usually methods such as cosine similarity are used.

[0137] Additionally, it should be noted that the historical abnormal behavior features are features extracted from the abnormal behaviors in the historical behavior records and are used as the benchmark for comparison.

[0138] It can be understood that since abnormal behaviors have certain patterns and features, step S44 can avoid missed detection of abnormal behaviors, thereby improving the reliability of abnormal behavior detection.

[0139] Step S45: If the similarity is greater than or equal to the preset threshold, output an abnormal behavior alarm.

[0140] It should be noted that the preset threshold is set according to the actual application scenario and the sensitivity to abnormal behaviors and is used to determine whether the behavior features are close enough to the historical abnormal behavior features.

[0141] Additionally, it should be noted that the abnormal behavior alarm can include alarm information, behavior description, relevant video clips, etc., and is used to notify the monitoring personnel in a timely manner.

[0142] It can be understood that since the timely discovery and handling of abnormal behaviors are crucial for security assurance, performing step S45 can avoid the delay in handling abnormal behaviors, thereby improving the security and response speed of the monitoring system.

[0143] Traditional tracking methods based on pure visual data usually only analyze the behavior features of the target based on the video data of the current day, lacking the recording and analysis of the target's past behaviors, resulting in the inability to establish effective behavior context associations. When traditional methods perform abnormal behavior detection on the target, it is difficult to distinguish the accidental actions of the target from the real abnormal behaviors, thereby reducing the judgment accuracy of the system and the reliability of behavior analysis.

[0144] In this application, the behavior description text is used as the input and sent to the natural language processing model for analysis. This text-based input method enables the system to call the historical behavior records of the tracked object and automatically associate the spatio-temporal context during real-time analysis, such as all the behavior data of the tracked object at present and in a certain past time period. At the same time, since the storage occupancy of text data is smaller than that of video streams, it significantly reduces the system's demand for computing resources and also improves the system's processing efficiency.

[0145] Exemplarily, to help understand the implementation process of the artificial intelligence-based target tracking method obtained by combining the above third embodiment, specifically:

[0146] The system detected that target B appeared in area P3 of the mall and stayed there for 30 minutes. During the stay in area P3, target B continuously looked in all directions of area P3 and made the gesture of looking down at the watch many times. The system input the video data corresponding to target B into the scene recognition model and determined that the current scene was a mall. Then, the system input the scene recognition result and the video data corresponding to target B into the behavior analysis model. Based on the input scene recognition result, the behavior analysis model loaded the corresponding atomic event recognition rules and identified that the atomic events in the current video data included "area P3", "stay", "30 minutes", "look up", "look down", and "look at the watch". Then, according to the temporal and spatial correlations of these atomic events, the behavior analysis model further analyzed and generated composite events, including "target B appeared in area P3", "target B stayed in area P3 for 30 minutes", "looked at area P3", and "target B looked down at the watch". Finally, based on the composite events and predefined sentence patterns, the behavior analysis model generated the behavior description text: "Target B appeared in area P3 of the mall at 1 pm and stayed for 30 minutes. During the stay, target B looked at area P3 many times and looked down at the watch." This behavior description text intuitively reflects the behavior characteristics of target B in area P3.

[0147] Furthermore, the system found in the historical monitoring results that the same target B had appeared in area P3 of the mall in the past two days, and each stay lasted about 30 minutes. Specifically, the behavior description text of target B on the first day in the past was: "Target B appeared in area P3 of the mall at around 1 pm, wandered for 25 minutes and then left." The behavior description text of target B on the second day in the past was: "Target B appeared in area P3 of the mall again at 1 pm, communicated with the staff in area P3, and left 20 minutes later."

[0148] According to the time sequence, the system synchronously input the historical behavior description texts generated by target B in the past two days and the current behavior description text into the natural language processing model, extracted the behavior characteristics of target B, and compared them with the historical abnormal behavior characteristics. The system found that the behavior pattern of target B exceeded the preset threshold in terms of similarity with the recorded historical abnormal behavior characteristics. These behavior characteristics specifically included wandering in area P3, communicating with the staff, looking at area P3 many times, and frequently looking down at the watch, etc. The behavior pattern of target B was repetitive and regular, and at the same time involved attention to time, which was highly similar to the theft behavior characteristics corresponding to the historical abnormal behavior characteristics. Therefore, the system determined that the behavior of target B was abnormal, triggered an abnormal behavior alarm, and marked and continuously monitored target B and its behavior pattern, prompting the staff to pay key attention to target B.

[0149] Based on the above embodiments of the present application, in the fifth embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 5 , after step S40, the target tracking method based on artificial intelligence further includes steps S401 to S405:

[0150] Step S401: If it is detected that the tracked object has left the monitoring range of the target camera;

[0151] Step S402: Determine the moving trend of the tracked object according to the moving trajectory of the tracked object within a preset period before leaving the monitoring range;

[0152] Step S403: Determine the handover camera according to the moving trend;

[0153] Step S404: Set the handover camera as the target camera and generate a target tracking process trigger instruction;

[0154] Step S405: Jump to execute the step of determining the human contour of the tracked object in the visual data collected by the target camera.

[0155] It should be noted that the monitoring range of the target camera refers to the monitoring area where the target camera can effectively collect visual data and thermal imaging data. The moving trajectory refers to the movement path of the tracked object within a preset period. The handover camera refers to the next camera that can continue to track the target.

[0156] It can be understood that during the monitoring process, the tracked object may move between multiple cameras, and it is difficult to achieve continuous tracking relying on a single camera. Through the above steps, the system can realize the handover between multiple target cameras according to the moving trajectory of the tracked object, ensuring the continuity of target tracking.

[0157] When the system detects that the tracked object has moved beyond the current monitoring area of the target camera, the system will detect the departure of the tracked target. At this time, the system determines a suitable handover camera according to the moving trajectory and moving trend of the tracked target before departure to ensure the continuity of target tracking. Specifically, the system first analyzes the moving trajectory of the tracked object within a preset period, which reflects the movement path of the tracked object within the monitoring range of the current camera. By the movement path of the tracked object before leaving the range of the current target camera, the system can predict the moving direction and speed of the tracked object, thereby determining the moving trend of the tracked object. According to the moving trend of the tracked object, the system selects a handover camera that can cover the predicted moving direction of the tracked object. The selection of the handover camera takes into account its position and monitoring range to ensure that the target can enter its monitoring area.

[0158] In an alternative embodiment, the system predicts the movement trend of the tracked object according to the current environmental state of the tracked object, such as the direction of crowd movement, the channel layout, etc.

[0159] After determining the handover camera, the system sets the handover camera as the new target camera and generates a target tracking process trigger instruction. The target tracking process trigger instruction includes information such as the identifier of the new target camera and the initial position of the tracked object within the monitoring range of the new target camera, ensuring that the tracking task can be seamlessly connected. Finally, the system jumps to execute step S10, re-determines the human contour of the tracked object in the visual data collected by the new target camera, and continues to perform target tracking.

[0160] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation on the object tracking method based on artificial intelligence of the present application. Any simple transformation in more forms based on this technical concept is within the protection scope of the present application.

[0161] The present application provides an object tracking system based on artificial intelligence. The object tracking system based on artificial intelligence includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the object tracking method based on artificial intelligence in the above first embodiment.

[0162] Refer to the following Figure 6 , which shows a schematic structural diagram of an object tracking system based on artificial intelligence suitable for implementing the embodiments of the present application. The object tracking system based on artificial intelligence in the embodiments of the present application may include, but is not limited to, video data acquisition terminals such as surveillance cameras, sensor devices, etc., data processing terminals such as video analysis servers, big data analysis platforms, etc., and fixed terminals such as monitoring center terminals, desktop computers, etc. Figure 6 The shown object tracking system based on artificial intelligence is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present application.

[0163] Such as Figure 6As shown in the figure, the artificial intelligence-based target tracking system may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the artificial intelligence-based target tracking system are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the artificial intelligence-based target tracking system to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an artificial intelligence-based target tracking system with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or had alternatively.

[0164] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0165] The artificial intelligence-based target tracking system provided by the present application adopts the artificial intelligence-based target tracking method in the above embodiments, and can solve the technical problem of tracking interruption or misjudgment caused by the target being blocked. Compared with the prior art, the beneficial effects of the artificial intelligence-based target tracking system provided by the present application are the same as those of the artificial intelligence-based target tracking method provided by the above embodiments, and other technical features in the artificial intelligence-based target tracking system are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0166] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0167] As mentioned above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0168] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the artificial intelligence-based target tracking method in the above embodiments.

[0169] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories (EPROMs), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0170] The above computer-readable storage medium can be included in an artificial intelligence-based target tracking system; or it can exist alone without being assembled into an artificial intelligence-based target tracking system.

[0171] The above computer-readable storage medium carries one or more programs, which, when executed by an artificial intelligence-based target tracking system, cause the artificial intelligence-based target tracking system to: in response to a target tracking process trigger instruction, determine the human body contour of a tracking object from the visual data collected by a target camera; if an occlusion area appears in the detected human body contour, obtain the thermal imaging data collected by the target camera; generate human body contour supplementary information corresponding to the occlusion area according to the spatio-temporal relationship between the visual data and the thermal imaging data, and the thermal imaging data; and simultaneously display the human body contour supplementary information and the visual data in the display interface of the visual data, wherein the human body contour supplementary information is displayed on the upper layer of the visual data.

[0172] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, through the Internet using an Internet service provider).

[0173] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0174] The modules involved in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0175] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned artificial intelligence-based target tracking method, which can solve the technical problem of tracking interruption or misjudgment caused by target occlusion. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the artificial intelligence-based target tracking method provided by the above embodiments, and will not be elaborated here.

[0176] The above are only some embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. An artificial intelligence-based target tracking method, characterized in that, The artificial intelligence-based target tracking method includes: In response to a target tracking process trigger instruction, determine the human contour of the tracking object in the visual data collected by the target camera; If an occlusion area appears in the detected human contour, obtain the thermal imaging data collected by the target camera; Generate human contour supplementary information corresponding to the occlusion area according to the spatio-temporal relationship between the visual data and the thermal imaging data, and the thermal imaging data; In the display interface of the visual data, simultaneously display the human contour supplementary information and the visual data, where the human contour supplementary information is displayed on the upper layer of the visual data.

2. The object tracking method based on artificial intelligence according to claim 1, characterized in that, The step of generating the human contour supplementary information corresponding to the occlusion area according to the spatio-temporal relationship between the visual data and the thermal imaging data, and the thermal imaging data includes: According to the spatio-temporal relationship, determine the human area corresponding to the tracking object in the thermal imaging frames corresponding to each moment; Generate a thermal imaging human contour according to the human area in the thermal imaging frame; Determine the human contour supplementary information according to the thermal imaging human contour, the human contour, and the occlusion area.

3. The object tracking method based on artificial intelligence according to claim 1, wherein, After the step of simultaneously displaying the human contour supplementary information and the visual data in the display interface of the visual data, the method further includes: Input the visual data into a scene recognition model to obtain a scene recognition result through the scene recognition model; Input the scene recognition result and the visual data into a behavior analysis model to obtain a behavior description text.

4. The object tracking method based on artificial intelligence according to claim 3, characterized in that, The step of inputting the scene recognition result and the visual data into a behavior analysis model to obtain a behavior description text includes: In response to the input data, the behavior analysis model loads the atomic event recognition rules corresponding to the target scene according to the scene recognition result; Generate atomic events according to the atomic event recognition rules and the visual data; Generate composite events based on the spatio-temporal combination of at least two of the atomic events; Generate the behavior description text based on the composite event and a predefined sentence pattern.

5. The object tracking method based on artificial intelligence according to claim 4, characterized in that, The step of generating atomic events according to the atomic event recognition rules and the visual data includes: Generate the movement trajectory of the tracking object according to the spatio-temporal position relationship between the tracking object and a preset reference object in the visual data, and the pre-stored geodetic coordinates associated with the preset reference object; Determine the recognition rules corresponding to the stay event, the area switching event, and / or the sensitive area entry event according to the atomic event recognition rules; When it is detected that the movement trajectory matches the recognition rules, generate the corresponding atomic event.

6. The object tracking method based on artificial intelligence according to claim 3, wherein After the step of inputting the scene recognition result and the visual data into a behavior analysis model to obtain a behavior description text, the method further includes: Input the behavior description text into a natural language processing model to obtain the behavior characteristics corresponding to the tracking object; Determine the similarity between the behavior characteristics and the historical abnormal behavior characteristics; If the similarity is greater than or equal to a preset threshold, output an abnormal behavior alarm.

7. The object tracking method based on artificial intelligence according to claim 1, wherein After the step of simultaneously displaying the human body contour supplementary information and the visual data in the display interface of the visual data, the method further includes: If it is detected that the tracked object exits the monitoring range of the target camera; Determine the moving trend of the tracked object according to the moving trajectory of the tracked object within a preset time period before the tracked object exits the monitoring range; Determine the handover camera according to the moving trend; Set the handover camera as the target camera and generate a target tracking process trigger instruction; Jump to execute the step of determining the human body contour of the tracked object in the visual data collected by the target camera.

8. The object tracking method based on artificial intelligence according to claim 1, characterized in that Before the step of determining the human body contour of the tracked object in the visual data collected by the target camera, the method further includes: In response to the control operation received by the monitoring video display interface, determine the tracked object and generate the target tracking process trigger instruction.

9. An artificial intelligence-based target tracking system, characterized in that, The system includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the artificial intelligence-based target tracking method according to any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the artificial intelligence-based target tracking method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Cleaning control method and device of cleaning equipment, medium and electronic equipment

    CN121455036A

  • Vehicle thermal imaging intelligent sensing method linked with aerial imaging system

    CN121937446A