Image identification and perception collaborative video linkage tracking method and system

Through the video linkage tracking method of coordinated image recognition and perception, combined with video enhancement and multi-dimensional perception processing, the accuracy and stability problems of traditional video tracking in complex scenes are solved, and high-precision target tracking effects are achieved.

CN120808237APending Publication Date: 2025-10-17HANGZHOU GUANGYU TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510993811.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In existing technologies, traditional video tracking relies on single image recognition and lacks effective fusion of multimodal perception information, resulting in insufficient target tracking accuracy and stability in complex scenarios such as lighting changes, target occlusion, rapid movement, or multi-view switching, and is prone to loss or misjudgment.

Method used

Through the video linkage tracking method that coordinates image recognition and perception, the monitoring video stream is preprocessed in combination with video enhancement strategy and multi-dimensional sensor, the target video segment is intercepted, enhanced processing and multi-dimensional perception monitoring are performed, and the dual cascade mechanism is used to fuse the image and perception information to achieve linkage tracking of the target.

Benefits of technology

It improves the stability and accuracy of target tracking, enhances adaptability in complex environments, and achieves high-precision dynamic linkage tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808237A_ABST
    Figure CN120808237A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition and perception collaborative video linkage tracking method and system, and relates to the technical field related to image recognition, and the method comprises the steps: carrying out the preprocessing of a monitoring video stream, obtaining a target video stream, and intercepting any video segment; introducing a video enhancement strategy to perform enhancement processing on any video segment to obtain any image; activating a multi-dimensional sensor arranged on any lens to perform dynamic continuous sensing monitoring on a predetermined target to obtain multi-dimensional sensing information; performing double-cascade processing on any image and the multi-dimensional perception information according to a double-cascade mechanism to obtain a target fusion feature; and performing linkage tracking of the predetermined target based on the target fusion feature. The technical problems that in the prior art, pure image recognition tracking is insufficient in accuracy and stability, and target tracking is prone to losing or misjudgment in a complex scene are solved, and the technical effects that the stability and accuracy of target tracking are improved, the adaptability to the complex environment is enhanced, and high-precision dynamic linkage tracking is achieved are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition, and particularly relates to a video linkage tracking method and system based on image recognition and perception coordination. BACKGROUND

[0002] Intelligent video analysis plays an important role in the fields of security and protection, traffic management, smart city, etc. However, traditional video tracking mainly relies on single image recognition, lacks effective fusion of multi-modal perception information, and often has difficulty in dealing with dynamic target tracking problems in complex scenes, especially under challenging conditions such as light changes, target occlusions, fast movements or multi-view switching, resulting in insufficient accuracy, continuity and stability of target tracking. In addition, existing video enhancement processing is independent of the tracking task and lacks coordination and optimization with the perception module, so that the enhanced information cannot fully serve the accuracy and real-time requirements of target tracking.

[0003] Therefore, in the related art, there are technical problems of insufficient accuracy and stability of pure image recognition tracking and easy loss or misjudgment of target tracking in complex scenes. SUMMARY

[0004] The present application provides a video linkage tracking method and system based on image recognition and perception coordination, which solves the technical problems of insufficient accuracy and stability of pure image recognition tracking and easy loss or misjudgment of target tracking in complex scenes in the prior art, and achieves the technical effects of improving the stability and accuracy of target tracking, enhancing the adaptability to complex environments, and realizing high-precision dynamic linkage tracking.

[0005] The present application provides a video linkage tracking method based on image recognition and perception coordination, which comprises: pre-processing a monitoring video stream to obtain a target video stream, and intercepting any video segment in the target video stream, wherein the any video segment corresponds to any shot and contains a predetermined target; introducing a video enhancement strategy to perform enhancement processing on the any video segment to obtain any image; activating a multi-dimensional perception device arranged in the any shot to perform dynamic and continuous perception monitoring on the predetermined target to obtain multi-dimensional perception information; performing double-level connection processing on the any image and the multi-dimensional perception information according to a double-level connection mechanism to obtain target fusion features; and performing linkage tracking on the predetermined target based on the target fusion features.

[0006] In a possible implementation, the image recognition and perception cooperative video linkage tracking method further performs the following processing: compressing the target video stream based on a dynamic image expert group to obtain a compressed video, wherein the compressed video includes a first encoding element and a second encoding element, and the first encoding element and the second encoding element are adjacent encoding elements; sequentially extracting a first intraframe image in the first encoding element and a second intraframe image in the second encoding element; determining whether a comparison result obtained by comparing the first intraframe image and the second intraframe image satisfies a predetermined segmentation constraint; if yes, taking the first encoding element and the second encoding element as a segmentation node; segmenting the compressed video based on the segmentation node to obtain a segmentation result, wherein the segmentation result includes a plurality of video segments; and randomly extracting an arbitrary video segment from the plurality of video segments, denoted as the arbitrary video segment.

[0007] In a possible implementation, the image recognition and perception cooperative video linkage tracking method further performs the following processing: obtaining a first feature coefficient of the first intraframe image; obtaining a second feature coefficient of the second intraframe image; taking a difference between the first feature coefficient and the second feature coefficient as the comparison result; and wherein the obtaining the first feature coefficient of the first intraframe image includes: obtaining a first tone feature value of the first intraframe image; obtaining a first texture feature value of the first intraframe image; and performing weighted calculation on the first tone feature value and the first texture feature value to obtain the first feature coefficient.

[0008] In a possible implementation, the image recognition and perception cooperative video linkage tracking method further performs the following processing: taking a first frame image in the arbitrary video segment as an enhancement reference; and performing alignment enhancement on a plurality of frame images in the arbitrary video segment based on the video enhancement strategy and taking the enhancement reference as a constraint to obtain the arbitrary image.

[0009] In a possible implementation, the image recognition and perception cooperative video linkage tracking method further performs the following processing: monitoring, by the infrared perception component, to obtain an infrared image of the predetermined target; monitoring, by the sonar perception component, to obtain a sonar image of the predetermined target; monitoring, by the radar perception component, to obtain a radar image of the predetermined target; and taking the infrared image, the sonar image, and the radar image to constitute the multi-dimensional perception information.

[0010] In a possible implementation, the image recognition and perception cooperative video linkage tracking method further performs the following processing: extracting a fractional order plan in the double cascade mechanism; processing the multi-dimensional perception information according to the fractional order plan to obtain target perception information; extracting a cross-modal fusion plan in the double cascade mechanism; processing the arbitrary image and the target perception information according to the cross-modal fusion plan to obtain the target fusion feature.

[0011] In a possible implementation, the image recognition and perception cooperative video linkage tracking method further performs the following processing: obtaining a target infrared image of the infrared image according to the fractional order plan; obtaining a target sonar image of the sonar image according to the fractional order plan; obtaining a target radar image of the radar image according to the fractional order plan; the target infrared image, the target sonar image and the target radar image constitute the target perception information; wherein obtaining the target infrared image of the infrared image according to the fractional order plan comprises: performing two-dimensional Fourier transform on the infrared image according to the fractional order plan to obtain an infrared fractional order domain; performing filtering processing on the infrared fractional order domain in combination with a predetermined rotation order, and inversely transforming to obtain the target infrared image.

[0012] In a possible implementation, the image recognition and perception cooperative video linkage tracking method further performs the following processing: obtaining a real-time linkage tracking image of the predetermined target; sequentially obtaining a front frame image and a rear frame image of the real-time linkage tracking image; performing matching analysis on the real-time linkage tracking image, the front frame image and the rear frame image based on a three-frame shadow matching strategy to obtain a matching result; if the matching result meets a predetermined matching condition, taking a target contour of the predetermined target in the real-time linkage tracking image as a meta-learning benchmark; wherein the meta-learning benchmark is used for continuously tracking the predetermined target.

[0013] In a possible implementation, the image recognition and perception cooperative video linkage tracking method further performs the following processing: obtaining a real-time centroid of the real-time linkage tracking image; obtaining a rear frame centroid of the rear frame image; comparing a centroid offset of the real-time centroid and the rear frame centroid; if the centroid offset is less than a predetermined threshold, determining that the predetermined target is in a stationary state, and stopping linkage tracking.

[0014] The application also provides a video linkage tracking system based on image recognition and perception cooperation, comprising: an arbitrary video segment intercepting module, configured to pre-process a monitoring video stream to obtain a target video stream, and intercept an arbitrary video segment in the target video stream, wherein the arbitrary video segment corresponds to an arbitrary shot and contains a predetermined target; an enhancement processing module, configured to introduce a video enhancement strategy to perform enhancement processing on the arbitrary video segment to obtain an arbitrary image; a multi-dimensional perception information obtaining module, configured to activate a multi-dimensional perception device arranged in the arbitrary shot to perform dynamic continuity perception monitoring on the predetermined target to obtain multi-dimensional perception information; a double cascade processing module, configured to perform double cascade processing on the arbitrary image and the multi-dimensional perception information according to a double cascade mechanism to obtain target fusion features; and a linkage tracking module, configured to perform linkage tracking on the predetermined target based on the target fusion features.

[0015] The application provides a video linkage tracking method and system based on image recognition and perception cooperation. The method comprises the following steps: pre-processing a monitoring video stream to obtain a target video stream, and intercepting an arbitrary video segment; introducing a video enhancement strategy to perform enhancement processing on the arbitrary video segment to obtain an arbitrary image; activating a multi-dimensional perception device arranged in the arbitrary shot to perform dynamic continuity perception monitoring on the predetermined target to obtain multi-dimensional perception information; performing double cascade processing on the arbitrary image and the multi-dimensional perception information according to a double cascade mechanism to obtain target fusion features; and performing linkage tracking on the predetermined target based on the target fusion features. The technical problems of insufficient tracking accuracy and stability of pure image recognition and easy loss or misjudgment of a target in a complex scene are solved, and the technical effects of improving the stability and accuracy of target tracking, enhancing the adaptability to a complex environment, and realizing high-precision dynamic linkage tracking are achieved. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments of the present disclosure will be briefly introduced below. The flowcharts are used to illustrate the operations performed by the system according to the embodiments of the present disclosure. It should be understood that the foregoing or the following operations are not necessarily performed in sequence. On the contrary, various steps can be processed in reverse order or simultaneously according to needs. Meanwhile, other operations can be added to these processes, or one or more steps can be removed from these processes.

[0017] Figure 1 A flowchart of a video linkage tracking method based on image recognition and perception cooperation provided by the embodiments of the present disclosure.

[0018] Figure 2 A structural diagram of a video linkage tracking system based on image recognition and perception cooperation provided by the embodiments of the present disclosure.

[0019] Reference signs: arbitrary video segment interception module 10, enhancement processing module 20, multi-dimensional perception information obtaining module 30, double cascade processing module 40, linkage tracking module 50. DETAILED DESCRIPTION

[0020] The above description is only a summary of the technical scheme of the present application. In order to make the technical means of the present application more clear, the present application can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described.

[0021] In order to make the purposes, technical schemes and advantages of the present application more clear, the following will combine the drawings to further describe the present application in detail. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0022] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict. The term "first\second" is only to distinguish similar objects, and does not represent the specific order of the objects. The terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application.

[0023] The embodiments of the present application provide a video linkage tracking method based on image recognition and perception cooperation, as shown in Figure 1 The method comprises the following steps:

[0024] In step S100, the target video stream is obtained by preprocessing the monitoring video stream, and an arbitrary video segment in the target video stream is intercepted, wherein the arbitrary video segment corresponds to an arbitrary shot and contains a predetermined target.

[0025] Preferably, in the video monitoring system, the original monitoring video stream usually contains a large amount of redundant information such as static background, irrelevant moving objects, noise, etc., and direct processing will reduce the computational efficiency and affect the tracking accuracy. Therefore, the monitoring video stream is pre-processed to extract the target video stream. Specifically, the monitoring video stream is denoised by using Gaussian filtering, median filtering and the like to reduce the noise interference of the monitoring video, and then the dynamic target and the static background are separated by using background subtraction or optical flow method to retain the video frames containing the moving target. Then, a target detection algorithm such as YOLO or Faster R-CNN is used to identify specific moving targets in the video, such as cars, pedestrians, etc., and filter out irrelevant content to ensure that the predetermined target is processed. Then, the video is normalized in size, such as uniformly adjusted to 1080p, and then the target video stream is obtained, mainly containing high signal-to-noise ratio clear pictures and standardized video data of the predetermined target.

[0026] Further, step S100 further comprises step S110, compressing the target video stream based on a dynamic image expert group to obtain a compressed video, wherein the compressed video comprises a first encoding element and a second encoding element, and the first encoding element and the second encoding element are adjacent encoding elements; step S120, sequentially extracting a first intra-frame image in the first encoding element and a second intra-frame image in the second encoding element; step S130, judging whether a comparison result obtained by comparing the first intra-frame image and the second intra-frame image satisfies a predetermined segmentation constraint; step S140, if yes, taking the first encoding element and the second encoding element as a segmentation node; step S150, segmenting the compressed video based on the segmentation node to obtain a segmentation result, wherein the segmentation result comprises a plurality of video segments; and step S160, randomly extracting any one of the plurality of video segments as the arbitrary video segment.

[0027] Preferably, the dynamic image expert group MPEG is a general video compression standard. First, the pre-processed target video stream is MPEG encoded and compressed to generate two image groups, i.e. a first encoding element and a second encoding element, and then the two encoding elements form a compressed video. The first encoding element and the second encoding element are adjacent GOP units, i.e. the last frame of the first encoding element and the first frame of the second encoding element are continuous in time, and the logical segmentation of the video stream is realized through the correlation between the encoding elements. Each encoding element contains three frame types, i.e. I frame, P frame and B frame. The I frame is an intra-frame encoding frame, which is a key frame for independent compression. The P frame is a prediction frame, which is compressed based on the previous I frame or P frame. The B frame is a bidirectional prediction frame, which refers to the previous and subsequent frames for higher compression.

[0028] Preferably, the key frames in the adjacent coding units are extracted in sequence for content analysis, i.e. the first intra-frame image is extracted from the first coding unit, the second intra-frame image is extracted from the second coding unit, which may correspond to the I-frame in the coding unit, and then it is determined whether the comparison result obtained by comparing the first intra-frame image and the second intra-frame image satisfies the predetermined segmentation constraint. Specifically, the structural similarity or feature point matching degree of the first intra-frame image and the second intra-frame image is calculated to obtain the difference between the two intra-frame images, the comparison result is determined and compared with the predetermined segmentation constraint, which may be scene switching or target hour. If the comparison result satisfies the predetermined segmentation constraint, i.e. the segmentation condition is met, the connection between the first coding unit and the second coding unit is marked as a segmentation node, and the compressed video is segmented based on the segmentation node. For example, when the camera view angle is switched, the I-frames of adjacent coding units are significantly different, which automatically triggers segmentation, and then the segmentation result is obtained, i.e. the compressed video is divided into multiple video segments, each segment contains several complete coding units, and the visual content in the segment is consistent. Finally, any one of the multiple video segments is randomly extracted and recorded as an arbitrary video segment.

[0029] Further, step S130 further comprises step S131 of obtaining a first feature coefficient of the first intra-frame image; step S132 of obtaining a second feature coefficient of the second intra-frame image; and step S133 of taking the difference between the first feature coefficient and the second feature coefficient as the comparison result. Step S131 comprises: step A of obtaining a first tone feature value of the first intra-frame image; step B of obtaining a first texture feature value of the first intra-frame image; and step C of performing weighted calculation on the first tone feature value and the first texture feature value to obtain the first feature coefficient.

[0030] Preferably, the key frame color and texture features in the adjacent image group are extracted, and the weighted feature coefficients are calculated, and the feature difference is used to determine whether the video needs to be segmented, wherein the color feature reflects the overall color distribution of the image, which is used to detect the scene light change or target color mutation, and the texture feature describes the local structure information of the image, which is used to identify the scene content change, such as from a smooth wall to a complex street. Specifically, for the first frame image, the pixel distribution of the RGB / HSV channel is counted, the mean or entropy value is calculated, and the main color component is obtained by clustering as the first color feature value; the gray level co-occurrence matrix is generated for the first frame image, and the contrast, energy, homogeneity and other indicators are calculated, and then the statistical features of the texture pattern are extracted to capture the texture edge and shape information to obtain the first texture feature value; finally, the first color feature value and the first texture feature value are weighted to obtain the first feature coefficient corresponding to the first frame image, wherein the weight coefficient is set according to the experimental data. Similarly, the second color feature value and the second texture feature value are obtained and weighted to obtain the second feature coefficient corresponding to the second frame image, and then the absolute value of the difference between the first feature coefficient and the second feature coefficient is calculated as the comparison result. If the comparison result meets the predetermined segmentation constraint, it means that there is a significant content change between the two coding units, and the compressed video is segmented, for example, when the camera switches from indoor to outdoor, the feature coefficient difference far exceeds the predetermined segmentation constraint, triggering the video segmentation.

[0031] Step S200, introducing a video enhancement strategy to enhance the processing of the arbitrary video segment to obtain an arbitrary image.

[0032] Step S200 further includes step S210, taking the first frame image in the arbitrary video segment as an enhancement reference; step S220, according to the video enhancement strategy, aligning and enhancing the multiple frame images in the arbitrary video segment with the enhancement reference as a constraint to obtain the arbitrary image.

[0033] Preferably, when video monitoring and target tracking are performed, due to factors such as camera shaking, light change or target motion, the multiple frame images in the monitoring video stream may have spatial inconsistencies such as offset and blur, and directly enhancing a single frame image may easily lead to incoherent information between image frames, affecting the target tracking accuracy. Then, the video enhancement strategy is introduced to enhance the processing of the arbitrary video segment, specifically, the first frame image in the arbitrary video segment is taken as an enhancement reference, wherein the first frame image is usually the initial stable state of the scene, such as the clearest picture when the camera is just started, which is suitable as a reference reference. The reference content of the first frame image includes spatial reference and color reference, the spatial reference refers to the composition of the first frame image, such as target position and background structure, and the color reference refers to the brightness, contrast, white balance and other parameters of the first frame.

[0034] Preferably, the video enhancement strategy can include super-resolution reconstruction, i.e. improving resolution based on deep learning; using non-local mean or generative adversarial network to denoise and deblur; adaptive histogram equalization for illumination correction. Then according to the video enhancement strategy, the enhancement benchmark is constrained, and the multi-frame images in the arbitrary video segment are aligned and enhanced. Alignment refers to calculating the pixel-level motion vector of the subsequent frame and the first frame by the optical flow method and performing deformation correction, or detecting key points by feature matching and performing perspective transformation. Specifically, first, align the multi-frame images in the arbitrary video segment with the first frame image by optical flow or feature matching, then apply super-resolution reconstruction and other enhancement processing algorithms to the aligned images to ensure that the enhancement results meet the enhancement benchmark, and finally output the enhanced image sequence, i.e. obtain an arbitrary image.

[0035] Step S300, activating the multi-dimensional sensor arranged in the arbitrary lens to perform dynamic continuity perception monitoring of the predetermined target, and obtaining multi-dimensional perception information.

[0036] Step S300 further includes that the multi-dimensional sensor includes an infrared perception component, a sonar perception component, and a radar perception component; step S310, monitoring to obtain an infrared image of the predetermined target through the infrared perception component; step S320, monitoring to obtain a sonar image of the predetermined target through the sonar perception component; step S330, monitoring to obtain a radar image of the predetermined target through the radar perception component; and step S340, the infrared image, the sonar image, and the radar image constitute the multi-dimensional perception information.

[0037] Preferably, in complex environments such as night, haze, and occlusion scenes, visual sensors are difficult to ensure the continuity of target tracking. Through multi-dimensional sensor fusion, the complementary advantages of infrared, sonar, and radar are combined to realize all-weather and all-scene target perception. The multi-dimensional sensor includes an infrared perception component, a sonar perception component, and a radar perception component. The infrared perception component generates an infrared image by detecting the thermal radiation of the target and is not affected by the visible light illumination condition. It outputs a thermal imaging image and a target contour, and is suitable for night monitoring or occlusion penetration. The sonar perception component generates a sonar image by emitting ultrasonic waves and receiving echoes, and is suitable for short-range high-precision detection. It outputs a range-azimuth image and target motion speed, and is suitable for underwater target tracking or indoor close-range obstacle avoidance. The radar perception component generates a radar image by emitting millimeter waves / microwaves and analyzing reflected signals, and has long-range detection capability. It outputs point cloud data and target micro-motion features, and is suitable for traffic monitoring and severe weather target detection.

[0038] Preferably, when a predetermined target is detected in the video stream, the infrared sensing component, the sonar sensing component and the radar sensing component arranged in any lens are synchronously activated to dynamically and continuously monitor the predetermined target, including monitoring the infrared image of the predetermined target through the infrared sensing component, which contains the thermal features of the target; monitoring the sonar image of the predetermined target through the sonar sensing component, which contains the range-azimuth features of the target; and monitoring the radar image of the predetermined target through the radar sensing component, which contains the speed and three-dimensional position information of the target. Finally, the infrared image, the sonar image and the radar image are time-aligned and combined into multi-dimensional perception information to form multi-modal representation of the predetermined target, thereby ensuring high-precision dynamic linkage tracking of the target.

[0039] Step S400: performing double-cascade processing on the arbitrary image and the multi-dimensional perception information according to a double-cascade mechanism to obtain target fusion features.

[0040] Step S400 further includes the following steps: Step S410: extracting a fractional order plan in the double-cascade mechanism; Step S420: processing the multi-dimensional perception information according to the fractional order plan to obtain target perception information; Step S430: extracting a cross-modal fusion plan in the double-cascade mechanism; and Step S440: processing the arbitrary image and the target perception information according to the cross-modal fusion plan to obtain the target fusion features.

[0041] Preferably, for target tracking in a complex scene, the modal data greatly differ, in which the visual data is a pixel matrix, the radar is a point cloud, and the sonar is a distance graph, and the image and the multi-dimensional perception information need to be fused, i.e., the double-cascade processing is performed on the arbitrary image and the multi-dimensional perception information through a double-cascade mechanism, in which the double-cascade mechanism includes a fractional order plan and a cross-modal fusion plan. The fractional order plan is used to optimize the quality of single-modal data, such as removing noise, enhancing resolution and signal features, etc. The cross-modal fusion plan is used to dynamically and cooperatively fuse multi-modal features. Moreover, the reliability of each modality is different in different scenes, such as the weight of the infrared image is greater at night, and the weight of the radar image is greater in rainy days.

[0042] Preferably, the fractional order plan in the double cascade mechanism is extracted to process the multi-dimensional perception information. Specifically, the multi-dimensional perception information is processed by fractional order filtering, i.e., fractional order Gaussian filtering is applied to the radar point cloud to retain the characteristics of micro motion; sonar image enhancement is performed, i.e., a fractional order differential operator is used to strengthen the edges; infrared image correction is performed, i.e., a fractional order heat conduction model is used to compensate for temperature drift; and finally, the denoised target perception information is output, including enhanced radar, sonar and infrared data. Then, the cross-modal fusion plan in the double cascade mechanism is extracted to process any image and target perception information. According to the environmental credibility, the weight coefficients of each modal data are dynamically adjusted, and then the different modal data are mapped to a unified feature space for weighted fusion. Finally, the target fusion feature is obtained, including multi-dimensional information such as target shape, motion and thermal radiation, thereby significantly improving the target representation ability in complex scenes.

[0043] Further, step S420 further comprises step S421, obtaining a target infrared image of the infrared image according to the fractional order plan; step S422, obtaining a target sonar image of the sonar image according to the fractional order plan; step S423, obtaining a target radar image of the radar image according to the fractional order plan; and step S424, the target infrared image, the target sonar image and the target radar image constitute the target perception information; wherein, obtaining a target infrared image of the infrared image according to the fractional order plan comprises: step a, performing two-dimensional Fourier transform on the infrared image according to the fractional order plan to obtain an infrared fractional order domain; and step b, filtering the infrared fractional order domain in combination with a predetermined rotation order, and inversely transforming to obtain the target infrared image.

[0044] Preferably, the infrared image is processed according to the fractional order plan to obtain a corresponding target infrared image, that is, the fractional order calculus theory is adopted, and the fractional order Fourier transform and fractional order filtering are performed to significantly improve the signal-to-noise ratio while retaining the essential characteristics of the target. Specifically, the infrared image is subjected to two-dimensional Fourier transform according to the fractional order plan, the image is converted from the spatial domain to the time-frequency joint domain, while retaining the spatial position information and frequency characteristic information, and then the infrared fractional order domain is obtained, which is a hybrid domain between the time domain and the frequency domain; then, the infrared fractional order domain is filtered in combination with a predetermined rotation order, wherein the predetermined rotation order represents the best time-frequency resolution, and the optimal order is set according to the target characteristics, which can be 0.5, and the predetermined rotation order can be dynamically optimized according to the image signal-to-noise ratio, and the fractional order filtering is performed on the infrared fractional order domain to retain the weak target, and finally the filtered fractional order domain data is inversely converted back to the spatial domain to generate the target infrared image, which ensures the enhancement of the target contour and the smoothing of the non-uniform thermal noise. Similarly, the sonar image is processed according to the fractional order plan to strengthen the echo edge of the underwater target, and the target sonar image is obtained; the radar image is processed according to the fractional order plan, that is, the fractional order calculus is used to enhance the micro-Doppler characteristics of slow small targets, and the target radar image is obtained; and finally, the target infrared image, the target sonar image and the target radar image are spatio-temporally aligned to form the target perception information.

[0045] In step S500, the target fusion feature is used to perform the linkage tracking of the predetermined target.

[0046] Preferably, the target fusion feature obtained through multi-modal fusion is used in combination with a cross-sensor cooperation mechanism to realize the linkage tracking of the predetermined target. Specifically, the visual image is fused with the radar point cloud and the sonar range image to form a representation containing multi-dimensional information such as target appearance, motion, thermal radiation and spatial position, which retains the detailed recognition ability of vision and has the anti-interference characteristics of non-visual sensors. At the same time, the contribution weights of different sensors are dynamically allocated according to the environmental conditions, for example, infrared and radar are used as the main sensors at night, and radar and sonar are used in foggy weather; the target motion trajectory is predicted through Kalman filtering or LSTM time series modeling, and the spatial relationship of multiple cameras or multiple perception nodes is used to realize seamless handover of the target between different viewing angles and different sensors, automatically associate the target identity and continue tracking when the target enters the monitoring range of another camera from the field of view of one camera, and quickly re-identify based on the historical data of the fusion feature when the target is temporarily lost due to occlusion, thereby reducing the tracking interruption and ensuring the stability and accuracy of the target tracking.

[0047] Further, step S500 further comprises step S510 of acquiring a real-time linkage tracking image of the predetermined target; step S520 of sequentially acquiring a front frame image and a rear frame image of the real-time linkage tracking image; step S530 of performing matching analysis on the real-time linkage tracking image, the front frame image and the rear frame image based on a three-frame shadow matching strategy to obtain a matching result; and step S540 of taking a target contour of the predetermined target in the real-time linkage tracking image as a meta-learning benchmark if the matching result meets a predetermined matching condition, wherein the meta-learning benchmark is used for continuous tracking of the predetermined target.

[0048] Preferably, a tracking picture of the predetermined target is output in real time based on target fusion features as a current frame, i.e., a real-time linkage tracking image, and a front frame image and a rear frame image adjacent to the real-time linkage tracking image are sequentially cached synchronously to form three frames of time sequence images; then matching analysis is performed on the real-time linkage tracking image, the front frame image and the rear frame image based on a three-frame shadow matching strategy, specifically, a target-shadow contour pair in the three frames is extracted by using Hu moment invariants matching, and a cross matching is performed to calculate Hu moments thereof, containing 7 translation / rotation / scaling invariant moment features, if the three continuous frames all meet a Hu moment similarity threshold, i.e., successful pairing, it is determined as a first suspected target, a matching result is obtained, and an outer contour template and an initial centroid coordinate of the target are saved. Then a target contour of the predetermined target in the real-time linkage tracking image is taken as a meta-learning benchmark, containing geometric features such as Hu moments and contour templates, and space-time features such as initial positions and motion trends, and the meta-learning benchmark is used for continuous tracking of the predetermined target, for example, if a matching degree between a real-time contour and the meta-learning benchmark in a subsequent frame image is lower than a threshold, the benchmark template is re-matched in a peripheral region of a predicted position, and when the re-matching fails, the meta-learning benchmark state is traced back to reset a tracking starting point, thereby significantly improving tracking reliability in a complex scene.

[0049] Further, step S520 further comprises step S521 of acquiring a real-time centroid of the real-time linkage tracking image; step S522 of acquiring a rear frame centroid of the rear frame image; step S523 of comparing a centroid offset between the real-time centroid and the rear frame centroid; and step S524 of determining that the predetermined target is in a stationary state and stopping linkage tracking if the centroid offset is less than a predetermined threshold.

[0050] Preferably, during the continuous target tracking process, when the target is stationary for a long time, such as a vehicle stopping or a pedestrian standing, continuing full-function tracking will cause unnecessary consumption of computing resources. Through real-time centroid displacement analysis, intelligent identification and energy-saving control of the motion state are achieved. Specifically, the geometric center coordinates of the target bounding rectangle are extracted from the real-time linked tracking image as the real-time centroid, and the center coordinates of the target are extracted from the next frame image as the post-frame centroid. Then, the centroid displacement of the real-time centroid and the post-frame centroid is compared, that is, the centroid displacement of the real-time centroid and the post-frame centroid is calculated using the Euclidean distance. If the centroid displacement is less than a predetermined threshold (such as 3 pixels), it is determined that the predetermined target is in a stationary state, and the linked tracking is stopped, including stopping sensor data fusion and feature calculation or switching to an interval sampling mode. When the centroid displacement exceeds the threshold again or the target shape changes suddenly, full-function tracking is resumed to ensure the energy efficiency of the monitoring system.

[0051] In the foregoing, reference is made to Figure 1 A video linked tracking method based on image recognition and perception cooperation is described in detail according to an embodiment of the present application. Next, a video linked tracking system based on image recognition and perception cooperation according to an embodiment of the present application will be described with reference to Figure 2 A video linked tracking system based on image recognition and perception cooperation according to an embodiment of the present application will be described with reference to

[0052] The video linked tracking system based on image recognition and perception cooperation according to an embodiment of the present application is used to solve the technical problems of insufficient accuracy and stability of pure image recognition tracking and easy loss or misjudgment of target tracking in a complex scene in the prior art, and achieves the technical effects of improving the stability and accuracy of target tracking, enhancing the adaptability to complex environments, and realizing high-precision dynamic linked tracking. As shown in Figure 2 A video linked tracking system based on image recognition and perception cooperation includes an arbitrary video segment interception module 10, an enhancement processing module 20, a multi-dimensional perception information obtaining module 30, a double cascade processing module 40, and a linked tracking module 50.

[0053] The arbitrary video segment interception module 10 is used to pre-process a monitoring video stream to obtain a target video stream, and intercept an arbitrary video segment in the target video stream, wherein the arbitrary video segment corresponds to an arbitrary shot and contains a predetermined target. The enhancement processing module 20 is used to introduce a video enhancement strategy to perform enhancement processing on the arbitrary video segment to obtain an arbitrary image. The multi-dimensional perception information obtaining module 30 is used to activate a multi-dimensional perception device arranged in the arbitrary shot to perform dynamic and continuous perception monitoring on the predetermined target to obtain multi-dimensional perception information. The double cascade processing module 40 is used to perform double cascade processing on the arbitrary image and the multi-dimensional perception information according to a double cascade mechanism to obtain target fusion features. The linked tracking module 50 is used to perform linked tracking of the predetermined target based on the target fusion features.

[0054] In the following, the specific configuration of the arbitrary video segment intercepting module 10 will be described in detail. The arbitrary video segment intercepting module 10 further comprises: compressing the target video stream based on a dynamic image expert group to obtain a compressed video, wherein the compressed video comprises a first encoding element and a second encoding element, and the first encoding element and the second encoding element are adjacent encoding elements; sequentially extracting a first intra-frame image in the first encoding element and a second intra-frame image in the second encoding element; judging whether a comparison result obtained by comparing the first intra-frame image and the second intra-frame image satisfies a predetermined segmentation constraint; if yes, taking the first encoding element and the second encoding element as a segmentation node; segmenting the compressed video based on the segmentation node to obtain a segmentation result, wherein the segmentation result comprises a plurality of video segments; and randomly extracting an arbitrary one of the plurality of video segments as the arbitrary video segment.

[0055] In the following, the specific configuration of the arbitrary video segment intercepting module 10 will be described in detail. The arbitrary video segment intercepting module 10 further comprises: obtaining a first feature coefficient of the first intra-frame image; obtaining a second feature coefficient of the second intra-frame image; taking a difference between the first feature coefficient and the second feature coefficient as the comparison result; wherein it comprises: obtaining a first hue feature value of the first intra-frame image; obtaining a first texture feature value of the first intra-frame image; and performing weighted calculation on the first hue feature value and the first texture feature value to obtain the first feature coefficient.

[0056] In the following, the specific configuration of the enhancement processing module 20 will be described in detail. The enhancement processing module 20 further comprises: taking a first frame image in the arbitrary video segment as an enhancement reference; and performing alignment enhancement on a plurality of frame images in the arbitrary video segment based on the video enhancement strategy and taking the enhancement reference as a constraint to obtain the arbitrary image.

[0057] In the following, the specific configuration of the multi-dimensional perception information obtaining module 30 will be described in detail. The multi-dimensional perception information obtaining module 30 further comprises: monitoring an infrared image of the predetermined target through the infrared perception component; monitoring a sonar image of the predetermined target through the sonar perception component; monitoring a radar image of the predetermined target through the radar perception component; and the infrared image, the sonar image and the radar image constitute the multi-dimensional perception information.

[0058] Hereinafter, the specific configuration of the dual-cascade processing module 40 will be described in detail. The dual-cascade processing module 40 further comprises: extracting a fractional order plan in the dual-cascade mechanism; processing the multi-dimensional perception information according to the fractional order plan to obtain target perception information; extracting a cross-modal fusion plan in the dual-cascade mechanism; processing the arbitrary image and the target perception information according to the cross-modal fusion plan to obtain the target fusion feature.

[0059] Hereinafter, the specific configuration of the dual-cascade processing module 40 will be described in detail. The dual-cascade processing module 40 further comprises: extracting a fractional order plan in the dual-cascade mechanism; processing the multi-dimensional perception information according to the fractional order plan to obtain target perception information; extracting a cross-modal fusion plan in the dual-cascade mechanism; processing the arbitrary image and the target perception information according to the cross-modal fusion plan to obtain the target fusion feature.

[0060] Hereinafter, the specific configuration of the dual-cascade processing module 40 will be described in detail. The dual-cascade processing module 40 further comprises: extracting a fractional order plan in the dual-cascade mechanism; processing the multi-dimensional perception information according to the fractional order plan to obtain target perception information; extracting a cross-modal fusion plan in the dual-cascade mechanism; processing the arbitrary image and the target perception information according to the cross-modal fusion plan to obtain the target fusion feature.

[0061] Hereinafter, the specific configuration of the dual-cascade processing module 40 will be described in detail. The dual-cascade processing module 40 further comprises: extracting a fractional order plan in the dual-cascade mechanism; processing the multi-dimensional perception information according to the fractional order plan to obtain target perception information; extracting a cross-modal fusion plan in the dual-cascade mechanism; processing the arbitrary image and the target perception information according to the cross-modal fusion plan to obtain the target fusion feature.

[0062] The image recognition and perception coordination video linkage tracking system provided by the embodiment of the present application can execute the image recognition and perception coordination video linkage tracking method provided by any embodiment of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0063] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or the server, the various units and modules included are only divided according to the functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific name of each functional unit is only for the convenience of mutual differentiation, and is not used to limit the protection scope of the present application.

[0064] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A video linkage tracking method based on image recognition and perception collaboration, characterized in that: include: Preprocessing the monitoring video stream to obtain a target video stream, and intercepting any video segment in the target video stream, wherein the arbitrary video segment corresponds to any shot and contains a predetermined target; Introducing a video enhancement strategy to enhance the arbitrary video segment to obtain an arbitrary image; activating a multidimensional sensor disposed on the arbitrary lens to perform dynamic and continuous perception monitoring of the predetermined target to obtain multidimensional perception information; Performing double cascade processing on the arbitrary image and the multi-dimensional perception information according to a double cascade mechanism to obtain a target fusion feature; The predetermined target is tracked in a linked manner based on the target fusion feature.

2. The video linkage tracking method of image recognition and perception collaboration as claimed in claim 1, characterized in that: Preprocessing the monitoring video stream to obtain a target video stream, and intercepting any video segment in the target video stream, including: Compressing the target video stream based on the Moving Picture Experts Group to obtain a compressed video, wherein the compressed video includes a first coding element and a second coding element, and the first coding element and the second coding element are adjacent coding elements to each other; sequentially extracting a first intra-frame image from the first coding element and a second intra-frame image from the second coding element; determining whether a comparison result obtained by comparing the first intra-frame image with the second intra-frame image satisfies a predetermined segmentation constraint; If so, the node between the first coding element and the second coding element is used as a split node; Segmenting the compressed video based on the segmentation node to obtain a segmentation result, wherein the segmentation result includes a plurality of video segments; Any one video segment from the multiple video segments is randomly extracted and recorded as the arbitrary video segment.

3. The video linkage tracking method of image recognition and perception collaboration as claimed in claim 2, characterized in that: Determining whether a comparison result obtained by comparing the first intra-frame image with the second intra-frame image satisfies a predetermined segmentation constraint includes: Obtaining a first characteristic coefficient of the first intra-frame image; Obtaining a second characteristic coefficient of the image within the second frame; taking the difference between the first characteristic coefficient and the second characteristic coefficient as the comparison result; Among them, include: Acquire a first hue characteristic value of the image within the first frame; Obtaining a first texture feature value of the image within the first frame; A weighted calculation is performed on the first hue feature value and the first texture feature value to obtain the first feature coefficient.

4. The video linkage tracking method of image recognition and perception collaboration as claimed in claim 1, characterized in that: Introducing a video enhancement strategy to enhance the arbitrary video segment to obtain an arbitrary image, including: Taking the first frame image in the arbitrary video segment as an enhancement benchmark; According to the video enhancement strategy, with the enhancement benchmark as a constraint, multiple frames of images in the arbitrary video segment are aligned and enhanced to obtain the arbitrary image.

5. The video linkage tracking method of image recognition and perception collaboration as claimed in claim 1, characterized in that: The multi-dimensional sensor includes an infrared sensor component, a sonar sensor component, and a radar sensor component. The multi-dimensional sensor deployed on any lens is activated to perform dynamic and continuous sensory monitoring of the predetermined target to obtain multi-dimensional sensory information, including: Obtaining an infrared image of the predetermined target through monitoring by the infrared sensing component; Obtaining a sonar image of the predetermined target through monitoring by the sonar sensing component; Obtaining a radar image of the predetermined target through monitoring by the radar sensing component; The infrared image, the sonar image, and the radar image constitute the multi-dimensional perception information.

6. The video linkage tracking method of image recognition and perception collaboration as claimed in claim 5, characterized in that: Performing double cascade processing on the arbitrary image and the multi-dimensional perception information according to a double cascade mechanism to obtain target fusion features, including: Extracting the fractional order plan in the dual cascade mechanism; Processing the multi-dimensional perception information according to the fractional-order plan to obtain target perception information; Extracting the cross-modal fusion plan in the dual-cascade mechanism; The arbitrary image and the target perception information are processed according to the cross-modal fusion plan to obtain the target fusion feature.

7. The video linkage tracking method of image recognition and perception collaboration as claimed in claim 6, characterized in that: Processing the multi-dimensional perception information according to the fractional-order plan to obtain target perception information includes: Acquire a target infrared image of the infrared image according to the fractional order plan; Acquire a target sonar image of the sonar image according to the fractional order plan; Acquire a target radar image of the radar image according to the fractional order plan; The target infrared image, the target sonar image and the target radar image constitute the target perception information; Wherein, acquiring the target infrared image of the infrared image according to the fractional order plan includes: Performing a two-dimensional Fourier transform on the infrared image according to the fractional order plan to obtain an infrared fractional order domain; The infrared fractional-order domain is filtered in combination with a predetermined rotation order, and the result is inversely transformed into the target infrared image.

8. The video linkage tracking method of image recognition and perception collaboration as claimed in claim 1, characterized in that: After performing linkage tracking of the predetermined target based on the target fusion feature, the method further includes: Acquiring a real-time linkage tracking image of the predetermined target; Sequentially acquiring a preceding frame image and a succeeding frame image of the real-time linkage tracking image; Perform matching analysis on the real-time linkage tracking image, the previous frame image, and the next frame image based on a three-frame object-shadow matching strategy to obtain a matching result; If the matching result meets a predetermined matching condition, the target contour of the predetermined target in the real-time linkage tracking image is used as a meta-learning benchmark; The meta-learning benchmark is used to continuously track the predetermined target.

9. The method for video linkage tracking based on image recognition and perception collaboration as claimed in claim 8, characterized in that: After sequentially acquiring the preceding frame image and the following frame image of the real-time linkage tracking image, the method further includes: Obtaining the real-time centroid of the real-time linkage tracking image; Obtaining a subsequent frame centroid of the subsequent frame image; Comparing the center of mass offset between the real-time center of mass and the subsequent frame center of mass; If the center of mass offset is less than a predetermined threshold, it is determined that the predetermined target is in a stationary state, and the linkage tracking is stopped.

10. A video linkage tracking system that combines image recognition and perception, characterized in that: The system is used to implement the video linkage tracking method of image recognition and perception collaboration according to any one of claims 1 to 9, and the system includes: An arbitrary video segment interception module is used to pre-process the monitoring video stream to obtain a target video stream, and intercept an arbitrary video segment in the target video stream, wherein the arbitrary video segment corresponds to an arbitrary shot and contains a predetermined target; An enhancement processing module, configured to introduce a video enhancement strategy to enhance the arbitrary video segment to obtain an arbitrary image; a multi-dimensional perception information acquisition module, configured to activate a multi-dimensional sensor disposed on any of the lenses to perform dynamic and continuous perception monitoring of the predetermined target, thereby obtaining multi-dimensional perception information; A dual-cascade processing module, configured to perform dual-cascade processing on the arbitrary image and the multi-dimensional perception information according to a dual-cascade mechanism to obtain a target fusion feature; A linkage tracking module is used to perform linkage tracking of the predetermined target based on the target fusion feature.

Citation Information

Patent Citations

  • Digital video processing method oriented to social security monitoring and apparatus thereof

    CN101483763A

  • Multi-attention RGBT target tracking method based on visible light guidance

    CN118365675A

  • Aviation video stream target identification processing method and system

    CN119151984A