Target Tracking Data Processing Method and Gigapixel Computational Imaging System

By introducing a tracking service module in the million-pixel computing imaging system to track local videos, generate fused position information and local fused position information, the problems of wasted computing resources and insufficient real-time performance in the existing system are solved, and the target detection box display on display terminals of different resolutions and types is realized.

CN118505750BActive Publication Date: 2025-08-05SUZHOU YIJI INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410659876.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-08-05
Estimated Expiration
2044-05-27

AI Technical Summary

Technical Problem

When the existing million-pixel computing imaging system outputs video data, it will waste computing resources for target detection and tracking of different types of display terminals, and it is difficult to meet real-time requirements, resulting in the drawing of the target detection box being unable to synchronize with video playback.

Method used

By introducing a tracking service module in the million-pixel computing imaging system to track local videos, generate fusion position information and local fusion position information, and send the target tracking information to the fusion service module and the on-screen service module, parallel processing of the target detection box is realized, suitable for display terminals of different resolutions and types.

Benefits of technology

It realizes timely and accurately displaying the object detection box on display terminals of different resolutions and types, saving computing resources, improving the efficiency of target tracking data processing, and meeting real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118505750B_ABST
    Figure CN118505750B_ABST
Patent Text Reader

Abstract

The present application relates to a target tracking data processing method and a billion-pixel computational imaging system. The method comprises: shooting multiple local videos through an array camera; a tracking service module performs target tracking on each frame of the local video image of each local video to obtain the local position information of the target object in the local video image; based on the coordinate mapping relationship corresponding to the local video image, the local position information is converted into fusion position information and local fusion position information, and according to the fusion position information and the local fusion position information, the target tracking information corresponding to the local video image is generated; the target tracking information is sent to the fusion service module and / or the upper screen service module. The present method can be applied to a billion-pixel computational imaging system that supports multi-resolution fusion video playback, so as to achieve the purpose of timely and accurate display of target detection frames on display terminals of different resolutions and different types, and saving computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computational imaging technology, and in particular to a target tracking data processing method and a billion-pixel computational imaging system. Background Art

[0002] In order to overcome the problem that a single camera device cannot take into account both a large field of view and high definition, a billion-pixel computational imaging system based on array cameras has emerged. It uses an array camera containing multiple lenses to collect multiple local videos, and then uses relevant algorithms to perform computational processing such as stitching and fusion on each local video. It can obtain ultra-high-definition fused video with a large field of view, billion-pixel level or even billion-pixel level. It can be applied to technical fields with large-scene ultra-high-definition video requirements, such as video surveillance in the security field, digital cultural and creative technology fields, augmented reality technology fields, etc.

[0003] In actual applications, the display terminals used in different video playback scenarios have different resolutions and types (such as spliced screens or single screens), and the processing power of a single graphics card is limited (billion-pixel fusion video has exceeded the current encoding, decoding and rendering capabilities of a single graphics card). Therefore, a billion-pixel computational imaging system that can support multi-resolution fusion video playback has emerged. It can obtain fusion video data through the fusion service module, which is suitable for playback on a single-screen display terminal with lower resolution. It can also obtain multi-channel local fusion video data through the upper screen service module, which is suitable for the display modules of the spliced display terminal to splice and play back complete, ultra-high-resolution billion-pixel fusion video.

[0004] When using billion-pixel computational imaging systems for video surveillance in the security field, there is usually a need for intelligent target tracking so that the detection frame of each target can be marked in the fused video played on the display terminal to improve the intelligent monitoring effect. If target detection and tracking are performed separately for each type of video data (including fused video and local fused video) output by the above-mentioned billion-pixel computational imaging system, a lot of computing resources will be wasted, and the processing efficiency will be difficult to meet the real-time requirements, resulting in the inability to draw the target detection frame in sync with the video playback. Therefore, there is an urgent need for a target tracking data processing method that is compatible with the above-mentioned billion-pixel computational imaging system that supports multi-resolution fused video playback, so as to achieve the purpose of timely and accurate display of target detection frames on display terminals of different resolutions and types, and to save computing resources. Summary of the Invention

[0005] Based on this, the present application provides a target tracking data processing method and a billion-pixel computational imaging system, which can be applied to a billion-pixel computational imaging system that supports multi-resolution fusion video playback, so as to achieve the purpose of timely and accurate display of target detection frames on display terminals of different resolutions and types, and save computing resources.

[0006] In a first aspect, the present application provides a target tracking data processing method. The method is applied to a billion-pixel computational imaging system, wherein the billion-pixel computational imaging system includes an array camera, a fusion service module, an on-screen service module, and a tracking service module. The method includes:

[0007] Shooting multiple local videos by using multiple local cameras of the array camera;

[0008] The tracking service module performs target tracking on each frame of the local video image of each local video to obtain local position information of the target object in the local video image;

[0009] Based on the coordinate mapping relationship corresponding to the local video image, the local position information is converted into fused position information and local fused position information, and target tracking information corresponding to the local video image is generated according to the fused position information and the local fused position information;

[0010] The target tracking information is sent to the fusion service module and / or the upper screen service module, so that the fusion service module adds the target detection frame of the target object in the fused video according to the fusion position information contained in the target tracking information, and the upper screen service module adds the target detection frame of the target object in the corresponding local fusion video according to the local fusion position information contained in the target tracking information.

[0011] In one embodiment, the target tracking information includes main image target information and cross-screen target information; and generating the target tracking information corresponding to the local video image based on the fusion position information and the local fusion position information includes:

[0012] Determine the local camera identifier corresponding to the local video image, the timestamp of the local video image, the fusion position information of the target object, and the local fusion position information as the main image target information;

[0013] In the case where the target object spans other local fusion video images, determining the local camera identifiers corresponding to the other local fusion video images spanned by the target object as the cross-screen camera identifier, converting the local fusion position information of the target object into cross-screen position information in the coordinate system of the other local fusion video images spanned by the target object, and determining the cross-screen camera identifier and the cross-screen position information as the cross-screen target information;

[0014] Target tracking information corresponding to the local video image is generated based on the main image target information and the cross-screen target information.

[0015] In one embodiment, the method further comprises:

[0016] The fusion service module receives the target tracking information sent by the tracking service module;

[0017] According to the local camera identifier and timestamp in the main image target information contained in the target tracking information, the target local video image to which the target tracking information belongs is determined, and according to the fusion position information of the target object in the main image target information, the target detection frame of the target object is added in the fusion video frame corresponding to the target local video image.

[0018] In one embodiment, the method further comprises:

[0019] The upper screen service module receives the target tracking information sent by the tracking service module;

[0020] Determine the target rendering thread corresponding to the target tracking information according to the local camera identifier in the main image target information contained in the target tracking information, and determine the cross-screen rendering thread corresponding to the target tracking information according to the cross-screen camera identifier in the cross-screen target information contained therein;

[0021] Adding, by the target rendering thread, a main image target detection frame of the target object in the local fusion video frame corresponding to the local camera identifier and the timestamp according to the local fusion position information of the target object in the main image target information;

[0022] The cross-screen rendering thread adds a cross-screen target detection frame of the target object in the local fusion video frame in which the cross-screen camera identifier and the timestamp match according to the cross-screen position information of the target object in the cross-screen target information.

[0023] In one embodiment, the on-screen service module includes a receiving thread, a transmission thread, and multiple rendering threads, one of the rendering threads corresponding to processing one channel of the local video and corresponding to one display module of the splicing display terminal; the method further includes:

[0024] The receiving thread receives the target tracking information sent by the tracking service module, and stores the target tracking information in a tracking information queue corresponding to the local camera identifier contained therein;

[0025] The transmission thread obtains target tracking information from the tracking information queue, and stores the target tracking information in a first buffer of a target rendering thread corresponding to the tracking information queue, and stores the target tracking information in a second buffer of a cross-screen rendering thread corresponding to a cross-screen camera identifier contained therein;

[0026] The target rendering thread obtains the target tracking information from the first buffer, and adds a main image target detection frame of the target object in the local fusion video frame corresponding to the local camera identifier and the timestamp according to the local fusion position information of the target object in the main image target information contained therein;

[0027] The cross-screen rendering thread obtains the target tracking information from the second buffer, and adds a cross-screen target detection frame of the target object in the local fusion video frame that matches the cross-screen camera identifier and the timestamp based on the cross-screen position information of the target object contained therein.

[0028] In a second aspect, the present application also provides a billion-pixel computational imaging system. The billion-pixel computational imaging system includes an array camera, a fusion service module, an on-screen service module, and a tracking service module;

[0029] The array camera is used to shoot multiple local videos through multiple local cameras, and send the local videos to the fusion service module, the on-screen service module and the tracking service module;

[0030] The tracking service module is used to perform target tracking on each frame of the local video image of each local video, obtain local position information of the detected target object in the local video image, convert the local position information into fusion position information and local fusion position information based on the coordinate mapping relationship corresponding to the local video image, generate target tracking information corresponding to the local video image according to the fusion position information and the local fusion position information, and send the target tracking information to the fusion service module and the on-screen service module;

[0031] The fusion service module is used to splice and fuse the multiple local videos to obtain a fused video, and add a target detection frame of the target object in the fused video according to the fusion position information included in the target tracking information;

[0032] The upper screen service module is used to perform video frame conversion and rendering on each of the local videos through multiple graphics cards to obtain a local fusion video corresponding to each of the local videos, and add a target detection frame of the target object in the corresponding local fusion video based on the local fusion position information contained in the target tracking information.

[0033] In one embodiment, the target tracking information includes main image target information and cross-screen target information;

[0034] The tracking service module is further configured to determine the local camera identifier corresponding to the local video image, the timestamp of the local video image, the fusion position information of the target object, and the local fusion position information as the main image target information;

[0035] In the case where the target object spans other local fusion video images, determining the local camera identifiers corresponding to the other local fusion video images spanned by the target object as the cross-screen camera identifier, converting the local fusion position information of the target object into cross-screen position information in the coordinate system of the other local fusion video images spanned by the target object, and determining the cross-screen camera identifier and the cross-screen position information as the cross-screen target information;

[0036] Target tracking information corresponding to the local video image is generated based on the main image target information and the cross-screen target information.

[0037] In one of the embodiments, the fusion service module is also used to determine the target local video image to which the target tracking information belongs based on the local camera identifier and timestamp in the main image target information contained in the target tracking information, and add the target detection frame of the target object in the fused video frame corresponding to the target local video image based on the fusion position information of the target object in the main image target information.

[0038] In one embodiment, the upper screen service module is further used to determine the target rendering thread corresponding to the target tracking information according to the local camera identifier in the main image target information contained in the target tracking information, and determine the cross-screen rendering thread corresponding to the target tracking information according to the cross-screen camera identifier in the cross-screen target information contained therein;

[0039] Adding, by the target rendering thread, a main image target detection frame of the target object in the local fusion video frame corresponding to the local camera identifier and the timestamp according to the local fusion position information of the target object in the main image target information;

[0040] The cross-screen rendering thread adds a cross-screen target detection frame of the target object in the local fusion video frame in which the cross-screen camera identifier and the timestamp match according to the cross-screen position information of the target object in the cross-screen target information.

[0041] In one embodiment, the on-screen service module includes a receiving thread, a transmission thread, and multiple rendering threads, one of the rendering threads processes one channel of the local video and corresponds to one display module of the splicing display terminal;

[0042] The receiving thread is used to receive the target tracking information sent by the tracking service module, and store the target tracking information in a tracking information queue corresponding to the local camera identifier contained therein;

[0043] The transmission thread is used to obtain target tracking information from the tracking information queue, and store the target tracking information in a first buffer of a target rendering thread corresponding to the tracking information queue, and store the target tracking information in a second buffer of a cross-screen rendering thread corresponding to a cross-screen camera identifier contained therein;

[0044] The target rendering thread is configured to obtain the target tracking information from the first buffer, and add a main image target detection frame of the target object in the local fusion video frame corresponding to the local camera identifier and the timestamp based on the local fusion position information of the target object in the main image target information contained therein;

[0045] The cross-screen rendering thread is used to obtain the target tracking information from the second buffer, and add a cross-screen target detection box of the target object in the local fusion video frame that matches the cross-screen camera identifier and the timestamp based on the cross-screen position information of the target object in the cross-screen target information contained therein.

[0046] The above-mentioned target tracking data processing method and billion-level pixel computational imaging system, through the tracking service module, performs target tracking on each local video image captured by the array camera, and obtains the target tracking information corresponding to each local video image. The target tracking information includes the fusion position information and local fusion position information of each target object in the local video image, which can simultaneously meet the tracking data usage requirements of the fusion service module and the upper screen service module, and realize the decoupling of target tracking data processing and fusion video data processing. The two can be processed in parallel, improving the efficiency of target tracking data processing, and avoiding the need to perform target detection and tracking separately for different types of video data output by the billion-level pixel computational imaging system, saving computing resources. Therefore, this method can be applied to billion-level pixel computational imaging systems that support multi-resolution fusion video playback, and can achieve the purpose of timely and accurate display of target detection frames on display terminals of different resolutions and different types, and saving computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 A schematic diagram of the structure of a billion-pixel computational imaging system in an example;

[0048] Figure 2 1 is a flow chart of a target tracking data processing method according to an embodiment;

[0049] Figure 3 is a position relationship diagram of each local video image in an example;

[0050] Figure 4a A schematic diagram of an example in which a local video image includes a cross-screen target;

[0051] Figure 4b A schematic diagram of display modules stitching together and displaying a cross-screen target in an example;

[0052] Figure 5 FIG. 4 is a flow chart of a target tracking data processing method in another embodiment. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0054] The target tracking data processing method provided in the embodiment of the present application can be applied to Figure 1 Exascale pixel computational imaging system 100 is shown. Exascale pixel computational imaging system 100 includes an array camera 102, a fusion service module 104, an on-screen service module 106, and a tracking service module 108. Array camera 102 communicates with fusion service module 104, on-screen service module 106, and tracking service module 108, respectively, to exchange data. Fusion service module 104, on-screen service module 106, and tracking service module 108 can be implemented via a server.

[0055] The first display terminal 200 can be a single-screen display terminal with a relatively small resolution (such as not more than the target resolution of 8K, and the target resolution can be specifically determined based on the current mainstream single graphics card encoding and rendering capabilities), such as a personal computer, a mobile terminal, etc. The fusion service module 104 can be connected to the first display terminal 200 through a network. The second display terminal 300 can be a spliced display screen with a resolution comparable to or even greater than the resolution of the billion-pixel fusion video, so as to fully play the un-downsampled ultra-high-definition fusion video. The display modules included in the second display terminal 300 can correspond one-to-one to the local cameras or local videos of the array camera 102, and the resolution of each display module can be the same as the resolution of the local video (such as both are 4K). The upper screen service module 106 can be electrically connected to the second display terminal 300 through a data connection line, so that the upper screen service module 106 can directly transmit the rendered local fusion video (without encoding) to the corresponding display module for playback.

[0056] The fusion service module 104 can obtain all or part of the local video captured by the array camera 102 (for example, the array camera includes 3 rows and 6 columns with a total of 18 telephoto lenses, but only the images captured by the first 5 columns of lenses are needed, so only the local video captured by the 5 columns of lenses can be obtained), and then use the image stitching algorithm or the pre-configured image stitching model to perform stitching and fusion processing on the acquired local videos to obtain a large-scale billion-pixel fused video. Then, the fusion service module 104 can perform corresponding sampling processing (such as the following sampling) on the billion-pixel fused video according to the video resolution requirements of the first display terminal 200 to obtain a first fused video that matches the resolution of the first display terminal 200 (usually a fused video with a resolution less than 8K that can be processed by a single graphics card). The fusion service module 104 can render and encode the sampled fused video data (such as using the H264 encoding format), and send the encoded data to the first display terminal 200 via the network for decoding and playback.

[0057] The on-screen service module 106 can use multiple graphics cards to perform video frame transformation and rendering on each local video. One graphics card can process one or more (e.g., four) local videos to obtain the local fused video corresponding to each local video in the billion-pixel fused video. The video frame transformation includes a homography transformation, which transforms each local video image to the same field of view plane, and then crops each transformed image to obtain a local fused video image. Each local fused video image is sequentially spliced according to its relative position (e.g., the relative position of the local video image corresponding to the local camera A11 located in row 1 and column 1 is adjacent to the local camera A12 located in row 1 and column 2 is left and right), thereby obtaining a complete billion-pixel fused video. An image splicing model can be constructed in advance based on each local sample image captured by the array camera to obtain the homography transformation matrix and cropping parameters corresponding to each local camera. Thus, the on-screen service module 106 can use the homography transformation matrix and cropping parameters corresponding to the local video to perform video frame transformation on each frame of the local video. The transformed video frame image is then rendered as a local fusion video and transmitted to the display module corresponding to the local video in the second display terminal 300 for playback. Each display module can be spliced according to the relative position of each local video, thereby realizing the complete picture of the spliced playback of the billion-pixel fusion video.

[0058] Therefore, the billion-pixel computational imaging system can output fused video or partially fused video according to the video playback requirements of different resolutions and different types of display terminals, thereby achieving the purpose of simultaneous or time-sharing playback of fused videos of different resolutions.

[0059] However, due to the different types of video data output by the billion-pixel computational imaging system 100, target tracking data that can adapt to different types of video data is required so that target detection frames can be marked promptly and accurately in videos played on display terminals of different resolutions and types. If target detection and tracking are performed separately for each type of video data (including fused video and partially fused video) output by the billion-pixel computational imaging system, such as using a target detection algorithm and a target tracking algorithm to perform target detection and tracking on the fused video output by the fusion service module and each partially fused video output by the upper screen service module, a lot of computing resources will be wasted, and target detection and tracking will need to be performed on the output video data after the video data processing is completed. The target tracking data processing efficiency is low, and it is difficult to meet the real-time requirements of video playback and target detection frame marking.

[0060] In order to be able to timely and accurately mark target detection frames when playing videos on display terminals of different resolutions and types, and to save computing resources, the present application proposes a target tracking data processing method. By decoupling video data processing and target tracking data processing, the two can be processed in parallel, thereby improving the target tracking data processing efficiency. The processed target tracking data can simultaneously meet the usage requirements of various types of video data, effectively save computing resources, and support timely and accurate marking of target detection frames when playing videos on display terminals of different resolutions and types.

[0061] In one embodiment, Figure 2 As shown, a target tracking data processing method is provided, which includes the following steps:

[0062] Step 201: shoot multiple local videos using multiple local cameras of an array camera.

[0063] Among them, the array camera 102 may include multiple local cameras (usually telephoto lens cameras) with different shooting angles, such as the common 3 rows and 5 columns, 3 rows and 6 columns, 1 row and 10 columns, etc. array cameras. Each local camera shoots a local video (usually 4K video) to collect local area images in the shooting scene. Usually, the local video images shot by two lenses with adjacent shooting fields have overlapping areas. An image stitching algorithm can be used to stitch and fuse the local video images based on the overlapping areas to obtain a panoramic fusion image with a larger field of view. The array camera 104 may also include a short-focus lens (usually located in the middle) whose shooting field of view can cover the entire shooting scene to shoot panoramic images to assist in building image stitching models and color correction, etc. The processor of the array camera 102 can encode each local video and send it to the fusion service module 104, the on-screen service module 106 and the tracking service module 108 for subsequent processing. Figure 3The figure shows the relative position relationship of each local video image captured by the array camera with 3 rows and 6 columns. The local video images in adjacent positions have overlapping areas. A11 represents the local video image corresponding to the local camera in the first row and first column. For the convenience of corresponding description, A11 can be used as the local camera identifier. In addition, the local camera identifier and the timestamp of the local video image can be used as the index information of a frame of local video image to uniquely identify a frame of image. It can be understood that Figure 3 The overlapping areas shown are examples only, and the actual overlapping areas may have irregular shapes.

[0064] In step 202 , the tracking service module performs target tracking on each frame of the local video image of each local video to obtain local position information of the target object in the local video image.

[0065] In implementation, the tracking service module 108 may use a target tracking algorithm to track the target in each frame of the local video image. For example, the tracking service module 108 may first use a target detection algorithm (YOLO, Faster R-CNN, etc.) to perform target detection on each acquired frame of the local video image, and then use a data association algorithm (such as the Hungarian algorithm) to associate the target object detected in the current frame with the target object detected in the historical frame to obtain the local position information of each target object contained in the current local video image (using the globally unique identifier of the tracking target) in the current local video image.

[0066] Step 203 : based on the coordinate mapping relationship corresponding to the local video image, convert the local position information into fused position information and local fused position information, and generate target tracking information corresponding to the local video image according to the fused position information and the local fused position information.

[0067] Among them, the fusion position information refers to the position information in the fusion video image coordinate system, and the local fusion position information refers to the position information in the coordinate system of the local fusion video image corresponding to the local video image in the fusion video image, and normalized coordinates are usually used. For example, for the fusion video image coordinate system, the coordinates of the upper left corner point of the fusion video image (obtained by splicing and fusing the local video images) are (0, 0), and the coordinates of the lower right corner point are (1, 1). For the local fusion video image coordinate system, the coordinates of the upper left corner point of the local fusion video image (obtained after homography transformation and cropping the local video image) are (0, 0), and the coordinates of the lower right corner point are (1, 1). Each local fusion video image has its own coordinate system. Usually, the upper left corner point of the local fusion video image is used as the coordinate system. The tracking service module 108 can convert the local position information of each target object into fusion position information and local fusion position information according to the pre-established coordinate mapping relationship, and then generate the target tracking information corresponding to the local video image. The target tracking information can include the local camera identifier of the local video image, timestamp, fusion position information of each target object, and local fusion position information. For example, target tracking information can be expressed as:

[0068] {matID:A11,timestamp:t1,obj:ID1,position1:(x1′,y1′,w′,h′),position2:(x1″,y1″,w″,h″)}.

[0069] Where matID is the local camera ID, timestamp is the timestamp, obj is the target object ID, position1 is the fused position information corresponding to the target object ID, and position2 is the local fused position information corresponding to the target object ID. This example indicates that the fused position information of target object ID1 detected in the local video image captured by local camera A11 at timestamp t1 is (x1′, y1′, w′, h′), and the local fused position information is (x1″, y1″, w″, h″).

[0070] Step 204: Send the target tracking information to the fusion service module and / or the on-screen service module.

[0071] In implementation, the tracking service module 108 may track each local video image separately to obtain target tracking information corresponding to each local video image. Upon receiving a request from the fusion service module 104 and / or the on-screen service module 106 for obtaining target tracking information, the tracking service module 108 may send the currently obtained target tracking information to the fusion service module 104 and / or the on-screen service module 106.

[0072] The fusion service module 104 receives the target tracking information and can determine the fusion position information of each target object contained in the target tracking information, and then add a target detection frame to the corresponding position in the currently processed fusion video image, render the fusion video image and the target detection frame together, obtain a fusion video containing the target detection frame, and send it to the display terminal for playback.

[0073] The upper screen service module 106 receives the target tracking information and can determine the local fusion position information of each target object contained in the target tracking information, and then add a target detection frame to the corresponding position in the currently processed local fusion video image, render the local fusion video image and the target detection frame together, and obtain a local fusion video containing the target detection frame, and then send each local fusion video to each display module of the splicing display terminal for splicing and playback.

[0074] Among them, since the speed at which the tracking service module processes the target tracking data is not completely consistent with the speed at which the fusion service module and the upper screen service module process the video image, there may be a sequence, so the target tracking data may also include the timestamp of the local video image. After the fusion service module and / or the upper screen service module receives the target tracking information, it can first store the target tracking information in a queue. When the fusion service module processes the current frame fused video image, or when the upper screen service module processes the current frame local fused video image, it can obtain the target tracking information with a timestamp matching from the queue based on the timestamp of the local video image corresponding to the current frame video image, so that the target tracking information matches the current fused video frame and / or the current frame local fused video image, thereby ensuring the accuracy of the target detection frame annotation position.

[0075] The above-mentioned target tracking data processing method, through the tracking service module, performs target tracking on each local video image captured by the array camera, and obtains the target tracking information corresponding to each local video image. The target tracking information includes the fusion position information and local fusion position information of each target object in the local video image, which can simultaneously meet the tracking data usage requirements of the fusion service module and the upper screen service module, and realizes the decoupling of target tracking data processing and fusion video data processing. The two can be processed in parallel, improving the efficiency of target tracking data processing, and avoiding the need to perform target detection and tracking separately for different types of video data output by the billion-pixel computing imaging system, saving computing resources. This method performs target detection and tracking on the local video images captured by the array camera. Compared with performing target detection and tracking on the fusion video or local fusion video that may be deformed after processing, and the number of effective pixels of the target object in the fusion video after sampling processing is smaller, the target detection and tracking accuracy of this method is higher. Therefore, this method can be applied to the billion-pixel computing imaging system that supports multi-resolution fusion video playback, and can achieve the purpose of timely and accurate display of target detection frames on display terminals of different resolutions and different types, and saving computing resources.

[0076] During the research and development process, the applicant discovered that when displaying each local fusion video through a spliced display terminal, the target detection frame of the target object that spans adjacent screens (which can be called a cross-screen target) is not displayed completely. For example, Figure 4a The figure shows the local video images taken by four local cameras, among which the local video image taken by local camera A11 contains the complete image of target object T1, and the local video image taken by local camera A12 contains the partial image of target object T1. The dotted lines in the figure represent the cropping lines or stitching lines. The upper screen service module performs homography transformation and cropping on each local video image to obtain a fused video image, and obtains the target tracking information sent by the tracking service module. Based on the local fused position information of target object T1 in the target tracking information and the fused video image, a local fused video is rendered and sent to the corresponding display module for playback. The playback effect is as follows: Figure 4b As shown, target object T1 spans display modules S11 and S12, representing a cross-screen target. However, only a portion of the detection frame of target object T1 is annotated in the local fusion video played by display module S11, corresponding to local camera A11. However, the other portion of the detection frame of target object T1 is not annotated in the local fusion video played by display module S12, corresponding to local camera A12. This results in an incomplete display of the detection frame of target object T1, affecting target monitoring effectiveness.

[0077] The possible reasons for this problem are: first, when the local video image corresponding to the cross-screen image contains fewer valid pixels of the target object, the target object is not detected in the local video image corresponding to the cross-screen image; second, even if the target object located in the overlapping area is detected in multiple local video images, during target tracking, the detection frames of the same target object are usually deduplicated and merged (to avoid the fusion service module repeatedly marking multiple detection frames for the same target object in the fused video), so that the position information of the target object (such as the position information of the merged detection frame) is retained only in the target tracking information of one local video image, usually the local video image to which the center point of the detection frame (merged detection frame) of the target object belongs.

[0078] In order to solve the problem of incomplete display of cross-screen target detection frame, the present application also provides another embodiment of the target tracking data processing method, which differs from the previous embodiment in the process of generating target tracking information in step 203. In this embodiment, the target tracking information includes the main image target information and the cross-screen target information, such as Figure 5 As shown, the process of generating target tracking information in step 203 specifically includes the following steps:

[0079] Step 501 : The local camera identifier corresponding to the local video image, the timestamp of the local video image, the fusion position information of the target object and the local fusion position information are determined as the main image target information.

[0080] Step 502: When the target object spans other local fusion video images, the local camera identifier corresponding to the other local fusion video images it spans is determined as the cross-screen camera identifier, and the local fusion position information of the target object is converted into cross-screen position information in the coordinate system of the other local fusion video images it spans, and the cross-screen camera identifier and cross-screen position information are determined as cross-screen target information.

[0081] In implementation, the tracking service module 108 can determine whether the target object is a cross-screen target, that is, whether it spans other local fusion video images. For example, the tracking service module can determine whether it is a cross-screen target based on the local fusion position information of the target object. For example, the coordinate range of the local fusion video image is (0,0,1,1). If the local fusion position information of the target object in the local fusion video image coordinate system is (0.6,0.2,1.1,0.4), where (0.6,0.2) is the coordinate of the upper left corner point of the target object's detection frame, and (1.1,0.4) is the coordinate of the lower right corner point of the target object's detection frame, it can be seen that the coordinates of some pixel points of the target object's detection frame exceed the coordinate range of the local fusion video image (where the horizontal coordinate 1.1 of the lower right corner point is greater than the maximum value of the horizontal coordinate 1). Therefore, it can be determined that the target object is a cross-screen target.

[0082] When the target object crosses other local fusion video images, the tracking service module can determine the local camera identifiers corresponding to the other local fusion video images it crosses as cross-screen camera identifiers. For example, if the minimum horizontal coordinate of the target object is less than the minimum horizontal coordinate value of the coordinate system (less than 0, a negative number), and the vertical coordinate does not exceed the coordinate range, then the target object crosses the left adjacent image of the local fusion video image; if the maximum horizontal coordinate of the target object is greater than the maximum horizontal coordinate value of the coordinate system (greater than 1), and the vertical coordinate does not exceed the range, then the target object crosses the right adjacent image of the local fusion video image; if the minimum vertical coordinate of the target object is less than the minimum vertical coordinate value of the coordinate system (less than 0, a negative number), and the horizontal coordinate does not exceed the range, then the target object crosses the upper adjacent image of the local fusion video image; if the maximum vertical coordinate of the target object is greater than the maximum vertical coordinate value of the coordinate system (greater than 1), and the horizontal coordinate does not exceed the range, then the target object crosses the lower adjacent image of the local fusion video image. The target object may also cross multiple images in the four neighboring areas, which can be specifically determined based on the local fusion position information of the target object.

[0083] Then, the local fusion position information of the target object can be converted into cross-screen position information in the coordinate system of other local fusion video images it spans, and the cross-screen camera identifier and cross-screen position information are determined as cross-screen target information. Figure 4a In the example, the local fusion position information (0.6, 0.2, 1.1, 0.4) of the target object T1 in the local fusion video image A11 coordinate system can be converted to the cross-screen position information (-0.4, 0.2, 0.1, 0.4) in the local fusion video image A12 coordinate system, that is, converted to the coordinate system with the upper left corner point of the local fusion video image A12 as (0, 0) and the lower right corner point as (1, 1).

[0084] Step 503: Generate target tracking information corresponding to the partial video image based on the main image target information and the cross-screen target information.

[0085] During implementation, the tracking service module can generate target tracking information corresponding to the local video image A11 together with the main image target information and the cross-screen target information, so that the fusion service module can mark the target detection frame in the fused video according to the main image target information contained therein, and the upper screen service module can mark the target detection frame in the corresponding local fused video image according to the main image target information and the cross-screen target information contained therein. It can be understood that the local fusion position information and the cross-screen position information of the cross-screen target both contain some pixel coordinates that exceed the coordinate range of this coordinate system. For pixels that exceed the range, they will not be rendered and displayed in this local fusion video. Only the pixel coordinates within the coordinate range are displayed in this local fusion video. Therefore, the main image target detection frame and the cross-screen target detection frame are only partial detection frames of the target object. When the two are spliced and displayed through the corresponding display modules, a complete target detection frame can be synthesized.

[0086] In some embodiments, after the fusion service module receives the target tracking information sent by the tracking service module, it can determine the target local video image to which the target tracking information belongs based on the local camera identifier and timestamp in the main image target information contained in the target tracking information, and add the target detection box of the target object in the fused video frame corresponding to the target local video image (usually the currently processed fused video frame) based on the fusion position information of the target object in the main image target information.

[0087] After the upper-screen service module receives the target tracking information sent by the tracking service module, it can determine the target rendering thread corresponding to the target tracking information based on the local camera identifier in the main-image target information contained in the target tracking information, and determine the cross-screen rendering thread corresponding to the target tracking information based on the cross-screen camera identifier in the cross-screen target information contained therein. Then, the target rendering thread adds the main-image target detection frame of the target object to the local fusion video frame corresponding to the local camera identifier and timestamp based on the local fusion position information of the target object in the main-image target information; and adds the cross-screen target detection frame of the target object to the local fusion video frame whose cross-screen camera identifier and timestamp match based on the cross-screen position information of the target object in the cross-screen target information through the cross-screen rendering thread. In this way, each display module can splice and display the complete fusion video and the target detection frame of each target object.

[0088] In some embodiments, the upper screen service module includes a receiving thread, a transmission thread, and multiple rendering threads. Each rendering thread processes a local video channel and corresponds to a display module of the spliced display terminal. The receiving thread receives target tracking information sent by the tracking service module and stores the target tracking information in a tracking information queue corresponding to the local camera identifier contained therein. Each local video channel corresponds to a single tracking information queue.

[0089] The transmission thread obtains target tracking information from each tracking information queue, stores the target tracking information in the first buffer of the target rendering thread corresponding to the tracking information queue, and stores the target tracking information in the second buffer of the cross-screen rendering thread corresponding to the cross-screen camera identifier contained in the tracking information queue. For example, the transmission thread can store the target tracking information obtained from the tracking message queue corresponding to local camera A11 in the first buffer of the rendering thread corresponding to local camera A11, and store it in the second buffer of the rendering thread corresponding to cross-screen local camera A12.

[0090] The target rendering thread (such as A11) obtains the target tracking information from the first buffer, and adds the main image target detection box of the target object to the local fusion video frame corresponding to the local camera identifier and timestamp based on the local fusion position information of the target object in the main image target information contained therein.

[0091] The cross-screen rendering thread (such as A12) obtains the target tracking information from the second buffer, and adds the cross-screen target detection box of the target object in the local fusion video frame whose cross-screen camera identifier and timestamp match based on the cross-screen position information of the target object contained in the cross-screen target information.

[0092] In one implementation, in order to improve processing efficiency, the transmission thread can store the target tracking information in the second buffer of all other rendering threads (non-target rendering threads), so that each rendering thread can determine by itself whether the cross-screen camera identifier it contains is consistent with the camera identifier corresponding to the current rendering thread. If consistent, a cross-screen target detection frame is added to the local fusion video processed by the rendering thread according to the cross-screen position information corresponding to the cross-screen camera identifier. Compared with the transmission thread's cross-screen camera identifier recognition and directional transmission of all target tracking information, the distributed recognition processing efficiency of each rendering thread is higher.

[0093] It is understandable that each rendering thread may have a first buffer and a second buffer. The first buffer is used to store target tracking information corresponding to the local video processed by the rendering thread, that is, the target tracking information stored in the first buffer has the same local camera identifier in the main image target information as the camera identifier of the local video processed by the rendering thread. The second buffer is used to store target tracking information corresponding to other local videos, that is, the target tracking information stored in the second buffer has the same local camera identifier in the main image target information as the camera identifier of the local video processed by the rendering thread, and the cross-screen camera identifier in the cross-screen target information may be the same as or different from the camera identifier of the local video processed by the rendering thread.

[0094] Specifically, the transmission thread is used to transmit the local video image data in the local decoded frame queue to the rendering thread for image transformation, cropping, rendering, and other processing, and store the target tracking information corresponding to the local video image data in the tracking information queue in the first buffer of the target rendering thread and the second buffer of other rendering threads. During (or after) processing the local video image of the current frame, the rendering thread can obtain the target tracking information of the local video image of the current frame from the first buffer, parse and obtain the local fusion position information of each target object contained in its main image target information, and obtain the target tracking information of other associated images of the local video image of the current frame (other local video images belonging to the same frame fusion video image) from the second buffer. If the cross-screen camera identifier contained in the other target tracking information is the same as the camera identifier corresponding to the current rendering thread, the cross-screen position information corresponding to the cross-screen camera identifier is determined, and then based on the local fusion position information and the cross-screen position information, the target detection box is added to the corresponding position in the local fusion video processed by the rendering thread. Among them, the local video image data stored in the local decoded frame queue is the local video image obtained by the decoding thread of the upper screen service module and stored in the queue after decoding.

[0095] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0096] An embodiment of the present application also provides a billion-pixel computational imaging system, including an array camera, a fusion service module, an on-screen service module, and a tracking service module.

[0097] The array camera is used to shoot multiple local videos through multiple local cameras and send the local videos to the fusion service module, the on-screen service module and the tracking service module.

[0098] The tracking service module is used to track the target of each frame of the local video image of each local video, obtain the local position information of the detected target object in the local video image, and convert the local position information into fusion position information and local fusion position information based on the coordinate mapping relationship corresponding to the local video image. According to the fusion position information and the local fusion position information, the target tracking information corresponding to the local video image is generated, and the target tracking information is sent to the fusion service module and the upper screen service module.

[0099] The fusion service module is used to stitch and fuse multiple local videos to obtain a fused video, and add a target detection frame of the target object in the fused video based on the fusion position information contained in the target tracking information.

[0100] The upper screen service module is used to perform video frame transformation and rendering on each local video through multiple graphics cards to obtain the local fusion video corresponding to each local video, and add the target detection frame of the target object in the corresponding local fusion video according to the local fusion position information contained in the target tracking information.

[0101] In one embodiment, the target tracking information includes main image target information and cross-screen target information.

[0102] The tracking service module is further used to determine the local camera identifier corresponding to the local video image, the timestamp of the local video image, the fusion position information of the target object and the local fusion position information as the main image target information.

[0103] When the target object spans other local fusion video images, the local camera identifier corresponding to the other local fusion video images it spans is determined as the cross-screen camera identifier, and the local fusion position information of the target object is converted into cross-screen position information in the coordinate system of the other local fusion video images it spans, and the cross-screen camera identifier and cross-screen position information are determined as cross-screen target information.

[0104] According to the main image target information and the cross-screen target information, the target tracking information corresponding to the local video image is generated.

[0105] In one embodiment, the fusion service module is also used to determine the target local video image to which the target tracking information belongs based on the local camera identifier and timestamp in the main image target information contained in the target tracking information, and add a target detection frame of the target object in the fused video frame corresponding to the target local video image based on the fusion position information of the target object in the main image target information.

[0106] In one embodiment, the upper-screen service module is also used to determine the target rendering thread corresponding to the target tracking information based on the local camera identifier in the main image target information contained in the target tracking information, and determine the cross-screen rendering thread corresponding to the target tracking information based on the cross-screen camera identifier in the cross-screen target information contained therein.

[0107] The target rendering thread adds the main image target detection frame of the target object to the local fusion video frame corresponding to the local camera identifier and timestamp based on the local fusion position information of the target object in the main image target information.

[0108] The cross-screen rendering thread adds a cross-screen target detection frame of the target object to the local fusion video frame whose cross-screen camera identifier and timestamp match according to the cross-screen position information of the target object in the cross-screen target information.

[0109] In one embodiment, the upper screen service module includes a receiving thread, a transmission thread and multiple rendering threads. One rendering thread processes one local video and corresponds to one display module of the splicing display terminal.

[0110] The receiving thread is used to receive the target tracking information sent by the tracking service module, and store the target tracking information in the tracking information queue corresponding to the local camera identifier contained therein.

[0111] The transmission thread is used to obtain target tracking information from the tracking information queue, and store the target tracking information in the first buffer of the target rendering thread corresponding to the tracking information queue, and store the target tracking information in the second buffer of the cross-screen rendering thread corresponding to the cross-screen camera identifier contained therein.

[0112] The target rendering thread is used to obtain target tracking information from the first buffer, and add the main image target detection box of the target object to the local fusion video frame corresponding to the local camera identifier and timestamp based on the local fusion position information of the target object in the main image target information contained therein.

[0113] The cross-screen rendering thread is used to obtain target tracking information from the second buffer, and add a cross-screen target detection box of the target object in the local fusion video frame whose cross-screen camera identifier and timestamp match based on the cross-screen position information of the target object contained in the cross-screen target information.

[0114] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0115] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A target tracking data processing method, characterized in that: The method is applied to a billion-pixel computing imaging system, which includes an array camera, a fusion service module, an on-screen service module, and a tracking service module. The method includes: Shooting multiple local videos by using multiple local cameras of the array camera; The tracking service module performs target tracking on each frame of the local video image of each local video to obtain local position information of the target object in the local video image; Based on the coordinate mapping relationship corresponding to the partial video image, the local position information is converted into fusion position information and local fusion position information, and target tracking information corresponding to the partial video image is generated according to the fusion position information and the local fusion position information; wherein the target tracking information includes main image target information and cross-screen target information; generating the target tracking information corresponding to the partial video image according to the fusion position information and the local fusion position information includes: determining the local camera identifier corresponding to the partial video image, the timestamp of the partial video image, the fusion position information and the local fusion position information of the target object as the main image target information; when the target object spans other partial fusion video images, determining the local camera identifier corresponding to the other partial fusion video images spanned by it as the cross-screen camera identifier, and converting the local fusion position information of the target object into cross-screen position information in the coordinate system of the other partial fusion video images spanned by it, and determining the cross-screen camera identifier and the cross-screen position information as the cross-screen target information; generating the target tracking information corresponding to the partial video image according to the main image target information and the cross-screen target information; Sending the target tracking information to the fusion service module and / or the on-screen service module, so that the fusion service module adds a target detection frame of the target object in the fused video according to the fusion position information contained in the target tracking information, and the on-screen service module adds a target detection frame of the target object in the corresponding local fused video according to the local fusion position information contained in the target tracking information; The on-screen service module includes a receiving thread, a transmission thread, and multiple rendering threads, wherein one rendering thread processes one channel of the local video and corresponds to one display module of the spliced display terminal. The method further includes: the receiving thread receives the target tracking information sent by the tracking service module and stores the target tracking information in a tracking information queue corresponding to the local camera identifier contained therein; the transmission thread obtains the target tracking information from the tracking information queue and stores the target tracking information in a first buffer of a target rendering thread corresponding to the tracking information queue, and stores the target tracking information in a second buffer of a cross-screen rendering thread corresponding to the cross-screen camera identifier contained therein; the target rendering thread obtains the target tracking information from the first buffer and, based on the local fusion position information of the target object in the main image target information contained therein, adds a main image target detection frame of the target object to the local fusion video frame corresponding to the local camera identifier and the timestamp; the cross-screen rendering thread obtains the target tracking information from the second buffer and, based on the cross-screen position information of the target object in the cross-screen target information contained therein, adds a cross-screen target detection frame of the target object to the local fusion video frame whose cross-screen camera identifier and the timestamp match.

2. The method according to claim 1, characterized in that The method further comprises: The fusion service module receives the target tracking information sent by the tracking service module; According to the local camera identifier and timestamp in the main image target information contained in the target tracking information, the target local video image to which the target tracking information belongs is determined, and according to the fusion position information of the target object in the main image target information, the target detection frame of the target object is added in the fusion video frame corresponding to the target local video image.

3. The method according to claim 1, characterized in that The method further comprises: The upper screen service module receives the target tracking information sent by the tracking service module; Determine the target rendering thread corresponding to the target tracking information according to the local camera identifier in the main image target information contained in the target tracking information, and determine the cross-screen rendering thread corresponding to the target tracking information according to the cross-screen camera identifier in the cross-screen target information contained therein; Adding, by the target rendering thread, a main image target detection frame of the target object in the local fusion video frame corresponding to the local camera identifier and the timestamp according to the local fusion position information of the target object in the main image target information; The cross-screen rendering thread adds a cross-screen target detection frame of the target object in the local fusion video frame in which the cross-screen camera identifier and the timestamp match according to the cross-screen position information of the target object in the cross-screen target information.

4. A billion-pixel computational imaging system, characterized in that: The billion-pixel computational imaging system includes an array camera, a fusion service module, an on-screen service module, and a tracking service module; The array camera is used to shoot multiple local videos through multiple local cameras, and send the local videos to the fusion service module, the on-screen service module and the tracking service module; The tracking service module is used to perform target tracking on each frame of the local video image of each local video, obtain the local position information of the detected target object in the local video image, convert the local position information into fusion position information and local fusion position information based on the coordinate mapping relationship corresponding to the local video image, generate the target tracking information corresponding to the local video image according to the fusion position information and the local fusion position information, and send the target tracking information to the fusion service module and the upper screen service module; wherein, the target tracking information includes main image target information and cross-screen target information; the tracking service module is also used to convert the local video image into The corresponding local camera identifier, the timestamp of the local video image, the fusion position information of the target object and the local fusion position information are determined as the main image target information; in the case where the target object spans other local fusion video images, the local camera identifier corresponding to the other local fusion video images it spans is determined as the cross-screen camera identifier, and the local fusion position information of the target object is converted into cross-screen position information in the coordinate system of the other local fusion video images it spans, and the cross-screen camera identifier and the cross-screen position information are determined as cross-screen target information; according to the main image target information and the cross-screen target information, target tracking information corresponding to the local video image is generated; The fusion service module is used to splice and fuse the multiple local videos to obtain a fused video, and add a target detection frame of the target object in the fused video according to the fusion position information included in the target tracking information; The upper screen service module is used to perform video frame conversion and rendering on each of the local videos through multiple graphics cards to obtain a local fusion video corresponding to each of the local videos, and add a target detection frame of the target object in the corresponding local fusion video according to the local fusion position information contained in the target tracking information; wherein, the upper screen service module includes a receiving thread, a transmission thread and multiple rendering threads, one of the rendering threads processes one of the local videos and corresponds to a display module of the splicing display terminal; the receiving thread is used to receive the target tracking information sent by the tracking service module, and store the target tracking information in the tracking information queue corresponding to the local camera identifier contained therein; the transmission thread is used to obtain the target tracking information from the tracking information queue, and store the target detection frame in the tracking information queue corresponding to the local camera identifier contained therein; The target tracking information is stored in the first buffer of the target rendering thread corresponding to the tracking information queue, and the target tracking information is stored in the second buffer of the cross-screen rendering thread corresponding to the cross-screen camera identifier contained therein; the target rendering thread is used to obtain the target tracking information from the first buffer, and add the main image target detection frame of the target object in the local fusion video frame corresponding to the local camera identifier and the timestamp according to the local fusion position information of the target object in the main image target information contained therein; the cross-screen rendering thread is used to obtain the target tracking information from the second buffer, and add the cross-screen target detection frame of the target object in the local fusion video frame matching the cross-screen camera identifier and the timestamp according to the cross-screen position information of the target object in the cross-screen target information contained therein.

5. The system according to claim 4, characterized in that The fusion service module is also used to determine the target local video image to which the target tracking information belongs based on the local camera identifier and timestamp in the main image target information contained in the target tracking information, and add the target detection frame of the target object in the fused video frame corresponding to the target local video image based on the fusion position information of the target object in the main image target information.

6. The system according to claim 4, characterized in that The upper-screen service module is further configured to determine a target rendering thread corresponding to the target tracking information based on a local camera identifier in the main image target information included in the target tracking information, and determine a cross-screen rendering thread corresponding to the target tracking information based on a cross-screen camera identifier in the cross-screen target information included therein; Adding, by the target rendering thread, a main image target detection frame of the target object in the local fusion video frame corresponding to the local camera identifier and the timestamp according to the local fusion position information of the target object in the main image target information; The cross-screen rendering thread adds a cross-screen target detection frame of the target object in the local fusion video frame in which the cross-screen camera identifier and the timestamp match according to the cross-screen position information of the target object in the cross-screen target information.

Citation Information

Patent Citations

  • Message sending method and device, computer equipment, storage medium and computer program product

    CN116661657A

  • Message card generation method and device, computer equipment and storage medium

    CN117010358A