Compressed outline image-based bandwidth optimization object detection method and system

By generating compressed outline diagrams at the edge and encoding and transmission, combined with cloud condition detection model, the video data streaming delay problem is solved, and efficient object detection is achieved.

WO2025145499A1PCT designated stage expired Publication Date: 2025-07-10PENG CHENG LAB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/081887
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-03
Filing Date
2024-03-15
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

In the prior art, directly streaming video data to the cloud for inference will result in high latency, especially in the case of poor network, edge devices are limited in computing power and cannot deploy deep learning models with good results.

Method used

By generating compressed outline diagrams at the edge, inter-frame and intra-frame compression are performed, encoded data is generated, and inference is used in the cloud using the conditional object detection model, and the edge and cloud work together.

Benefits of technology

It reduces the bandwidth consumption of data transmission from the edge to the cloud, while ensuring the inference accuracy of the target video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024081887_10072025_PF_FP_ABST
    Figure CN2024081887_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of object detection. Disclosed is a compressed outline image-based bandwidth optimization object detection method and system. The method comprises: acquiring the current image frame of a target video; extracting a target straight line segment from the current image frame to generate a current outline image; performing inter-frame compression and intra-frame compression on the current outline image to obtain compressed outline data; and determining coded data of the compressed outline data, such that upon receiving the coded data, a cloud inputs into a conditional object detection model the current outline image corresponding to the coded data, an outline image of a previous frame and a detection result mask of the previous frame, so as to obtain a current frame detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Bandwidth-optimized target detection method and system based on compressed contour graph

[0001] Related applications

[0002] This application claims priority to Chinese patent application No. 202410024948.4 filed on January 3, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present application relates to the technical field of target detection, and in particular to a bandwidth-optimized target detection method and system based on a compressed contour map. Background Art

[0004] Real-time video analytics has revolutionized various fields, with applications ranging from smart cameras in homes to video surveillance, autonomous driving, and smart manufacturing. Most of these applications are enabled by deep learning models. However, analyzing large amounts of video data presents challenges. Limited computing power on edge devices prevents the deployment of effective state-of-the-art models. Consequently, video streams must be uploaded to the cloud for inference. However, directly transmitting video to the cloud for inference results in high latency, especially over poor network conditions.

[0005] Summary of the Invention

[0006] The main purpose of this application is to provide a bandwidth-optimized target detection method and system based on compressed contour maps, aiming to solve the technical problem in the existing technology that directly transmitting video data streams to the cloud for inference will result in very high latency.

[0007] To achieve the above objectives, the present application provides a bandwidth optimization target detection method based on a compressed contour map. The bandwidth optimization target detection system based on a compressed contour map includes an edge end and a cloud end. The bandwidth optimization target detection method based on a compressed contour map is applied to the edge end. The method includes the following steps:

[0008] Get the current frame image of the target video;

[0009] Extracting a target straight line segment from the current frame image to generate a current contour map;

[0010] Performing inter-frame compression and intra-frame compression on the current contour image to obtain compressed contour data;

[0011] Determine the encoding data of the compressed contour data so that after the cloud receives the encoding data, the current contour map, the previous frame contour map and the previous frame detection result mask corresponding to the encoding data are input into the conditional target detection model to obtain the current frame detection result.

[0012] In one embodiment, extracting a target straight line segment from the current frame image to generate a current contour map includes:

[0013] Set up coarse-grained line segment detection;

[0014] determining a first target straight line segment of the current frame image according to the coarse-grained line segment detection, and determining an initial contour map according to the first target straight line segment;

[0015] determining a target position in the current frame image according to a target recorder and a presence detector;

[0016] Performing fine-grained line segment detection on the target position to obtain a second target straight line segment of the current frame image;

[0017] A current contour map is generated based on the initial contour map and the second target straight line segment.

[0018] In one embodiment, determining the target position in the current frame image based on the target recorder and the presence detector includes:

[0019] Recording the frequency of target appearance in each area of ​​the target video in the past time period by the target recorder, and determining a target recording matrix according to the frequency;

[0020] Determining the probability of a new target appearing in each area of ​​the current frame image by the presence detector, and determining a presence detection matrix based on the probability;

[0021] Determining the total probability of each region in the current frame image according to the target record matrix and the presence detection matrix;

[0022] The target position in the current frame image is determined according to the total probability.

[0023] In one embodiment, generating a current contour map based on the initial contour map and the second target straight line segment includes:

[0024] Adding the second target straight line segment to the initial contour map to obtain a summary contour map;

[0025] determining similar line segments of the summary contour map;

[0026] Removing other line segments in the summary contour map to obtain a fine-tuned contour map, wherein the other line segments are line segments that are not retained line segments among the similar line segments;

[0027] The free line segments in the fine-tuned contour map are removed to generate a current contour map.

[0028] In one embodiment, performing inter-frame compression and intra-frame compression on the current contour image to obtain compressed contour data includes:

[0029] Determine the same line segments between the current contour image and the previous frame contour image, and use the line segment index of the previous frame image to represent the same line segments in the current contour image, to obtain inter-frame compressed contour data;

[0030] Determining a matching failure line segment of the current contour image, and converting the coordinates of the matching failure line segment into a difference between the coordinates of the matching failure line segment and an adjacent line segment to obtain intra-frame compressed contour data;

[0031] Compression profile data is determined based on the inter-frame compression profile data and the intra-frame compression profile data.

[0032] In addition, the present application also provides a bandwidth optimization target detection method based on a compressed contour map. The bandwidth optimization target detection system based on a compressed contour map includes an edge end and a cloud end. The bandwidth optimization target detection method based on a compressed contour map is applied to the cloud end. The method includes the following steps:

[0033] After receiving the encoded data sent by the edge end, the current contour map, the previous frame contour map and the previous frame detection result mask corresponding to the encoded data are input into the conditional target detection model to obtain the current frame detection result, wherein the encoded data is the edge end extracting the target straight line segment from the current frame image to generate the current contour map, and the current contour map is compressed between frames and within frames to obtain compressed contour data, which is determined by the compressed contour data.

[0034] In one embodiment, after receiving the encoded data sent by the edge end, the current contour map, the previous frame contour map, and the previous frame detection result mask corresponding to the encoded data are input into the conditional target detection model, and before obtaining the current frame detection result, the method further includes:

[0035] Train an unconditional basic detection model based on a large-scale target detection dataset to obtain an unconditional target detection model;

[0036] The unconditional detection model is fine-tuned using a dedicated dataset to obtain a conditional detection model.

[0037] In one embodiment, after receiving the encoded data sent by the edge terminal, the current contour image corresponding to the encoded data is input into the conditional target detection model to obtain the current frame detection result, including:

[0038] Determine the picture sending mode of the current frame image;

[0039] If the picture sending mode is the first mode, after receiving the current frame image sent by the edge end, inputting the current frame image into the target detection model to obtain the current frame detection result;

[0040] If the image sending mode is the second mode, after receiving the current frame image and encoded data sent by the edge end, determine the current frame detection result and the inference result according to the target detection model and the conditional target detection model, and fine-tune the conditional target detection model according to the current frame detection result and the inference result;

[0041] If the image sending mode is the third mode, after receiving the encoded data sent by the edge end, the current contour map, the previous frame contour map and the previous frame detection result mask corresponding to the encoded data are input into the conditional target detection model to obtain the current frame detection result.

[0042] In one embodiment, if the picture sending mode is the second mode, after receiving the current frame image and encoded data sent by the edge end, determining the current frame detection result and the inference result according to the target detection model and the conditional target detection model, and fine-tuning the conditional target detection model according to the current frame detection result and the inference result, including:

[0043] If the picture sending mode is the second mode, receiving the current frame image and encoded data sent by the edge end;

[0044] Inputting the current frame image into the target detection model to obtain the current frame detection result;

[0045] Inputting the current contour map, the previous frame contour map and the previous frame detection result mask corresponding to the encoded data into the conditional target detection model to obtain an inference result;

[0046] Determining an error between the current frame detection result and the inference result;

[0047] When the error is greater than a preset error, the conditional target detection model is fine-tuned according to the current frame detection result, the previous frame detection result and the current contour map.

[0048] In addition, to achieve the above-mentioned purpose, the present application also proposes a bandwidth-optimized target detection system based on a compressed contour map, wherein the bandwidth-optimized target detection system based on a compressed contour map includes an edge end and a cloud end, wherein the edge end executes the method described above, and the cloud end executes the method described above.

[0049] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, on which a bandwidth optimization target detection program based on a compressed contour map is stored. When the bandwidth optimization target detection program based on a compressed contour map is executed by a processor, the steps of the bandwidth optimization target detection method based on a compressed contour map as described above are implemented.

[0050] The bandwidth-optimized target detection method and system based on compressed contour maps proposed in this application obtains the current frame image of the target video; extracts the target straight line segment from the current frame image to generate the current contour map; performs inter-frame compression and intra-frame compression on the current contour map to obtain compressed contour data; determines the encoding data of the compressed contour data so that after the cloud receives the encoding data, the current contour map, the previous frame contour map, and the previous frame detection result mask corresponding to the encoding data are input into the conditional target detection model to obtain the current frame detection result. Through the above method, the straight line contour can be extracted from the RGB image to convey the semantic information in the original image to obtain the contour map, and the contour map can be further encoded and compressed and transmitted to the cloud. The conditional detection model on the cloud can be used to realize the reasoning of the RGB image based on the contour map, which can not only greatly reduce the bandwidth consumption of transmitting data from the edge to the cloud, but also ensure the reasoning accuracy of the target video. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] FIG1 is a schematic structural diagram of a bandwidth-optimized target detection device based on a compressed profile graph in a hardware operating environment according to an embodiment of the present application;

[0052] FIG2 is a flow chart of a first embodiment of a bandwidth-optimized target detection method based on a compressed profile graph of the present application;

[0053] FIG3 is a schematic structural diagram of a bandwidth-optimized target detection device based on a compressed profile graph in a hardware operating environment according to an embodiment of the present application;

[0054] FIG4 is a flow chart of a second embodiment of the bandwidth optimization target detection method based on a compressed profile graph of the present application;

[0055] FIG5 is a schematic structural diagram of a bandwidth-optimized target detection system based on a compressed contour map of the present application.

[0056] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0057] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0058] Refer to Figure 1, which is a structural diagram of a bandwidth-optimized target detection device based on a compressed contour map in a hardware operating environment according to an embodiment of the present application.

[0059] As shown in FIG1 , the bandwidth-optimized target detection device based on a compressed profile may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to implement communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard. The user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may include a standard wired interface or a wireless interface (such as a wireless fidelity (WI-FI) interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk storage device. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0060] Those skilled in the art will understand that the structure shown in FIG1 does not constitute a limitation on the bandwidth-optimized target detection device based on the compressed contour map, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0061] As shown in FIG1 , the memory 1005 as a storage medium may include an operating system, a network communication module, a user interface module, and a bandwidth-optimized target detection program based on a compression profile.

[0062] In the bandwidth optimization target detection device based on the compressed contour map shown in Figure 1, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the bandwidth optimization target detection system based on the compressed contour map of the present application can be set in the bandwidth optimization target detection device based on the compressed contour map, and the bandwidth optimization target detection device based on the compressed contour map calls the bandwidth optimization target detection program based on the compressed contour map stored in the memory 1005 through the processor 1001, and executes the bandwidth optimization target detection method based on the compressed contour map provided in the embodiment of the present application.

[0063] Based on the above hardware structure, an embodiment of the bandwidth optimized target detection method based on compressed contour graph of the present application is proposed.

[0064] Refer to FIG. 2 , which is a flow chart of a first embodiment of a bandwidth-optimized target detection method based on a compressed contour graph of the present application.

[0065] The bandwidth optimization target detection system based on the compressed contour map includes an edge end and a cloud end. In this embodiment, the bandwidth optimization target detection method based on the compressed contour map is applied to the edge end. The bandwidth optimization target detection method based on the compressed contour map includes:

[0066] Step S10: Acquire the current frame image of the target video.

[0067] It should be noted that the execution entity of this embodiment can be a computing service device with data processing, network communication, and program execution capabilities, such as a mobile phone, tablet computer, personal computer, etc., or an electronic device capable of implementing the above functions or a bandwidth-optimized target detection device based on a compressed profile. This embodiment and the following embodiments are described below using the edge as an example.

[0068] It should be noted that the target video refers to the video data that needs to be uploaded to the cloud for inference, and the current frame image is one of the frames in the target video, which is an RGB image.

[0069] Step S20: extracting target straight line segments from the current frame image to generate a current contour map.

[0070] It should be noted that the target straight line segment is a straight line segment containing key semantics in the RGB image; the current contour map is composed only of straight line segments; when generating the current contour map, straight line segments containing key semantic information are extracted from the RGB image to form the contour. The number of these target straight line segments should be as small as possible, and these target straight line segments also need to be able to convey the semantic information of the original image well to ensure accurate reasoning; for the detection of straight line segments, the LSD algorithm in OpenCV can be used.

[0071] In the specific implementation, a two-stage line segment detection algorithm can be designed based on the LSD algorithm to optimize the generation of the current contour map. Specifically, in the first stage, we use a coarse-grained line segment detection algorithm to detect the contour and obtain a rough contour generation result; in the second stage, the output results of the target recorder and the presence detector are used to determine the positions where the target may appear in the picture, and fine-grained line segment detection is performed on these positions. Finally, the current contour map is generated based on the line segment detection results of these two stages.

[0072] In one embodiment, extracting a target straight line segment from the current frame image to generate a current contour map includes:

[0073] Set up coarse-grained line segment detection;

[0074] determining a first target straight line segment of the current frame image according to the coarse-grained line segment detection, and determining an initial contour map according to the first target straight line segment;

[0075] determining a target position in the current frame image according to a target recorder and a presence detector;

[0076] Performing fine-grained line segment detection on the target position to obtain a second target straight line segment of the current frame image;

[0077] A current contour map is generated based on the initial contour map and the second target straight line segment.

[0078] It should be noted that in the first stage of generating the current contour map, a rough contour map (i.e., the initial contour map) is generated by setting coarse-grained line segment detection to convey rough semantic information; in the second stage of generating the current contour map, we perform finer-grained line segment detection on areas where targets may appear, increase the detail information of specific areas, and obtain a more refined contour map (i.e., the current contour map). These areas that require fine-grained detection are jointly determined by the target recorder and presence detector modules.

[0079] In one embodiment, determining the target position in the current frame image according to the target recorder and the presence detector includes:

[0080] Recording the frequency of target appearance in each area of ​​the target video in the past time period by the target recorder, and determining a target recording matrix according to the frequency;

[0081] Determining the probability of a new target appearing in each area of ​​the current frame image by the presence detector, and determining a presence detection matrix based on the probability;

[0082] Determining the total probability of each region in the current frame image according to the target record matrix and the presence detection matrix;

[0083] The target position in the current frame image is determined according to the total probability.

[0084] It should be noted that the design idea of ​​the target recorder is to record the number of times a target appears in each area of ​​the video data of a given camera in the past period of time, which is used to indicate the probability that a target may appear in each area in the future. Specifically, given an input image I∈R H×W×3 , the target recorder maintains a matrix M ol (i.e., target record matrix), where each element represents the probability that a target appeared in the N×N region of the original image in the past P frames, and is updated each time the target detection result is returned.

[0085] It should be noted that the presence detector also maintains a matrix M pd (i.e., the presence detection matrix), where each element represents the probability that a new target may appear in the N×N region of the original image (i.e., the current frame image). Specifically, given the contour map I of the current frame c and the outline of the previous frame I p , we first calculate the difference between them Where i represents the i-th row, j represents the j-th column, and D ij Then a convolution operation is performed to obtain the final M pd (and M ol same size).

[0086] It should be noted that, given the target record matrix and the presence detection matrix, we first calculate the total probability M that a target may appear in each N×N area (i.e., a macroblock) prob =M ol +λ·M pd , where λ is the weight parameter, and then from M prob The rectangular areas with a probability greater than γ (i.e., the target position, the area with a total probability greater than the preset probability) are taken out, and these areas are detected in a finer-grained manner to obtain a contour map with increased details.

[0087] In this embodiment, the target position in the current contour map is jointly determined by the target recorder and the presence detector, which facilitates subsequent finer-grained line segment detection for the target position and ensures that the original semantic information is retained in the current contour map.

[0088] In one embodiment, generating a current contour map based on the initial contour map and the second target straight line segment includes:

[0089] Adding the second target straight line segment to the initial contour map to obtain a summary contour map;

[0090] determining similar line segments of the summary contour map;

[0091] Removing other line segments in the summary contour map to obtain a fine-tuned contour map, wherein the other line segments are line segments that are not retained line segments among the similar line segments;

[0092] The free line segments in the fine-tuned contour map are removed to generate a current contour map.

[0093] It should be noted that the summary contour map refers to the contour map obtained by performing coarse-grained line segment detection and fine-grained detection on the current frame image; in order to remove some irrelevant fields in the summary contour map as much as possible and retain the original semantic information, it is necessary to continue to fine-tune the summary contour map. Specifically, the fine-tuning includes two steps: 1) Removing similar line segments: If the starting point and end point positions of any line segments in the summary contour map are close, these line segments can be regarded as similar line segments. The number of line segments in the same similar line segment is greater than or equal to two, then a retained line segment can be determined from the similar line segments (the retained segment is the line segment that is not removed), and the other line segments of the similar line segment are removed. Only the retained line segment is retained to obtain the fine-tuned contour map. In order to improve computational efficiency, we use KDTree to search for similar line segments; 2) Removing free line segments: If there is no line segment in the vicinity of a line segment, the line segment is determined as a free line segment and removed.

[0094] In this embodiment, similar segments and loose segments in the summary contour graph are removed, which can remove as many irrelevant segments as possible while retaining the original voice information, thereby effectively reducing the amount of data compression and thus reducing the transmission delay of the target video.

[0095] Step S30: performing inter-frame compression and intra-frame compression on the current contour image to obtain compressed contour data.

[0096] In one embodiment, performing inter-frame compression and intra-frame compression on the current contour image to obtain compressed contour data includes:

[0097] Determine the same line segments between the current contour image and the previous frame contour image, and use the line segment index of the previous frame image to represent the same line segments in the current contour image, to obtain inter-frame compressed contour data;

[0098] Determining a matching failure line segment of the current contour image, and converting the coordinates of the matching failure line segment into a difference between the coordinates of the matching failure line segment and an adjacent line segment to obtain intra-frame compressed contour data;

[0099] Compression profile data is determined based on the inter-frame compression profile data and the intra-frame compression profile data.

[0100] It should be noted that the previous frame contour image refers to the contour image of the previous frame image of the current frame image. To optimize the transmission of the contour image, we compress and encode the current contour image. Since the current contour image is composed of line segments, each line segment is represented by the horizontal and vertical coordinates of the starting and ending points. Similar to traditional video encoding algorithms, contour image compression also includes two stages: inter-frame compression and intra-frame compression:

[0101] For inter-frame compression, we use the previous frame contour map as a reference to determine the same line segment between the current contour map and the previous frame contour map, and use the index of the line segment in the previous frame contour map to represent the same line segment. For the line segments in the current frame image that fail to match (i.e., other line segments that are not the same line segments in the current contour map), intra-frame compression is performed later. Due to the existence of detection errors and image jitter, the same line segments between the contour maps of adjacent frames may not be completely consistent (consistent means that the horizontal and vertical coordinates of the starting point and the end point are completely equal), so we allow a matching error e when matching. For two line segments, given the horizontal and vertical coordinates of the starting point and the end point (x1, y1), (x2, y2) and if These two line segments are considered to be the same line segment.

[0102] After the inter-frame compression is completed, the failed matching line segments in the current frame image (that is, the line segments in the current contour map that fail to match the line segments in the previous frame contour map) are further compressed within the frame. The method of intra-frame compression is mainly to convert the original line segment coordinates into the difference with the coordinates of the adjacent line segments, so that they can be encoded more efficiently later. Specifically, we first sort all the failed matching line segments according to one of the four coordinates (that is, the horizontal and vertical coordinates of the starting point and the end point) (in fact, we found that the effect of sorting by each coordinate is similar), and then subtract the coordinates of each line segment from the coordinates of the previous line segment to obtain the difference, and retain each difference and the coordinates of the first line segment for subsequent encoding operations (that is, retain the coordinates of the first line segment and the difference between the other failed matching line segments except the first line segment).

[0103] In this embodiment, the coordinates of the failed matching line segment are replaced by the difference with the previous failed matching line segment, so that the line segment coordinate difference can be more effectively compressed by the variable length coding algorithm, thereby further reducing the amount of data while ensuring complete recovery.

[0104] Step S40: Determine the encoding data of the compressed contour data so that after the cloud receives the encoding data, the current contour map, the previous frame contour map and the previous frame detection result mask corresponding to the encoding data are input into the conditional target detection model to obtain the current frame detection result.

[0105] It should be noted that the detection result mask is also an image with a pixel value of 0 or 255. 255 means that the target is detected at the corresponding position, and 0 means that the target is not detected at the corresponding position.

[0106] It should be noted that the compressed contour data obtained after compression will be used for encoding. The goal of the encoding stage is to further reduce the amount of data from the perspective of data representation. Specifically, given a contour map I∈R H×W×3 And the corresponding line segment data L∈RN×4 , where N is the total number of line segments in I. Each row in L stores the difference between the current line segment and the previous line segment (except the first row). The value range of each element in L is [-M, M], where M = max(H, W) - 1. First, we round the floating-point numbers in L to integers, because the sub-pixel error of line segment detection can be ignored. After rounding, the value of each element can be encoded using d binary bits, where Since directly using d bits to encode a number may not be supported by existing programming languages, we first convert the numbers in L into binary codes (represented by 0 and 1), and then encode L using the bool type. After the Boolean encoding is completed, use the existing lossless data compression algorithm to convert L into a byte stream for transmission (i.e., encoded data).

[0107] It should be noted that the cloud receives the encoded byte stream (i.e., encoded data), and then performs the reverse operation to decode and decompress the encoded data to obtain the original line segment data (i.e., the current contour map). The current contour map and the detection results of the previous frame are input into the conditional target detection model to obtain the detection results of the current frame.

[0108] This embodiment obtains the current frame image of the target video; extracts the target straight line segment from the current frame image to generate the current contour map; performs inter-frame compression and intra-frame compression on the current contour map to obtain compressed contour data; and determines the encoding data of the compressed contour data so that after the cloud receives the encoding data, the current contour map, the previous frame contour map, and the previous frame detection result mask corresponding to the encoding data are input into the conditional target detection model to obtain the current frame detection result. Through the above method, the straight line contour can be extracted from the RGB image to convey the semantic information in the original image to obtain the contour map, and the contour map is further encoded and compressed before being transmitted to the cloud. The conditional detection model on the cloud is then used to realize the reasoning of the RGB image based on the contour map. This can not only greatly reduce the bandwidth consumption of transmitting data from the edge to the cloud, but also ensure the reasoning accuracy of the target video.

[0109] Refer to Figure 3, which is a structural diagram of a bandwidth-optimized target detection device based on a compressed profile graph in a hardware operating environment according to an embodiment of the present application.

[0110] As shown in FIG3 , the bandwidth-optimized target detection device based on a compressed profile may include: a processor 1001′, such as a central processing unit (CPU), a communication bus 1002′, a user interface 1003′, a network interface 1004′, and a memory 1005′. The communication bus 1002′ is used to implement communication between these components. The user interface 1003′ may include a display screen and an input unit such as a keyboard. The user interface 1003′ may also include a standard wired interface or a wireless interface. The network interface 1004′ may include a standard wired interface or a wireless interface (such as a wireless fidelity (WI-FI) interface). The memory 1005′ may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk storage device. The memory 1005′ may also be a storage device independent of the aforementioned processor 1001.

[0111] Those skilled in the art will understand that the structure shown in FIG3 does not constitute a limitation on the bandwidth-optimized target detection device based on the compressed contour map, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0112] As shown in FIG3 , the memory 1005 ′ as a storage medium may include an operating system, a network communication module, a user interface module, and a bandwidth-optimized target detection program based on a compression profile.

[0113] In the bandwidth optimization target detection system based on the compressed profile shown in Figure 3, the network interface 1004' is mainly used for data communication with the network server; the user interface 1003' is mainly used for data interaction with the user; the processor 1001' and the memory 1005 in the bandwidth optimization target detection device based on the compressed profile of the present application can be set in the bandwidth optimization target detection system based on the compressed profile, and the bandwidth optimization target detection device based on the compressed profile calls the bandwidth optimization target detection program based on the compressed profile stored in the memory 1005' through the processor 1001', and executes the bandwidth optimization target detection method based on the compressed profile provided in the embodiment of the present application.

[0114] Based on the above hardware structure, an embodiment of the bandwidth optimized target detection method based on compressed contour graph of the present application is proposed.

[0115] The bandwidth optimization target detection method based on the compressed contour map of this embodiment is applied to the cloud. The bandwidth optimization target detection method based on the compressed contour map includes:

[0116] Step S110: After receiving the encoded data sent by the edge end, the current contour map, the previous frame contour map and the previous frame detection result mask corresponding to the encoded data are input into the conditional target detection model to obtain the current frame detection result, wherein the encoded data is the edge end extracting the target straight line segment from the current frame image to generate the current contour map, and the current contour map is compressed between frames and within frames to obtain compressed contour data, which is determined by the compressed contour data.

[0117] It should be noted that the detection result mask is also an image with a pixel value of 0 or 255. 255 means that the target is detected at the corresponding position, and 0 means that the target is not detected at the corresponding position.

[0118] In one embodiment, after receiving the encoded data sent by the edge terminal, the current contour map, the previous frame contour map, and the previous frame detection result mask corresponding to the encoded data are input into the conditional target detection model, and before obtaining the current frame detection result, the method further includes:

[0119] Train an unconditional basic detection model based on a large-scale target detection dataset to obtain an unconditional target detection model;

[0120] The unconditional detection model is fine-tuned using a dedicated dataset to obtain a conditional detection model.

[0121] It should be noted that training the conditional landmark detection model consists of two steps: 1) training an unconditional base object detection model to obtain the unconditional object detection model SketchDet_uc; 2) fine-tuning SketchDet_uc to obtain the conditional object detection model SketchDet. The only difference between SketchDet_uc and conventional object detection models is the input data: conventional models use RGB images as input, while SketchDet_uc uses contour maps as input. SketchDet_uc maintains the same model structure as conventional models, making it compatible with existing object detection model architectures. SketchDet_uc is trained on the COCO large-scale object detection dataset. First, the LSD algorithm is used to convert COCO RGB images into contour maps. Since SketchDet_uc's input consists of three channels, while contour maps only have one, the contour map is fed into the first channel for training. The remaining two channels are filled with all-zero values ​​for subsequent fine-tuning to obtain the conditional object detection model SketchDet. SketchDet adds conditional information on the basis of SketchDet_uc. The conditional information here is the contour map of the previous frame and the detection result mask of the previous frame. The detection result mask is also a picture with a pixel value of 0 or 255. 255 means that the target is detected at the corresponding position, and 0 means that the target is not detected at the corresponding position. Given a video dataset, SketchDet_uc is fine-tuned, where the first input channel is the current contour map (i.e., the contour map of the current frame image), the second channel is the contour map of the previous frame (i.e., the contour map of the previous frame image), and the third channel is the detection result mask of the previous frame. The output of the model is the detection result of the current frame image. By fine-tuning on a given video dataset, the model can learn the pattern of target movement between adjacent frames, which can improve the effect of the target detection model.

[0122] It should be noted that the accuracy of the conditional target detection model that incorporates conditional information is affected by two factors: the accuracy of the mask of the previous frame's detection result, and the motion characteristics of the target in the video. Over time, the accuracy of the conditional target detection model SketchDet will gradually decrease. In order to maintain the accuracy of the model, a continuous learning method is proposed. Furthermore, after receiving the encoded data sent by the edge end, the current contour map corresponding to the encoded data is input into the conditional target detection model to obtain the current frame detection result, including:

[0123] Determine the picture sending mode of the current frame image;

[0124] If the picture sending mode is the first mode, after receiving the current frame image sent by the edge end, inputting the current frame image into the target detection model to obtain the current frame detection result;

[0125] If the image sending mode is the second mode, after receiving the current frame image and encoded data sent by the edge end, determine the current frame detection result and the inference result according to the target detection model and the conditional target detection model, and fine-tune the conditional target detection model according to the current frame detection result and the inference result;

[0126] If the image sending mode is the third mode, after receiving the encoded data sent by the edge end, the current contour map, the previous frame contour map and the previous frame detection result mask corresponding to the encoded data are input into the conditional target detection model to obtain the current frame detection result.

[0127] It should be noted that when the camera captures an image, a mode controller at the edge determines the mode in which the image is sent to the cloud for inference. There are three modes to choose from:

[0128] The first mode: The edge sends the image in RGB format to the cloud. The cloud can directly use the RGB object detection model for inference. This mode is used in the K1Nth frame, where K1∈N and N is the anchor period.

[0129] The second mode: The edge sends the image to the cloud in RGB format. The edge also sends the image in outline format to the cloud. The cloud uses the RGB target detection model for inference and the conditional target detection model SketchDet for outline inference. This mode is used for accuracy checking and model fine-tuning and is used in the K2MN+1th frame, where K2∈N and M∈N + ;

[0130] The third mode: For the remaining frames within the anchor period, they are sent to the cloud in the form of silhouette images, and the cloud uses the conditional target detection model SketchDet for silhouette images for reasoning.

[0131] As you can understand, we call frames with known accurate detection results and used as conditional information for inference anchor frames (i.e., frames that are directly used for inference using the object detection model, and the detection results determined by the object detection model are accurate detection results). The mask of this detection result is called the anchor mask, and the distance between each video frame and the previous anchor frame is called the anchor frame distance. Every fixed period, the edge sends an anchor frame to the cloud for inference to obtain the anchor mask. This period is called the anchor period.

[0132] It should be noted that within an anchoring cycle, the edge end will reason about the first frame of the image (anchor frame) within the anchoring cycle in the first mode (that is, the edge end sends the first frame of the image to the cloud in RGB format, and then the cloud directly uses the RGB target detection model for reasoning). Then the edge end will reason about the second frame of the image within the anchoring cycle in the second mode (that is, the edge end will send the second frame of the image to the cloud in RGB format, and the edge end will also send the contour map format of the second frame of the image to the cloud. The cloud will use the RGB target detection model for reasoning, and will also use the conditional target detection model SketchDet for reasoning for the contour map). At this time, the inference result of the target detection model will be obtained, and the inference result of the conditional target detection model will also be obtained. When the two inference results are the same, When the difference is not big, there is no need to update the conditional detection model. When the two inference results differ greatly, it means that the target motion mode has changed and the conditional target detection model needs to be updated (because the detection results and contour maps of adjacent frames are required to fine-tune the conditional target detection model, if the conditional target detection model needs to be updated, the detection results and contour maps of the first and second modes need to be combined to achieve the update of the conditional target detection model); finally, the edge end will infer the third frame image within the anchor period in the third mode (that is, the edge end will send the third frame image in the contour map format to the cloud, and the cloud uses the conditional target detection model SketchDet for contour map for inference. The conditional target detection model is the conditional target detection model fine-tuned with the output results of the second mode).

[0133] It's understandable that within an anchor period, the anchor frame is sent only once (and the first mode is used for inference only once). The first mode is used to obtain an accurate anchor mask, which serves as input to the conditional object detection model for the next frame. Because RGB images have much larger data volumes than contour images, anchor frames are only sent every N frames (e.g., N = 10) and at a lower encoding quality. Therefore, after obtaining the anchor mask, subsequent frames are sent using the third mode. Since the third mode occupies the majority of the transmitted image, it significantly reduces bandwidth consumption.

[0134] It is understandable that the purpose of the second mode is to determine when to update the model and collect data for fine-tuning. The anchor frame distance of the frames using the second mode is 1. Generally speaking, the smaller the anchor frame distance, the higher the inference accuracy of the contour map. Therefore, by comparing the inference results of the contour map in the second mode with the inference results of the RGB map (i.e., performing an accuracy check), if the inference results of the contour map and the inference results of the RGB are far apart, it means that the target motion pattern has changed and the model needs to be updated. In addition, since the detection results and contour maps of adjacent frames are required for fine-tuning SketchDet, the data sent by the second mode can be combined with the data sent by the first mode to fine-tune the model.

[0135] In this embodiment, an accurate anchor mask is obtained by setting the first mode, and a second mode is set to determine whether the model is suitable for updating. When the model needs to be updated, the anchor mask obtained in the first mode is combined for updating. Finally, the updated conditional target detection model is used for model inference in the third mode, thereby realizing continuous learning and online updating during the target video analysis process. This not only maintains the accuracy of the model, but also greatly reduces bandwidth consumption.

[0136] In one embodiment, if the image sending mode is the second mode, after receiving the current frame image and encoded data sent by the edge end, determining the current frame detection result and the inference result according to the target detection model and the conditional target detection model, and fine-tuning the conditional target detection model according to the current frame detection result and the inference result, including:

[0137] If the picture sending mode is the second mode, receiving the current frame image and encoded data sent by the edge end;

[0138] Inputting the current frame image into the target detection model to obtain the current frame detection result;

[0139] Inputting the current contour map, the previous frame contour map and the previous frame detection result mask corresponding to the encoded data into the conditional target detection model to obtain an inference result;

[0140] Determining an error between the current frame detection result and the inference result;

[0141] When the error is greater than a preset error, the conditional target detection model is fine-tuned according to the current frame detection result, the previous frame detection result and the current contour map.

[0142] It should be noted that if the inference result of the contour map differs significantly from the inference result of the RGB image (i.e., the current frame image) (i.e., the current frame detection result), specifically, if the error between the current frame detection result and the inference result is greater than a preset error, it indicates that the target motion pattern has changed and the model needs to be updated. Furthermore, since fine-tuning the conditional target detection model requires the detection results and contour map of adjacent frames, the data sent in the second mode can be combined with the data sent in the first mode to fine-tune the model.

[0143] In this embodiment, by setting a preset error to determine whether the difference between the inference result of the contour map and the inference result of the RGB image is too large, and then determining whether the model needs to be updated, the accuracy check of the conditional target detection model can be quickly achieved.

[0144] This embodiment obtains the current frame detection result by inputting the current contour map, the previous frame contour map, and the previous frame detection result mask corresponding to the encoded data sent by the edge end after receiving the encoded data, wherein the encoded data is the target straight line segment extracted by the edge end from the current frame image to generate the current contour map, and the current contour map is compressed between frames and within frames to obtain compressed contour data, which is determined by the compressed contour data. In this way, after receiving the compressed map from the edge end, the conditional detection model can be used to realize the reasoning of the RGB image according to the contour map, which can not only greatly reduce the bandwidth consumption of transmitting data from the edge end to the cloud, but also ensure the reasoning accuracy of the target video.

[0145] Referring to Figure 5, Figure 5 is a schematic diagram of the structure of a bandwidth-optimized target detection system based on a compressed profile graph in a hardware operating environment according to an embodiment of the present application. The bandwidth-optimized target detection system based on a compressed profile graph includes an edge device and a cloud, wherein the edge device executes the method described above, and the cloud device executes the method described above, and the edge device and the cloud device can exchange information.

[0146] In addition, an embodiment of the present application also proposes a storage medium, on which a bandwidth optimization target detection program based on a compressed contour map is stored. When the bandwidth optimization target detection program based on a compressed contour map is executed by a processor, the steps of the bandwidth optimization target detection method based on a compressed contour map as described above are implemented.

[0147] It should be understood that the above is only an example and does not constitute any limitation to the technical solution of the present application. In specific applications, technicians in this field can make settings as needed, and the present application does not impose any restrictions on this.

[0148] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this application. In actual applications, technicians in this field can select part or all of it according to actual needs to achieve the purpose of this embodiment scheme, and no restrictions are imposed here.

[0149] In addition, for technical details not fully described in this embodiment, please refer to the bandwidth optimization target detection method based on compressed contour map provided in any embodiment of the present application, which will not be repeated here.

[0150] In addition, it should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0151] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0152] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus the necessary general hardware platform, or of course by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory (ROM) / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0153] The above are merely optional embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A bandwidth optimization object detection method based on a compressed contour map, wherein, The bandwidth-optimized object detection system based on compressed contour maps includes an edge side and a cloud side. The bandwidth-optimized object detection method based on compressed contour maps is applied to the edge side, and the method includes: Obtain the current frame image of the target video; Extract target straight line segments from the current frame image to generate a current contour map; Perform inter-frame compression and intra-frame compression on the current contour map to obtain compressed contour data; Determine the encoded data of the compressed contour data, so that after the cloud side receives the encoded data, it inputs the current contour map, the previous frame contour map, and the previous frame detection result mask corresponding to the encoded data into the conditional object detection model to obtain the current frame detection result.

2. The method according to claim 1, wherein The extracting target straight line segments from the current frame image to generate a current contour map includes: Set coarse-grained line segment detection; Determine the first target straight line segment of the current frame image according to the coarse-grained line segment detection, and determine an initial contour map according to the first target straight line segment; Determine the target positions in the current frame image according to the target recorder and the presence detector; Perform fine-grained line segment detection on the target positions to obtain the second target straight line segment of the current frame image; Generate a current contour map based on the initial contour map and the second target straight line segment.

3. The method according to claim 2, wherein The determining the target positions in the current frame image according to the target recorder and the presence detector includes: Record the frequency of the target appearing in each area of the target video in the past time period through the target recorder, and determine a target record matrix according to the frequency; Determine the probability of a new target appearing in each area of the current frame image through the presence detector, and determine a presence detection matrix according to the probability; Determine the total probability of each area in the current frame image according to the target record matrix and the presence detection matrix; Determine the target positions in the current frame image according to the total probability.

4. The method according to claim 2, wherein The generating a current contour map based on the initial contour map and the second target straight line segment includes: Add the second target straight line segment to the initial contour map to obtain a summary contour map; Determine the similar line segments of the summary contour map; Remove other line segments in the summary contour map to obtain a fine-tuned contour map, where the other line segments are the line segments that are not the retained line segments among the similar line segments; Remove the free line segments in the fine-tuned contour map to generate a current contour map.

5. The method according to claim 1, wherein The performing inter-frame compression and intra-frame compression on the current contour map to obtain compressed contour data includes: Determine the same line segments between the current contour map and the previous frame contour map, and use the line segment index of the previous frame image to represent the same line segments in the current contour map to obtain inter-frame compressed contour data; Determine the unmatched line segments of the current contour map, and convert the coordinates of the unmatched line segments into the difference between the coordinates of the unmatched line segments and the adjacent line segments to obtain intra-frame compressed contour data; Determine the compressed contour data according to the inter-frame compressed contour data and the intra-frame compressed contour data.

6. A bandwidth optimization object detection method based on a compressed contour map, wherein, The bandwidth-optimized object detection system based on compressed contour maps includes an edge side and a cloud side. The bandwidth-optimized object detection method based on compressed contour maps is applied to the cloud side. The bandwidth-optimized object detection method based on compressed contour maps includes: After receiving the encoded data sent by the edge side, input the current contour map, the previous frame contour map, and the previous frame detection result mask corresponding to the encoded data into the conditional object detection model to obtain the current frame detection result. Wherein, the encoded data is determined by the edge side extracting target straight line segments from the current frame image, generating the current contour map, performing inter-frame compression and intra-frame compression on the current contour map, and obtaining the compressed contour data.

7. The method according to claim 6, wherein Before inputting the current contour map, the previous frame contour map, and the previous frame detection result mask corresponding to the encoded data into the conditional object detection model after receiving the encoded data sent by the edge side to obtain the current frame detection result, it further includes: Training an unconditional basic detection model based on a large-scale object detection data set to obtain an unconditional object detection model; Fine-tuning the unconditional detection model using a dedicated data set to obtain a conditional detection model.

8. The method according to claim 6, wherein inputting the current contour map corresponding to the encoded data into the conditional object detection model after receiving the encoded data sent by the edge side to obtain the current frame detection result includes: Determine the picture sending mode of the current frame image; If the picture sending mode is the first mode, after receiving the current frame image sent by the edge side, input the current frame image into the object detection model to obtain the current frame detection result; If the picture sending mode is the second mode, after receiving the current frame image and the encoded data sent by the edge side, determine the current frame detection result and the inference result according to the object detection model and the conditional object detection model, and fine-tune the conditional object detection model according to the current frame detection result and the inference result; If the picture sending mode is the third mode, after receiving the encoded data sent by the edge side, input the current contour map, the previous frame contour map, and the previous frame detection result mask corresponding to the encoded data into the conditional object detection model to obtain the current frame detection result.

9. The method according to claim 8, wherein, The step of, if the picture sending mode is the second mode, after receiving the current frame image and the encoded data sent by the edge side, determining the current frame detection result and the inference result according to the object detection model and the conditional object detection model, and fine-tuning the conditional object detection model according to the current frame detection result and the inference result includes: If the picture sending mode is the second mode, receive the current frame image and the encoded data sent by the edge side; Input the current frame image into the object detection model to obtain the current frame detection result; Input the current contour map, the previous frame contour map, and the previous frame detection result mask corresponding to the encoded data into the conditional object detection model to obtain the inference result; Determine the error between the current frame detection result and the inference result; When the error is greater than a preset error, the conditional target detection model is fine-tuned according to the current frame detection result, the previous frame detection result, and the current contour map.

10. A bandwidth-optimized object detection system based on a compressed contour map, wherein, The bandwidth optimization target detection system based on the compressed contour map includes an edge side and a cloud side. The edge side executes the method according to any one of claims 1-5, and the cloud side executes the method according to any one of claims 6-9.

Citation Information

Patent Citations

  • Edge cooperative target detection method and device based on RoI coding

    CN112287803A

  • Edge cloud collaborative real-time video analysis method and system based on background removal

    CN114782872A

  • Video stream data identification method and processing system based on edge algorithm

    CN116229308A