A multi-channel video analysis method and device
By performing local picture frame splicing and fusion and dynamic scheduling analysis on the areas of interest in multi-channel video streams, the problem of insufficient intelligent analysis capabilities of back-end devices is solved, and efficient intelligent analysis of more channels is achieved.
Patent Information
- Application Number
- CN202111548297.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-17
AI Technical Summary
The intelligent analysis capabilities of back-end devices in the prior art are limited, resulting in a small number of supported channels and the inability to effectively utilize the potential of multi-channel video analysis.
By obtaining the dynamic detection results of video streams in multiple channels, local picture frames in the area of interest are intercepted and video stitching and fused to form the picture frame to be analyzed, and the dynamic scheduling analyzer processes it to optimize the allocation of intelligent analysis resources.
Without increasing the intelligent analysis capabilities of back-end devices, intelligent analysis of more channels is supported, improving video analysis efficiency and accuracy.
Smart Images

Figure CN114125400B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of video intelligent analysis, and in particular, to a multi-channel video analysis method and apparatus. Background Art
[0002] In the current video surveillance field, a backend device (such as a network video recorder) can access multiple frontend devices. The backend device accesses the video streams of different frontend devices by dividing different channels, and then performs different types of intelligent analysis on different channels respectively, generates intelligent analysis results and reports them. Currently, the allocation method for intelligent analysis of channels is based on the channel binding method, that is, if a certain type of intelligence is enabled for a certain channel, it will always occupy the intelligent decoding and analysis resources of that type. Moreover, due to the hardware limitations of the intelligent chip specifications of the backend device, the total intelligent analysis ability is limited, resulting in a small number of channels that the backend device can support for intelligent analysis.
[0003] Currently, there is an urgent need for a multi-channel video analysis method to support the intelligent analysis ability of more channels on the basis of the unchanged total intelligent analysis ability of the backend device. Summary of the Invention
[0004] Embodiments of this application provide a multi-channel video analysis method and apparatus to support the intelligent analysis ability of more channels on the basis of the unchanged total intelligent analysis ability of the backend device.
[0005] In a first aspect, embodiments of this application provide a multi-channel video analysis method. This method is applied to a backend device and includes: obtaining dynamic detection results of video streams of multiple channels, where the dynamic detection results of the video streams of each channel include coordinate information of regions of interest dynamically detected from the video stream of the channel and information of the picture frames where the regions of interest are located; for the video stream of each channel among the multiple channels, according to the coordinate information of the region of interest and the information of the picture frame where the region of interest is located, intercepting the picture containing the region of interest from the picture frame in the video stream as a partial picture frame; performing video splicing and fusion on the partial picture frames containing the regions of interest intercepted from the video streams of the multiple channels to obtain a picture frame to be analyzed; and inputting the picture frame to be analyzed into an analyzer with dynamic scheduling for analysis and processing.
[0006] In the above technical solution, on the one hand, the video splicing and fusion strategy is optimized. By performing video splicing and fusion on the partial picture frames where the regions of interest dynamically detected are located to obtain the picture frame to be analyzed, the intelligent analysis ability of the backend device can be reused with a smaller granularity, so as to support more video channels at the frontend. On the other hand, an analyzer used for the picture frame to be analyzed can be dynamically scheduled, so as to make full use of the intelligent analysis ability of the backend device.
[0007] In a possible design, for each of the multiple channels, obtaining the dynamic detection result of the video stream of the channel includes: receiving the dynamic detection result of the video stream of the channel from the front-end device connected to the channel; or, performing dynamic detection on the secondary code stream of the received video stream of the channel to obtain the dynamic detection result.
[0008] In a possible design, obtaining the dynamic detection result of the video stream of the channel further includes: if the front-end device connected to the channel does not support identifying the target type of the region of interest, performing dynamic detection on the secondary code stream of the video stream of the channel.
[0009] The above technical solution can execute different dynamic detection strategies based on different dynamic detection capabilities of the front-end device. For example, if the front-end device can detect the target type of the region of interest, dynamic detection can be performed on the primary code stream of the video stream in the front-end device. Detecting the region of interest in the front-end device can save the intelligent analysis capabilities of the back-end device. Optionally, the dynamic detection result may include the target type of the region of interest, which is convenient for the subsequent back-end device to perform further detection and processing on the video stream according to the target type; since the primary code stream has a higher resolution, performing dynamic detection on the primary code stream can make the dynamic detection result more accurate. If the front-end device cannot detect the target type of the region of interest, dynamic detection is performed on the secondary code stream of the video stream in the back-end device. The secondary code stream of the video stream has a lower resolution than the primary code stream. Performing dynamic detection on the secondary code stream of the video stream in the back-end device can avoid consuming too much intelligent analysis capabilities of the back-end device.
[0010] In a possible design, the coordinate information of the region of interest included in the dynamic detection result is the coordinate information in the coordinate system of the primary code stream of the video stream of the channel; the method further includes: after obtaining the region of interest by performing dynamic detection on the secondary code stream of the video stream of the channel, determining the coordinate information of the region of interest in the coordinate system of the primary code stream according to the coordinate information of the region of interest in the coordinate system of the secondary code stream and the conversion relationship between the coordinate system of the primary code stream and the coordinate system of the secondary code stream of the video stream of the channel.
[0011] The above technical solution converts the coordinate information of the region of interest in the coordinate system of the secondary code stream into the coordinate information in the coordinate system of the primary code stream, unifies the coordinate information, which is convenient for subsequently restoring the coordinate information of the detection result back to the coordinate system of the primary code stream to report the analysis result.
[0012] In a possible design, the dynamic detection result further includes the target type corresponding to each region of interest; the step of intercepting the local picture frames including the regions of interest from the video streams of the multiple channels and performing video stitching and fusion to obtain the picture frame to be analyzed includes: stitching and fusing the local picture frames of the regions of interest corresponding to the same target type in the video streams of the multiple channels into one or more video streams to be analyzed; inputting the video streams to be analyzed into the analyzer corresponding to the target type for analysis and processing.
[0013] The above technical solution can, according to the target type of the region of interest in the video to be analyzed, specifically activate the intelligent analysis capabilities corresponding to the target type in the analyzer. In this way, the range and computing power of internal model matching, recognition, and analysis in the intelligent analysis algorithm can be reduced, and the intelligent analysis efficiency of the backend device can be improved.
[0014] In a possible design, the backend device includes multiple target region cache lists, each target region cache list corresponding to a target type, and the target region cache list is used to store the local picture frames of the regions of interest including the corresponding target type.
[0015] The above technical solution sets up the target region cache list, which can cache the video frames detected by the front-end device into the target region cache list when the analyzer is full, facilitating the timely dynamic scheduling of the analyzer. Moreover, setting multiple target cache lists of different target types also improves the efficiency of stitching and fusing the local picture frames of the regions of interest with the same target type subsequently.
[0016] In a possible design, the dynamic detection result further includes the frequency information of the video stream of the channel triggering the dynamic detection; according to the frequency information of the video streams of the multiple channels triggering the dynamic detection, the frame rate of the video stream to be analyzed is determined.
[0017] The above technical solution determines the frame rate of the video stream to be analyzed according to the frequency information of the video stream triggering the dynamic detection. In this way, it helps to balance the effect of intelligent analysis and avoid the problem of omission in intelligent analysis.
[0018] In a possible design, according to the conversion relationship between the coordinate system of the main code stream of the video stream of the first channel and the coordinate system of the video stream to be analyzed, the coordinate information of the analysis result in the coordinate system of the main code stream is determined from the coordinate information of the analysis result obtained by analyzing and processing the picture frame to be analyzed.
[0019] In a second aspect, an embodiment of the present application provides a multi-channel video analysis device, including:
[0020] An acquisition module, configured to acquire the dynamic detection results of video streams of multiple channels. The dynamic detection results of the video stream of each channel include the coordinate information of one or more regions of interest dynamically detected from the video stream of the channel and the information of the picture frame where the region of interest is located.
[0021] A processing module, configured to, for the video stream of each channel among the multiple channels, intercept, according to the coordinate information of the region of interest and the information of the picture frame where the region of interest is located, the picture containing the region of interest from the picture frame in the video stream as a partial picture frame; splice and fuse the partial picture frames containing the region of interest dynamically detected from the video streams of the multiple channels according to the coordinate information of the region of interest to obtain a picture frame to be analyzed; and input the picture frame to be analyzed into an analyzer with dynamic scheduling for analysis and processing.
[0022] In a possible design, the processing module is specifically configured to: receive the dynamic detection results of the video stream of the channel from the front-end device connected to the channel; or perform dynamic detection on the secondary code stream of the received video stream of the channel to obtain the dynamic detection results. In a possible design, the device further includes a detection module, configured to perform dynamic detection on the secondary code stream of the video stream of the channel if the front-end device connected to the channel does not support identifying the target type of the region of interest.
[0023] In a possible design, the coordinate information of the region of interest included in the dynamic detection results is the coordinate information in the coordinate system of the primary code stream of the video stream of the channel; the processing module is specifically configured to: after obtaining the region of interest by performing dynamic detection on the secondary code stream of the video stream of the channel, determine the coordinate information of the region of interest in the coordinate system of the primary code stream according to the coordinate information of the region of interest in the coordinate system of the secondary code stream and the conversion relationship between the coordinate system of the primary code stream and the coordinate system of the secondary code stream of the video stream of the channel.
[0024] In a possible design, the dynamic detection results further include the target type corresponding to the region of interest; the processing module is specifically configured to: splice and fuse the partial picture frames of the regions of interest corresponding to the same target type in the video streams of the multiple channels into one or more video streams to be analyzed; and input the video stream to be analyzed into the analyzer corresponding to the target type for analysis and processing.
[0025] In a possible design, the backend device includes multiple target region cache lists, each target region cache list corresponding to a target type, and the target region cache list is used to store the partial picture frames containing the regions of interest corresponding to the target type.
[0026] In a possible design, the dynamic detection result further includes the frequency information of the video stream triggering dynamic detection of the channel; specifically, the processing module is configured to: determine the frame rate of the video stream to be analyzed according to the frequency information of the video streams of the multiple channels triggering dynamic detection.
[0027] In a possible design, specifically, the processing module is configured to: according to the conversion relationship between the coordinate system of the main code stream of the video stream of the first channel and the coordinate system of the video stream to be analyzed, determine the coordinate information of the analysis result in the coordinate system of the main code stream based on the coordinate information of the analysis result obtained by analyzing and processing the frame of the picture to be analyzed.
[0028] In a third aspect, an embodiment of the present application further provides a computing device, including:
[0029] A memory for storing program instructions;
[0030] A processor for calling the program instructions stored in the memory and executing the methods described in the various possible designs of the first aspect according to the obtained program instructions.
[0031] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which computer-readable instructions are stored. When a computer reads and executes the computer-readable instructions, the methods described in the first aspect or any possible design of the first aspect are implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 It is a schematic diagram of a video surveillance system applicable to an embodiment of the present application;
[0034] Figure 2 It is a schematic diagram of the intelligent analysis method of the backend device in the prior art;
[0035] Figure 3 It is a schematic diagram of the flow of a multi-channel video analysis method provided by an embodiment of the present application;
[0036] Figure 4 It is a specific example of a multi-channel video analysis method provided by an embodiment of the present application;
[0037] Figure 5Another specific example of a multi-channel video analysis method provided by an embodiment of the present application;
[0038] Figure 6 A schematic diagram of a specific working process of a multi-channel video analysis method provided by an embodiment of the present application;
[0039] Figure 7 A schematic diagram of another specific working process of a multi-channel video analysis method provided by an embodiment of the present application;
[0040] Figure 8 A schematic diagram of a multi-channel video analysis device provided by an embodiment of the present application. Detailed implementation manners
[0041] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0042] In the embodiments of the present application, "a plurality of" means two or more. Terms such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order.
[0043] In addition, the terms "include" and "have" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device including a series of components does not necessarily have to be limited to those components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0044] To better understand the embodiments of the present application, the nouns and functions involved in the prior art and the present application will be explained below.
[0045] As Figure 1 shown, in the current video surveillance field, backend devices such as network video recorders (NVRs) and digital video recorders (DVRs) can access multiple front-end devices such as IP cameras (IPCs) and PTZ cameras. As Figure 2As shown, the backend device can receive video streams from different front-end devices by dividing different channels, and then perform different types of intelligent analysis on the video streams in different channels respectively, and report the generated intelligent analysis results. Currently, the allocation method for intelligent analysis of channels is based on channel binding. That is, if a certain type of intelligence is enabled for a certain channel, it will always occupy the intelligent decoding and analysis resources of this channel. Moreover, due to the limitation of the intelligent chip specifications, the total intelligent analysis ability of each backend device has an upper limit, which results in a small number of channels for intelligent analysis that the backend device can support.
[0046] There are two types of dynamic detections that can be performed in the front-end device, namely ordinary motion detection and intelligent motion detection. Both are based on dividing the video into different blocks, and when the pixel point difference in a certain block reaches a certain threshold, a motion detection event is triggered. However, the difference is that intelligent motion detection is target-oriented, and on the basis of the original ordinary motion detection, some low-power computing algorithms are analyzed to identify relevant targets such as people or vehicles, which improves the application value in actual scenarios. Currently, most front-end devices support ordinary motion detection, while intelligent motion detection has certain requirements for the intelligent capabilities of the device, and only some front-end devices support it.
[0047] The video bitstreams of the front-end device are divided into a main bitstream and a secondary bitstream. The main bitstream has a larger resolution and a higher frame rate than the secondary bitstream, and is suitable for performing target-precise intelligent analysis. The secondary bitstream has a smaller resolution and occupies less bandwidth than the main bitstream, and is suitable for image transmission over a low-bandwidth network.
[0048] Generally, the consumption of the intelligent analysis ability of the backend device is directly related to parameters such as the resolution and frame rate of the front-end main bitstream. For example, assume that the backend device can support a maximum of 4 channels of face intelligent analysis at 1080P (1920×1080). If a front-end device with a main bitstream of 4096X2160 is connected to a channel, then actually only 1 channel of face intelligent analysis can be enabled on the backend device. Then, in the case where the main bitstream resolutions of multiple channels connected to the backend device are all large, if the main bitstream is directly analyzed on the backend device, there are certain limitations on the number of video channels that can be analyzed simultaneously.
[0049] Based on the above content, the present application proposes a multi-channel video analysis method to support the intelligent analysis ability of more channels on the basis of the unchanged total intelligent analysis ability of the backend device.
[0050] Figure 3 Exemplarily, a multi-channel video analysis method provided by an embodiment of the present application is shown. This method is applied to the backend device. As Figure 3 shown, this method includes the following steps:
[0051] Step 301: The backend device obtains the dynamic detection results of the video streams of multiple channels. The dynamic detection result of the video stream of each channel includes the coordinate information of the region of interest dynamically detected from the video stream of the channel and the information of the picture frame where the region of interest is located.
[0052] In this application, the backend device can receive the video streams of multiple different front-end devices in units of channels. The video streams may include the main stream and the sub-stream.
[0053] Taking the first channel among the multiple channels as an example, the backend device can obtain the dynamic detection result of the video stream of the first channel in the following ways: receiving the dynamic detection result of the video stream of the first channel from the front-end device connected to the first channel, or performing dynamic detection on the sub-stream of the received video stream of the first channel to obtain the dynamic detection result. The dynamic detection result includes the coordinate information of one or more regions of interest detected dynamically and the information of the picture frame where the region of interest is located. Optionally, it also includes the target type corresponding to each region of interest.
[0054] Specifically, the capabilities of dynamic detection supported by different front-end devices may vary. For example, some front-end devices can support intelligent dynamic detection, while some front-end devices can only support ordinary dynamic detection. The front-end device that supports intelligent dynamic detection can detect the target type of the region of interest. For example, it can identify whether the target in the region of interest is a target person, a target vehicle, or other targets; while the front-end device that does not support intelligent dynamic detection but only supports ordinary dynamic detection can only detect the region of interest and cannot detect the target type of the region of interest, that is, it cannot identify whether the target in the region of interest is a target person or a target vehicle, and can only know that the region of interest has triggered a dynamic detection event, and further analysis and processing are still required for this region of interest.
[0055] Furthermore, according to the differences in the intelligent chips used by the backend device, the backend device can be divided into those using high-end intelligent chips and those using low-end intelligent chips. The backend device using a high-end intelligent chip can support intelligent dynamic detection of the front-end video stream, while the backend device using a low-end intelligent chip does not support intelligent dynamic detection of the front-end video stream due to limited intelligent detection capabilities.
[0056] Exemplarily, if the front-end device connected to the first channel supports identifying the target type of the region of interest detected dynamically, the backend device can receive the dynamic detection result of the video stream of the first channel from the front-end device connected to the first channel. The dynamic detection result of the video stream of the first channel includes the coordinate information of one or more regions of interest and the target type corresponding to each region of interest.
[0057] If the front-end device connected to the first channel does not support identifying the target types of the regions of interest (ROIs) for dynamic detection, and the back-end device uses a high-end intelligent chip, then the back-end device can perform dynamic detection on the secondary code stream of the video stream of the first channel. Here, the secondary code stream is the video stream encoded by the front-end device connected to the first channel and transmitted to the back-end device. The reason for selecting the secondary code stream for dynamic detection is that generally, the resolution of the primary code stream is relatively large, and performing dynamic detection on the primary code stream by the back-end device will consume a large amount of intelligent analysis capabilities, while the resolution of the secondary code stream is relatively small, which is convenient for the back-end device to perform detection and processing.
[0058] If the front-end device connected to the first channel does not support identifying the target types of the regions of interest (ROIs) for dynamic detection, and the back-end device uses a low-end intelligent chip, then the back-end device receives the dynamic detection results of the video stream of the first channel from the front-end device connected to the first channel, and the dynamic detection results of the video stream of the first channel do not include the target types corresponding to each region of interest.
[0059] It should be noted that the coordinate information of the regions of interest in the dynamic detection results can refer to the coordinate information of the regions of interest in the coordinate system of the primary code stream of the video stream of the first channel. In this way, if the back-end device performs dynamic detection on the secondary code stream of the video stream of the first channel, then after obtaining the regions of interest, the back-end device can also determine the coordinate information of the regions of interest in the coordinate system of the primary code stream based on the coordinate information of the regions of interest in the coordinate system of the secondary code stream and the conversion relationship between the coordinate system of the primary code stream and the coordinate system of the secondary code stream of the video stream of the first channel, so as to extract the frame images in the primary code stream for subsequent analysis and processing.
[0060] The following combines Figure 4 The above steps are illustrated by way of example. The front-end device connected to Channel 1 does not support identifying the target types of the regions of interest (ROIs) for dynamic detection. The back-end device can perform dynamic detection on the secondary code stream of the video stream of Channel 1, and the obtained dynamic detection results include Region of Interest A and Region of Interest B. Among them, the coordinates of Region of Interest A and Region of Interest B in the coordinate system of the secondary code stream are A0(m1,n1,m2,n2) and B0(m3,n3,m4,n4) respectively. It should be noted that here, the two diagonal vertex coordinates of the region of interest are selected to represent the range of the region of interest, or other methods can be used to represent the range of the region of interest, and the present application does not make specific limitations on this.
[0061] According to the conversion relationship between the coordinate system of the primary code stream and the coordinate system of the secondary code stream of the video stream of Channel 1, A0(m1,n1,m2,n2) and B0(m3,n3,m4,n4) are converted into the coordinates A(x1,y1,x2,y2) and B(x3,y3,x4,y4) in the coordinate system of the primary code stream. The specific conversion process is as follows:
[0062] A(x1,y1,x2,y2) = (m1*W1 / w1,n1*H1 / h1,m2*W1 / w1,n2*H1 / h1)
[0063] B(x3,y3,x24,y4) = (m3*W1 / w1,n3*H1 / h1,m4*W1 / w1,n4*H1 / h1)
[0064] Among them, the resolution of the main video stream of the video stream of channel 1 is W1*H1, and the resolution of the sub-video stream is w1*h1.
[0065] The front-end device connected to channel 2 supports identifying the target types of the regions of interest for dynamic detection. Then, the back-end device can receive the dynamic detection results of the video stream of the first channel from the front-end device connected to channel 2, and the obtained dynamic detection results include the region of interest C and the region of interest D. Among them, the coordinates of the region of interest C and the region of interest D in the coordinate system of the main video stream are C(x5,y5,x6,y6) and D(x7,y7,x8,y8) respectively.
[0066] Step 302: For the video stream of each channel among multiple channels, the back-end device intercepts the picture containing the region of interest in the picture frame in the video stream as a local picture frame according to the coordinate information of the region of interest and the information of the picture frame where the region of interest is located.
[0067] Step 303: The back-end device performs video splicing and fusion on the local picture frames containing the regions of interest intercepted from the video streams of multiple channels to obtain a picture frame to be analyzed.
[0068] In this application, after the back-end device obtains the dynamic detection results, the back-end device can crop out the detected regions of interest from the picture frames of the video stream to obtain local picture frames of the regions of interest. If the dynamic detection results of the video stream of each channel also include the target types corresponding to each region of interest, the back-end device can splice and fuse the local picture frames of the regions of interest with the same target type in the video streams of multiple channels into multiple picture frames to be analyzed, and the multiple picture frames to be analyzed form one or more video streams to be analyzed.
[0069] It should be noted that the visual stitching and fusion of local picture frames containing regions of interest may include stitching the local picture frames of multiple regions of interest into a picture frame to be analyzed, and then forming a video stream to be analyzed by multiple picture frames to be analyzed. It should be noted that a picture frame to be analyzed is composed of the local pictures of several regions of interest, and the present application does not make specific limitations, and can be specifically adjusted according to the resolution of the local picture frames of the intercepted regions of interest or the analysis ability of the analyzer. When stitching and fusing the local picture frames of the regions of interest, the local picture frames can be appropriately compressed first to better adapt to the problem that the resolutions of the local picture frames of different channels are different. And the number of local picture frames stitched and fused for one or more video streams to be analyzed is not fixed, and can be adaptively adjusted according to the resolution of the local picture frames or the different analysis abilities of the analyzer.
[0070] Optionally, the backend device includes multiple target area cache lists of different target types, each target area cache list corresponds to a target type, and can be used to store the local picture frames of the regions of interest containing the corresponding target type. The local picture frames of the regions of interest dynamically detected in the video streams of multiple channels, and the intelligent frames carrying the dynamic detection results are stored in the target area cache lists of the corresponding target types. In this way, the backend device can stitch and fuse one or more local picture frames of the regions of interest of the target type stored in the target area cache list into one or more video streams to be analyzed. In one example, if only the target area cache list with the target type of person in the backend device stores the local picture frames of the regions of interest with the target type of person, then they can be stitched and fused into one video stream to be analyzed. In another example, if the target area cache list with the target type of person in the backend device stores the local picture frames of the regions of interest with the target type of person, and the target area cache list with the target type of vehicle stores the local picture frames of the regions of interest with the target type of vehicle, then they can be stitched and fused into one video stream to be analyzed respectively, and two video streams to be analyzed are obtained. In yet another example, if the size of the video stream to be analyzed stitched and fused from the local picture frames of the regions of interest with the target type of person stored in the target area cache list with the target type of person in the backend device is relatively large, then it can be stitched and fused into two or more video streams to be analyzed. It should be noted that for the situation where the target type of the region of interest cannot be recognized, the region of interest can be directly stored in the target cache list that does not divide the target type.
[0071] Optionally, after performing video stitching and fusion on local picture frames containing regions of interest intercepted from video streams of multiple channels to obtain the video stream to be analyzed, the backend device can also determine the coordinate information of the regions of interest in the coordinate system of the video stream to be analyzed after stitching and fusion based on the coordinate information of the regions of interest in the coordinate system of the main stream and the conversion relationship between the coordinate system of the main stream of the video stream of the first channel and the coordinate system of the video stream to be analyzed after stitching and fusion. The backend device packs and integrates the coordinate information of each set of regions of interest of each channel included in the video stream to be analyzed in the coordinate system of the video stream to be analyzed after stitching and fusion, and the coordinate information of the corresponding regions of interest in the coordinate system of the main stream of the original channel, and then inserts it into the intelligent frame of the video stream to be analyzed. This is to facilitate the subsequent restoration of coordinate information for the intelligent analysis results.
[0072] Combine Figure 4 The above steps are illustrated by examples. The backend device performs dynamic detection on the secondary stream of the video stream of Channel 1, and the obtained dynamic detection results include Region of Interest A and Region of Interest B, and it is detected that the target type of Region of Interest A is a person and the target type of Region of Interest B is a vehicle. The front-end device performs dynamic detection on the main stream of the video stream of Channel 2, and the obtained dynamic detection results include Region of Interest C and Region of Interest D, and it is detected that the target type of Region of Interest C is a vehicle and the target type of Region of Interest D is a person. After the backend device obtains the above dynamic detection results, it can stitch and fuse Region of Interest A and Region of Interest D with the target type of person into a video stream to be analyzed with the target type of person, and stitch and fuse Region of Interest B and Region of Interest C with the target type of vehicle into a video stream to be analyzed with the target type of vehicle; furthermore, the backend device can sequentially calculate the coordinate information of the regions of interest in the video stream to be analyzed with the target type of person and the video stream to be analyzed with the target type of vehicle after stitching and fusion, pack and integrate the coordinate information of each set of regions of interest in the coordinate system of the video stream to be analyzed after stitching and fusion, and the coordinate information of the corresponding regions of interest in the coordinate system of the main stream of the original channel, and then insert it into the intelligent frame of the video stream to be analyzed. Specifically, the intelligent frame can record the channel information of the region of interest, the coordinate information of the region of interest in the main stream of the original video stream, and the coordinate information of the region of interest in the video to be analyzed after fusion.
[0073] Step 304: The backend device inputs the picture frame to be analyzed into an analyzer with dynamic scheduling for analysis and processing.
[0074] In this application, the backend device can simultaneously activate multiple analyzers for intelligent analysis and processing, and dynamically allocate the analysis video stream composed of multiple frames of the to-be-analyzed images to a certain analyzer for processing. For example, if there is an empty analyzer, the backend device can retrieve some local frames of the regions of interest detected dynamically from the target area cache list for splicing and fusion, and input the spliced and fused analysis video stream into one of the analyzers for analysis; if all the analyzers are full, the backend device can pause retrieving the local frames of the regions of interest detected dynamically from the target cache list for splicing and fusion.
[0075] If the analysis video stream after splicing and fusion also includes the target types of the regions of interest, the backend device can activate the corresponding intelligent analysis capabilities for the target types in the analyzer according to the target types of the regions of interest in the analysis video. For example, if the target type is a person, the analyzer can perform refined analysis on the attributes of the person, such as detecting the gender of the target person, whether wearing a mask, whether wearing glasses and other attribute information; if the target type is a vehicle, the analyzer can perform refined analysis on the attributes of the vehicle, such as detecting the color of the target vehicle, license plate number, vehicle logo and other attribute information. It should be noted that the target types that each analyzer can analyze are not fixed, and different types of intelligent analysis capabilities can be activated according to the type of the analysis video stream.
[0076] Optionally, the dynamic detection result of the video stream of each channel also includes the frequency information of the video stream of this channel triggering dynamic detection. The backend device determines the frame rate of the analysis video stream according to the frequency information of the video streams of multiple channels triggering dynamic detection. Specifically, for the situation where the dynamic detection frequency values between channels vary greatly, that is, some channels trigger dynamic detection very frequently, while some channels only trigger a few times occasionally, then the channels with a very high dynamic detection trigger frequency can be fused into one or more analysis video streams with a larger frame rate and input into one or more analyzers with a larger intelligent analysis frame rate for processing, and the channels with a very low dynamic detection trigger frequency can be fused into an analysis video stream with a smaller frame rate and input into an analyzer with a smaller intelligent analysis frame rate for analysis.
[0077] In the embodiment of this application, after the backend device inputs the analysis video stream into the dynamically scheduled analyzer for analysis and processing, it is also necessary to determine the coordinate information of the analysis result in the coordinate system of the main stream according to the conversion relationship between the coordinate system of the main stream of the video stream of the first channel and the coordinate system of the analysis video stream.
[0078] Specifically, after the analyzer finishes analyzing and processing the video stream to be analyzed, the area corresponding to the analysis result may be a partial area in the local picture of the region of interest. For example, in the analyzer, a refined analysis is performed on whether a target person wears a mask, and the analysis result is that a certain target person does not wear a mask. Then, the area corresponding to the analysis result may be the head area of the target person. Therefore, it is necessary to further restore the partial area in the analysis result to the coordinate information in the coordinate system of the main code stream of the original channel, so as to know the specific position of the area corresponding to the analysis result in the main code stream of the original channel.
[0079] The specific conversion process is as follows: The backend device sequentially compares the coordinate information of the partial area in the analysis result with the coordinate information of each local picture of the region of interest in the frame. If the coordinate information area of a certain local picture of the region of interest contains the coordinate information area of this partial area, it is considered that the analysis result is the intelligent analysis result of this local picture of the region of interest. Then, combined with the coordinate information of the main code stream of the original video stream of this local picture of the region of interest recorded in the intelligent frame, the partial area of the analysis result is restored to the coordinate information in the coordinate system of the main code stream of the original video stream.
[0080] The following combines Figure 5 to illustrate the above steps with examples.
[0081] Taking the above Figure 4 as an example, the video stream to be analyzed with the target type of person after splicing and fusion is input into the analyzer with the target type of person for further analysis and processing. For example, further face detection can be performed on the video stream to be analyzed with the target type of person, such as detecting attributes such as gender, whether wearing a mask, and whether wearing glasses.
[0082] As Figure 5 shown, the coordinate of a partial area M(p1, q1, p2, q2) in the analysis result represents the head area of the local picture A of the region of interest. The coordinate information of the partial area M(p1, q1, p2, q2) in the analysis result is sequentially compared with the coordinate information of each local picture of the region of interest after fusion in the frame. If the coordinate information area of the coordinate A1(xx1, yy1, xx2, yy2) of a local picture of the region of interest contains the coordinate information area of the coordinate M(p1, q1, p2, q2), it is considered to be the intelligent analysis result of this local picture of the region of interest. Then, combined with the coordinate A(x1, y1, x2, y2) of the main code stream of the original video stream of this local picture of the region of interest recorded in the intelligent frame, the coordinate information of the partial area of the analysis result is restored to the coordinate O(xxx1, yyy1, xxx2, yyy2) in the coordinate system of the main code stream of the original video stream. The conversion relationship is as follows:
[0083] O(xxx1, yyy1, xxx2, yyy2) = (p1 * x1 / xx1, q1 * y1 / yy1, p2 * x2 / xx2, q2 * y2 / yy2). Similarly, Q(xxx7, yyy7, xxx8, yyy8) = (p3 * x7 / xx7, q3 * y7 / yy7, p4 * x8 / xx8, q4 * y8 / yy8).
[0084] Finally, pack the restored coordinate information and the intelligent analysis results and report them to the corresponding channel recorded within the frame. To understand the embodiments of the present application more clearly, the following combines Figure 6 to describe in detail the specific working process in the technical solution of the present application. Figure 6 It is applicable to the situation where the front-end device supports intelligent motion detection or the back-end device supports intelligent motion detection.
[0085] The specific steps are as follows:
[0086] Step 601: Determine whether the front-end device supports intelligent motion detection.
[0087] If the front-end device of a certain channel supports intelligent motion detection, receive the dynamic detection result of the video stream from the front-end device of this channel, and according to the target type of the region of interest in the dynamic detection result, store the local picture frames of the region of interest in the dynamic detection of the video stream of this channel, as well as the intelligent frames carrying the dynamic detection result, in the target region cache list corresponding to the target type; if not, execute step 602.
[0088] Step 602: The back-end device performs dynamic detection on the secondary code stream of the video stream of this channel.
[0089] According to the target type of the region of interest in the dynamic detection result, store the local picture frames of the region of interest in the dynamic detection of the video stream of this channel, as well as the intelligent frames carrying the dynamic detection result, in the target region cache list corresponding to the target type.
[0090] Step 603: Stitch and fuse the local picture frames of the regions of interest with the same target type in the video streams of multiple channels into multiple frames to be analyzed, and the multiple frames to be analyzed form one or more video streams to be analyzed.
[0091] Step 604: Input the stitched and fused video stream to be analyzed into the analyzer for analysis and processing.
[0092] Step 605: Convert the coordinate information in the analysis result of the analyzer into the coordinate information in the primary code stream coordinate system of the original video stream.
[0093] Step 606: Report the analysis result of the event.
[0094] The following combinesFigure 7 Describe in detail the specific work process in the technical solution of this application. Figure 7 It is applicable to the situation where the front-end device only supports ordinary motion detection and does not support intelligent motion detection, and the back-end device does not support intelligent motion detection.
[0095] Step 701: Receive the dynamic detection result of the video stream from the front-end device, and store the partial picture frames of the regions of interest in the video stream for dynamic detection, as well as the intelligent frames carrying the dynamic detection results, into the target area cache list.
[0096] Step 702: Stitch and fuse the partial picture frames of the regions of interest in the video streams of multiple channels into multiple frames to be analyzed, and the multiple frames to be analyzed form one or more video streams to be analyzed.
[0097] Step 703: Input the stitched and fused video stream to be analyzed into an analyzer for analysis and processing.
[0098] Step 704: Convert the coordinate information in the analysis result of the analyzer into the coordinate information in the main code stream coordinate system of the original video stream.
[0099] Step 705: Report the analysis result of the event.
[0100] A multi-channel video analysis method provided by an embodiment of this application first stitches and fuses the partial picture frames of the regions of interest in the video streams of multiple channels to obtain frames to be analyzed, and then inputs the frames to be analyzed into an analyzer for analysis and processing. Since the frames to be analyzed are obtained by stitching and fusing the partial picture frames of the regions of interest, the video stitching and fusion strategy is optimized, and the intelligent analysis capabilities of the back-end device can be reused with a smaller granularity. And the video stream to be analyzed composed of the frames to be analyzed is dynamically allocated to the analyzer. In this way, the intelligent analysis capabilities of the back-end device can be adjusted in a timely manner.
[0101] Based on the same technical concept, Figure 8 Exemplarily, an embodiment of this application provides a multi-channel video analysis device, which is used to implement the multi-channel video analysis method in the above embodiment.
[0102] As Figure 8 shown, the device 800 includes:
[0103] An acquisition module 801, configured to acquire the dynamic detection results of the video streams of multiple channels. The dynamic detection result of each channel's video stream includes the coordinate information of one or more regions of interest dynamically detected in the video stream of the channel and the information of the picture frame where the region of interest is located;
[0104] A processing module 802 is configured to, for the video stream of each of the multiple channels, intercept, from the picture frames in the video stream, a picture including the region of interest as a partial picture frame according to the coordinate information of the region of interest and the information of the picture frame where the region of interest is located; dynamically detect and splice and fuse the intercepted partial picture frames including the region of interest from the video streams of the multiple channels according to the coordinate information of the region of interest to obtain a picture frame to be analyzed; and input the picture frame to be analyzed into an analyzer with dynamic scheduling for analysis and processing.
[0105] In a possible design, the processing module 802 is specifically configured to: receive the dynamic detection result of the video stream of the channel from the front-end device connected to the channel; or perform dynamic detection on the secondary code stream of the received video stream of the channel to obtain the dynamic detection result.
[0106] In a possible design, the apparatus further includes a detection module 803 configured to perform dynamic detection on the secondary code stream of the video stream of the channel if the front-end device connected to the channel does not support identifying the target type of the region of interest.
[0107] In a possible design, the coordinate information of the region of interest included in the dynamic detection result is the coordinate information in the coordinate system of the primary code stream of the video stream of the channel; the processing module 802 is specifically configured to: after obtaining the region of interest by performing dynamic detection on the secondary code stream of the video stream of the channel, determine the coordinate information of the region of interest in the coordinate system of the primary code stream according to the coordinate information of the region of interest in the coordinate system of the secondary code stream and the conversion relationship between the coordinate system of the primary code stream and the coordinate system of the secondary code stream of the video stream of the channel.
[0108] In a possible design, the dynamic detection result further includes the target type corresponding to the region of interest; the processing module 802 is specifically configured to: splice and fuse the partial picture frames of the regions of interest corresponding to the same target type in the video streams of the multiple channels into one or more video streams to be analyzed; and input the video stream to be analyzed into the analyzer corresponding to the target type for analysis and processing.
[0109] In a possible design, the backend device includes a plurality of target region cache lists, each target region cache list corresponding to a target type, and the target region cache list is used to store the partial picture frames including the regions of interest corresponding to the target type.
[0110] In a possible design, the dynamic detection result further includes the frequency information of the video stream triggering the dynamic detection of the channel; specifically, the processing module 802 is configured to: determine the frame rate of the video stream to be analyzed according to the frequency information of the video streams of the multiple channels triggering the dynamic detection.
[0111] In a possible design, specifically, the processing module 802 is configured to: according to the conversion relationship between the coordinate system of the main code stream of the video stream of the first channel and the coordinate system of the video stream to be analyzed, determine the coordinate information of the analysis result in the coordinate system of the main code stream for the coordinate information of the analysis result obtained by analyzing and processing the frame of the picture to be analyzed.
[0112] Based on the same inventive concept, an embodiment of the present application further provides a computing device, including:
[0113] A memory for storing program instructions;
[0114] A processor for calling the program instructions stored in the memory and executing the above multi-channel video analysis method according to the obtained program instructions.
[0115] Based on the same inventive concept, an embodiment of the present application further provides a computer-readable storage medium, in which computer-readable instructions are stored, and when a computer reads and executes the computer-readable instructions, the above multi-channel video analysis method is implemented.
[0116] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0117] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or blocks. Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for implementing the functions specified in one block or a plurality of blocks.
[0119] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0120] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A multi-channel video analysis method, characterized in that, The method is applied to a backend device, and the method includes: Obtaining dynamic detection results of video streams of multiple channels, where the dynamic detection results of the video stream of each channel include coordinate information of the region of interest dynamically detected from the video stream of the channel and information of the picture frame where the region of interest is located; For the video stream of each channel among the multiple channels, according to the coordinate information of the region of interest and the information of the picture frame where the region of interest is located, intercepting the picture including the region of interest from the picture frame in the video stream as a partial picture frame; Performing video splicing and fusion on the partial picture frames including the region of interest intercepted from the video streams of the multiple channels to obtain a picture frame to be analyzed; Inputting the picture frame to be analyzed into an analyzer with dynamic scheduling for analysis and processing; For each channel among the multiple channels, obtaining the dynamic detection result of the video stream of the channel includes: receiving the dynamic detection result of the video stream of the channel from the front-end device connected to the channel; or, performing dynamic detection on the secondary code stream of the received video stream of the channel to obtain the dynamic detection result; If the front-end device connected to the channel does not support identifying the target type of the region of interest, then perform dynamic detection on the secondary code stream of the video stream of the channel.
2. The method according to claim 1, wherein The coordinate information of the region of interest included in the dynamic detection result is the coordinate information in the coordinate system of the primary code stream of the video stream of the channel; The method further includes: After performing dynamic detection on the secondary code stream of the video stream of the channel to obtain the region of interest, determining the coordinate information of the region of interest in the coordinate system of the primary code stream according to the coordinate information of the region of interest in the coordinate system of the secondary code stream and the conversion relationship between the coordinate system of the primary code stream and the coordinate system of the secondary code stream of the video stream of the channel.
3. The method according to claim 1, characterized in that, The dynamic detection result further includes the target type corresponding to the region of interest; The performing video splicing and fusion on the partial picture frames including the region of interest intercepted from the video streams of the multiple channels to obtain a picture frame to be analyzed includes: Splicing and fusing the partial picture frames of the regions of interest corresponding to the same target type in the video streams of the multiple channels into one or more video streams to be analyzed; Inputting the video stream to be analyzed into the analyzer corresponding to the target type for analysis and processing.
4. The method according to claim 3, characterized in that, The backend device includes multiple target region cache lists, each target region cache list corresponding to a target type, and the target region cache list is used to store the partial picture frames including the regions of interest corresponding to the target type.
5. The method according to claim 1 or 3, characterized in that, The dynamic detection result further includes the frequency information of the video stream of the channel triggering dynamic detection; Determining the frame rate of the video stream to be analyzed according to the frequency information of the video streams of the multiple channels triggering dynamic detection.
6. The method according to claim 1, wherein After inputting the picture frame to be analyzed into an analyzer with dynamic scheduling for analysis and processing, the method further includes: Based on the conversion relationship between the coordinate system of the main stream of the video stream of the channel and the coordinate system of the video stream to be analyzed, determine the coordinate information of the analysis result in the coordinate system of the main stream according to the coordinate information of the analysis result obtained by analyzing and processing the frame of the video stream to be analyzed.
7. A multi-channel video frame analysis device, characterized in that, It includes: An acquisition module, configured to acquire the dynamic detection results of the video streams of multiple channels. The dynamic detection result of the video stream of each channel includes the coordinate information of one or more regions of interest dynamically detected from the video stream of the channel and the information of the frame of the picture where the region of interest is located; A processing module, configured to, for the video stream of each channel among the multiple channels, intercept the picture including the region of interest from the frame in the video stream as a partial frame according to the coordinate information of the region of interest and the information of the frame of the picture where the region of interest is located; perform video stitching and fusion on the partial frames including the region of interest dynamically detected from the video streams of the multiple channels according to the coordinate information of the region of interest to obtain a frame of the video stream to be analyzed; input the frame of the video stream to be analyzed into an analyzer with dynamic scheduling for analysis and processing; The processing module is further configured to receive the dynamic detection result of the video stream of the channel from the front-end device connected to the channel; or perform dynamic detection on the secondary stream of the video stream of the channel received to obtain the dynamic detection result; If the front-end device connected to the channel does not support identifying the target type of the region of interest, perform dynamic detection on the secondary stream of the video stream of the channel.
8. A computing device, characterized in that, It includes: A memory, configured to store program instructions; A processor, configured to call the program instructions stored in the memory and execute the method according to any one of claims 1 to 6 according to the obtained program instructions.
9. A computer-readable storage medium, characterized in that, It includes computer-readable instructions. When a computer reads and executes the computer-readable instructions, the computer is caused to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method for implementing multichannel combined interested area video coding and transmission
CN101252687A
Video stream processing method and system, electronic device and storage medium
CN112954449A