Video recording device, off-line video analysis method, electronic device and storage medium
By performing parallel decoding and analysis of offline videos, the problem of low speed and accuracy of offline video analysis in the prior art is solved, especially in the cross section of the video segment, multiple crawlings of the same target are avoided, and more efficient video analysis is achieved.
Patent Information
- Application Number
- CN202211109765.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-09-13
AI Technical Summary
The prior art cannot effectively improve the analysis rate and accuracy when analyzing offline videos, especially when the same target is crawled multiple times in the intersection of adjacent video segments.
Multiple decoders are used to decode the multiple code stream of the target offline video in parallel to obtain multiple independent video sequences; then, based on the task to be analyzed, multiple intelligent analysis units are used to analyze these video sequences in parallel, and finally splicing them according to the analysis results of each video sequence to obtain the analysis results of the target offline video.
It effectively solves the problem of multiple crawlings for the same target, improves the accuracy of video analysis, and significantly improves the speed of video analysis through parallel decoding and analysis.
Smart Images

Figure CN115460369B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video analysis technology, and in particular to video recording equipment, offline video analysis methods, electronic equipment and storage media. Background Art
[0002] In the field of video analysis, it is usually possible to analyze the real-time video stream collected by a camera or a capture machine. For example, in some specific industry applications, some video sources (such as cameras or capture machines) can be connected to the video analysis system. When the video analysis system obtains the real-time video stream from the video source, it can analyze the real-time video stream; however, for offline videos, it is impossible to obtain the real-time video stream, and the offline video can only be viewed manually, resulting in a lot of time and human resources. For example, some offline videos are not intelligently analyzed when they are stored, so when retrieving the offline video, it is necessary to manually view it frame by frame, which is inefficient; for another example, after a user uploads a video recording, if he wants to extract the target information in the recording, he can only manually view it frame by frame, and cannot quickly obtain the desired information.
[0003] In the related art, offline videos can be segmented (for example, into 4 20-second video streams), with some overlap between two adjacent video segments, and then multiple video segments can be analyzed simultaneously to increase the analysis speed. However, the problem with the related art is that when there is an overlap between two adjacent video segments, the same target will be captured multiple times, resulting in inaccurate analysis results. Summary of the invention
[0004] The embodiments of the present application provide a video recording device, an offline video analysis method, an electronic device and a storage medium for improving the rate and accuracy of offline video analysis.
[0005] In a first aspect, the present application provides a video recording device, comprising: a code stream parsing module, a decoding module and an intelligent analysis module; the code stream parsing module is used to obtain a package of a target offline video from a storage space, and perform code stream parsing on the package of the target offline video to obtain multiple code streams of the target offline video; the decoding module is used to use multiple decoders to decode the multiple code streams of the target offline video in parallel to obtain multiple video sequences of the target offline video; wherein one code stream of the target offline video is decoded by one decoder; and the multiple video sequences are independent of each other; the intelligent analysis module is used to use multiple intelligent analysis units to perform parallel analysis on the multiple video sequences of the target offline video based on the task to be analyzed, The analysis result of each video sequence in multiple video sequences is obtained; wherein a video sequence of the target offline video is analyzed by an intelligent analysis unit; the task to be analyzed is a task of analyzing a target object in the target offline video; wherein the number of multiple decoders and the number of multiple intelligent analysis units are related to the decoding speed of the decoder and the analysis speed of the intelligent analysis unit; the intelligent analysis module is also used to splice the analysis result of each video sequence according to the sequence identifier of each video sequence to obtain the analysis result of the target offline video; wherein the sequence identifier of each video sequence includes: the frame number of each video frame included in each video sequence; or, the time when each video sequence is analyzed by the intelligent analysis unit.
[0006] It can be understood that the video recording device provided by the present application: uses multiple decoders to decode the multiple code streams of the target offline video in parallel to obtain multiple video sequences; then based on the task to be analyzed, uses multiple intelligent analysis units to analyze the multiple video sequences in parallel to obtain the analysis results of the multiple video sequences; finally, according to the analysis results of each video sequence in the multiple video sequences, the analysis results of the target offline video are obtained. It can be seen that compared with the method of segmenting offline videos and decoding and analyzing multiple video segments in the related art, the embodiment of the present application uses the video sequence as the smallest unit to perform parallel decoding and parallel analysis on multiple video sequences of a video, which effectively solves the phenomenon of multiple captures of the same target caused by analyzing multiple video segments at the same time in the related art, and effectively improves the accuracy of video analysis.
[0007] In addition, since decoding and intelligent analysis take the longest time in the entire video analysis process (i.e., bitstream parsing, decoding, intelligent analysis, and result integration), the embodiment of the present application uses multiple decoders for parallel decoding and multiple intelligent analysis units for parallel analysis, which effectively increases the speed of video analysis.
[0008] In one possible implementation, the intelligent analysis unit is specifically used to independently analyze each video frame of multiple video frames in a target video sequence based on a task to be analyzed to obtain an analysis result of each video frame; according to a sequence identifier of each video frame, the analysis result of each video frame is spliced to obtain an analysis result of the target video sequence; wherein the target video sequence is any one of the multiple video sequences; the sequence identifier of each video frame includes: a frame number of each video frame; or, the time when each video frame is analyzed by the intelligent analysis unit.
[0009] In another possible implementation, the task to be analyzed includes at least one of the following: a target detection task, a target classification task, or a target attribute recognition task; the analysis result of each video sequence includes at least one of the following: the target object in each video sequence and the location information of the target object, the category of the target object in each video sequence, or the attribute of the target object in each video sequence.
[0010] In another possible implementation, when the analysis speed of the intelligent analysis unit is greater than the decoding speed of the decoder, the number of the plurality of decoders is greater than the number of the plurality of intelligent analysis units; or,
[0011] When the analysis speed of the intelligent analysis unit is lower than the decoding speed of the decoder, the number of the plurality of decoders is less than the number of the plurality of intelligent analysis units.
[0012] In a second aspect, the present application provides an offline video analysis method, which is applied to the video recording device provided in the first aspect, the method comprising: parsing a code stream of an encapsulated packet of a target offline video to obtain multiple code streams of the target offline video; using multiple decoders to decode the multiple code streams of the target offline video in parallel to obtain multiple video sequences of the target offline video; wherein one code stream of the target offline video is decoded by one decoder; the multiple video sequences are independent of each other; based on the task to be analyzed, using multiple intelligent analysis units to analyze the multiple video sequences of the target offline video in parallel to obtain an analysis result of each video sequence in the multiple video sequences; wherein one video sequence in the target offline video is analyzed by one intelligent analysis unit; the task to be analyzed is a task of analyzing a target object in the target offline video; wherein the number of the multiple decoders and the number of the multiple intelligent analysis units are related to the decoding speed of the decoder and the analysis speed of the intelligent analysis unit; according to the sequence identifier of each video sequence, splicing the analysis results of each video sequence to obtain the analysis result of the target offline video; wherein the sequence identifier of each video sequence includes: the frame number of each video frame included in each video sequence; or the time when each video sequence is analyzed by the intelligent analysis unit.
[0013] In a possible implementation, based on the task to be analyzed, multiple intelligent analysis units are used to perform parallel analysis on multiple video sequences of the target offline video to obtain analysis results of each video sequence in the multiple video sequences, including: based on the task to be analyzed, each video frame in the multiple video frames of the target video sequence is independently analyzed to obtain analysis results of each video frame; according to the sequence identifier of each video frame, the analysis results of each video frame are spliced to obtain the analysis results of the target video sequence; wherein, the target video sequence is any one of the multiple video sequences; the sequence identifier of each video frame includes: the frame number of each video frame; or, the time when each video frame is analyzed by the intelligent analysis unit.
[0014] In another possible implementation, the task to be analyzed includes at least one of the following: a target detection task, a target classification task, or a target attribute recognition task; the analysis result of each video sequence includes at least one of the following: the target object in each video sequence and the location information of the target object, the category of the target object in each video sequence, or the attribute of the target object in each video sequence.
[0015] In another possible implementation, when the analysis speed of the intelligent analysis unit is greater than the decoding speed of the decoder, the number of multiple decoders is greater than the number of multiple intelligent analysis units; or, when the analysis speed of the intelligent analysis unit is less than the decoding speed of the decoder, the number of multiple decoders is less than the number of multiple intelligent analysis units.
[0016] In a third aspect, the present application provides an electronic device comprising: one or more processors; one or more memories; wherein the one or more memories are used to store computer program codes, the computer program codes include computer instructions, and when the one or more processors execute the computer instructions, the electronic device executes any one of the offline video analysis methods provided in the second aspect above.
[0017] In a fourth aspect, the present application provides a computer-readable storage medium, which stores computer execution instructions. When the computer execution instructions are executed on a computer, the computer executes any one of the offline video analysis methods provided in the second aspect above.
[0018] The specific description of the second to fourth aspects and their various implementations in this application can refer to the detailed description of the first aspect and its various implementations. The beneficial effects of the second to fourth aspects and their various implementations can refer to the beneficial effect analysis of the first aspect and its various implementations, which will not be repeated here.
[0019] These and other aspects of the present application will become more apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A schematic diagram of a video sequence provided in an embodiment of the present application;
[0021] Figure 2 A schematic diagram of a segmented video provided in an embodiment of the present application;
[0022] Figure 3 Schematic diagram of an implementation environment involved in an offline video analysis method provided in an embodiment of the present application Figure 1 ;
[0023] Figure 4 Schematic diagram of an implementation environment involved in an offline video analysis method provided in an embodiment of the present application Figure 2 ;
[0024] Figure 5 A schematic diagram of the structure of a video recording device provided in an embodiment of the present application Figure 1 ;
[0025] Figure 6 A schematic diagram of the structure of a video recording device provided in an embodiment of the present application Figure 2 ;
[0026] Figure 7 The process of an offline video analysis method provided in the embodiment of the present application is Figure 1 ;
[0027] Figure 8 A schematic diagram of a decoder performing a decoding operation provided in an embodiment of the present application Figure 1 ;
[0028] Fig. 9 A schematic diagram of a decoder performing a decoding operation provided in an embodiment of the present application Figure 2 ;
[0029] Fig.10 A schematic diagram of an intelligent analysis unit performing an analysis operation provided in an embodiment of the present application Figure 1 ;
[0030] Fig.11 A schematic diagram of an intelligent analysis unit performing an analysis operation provided in an embodiment of the present application Figure 2 ;
[0031] Fig.12 A schematic diagram of an application scenario of an offline video analysis method provided in an embodiment of the present application;
[0032] Fig.13 The process of an offline video analysis method provided in the embodiment of the present application is Figure 2 ;
[0033] Fig.14 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0034] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0035] The terms "first" and "second" and the like in the specification and drawings of this application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.
[0036] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.
[0037] It should be noted that, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0038] In the description of the present application, unless otherwise specified, “plurality” means two or more.
[0039] 1. Digital video recorder (DVR), DVR is a type of video recording device used in conjunction with analog cameras. The main working mode of DVR is to access analog audio and video signals and record audio and video through a hard disk. The core of DVR lies in hard disk recording, so DVR is also called a hard disk recorder.
[0040] DVR can integrate cameras, mice, remote controls, remote terminal devices, etc. to form a complete monitoring system, realizing the functions of long-term recording, remote monitoring, remote control, playback, intelligent analysis and backup of audio and video signal data.
[0041] 2. Network video recorder (NVR): NVR is a type of video recording device that is used in conjunction with a network camera or video encoder to record digital videos transmitted over the network.
[0042] The main function of NVR is to receive digital video code streams transmitted by network camera devices through the network, store and manage them, thereby realizing the distributed architecture advantage brought by networking. Through NVR, you can simultaneously watch, browse, play back, manage, intelligently analyze and store digital video code streams transmitted by multiple network camera devices.
[0043] 3. IP CAMERA (IPC), IP camera is a new generation of camera that combines traditional camera with network technology. It can transmit video images to the other side of the earth through the network, and the remote viewer does not need any professional software, just a standard network browser (such as Microsoft IE or Netscape) to monitor its video images. IPC is generally composed of lens, image sensor, sound sensor, signal processor, A / D converter, encoding chip, main control chip, network and control interface, etc.
[0044] 4. Program stream (PS) encapsulation: MPEG2-PS is a container for multiplexing digital audio, video, etc. The PS stream is obtained by encapsulating the basic bit stream output from the encoder in PS; the PS stream consists of PS packets.
[0045] 5. Group of pictures (GOP), also called video sequence, such as Figure 1 As shown, GOP is a group of continuous pictures, consisting of an I frame and several P frames.
[0046] Among them, I frame is an intra-coded frame (also called key frame), and P frame is a forward prediction frame (forward reference frame). Simply put, I frame is a complete picture, while P frame records the changes relative to I frame. Without I frame, P frame cannot be decoded.
[0047] In the H.264 compression standard, I frames and P frames are used to represent the transmitted video images. The encoder encodes multiple images into one or more GOP segments, and the decoder reads one or more GOP segments for decoding and then renders the images for display.
[0048] 6. Video frame: Video is a seemingly connected image composed of independent pictures, each of which is called a video frame. In order to ensure continuity and smoothness, the number of frames per second of the video is fixed, called the frame rate, for example: 25 frames / S, 30 frames / S, 50 frames / S, etc.
[0049] The above is an introduction to some concepts involved in the embodiments of the present application, which will not be repeated below.
[0050] As described in the background technology, in the field of video analysis, it is usually possible to analyze the real-time video stream collected by the camera or the capture machine. For example, in the application of some specific industries, some video sources (such as cameras or capture machines) can be connected to the video analysis system. When the video analysis system obtains the real-time video stream from the video source, the real-time video stream can be analyzed; however, for offline videos, it is impossible to obtain the real-time video stream, and the offline video can only be viewed manually, resulting in a large amount of time and human resources. For example, some offline videos are not intelligently analyzed when they are stored, so when the offline video is retrieved, it is necessary to manually view it frame by frame, which is inefficient; for another example, after a user uploads a video recording, if he wants to extract the target information in the recording, he can only manually view it frame by frame, and cannot quickly obtain the desired information.
[0051] In the related art, offline videos can be segmented, with some overlapping areas between two adjacent videos (if there is no overlapping area between the two videos, the target may be missed), and then multiple videos can be analyzed at the same time to improve the analysis speed. Figure 2 As shown in FIG. 1 , assuming that the offline video is 78 seconds long, the offline video can be divided into four videos of 20 seconds each, wherein there is a partial overlap between two adjacent videos (e.g. Figure 2 ), the overlapping part is analyzed twice during video analysis (for example, when analyzing the first video, the overlapping area from 18s to 20s is analyzed; when analyzing the second video, the overlapping area from 18s to 20s is analyzed repeatedly). Therefore, the same target may be captured multiple times, resulting in inaccurate analysis results.
[0052] In response to the above technical problems, an embodiment of the present application provides an offline video analysis method, the idea of which is: using multiple decoders to decode the multiple code streams of the target offline video in parallel to obtain multiple video sequences; then based on the task to be analyzed, using multiple intelligent analysis units to analyze the multiple video sequences in parallel to obtain the analysis results of the multiple video sequences; finally, according to the analysis results of each video sequence in the multiple video sequences, the analysis result of the target offline video is obtained. It can be seen that compared with the method of segmenting offline videos and decoding and analyzing multiple video segments in the related technology, the embodiment of the present application uses the video sequence as the smallest unit, and performs parallel decoding and parallel analysis on multiple video sequences of a video, which effectively solves the phenomenon of multiple captures of the same target caused by the simultaneous analysis of multiple video segments in the related technology, and effectively improves the accuracy of video analysis.
[0053] In addition, since decoding and intelligent analysis take the longest time in the entire video analysis process (i.e., bitstream parsing, decoding, intelligent analysis, and result integration), the embodiment of the present application uses multiple decoders for parallel decoding and multiple intelligent analysis units for parallel analysis, which effectively increases the speed of video analysis.
[0054] The embodiments provided in this application are described in detail below in conjunction with the accompanying drawings.
[0055] Please refer to Figure 3 , which shows a schematic diagram of an implementation environment involved in an offline video analysis method provided in an embodiment of the present application. Figure 3 As shown, the implementation environment may include: a video recording device 10 and a terminal device 20.
[0056] The video recording device 10 is used for performing video storage, video calculation, video processing, etc. Exemplarily, the video recording device 10 may be a digital video recorder DVR device, or a network video recorder NVR device.
[0057] In some embodiments, the video recording device 10 can receive and store a video code stream transmitted from a camera device (such as an analog camera or an IPC), a video encoding device, or a user device. Specifically, after receiving the video code stream, the video recording device encapsulates the video code stream (such as PS encapsulation) to obtain an encapsulation package of the video code stream, and stores the encapsulation package in a storage space (such as a hard disk).
[0058] Exemplarily, the video recording device 10 can encapsulate the video code stream in the PS encapsulation format. During the PS encapsulation process, the video recording device 10 determines the information corresponding to each video frame (including I frame and P frame) in the video code stream, and encapsulates the information corresponding to each video frame in the PS packet. Exemplarily, the information corresponding to the I frame includes the timestamp and frame number of the I frame; the information corresponding to the P frame includes the timestamp and frame number of the P frame.
[0059] It is understandable that if the video recording device 10 receives the video code stream, performs real-time decoding and playback or real-time decoding and analysis, the video corresponding to the video code stream is an online video; if the video recording device 10 receives the video code stream, encapsulates and stores the video code stream, the video corresponding to the video code stream is an offline video. The method provided in the embodiment of the present application is an analysis method for offline videos.
[0060] In some embodiments, the video recording device 10 is specifically used to obtain a package of an offline video from a storage space, perform code stream analysis on the package of the offline video, and obtain multiple code streams corresponding to the offline video; then, decode the multiple code streams corresponding to the offline video to obtain multiple video sequences corresponding to the offline video (a video sequence includes multiple video frames); finally, perform video analysis on the multiple video sequences corresponding to the offline video to obtain analysis results of the offline video.
[0061] In some embodiments, the video recording device 10 is further used to receive an offline video analysis instruction; the offline video analysis instruction includes: an identifier of the offline video to be analyzed and a task to be analyzed.
[0062] In this way, the video recording device 10 can obtain the package of the offline video to be analyzed from the storage space according to the offline video analysis instruction, and then perform operations such as code stream parsing, decoding and video analysis.
[0063] Optionally, the offline video analysis instruction may be issued by a user through a human-computer interaction interface of the video recording device 10 ; or, the offline video analysis instruction may be issued by the terminal device 20 .
[0064] In some embodiments, Figure 4 As shown, the video recording device 10 includes a code stream parsing module 11 , a decoding module 12 and an intelligent analysis module 13 .
[0065] The code stream parsing module 11 is used to obtain the encapsulation package of the target offline video from the storage space, and perform code stream parsing on the encapsulation package of the target offline video to obtain multiple code streams of the target offline video.
[0066] The decoding module 12 is used to use the multiple decoders to decode the multiple code streams of the target offline video in parallel to obtain multiple video sequences of the target offline video.
[0067] Wherein, one code stream of the target offline video is decoded by one decoder; and the multiple video sequences are independent of each other.
[0068] Optionally, the decoder may be a software decoder or a hardware decoder, wherein the hardware decoder is decoded by a graphics processing unit (GPU), and the software decoder is decoded by a central processing unit (CPU).
[0069] Optionally, the decoder may be a decoder provided by the video recording device 10 ; or, the decoder may be a device connected to the video recording device 10 and having a decoding function.
[0070] Optional, such as Figure 5 As shown, the above decoding module 12 may include multiple decoders; or, as shown in Figure 6 As shown, the above decoding module 12 can call multiple decoders.
[0071] The intelligent analysis module 13 is used to use multiple intelligent analysis units to perform parallel analysis on multiple video sequences of the target offline video based on the task to be analyzed, so as to obtain the analysis result of each video sequence in the multiple video sequences.
[0072] Wherein, a video sequence of the target offline video is analyzed by an intelligent analysis unit.
[0073] In some embodiments, the task to be analyzed is a task of analyzing a target object in a target offline video. Exemplarily, the task to be analyzed includes at least one of the following: a target detection task, a target classification task, or a target attribute recognition task. Therefore, the analysis result of each video sequence includes at least one of the following: a target object in each video sequence and the location information of the target object, a category of the target object in each video sequence, or an attribute of the target object in each video sequence.
[0074] In some embodiments, the intelligent analysis module 13 is specifically used to independently analyze each video frame of multiple video frames in the target video sequence based on the task to be analyzed to obtain the analysis result of each video frame; according to the sequence identification of each video frame, the analysis result of each video frame is spliced to obtain the analysis result of the target video sequence.
[0075] The target video sequence is any one of the multiple video sequences; the sequence identifier of each video frame includes: the frame number of each video frame; or the time when each video frame is analyzed by the intelligent analysis unit.
[0076] In some embodiments, the intelligent analysis module 13 is further used to splice the analysis results of each video sequence according to the sequence identifier of each video sequence to obtain the analysis results of the target offline video.
[0077] The sequence identifier of each video sequence includes: the frame number of each video frame included in each video sequence; or the time when each video sequence is analyzed by the intelligent analysis unit.
[0078] Optionally, the intelligent analysis unit may be an intelligent analysis resource provided by the video recording device 10 ; or, the intelligent analysis unit may be a device connected to the video recording device 10 and having an intelligent analysis function.
[0079] Optional, such as Figure 5 As shown, the intelligent analysis module 13 may include multiple intelligent analysis units; or Figure 6 As shown, the above-mentioned intelligent analysis module 13 can call multiple intelligent analysis units.
[0080] Exemplarily, the intelligent analysis unit includes but is not limited to analysis resources such as a graphics card, a graphics processing unit (GPU), a central processing unit (CPU), etc.
[0081] In some embodiments, the number of the plurality of decoders and the number of the plurality of intelligent analysis units are related to the decoding speed of the decoders and the analysis speed of the intelligent analysis units.
[0082] Exemplarily, when the analysis speed of the intelligent analysis unit is greater than the decoding speed of the decoder, the number of multiple decoders is greater than the number of multiple intelligent analysis units; or, when the analysis speed of the intelligent analysis unit is less than the decoding speed of the decoder, the number of multiple decoders is less than the number of multiple intelligent analysis units; or, when the analysis speed of the intelligent analysis unit is equal to the decoding speed of the decoder, the number of multiple decoders is equal to the number of multiple intelligent analysis units.
[0083] The terminal device 20 is used to remotely access the video recording device 10 .
[0084] In some embodiments, the terminal device 20 can remotely access the video recording device 10 and download offline video files from the video recording device 10 to analyze the offline video files.
[0085] In some embodiments, the user can initiate an offline video analysis instruction to the video recording device 10 through the terminal device 20, so that the video recording device 10 obtains the package of the offline video to be analyzed from the storage space, and then performs operations such as bitstream parsing, decoding and video analysis.
[0086] In some embodiments, the user can obtain the progress of the video recording device 10 analyzing the offline video through the terminal device 20; and the terminal device 20 can also control the progress of the video recording device 10 analyzing the offline video.
[0087] Exemplarily, the terminal device 20 may be an electronic device, such as a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) or virtual reality (VR) device, etc. The present disclosure does not impose any particular limitation on the specific form of the terminal device 20.
[0088] The offline video analysis method provided in this application is applied to the scenario of performing intelligent analysis on offline videos stored in the storage space of a video recording device, or on offline videos uploaded by users. Optionally, the offline video analysis method provided in this application can also be applied to other scenarios of performing intelligent analysis on offline videos. Examples will not be given one by one here.
[0089] The offline video analysis method provided in this application is applied to scenarios where intelligent analysis does not rely on the analysis results of other video frames. That is, the result of the current video frame does not rely on the analysis results of the previous video frame. Exemplarily, the offline video analysis method provided in this application is applied to analyze whether a target object appears in an offline video, or to analyze text in an offline video, etc.
[0090] It should be noted that the execution subject of the offline video analysis method provided in the present application is not limited. For example, the method can be executed by the video recording device itself, by the terminal device, or by an external device (for example, a processing device or an analysis server, etc.). For the convenience of subsequent description, the following embodiment is described by taking the method executed by the video recording device as an example.
[0091] The following is a detailed introduction to an offline video analysis method provided in an embodiment of the present application.
[0092] The offline video analysis method provided in the embodiment of the present application can be performed by a video recording device. Figure 7 As shown, the method comprises the following steps:
[0093] S101, performing code stream parsing on an encapsulation package of a target offline video to obtain multiple code streams of the target offline video.
[0094] In some embodiments, the above step S101 can be implemented as follows: in response to an offline video analysis instruction, obtaining a package of the target offline video from a storage space, and then performing code stream parsing on the package of the target offline video to obtain multiple code streams of the target offline video.
[0095] The target offline video is the offline video indicated by the offline video analysis instruction. Exemplarily, the offline video analysis instruction includes: the identification of the target offline video to be analyzed and the task to be analyzed. Optionally, the offline video can be a video file uploaded by the user to the video recording device; or, the offline video can be a video file stored in the storage space (such as a hard disk) of the video recording device.
[0096] Optionally, the above-mentioned offline video analysis instruction is initiated by a user through a human-computer interaction interface of a video recording device; or, the above-mentioned offline video analysis instruction is initiated by a user remotely to a video recording device through a terminal device.
[0097] In some embodiments, the task to be analyzed in the offline video analysis instruction is a task of analyzing a target object in a target offline video. Exemplarily, the task to be analyzed includes but is not limited to: a target detection task, a target classification task, or a target attribute recognition task.
[0098] The target detection task is used to detect the target object of interest from each video frame of the target offline video. For example, the target detection task may include: detecting faces in the target offline video, detecting objects in the target offline video, detecting text in the target offline video, etc.
[0099] The target classification task is used to determine the category of the target object identified from each video frame of the target offline video. For example, if the target object is a person, the target object includes two categories: staff and intruder; then the target classification task may include: determining whether the person identified from each video frame of the target offline video belongs to a staff member or an intruder.
[0100] The target attribute recognition task is used to determine the attributes of the target object in each video frame of the target offline video. For example, the target attribute recognition task may include: target height detection, target color detection, target occlusion detection, etc.
[0101] It is understandable that after receiving the video code stream, the video recording device needs to encapsulate the video code stream according to a certain encapsulation format to obtain an encapsulation package, and store the encapsulation package in a storage space (such as a hard disk). Among them, video encapsulation refers to putting the encoded and compressed video code stream into a file according to a certain format. There are many video encapsulation formats, such as: PS encapsulation, flash video (FLV) encapsulation format, multimedia (MKVToolNix, MKV) encapsulation format, digital multimedia (MPEG-4Part 14, MP4) encapsulation format, etc.
[0102] Therefore, after receiving the offline video analysis instruction, the video recording device needs to obtain the encapsulation package of the target offline video from the storage space, and obtain multiple streams of the target offline video after stream parsing.
[0103] S102: using multiple decoders to decode multiple code streams of the target offline video in parallel to obtain multiple video sequences of the target offline video.
[0104] Among them, one bit stream of the target offline video is decoded by one decoder.
[0105] The above-mentioned multiple video sequences are independent of each other. It can be understood that a video sequence consists of an I frame and at least one P frame. The I frame is a complete picture, and the P frame is used to record the changes relative to the previous frame (which can be an I frame or a P frame). In other words, a video sequence includes a complete action. Therefore, the multiple video sequences are independent of each other.
[0106] In some embodiments, the video recording device includes a decoder resource pool (the decoder resource pool includes multiple decoders). When decoding, the video recording device can obtain a certain number of decoders from the decoder resource pool and decode multiple code streams at the same time.
[0107] In other embodiments, the video recording device is connected to a device with a decoding function, so that when decoding, the video recording device can call a certain number of decoders to decode multiple code streams at the same time.
[0108] As a possible implementation, the number of decoders may be the same as the number of multiple streams of the target offline video. Figure 8 As shown, assuming that the number of multiple code streams of the target offline video is N (N is an integer greater than 0), the number of multiple decoders can be N. In this way, multiple decoders can be implemented in parallel to decode the multiple code streams of the target offline video at the same time, which greatly improves the decoding rate.
[0109] As another possible implementation, when the number of decoders is determined, the number of code stream paths during each decoding is the same as the number of decoders. Fig. 9 As shown, assuming that there are 3 decoders and the target offline video has 9 bitstreams, 3 decoders are used in parallel to decode the 3 bitstreams each time the decoding is performed. In this way, three decodings are required to complete the decoding of the target offline video.
[0110] As another possible implementation, the number of decoders is determined by the decoding speed of the decoders. For example, if the decoding speed of the decoders is faster, the number of decoders can be reduced accordingly; if the decoding speed of the decoders is slower, the number of decoders can be increased accordingly.
[0111] It is understandable that the decoder provided in the embodiment of the present application is used to decode the multiple code streams of the target offline video to obtain multiple video sequences of the target offline video (wherein one code stream corresponds to one video sequence). Since the video sequence is shorter than the video clip, compared with the technology of using a decoder to decode the video clip in the related art, the embodiment of the present application uses a decoder to decode the code stream corresponding to a video sequence, which can improve the decoding speed.
[0112] S103: Based on the task to be analyzed, multiple intelligent analysis units are used to perform parallel analysis on multiple video sequences of the target offline video to obtain an analysis result of each video sequence in the multiple video sequences.
[0113] Among them, a video sequence in the target offline video is analyzed by an intelligent analysis unit. It can be understood that since each intelligent analysis unit is independent, the analysis results of each video sequence are independent of each other and do not depend on each other, that is, the analysis result of the current video sequence is not dependent on the analysis result of the previous video sequence.
[0114] In some embodiments, the task to be analyzed is a task of analyzing a target object in a target offline video. The task to be analyzed includes at least one of the following: a target detection task, a target classification task, or a target attribute recognition task; therefore, the analysis result of each video sequence includes at least one of the following: a target object in each video sequence and the location information of the target object, a category of the target object in each video sequence, or an attribute of the target object in each video sequence. For example, assuming that the task to be analyzed is to identify a vehicle in a target offline video and detect the color information of the vehicle; then the analysis result of each video sequence includes: the vehicle in each video frame of each video sequence, and the color information of the vehicle.
[0115] Specifically, according to the task to be analyzed, the analysis algorithm corresponding to the task to be analyzed is determined, and then multiple intelligent analysis units simultaneously use the analysis algorithm corresponding to the task to be analyzed to analyze multiple video sequences to obtain the analysis result of each video sequence in the multiple video sequences. Exemplarily, if the task to be analyzed is a face recognition task, the analysis algorithm corresponding to the task to be analyzed is a face recognition algorithm, and multiple intelligent analysis units simultaneously use the face recognition algorithm to perform face recognition on multiple video sequences to obtain the face recognition result of each video sequence in the multiple video sequences.
[0116] It is understandable that the intelligent analysis unit is configured with a variety of intelligent analysis algorithms (for example, target recognition algorithm, target detection algorithm, etc.). When performing video analysis, the corresponding analysis algorithm can be selected according to the task to be analyzed.
[0117] In some embodiments, the video recording device includes an intelligent analysis unit resource pool (the intelligent analysis unit pool includes multiple intelligent analysis units). When analyzing offline videos, the video recording device can obtain a certain number of intelligent analysis units from the intelligent analysis unit resource pool and analyze multiple video sequences at the same time.
[0118] In other embodiments, the video recording device is connected to a device having an intelligent analysis function, so that when the video recording device performs intelligent analysis, it can call a certain number of intelligent analysis units to analyze multiple video sequences at the same time.
[0119] As a possible implementation, the number of intelligent analysis units may be the same as the number of multiple video sequences of the target offline video. Fig.10 As shown, assuming that the number of multiple video sequences of the target offline video is N (N is an integer greater than 0), the number of multiple intelligent analysis units can be N. In this way, multiple intelligent analysis units can be implemented in parallel to analyze multiple video sequences of the target offline video at the same time, which greatly improves the video analysis rate.
[0120] As another possible implementation, when the number of intelligent analysis units is determined, the number of the plurality of video sequences in each analysis is the same as the number of the intelligent analysis units. Fig.11 As shown, assuming that the number of intelligent analysis units is 3 and the number of video sequences of the target offline video is 9, then in each analysis, 3 intelligent analysis units are used in parallel to analyze 3 video sequences, and the video analysis of the target offline video needs to be completed three times.
[0121] As another possible implementation, the number of intelligent analysis units is determined by the analysis speed of the intelligent analysis units. For example, if the analysis speed of the intelligent analysis units is faster, the number of intelligent analysis units can be reduced accordingly; if the analysis speed of the intelligent analysis units is slower, the number of intelligent analysis units can be increased accordingly.
[0122] Optionally, the intelligent analysis unit may be a processor such as a graphics card, a GPU, a CPU, etc. It is understandable that different types of processors have different analysis capabilities, analysis speeds, etc. Therefore, the analysis speed of the intelligent analysis unit may be determined by the type of processor.
[0123] In some embodiments, the number of the plurality of decoders and the number of the plurality of intelligent analysis units are related to the decoding speed of the decoders and the analysis speed of the intelligent analysis units.
[0124] Exemplarily, when the analysis speed of the intelligent analysis unit is equal to the decoding speed of the decoder, the number of the plurality of decoders is equal to the number of the plurality of intelligent analysis units. Fig.12 As shown, when the analysis speed of the intelligent analysis unit is the same as the decoding speed of the decoder, and the number of decoders is the same as the number of intelligent analysis units, the decoder decodes multiple streams to obtain multiple video sequences. The video sequences can be directly input into the corresponding intelligent analysis unit for analysis without queuing for decoding or queuing for analysis, so that the decoder and the intelligent analysis unit can cooperate better.
[0125] In another exemplary embodiment, when the analysis speed of the intelligent analysis unit is greater than the decoding speed of the decoder, the number of the plurality of decoders is greater than the number of the plurality of intelligent analysis units.
[0126] It is understandable that if the analysis speed of the intelligent analysis unit is fast, while the decoding speed of the decoder is slow, there may be a situation where the decoding and analysis are "in short supply" (i.e., the intelligent analysis unit is idle and waiting for the decoder to decode). Therefore, the number of intelligent analysis units can be set to be less than the number of decoders, so that the decoder and the intelligent analysis unit can cooperate better.
[0127] As another example, when the analysis speed of the intelligent analysis unit is lower than the decoding speed of the decoder, the number of the plurality of decoders is less than the number of the plurality of intelligent analysis units.
[0128] It is understandable that if the analysis speed of the intelligent analysis unit is slow and the decoding speed of the decoder is fast, a large number of video sequences may be queued up waiting for analysis by the intelligent analysis unit. Therefore, the number of decoders set can be less than the number of intelligent analysis units, so that the decoder and the intelligent analysis unit can cooperate better.
[0129] In summary, the embodiment of the present application determines the number of decoders and the number of intelligent analysis units according to the decoding speed of the decoder and the analysis speed of the intelligent analysis unit, which can ensure the orderly cooperation between the decoder and the intelligent analysis unit, that is, the video sequence decoded by the decoder can directly enter the corresponding intelligent analysis unit for analysis, and there will be no queue waiting for decoding or queuing for analysis, which effectively improves the speed of decoding and analysis of the target offline video.
[0130] In some embodiments, Fig.13 As shown, the above step S103 can be specifically implemented as follows:
[0131] S1031. Based on the task to be analyzed, independently analyze each of the multiple video frames in the target video sequence to obtain an analysis result of each video frame.
[0132] The analysis result of each video frame is independent, and there is no dependence between the analysis results of each video frame, that is, the analysis result of the current video frame is not dependent on the analysis result of the previous video frame.
[0133] Exemplarily, if the task to be analyzed is a face recognition task, the intelligent analysis unit uses a face recognition algorithm to perform face recognition on multiple video frames of the target video sequence respectively, and obtains a face recognition result for each of the multiple video frames.
[0134] As another example, if the task to be analyzed is a vehicle recognition task and a vehicle color detection task, the intelligent analysis unit adopts a vehicle recognition algorithm and a vehicle color detection algorithm to respectively detect vehicle recognition and vehicle color detection on multiple video frames of the target video sequence, and obtains the vehicle position information and vehicle color information in each of the multiple video frames.
[0135] S1032. According to the sequence identifier of each video frame, the analysis result of each video frame is spliced to obtain the analysis result of the target video sequence.
[0136] The target video sequence is any one of the multiple video sequences; the sequence identifier of each video frame includes: the frame number of each video frame; or the time when each video frame is analyzed by the intelligent analysis unit.
[0137] It is understandable that the intelligent analysis unit analyzes the video frames one by one, so the time when each video frame of the target video sequence is analyzed by the intelligent analysis unit is different, which can be used to indicate the order in which the intelligent analysis unit analyzes the video frames. Therefore, according to the time when each video frame is analyzed by the intelligent analysis unit, the positions of multiple video frames in the target video sequence can be determined, and then the order of the analysis results of the multiple video frames can be determined to obtain the analysis results of the target video sequence.
[0138] Exemplarily, assuming that the multiple video frames of the target video sequence include: a first video frame, a second video frame, a third video frame, and a fourth video frame, wherein the first video frame is analyzed by the intelligent analysis unit at the 20th second, the second video frame is analyzed by the intelligent analysis unit at the 60th second, the third video frame is analyzed by the intelligent analysis unit at the 40th second, and the fourth video frame is analyzed by the intelligent analysis unit at the 80th second, then the order of the video frames in the target video sequence should be: the first video frame, the third video frame, the second video frame, and the fourth video frame. Therefore, the order of the analysis results of the multiple video frames of the target video sequence is: the analysis result of the first video frame, the analysis result of the third video frame, the analysis result of the second video frame, and the analysis result of the fourth video frame.
[0139] The frame number of the above-mentioned video frame can represent the position of the video frame in the video sequence. Therefore, according to the frame number of each video frame in the multiple video frames, the position of the multiple video frames in the target video sequence can be determined, and then the order of the analysis results of the multiple video frames can be determined to obtain the analysis results of the target video sequence.
[0140] In another exemplary embodiment, assuming that the multiple video frames of the target video sequence include: a first video frame, a second video frame, a third video frame, and a fourth video frame, wherein the frame number of the first video frame is 1, the frame number of the second video frame is 3, the frame number of the third video frame is 2, and the frame number of the fourth video frame is 4, then the order of the video frames in the target video sequence should be: the first video frame, the third video frame, the second video frame, and the fourth video frame. Therefore, the order of the analysis results of the multiple video frames of the target video sequence is: the analysis result of the first video frame, the analysis result of the third video frame, the analysis result of the second video frame, and the analysis result of the fourth video frame.
[0141] S104: splicing the analysis results of each video sequence according to the sequence identifier of each video sequence to obtain the analysis result of the target offline video.
[0142] The sequence identifier of each video sequence includes: the frame number of each video frame included in each video sequence; or the time when each video sequence is analyzed by the intelligent analysis unit.
[0143] In some embodiments, according to the time when one or more video frames in each video sequence are analyzed by the intelligent analysis unit, the analysis results of multiple video sequences are spliced to obtain the analysis result of the target offline video. Exemplarily, according to the time when the I frame of each video sequence is analyzed by the intelligent analysis unit, the analysis results of multiple video sequences are spliced to obtain the analysis result of the target offline video.
[0144] In some embodiments, the analysis results of multiple video sequences are spliced according to the frame numbers of one or more video frames in each video sequence to obtain the analysis results of the target offline video. Exemplarily, the analysis results of multiple video sequences are spliced according to the frame numbers of I frames in each video sequence to obtain the analysis results of the target offline video.
[0145] Based on the technical solution provided by the embodiment of the present application, at least the following beneficial effects can be produced: multiple decoders are used to decode the multiple code streams of the target offline video in parallel to obtain multiple video sequences; then based on the task to be analyzed, multiple intelligent analysis units are used to analyze the multiple video sequences in parallel to obtain the analysis results of the multiple video sequences; finally, according to the analysis results of each video sequence in the multiple video sequences, the analysis results of the target offline video are obtained. It can be seen that compared with the method of segmenting offline videos and decoding and analyzing multiple video segments in the related art, the embodiment of the present application uses the video sequence as the smallest unit to perform parallel decoding and parallel analysis on multiple video sequences of a video, which effectively solves the phenomenon of multiple captures of the same target caused by analyzing multiple video segments at the same time in the related art, and effectively improves the accuracy of video analysis.
[0146] In addition, since decoding and intelligent analysis take the longest time in the entire video analysis process (i.e., bitstream parsing, decoding, intelligent analysis, and result integration), the embodiment of the present application uses multiple decoders for parallel decoding and multiple intelligent analysis units for parallel analysis, which effectively increases the speed of video analysis.
[0147] The present application embodiment provides a schematic diagram of the structure of an electronic device, which is used to perform the offline video analysis method provided in the above embodiment. Fig.14 As shown, the electronic device 400 includes: a processor 402 , a communication interface 403 , and a bus 404 . Optionally, the electronic device 400 may further include a memory 401 .
[0148] The processor 402 may be a processor that implements or executes various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor 402 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor 402 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0149] The communication interface 403 is used to connect with other devices via a communication network, such as Ethernet, wireless access network, wireless local area network (WLAN), etc.
[0150] The memory 401 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program codes in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0151] As a possible implementation, the memory 401 may exist independently of the processor 402, and the memory 401 may be connected to the processor 402 via a bus 404 for storing instructions or program codes. When the processor 402 calls and executes the instructions or program codes stored in the memory 401, the offline video analysis method provided in the embodiment of the present application can be implemented.
[0152] In another possible implementation, the memory 401 may also be integrated with the processor 402 .
[0153] The bus 404 may be an extended industry standard architecture (EISA) bus, etc. The bus 404 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.14 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0154] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the electronic device can be divided into different functional modules to complete all or part of the functions described above.
[0155] The embodiment of the present application also provides a computer-readable storage medium. All or part of the processes in the above method embodiment can be completed by computer instructions to instruct the relevant hardware, and the program can be stored in the above computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. The computer-readable storage medium can be any of the above embodiments or memory. The above computer-readable storage medium can also be an external storage device of the above electronic device, such as a plug-in hard disk, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. equipped on the above electronic device. Further, the above computer-readable storage medium can also include both the internal storage unit of the above electronic device and an external storage device. The above computer-readable storage medium is used to store the above computer program and other programs and data required by the above electronic device. The above computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.
[0156] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program product is run on a computer, the computer is enabled to execute any one of the offline video analysis methods provided in the above embodiments.
[0157] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other changes to the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "one" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0158] Although the present application has been described in conjunction with specific features and embodiments thereof, it is obvious that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely exemplary illustrations of the present application as defined by the appended claims, and are deemed to have covered any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
[0159] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A video recording device, characterized in that: include: Code stream parsing module, decoding module and intelligent analysis module; The code stream parsing module is used to obtain the encapsulation package of the target offline video from the storage space, and perform code stream parsing on the encapsulation package of the target offline video to obtain the multi-channel code stream of the target offline video; The decoding module is used to use multiple decoders to decode multiple code streams of the target offline video in parallel to obtain multiple video sequences of the target offline video; wherein one code stream of the target offline video is decoded by one decoder; and the multiple video sequences are independent of each other; The intelligent analysis module is used to use multiple intelligent analysis units to perform parallel analysis on multiple video sequences of the target offline video based on the task to be analyzed, and obtain the analysis result of each video sequence in the multiple video sequences; wherein one video sequence of the target offline video is analyzed by one intelligent analysis unit; the task to be analyzed is a task of analyzing the target object in the target offline video; wherein the number of the multiple decoders and the number of the multiple intelligent analysis units are related to the decoding speed of the decoder and the analysis speed of the intelligent analysis unit; The intelligent analysis module is also used to splice the analysis results of each video sequence according to the sequence identifier of each video sequence to obtain the analysis results of the target offline video; wherein the sequence identifier of each video sequence includes: the frame number of each video frame included in each video sequence; or, the time when each video sequence is analyzed by the intelligent analysis unit.
2. The device according to claim 1, characterized in that The intelligent analysis unit is specifically used to independently analyze each video frame of the multiple video frames of the target video sequence based on the task to be analyzed to obtain the analysis result of each video frame; splicing the analysis result of each video frame according to the sequence identification of each video frame to obtain the analysis result of the target video sequence; wherein the target video sequence is any one of the multiple video sequences; the sequence identification of each video frame includes: the frame number of each video frame; or, the time when each video frame is analyzed by the intelligent analysis unit.
3. The device according to claim 1 or 2, characterized in that The task to be analyzed includes at least one of the following: a target detection task, a target classification task or a target attribute recognition task; the analysis result of each video sequence includes at least one of the following: the target object in each video sequence and the location information of the target object, the category of the target object in each video sequence or the attribute of the target object in each video sequence.
4. The device according to claim 1 or 2, characterized in that In the case where the analysis speed of the intelligent analysis unit is greater than the decoding speed of the decoder, the number of the plurality of decoders is greater than the number of the plurality of intelligent analysis units; or, In the case where the analysis speed of the intelligent analysis unit is lower than the decoding speed of the decoder, the number of the plurality of decoders is less than the number of the plurality of intelligent analysis units.
5. An offline video analysis method, characterized in that: The method comprises: Performing code stream parsing on the encapsulation package of the target offline video to obtain multiple code streams of the target offline video; Using multiple decoders to decode multiple code streams of the target offline video in parallel to obtain multiple video sequences of the target offline video; wherein one code stream of the target offline video is decoded by one decoder; and the multiple video sequences are independent of each other; Based on the task to be analyzed, multiple intelligent analysis units are used to analyze multiple video sequences of the target offline video in parallel to obtain analysis results of each video sequence in the multiple video sequences; wherein one video sequence in the target offline video is analyzed by one intelligent analysis unit; the task to be analyzed is a task of analyzing a target object in the target offline video; wherein the number of the multiple decoders and the number of the multiple intelligent analysis units are related to the decoding speed of the decoders and the analysis speed of the intelligent analysis units; According to the sequential identifier of each video sequence, the analysis results of each video sequence are spliced to obtain the analysis results of the target offline video; wherein the sequential identifier of each video sequence includes: the frame number of each video frame included in each video sequence; or, the time when each video sequence is analyzed by the intelligent analysis unit.
6. The method according to claim 5, characterized in that Based on the task to be analyzed, multiple intelligent analysis units are used to perform parallel analysis on multiple video sequences of the target offline video to obtain analysis results of each video sequence in the multiple video sequences, including: Based on the task to be analyzed, each of the multiple video frames of the target video sequence is independently analyzed to obtain an analysis result of each of the video frames; According to the sequence identifier of each video frame, the analysis results of each video frame are spliced to obtain the analysis results of the target video sequence; wherein the target video sequence is any one of the multiple video sequences; the sequence identifier of each video frame includes: the frame number of each video frame; or, the time when each video frame is analyzed by the intelligent analysis unit.
7. The method according to claim 5 or 6, characterized in that: The task to be analyzed includes at least one of the following: a target detection task, a target classification task or a target attribute recognition task; the analysis result of each video sequence includes at least one of the following: the target object in each video sequence and the location information of the target object, the category of the target object in each video sequence or the attribute of the target object in each video sequence.
8. The method according to claim 5 or 6, characterized in that: In the case where the analysis speed of the intelligent analysis unit is greater than the decoding speed of the decoder, the number of the plurality of decoders is greater than the number of the plurality of intelligent analysis units; or, In the case where the analysis speed of the intelligent analysis unit is lower than the decoding speed of the decoder, the number of the plurality of decoders is less than the number of the plurality of intelligent analysis units.
9. An electronic device, characterized in that: include: one or more processors; one or more memories; Wherein, the one or more memories are used to store computer program codes, and the computer program codes include computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the offline video analysis method according to any one of claims 5 to 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed on a computer, the computer executes the offline video analysis method according to any one of claims 5 to 8.
Citation Information
Patent Citations
Lossless video acceleration analysis method
CN107888924A
Method, device and equipment for optimizing intelligent video analysis performance
CN109769115A