An image transmission and storage method, device, electronic equipment, and storage medium

By encoding the target frame image into a switching prediction frame image and encapsulating it together with other images into a data packet for transmission, the bandwidth consumption problem when intelligent events are frequently triggered is solved, achieving fast decoding and efficient transmission, and improving device performance.

CN122293867APending Publication Date: 2026-06-26ZHEJIANG UNIVIEW TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIVIEW TECH CO LTD
Filing Date
2024-12-25
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In scenarios where smart events are frequently triggered or when a large number of shooting devices are connected, the transmission of images and target object data triggered by smart events consumes a large amount of network bandwidth, leading to a decrease in the performance of shooting devices and image management devices.

Method used

The target frame image is encoded into a switching prediction frame image, and then encapsulated with other encoded images into a data packet and sent to the image management device. The switching prediction frame image can be used to quickly decode the key frame image, reducing the transmission bandwidth usage.

Benefits of technology

By reducing transmission bandwidth usage, the performance of both the front-end shooting equipment and the back-end image management equipment was ensured, and transmission efficiency and equipment processing capabilities were improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122293867A_ABST
    Figure CN122293867A_ABST
Patent Text Reader

Abstract

This invention discloses an image transmission and storage method, apparatus, electronic device, and storage medium. The image transmission method is executed by a shooting device and includes: if it is determined that at least one target object exists in a target frame image, encoding the target frame image into a switching prediction frame image; encapsulating the switching prediction frame image and other encoded images into a data packet, and sending the data packet to an image management device to achieve the transmission of the target frame image. This invention can reduce the bandwidth occupied by image transmission and improve the performance of the front-end shooting device and the back-end image management device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image transmission and storage method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of image processing technology and the popularization of intelligent shooting devices, shooting devices have begun to support more and more types of intelligent event detection, such as pedestrian detection, vehicle detection, and detection of not wearing masks. When the shooting device detects a triggering intelligent event, it will upload the frame image that triggered the intelligent event, the thumbnail of the target object in that frame image, and the target object's related data (such as quantity, type, location, and attribute information) to the backend NVR (Network Video Recorder) device or server.

[0003] However, in scenarios where intelligent events are frequently triggered, or when a large number of shooting devices are connected, the transmission of the aforementioned data will consume a significant amount of network bandwidth, leading to a decrease in the video transmission bandwidth of the shooting devices and affecting the performance of both the front-end shooting devices and the back-end image management devices. Summary of the Invention

[0004] This invention provides an image transmission and storage method, apparatus, electronic device, and storage medium to reduce the bandwidth occupied by image transmission and improve the performance of front-end shooting equipment and back-end image management equipment.

[0005] In a first aspect, embodiments of the present invention provide an image transmission method, which is executed by a shooting device, and the method includes:

[0006] If it is determined that there is at least one target object in the target frame image, then the target frame image is encoded as a switching prediction frame image;

[0007] The switching prediction frame image and other encoded images are encapsulated into data packets, and the data packets are sent to the image management device to achieve the transmission of the target frame image.

[0008] Secondly, embodiments of the present invention also provide an image storage method, which is executed by an image management device, and the method includes:

[0009] Receive data packets sent by the imaging device and determine the switching prediction frame image in the data packets;

[0010] Store each frame of the image in the data packet;

[0011] The storage locations of the switch prediction frame image and the keyframe image that matches the switch prediction frame image are stored so that the switch prediction frame image can be decoded based on the storage locations.

[0012] Thirdly, embodiments of the present invention also provide an image transmission device, which is deployed on a shooting device, and the device includes:

[0013] The target frame image encoding module is used to determine, if it is determined that there is at least one target object in the target frame image, to encode the target frame image into a switching prediction frame image;

[0014] The data packet sending module is used to encapsulate the switching prediction frame image and other encoded images into data packets and send the data packets to the image management device to realize the transmission of the target frame image.

[0015] Fourthly, embodiments of the present invention also provide an image storage device, which is deployed in an image management device, and the device includes:

[0016] The switching prediction frame image determination module is used to receive data packets sent by the shooting device and determine the switching prediction frame images in the data packets;

[0017] The data storage module is used to store the images of each frame in the data packet;

[0018] The storage location module is used to store the storage location of the switching prediction frame image and the storage location of the key frame image that matches the switching prediction frame image, so as to decode the switching prediction frame image according to the storage location.

[0019] Fifthly, embodiments of the present invention also provide an image transmission system, which includes an image capturing device and an image management device;

[0020] The shooting device is used to encode a target frame image containing at least one target object into a switching prediction frame image;

[0021] The image of the switched prediction frame image is encapsulated into a data packet and sent to the image management device.

[0022] Other encoded images include keyframe images and forward prediction frame images;

[0023] The image management device is used to receive data packets sent by the shooting device and determine the switching prediction frame image in the data packets;

[0024] Store each frame image in the data packet, and store the storage location of the switching prediction frame image and the storage location of the key frame image that matches the switching prediction frame image.

[0025] The image management device is also used to decode based on the storage location of the switching prediction frame image and the storage location of the key frame image that matches the switching prediction frame image.

[0026] In a sixth aspect, embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the image transmission method as described in any of the embodiments of the present invention, or implements the image storage method as described in any of the embodiments of the present invention.

[0027] In a seventh aspect, embodiments of the present invention also provide a storage medium for storing computer-executable instructions, which, when executed by a computer processor, are used to perform an image transmission method as described in any of the embodiments of the present invention, or to perform an image storage method as described in any of the embodiments of the present invention.

[0028] The technical solution of this invention encapsulates the target frame image into a switching prediction frame image when a target object exists in the target frame image. This switching prediction frame image is then encapsulated together with other encoded images into a data packet, which is then sent to the image management device. This solves the problem in existing technologies where, in scenarios with frequent intelligent event triggering or when multiple shooting devices are connected, the transmission of images triggered by intelligent events and target object data consumes significant network bandwidth, leading to a decrease in the video transmission bandwidth of the shooting devices. Encapsulating the target frame image containing the target object into a switching prediction frame image allows for fast decoding by referencing keyframe images. Therefore, it eliminates the need for additional transmission of the target frame image, reducing bandwidth consumption and ensuring the performance of both the front-end shooting device and the back-end image management device.

[0029] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a flowchart of an image transmission method provided in Embodiment 1 of the present invention;

[0032] Figure 2 This is a schematic diagram of an intelligent event detection and video encoding process in the prior art, provided in Embodiment 1 of the present invention;

[0033] Figure 3 This is a schematic diagram of intelligent event image and target object information transmission based on video encoding, provided in Embodiment 1 of the present invention;

[0034] Figure 4 This is a flowchart of an image storage method provided in Embodiment 2 of the present invention;

[0035] Figure 5 This is a schematic diagram of a target object information storage format provided in Embodiment 2 of the present invention;

[0036] Figure 6 This is a schematic diagram of the structure of an image transmission device provided in Embodiment 3 of the present invention;

[0037] Figure 7 This is a schematic diagram of the structure of an image storage device provided in Embodiment 4 of the present invention;

[0038] Figure 8 This is a schematic diagram of the structure of an image transmission system provided in Embodiment 5 of the present invention;

[0039] Figure 9 This is a schematic diagram of the structure of an intelligent event data management system provided in Embodiment 5 of the present invention;

[0040] Figure 10 This is a schematic diagram of the structure of an electronic device provided in Embodiment 5 of the present invention. Detailed Implementation

[0041] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0042] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices. In the embodiments of this application, certain software, components, models, and other existing industry solutions may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solutions of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0043] The acquisition, transmission, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.

[0044] Example 1

[0045] Figure 1 The flowchart of an image transmission method provided in Embodiment 1 of the present invention is applicable to the situation where, when an intelligent event is triggered, the image triggering the intelligent event and its target object information are sent to an image management device. The method can be executed by an image transmission device, which can be implemented in hardware and / or software. The image transmission device can be configured in a shooting device, especially a network camera (IP camera) or other shooting device with intelligent analysis function, and can be used in conjunction with a backend image management device.

[0046] like Figure 1 As shown, the method includes:

[0047] S110. If it is determined that there is at least one target object in the target frame image, then the target frame image is encoded into a switching prediction frame image.

[0048] The target frame image refers to the video frame image on which intelligent event detection is performed. Intelligent events can be face detection, motor vehicle detection, license plate detection, no mask detection, no work uniform detection, etc. This embodiment does not limit the type of intelligent event or the specific detection method.

[0049] The type of the target object matches the type of the smart event. For example, if the smart event is motion detection, the target object is a moving object; if the smart event is vehicle detection, the target object is a vehicle; and if the smart event is maskless detection, the target object is a face or body without a mask.

[0050] In this embodiment, the presence of at least one target object in the target frame image indicates that the target frame image has triggered a smart event and requires further processing.

[0051] Furthermore, in this embodiment, a target frame image is determined from each captured frame image based on a pre-set target object detection frame rate and encoding frame rate, and target object detection is performed on the target frame image.

[0052] Understandably, the target object detection frame rate (i.e. the frame rate for intelligent event detection) may not be the same as the encoding frame rate. Therefore, target object detection may not process all video frames captured by the camera, but rather select video frames as target frames at certain intervals for target object detection.

[0053] For example, if the target object detection frame rate, which is also the smart event detection frame rate, is 5fps (i.e., 5 frames / s), and the encoding frame rate is 30fps, then every 5 frames, one video frame image can be taken as the target frame image for smart event detection, so as to maintain the synchronization between smart event detection and video encoding.

[0054] In this embodiment, the process of intelligent event detection for the target frame image is similar to that of the prior art, and will not be described in detail here. However, the difference between this embodiment and the prior art lies in that the encoding process of the target frame image is controlled when it is determined that the target frame image triggers an intelligent event.

[0055] In existing technologies, intelligent event detection and video encoding are usually performed relatively independently. Figure 2 A schematic diagram of the intelligent event detection and video encoding process in the prior art is provided, such as... Figure 2As shown, the result of intelligent event detection is the target frame image that triggered the intelligent event and its target object information, while the result of video encoding is a bitstream represented in the form of GOP (Group of Pictures). A frame group is a sequence of frames consisting of one I-frame (Intra-coded Frame) and several P-frames (Predicted-coded Frames). The intra-coded frame image, also known as the keyframe image, can be decoded independently, while the predicted-coded frame image must rely on a reference frame image for decoding. The reference frame image can be either the keyframe image or the previous predicted-coded frame image.

[0056] according to Figure 2 It is known that, using existing technical solutions, when a smart event is triggered, in addition to the normal transmission of the bitstream, the target frame image and target object information that triggered the smart event also need to be sent to the image management device. This data requires at least several hundred KB in size. When smart events are triggered frequently, or when there are many shooting devices connected to the image management device, this data consumes a significant amount of network bandwidth, which in turn consumes bandwidth for bitstream transmission, affecting the transmission performance of both the front-end shooting device and the back-end image management device.

[0057] Switching prediction frame images refer to SP (Switching P-frame) images, which are a special type of frame in the video coding standard. Unlike ordinary P-frame images, switching prediction frame images are key P-frame images.

[0058] The target frame image is encoded into a switching prediction frame image. Specifically, the key frame image corresponding to the target frame image can be determined, the difference information between the target frame image and the key frame image can be determined, the difference information can be quantized and entropy encoded, and the switching prediction frame image can be generated. The specific process can adopt the conventional SP frame encoding method, and this embodiment does not limit it.

[0059] In this embodiment, the target frame image that triggers the intelligent event is encoded as a switching prediction frame image. One reason for this is that the switching prediction frame image does not rely on the decoding result of its previous frame image; it can be directly decoded by referencing the keyframe image, achieving fast decoding. Therefore, even without directly transmitting the target frame image, the target frame image can be decoded by the backend image management device. Another reason is that the size of the encoded switching prediction frame image is smaller than the target frame image. Therefore, in this embodiment, the target frame image is encoded as a switching prediction frame image and sent to the backend image management device along with the bitstream, eliminating the need for additional transmission of the target frame image and reducing bandwidth consumption.

[0060] Furthermore, in the above embodiment, when a smart event is triggered by the target frame image and a target object is detected, the target frame image is encoded as a switching prediction frame image. If it is determined that no target object exists in the target frame image, the target frame image is encoded as a forward prediction frame image or a keyframe image.

[0061] Understandably, the transmission of the target frame image and its target object information is only necessary when the target frame image triggers a smart event. If the target frame image does not contain a target object, i.e., when the target frame image does not trigger a smart event, there is no need to transmit images and information related to the smart event. In this case, the target frame image without a target object can be directly encoded as a forward prediction frame image or a keyframe image. Whether it is specifically encoded as a forward prediction frame image or a keyframe image, and the specific encoding process, can all be implemented using conventional video encoding methods, which will not be elaborated upon in this embodiment.

[0062] Furthermore, in an optional embodiment, after determining that at least one target object exists in the target frame image, the method further includes: determining target object information in the target frame image and encapsulating the target object information into supplementary enhancement information.

[0063] This embodiment provides a method for transmitting target object information when it is necessary to transmit target object information in a target frame image to a backend image management device.

[0064] Target object information refers to the quantity, identification, type, location coordinates, and attribute information of target objects. For example, taking motor vehicle detection, target object information may include the number of motor vehicles, the identification added to each vehicle when there are multiple vehicles, the location coordinates of each vehicle's area, basic vehicle attribute information (model, color, license plate number, etc.), and vehicle status information (parked or moving; in the moving state, speed may also be included). Transmitting this target object information to the backend image management device allows it to quickly locate the target object in the target frame image during further recognition and analysis.

[0065] Supplemental enhancement information (SEI) is used to carry auxiliary information about the bitstream. SEI does not affect the video encoding and decoding process, but it can provide additional enhancements for video display, processing, and interaction.

[0066] The target object information is encapsulated into supplementary enhancement information. Specifically, the target object information is encapsulated according to the format of supplementary enhancement information, and a header for supplementary enhancement information is added. The header may contain a supplementary enhancement information identifier to indicate that this is supplementary enhancement information. The target object information is then filled into the data portion of the supplementary enhancement information according to the format of the supplementary enhancement information. Furthermore, verification information can be added to the end of the supplementary enhancement information to ensure the integrity and accuracy of the target object information. Similarly, the target object information is encapsulated into supplementary enhancement information using conventional supplementary enhancement information encapsulation methods, and this embodiment does not impose any restrictions on this.

[0067] In this embodiment, the target object information is encapsulated as supplementary enhancement information, so that the target object information can be sent along with the bitstream without the need for additional transmission of the target object information, thereby reducing the consumption of transmission bandwidth and improving transmission efficiency.

[0068] Furthermore, encoding the target frame image into a switching prediction frame image and encapsulating the target object information into supplementary enhancement information allows the target frame image and target object information, which were originally sent separately from the bitstream, to be sent together with the bitstream. This not only achieves the aforementioned effects of reducing transmission bandwidth and improving transmission efficiency, but also simplifies the data reception process of the backend image management device and avoids the loss or out-of-order transmission of images and information related to intelligent events.

[0069] Furthermore, in another optional embodiment, when it is determined that a target object exists in the target frame image, the target frame image can be cropped to obtain a target object image, and the target object image can be stitched together with the target frame image before being encoded into a switching prediction frame image.

[0070] The advantage of this setup is that it enables fast decoding and reduces bandwidth usage, allowing the backend image management device to directly obtain the target object image from the decoded image, thereby improving the target object recognition and analysis efficiency of the backend image management device.

[0071] Furthermore, in another optional embodiment, the target frame image can be encoded into a switching prediction frame image, which is then decoded by the backend image management device, and the encoded image is used for target object detection and recognition. The advantage of this setup is that, due to the more powerful processing capabilities of the backend image management device, more accurate target object information can be obtained.

[0072] S120. The switching prediction frame image and other encoded images are encapsulated into a data packet, and the data packet is sent to the image management device to realize the transmission of the target frame image.

[0073] Other encoded images may include keyframe images and forward prediction frame images. In this embodiment, during the encoding stage, the target frame image that triggers the intelligent event is encoded to obtain the switching prediction frame image; the target frame image that does not trigger the intelligent event is encoded to obtain the forward prediction frame image or the keyframe image; other video frame images besides the target frame image are encoded to obtain the forward prediction frame image or the keyframe image. Data packets are segmented and encapsulated sequentially, and the bitstream is transmitted between the shooting device and the image management device in units of data packets.

[0074] Specifically, after encoding each keyframe image, forward prediction frame image, and switching prediction frame image, header information is added. This header information includes frame type identifier, timestamp, and sequence number. Then, based on the transmission protocol between the capturing device and the image management device, such as RTP (Real-Time Transport Protocol), UDP (User Datagram Protocol), or TCP (Transmission Control Protocol), the encoded image and supplementary enhancement information are encapsulated into data packets. For example, using the RTP protocol, the encoded image and supplementary enhancement information are encapsulated into appropriately sized RTP data packets, and RTP header information is added, including the data packet version number, sequence number, and timestamp. The data packets are then transmitted based on the transmission protocol between the capturing device and the image management device.

[0075] In this embodiment, since the target frame image that triggers the smart event is encoded as a switching prediction frame image, the target frame image that triggers the smart event can be transmitted along with the bitstream, reducing the bandwidth usage and improving transmission efficiency.

[0076] Furthermore, encapsulating the switch prediction frame image and other encoded images into a data packet can also include: sequentially encapsulating the switch prediction frame image and other encoded video images, and inserting supplementary enhancement information before the switch prediction frame image to obtain a data packet.

[0077] The supplementary enhancement information can be inserted either before or after its corresponding switching prediction frame image; this embodiment does not impose any restrictions on this.

[0078] This embodiment uses the example of inserting supplementary enhancement information before the switching prediction frame image. The advantage of this setting is that after the backend image management device parses the bitstream and detects the switching prediction frame image, it can determine that the content of the supplementary enhancement information before the switching prediction frame image is the target object information corresponding to the switching prediction frame image, thereby realizing the rapid storage of supplementary enhancement information.

[0079] Taking the acquisition of target object information from the target frame image and its encapsulation as supplementary enhancement information, which is then inserted into the bitstream for joint transmission as an example... Figure 3 A schematic diagram is provided for transmitting intelligent event images and their target object information based on video encoding, such as... Figure 3 As shown, the video encoding process is controlled by the intelligent analysis results of the target frame image. Specifically, for target frame images whose intelligent analysis results trigger an intelligent event, they are encoded as SP frame images, their target object information is encapsulated as SEI, and the SEI is inserted before the SP frame image; for target frame images whose intelligent analysis results do not trigger an intelligent event, they are encoded as P frame images.

[0080] In this embodiment, the GOP may include I-frame images, P-frame images, SP-frame images and their corresponding SEI, thereby realizing the encapsulation of the target frame image that triggers the smart event and its target object information into the bitstream and transmitting it with the bitstream.

[0081] The technical solution of this invention involves acquiring target object information when a target object is present in the target frame image, encapsulating the target frame image into a switching prediction frame image, and then encapsulating the switching prediction frame image together with other encoded images into a data packet, which is then sent to the image management device. This solves the problem in the prior art where, in scenarios with frequent intelligent event triggering or when multiple shooting devices are connected, the transmission of images triggered by intelligent events and target object data consumes a large amount of network bandwidth, leading to a decrease in the video transmission bandwidth of the shooting devices. Encapsulating the target frame image containing the target object into a switching prediction frame image allows for fast decoding by referencing keyframe images, eliminating the need for additional transmission of the target frame image, reducing transmission bandwidth consumption, and ensuring the performance of both the front-end shooting device and the back-end image management device.

[0082] Example 2

[0083] Figure 4The flowchart of an image storage method provided in Embodiment 2 of the present invention is applicable to the case where, upon receiving a data packet, the data packet and the target object information corresponding to the intelligent event triggered in the data packet are stored separately. This method can be executed by an image storage device, which can be implemented in hardware and / or software. The image storage device can be configured in an image management device, such as a server or an NVR (Network Video Recorder), and can be used in conjunction with front-end network cameras or other shooting devices with intelligent analysis functions.

[0084] like Figure 4 As shown, the method includes:

[0085] S210: Receive the data packet sent by the shooting device and determine the switching prediction frame image in the data packet.

[0086] In this embodiment, since the bitstream transmission between the imaging device and the image management device is in units of data packets, the image management device parses the data packets after receiving them to obtain each frame image. The switching prediction frame image can be identified through the frame type identifier in the header information of each frame image.

[0087] Furthermore, this embodiment also provides a processing method when the shooting device encapsulates the target object information in the target frame image as supplementary enhancement information and inserts it into the bitstream for joint transmission. Taking the insertion of supplementary enhancement information before the switching prediction frame image as an example, the supplementary enhancement information before the switching prediction frame image can be used as the supplementary enhancement information corresponding to the switching prediction frame image.

[0088] In this embodiment, since the target object information is encapsulated in the data part of the supplementary enhancement information when the shooting device encapsulates the supplementary enhancement information, the target object information can be obtained by parsing the supplementary enhancement information after it is determined.

[0089] The target object information may include the number of target objects, their identifier, type, location coordinates, and attribute information.

[0090] S220. Store each frame of the image in the data packet.

[0091] In this embodiment, each frame image (i.e., bitstream) in the data packet is stored. It is understood that the bitstream is stored during transmission. If the existing technology were used, where the target frame image triggering the smart event is transmitted separately, then the target frame image would need to be stored separately in the image management device.

[0092] The technical solution of this embodiment has two advantages. First, since the target frame image that triggers the intelligent event is already encapsulated in the bitstream, only the parsed frame images need to be stored; there is no need to store the target frame image separately, thus saving storage space in the image management device. Second, since switching the prediction frame image does not depend on the decoding result of the previous frame and can be directly decoded based on the keyframe image, switching the prediction frame image can achieve fast decoding to obtain the target frame image that triggers the intelligent event. Even without storing the target frame image separately, it can be quickly obtained for subsequent image management and analysis, achieving the same effect as directly transmitting the target frame image.

[0093] S230. Store the storage location of the switching prediction frame image and the storage location of the key frame image that matches the switching prediction frame image, so as to decode the switching prediction frame image according to the storage location.

[0094] The keyframe image that matches the switching prediction frame image refers to the reference frame image required for decoding the switching prediction frame image, which can be the most recent keyframe image before the switching prediction frame image.

[0095] The storage location mentioned in this embodiment refers to the storage location of the switching prediction frame image or the keyframe image matching the switching prediction frame image in the storage system of the image management device. Taking an NVR as an example, an NVR typically stores each frame image in a file system on a hard disk. During storage, files can be created according to certain rules to store each frame image; files can be divided according to date, channel, or other methods. When storing each frame image into a file, the position offset of each frame image can be obtained. Based on the position offset and the file location, the storage location of each frame image can be determined.

[0096] In this embodiment, while storing each frame image in the data packet, the storage locations of the switching prediction frame image and the key frame images required for decoding the switching frame image are also stored. The advantage of this setup is that it allows for rapid location of the switching prediction frame image and its corresponding key frame image, thereby enabling fast decoding to obtain the original target frame image that triggered the intelligent event.

[0097] Furthermore, after determining the handover prediction frame image in the data packet, and further determining and parsing supplementary enhancement information to obtain target object information, this target object information can be stored separately from each frame image. Simultaneously, the storage locations of the handover prediction frame image and the keyframe image matching the handover prediction frame image can be stored correspondingly with the target object information corresponding to the handover prediction frame image.

[0098] Figure 5 A schematic diagram of a target object information storage format is provided, such as... Figure 5 As shown, each frame image and target object information in the data packet are stored separately. Each frame image in the data packet is stored in the video data area, with a Group of Pictures (GOP) as the basic unit. Target object information is stored in the target object information area. In addition to the image sequence number, number of targets, target type, target location, and target attributes corresponding to the target object information, the target object information also includes the storage location of the SP frame image corresponding to the target object information and the storage location of the I frame image corresponding to the SP frame image when storing each frame image.

[0099] Furthermore, the restoration of relevant image and target object information triggered by intelligent events can be achieved through the following steps:

[0100] A1. Determine the storage location of the switching prediction frame image to be processed, and the storage location of the keyframe image that matches the switching prediction frame image to be processed.

[0101] A2. Determine the switching prediction frame image to be processed based on the storage location of the switching prediction frame image to be processed, and determine the key frame image based on the storage location of the key frame image that matches the switching prediction frame image.

[0102] A3. Decode the switch prediction frame image to be processed based on the keyframe image to obtain the frame image to be processed.

[0103] This embodiment also provides a method for retrieving images triggered by intelligent events based on the images of each frame stored in the image management device and the target object information.

[0104] The switching prediction frame images to be processed can be any switching prediction frame images required for intelligent event analysis. In this embodiment, when storing each frame image, each frame image can be stored separately according to different intelligent event types, different time periods, etc. When further analysis of the relevant intelligent images of the intelligent event to be processed is required, the storage location of the switching prediction frame images to be processed in the storage area corresponding to the intelligent event to be processed is determined.

[0105] Based on the storage location of the switch prediction frame image to be processed, read the switch prediction frame image to be processed. Based on the storage location of the keyframe image, read the keyframe image. Decode the switch prediction frame image to be processed based on the keyframe image to obtain the frame image to be processed.

[0106] In this embodiment, when storing the target object information of the switching prediction frame image separately, the switching prediction frame image to be processed can also be determined based on the target object information to be processed. The target object information to be processed is the target object information corresponding to this intelligent image and target object information retrieval task. After determining the corresponding frame image to be processed that triggers the intelligent event based on the target object information to be processed, subsequent image management and analysis are performed based on the frame image to be processed and its corresponding target object information to be processed.

[0107] Understandably, image management devices store a large amount of intelligent events of different times and types, or target object information from different shooting devices. The required target object information can be retrieved from these target object information sources based on search criteria. Search criteria can consist of date, intelligent event type, shooting device identifier, etc., and this embodiment does not impose any limitations on these criteria.

[0108] After identifying the target object information to be processed from the target object information, the storage locations of the handover prediction frame image to be processed and its corresponding keyframe image can be obtained from the target object information. Based on the storage location of the handover prediction frame image to be processed, the handover prediction frame image to be processed is read; based on the storage location of the keyframe image, the keyframe image is read. The handover prediction frame image to be processed is then decoded based on the keyframe image to obtain the frame image to be processed.

[0109] Specifically, the header information of the switch prediction frame image to be processed is parsed. Based on the motion vector information of the header of the switch prediction frame image to be processed, motion estimation and motion compensation are performed with the keyframe image as a reference to obtain the predicted values ​​of the corresponding pixel blocks of the motion vector in the switch prediction frame image to be processed. Based on the residual information of the header of the switch prediction frame image to be processed, the predicted values ​​are added to the residual pixel values ​​to reconstruct the pixels of the switch prediction frame image to be processed. The process of decoding the switch prediction frame image to be processed based on the keyframe image is implemented using a conventional decoding process, which will not be elaborated on in this embodiment.

[0110] In this embodiment, after obtaining the frame image to be processed, since the target object information to be processed includes the number and location of the target objects, the frame image to be processed is cropped according to the number and location of the target objects in the target object information to be processed, so that a corresponding number of target object images can be obtained.

[0111] The frame image to be processed, the target object image, and the target object information together constitute intelligent event-related data, which is used for subsequent image management and intelligent event analysis.

[0112] In this embodiment, the original target frame image that triggers the smart event and the target object information are encoded and encapsulated and sent with the data packet. The image management device only needs to receive the data packet sent by the shooting device to achieve normal reception of the bitstream and the smart event data, simplifying the data reception process of the image management device and avoiding the loss or out-of-order transmission of smart event data. The image management device can store only the bitstream and the target object information, without storing the original image that triggers the smart event, thus saving storage space. By storing the storage location of the switching prediction frame image and its corresponding keyframe image in the target object information when storing the bitstream, the image reading during subsequent retrieval can be performed based on the storage location of the switching prediction frame image and its corresponding keyframe image, quickly decoding the image that triggers the smart event and achieving the same effect as transmitting the image that triggers the smart event separately.

[0113] The technical solution of this embodiment receives data packets sent by the shooting device, parses the data packets to obtain the switching prediction frame image and its corresponding supplementary enhancement information, determines the target object information corresponding to the switching prediction frame image based on the supplementary enhancement information, and stores each frame image in the data packet and the target object information corresponding to the switching prediction frame image separately. This solves the problem in the prior art where the storage of images triggered by intelligent events and target object data occupies a large amount of memory in scenarios with frequent intelligent event triggering or when a large number of shooting devices are connected. Since the switching prediction frame image can be decoded with reference to the keyframe image to obtain the original intelligent event trigger frame image corresponding to the switching prediction frame image, it is not necessary to store the intelligent event trigger image separately, thus reducing storage capacity.

[0114] Example 3

[0115] Figure 6 This is a schematic diagram of an image transmission device according to Embodiment 3 of the present invention. The device is deployed in a shooting device, such as... Figure 6 As shown, the device includes:

[0116] The target frame image encoding module 310 is used to encode the target frame image into a switching prediction frame image if it is determined that there is at least one target object in the target frame image;

[0117] The data packet sending module 320 is used to encapsulate the switching prediction frame image and other encoded images into data packets and send the data packets to the image management device to realize the transmission of the target frame image.

[0118] The technical solution of this invention encapsulates the target frame image into a switching prediction frame image when a target object exists in the target frame image. This switching prediction frame image is then encapsulated together with other encoded images into a data packet, which is then sent to the image management device. This solves the problem in existing technologies where, in scenarios with frequent intelligent event triggering or when multiple shooting devices are connected, the transmission of images triggered by intelligent events and target object data consumes significant network bandwidth, leading to a decrease in the video transmission bandwidth of the shooting devices. Encapsulating the target frame image containing the target object into a switching prediction frame image allows for fast decoding by referencing keyframe images. Therefore, it eliminates the need for additional transmission of the target frame image, reducing bandwidth consumption and ensuring the performance of both the front-end shooting device and the back-end image management device.

[0119] Optionally, based on the above embodiments, the apparatus further includes:

[0120] The supplementary enhancement information encapsulation module is used to determine the target object information in the target frame image and encapsulate the target object information into supplementary enhancement information;

[0121] The data packet sending module 320 includes:

[0122] The data packet encapsulation unit is used to encapsulate the switching prediction frame image and other video encoded images sequentially, and insert supplementary enhancement information before the switching prediction frame image to obtain a data packet.

[0123] Optionally, based on the above embodiments, the apparatus further includes:

[0124] The target object detection module is used to determine the target frame image in each captured frame image according to the preset target object detection frame rate and encoding frame rate, and to perform target object detection on the target frame image;

[0125] The target frame image encoding module is used to encode the target frame image into a forward prediction frame image or a key frame image if it is determined that there is no target object in the target frame image.

[0126] The image transmission device provided in the embodiments of the present invention can execute the image transmission method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0127] Example 4

[0128] Figure 7 This is a schematic diagram of an image storage device according to Embodiment 4 of the present invention. This device is deployed in an image management device, such as... Figure 7 As shown, the device includes:

[0129] The switching prediction frame image determination module 410 is used to receive data packets sent by the shooting device and determine the switching prediction frame images in the data packets.

[0130] The data storage module 420 is used to store the images of each frame in the data packet;

[0131] The storage location storage module 430 is used to store the storage location of the switching prediction frame image and the storage location of the key frame image that matches the switching prediction frame image, so as to decode the switching prediction frame image according to the storage location.

[0132] The technical solution of this embodiment receives data packets sent by the shooting device, parses the data packets to obtain switching prediction frame images, and stores each frame image in the data packet. This solves the problem in the prior art where the storage of images triggered by intelligent events and target object data occupies a large amount of memory in scenarios with frequent intelligent event triggering or when a large number of shooting devices are connected. Since the switching prediction frame image can be decoded with reference to the keyframe image to obtain the original intelligent event trigger frame image corresponding to the switching prediction frame image, it is not necessary to store the intelligent event trigger image separately, thus reducing storage capacity.

[0133] Optionally, based on the above embodiments, the apparatus further includes:

[0134] The target object information determination module is used to determine the supplementary enhancement information corresponding to the switching prediction frame image, and determine the target object information corresponding to the switching prediction frame image based on the supplementary enhancement information.

[0135] Storage location storage module 430 also includes:

[0136] The data storage unit is used to store the storage location of the switching prediction frame image, the storage location of the key frame image that matches the switching prediction frame image, and the target object information corresponding to the switching prediction frame image.

[0137] Optionally, based on the above embodiments, the apparatus further includes:

[0138] The storage location determination module is used to determine the storage location of the switching prediction frame image to be processed, and the storage location of the key frame image that matches the switching prediction frame image to be processed.

[0139] The module for determining the switching prediction frame image to be processed is used to determine the switching prediction frame image to be processed based on the storage location of the switching prediction frame image to be processed, and to determine the key frame image based on the storage location of the key frame image that matches the switching prediction frame image to be processed.

[0140] The frame image determination module is used to decode the switch prediction frame image to be processed based on the key frame image to obtain the frame image to be processed.

[0141] The image storage device provided in the embodiments of the present invention can execute the image storage method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0142] Example 5

[0143] Figure 8 This embodiment provides a schematic diagram of the structure of an image transmission system, such as... Figure 8 As shown, the image transmission system includes a shooting device and an image management device.

[0144] The shooting device is used to encode a target frame image containing at least one target object into a switching prediction frame image;

[0145] The switched prediction frame image and other encoded images are encapsulated into data packets, which are then sent to the image management device.

[0146] Other encoded images include keyframe images and forward prediction frame images;

[0147] The image management device is used to receive data packets sent by the shooting device and determine the switching prediction frame image in the data packets;

[0148] Store each frame image in the data packet, and store the storage location of the switching prediction frame image and the storage location of the key frame image that matches the switching prediction frame image.

[0149] The image management device is also used to decode based on the storage location of the switching prediction frame image and the storage location of the key frame image that matches the switching prediction frame image.

[0150] In a specific applicable scenario Figure 9 This embodiment provides a structural diagram of an intelligent event data management system, such as... Figure 9 As shown, the intelligent event data management system includes front-end shooting equipment and back-end image management equipment.

[0151] The shooting device includes an intelligent analysis module and a video encoding module. The intelligent analysis module performs intelligent event detection on the target frame image. If it is determined that the target frame image triggers an intelligent event, it determines its target object information and controls the video encoding module to encode the target frame image into a switching prediction frame image, encapsulating the target object information into supplementary enhancement information. If it is determined that the target frame image does not trigger an intelligent event, it controls the video encoding module to encode the target frame image into a forward prediction frame image or a keyframe image.

[0152] When the intelligent analysis frame rate is inconsistent with the video encoding frame rate, the video encoding module performs normal encoding on video frame images other than the target frame image for intelligent analysis.

[0153] The imaging device encapsulates each encoded frame image into a data packet, in which supplementary enhancement information is inserted before the corresponding switching prediction frame image, and then transmits the encapsulated data packets to the image management device.

[0154] The image management device's storage area includes a video data area and a target object information area. Data packets are parsed, and each frame is stored normally in the video data area. Based on the switching prediction frame image in the data packet, its corresponding supplementary enhancement information is determined, and the target object information from the supplementary enhancement information is stored in the target object information area. Simultaneously, when storing each frame image in the video data area, the storage location of the switching prediction frame image and the storage location of the keyframe images required for decoding are added to the target object information and stored together.

[0155] When intelligent data retrieval is required, the image management device determines the target object information to be processed from the stored target object information based on the intelligent data retrieval conditions. According to the storage location of the switching prediction frame image and the storage location of the keyframe image required for decoding recorded in the target object information, the switching prediction frame image and the keyframe image are read respectively, and the switching prediction frame image is decoded to obtain the frame image to be processed. Then, based on the number and location of targets in the target object, the target object image is obtained from the frame image to be processed. The frame image to be processed, the target object image, and the target object information are used as intelligent event-related data.

[0156] The intelligent event data management system, through the cooperation between shooting equipment and image management equipment, can reduce the transmission bandwidth and storage capacity of intelligent event-related data and improve the transmission efficiency of the bitstream.

[0157] Figure 10 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0158] like Figure 10As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0159] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0160] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as image transmission methods or image storage methods.

[0161] In some embodiments, the image transmission method or image storage method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the image transmission method or image storage method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the image transmission method or image storage method by any other suitable means (e.g., by means of firmware).

[0162] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0163] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0164] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0165] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0166] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0167] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0168] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0169] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An image transmission method, characterized in that, The method is performed by a shooting device, and the method includes: If it is determined that there is at least one target object in the target frame image, then the target frame image is encoded as a switching prediction frame image; The switching prediction frame image and other encoded images are encapsulated into data packets, and the data packets are sent to the image management device to achieve the transmission of the target frame image.

2. The method according to claim 1, characterized in that, After determining that at least one target object exists in the target frame image, the process also includes: Determine the target object information in the target frame image, and encapsulate the target object information into supplementary enhancement information; The switched prediction frame image and other encoded images are encapsulated into data packets, including: The switching prediction frame image is sequentially encapsulated with other video encoded images, and supplementary enhancement information is inserted before the switching prediction frame image to obtain a data packet.

3. The method according to claim 1, characterized in that, The method further includes: Based on the pre-set target object detection frame rate and encoding frame rate, the target frame image is determined in each captured frame image, and target object detection is performed on the target frame image; If it is determined that the target object does not exist in the target frame image, the target frame image is encoded as a forward prediction frame image or a keyframe image.

4. An image storage method, characterized in that, The method is performed by an image management device, and the method includes: Receive data packets sent by the imaging device and determine the switching prediction frame image in the data packets; Store each frame of the image in the data packet; The storage locations of the switch prediction frame image and the keyframe image that matches the switch prediction frame image are stored so that the switch prediction frame image can be decoded based on the storage locations.

5. The method according to claim 4, characterized in that, After determining the handover prediction frame image in the data packet, the following is also included: Determine the supplementary enhancement information corresponding to the switching prediction frame image, and determine the target object information corresponding to the switching prediction frame image based on the supplementary enhancement information; The storage locations for the switched prediction frame image and the keyframe image matching the switched prediction frame image are stored, including: The storage locations of the switching prediction frame image and the keyframe image that matches the switching prediction frame image are stored together with the target object information corresponding to the switching prediction frame image.

6. The method according to claim 4, characterized in that, After storing the storage locations of the switching prediction frame image and the keyframe image matching the switching prediction frame image, the process also includes: Determine the storage location of the switching prediction frame image to be processed, and the storage location of the keyframe image that matches the switching prediction frame image to be processed; Based on the storage location of the switching prediction frame image to be processed, determine the switching prediction frame image to be processed, and based on the storage location of the keyframe image that matches the switching prediction frame image to be processed, determine the keyframe image. The keyframe image is used to decode the switching prediction frame image to be processed, thus obtaining the frame image to be processed.

7. An image transmission device, characterized in that, The device is deployed on the imaging equipment, and the device includes: The target frame image encoding module is used to encode the target frame image into a switching prediction frame image if it is determined that there is at least one target object in the target frame image; The data packet sending module is used to encapsulate the switching prediction frame image and other encoded images into data packets and send the data packets to the image management device to realize the transmission of the target frame image.

8. An image storage device, characterized in that, The device is deployed in an image management device, and the device includes: The switching prediction frame image determination module is used to receive data packets sent by the shooting device and determine the switching prediction frame images in the data packets; The data storage module is used to store the images of each frame in the data packet; The storage location module is used to store the storage location of the switching prediction frame image and the storage location of the key frame image that matches the switching prediction frame image, so as to decode the switching prediction frame image according to the storage location.

9. An image transmission system, characterized in that, The system includes a shooting device and an image management device; The shooting device is used to encode a target frame image containing at least one target object into a switching prediction frame image; The switched prediction frame image and other encoded images are encapsulated into data packets, which are then sent to the image management device. Other encoded images include keyframe images and forward prediction frame images; The image management device is used to receive data packets sent by the shooting device and determine the switching prediction frame image in the data packets; Store each frame image in the data packet, and store the storage location of the switching prediction frame image and the storage location of the key frame image that matches the switching prediction frame image. The image management device is also used to decode based on the storage location of the switching prediction frame image and the storage location of the key frame image that matches the switching prediction frame image.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the image transmission method as described in any one of claims 1-3, or the image storage method as described in any one of claims 4-6.

11. A storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the image transmission method as described in any one of claims 1-3, or to perform the image storage method as described in any one of claims 4-6.