Picture capturing method and device, electronic equipment, storage medium and program

By identifying and processing keyframe packets in the video stream, and directly generating target images, the problem of complex configuration of existing tools is solved and an efficient image capture process is achieved.

CN120238701APending Publication Date: 2025-07-01HONGHU WANLIAN (JIANGSU) TECH DEV CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510299809.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing real-time video streaming screenshot tools have many functions, are inconvenient to use and complex configuration, making it difficult to efficiently capture images.

Method used

By obtaining multiple data packets of the target video stream, identifying and obtaining the starting data packets of the keyframe and their subsequent data packets, removing the protocol header and splicing the payload data, decoding and format conversion, and generating the target picture.

Benefits of technology

It realizes simple and efficient capture of pictures from video streams, avoids complex configurations and real-time transmission delays, and is suitable for scenarios such as video surveillance, video conferencing, and online live broadcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238701A_ABST
    Figure CN120238701A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a picture capturing method and device, electronic equipment, a storage medium and a program, and the method comprises the steps: obtaining a plurality of data packets of a target video stream; under the condition that the acquired target data packet is determined to be the initial data packet of the key frame, acquiring a subsequent data packet of the key frame; acquiring original data of the key frame according to the initial data packet of the key frame and the subsequent data packet of the key frame; and converting the original data of the key frame into a target picture. According to the technical scheme provided by the embodiment of the invention, the picture can be simply and efficiently captured from the video stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of image processing, and in particular, to a method, apparatus, electronic device, storage medium, and program for capturing pictures. Background Art

[0002] With the rapid development of video technology, real-time video streams have been widely used in many fields such as video conferencing, security monitoring, and online live streaming. In these application scenarios, in addition to requiring the video stream to be transmitted stably and efficiently, it is often also required to capture screenshots of the video stream. However, currently, capturing screenshots of real-time video streams mostly relies on professional screenshot tools. These screenshot tools have a variety of functions, are inconvenient to use, and require users to configure them in advance before they can be used. Summary of the Invention

[0003] The embodiments of the present invention provide a method, apparatus, electronic device, storage medium, and program for capturing pictures, which can simply and efficiently capture pictures from a video stream.

[0004] According to one aspect of the present invention, there is provided a method for capturing pictures, including:

[0005] Obtaining a plurality of data packets of a target video stream;

[0006] When it is determined that the obtained target data packet is the starting data packet of a key frame, obtaining subsequent data packets of the key frame;

[0007] Obtaining the original data of the key frame according to the starting data packet of the key frame and the subsequent data packets of the key frame;

[0008] Converting the original data of the key frame into a target picture.

[0009] According to another aspect of the present invention, there is provided a picture capturing apparatus, including:

[0010] A data packet obtaining module, configured to obtain a plurality of data packets of a target video stream;

[0011] A key frame obtaining module, configured to obtain subsequent data packets of the key frame when it is determined that the obtained target data packet is the starting data packet of the key frame;

[0012] An original data obtaining module of the key frame, configured to obtain the original data of the key frame according to the starting data packet of the key frame and the subsequent data packets of the key frame;

[0013] A target picture obtaining module, configured to convert the original data of the key frame into a target picture.

[0014] According to another aspect of the present invention, there is provided an electronic device, where the electronic device includes:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the picture grabbing method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the picture grabbing method according to any embodiment of the present invention when executed.

[0019] According to another aspect of the present invention, there is also provided a computer program product including a computer program which implements the picture grabbing method according to any embodiment of the present invention when executed by a processor.

[0020] In the embodiment of the present invention, by obtaining a plurality of data packets of a target video stream and obtaining subsequent data packets of a key frame when it is determined that the obtained target data packet is the starting data packet of the key frame. Further, the original data of the key frame can be obtained according to the starting data packet of the key frame and the subsequent data packets of the key frame, so that the original data of the key frame can be converted into a target picture, solving the problems that existing screenshot tools are complex to use and cumbersome to configure, and being able to simply and efficiently grab pictures from a video stream.

[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0023] Figure 1 is a flowchart of a picture grabbing method provided in Embodiment 1 of the present invention;

[0024] Figure 2 is a schematic structural diagram of the starting data packet of a key frame provided in Embodiment 1 of the present invention;

[0025] Figure 3 It is a schematic structural diagram of subsequent data packets of a key frame provided in the first embodiment of the present invention;

[0026] Figure 4 It is a flowchart of a picture grabbing method provided in the second embodiment of the present invention;

[0027] Figure 5 It is a schematic diagram of the original data of a key frame provided in the second embodiment of the present invention;

[0028] Figure 6 It is a schematic diagram of a picture grabbing process provided in the second embodiment of the present invention;

[0029] Figure 7 It is a flowchart of a target picture generation method provided in the second embodiment of the present invention;

[0030] Figure 8 It is a flowchart of a target picture saving method provided in the second embodiment of the present invention;

[0031] Figure 9 It is a schematic diagram of a picture grabbing device provided in the third embodiment of the present invention;

[0032] Figure 10 It is a schematic structural diagram of an electronic device provided in the fourth embodiment of the present invention. Detailed implementation manners

[0033] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0034] It should be noted that the terms "target", "original", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0035] Embodiment 1

[0036] Figure 1 This is a flowchart of a method for capturing images provided by the first embodiment of the present invention. This embodiment is applicable to the case where images are captured based on key frames in a video stream. The method can be executed by an image capture device, which can be implemented by software and / or hardware and can generally be integrated in an electronic device. The electronic device can be a terminal device or a server device. As long as the image capture method can be executed, the embodiment of the present invention does not limit the specific device type of the electronic device. Figure 1 As shown, the method includes the following operations:

[0037] S110: Acquire multiple data packets of a target video stream.

[0038] The target video stream may be a video to be captured. For example, the target video stream may include, but is not limited to, indoor surveillance video, traffic surveillance video, and sports event video, etc., as long as it is a video with a need for image capture in a related scene. The embodiment of the present invention does not limit the specific type of the target video stream. The data packet may be a basic unit for encapsulating and transmitting real-time data when the target video stream is transmitted.

[0039] In the embodiment of the present invention, when it is necessary to capture images of a video in a certain scene, the video can be used as a target video stream, and multiple data packets of the target video stream can be obtained to capture images according to the data packets.

[0040] In a specific example, assuming that the target video stream is a traffic surveillance video, key information in the traffic surveillance video can be obtained by capturing video frames at a fixed time. For example, a picture of the traffic surveillance video can be captured every 5 seconds, and multiple data packets of the current video frame in the traffic surveillance video can be extracted, so that the license plate information of the illegal vehicle can be recorded in time, which is convenient for the traffic management department to enforce the law and manage.

[0041] S120: When it is determined that the acquired target data packet is a start data packet of a key frame, acquire subsequent data packets of the key frame.

[0042] It is understandable that the target video stream is usually composed of a series of video frames, and the video frames usually include but are not limited to key frames, P frames (Predictive-coded Frames) and B frames (Bidirectional Predictive-coded Frames).

[0043] Among them, the key frame is also called the I frame (Intra-coded Frame), and the I frame uses the intra-frame prediction coding method, encoding only with the data within this frame, so it includes complete picture information. Therefore, when decoding according to the I frame, a complete image can be decoded independently without referring to other frames. The P frame uses the inter-frame prediction coding method, predicting and encoding using the previous I frame or P frame, and only storing the difference information from the reference frame. Therefore, when decoding according to the P frame, the previous I frame or P frame needs to be referred to in order to decode a complete image. The B frame also uses the inter-frame prediction coding method, but it needs to refer to the previous frame and the next frame for prediction coding, and stores the difference data between the two frames. Therefore, when decoding according to the B frame, the previous frame and the next frame need to be referred to in order to decode a complete image. Thus, it can be seen that if you want to capture a complete picture from the target video stream, the I frames in the target video stream can be identified and the I frames can be converted into pictures.

[0044] Among them, the target data packet can be the starting data packet of the key frame. The starting data packet of the key frame can be the first data packet encapsulating the key frame data in the target video stream. The subsequent data packets of the key frame can be other data packets in the target video stream after the starting data packet encapsulating the key frame data. The subsequent data packets of the key frame are usually used to transmit the remaining part of the key frame except the starting data packet. It can be understood that the subsequent data packets of the key frame can include at least one data packet.

[0045] Correspondingly, after obtaining multiple data packets of the current frame in the target video stream, the multiple data packets can be detected to determine which data packets belong to the key frame. Specifically, first, the starting data packet of the key frame can be determined and used as the target data packet. For example, it can be detected through features such as specific protocol header identifiers and timestamp rules. After determining the starting data packet of the key frame, the data packets subsequent to the starting data packet can be saved in sequence until a data packet including a frame end flag is detected, so as to efficiently and accurately obtain the complete key frame data and provide a basis for subsequent processing. The so-called frame end flag can be a specific flag or sequence used to identify the end of a video frame. Exemplarily, the end of a video frame can be identified by a frame check sequence.

[0046] Figure 2 It is a schematic structural diagram of the starting data packet of a key frame provided in Embodiment 1 of the present invention. Figure 3 It is a schematic structural diagram of the subsequent data packet of a key frame provided in Embodiment 1 of the present invention. In a specific example, such as Figure 2 and Figure 3As shown in the figure, assume that the encoding format of the target video stream is H.264 and it is transmitted using RTP (Real-time Transport Protocol). The starting data packet of a key frame may include, but is not limited to, RTP header (Real-time Transport Protocol header), PSH (Program Stream) header, SYS (System) header, PSM (Program Stream Map) header, PES (Packetized Elementary Stream) header, SPS NALU (Sequence Parameter Set Network Abstraction Layer Unit), PPS NALU (Picture Parameter Set Network Abstraction Layer Unit), and H.264 Instantaneous Decoding Refresh Frame (IDR frame), etc. The subsequent data packets of the key frame only include the RTP header and the remaining part of the H.264 IDR frame. Therefore, the data packet including the combined protocol header of PSH+SYS+PSM can be used as the starting data packet of the key frame.

[0047] It can be understood that the sequence numbers of the subsequent data packets of the key frame are consecutive with those of the starting data packet of the key frame, and the subsequent data packets of the key frame have the same timestamp information as the starting data packet of the key frame. Therefore, the subsequent data packets of the key frame can be determined according to the sequence numbers and timestamp information of the data packets.

[0048] S130. Obtain the original data of the key frame according to the starting data packet of the key frame and the subsequent data packets of the key frame.

[0049] Among them, the original data of the key frame can be the complete key frame data that has been encoded but not encapsulated at all during the encoding process of the target video stream.

[0050] Correspondingly, after determining the starting data packet of the key frame and the subsequent data packets of the key frame, the starting data packet and the subsequent data packets can be spliced together according to the sequence numbers of these data packets, and it is ensured that the data of the starting data packet is in the front and the data of the subsequent data packets follows in sequence. Optionally, after the merging is completed, the integrity of the original data of the key frame can be verified by checking the frame end flag and the data length.

[0051] In a specific example, assuming that the encoding format of the target video stream is H.264, during the process of determining the original data of the key frame, it is necessary to accurately splice the macroblock data and motion vector data in each data packet according to its encoding specification, so as to completely restore the original data of the key frame. These original data contain the complete image information of the key frame, which is the core basis for subsequent operations such as image decoding, image analysis, and image conversion. At the same time, it also provides the necessary data support for functions such as extracting the target image from the target video stream and understanding the video content.

[0052] S140. Convert the original data of the key frame into a target picture.

[0053] Among them, the target picture can be a picture obtained by converting the original data of the key frame.

[0054] Correspondingly, after determining the original data of the key frame, these data can be processed to realize the visualization of the key frame. Specifically, the original data of the key frame can be decoded and converted, and its data format after encoding can be converted into a displayable image format, such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), or BMP (Bitmap) format, etc. Through the above method, a target picture can be generated. This process not only realizes the effective utilization of the key frame data, but also provides a convenient picture capture function for application scenarios such as video surveillance, video conferencing, and online live broadcast, facilitating users to obtain and save important video moments at any time.

[0055] It can be seen that the above solution identifies the key frame in the target video stream and converts the original data of the key frame into a target picture. The whole process is simple to operate, without complex configuration and manual intervention, and can simply and efficiently capture pictures from the real-time video stream. In addition, during the picture capture process, it will not cause an additional burden on the real-time transmission of the video stream, nor will it introduce significant delay or performance loss, ensuring the smoothness and real-time nature of the target video stream, and providing strong and reliable support for various application scenarios that rely on the real-time video stream.

[0056] In the embodiment of the present invention, multiple data packets of the target video stream are obtained, and when it is determined that the obtained target data packet is the starting data packet of the key frame, subsequent data packets of the key frame are obtained. Further, the original data of the key frame can be obtained according to the starting data packet of the key frame and the subsequent data packets of the key frame, so that the original data of the key frame can be converted into a target picture, solving the problem that the existing picture capture tools are complex to use and cumbersome to configure, and being able to simply and efficiently capture pictures from the video stream.

[0057] Embodiment 2

[0058] Figure 4 is a flowchart of a picture grabbing method provided by Embodiment 2 of the present invention. This embodiment is a specific implementation based on the above embodiment. In this embodiment, subsequent data packets of the key frame are obtained, and various specific and optional implementation manners for obtaining the original data of the key frame based on the start data packet of the key frame and the subsequent data packets of the key frame and converting the original data of the key frame into a target picture are given. Correspondingly, as Figure 4 shown, the method of this embodiment may include:

[0059] S210. Obtain multiple data packets of the target video stream.

[0060] S220. When it is determined that the obtained target data packet is the start data packet of the key frame, obtain the subsequent data packets of the key frame.

[0061] S230. Remove the protocol headers of the start data packet of the key frame and the subsequent data packets of the key frame, and obtain the payload data of the key frame.

[0062] Among them, the protocol header may be fixed-format information attached to the front of the data packet, which can be used to describe the attributes and transmission requirements of the data packet. The payload data of the key frame may be the actual video data part encapsulated in the data packet.

[0063] Correspondingly, after determining the start data packet of the key frame and the subsequent data packets of the key frame, in order to obtain the payload data of the key frame, redundant information related to data transmission control in the start data packet and the subsequent data packets, that is, the protocol header, can be removed. It can be understood that the protocol headers and payload data of the start data packet and the subsequent data packets can be separated according to the protocol header format specifications defined by different protocols.

[0064] In a specific example, assume that the target video stream is transmitted using RTP. The protocol header of the RTP data packet generally starts from the start position of the RTP data packet and has a length of 12 bytes. Therefore, the protocol header of the RTP data packet can be extracted through slicing operations. After removing the protocol header, the remaining part is the payload data of the RTP data packet.

[0065] S240. Combine the payload data of the key frame to obtain the original data of the key frame.

[0066] Correspondingly, after obtaining the payload data of the key frame, based on the sequence identification information carried during the transmission of the start data packet and the subsequent data packets, such as the continuity of the data packet sequence number and the sameness of the timestamp, etc., the payload data of the key frame can be accurately combined to obtain the original data of the key frame.

[0067] Figure 5 It is a schematic diagram of the original data of a key frame provided in the second embodiment of the present invention. Continuing with the above example for illustration, as Figure 5 shown, the protocol header is only used to distinguish frame types and save frame sequence numbers during data transmission. The real original data of the key frame does not include the protocol header. Therefore, after saving a set of complete data of the key frame, all protocol header parts can be removed first. The remaining data after removing all protocol headers is the payload data of the key frame. Further, the payload data of the key frame can be sorted to obtain the original data of the key frame. Exemplarily, the original data of the key frame may include, but is not limited to, SPS NALU, PPS NALU, H.264 IDR frame, and the remaining part of the H.264 IDR frame, etc.

[0068] S250. Decode the original data of the key frame to obtain the original image data.

[0069] Among them, decoding can be a process of restoring the compressed original data of the key frame to the original image data. The original image data can be the image data obtained after decoding the original data of the key frame.

[0070] Correspondingly, different formats of target video streams and data of different frame types correspond to different decoders. Therefore, after determining the original data of the key frame, a corresponding decoder can be matched for the original data of the key frame. Further, the original data of the key frame can be decoded by the decoder to obtain the original image data.

[0071] In an optional embodiment of the present invention, the decoding the original data of the key frame to obtain the original image data may include: performing an inverse transformation operation on the original data of the key frame to determine the original image pixel values; combining the original image pixel values to obtain the original image data.

[0072] Among them, the inverse transformation operation can be a process of restoring data or signals after a certain transformation to their original form. The original image pixel values can be the actual values of each pixel point in the original image, and these values can reflect information such as the brightness and color of the original image.

[0073] In an embodiment of the present invention, in the process of decoding the original data of a key frame to obtain the original image, first, a corresponding decoder can be matched for the original data of the key frame. For example, when the target video stream is encoded in the H.264 format, an H.264 decoder can be selected. During the encoding process, in order to reduce the data volume and improve the transmission efficiency, various transformations are usually performed on the original data of the key frame to convert the original data of the key frame from the spatial domain to the frequency domain. Therefore, in the decoding stage, the original data of the key frame that has been transformed needs to be converted back to the spatial domain to restore the pixel values of the original image. The decoder can perform an inverse operation on the original data of the key frame according to the transformation algorithm and parameters used during encoding. For example, when processing the original data of a key frame that has undergone a discrete cosine transform, the decoder can use the inverse discrete cosine transform to convert the coefficients in the frequency domain into pixel values in the spatial domain. After decoding is completed, the decoder can combine the restored pixel values of the original image to obtain the complete original image data. The inverse transformation operation can process a large amount of data in a short time, ensuring that the image can be decoded and presented in a timely and accurate manner. At the same time, the inverse transformation operation can effectively retain the details and features of the image, avoiding a decrease in image quality due to decoding, thereby providing users with a high-quality visual experience.

[0074] In a specific example, Tool A with data format conversion capabilities can provide a set of decoding interfaces. The interfaces of Tool A can be called in sequence, and parameters can be passed into the original data of the key frame to obtain the original image data. Specifically, first, avcodec_find_decoder(AVCodecID) can be called to find and obtain a specific decoder. The parameter ID (Identifier) is an enumerated type value representing the identifier of the decoder to be found. For example, if the original data of the key frame uses the H.264 encoding format, the decoder can be an H.264 decoder, and the parameter ID is AV_CODEC_ID_H264. Further, avcodec_alloc_context3() can be called to allocate a buffer for the decoder and provide the context for the decoder to run. Further, avcodec_open2() can be called to initialize the decoding buffer, open the associated decoder, and set all the internal structures and parameters required by the decoder. Finally, avcodec_decode_video2() can be called to receive the original data of the key frame and decode it into uncompressed original image data.

[0075] Optionally, to further improve the quality of the original image data, operations such as deblocking, color space conversion, sharpening, and noise reduction can also be performed on the decoded image data. Through the above operations, not only can the correct display of the decoded image be ensured, but also better visual effects can be provided to meet the requirements of different application scenarios.

[0076] S260. Encode the original image data to obtain the target picture.

[0077] Specifically, after obtaining the original image data, a suitable picture encoding format can be selected according to the application scenario and requirements to encode the original image data to obtain the target picture.

[0078] Figure 6 is a schematic diagram of a picture grabbing process provided in the second embodiment of the present invention. In a specific example, as Figure 6 shown, to grab a picture from the target video stream, first, multiple data packets can be intercepted from the target video stream. Further, the multiple data packets can be detected to find the starting data packet of the key frame. If the data packet is not the starting data packet of the key frame, the data packet is discarded; if the data packet is the starting data packet of the key frame, the subsequent data packets of the data packet can be further obtained as the subsequent data packets of the key frame. Further, the protocol headers of the starting data packet and the subsequent data packets can be removed, and the starting data packet and the subsequent data packets after removing the protocol headers can be combined to obtain the original data of the key frame. Further, after obtaining the original data of the key frame, a tool with data format conversion ability can be used to convert the original data of the key frame into the target picture.

[0079] In an optional embodiment of the present invention, the encoding the original image data to obtain the target picture and saving may include: performing color space conversion on the original image data to obtain converted image data; performing intra-frame compression on the converted image data to obtain the target picture.

[0080] Among them, color space conversion can be a process of converting image or video data from one color space representation form to another color space representation form. Common color spaces can include but are not limited to RGB (Red-Green-Blue), YUV (Luminance-Chrominance), and HSV (Hue-Saturation-Value), etc. Intra-frame compression can be a video compression technology that only uses the spatial redundancy within a single video frame for compression without relying on the information of other frames.

[0081] In an embodiment of the present invention, in the process of encoding the original image data to obtain the target picture, first, factors such as the specific application scenario, device compatibility, compression efficiency, and quality requirements of the target picture can be comprehensively considered to match the corresponding encoder for the original image data. Further, the color space of the original image data can be converted to obtain the converted image data. For example, the original image data can be converted from the RGB color space to the YUV color space. Since the human eye is more sensitive to luminance information, therefore, converting the color space of the original image data can effectively reduce the data volume while ensuring the visual effect, and at the same time improve the processing efficiency of subsequent operations. In a specific example, a specific conversion matrix can be used to accurately calculate the corresponding YUV component values based on the values of the three RGB components in the original image data, thereby completing the conversion of the color space of the original image data. Further, an intra-frame compression algorithm, such as a compression algorithm based on the discrete cosine transform, can be used to perform intra-frame compression on the converted image data, so as to obtain the target picture.

[0082] In a specific example, a tool A with data format conversion capabilities can provide a set of encoding interfaces, and the tool A interfaces can be called in sequence to obtain the target image data. Specifically, first, avcodec_find_encoder(AVCodecID) can be called to search for and obtain a specific picture encoder. The parameter ID is an enumerated type value representing the identifier of the encoder to be searched. For example, if the original image data is to be encoded in the M-JPEG (Motion Joint Photographic Experts Group) format, the decoder can be the M-JPEG encoder, and the parameter ID is AV_CODEC_ID_MJPEG. Further, avcodec_alloc_context3() can be called to allocate a buffer for the encoder and provide the context for the encoder to run. Further, avcodec_open2() can be called to initialize the encoding buffer, open the associated encoder, and set all the internal structures and parameters required by the encoder. Finally, avcodec_encode_video2() can be called to receive the original image data and encode it into M-JPEG graphic data.

[0083] Figure 7 is a flowchart of a method for generating a target picture provided in the second embodiment of the present invention. In a specific example, such as Figure 7As shown, assuming that the target picture is in the M-JPEG format, in the process of encoding the original image data to obtain the target picture, since the M-JPEG encoding process is based on the JPEG standard, it is necessary to match an encoder that supports JPEG encoding. Further, the encoder can perform format conversion on the original image data. For example, the original image data can be converted from the RGB color space to the YUV color space. After completing the format conversion, the encoder can perform intra-frame compression on the converted original image data to obtain an M-JPEG format image.

[0084] In an alternative embodiment of the present invention, after converting the original data of the key frame into the target picture, it may further include: creating a file input-output stream according to the target header file; creating a file creation object according to the file input-output stream; creating a new file through the file creation object and detecting the validity of the new file; when determining that the new file is valid, storing the data of the target picture into the new file through the file input-output stream; modifying the file suffix name of the new file to the suffix name of the picture format.

[0085] Among them, the target header file can be a header file used to define system-related characteristics and configurations during software development and compilation. The file input-output stream can be an input-output stream used to read and write files in computer programming. The file creation object can be a class or structure used to encapsulate the file creation and initialization logic. The new file can be a new file created through the file creation object. The file suffix name of the new file can be a specific suffix name added to the new file when creating the new file, which can be used to identify the type or purpose of the new file. Exemplarily, if the suffix name of the new file is ".txt", it means that the new file is a text file; if the suffix name of the new file is ".jpg", it means that the new file is a JPEG image file.

[0086] In an embodiment of the present invention, after converting the original data of the key frame into the target picture, to save the target picture, first, a file input-output stream can be created according to the target header file. Further, a file creation object can be created according to the file input-output stream, so that a new file can be created through the file creation object. Before writing data, it is also necessary to check the validity of the new file, so as to avoid the failure of creating the new file due to reasons such as the disk being full or having no permission when creating the new file. When determining that the new file is valid, the data of the target picture can be stored into the new file through the file input-output stream, and the file suffix name of the new file can be modified to the suffix name of the picture format.

[0087] Figure 8It is a flowchart of a method for saving a target picture provided in the second embodiment of the present invention. In a specific example, as Figure 8 shown, during the process of saving the target picture, it is first necessary to include the necessary header files to correctly use the file input / output stream. Further, a file object can be created using the file input / output stream, and this object can generate a new file as an additional file. It should be noted that the suffix name of the additional file can be set to a picture format. For example, the suffix name can be.jpeg or.png, etc. Setting the suffix name of the additional file to a picture format allows the picture to be viewed by clicking on the additional file. Further, the target picture can be written into the additional file using the file input / output stream. The above method can save the target picture in a specified format to the additional file and ensure that the additional file can be correctly recognized and viewed.

[0088] Optionally, in the implementation process of the embodiments of the present invention, a suitable data structure can be adopted, such as map (mapping), filter (filtering), and lambda (anonymous function), etc., to simplify the complexity of the code. At the same time, the system methods of the data structure can be called to process and parse the data. These system methods can not only ensure the conciseness of the code, but also significantly reduce the error rate of the code, and are also convenient for the subsequent maintenance and update of the code.

[0089] In the embodiments of the present invention, multiple data packets of the target video stream are obtained, and when it is determined that the obtained target data packet is the starting data packet of the key frame, the subsequent data packets of the key frame are obtained. Further, the protocol headers of the starting data packet and the subsequent data packets of the key frame can be removed to obtain the payload data of the key frame, so that the payload data of the key frame can be combined to obtain the original data of the key frame. After obtaining the original data of the key frame, the original data of the key frame can be decoded to obtain the original image data. Further, the original image data can be encoded to obtain the target picture, which solves the problems of complex use and cumbersome configuration of existing screen capture tools, and can simply and efficiently capture pictures from the video stream.

[0090] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0091] It should be noted that any permutation and combination of the technical features in the above embodiments also belong to the protection scope of the present invention.

[0092] Embodiment 3

[0093] Figure 9 It is a schematic diagram of a picture capture device provided in the third embodiment of the present invention, as Figure 9As shown in the figure, the device includes: a data packet acquisition module 310, a key frame acquisition module 320, a key frame original data acquisition module 330, and a target picture acquisition module 340, where:

[0094] The data packet acquisition module 310 is configured to acquire multiple data packets of a target video stream.

[0095] The key frame acquisition module 320 is configured to acquire subsequent data packets of the key frame when it is determined that the acquired target data packet is the starting data packet of the key frame.

[0096] The key frame original data acquisition module 330 is configured to acquire the original data of the key frame according to the starting packet of the key frame and the subsequent data packets of the key frame.

[0097] The target picture acquisition module 340 is configured to convert the original data of the key frame into a target picture.

[0098] In an embodiment of the present invention, by acquiring multiple data packets of a target video stream and acquiring subsequent data packets of a key frame when it is determined that the acquired target data packet is the starting data packet of the key frame. Further, the original data of the key frame can be acquired according to the starting data packet of the key frame and the subsequent data packets of the key frame, so that the original data of the key frame can be converted into a target picture, solving the problems that existing screen capture tools are complex to use and cumbersome to configure, and being able to simply and efficiently capture pictures from a video stream.

[0099] Optionally, the original data acquisition module 330 is specifically configured to: remove the protocol headers of the starting data packet of the key frame and the subsequent data packets of the key frame, and acquire the payload data of the key frame; combine the payload data of the key frame to obtain the original data of the key frame.

[0100] Optionally, the target picture acquisition module 340 is specifically configured to: decode the original data of the key frame to obtain original image data; encode the original image data to obtain the target picture.

[0101] Optionally, the target picture acquisition module 340 is further configured to: perform an inverse transformation operation on the original data of the key frame to determine the original image pixel values; combine the original image pixel values to obtain the original image data.

[0102] Optionally, the target picture acquisition module 340 is further configured to: perform a color space conversion on the original image data to obtain converted image data; perform intra-frame compression on the converted image data to obtain the target picture.

[0103] Optionally, the above device may further include a target picture saving module, configured to: create a file input / output stream according to the target header file; create a file creation object according to the file input / output stream; create a new file through the file creation object and detect the validity of the new file; in the case of determining that the new file is valid, store the data of the target picture into the new file through the file input / output stream; modify the file suffix name of the new file to the suffix name of the picture format.

[0104] The above picture capturing device can execute the picture capturing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the picture capturing method provided in any embodiment of the present invention.

[0105] Since the above-introduced picture capturing device is a device that can execute the picture capturing method in the embodiments of the present invention, based on the picture capturing method introduced in the embodiments of the present invention, those skilled in the art can understand the specific implementation manners and various variations of the picture capturing device in this embodiment. Therefore, the details of how the picture capturing device implements the picture capturing method in the embodiments of the present invention will not be described in detail here. As long as the device adopted by those skilled in the art to implement the picture capturing method in the embodiments of the present invention belongs to the scope protected by this application.

[0106] Embodiment 4

[0107] Figure 10 The structural schematic diagram of an electronic device 10 that can be used to implement the embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0108] As Figure 10As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0109] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0110] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the picture grabbing method.

[0111] In some embodiments, the picture grabbing method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the picture grabbing method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the picture grabbing method by any other appropriate means (e.g., by means of firmware).

[0112] Optionally, the picture capturing method may include: obtaining a plurality of data packets of a target video stream; obtaining subsequent data packets of the key frame when it is determined that the obtained target data packet is the starting data packet of the key frame; obtaining the original data of the key frame according to the starting data packet of the key frame and the subsequent data packets of the key frame; and converting the original data of the key frame into a target picture.

[0113] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0114] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processors of general-purpose computers, special-purpose computers, or other programmable data processing devices, such that when the computer programs are executed by the processors, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0115] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.

[0116] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).

[0117] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0118] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0119] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is made herein.

[0120] The above specific embodiments do not limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. A method for capturing images, characterized in that: include: Get multiple data packets of the target video stream; When it is determined that the acquired target data packet is a starting data packet of a key frame, acquiring subsequent data packets of the key frame; Acquire original data of the key frame according to the starting data packet of the key frame and the subsequent data packets of the key frame; The original data of the key frame is converted into a target image.

2. The method according to claim 1, characterized in that The obtaining the original data of the key frame according to the starting data packet of the key frame and the subsequent data packets of the key frame includes: Removing the protocol headers of the starting data packet of the key frame and the subsequent data packets of the key frame to obtain the payload data of the key frame; The payload data of the key frames are combined to obtain the original data of the key frames.

3. The method according to claim 1, characterized in that The converting the original data of the key frame into a target picture includes: Decoding the original data of the key frame to obtain original image data; The original image data is encoded to obtain the target image.

4. The method according to claim 3, characterized in that The decoding of the original data of the key frame to obtain the original image data includes: Performing an inverse transformation operation on the original data of the key frame to determine the original image pixel value; The original image pixel values ​​are combined to obtain the original image data.

5. The method according to claim 3, characterized in that: The encoding of the original image data to obtain the target image includes: Performing color space conversion on the original image data to obtain converted image data; The converted image data is intra-frame compressed to obtain the target image.

6. The method according to claim 1, characterized in that After converting the original data of the key frame into the target picture, the method further includes: Create file input and output streams based on the target header file; Create a file creation object according to the file input and output stream; Creating a new file through the file creation object, and detecting the validity of the new file; When it is determined that the newly added file is valid, storing the data of the target image into the newly added file through the file input and output stream; Change the file suffix of the newly added file to the suffix of the picture format.

7. A picture capture device, characterized in that: include: A data packet acquisition module, used to acquire multiple data packets of a target video stream; A key frame acquisition module, used to acquire subsequent data packets of the key frame when determining that the acquired target data packet is a starting data packet of the key frame; A key frame original data acquisition module, used for acquiring the original data of the key frame according to the starting data packet of the key frame and the subsequent data packets of the key frame; The target image acquisition module is used to convert the original data of the key frame into a target image.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the image capture method described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the image capture method described in any one of claims 1-6 when executed.

10. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the image capture method according to any one of claims 1 to 6 is implemented.