Image processing method and device based on key frame analysis

Through the image processing method based on keyframe analysis, the neural network model is used to process and restore images, and the problem of low efficiency of high-resolution image processing is solved, and efficient image data transmission and processing is achieved.

CN120014526BActive Publication Date: 2025-08-08WEIHAI KAISI INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510494922.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-08
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing image processing technology is inefficient in high-resolution image processing and cannot meet the processing requirements of real-time video streams.

Method used

The image processing method based on keyframe analysis is adopted, and continuous K-frame images are processed through a neural network model. All processing results are transmitted for keyframe images, while non-keyframe images only transmit some processing results, and unreleased processing results are restored using correlation.

Benefits of technology

It improves image processing efficiency, reduces data transmission volume, reduces bandwidth requirements, and improves the accuracy and efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014526B_ABST
    Figure CN120014526B_ABST
Patent Text Reader

Abstract

The present application provides an image processing method and device based on key frame analysis, which belongs to the field of image data processing technology and is used to improve image processing and transmission efficiency. The method includes: a first device obtains K consecutive frame images, the first frame image in the K frame images is a key frame image, and the K-1 frame images after the first frame image in the K frame images are K-1 frame non-key frame images, where K is an integer greater than 2; the first device processes the K frame images through a neural network model to obtain the processing results of the key frame images and the processing results of the K-1 frame non-key frame images; the first device sends the processing results of the key frame images and part of the processing results of the K-1 frame non-key frame images to the second device, the part of the processing results is determined according to the processing results of the key frame images, and the processing results of the key frame images and the part of the processing results are used to restore the K-1 frame non-key frame images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image data processing, and in particular to an image processing device based on key frame analysis. Background Art

[0002] With the rapid development of information technology, image processing technology has been widely applied in various fields. In particular, the processing and transmission of real-time video streams has become a key technology in areas such as video surveillance, remote medical diagnosis, and intelligent transportation. To meet the requirements of these applications, field equipment must efficiently and stably transmit captured video streams to remote devices for processing and analysis. Furthermore, to improve data processing efficiency and accuracy, field equipment typically employs advanced neural network models to process each frame in the video stream and extract feature information. This technology has the advantage of reducing data volume. Specifically, through feature extraction and data compression techniques, the amount of data transmitted can be significantly reduced, thereby lowering bandwidth requirements.

[0003] However, in future application scenarios, the image resolution may be larger, and the current image processing method may not be able to meet the image processing requirements. Summary of the Invention

[0004] The embodiments of the present application provide an image processing method and device based on key frame analysis to improve image processing efficiency.

[0005] To achieve the above objectives, this application adopts the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides an image processing method based on key frame analysis, which is applied to a first device, and includes: the first device acquires K consecutive frame images, the first frame image in the K frame images is a key frame image, the K-1 frame image after the first frame image in the K frame images is a K-1 frame non-key frame image, and K is an integer greater than 2; the first device processes the K frame images through a neural network model to obtain the processing results of the key frame images and the processing results of the K-1 frame non-key frame images; the first device sends the processing results of the key frame images and part of the processing results of the K-1 frame non-key frame images to the second device, the part of the processing results is determined according to the processing results of the key frame images, and the processing results of the key frame images and the part of the processing results are used to restore the K-1 frame non-key frame images.

[0007] Optionally, the first device processes K frame images through a neural network model to obtain processing results of key frame images and processing results of K-1 frame non-key frame images, including: the first device processes the key frame images through a neural network model to obtain processing results of key frame images, and the processing results of key frame images include multiple feature sets divided according to feature similarity; the first device processes K-1 frame non-key frame images respectively through a neural network model to obtain processing results of K-1 frame non-key frame images, and the processing results of K-1 frame non-key frame images respectively include multiple features of K-1 frame non-key frame images respectively.

[0008] Optionally, the first device processes the key frame image through a neural network model to obtain a processing result of the key frame image, including: the first device divides the key frame image into M1×N1 grids, M1 and N1 are both integers greater than 3, M1 is the number of grids in a row, and N1 is the number of grids in a column; the first device processes the M1×N1 grids in sequence through the neural network model to obtain M1×N1 processing results, and each processing result in the M1×N1 processing results includes a bit sequence; the first device divides the M1×N1 processing results into P1 feature sets by calculating the correlation of the bit sequences between the M1×N1 processing results, P1 is an integer greater than 1 and less than M1×N1, and the processing results belonging to the same feature set in the M1×N1 processing results are bit sequences. The processing results of the sequence are related, and the processing results belonging to different feature sets in the M1×N1 processing results are processing results that are unrelated to the bit sequence; accordingly, the first device processes the K-1 frames of non-key frame images respectively through the neural network model to obtain the processing results of the K-1 frames of non-key frame images, including: for the i-th frame of non-key frame image in the K-1 frames of non-key frame images, i is any integer from 2 to K: the first device divides the i-th frame of non-key frame image into Mi×Ni grids, Mi=M1, Ni=N1, Mi is the number of grids in a row, and Ni is the number of grids in a column; the first device processes the Mi×Ni grids in sequence through the neural network model to obtain Mi×Ni processing results, and each of the Mi×Ni processing results includes a bit sequence.

[0009] Optionally, the first device sends the processing results of the key frame image and part of the processing results of the K-1 frames of non-key frame images to the second device, including: the first device sends P1 feature sets to the second device; and, when i traverses 2 to K, the first device determines, based on the P1 feature sets, Pi processing results that correspond one-to-one and are related to the P1 feature sets among the Mi×Ni processing results, and sends the Pi processing results to the second device.

[0010] Optionally, the first device determines, based on the P1 feature set, Pi processing results among the Mi×Ni processing results that correspond one-to-one and are related to the P1 feature set, and sends the Pi processing results to the second device, including: the first device divides the Mi×Ni processing results into Pi feature sets according to the position of the grid, the position of the grid corresponding to the jth feature set in the Pi feature sets in the non-key frame image of the i-th frame is the same as the position of the grid corresponding to the jth feature set in the P1 feature sets in the key frame image, Pi=P1, j is any integer from 1 to Pi, the grid corresponding to the feature set means that the bit sequence in the feature set is obtained by processing the grid corresponding to the feature set; the first device determines a bit sequence from the jth feature set of the Pi feature sets, and when j traverses from 1 to Pi, Pi bit sequences are obtained, and the Pi bit sequences are the Pi processing results; the first device sends the Pi bit sequences to the second device.

[0011] Optionally, the first device determines a bit sequence from the jth feature set of Pi feature sets, and obtains Pi bit sequences when j traverses 1 to Pi, including: when i=2, the first device determines a bit sequence with the highest correlation with the P1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi; when i>2, the first device determines a bit sequence with the second highest correlation with the Pi-1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi, and the Pi-1 feature set corresponds to the i-1th non-key frame image in the K-1 non-key frame images.

[0012] Optionally, the first device sends Pi bit sequences to the second device, including: for the j-th bit sequence in the Pi bit sequences; the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, and sends the flipped j-th bit sequence to the second device. When j traverses from 1 to Pi, a total of Pi flipped bit sequences are sent.

[0013] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device fills the j-th bit sequence with a bit sequence of length L×L, and converts the bit sequence of length L×L into a matrix with L×L number of rows and columns; L is an odd number greater than or equal to 3; for any bit on the diagonal of the matrix, the first device flips the value of any bit from 1 to 0, or flips the value of any bit from 0 to 1, to obtain a flipped matrix, and the bits on the diagonal of the matrix are structural key bits; the first device restores the flipped matrix to a sequence to obtain the flipped j-th bit sequence.

[0014] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device divides the j-th bit sequence into multiple sub-bit sequences, and the lengths of the multiple sub-bit sequences are incremented by 2 starting from 1; for any sub-bit sequence in the multiple sub-bit sequences, the first device flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 1 to 0, or flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 0 to 1, to obtain multiple flipped sub-bit sequences, where the bits located at both ends of the column are structural key bits; the first device splices the multiple flipped sub-bit sequences to obtain the flipped j-th bit sequence.

[0015] In a second aspect, an image processing device based on key frame analysis is provided, which is applied to a first device, and the first device is configured as follows: the first device obtains K consecutive frame images, the first frame image in the K frame images is a key frame image, the K-1 frame image after the first frame image in the K frame images is a K-1 frame non-key frame image, and K is an integer greater than 2; the first device processes the K frame images through a neural network model to obtain the processing results of the key frame images and the processing results of the K-1 frame non-key frame images; the first device sends the processing results of the key frame images and part of the processing results of the K-1 frame non-key frame images to the second device, the part of the processing results is determined according to the processing results of the key frame images, and the processing results of the key frame images and the part of the processing results are used to restore the K-1 frame non-key frame images.

[0016] Optionally, the first device processes K frame images through a neural network model to obtain processing results of key frame images and processing results of K-1 frame non-key frame images, including: the first device processes the key frame images through a neural network model to obtain processing results of key frame images, and the processing results of key frame images include multiple feature sets divided according to feature similarity; the first device processes K-1 frame non-key frame images respectively through a neural network model to obtain processing results of K-1 frame non-key frame images, and the processing results of K-1 frame non-key frame images respectively include multiple features of K-1 frame non-key frame images respectively.

[0017] Optionally, the first device processes the key frame image through a neural network model to obtain a processing result of the key frame image, including: the first device divides the key frame image into M1×N1 grids, M1 and N1 are both integers greater than 3, M1 is the number of grids in a row, and N1 is the number of grids in a column; the first device processes the M1×N1 grids in sequence through the neural network model to obtain M1×N1 processing results, and each processing result in the M1×N1 processing results includes a bit sequence; the first device divides the M1×N1 processing results into P1 feature sets by calculating the correlation of the bit sequences between the M1×N1 processing results, P1 is an integer greater than 1 and less than M1×N1, and the processing results belonging to the same feature set in the M1×N1 processing results are bit sequences. The processing results of the sequence are related, and the processing results belonging to different feature sets in the M1×N1 processing results are processing results that are unrelated to the bit sequence; accordingly, the first device processes the K-1 frames of non-key frame images respectively through the neural network model to obtain the processing results of the K-1 frames of non-key frame images, including: for the i-th frame of non-key frame image in the K-1 frames of non-key frame images, i is any integer from 2 to K: the first device divides the i-th frame of non-key frame image into Mi×Ni grids, Mi=M1, Ni=N1, Mi is the number of grids in a row, and Ni is the number of grids in a column; the first device processes the Mi×Ni grids in sequence through the neural network model to obtain Mi×Ni processing results, and each of the Mi×Ni processing results includes a bit sequence.

[0018] Optionally, the first device sends the processing results of the key frame image and part of the processing results of the K-1 frames of non-key frame images to the second device, including: the first device sends P1 feature sets to the second device; and, when i traverses 2 to K, the first device determines, based on the P1 feature sets, Pi processing results that correspond one-to-one and are related to the P1 feature sets among the Mi×Ni processing results, and sends the Pi processing results to the second device.

[0019] Optionally, the first device determines, based on the P1 feature set, Pi processing results among the Mi×Ni processing results that correspond one-to-one and are related to the P1 feature set, and sends the Pi processing results to the second device, including: the first device divides the Mi×Ni processing results into Pi feature sets according to the position of the grid, the position of the grid corresponding to the jth feature set in the Pi feature sets in the non-key frame image of the i-th frame is the same as the position of the grid corresponding to the jth feature set in the P1 feature sets in the key frame image, Pi=P1, j is any integer from 1 to Pi, the grid corresponding to the feature set means that the bit sequence in the feature set is obtained by processing the grid corresponding to the feature set; the first device determines a bit sequence from the jth feature set of the Pi feature sets, and when j traverses from 1 to Pi, Pi bit sequences are obtained, and the Pi bit sequences are the Pi processing results; the first device sends the Pi bit sequences to the second device.

[0020] Optionally, the first device determines a bit sequence from the jth feature set of Pi feature sets, and obtains Pi bit sequences when j traverses 1 to Pi, including: when i=2, the first device determines a bit sequence with the highest correlation with the P1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi; when i>2, the first device determines a bit sequence with the second highest correlation with the Pi-1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi, and the Pi-1 feature set corresponds to the i-1th non-key frame image in the K-1 non-key frame images.

[0021] Optionally, the first device sends Pi bit sequences to the second device, including: for the j-th bit sequence in the Pi bit sequences; the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, and sends the flipped j-th bit sequence to the second device. When j traverses from 1 to Pi, a total of Pi flipped bit sequences are sent.

[0022] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device fills the j-th bit sequence with a bit sequence of length L×L, and converts the bit sequence of length L×L into a matrix with L×L number of rows and columns; L is an odd number greater than or equal to 3; for any bit on the diagonal of the matrix, the first device flips the value of any bit from 1 to 0, or flips the value of any bit from 0 to 1, to obtain a flipped matrix, and the bits on the diagonal of the matrix are structural key bits; the first device restores the flipped matrix to a sequence to obtain the flipped j-th bit sequence.

[0023] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device divides the j-th bit sequence into multiple sub-bit sequences, and the lengths of the multiple sub-bit sequences are incremented by 2 starting from 1; for any sub-bit sequence in the multiple sub-bit sequences, the first device flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 1 to 0, or flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 0 to 1, to obtain multiple flipped sub-bit sequences, where the bits located at both ends of the column are structural key bits; the first device splices the multiple flipped sub-bit sequences to obtain the flipped j-th bit sequence.

[0024] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having program code stored thereon. When the program code is run by the computer, the method described in the first aspect is executed.

[0025] In summary, the above method and device have the following technical effects:

[0026] Since the contents of consecutive K-frame images usually have continuity and correlation, the first frame image in the K-frame image can be regarded as a key frame image, and the K-1 frame image after the first frame image in the K-frame image can be regarded as a K-1 non-key frame image, where K is an integer greater than 2. At this time, when the K-frame image is processed by the neural network model, the first device sends the processing results of the key frame image to the second device, and only sends part of the processing results of the K-1 non-key frame images. At this time, the device on the opposite side can predict or restore the content of the image corresponding to the unsent processing result based on the correlation between the partial processing results and the processing results of the key frame image, thereby improving the image processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 A schematic diagram of the architecture of an image processing system provided in an embodiment of the present application;

[0028] Figure 2 A flowchart of an image processing method based on key frame analysis provided in an embodiment of the present application;

[0029] Figure 3 Schematic diagram of the application scenario of the method provided in the embodiment of the present application Figure 1 ;

[0030] Figure 4 Schematic diagram of the application scenario of the method provided in the embodiment of the present application Figure 2 ;

[0031] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] In the embodiment of the present invention, "indication" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. The information indicated by a certain information is called information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as but not limited to, the information to be indicated can be directly indicated, such as the information to be indicated itself or the index of the information to be indicated. The information to be indicated can also be indirectly indicated by indicating other information, wherein the other information and the information to be indicated have an association relationship. It is also possible to indicate only a part of the information to be indicated, while the other parts of the information to be indicated are known or agreed in advance. For example, the indication of specific information can be achieved by means of the arrangement order of each piece of information that is agreed upon in advance (for example, stipulated by the protocol), thereby reducing the indication overhead to a certain extent. At the same time, the common parts of each piece of information can be identified and indicated uniformly to reduce the indication overhead caused by indicating the same information separately.

[0033] In addition, the specific indication method can also be various existing indication methods, such as but not limited to the above-mentioned indication methods and various combinations thereof. The specific details of the various indication methods can refer to the existing technology and will not be repeated in this article. As can be seen from the above, for example, when it is necessary to indicate multiple information of the same type, there may be a situation where the indication methods for different information are different. In the specific implementation process, the required indication method can be selected according to specific needs. The embodiment of the present invention does not limit the selected indication method. In this way, the indication method involved in the embodiment of the present invention should be understood to cover various methods that can enable the party to be indicated to obtain the information to be indicated.

[0034] It should be understood that the information to be indicated can be sent as a whole or divided into multiple sub-information and sent separately, and the sending period and / or sending timing of these sub-information can be the same or different. The specific sending method is not limited by the embodiment of the present invention. The sending period and / or sending timing of these sub-information can be predefined, for example, predefined according to a protocol, or can be configured by the transmitting device through sending configuration information to the receiving device.

[0035] "Pre-definition" or "pre-configuration" can be achieved by pre-saving corresponding codes, tables or other methods that can be used to indicate relevant information in the device, and the embodiments of the present invention do not limit the specific implementation method. Among them, "saving" can mean saving in one or more memories. The one or more memories can be set separately or integrated in a processor or decoder, a processor, or an electronic device. The one or more memories can also be partially set separately and partially integrated in a decoder, a processor, or an electronic device. The type of memory can be any form of storage medium, and the embodiments of the present invention do not limit this.

[0036] The "protocol" involved in the embodiments of the present invention may refer to a protocol family in the communication field, a standard protocol with a similar protocol family frame structure, or a related protocol in a reliable access method system for future Internet of Things devices. The embodiments of the present invention do not specifically limit this.

[0037] In the embodiments of the present invention, descriptions such as "when...", "in the case of...", "if", and "if" all mean that the device will perform corresponding processing under certain objective circumstances. They do not limit the time, nor do they require the device to perform judgment actions during implementation, nor do they mean that there are other limitations.

[0038] In the description of the embodiments of the present invention, unless otherwise specified, " / " indicates that the associated objects are in an "or" relationship. For example, A / B can mean either A or B. "And / or" in the embodiments of the present invention merely describes an association relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, in the description of the embodiments of the present invention, unless otherwise specified, "multiple" refers to two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural. Furthermore, to facilitate the clear description of the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish between identical or similar items with substantially the same function or effect. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences. At the same time, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or design. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.

[0039] The network architecture and business scenarios described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Ordinary technicians in this field can know that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.

[0040] The technical solution in this application will be described below with reference to the accompanying drawings.

[0041] See also Figure 1 , an embodiment of the present application provides an image processing system, which may include a first device and a second device.

[0042] Both the first device and the second device may be devices in the form of a terminal, which may also be referred to as user equipment (UE), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user device. The terminal device in the embodiments of the present application may be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical care, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc.

[0043] The interaction between the first device and the second device in the above system will be described in detail below in conjunction with the method.

[0044] See also Figure 2 , the embodiment of the present application provides an image processing method based on key frame analysis, the process of the method is as follows:

[0045] S201, a first device obtains K consecutive frames of images.

[0046] K is an integer greater than 2. For example, considering that if the K value is too large, the continuity of the image content will be affected, the value of K can be 5-10.

[0047] The first frame in a K-frame image is a keyframe image. The K-1 frame image following the first frame in a K-frame image is a K-1 non-keyframe image. Keyframe images can transmit all processed content, while non-keyframe images can only transmit part of the processed content. This allows the full processed content of the keyframe images to be used to predict and recover the untransmitted processed content of the non-keyframe images.

[0048] The embodiments of the present application can be applied to ocean scenes. For example, the K-frame image can be an image of waves taken of the ocean, which can be used to analyze sea ripples in actual application scenarios, such as determining the current sea conditions and presetting future sea conditions. The above application scenarios are only some examples and are not specifically limited.

[0049] The embodiment of the present application does not limit the way in which the first device obtains continuous K-frame images. For example, the video stream captured by the shooting device can be transmitted to the first device. The first device can segment the video stream with K frames, and each segment includes continuous K-frame images. The continuous K-frame images in the embodiment of the present application are the continuous K-frame images contained in any segment.

[0050] S202, the first device processes K frame images through a neural network model to obtain processing results of the key frame images and processing results of K-1 frame non-key frame images.

[0051] The neural network model can be a deep neural network model, such as a convolutional neural network model or a fast convolutional neural network model, and there is no restriction on the specific model type.

[0052] The first device can process the key frame image through a neural network model to obtain a processing result of the key frame image, and the processing result of the key frame image includes multiple feature sets divided according to feature similarity.

[0053] For example, the first device can process with a grid as the granularity. For example, the first device can divide the key frame image into M1×N1 grids, where M1 and N1 are both integers greater than 3, M1 is the number of grids in a row, and N1 is the number of grids in a column. The values of M1 and N1 can be selected based on actual conditions. For example, if the performance of the transmitter and receiver are both relatively good, and prediction and recovery can be performed using fewer grids, then the values of M1 and N1 can be relatively large to further reduce transmission overhead. In other words, the better the performance of the device, the lower the processing overhead, such as M1=N1=20. For example, if the performance of the transmitter and receiver are both relatively poor, the values of M1 and N1 can be relatively small, such as M1=N1=8, to ensure performance. The first device can process M1×N1 grids sequentially using a neural network model to obtain M1×N1 processing results, each of which includes a bit sequence. Specifically, the first device can input each grid into the neural network model, which performs convolution on the grid to obtain a feature vector. The feature vector is then processed using a CABAC algorithm and an arithmetic algorithm to obtain a binary bit sequence. Thus, for each M1×N1 grid, M1×N1 bit sequences, namely, M1×N1 processing results, can be obtained one by one. The first device can calculate the correlation of the bit sequences between the M1×N1 processing results and divide the M1×N1 processing results into P1 feature sets, where P1 is an integer greater than 1 and less than M1×N1. Among the M1×N1 processing results, processing results belonging to the same feature set are bit sequence-correlated processing results, and processing results belonging to different feature sets among the M1×N1 processing results are bit sequence-uncorrelated processing results. For example, the first device can calculate the correlation between each two processing results among the M1×N1 processing results, such as the Euclidean distance or Manhattan distance between each two bit sequences, and thus determine that two processing results with a correlation greater than a preset threshold belong to the same feature set. In this way, processing results of the same feature set reflected in the grids indicate that the image content contained in the grids is relatively similar, while processing results of different feature sets reflected in the grids indicate that the image content contained in the grids is dissimilar.

[0054] For example, Figure 3 As shown in the figure, using an example, for a keyframe image, M1=N1=4, the six processing results corresponding to the six grids of pattern 1 all have correlations greater than the preset threshold, belonging to feature set #1. The eight processing results corresponding to the eight grids of pattern 2 all have correlations greater than the preset threshold, belonging to feature set #2. The two processing results corresponding to the two grids of pattern 3 all have correlations greater than the preset threshold, belonging to feature set #3, that is, P1=3.

[0055] The first device can also process K-1 frames of non-key frame images separately through a neural network model to obtain processing results of each K-1 frame of non-key frame images, and the processing results of each K-1 frame of non-key frame images include multiple features of each K-1 frame of non-key frame images.

[0056] For example, for the i-th non-key frame image in the K-1 frames of non-key frame images, i is any integer from 2 to K: the first device divides the i-th non-key frame image into Mi×Ni grids, Mi=M1, Ni=N1, Mi is the number of grids in a row, and Ni is the number of grids in a column; the first device then processes the Mi×Ni grids in sequence through the neural network model to obtain Mi×Ni processing results, each of the Mi×Ni processing results includes a bit sequence, that is, the way of dividing the grids is consistent with the key frame image, and the processing method is also consistent. The difference is that, at this time, for the Mi×Ni processing results, the first device does not first divide the feature set according to correlation.

[0057] S203: The first device sends the processing result of the key frame image and part of the processing results of the K-1 frames of non-key frame images to the second device.

[0058] The above partial processing results are determined based on the processing results of the key frame images. The partial processing results and the processing results of the key frame images can be used to restore the K-1 frame non-key frame image, which is described in detail below.

[0059] Based on S202 , it can be known that the first device sends P1 feature sets to the second device.

[0060] Also, when i traverses 2 to K, the first device can also determine Pi processing results among the Mi×Ni processing results that correspond one-to-one and are related to the P1 feature set based on the P1 feature set, and send the Pi processing results to the second device, which is described in detail below.

[0061] For example, the first device can divide Mi×Ni processing results into Pi feature sets according to the position of the grids, and the position of the grid corresponding to the jth feature set in the Pi feature sets in the non-key frame image of the i-th frame is the same as the position of the grid corresponding to the jth feature set in the P1 feature sets in the key frame image, Pi=P1, j is any integer from 1 to Pi. The grid corresponding to the feature set means that the bit sequence in the feature set is obtained by processing the grid corresponding to the feature set. Therefore, it can be seen that the above is to determine the processing results of the grids at the same position in the images of different frames as belonging to the same feature set. For example, if the processing results of a part of the grids in the M1×N1 grids are divided into the same feature set, then the processing results of the grids in the Mi×Ni grids of the non-key frame image that are consistent with the position of the part of the grids are also divided into the same feature set, such as the jth feature set.

[0062] On this basis, the first device can determine a bit sequence from the jth feature set of Pi feature sets, and when j traverses from 1 to Pi, Pi bit sequences are obtained, and the Pi bit sequences are Pi processing results;

[0063] Specifically, when i=2, the first device determines a bit sequence in the jth feature set of Pi feature sets that has the highest correlation with the P1 feature set by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the P1 feature set, and obtains Pi bit sequences when j traverses from 1 to Pi. In other words, the first device can specifically calculate the correlation between each bit sequence in the jth feature set and each bit sequence in the P1 feature set, such as the above-mentioned Euclidean distance or Manhattan distance, to obtain a total of multiple correlations, and select the bit sequence with the highest correlation corresponding to the jth feature set. In other words, for the second frame image in the K frame image, or the first non-key frame image, the first device needs to select the grid that is most similar to the key frame image to ensure the accuracy of recovery or prediction during subsequent decoding.

[0064] In the case where i>2, the first device sequentially calculates the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, determines a bit sequence in the jth feature set of Pi feature sets that has the second highest correlation with the Pi-1 feature set, and when j traverses from 1 to Pi, Pi bit sequences are obtained, and the Pi-1 feature set corresponds to the i-1th non-key frame image in the K-1 non-key frame images. In other words, the first device can specifically calculate the correlation between each bit sequence in the jth feature set of the next non-key frame image and each bit sequence in the previous non-key frame image, such as the above-mentioned Euclidean distance or Manhattan distance, to obtain multiple correlations in total, and select the bit sequence corresponding to the second highest correlation in the jth feature set. In other words, for the image of the third frame or later in the K frame image, or the non-key frame image of the second frame or later, considering that the image is gradually changing, the first device needs to select a grid that is most similar to the previous frame image, so that when performing recovery or prediction in decoding, the gradual change of the image content can be taken into account more, thereby achieving better recovery or prediction effect.

[0065] For ease of understanding, continue with the above Figure 3 In the example shown, for the first non-keyframe image, i.e., i=2, M2=N2=4, the six processing results corresponding to the six grids of pattern 4 all have a correlation greater than a preset threshold, belonging to feature set #4. The positions of the six grids of pattern 4 are the same as the positions of the six grids of pattern 1. The eight processing results corresponding to the eight grids of pattern 5 all have a correlation greater than a preset threshold, belonging to feature set #5. The positions of the eight grids of pattern 5 are the same as the positions of the eight grids of pattern 2. The two processing results corresponding to the two grids of pattern 6 all have a correlation greater than a preset threshold, belonging to feature set #6. The positions of the two grids of pattern 6 are the same as the positions of the two grids of pattern 3. In this case, P2=3. Similarly, for the second non-keyframe image, i.e., i=3, M3=N3=4, the six processing results corresponding to the six grids of pattern 7 all have correlations greater than the preset threshold and belong to feature set #7. The positions of the six grids of pattern 7 are the same as those of the six grids of pattern 1. The eight processing results corresponding to the eight grids of pattern 8 all have correlations greater than the preset threshold and belong to feature set #8. The positions of the eight grids of pattern 8 are the same as those of the eight grids of pattern 2. The two processing results corresponding to the two grids of pattern 9 all have correlations greater than the preset threshold and belong to feature set #9. The positions of the two grids of pattern 9 are the same as those of the two grids of pattern 3. In this case, P3=3.

[0066] On this basis, for feature set #4, the first processing result has the highest correlation with the second processing result in feature set #1. Therefore, the first processing result in feature set #4 needs to be sent to the second device, and the remaining five processing results in feature set #4 do not need to be sent to the second device. In addition, the implementation of feature set #5 and feature set #6 is similar to feature set #4 and will not be repeated. For feature set #7, the third processing result has the second highest correlation with the fifth processing result in feature set #4. Therefore, the third processing result in feature set #7 needs to be sent to the second device, and the remaining five processing results in feature set #7 do not need to be sent to the second device. In addition, the implementation of feature set #8 and feature set #9 is similar to feature set #7 and will not be repeated.

[0067] Finally, the first device sends the Pi bit sequence to the second device.

[0068] For example, for the jth bit sequence among the Pi bit sequences;

[0069] The first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, and sends the flipped j-th bit sequence to the second device. When j traverses from 1 to Pi, a total of Pi flipped bit sequences are sent.

[0070] In one possible manner, the first device fills the j-th bit sequence with a bit sequence of length L×L, and converts the bit sequence of length L×L into a matrix with L×L rows and columns; L is an odd number greater than or equal to 3; for any bit on the diagonal of the matrix, the first device flips the value of any bit from 1 to 0, or flips the value of any bit from 0 to 1, to obtain a flipped matrix, where the bits on the diagonal of the matrix are structural key bits; the first device restores the flipped matrix to a sequence to obtain a flipped j-th bit sequence.

[0071] like Figure 4 As shown, through an example, assuming that the length of the j-th bit sequence is 33, the length of each bit sequence is fixed, such as 110011101010001001000111110110011, the first device can fill 3 bits with a value of 0 at the end of the j-th bit sequence to obtain 110011101010001001000111110110011000. At this time, it can be divided into a 6×6 matrix, and the diagonal elements are 101110 and 110010, which are 010001 and 001101 after flipping.

[0072] Alternatively, in another possible manner, the first device divides the j-th bit sequence into multiple sub-bit sequences, and the lengths of the multiple sub-bit sequences increase from 1 to 2 in sequence; for any sub-bit sequence among the multiple sub-bit sequences, the first device flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 1 to 0, or flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 0 to 1, to obtain multiple flipped sub-bit sequences, where the bits located at both ends of the column are structural key bits; the first device splices the multiple flipped sub-bit sequences to obtain a flipped j-th bit sequence.

[0073] For ease of understanding, an example is also used to illustrate. Assume that the length of the j-th bit sequence is 33, and the length of each bit sequence is fixed, such as 110011101010001001000111110110011. The first device can fill 3 bits with the value of 0 at the end of the j-th bit sequence to obtain 11001110101000 1001000111110110011000. At this time, it can be divided into sub-bit sequences with lengths of 6, 8, 10, and 12, respectively, which are 110011, 10101000, 1001000111, and 1001000111, respectively. After flipping, they are 010010, 00101001, 0001000110, and 0001000110, respectively. Of course, the above example takes the flipping of one bit at each end as an example, and it may also be a plurality of consecutive bits.

[0074] It can be seen that the above method flips some bits in the j-th bit sequence, making the j-th bit sequence de-informatized. Since the positions of these flipped bits are structured, such as being located on the diagonal of the matrix or at both ends of the sequence, the processing end can perform the inverse process of the above method after receiving the flipped j-th bit sequence, thereby restoring the flipped bits. However, if the above flipped j-th bit sequence is stolen by an attacker, the attacker cannot restore the flipped bits because they do not know the specific processing process. In this case, the flipped j-th bit sequence is invalid information to the attacker, thereby improving the data security of the image data.

[0075] For the second device, after receiving part of the processing results of the respective processing results of the K-1 frames of non-key frame images, the second device can perform the inverse process of the above-mentioned S203, that is, flip the corresponding bits according to the above-mentioned rules. Then, the second device inputs the processing results of the key frame image and the part of the processing results of the K frames of non-key frame images into a neural network model of the second device (such as neural network model #1), and the neural network model #1 performs the inverse process of the above-mentioned convolution, that is, deconvolution, to obtain the grids of the key frame image and the K frames of non-key frame images. At this time, the second device can input the grids of the key frame image and the second frame of non-key frame image, as well as the information indicating the area corresponding to each grid of the second frame of non-key frame image in the key frame image, into another neural network model of the second device (such as neural network model #2). The corresponding area in the key frame image, that is, the area of the grid corresponding to each of the above-mentioned feature sets in the key frame image, as Figure 3 The area occupied by pattern 1 or pattern 4 in the image. In this way, the neural network model #2 can predict the image content of the other parts of the area outside the grid in the second non-key frame image based on the image content in the area and the grid of the second non-key frame image. For example, Figure 3 As shown, neural network model #2 uses the image content of the area occupied by pattern 1 and the image content of a grid in pattern 4 to predict and restore the image content of the area occupied by the grid of the entire pattern 4, thereby achieving image restoration with higher precision and accuracy based on the region granularity. Subsequently, the second device can input the grids of the keyframe image and the third non-keyframe image, as well as information indicating the region in the keyframe image corresponding to each grid in the third non-keyframe image, into neural network model #2, and so on.

[0076] In summary, since the content of continuous K frame images usually has continuity and correlation, the first frame image in the K frame image can be regarded as the key frame image, and the K-1 frame image after the first frame image in the K frame image can be regarded as the K-1 frame non-key frame image, where K is an integer greater than 2; at this time, when the K frame image is processed by the neural network model, the first device sends the processing result of the key frame image to the second device, and only sends part of the processing results of each of the K-1 frame non-key frame images. At this time, the device on the opposite side can predict or restore the content of the image corresponding to the unsent processing result based on the correlation between the partial processing result and the processing result of the key frame image, thereby improving the image processing efficiency.

[0077] Combination of the above Figure 3The method provided by the embodiment of the present application is described in detail. The following introduces an image processing device based on key frame analysis for executing the method provided by the embodiment of the present application, which is applied to a first device, and the first device is configured as follows: the first device obtains K consecutive frame images, the first frame image in the K frame images is a key frame image, the K-1 frame image after the first frame image in the K frame images is a K-1 frame non-key frame image, and K is an integer greater than 2; the first device processes the K frame images through a neural network model to obtain the processing results of the key frame images and the processing results of the K-1 frame non-key frame images; the first device sends the processing results of the key frame images and part of the processing results of the K-1 frame non-key frame images to the second device, the part of the processing results is determined according to the processing results of the key frame images, and the processing results of the key frame images and the part of the processing results are used to restore the K-1 frame non-key frame images.

[0078] Optionally, the first device processes K frame images through a neural network model to obtain processing results of key frame images and processing results of K-1 frame non-key frame images, including: the first device processes the key frame images through a neural network model to obtain processing results of key frame images, and the processing results of key frame images include multiple feature sets divided according to feature similarity; the first device processes K-1 frame non-key frame images respectively through a neural network model to obtain processing results of K-1 frame non-key frame images, and the processing results of K-1 frame non-key frame images respectively include multiple features of K-1 frame non-key frame images respectively.

[0079] Optionally, the first device processes the key frame image through a neural network model to obtain a processing result of the key frame image, including: the first device divides the key frame image into M1×N1 grids, M1 and N1 are both integers greater than 3, M1 is the number of grids in a row, and N1 is the number of grids in a column; the first device processes the M1×N1 grids in sequence through the neural network model to obtain M1×N1 processing results, and each processing result in the M1×N1 processing results includes a bit sequence; the first device divides the M1×N1 processing results into P1 feature sets by calculating the correlation of the bit sequences between the M1×N1 processing results, P1 is an integer greater than 1 and less than M1×N1, and the processing results belonging to the same feature set in the M1×N1 processing results are bit sequences. The processing results of the sequence are related, and the processing results belonging to different feature sets in the M1×N1 processing results are processing results that are unrelated to the bit sequence; accordingly, the first device processes the K-1 frames of non-key frame images respectively through the neural network model to obtain the processing results of the K-1 frames of non-key frame images, including: for the i-th frame of non-key frame image in the K-1 frames of non-key frame images, i is any integer from 2 to K: the first device divides the i-th frame of non-key frame image into Mi×Ni grids, Mi=M1, Ni=N1, Mi is the number of grids in a row, and Ni is the number of grids in a column; the first device processes the Mi×Ni grids in sequence through the neural network model to obtain Mi×Ni processing results, and each of the Mi×Ni processing results includes a bit sequence.

[0080] Optionally, the first device sends the processing results of the key frame image and part of the processing results of the K-1 frames of non-key frame images to the second device, including: the first device sends P1 feature sets to the second device; and, when i traverses 2 to K, the first device determines, based on the P1 feature sets, Pi processing results that correspond one-to-one and are related to the P1 feature sets among the Mi×Ni processing results, and sends the Pi processing results to the second device.

[0081] Optionally, the first device determines, based on the P1 feature set, Pi processing results among the Mi×Ni processing results that correspond one-to-one and are related to the P1 feature set, and sends the Pi processing results to the second device, including: the first device divides the Mi×Ni processing results into Pi feature sets according to the position of the grid, the position of the grid corresponding to the jth feature set in the Pi feature sets in the non-key frame image of the i-th frame is the same as the position of the grid corresponding to the jth feature set in the P1 feature sets in the key frame image, Pi=P1, j is any integer from 1 to Pi, the grid corresponding to the feature set means that the bit sequence in the feature set is obtained by processing the grid corresponding to the feature set; the first device determines a bit sequence from the jth feature set of the Pi feature sets, and when j traverses from 1 to Pi, Pi bit sequences are obtained, and the Pi bit sequences are the Pi processing results; the first device sends the Pi bit sequences to the second device.

[0082] Optionally, the first device determines a bit sequence from the jth feature set of Pi feature sets, and obtains Pi bit sequences when j traverses 1 to Pi, including: when i=2, the first device determines a bit sequence with the highest correlation with the P1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi; when i>2, the first device determines a bit sequence with the second highest correlation with the Pi-1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi, and the Pi-1 feature set corresponds to the i-1th non-key frame image in the K-1 non-key frame images.

[0083] Optionally, the first device sends Pi bit sequences to the second device, including: for the j-th bit sequence in the Pi bit sequences; the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, and sends the flipped j-th bit sequence to the second device. When j traverses from 1 to Pi, a total of Pi flipped bit sequences are sent.

[0084] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device fills the j-th bit sequence with a bit sequence of length L×L, and converts the bit sequence of length L×L into a matrix with L×L number of rows and columns; L is an odd number greater than or equal to 3; for any bit on the diagonal of the matrix, the first device flips the value of any bit from 1 to 0, or flips the value of any bit from 0 to 1, to obtain a flipped matrix, and the bits on the diagonal of the matrix are structural key bits; the first device restores the flipped matrix to a sequence to obtain the flipped j-th bit sequence.

[0085] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device divides the j-th bit sequence into multiple sub-bit sequences, and the lengths of the multiple sub-bit sequences are incremented by 2 starting from 1; for any sub-bit sequence in the multiple sub-bit sequences, the first device flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 1 to 0, or flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 0 to 1, to obtain multiple flipped sub-bit sequences, where the bits located at both ends of the column are structural key bits; the first device splices the multiple flipped sub-bit sequences to obtain the flipped j-th bit sequence.

[0086] The following combination Figure 5 The components of the electronic device 500 are described in detail.

[0087] The processor 501 is the control center of the electronic device 500 and can be a single processor or a collective term for multiple processing elements. For example, the processor 501 can be one or more central processing units (CPUs), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application, such as one or more microprocessors (digital signal processors, DSPs) or one or more field programmable gate arrays (FPGAs).

[0088] Optionally, the processor 501 can execute various functions of the electronic device 500 by running or executing the software program stored in the memory 502 and calling the data stored in the memory 502, as described above. Figure 2 Function in the method shown.

[0089] In a specific implementation, as an embodiment, the processor 501 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 are shown in FIG.

[0090] In a specific implementation, as an example, the electronic device 500 may also include multiple processors. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0091] Among them, the memory 502 is used to store the software program for executing the solution of the present application, and the execution is controlled by the processor 501. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0092] Alternatively, the memory 502 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 502 may be integrated with the processor 501 or exist independently and be connected to the interface circuit ( Figure 5 (not shown) is coupled to the processor 501, which is not specifically limited in this embodiment of the present application.

[0093] The transceiver 503 is used for communicating with other devices. For example, if the multi-beam positioning device is a terminal, the transceiver 503 can be used to communicate with a network device or another terminal.

[0094] Optionally, the transceiver 503 may include a receiver and a transmitter ( Figure 5 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0095] Optionally, the transceiver 503 may be integrated with the processor 501 or may exist independently and communicate with the electronic device 500 through the interface circuit ( Figure 5 (not shown) is coupled to the processor 501, which is not specifically limited in this embodiment of the present application.

[0096] It should be noted that Figure 5 The structure of the electronic device 500 shown in the figure does not constitute a limitation on the device. The actual electronic device 500 may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0097] In addition, the technical effects based on the electronic device 500 can refer to the technical effects of the method in the above method embodiment, which will not be repeated here.

[0098] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), but may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0099] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0100] The above embodiments can be implemented in whole or in part via software, hardware (e.g., circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product comprises one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are fully or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). A computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.

[0101] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0102] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0103] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0104] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0105] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0106] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some feature fields can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0107] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0108] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0109] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0110] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. An image processing method based on key frame analysis, characterized in that: Applied to a first device, the method includes: The first device acquires K consecutive frames of images, wherein the first frame of the K frames is a key frame image, and the K-1 frames of the K frames following the first frame are K-1 non-key frame images, where K is an integer greater than 2; The first device processes the K frames of images using a neural network model to obtain processing results of the key frame images and processing results of the K-1 frames of non-key frame images; The first device sends a processing result of the key frame image and a partial processing result of each processing result of the K-1 frame non-key frame image to the second device, where the partial processing result is determined based on the processing result of the key frame image, and the processing result of the key frame image and the partial processing result are used to restore the K-1 frame non-key frame image; The first device processes the K frame images through a neural network model to obtain processing results of the key frame images and processing results of the K-1 frame non-key frame images, including: The first device processes the key frame image through the neural network model to obtain a processing result of the key frame image, wherein the processing result of the key frame image includes a plurality of feature sets divided according to feature similarity; The first device processes the K-1 frames of non-key frame images separately through the neural network model to obtain processing results of each of the K-1 frames of non-key frame images. The processing results of each of the K-1 frames of non-key frame images include multiple features of each of the K-1 frames of non-key frame images. Some of the processing results of each of the K-1 frames of non-key frame images are some features of the multiple features that are related to the multiple feature sets.

2. The method according to claim 1, characterized in that The first device processes the key frame image through the neural network model to obtain a processing result of the key frame image, including: The first device divides the key frame image into M1×N1 grids, where M1 and N1 are both integers greater than 3, M1 is the number of grids in a row, and N1 is the number of grids in a column; The first device processes the M1×N1 grids in sequence through the neural network model to obtain M1×N1 processing results, each of the M1×N1 processing results including a bit sequence; The first device divides the M1×N1 processing results into P1 feature sets by calculating the correlation of bit sequences between the M1×N1 processing results, where P1 is an integer greater than 1 and less than M1×N1. Processing results belonging to the same feature set among the M1×N1 processing results are bit sequence-correlated processing results, and processing results belonging to different feature sets among the M1×N1 processing results are bit sequence-independent processing results; Accordingly, the first device processes the K-1 frames of non-key frame images respectively through the neural network model to obtain processing results of the K-1 frames of non-key frame images, including: For the i-th non-key frame image in the K-1 non-key frame images, i is any integer from 2 to K: The first device divides the i-th non-key frame image into Mi×Ni grids, Mi=M1, Ni=N1, Mi is the number of grids in a row, and Ni is the number of grids in a column; The first device processes the Mi×Ni grids in sequence through the neural network model to obtain Mi×Ni processing results, each of the Mi×Ni processing results including a bit sequence.

3. The method according to claim 2, characterized in that The first device sends a processing result of the key frame image and a portion of the processing results of the K-1 frame non-key frame image to the second device, including: The first device sends the P1 feature sets to the second device; And, when i traverses 2 to K, the first device determines Pi processing results among the Mi×Ni processing results that correspond one-to-one and are related to the P1 feature sets based on the P1 feature sets, and sends the Pi processing results to the second device.

4. The method according to claim 3, characterized in that The first device determines, based on the P1 feature sets, Pi processing results among the Mi×Ni processing results that correspond one-to-one to and are related to the P1 feature sets, and sends the Pi processing results to the second device, including: The first device divides the Mi×Ni processing results into Pi feature sets according to the positions of the grids, the position of the grid corresponding to the jth feature set in the Pi feature sets in the i-th non-key frame image is the same as the position of the grid corresponding to the jth feature set in the P1 feature sets in the key frame image, Pi=P1, j is any integer from 1 to Pi, and the grid corresponding to the feature set means that the bit sequence in the feature set is obtained by processing the grid corresponding to the feature set; The first device determines a bit sequence from the jth feature set of the Pi feature sets, and obtains Pi bit sequences when j traverses from 1 to Pi, and the Pi bit sequences are the Pi processing results; The first device sends the Pi bit sequences to the second device.

5. The method according to claim 4, characterized in that The first device determines a bit sequence from the jth feature set of the Pi feature sets, and obtains Pi bit sequences when j traverses from 1 to Pi, including: When i=2, the first device sequentially calculates the correlation between each bit sequence in the j-th feature set of the Pi feature sets and the P1 feature set, determines a bit sequence in the j-th feature set of the Pi feature sets that has the highest correlation with the P1 feature set, and obtains the Pi bit sequences when j traverses from 1 to Pi; When i>2, the first device determines a bit sequence in the jth feature set of the Pi feature sets that has the second highest correlation with the Pi-1 feature set by successively calculating the correlation between each bit sequence in the jth feature set of the Pi feature sets and the Pi-1 feature set. When j traverses from 1 to Pi, the Pi bit sequences are obtained, and the Pi-1 feature set corresponds to the i-1th non-key frame image in the K-1 non-key frame images.

6. The method according to claim 4 or 5, characterized in that The first device sending the Pi bit sequences to the second device includes: For the j-th bit sequence among the Pi bit sequences; The first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, and sends the flipped j-th bit sequence to the second device. When j traverses from 1 to Pi, a total of Pi flipped bit sequences are sent.

7. The method according to claim 6, characterized in that The first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: The first device pads the j-th bit sequence into a bit sequence of length L×L, and converts the bit sequence of length L×L into a matrix of L×L rows and columns; L is an odd number greater than or equal to 3; For any bit on the diagonal of the matrix, the first device flips the value of the any bit from 1 to 0, or flips the value of the any bit from 0 to 1, to obtain a flipped matrix, where the bits on the diagonal of the matrix are the structural key bits; The first device restores the flipped matrix to a sequence to obtain the flipped j-th bit sequence.

8. The method according to claim 6, characterized in that The first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: The first device divides the j-th bit sequence into a plurality of sub-bit sequences, wherein the lengths of the plurality of sub-bit sequences are sequentially increased from 1 to 2; For any sub-bit sequence among the multiple sub-bit sequences, the first device flips the values of bits located at both ends of the sub-bit sequence from 1 to 0, or flips the values of bits located at both ends of the sub-bit sequence from 0 to 1, to obtain multiple flipped sub-bit sequences, where the bits located at both ends of the sub-bit sequence are the structural key bits; The first device concatenates the flipped multiple sub-bit sequences to obtain the flipped j-th bit sequence.

Citation Information

Patent Citations

  • Video semantic segmentation method and device

    CN112465826A