Image processing method and device based on key frame analysis
Through the image processing method based on keyframe analysis, the neural network model is used to process continuous K-frame images and send keyframe and partial non-keyframe processing results, the problem of greater resolution image processing needs in the future is solved and image processing efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510494922.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The current image processing method is difficult to meet the processing needs of larger-resolution images in the future, resulting in insufficient processing efficiency and accuracy.
The image processing method based on keyframe analysis is adopted, and the continuous K-frame images are processed through a neural network model, and the processing results of the keyframe images and the processing results of some non-keyframe images are sent to the opposite device to restore the unsent image content.
The image processing efficiency is improved, and the correlation between the processing results of the keyframe image and the partial processing results of the non-keyframe image is achieved, and the prediction and recovery of the content of the unspread image is achieved.
Smart Images

Figure CN120014526A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image data processing, and in particular to an image processing and device based on key frame analysis. Background Art
[0002] With the rapid development of information technology, image processing technology has been widely used in various fields. Especially in the fields of video surveillance, remote medical diagnosis, intelligent transportation, etc., the processing and transmission of real-time video streams has become one of the key technologies. In order to meet the needs of these applications, field equipment needs to transmit the collected video streams to remote devices for processing and analysis in an efficient and stable manner. At the same time, in order to improve data processing efficiency and accuracy, field equipment usually uses advanced neural network models to process each frame of the video stream and extract feature information. The advantage of this technology is that it can reduce the amount of data, that is, through feature extraction and data compression technology, the amount of data in the transmission process can be greatly reduced, reducing bandwidth requirements.
[0003] However, in future application scenarios, the image resolution may be larger, and the current image processing method may not be able to meet the image processing requirements. Summary of the invention
[0004] The embodiments of the present application provide an image processing method and device based on key frame analysis to improve image processing efficiency.
[0005] In order to achieve the above objectives, this application adopts the following technical solutions: In a first aspect, an embodiment of the present application provides an image processing method based on key frame analysis, which is applied to a first device, and includes: the first device acquires K consecutive frame images, the first frame image in the K frame images is a key frame image, the K-1 frame image after the first frame image in the K frame images is a K-1 frame non-key frame image, and K is an integer greater than 2; the first device processes the K frame images through a neural network model to obtain processing results of the key frame images and processing results of the K-1 frame non-key frame images; the first device sends the processing results of the key frame images and part of the processing results of the K-1 frame non-key frame images to the second device, the part of the processing results is determined according to the processing results of the key frame images, and the processing results of the key frame images and the part of the processing results are used to restore the K-1 frame non-key frame images.
[0006] Optionally, the first device processes K frame images through a neural network model to obtain processing results of the key frame images and processing results of K-1 frame non-key frame images, including: the first device processes the key frame images through a neural network model to obtain processing results of the key frame images, and the processing results of the key frame images include multiple feature sets divided according to feature similarity; the first device processes K-1 frame non-key frame images respectively through a neural network model to obtain processing results of the K-1 frame non-key frame images, and the processing results of the K-1 frame non-key frame images include multiple features of the K-1 frame non-key frame images.
[0007] Optionally, the first device processes the key frame image through a neural network model to obtain a processing result of the key frame image, including: the first device divides the key frame image into M1×N1 grids, M1 and N1 are both integers greater than 3, M1 is the number of grids in a row, and N1 is the number of grids in a column; the first device processes the M1×N1 grids in sequence through the neural network model to obtain M1×N1 processing results, each of the M1×N1 processing results including a bit sequence; the first device divides the M1×N1 processing results into P1 feature sets by calculating the correlation of the bit sequences between the M1×N1 processing results, P1 is an integer greater than 1 and less than M1×N1, and the processing results belonging to the same feature set among the M1×N1 processing results are bit sequences. The processing results are sequence-related, and the processing results belonging to different feature sets among the M1×N1 processing results are bit sequence-independent processing results; accordingly, the first device processes the K-1 frames of non-key frame images respectively through the neural network model to obtain the processing results of the K-1 frames of non-key frame images, including: for the i-th frame of non-key frame image in the K-1 frames of non-key frame images, i is any integer from 2 to K: the first device divides the i-th frame of non-key frame image into Mi×Ni grids, Mi=M1, Ni=N1, Mi is the number of grids in a row, and Ni is the number of grids in a column; the first device processes the Mi×Ni grids in sequence through the neural network model to obtain Mi×Ni processing results, and each of the Mi×Ni processing results includes a bit sequence.
[0008] Optionally, the first device sends the processing results of the key frame image and part of the processing results of the K-1 frame non-key frame images to the second device, including: the first device sends P1 feature sets to the second device; and, when i traverses from 2 to K, the first device determines, based on the P1 feature sets, Pi processing results among the Mi×Ni processing results that correspond one-to-one and are related to the P1 feature sets, and sends the Pi processing results to the second device.
[0009] Optionally, the first device determines, based on P1 feature sets, Pi processing results among Mi×Ni processing results that correspond one-to-one and are related to the P1 feature sets, and sends the Pi processing results to the second device, including: the first device divides the Mi×Ni processing results into Pi feature sets according to the positions of the grids, the position of the grid corresponding to the jth feature set among the Pi feature sets in the non-key frame image of the i-th frame is the same as the position of the grid corresponding to the jth feature set among the P1 feature sets in the key frame image, Pi=P1, j is any integer from 1 to Pi, the grid corresponding to the feature set means that the bit sequence in the feature set is obtained by processing the grid corresponding to the feature set; the first device determines a bit sequence from the jth feature set of the Pi feature sets, and when j traverses from 1 to Pi, Pi bit sequences are obtained, and the Pi bit sequences are the Pi processing results; the first device sends the Pi bit sequences to the second device.
[0010] Optionally, the first device determines a bit sequence from the jth feature set of Pi feature sets, and obtains Pi bit sequences when j traverses 1 to Pi, including: when i=2, the first device determines a bit sequence with the highest correlation with the P1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi; when i>2, the first device determines a bit sequence with the second highest correlation with the Pi-1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi, and the Pi-1 feature set corresponds to the i-1th non-key frame image in the K-1 non-key frame images.
[0011] Optionally, the first device sends Pi bit sequences to the second device, including: for the j-th bit sequence among the Pi bit sequences; the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, and sends the flipped j-th bit sequence to the second device, and when j traverses from 1 to Pi, a total of Pi flipped bit sequences are sent.
[0012] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device fills the j-th bit sequence with a bit sequence of length L×L, and converts the bit sequence of length L×L into a matrix with L×L number of rows and columns; L is an odd number greater than or equal to 3; for any bit on the diagonal of the matrix, the first device flips the value of any bit from 1 to 0, or flips the value of any bit from 0 to 1, to obtain a flipped matrix, the bits on the diagonal of the matrix are structural key bits; the first device restores the flipped matrix to a sequence to obtain the flipped j-th bit sequence.
[0013] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device divides the j-th bit sequence into multiple sub-bit sequences, and the lengths of the multiple sub-bit sequences are incremented from 1 to 2; for any sub-bit sequence in the multiple sub-bit sequences, the first device flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 1 to 0, or flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 0 to 1, to obtain multiple flipped sub-bit sequences, and the bits located at both ends of the column are structural key bits; the first device splices the multiple flipped sub-bit sequences to obtain the flipped j-th bit sequence.
[0014] In a second aspect, an image processing device based on key frame analysis is provided, which is applied to a first device, and the first device is configured as follows: the first device obtains K consecutive frame images, the first frame image in the K frame images is a key frame image, the K-1 frame image after the first frame image in the K frame images is a K-1 frame non-key frame image, and K is an integer greater than 2; the first device processes the K frame images through a neural network model to obtain processing results of the key frame images and processing results of the K-1 frame non-key frame images; the first device sends the processing results of the key frame images and part of the processing results of the K-1 frame non-key frame images to the second device, the part of the processing results is determined according to the processing results of the key frame images, and the processing results of the key frame images and the part of the processing results are used to restore the K-1 frame non-key frame images.
[0015] Optionally, the first device processes K frame images through a neural network model to obtain processing results of the key frame images and processing results of K-1 frame non-key frame images, including: the first device processes the key frame images through a neural network model to obtain processing results of the key frame images, and the processing results of the key frame images include multiple feature sets divided according to feature similarity; the first device processes K-1 frame non-key frame images respectively through a neural network model to obtain processing results of the K-1 frame non-key frame images, and the processing results of the K-1 frame non-key frame images include multiple features of the K-1 frame non-key frame images.
[0016] Optionally, the first device processes the key frame image through a neural network model to obtain a processing result of the key frame image, including: the first device divides the key frame image into M1×N1 grids, M1 and N1 are both integers greater than 3, M1 is the number of grids in a row, and N1 is the number of grids in a column; the first device processes the M1×N1 grids in sequence through the neural network model to obtain M1×N1 processing results, each of the M1×N1 processing results including a bit sequence; the first device divides the M1×N1 processing results into P1 feature sets by calculating the correlation of the bit sequences between the M1×N1 processing results, P1 is an integer greater than 1 and less than M1×N1, and the processing results belonging to the same feature set among the M1×N1 processing results are bit sequences. The processing results are sequence-related, and the processing results belonging to different feature sets among the M1×N1 processing results are bit sequence-independent processing results; accordingly, the first device processes the K-1 frames of non-key frame images respectively through the neural network model to obtain the processing results of the K-1 frames of non-key frame images, including: for the i-th frame of non-key frame image in the K-1 frames of non-key frame images, i is any integer from 2 to K: the first device divides the i-th frame of non-key frame image into Mi×Ni grids, Mi=M1, Ni=N1, Mi is the number of grids in a row, and Ni is the number of grids in a column; the first device processes the Mi×Ni grids in sequence through the neural network model to obtain Mi×Ni processing results, and each of the Mi×Ni processing results includes a bit sequence.
[0017] Optionally, the first device sends the processing results of the key frame image and part of the processing results of the K-1 frame non-key frame images to the second device, including: the first device sends P1 feature sets to the second device; and, when i traverses from 2 to K, the first device determines, based on the P1 feature sets, Pi processing results among the Mi×Ni processing results that correspond one-to-one and are related to the P1 feature sets, and sends the Pi processing results to the second device.
[0018] Optionally, the first device determines, based on P1 feature sets, Pi processing results among Mi×Ni processing results that correspond one-to-one and are related to the P1 feature sets, and sends the Pi processing results to the second device, including: the first device divides the Mi×Ni processing results into Pi feature sets according to the positions of the grids, the position of the grid corresponding to the jth feature set among the Pi feature sets in the non-key frame image of the i-th frame is the same as the position of the grid corresponding to the jth feature set among the P1 feature sets in the key frame image, Pi=P1, j is any integer from 1 to Pi, the grid corresponding to the feature set means that the bit sequence in the feature set is obtained by processing the grid corresponding to the feature set; the first device determines a bit sequence from the jth feature set of the Pi feature sets, and when j traverses from 1 to Pi, Pi bit sequences are obtained, and the Pi bit sequences are the Pi processing results; the first device sends the Pi bit sequences to the second device.
[0019] Optionally, the first device determines a bit sequence from the jth feature set of Pi feature sets, and obtains Pi bit sequences when j traverses 1 to Pi, including: when i=2, the first device determines a bit sequence with the highest correlation with the P1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi; when i>2, the first device determines a bit sequence with the second highest correlation with the Pi-1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi, and the Pi-1 feature set corresponds to the i-1th non-key frame image in the K-1 non-key frame images.
[0020] Optionally, the first device sends Pi bit sequences to the second device, including: for the j-th bit sequence among the Pi bit sequences; the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, and sends the flipped j-th bit sequence to the second device, and when j traverses from 1 to Pi, a total of Pi flipped bit sequences are sent.
[0021] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device fills the j-th bit sequence with a bit sequence of length L×L, and converts the bit sequence of length L×L into a matrix with L×L number of rows and columns; L is an odd number greater than or equal to 3; for any bit on the diagonal of the matrix, the first device flips the value of any bit from 1 to 0, or flips the value of any bit from 0 to 1, to obtain a flipped matrix, the bits on the diagonal of the matrix are structural key bits; the first device restores the flipped matrix to a sequence to obtain the flipped j-th bit sequence.
[0022] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device divides the j-th bit sequence into multiple sub-bit sequences, and the lengths of the multiple sub-bit sequences are incremented from 1 to 2; for any sub-bit sequence in the multiple sub-bit sequences, the first device flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 1 to 0, or flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 0 to 1, to obtain multiple flipped sub-bit sequences, and the bits located at both ends of the column are structural key bits; the first device splices the multiple flipped sub-bit sequences to obtain the flipped j-th bit sequence.
[0023] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having program code stored thereon. When the program code is executed by the computer, the method described in the first aspect is executed.
[0024] In summary, the above method and device have the following technical effects: Since the contents of consecutive K frame images usually have continuity and correlation, the first frame image in the K frame image can be regarded as a key frame image, and the K-1 frame image after the first frame image in the K frame image can be regarded as a K-1 frame non-key frame image, where K is an integer greater than 2; at this time, when the K frame image is processed by a neural network model, the first device sends the processing result of the key frame image to the second device, and only sends part of the processing results of each of the K-1 frame non-key frame images. At this time, the device on the opposite side can predict or restore the content of the image corresponding to the unsent processing result based on the correlation between the partial processing result and the processing result of the key frame image, thereby improving the image processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A schematic diagram of the architecture of an image processing system provided in an embodiment of the present application; Figure 2A flowchart of an image processing method based on key frame analysis provided in an embodiment of the present application; Figure 3 The application scenario of the method provided in the embodiment of the present application is shown as follows: Figure 1 ; Figure 4 The application scenario of the method provided in the embodiment of the present application is shown as follows: Figure 2 ; Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] In the embodiment of the present invention, "indication" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. The information indicated by a certain information is called information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as but not limited to, the information to be indicated can be directly indicated, such as the information to be indicated itself or the index of the information to be indicated. The information to be indicated can also be indirectly indicated by indicating other information, wherein there is an association relationship between the other information and the information to be indicated. It is also possible to indicate only a part of the information to be indicated, while the other parts of the information to be indicated are known or agreed in advance. For example, the indication of specific information can also be achieved by means of the arrangement order of each information agreed in advance (for example, specified by the protocol), thereby reducing the indication overhead to a certain extent. At the same time, the common parts of each information can also be identified and indicated uniformly to reduce the indication overhead caused by indicating the same information separately.
[0027] In addition, the specific indication method may also be various existing indication methods, such as but not limited to the above-mentioned indication methods and various combinations thereof. The specific details of the various indication methods can refer to the prior art and will not be repeated herein. As can be seen from the above, for example, when it is necessary to indicate multiple information of the same type, different indication methods may be used for different information. In the specific implementation process, the desired indication method can be selected according to specific needs. The embodiment of the present invention does not limit the selected indication method. In this way, the indication method involved in the embodiment of the present invention should be understood to cover various methods that can enable the party to be indicated to obtain the information to be indicated.
[0028] It should be understood that the information to be indicated can be sent as a whole or divided into multiple sub-information and sent separately, and the sending period and / or sending time of these sub-information can be the same or different. The specific sending method is not limited in the embodiment of the present invention. Among them, the sending period and / or sending time of these sub-information can be pre-defined, for example, pre-defined according to a protocol, or can be configured by the sending end device by sending configuration information to the receiving end device.
[0029] "Pre-definition" or "pre-configuration" can be implemented by pre-saving corresponding codes, tables or other methods that can be used to indicate relevant information in the device, and the embodiments of the present invention do not limit the specific implementation method. Among them, "saving" can mean saving in one or more memories. The one or more memories can be set separately or integrated in a processor or decoder, a processor, or an electronic device. The one or more memories can also be partially set separately and partially integrated in a decoder, a processor, or an electronic device. The type of memory can be any form of storage medium, which is not limited by the embodiments of the present invention.
[0030] The "protocol" involved in the embodiments of the present invention may refer to a protocol family in the communication field, a standard protocol with a similar protocol family frame structure, or a related protocol in a reliable access method system for future Internet of Things devices, and the embodiments of the present invention do not specifically limit this.
[0031] In the embodiments of the present invention, descriptions such as "when...", "in the case of...", "if" and "if" all mean that the device will perform corresponding processing under certain objective circumstances, but do not limit the time, nor do they require the device to perform judgment actions when implementing, nor do they mean the existence of other limitations.
[0032] In the description of the embodiments of the present invention, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in the embodiments of the present invention is only a kind of association relationship describing the associated objects, indicating that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. In addition, in the description of the embodiments of the present invention, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple. In addition, in order to facilitate the clear description of the technical solution of the embodiments of the present invention, in the embodiments of the present invention, the words "first" and "second" are used to distinguish the same or similar items with basically the same functions and effects. Those skilled in the art will appreciate that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit the difference. At the same time, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.
[0033] The network architecture and business scenarios described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. A person of ordinary skill in the art can appreciate that with the evolution of the network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.
[0034] The technical solution in this application will be described below in conjunction with the accompanying drawings.
[0035] See also Figure 1 An embodiment of the present application provides an image processing system, which may include a first device and a second device.
[0036] Both the first device and the second device can be devices in the form of terminals, and the terminal can also be called user equipment (UE), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device. The terminal device in the embodiment of the present application can be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc.
[0037] The interaction between the first device and the second device in the above system will be described in detail below in conjunction with the method.
[0038] See also Figure 2 The present application embodiment provides an image processing method based on key frame analysis, and the process of the method is as follows: S201, the first device obtains K consecutive image frames.
[0039] K is an integer greater than 2. For example, if the K value is too large, the continuity of the image content will be affected. Therefore, the value of K can be 5-10.
[0040] The first frame image in the K frame image is a key frame image. The K-1 frame image after the first frame image in the K frame image is a K-1 frame non-key frame image. The key frame image can transmit all the processing content, while the non-key frame image can only transmit part of the processing content, so as to rely on the full processing content of the key frame image to predict and restore the processing content that is not transmitted by the non-key frame image.
[0041] The embodiments of the present application can be applied to ocean scenes. For example, the K-frame image can be an image of waves taken of the ocean, which is used to analyze sea ripples in actual application scenarios, such as determining the current sea conditions and presetting future sea conditions. The above application scenarios are only some examples and are not specifically limited.
[0042] The embodiment of the present application does not limit the way in which the first device obtains continuous K frame images. For example, a video stream captured by a shooting device can be transmitted to the first device. The first device can segment the video stream with K frames, and each segment includes continuous K frame images. The continuous K frame images in the embodiment of the present application are continuous K frame images contained in any segment.
[0043] S202, the first device processes K frame images through a neural network model to obtain processing results of the key frame images and processing results of K-1 frame non-key frame images.
[0044] The neural network model may be a deep neural network model, such as a convolutional neural network model or a fast convolutional neural network model, and the specific model type is not limited.
[0045] The first device can process the key frame image through a neural network model to obtain a processing result of the key frame image, and the processing result of the key frame image includes a plurality of feature sets divided according to feature similarity.
[0046] For example, the first device can process with a grid as the granularity, such as the first device can divide the key frame image into M1×N1 grids, where M1 and N1 are both integers greater than 3, M1 is the number of grids in a row, and N1 is the number of grids in a column. The values of M1 and N1 can be selected according to actual conditions. For example, if the performance of the transmitter and the receiver are both good, and prediction and recovery can be performed using fewer grids, then the values of M1 and N1 can be relatively large to further reduce the transmission overhead. In other words, the better the performance of the device, the smaller the processing overhead, such as M1=N1=20. If the performance of the transmitter and the receiver is both poor, the values of M1 and N1 can be relatively small, such as M1=N1=8, to ensure performance. The first device can process M1×N1 grids in sequence through a neural network model to obtain M1×N1 processing results, each of which includes a bit sequence, that is, the first device can input each grid into the neural network model separately, and the neural network model convolves the grid to obtain a feature vector, and then processes the feature vector through a CABAC algorithm and an arithmetic algorithm to obtain a binary bit sequence. Thus, for M1×N1 grids, M1×N1 bit sequences, that is, M1×N1 processing results, can be obtained one by one. The first device can divide the M1×N1 processing results into P1 feature sets by calculating the correlation of the bit sequences between the M1×N1 processing results, where P1 is an integer greater than 1 and less than M1×N1, and the processing results belonging to the same feature set in the M1×N1 processing results are bit sequence-related processing results, and the processing results belonging to different feature sets in the M1×N1 processing results are bit sequence-independent processing results. For example, the first device can determine that two processing results with a correlation greater than a preset threshold value belong to the same feature set by calculating the correlation between each two processing results among the M1×N1 processing results, such as the Euclidean distance or the Manhattan distance between each two bit sequences. In this way, the processing results of the same feature set are reflected in the grids, that is, the image contents contained in the grids are relatively similar, while the processing results of different feature sets are reflected in the grids, that is, the image contents contained in the grids are dissimilar.
[0047] For example, Figure 3 As shown, through an example, for the key frame image, M1=N1=4, the 6 processing results corresponding to the 6 grids of pattern 1 are all processing results with correlations greater than the preset threshold, belonging to feature set #1. The 8 processing results corresponding to the 8 grids of pattern 2 are all processing results with correlations greater than the preset threshold, belonging to feature set #2. The 2 processing results corresponding to the 2 grids of pattern 3 are all processing results with correlations greater than the preset threshold, belonging to feature set #3, that is, P1=3.
[0048] The first device can also process K-1 frames of non-key frame images separately through a neural network model to obtain processing results of the K-1 frames of non-key frame images respectively, and the processing results of the K-1 frames of non-key frame images respectively include multiple features of the K-1 frames of non-key frame images respectively.
[0049] For example, for the i-th non-key frame image in the K-1 frames of non-key frame images, i is any integer from 2 to K: the first device divides the i-th non-key frame image into Mi×Ni grids, Mi=M1, Ni=N1, Mi is the number of grids in a row, and Ni is the number of grids in a column; the first device then processes the Mi×Ni grids in sequence through the neural network model to obtain Mi×Ni processing results, each of the Mi×Ni processing results includes a bit sequence, that is, the way of dividing the grids is consistent with the key frame image, and the processing method is also consistent. The difference is that, at this time, for the Mi×Ni processing results, the first device does not divide the feature set according to correlation.
[0050] S203: The first device sends the processing result of the key frame image and part of the processing results of the K-1 frame non-key frame image to the second device.
[0051] The above-mentioned partial processing results are determined according to the processing results of the key frame images. The partial processing results and the processing results of the key frame images can be used to restore the K-1 frame non-key frame image, which is described in detail below.
[0052] Based on S202, it can be known that the first device sends P1 feature sets to the second device.
[0053] Also, when i traverses 2 to K, the first device can also determine, based on the P1 feature set, Pi processing results among the Mi×Ni processing results that correspond one-to-one and are related to the P1 feature set, and send the Pi processing results to the second device, as described in detail below.
[0054] For example, the first device can divide Mi×Ni processing results into Pi feature sets according to the position of the grid, the position of the grid corresponding to the jth feature set in the Pi feature sets in the i-th non-key frame image is the same as the position of the grid corresponding to the jth feature set in the P1 feature sets in the key frame image, Pi=P1, j is any integer from 1 to Pi. The grid corresponding to the feature set means that the bit sequence in the feature set is obtained by processing the grid corresponding to the feature set. It can be seen that the above is to determine the processing results of the grids at the same position in the images of different frames as belonging to the same feature set. For example, if the processing results of a part of the grids in the M1×N1 grids are divided into the same feature set, then the processing results of the grids in the Mi×Ni grids of the non-key frame image that are consistent with the position of the part of the grids are also divided into the same feature set, such as the jth feature set.
[0055] On this basis, the first device can determine a bit sequence from the jth feature set of the Pi feature sets, and when j traverses from 1 to Pi, Pi bit sequences are obtained, and the Pi bit sequences are Pi processing results; Specifically, when i=2, the first device determines a bit sequence in the jth feature set of Pi feature sets that has the highest correlation with the P1 feature set by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the P1 feature set, and obtains Pi bit sequences when j traverses from 1 to Pi. In other words, the first device can specifically calculate the correlation between each bit sequence in the jth feature set and each bit sequence in the P1 feature set, such as the above-mentioned Euclidean distance or Manhattan distance, to obtain a total of multiple correlations, and select the bit sequence corresponding to the jth feature set with the highest correlation. In other words, for the second frame image in the K frame image, or the first frame non-key frame image, the first device needs to select the grid that is most similar to the key frame image to ensure the accuracy of recovery or prediction during subsequent decoding.
[0056] In the case of i>2, the first device determines a bit sequence in the jth feature set of Pi feature sets that has the second highest correlation with the Pi-1 feature set by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses from 1 to Pi, and the Pi-1 feature set corresponds to the i-1th non-key frame image in the K-1 non-key frame images. In other words, the first device can specifically calculate the correlation between each bit sequence in the jth feature set of the next non-key frame image and each bit sequence in the previous non-key frame image, such as the above-mentioned Euclidean distance or Manhattan distance, to obtain a total of multiple correlations, and select the bit sequence corresponding to the second highest correlation in the jth feature set. In other words, for the image of the third frame or later frames in the K frame image, or the non-key frame image of the second frame or later frames, considering that the image is gradually changing, the first device needs to select a grid that is most similar to the previous frame image, so that when performing recovery or prediction in decoding, the gradual change of the image content can be taken into account more, thereby achieving better recovery or prediction effects.
[0057] For ease of understanding, continue with the above Figure 3 In the example shown, for the first non-key frame image, i.e., i=2, M2=N2=4, the six processing results corresponding to the six grids of pattern 4 are all processing results with a correlation greater than a preset threshold, belonging to feature set #4, and the positions of the six grids of pattern 4 are the same as the positions of the six grids of pattern 1. The eight processing results corresponding to the eight grids of pattern 5 are all processing results with a correlation greater than a preset threshold, belonging to feature set #5, and the positions of the eight grids of pattern 5 are the same as the positions of the eight grids of pattern 2. The two processing results corresponding to the two grids of pattern 6 are all processing results with a correlation greater than a preset threshold, belonging to feature set #6, and the positions of the two grids of pattern 6 are the same as the positions of the two grids of pattern 3. At this time, P2=3. Similarly, for the second non-key frame image, i.e., i=3, M3=N3=4, the 6 processing results corresponding to the 6 grids of pattern 7 are all processing results with a correlation greater than the preset threshold, belonging to feature set #7, and the positions of the 6 grids of pattern 7 are the same as the positions of the 6 grids of pattern 1. The 8 processing results corresponding to the 8 grids of pattern 8 are all processing results with a correlation greater than the preset threshold, belonging to feature set #8, and the positions of the 8 grids of pattern 8 are the same as the positions of the 8 grids of pattern 2. The 2 processing results corresponding to the 2 grids of pattern 9 are all processing results with a correlation greater than the preset threshold, belonging to feature set #9, and the positions of the 2 grids of pattern 9 are the same as the positions of the 2 grids of pattern 3. At this time, P3=3.
[0058] On this basis, for feature set #4, the first processing result has the highest correlation with the second processing result in feature set #1, so the first processing result in feature set #4 needs to be sent to the second device, and the remaining 5 processing results in feature set #4 do not need to be sent to the second device. In addition, the implementation of feature set #5 and feature set #6 is similar to feature set #4 and will not be repeated. For feature set #7, the third processing result has the second highest correlation with the fifth processing result in feature set #4, so the third processing result in feature set #7 needs to be sent to the second device, and the remaining 5 processing results in feature set #7 do not need to be sent to the second device. In addition, the implementation of feature set #8 and feature set #9 is similar to feature set #7 and will not be repeated.
[0059] Finally, the first device sends Pi bit sequences to the second device.
[0060] For example, for the jth bit sequence among the Pi bit sequences; The first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, and sends the flipped j-th bit sequence to the second device. When j traverses from 1 to Pi, a total of Pi flipped bit sequences are sent.
[0061] In one possible manner, the first device fills the j-th bit sequence with a bit sequence of length L×L, and converts the bit sequence of length L×L into a matrix with L×L number of rows and columns; L is an odd number greater than or equal to 3; for any bit on the diagonal of the matrix, the first device flips the value of any bit from 1 to 0, or flips the value of any bit from 0 to 1, to obtain a flipped matrix, and the bits on the diagonal of the matrix are structural key bits; the first device restores the flipped matrix to a sequence to obtain the flipped j-th bit sequence.
[0062] like Figure 4 As shown, through an example, assuming that the length of the j-th bit sequence is 33, the length of each bit sequence is fixed, such as 110011101010001001000111110110011, the first device can fill 3 bits with a value of 0 at the end of the j-th bit sequence to obtain 110011101010001001000111110110011000. At this time, it can be divided into a 6×6 matrix, and the diagonal elements are 101110 and 110010, which are 010001 and 001101 after flipping.
[0063] Alternatively, in another possible manner, the first device divides the j-th bit sequence into multiple sub-bit sequences, and the lengths of the multiple sub-bit sequences are incremented from 1 to 2 in sequence; for any sub-bit sequence among the multiple sub-bit sequences, the first device flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 1 to 0, or flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 0 to 1, to obtain multiple flipped sub-bit sequences, where the bits located at both ends of the column are structural key bits; the first device splices the multiple flipped sub-bit sequences to obtain the j-th bit sequence after flipping.
[0064] For ease of understanding, an example is also used to introduce that, assuming that the length of the j-th bit sequence is 33, the length of each bit sequence is fixed, such as 110011101010001001000111110110011, the first device can fill 3 bits with the value of 0 at the end of the j-th bit sequence to obtain 11001110101000 1001000111110110011000. At this time, it can be divided into sub-bit sequences with lengths of 6, 8, 10, and 12, respectively, namely 110011, 10101000, 1001000111, and 1001000111, respectively. After flipping, they are 010010, 00101001, 0001000110, and 0001000110, respectively. Of course, the above is based on the example of flipping one bit at each end, and it can also be multiple consecutive bits.
[0065] It can be seen that the above method is to flip some bits in the j-th bit sequence, so that the j-th bit sequence is de-informed. Since the positions of these flipped bits are structured, such as being located on the diagonal of the matrix or at both ends of the sequence, the processing end can perform the inverse process of the above method after receiving the flipped j-th bit sequence, thereby restoring the flipped bits. However, if the above flipped j-th bit sequence is stolen by an attacker, the attacker cannot restore the flipped bits because he cannot know the specific processing process. At this time, the flipped j-th bit sequence is invalid information for the attacker, so the data security of the image data can be improved.
[0066] For the second device, after receiving part of the processing results of the respective processing results of the K-1 frames of non-key frame images, the second device can perform the inverse process of the above-mentioned S203, that is, flip the corresponding bits according to the above-mentioned rules. Then, the second device inputs the processing results of the key frame image and the part of the processing results of the K frames of non-key frame images into a neural network model of the second device (such as neural network model #1), and the neural network model #1 performs the inverse process of the above-mentioned convolution, that is, deconvolution, to obtain the grids of the key frame image and the K frames of non-key frame images. At this time, the second device can input the grids of the key frame image and the second frame of non-key frame image, as well as information indicating the area corresponding to each grid of the second frame of non-key frame image in the key frame image, into another neural network model of the second device (such as neural network model #2). The corresponding area in the key frame image, that is, the area of the grid corresponding to each of the above-mentioned feature sets in the key frame image, as Figure 3 In this way, the neural network model #2 can predict the image content of other parts of the non-key frame image in the region outside the grid according to the image content in the region and the grid of the non-key frame image in the second frame. For example, Figure 3 As shown, neural network model #2 uses the image content of the area occupied by pattern 1 and the image content of a grid in pattern 4 to predict and restore the image content of the area occupied by the grid of the entire pattern 4, that is, to achieve image preset restoration with region as the granularity, with higher precision and accuracy. Afterwards, the second device can input the grid of the key frame image and the third non-key frame image, as well as the information indicating the area corresponding to each grid of the third non-key frame image in the key frame image into neural network model #2, and then the same logic is applied, and so on.
[0067] In summary, since the contents of consecutive K frame images usually have continuity and correlation, the first frame image in the K frame image can be regarded as a key frame image, and the K-1 frame image after the first frame image in the K frame image can be regarded as a K-1 frame non-key frame image, where K is an integer greater than 2; at this time, when the K frame image is processed by a neural network model, the first device sends the processing result of the key frame image to the second device, and only sends part of the processing results of each of the K-1 frame non-key frame images. At this time, the device on the opposite side can predict or restore the content of the image corresponding to the unsent processing result based on the correlation between the partial processing result and the processing result of the key frame image, thereby improving the image processing efficiency.
[0068] Combination of the above Figure 3The method provided by the embodiment of the present application is described in detail. The following introduces an image processing device based on key frame analysis for executing the method provided by the embodiment of the present application, which is applied to a first device, and the first device is configured as follows: the first device obtains a continuous K frame image, the first frame image in the K frame image is a key frame image, the K-1 frame image after the first frame image in the K frame image is a K-1 frame non-key frame image, and K is an integer greater than 2; the first device processes the K frame image through a neural network model to obtain the processing result of the key frame image and the processing result of each K-1 frame non-key frame image; the first device sends the processing result of the key frame image and part of the processing result of each K-1 frame non-key frame image to the second device, and the part of the processing result is determined according to the processing result of the key frame image, and the processing result of the key frame image and the part of the processing result are used to restore the K-1 frame non-key frame image.
[0069] Optionally, the first device processes K frame images through a neural network model to obtain processing results of the key frame images and processing results of K-1 frame non-key frame images, including: the first device processes the key frame images through a neural network model to obtain processing results of the key frame images, and the processing results of the key frame images include multiple feature sets divided according to feature similarity; the first device processes K-1 frame non-key frame images respectively through a neural network model to obtain processing results of the K-1 frame non-key frame images, and the processing results of the K-1 frame non-key frame images include multiple features of the K-1 frame non-key frame images.
[0070] Optionally, the first device processes the key frame image through a neural network model to obtain a processing result of the key frame image, including: the first device divides the key frame image into M1×N1 grids, M1 and N1 are both integers greater than 3, M1 is the number of grids in a row, and N1 is the number of grids in a column; the first device processes the M1×N1 grids in sequence through the neural network model to obtain M1×N1 processing results, each of the M1×N1 processing results including a bit sequence; the first device divides the M1×N1 processing results into P1 feature sets by calculating the correlation of the bit sequences between the M1×N1 processing results, P1 is an integer greater than 1 and less than M1×N1, and the processing results belonging to the same feature set among the M1×N1 processing results are bit sequences. The processing results are sequence-related, and the processing results belonging to different feature sets among the M1×N1 processing results are bit sequence-independent processing results; accordingly, the first device processes the K-1 frames of non-key frame images respectively through the neural network model to obtain the processing results of the K-1 frames of non-key frame images, including: for the i-th frame of non-key frame image in the K-1 frames of non-key frame images, i is any integer from 2 to K: the first device divides the i-th frame of non-key frame image into Mi×Ni grids, Mi=M1, Ni=N1, Mi is the number of grids in a row, and Ni is the number of grids in a column; the first device processes the Mi×Ni grids in sequence through the neural network model to obtain Mi×Ni processing results, and each of the Mi×Ni processing results includes a bit sequence.
[0071] Optionally, the first device sends the processing results of the key frame image and part of the processing results of the K-1 frame non-key frame images to the second device, including: the first device sends P1 feature sets to the second device; and, when i traverses from 2 to K, the first device determines, based on the P1 feature sets, Pi processing results among the Mi×Ni processing results that correspond one-to-one and are related to the P1 feature sets, and sends the Pi processing results to the second device.
[0072] Optionally, the first device determines, based on P1 feature sets, Pi processing results among Mi×Ni processing results that correspond one-to-one and are related to the P1 feature sets, and sends the Pi processing results to the second device, including: the first device divides the Mi×Ni processing results into Pi feature sets according to the positions of the grids, the position of the grid corresponding to the jth feature set among the Pi feature sets in the non-key frame image of the i-th frame is the same as the position of the grid corresponding to the jth feature set among the P1 feature sets in the key frame image, Pi=P1, j is any integer from 1 to Pi, the grid corresponding to the feature set means that the bit sequence in the feature set is obtained by processing the grid corresponding to the feature set; the first device determines a bit sequence from the jth feature set of the Pi feature sets, and when j traverses from 1 to Pi, Pi bit sequences are obtained, and the Pi bit sequences are the Pi processing results; the first device sends the Pi bit sequences to the second device.
[0073] Optionally, the first device determines a bit sequence from the jth feature set of Pi feature sets, and obtains Pi bit sequences when j traverses 1 to Pi, including: when i=2, the first device determines a bit sequence with the highest correlation with the P1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi; when i>2, the first device determines a bit sequence with the second highest correlation with the Pi-1 feature set in the jth feature set of Pi feature sets by successively calculating the correlation between each bit sequence in the jth feature set of Pi feature sets and the Pi-1 feature set, and obtains Pi bit sequences when j traverses 1 to Pi, and the Pi-1 feature set corresponds to the i-1th non-key frame image in the K-1 non-key frame images.
[0074] Optionally, the first device sends Pi bit sequences to the second device, including: for the j-th bit sequence among the Pi bit sequences; the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, and sends the flipped j-th bit sequence to the second device, and when j traverses from 1 to Pi, a total of Pi flipped bit sequences are sent.
[0075] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device fills the j-th bit sequence with a bit sequence of length L×L, and converts the bit sequence of length L×L into a matrix with L×L number of rows and columns; L is an odd number greater than or equal to 3; for any bit on the diagonal of the matrix, the first device flips the value of any bit from 1 to 0, or flips the value of any bit from 0 to 1, to obtain a flipped matrix, the bits on the diagonal of the matrix are structural key bits; the first device restores the flipped matrix to a sequence to obtain the flipped j-th bit sequence.
[0076] Optionally, the first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: the first device divides the j-th bit sequence into multiple sub-bit sequences, and the lengths of the multiple sub-bit sequences are incremented from 1 to 2; for any sub-bit sequence in the multiple sub-bit sequences, the first device flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 1 to 0, or flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 0 to 1, to obtain multiple flipped sub-bit sequences, and the bits located at both ends of the column are structural key bits; the first device splices the multiple flipped sub-bit sequences to obtain the flipped j-th bit sequence.
[0077] Combine the following Figure 5 The components of the electronic device 500 are described in detail: The processor 501 is the control center of the electronic device 500, and may be a processor or a general term for multiple processing elements. For example, the processor 501 is one or more central processing units (CPUs), or may be application specific integrated circuits (ASICs), or may be one or more integrated circuits configured to implement the embodiments of the present application, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (FPGAs).
[0078] Optionally, the processor 501 can execute various functions of the electronic device 500 by running or executing the software program stored in the memory 502 and calling the data stored in the memory 502, such as the above Figure 2 Functionality in the method shown.
[0079] In a specific implementation, as an embodiment, the processor 501 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 are shown in FIG.
[0080] In a specific implementation, as an embodiment, the electronic device 500 may also include multiple processors. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0081] Among them, the memory 502 is used to store the software program for executing the solution of the present application, and the execution is controlled by the processor 501. The specific implementation method can refer to the above method embodiment and will not be repeated here.
[0082] Optionally, the memory 502 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 502 may be integrated with the processor 501, or may exist independently, and the interface circuit ( Figure 5 (not shown) is coupled to the processor 501, which is not specifically limited in this embodiment of the present application.
[0083] The transceiver 503 is used for communication with other devices. For example, if the multi-beam based positioning device is a terminal, the transceiver 503 can be used for communication with a network device or another terminal.
[0084] Optionally, the transceiver 503 may include a receiver and a transmitter ( Figure 5 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.
[0085] Optionally, the transceiver 503 may be integrated with the processor 501, or may exist independently and communicate with the processor 501 through the interface circuit ( Figure 5 (not shown) is coupled to the processor 501, which is not specifically limited in this embodiment of the present application.
[0086] It should be noted that Figure 5 The structure of the electronic device 500 shown in the figure does not constitute a limitation on the device, and the actual electronic device 500 may include more or less components than those shown in the figure, or combine certain components, or arrange the components differently.
[0087] In addition, the technical effects based on the electronic device 500 can refer to the technical effects of the method in the above method embodiment, which will not be repeated here.
[0088] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0089] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0090] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When a computer instruction or computer program is loaded or executed on a computer, a process or function according to an embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0091] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0092] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0093] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0094] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0095] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0096] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some feature fields can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0097] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0098] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0099] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0100] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An image processing method based on key frame analysis, characterized in that: Applied to a first device, the method comprises: The first device acquires K consecutive frames of images, wherein the first frame of the K frames of images is a key frame image, and the K-1 frames of images after the first frame of the K frames of images are K-1 non-key frame images, where K is an integer greater than 2; The first device processes the K frame images through a neural network model to obtain processing results of the key frame images and processing results of the K-1 frame non-key frame images; The first device sends the processing results of the key frame image and part of the processing results of the K-1 frame non-key frame image to the second device, and the part of the processing results is determined based on the processing results of the key frame image. The processing results of the key frame image and the part of the processing results are used to restore the K-1 frame non-key frame image.
2. The method according to claim 1, characterized in that The first device processes the K frame images through a neural network model to obtain processing results of the key frame images and processing results of the K-1 frame non-key frame images, including: The first device processes the key frame image through the neural network model to obtain a processing result of the key frame image, wherein the processing result of the key frame image includes a plurality of feature sets divided according to feature similarity; The first device processes the K-1 frame non-key frame images respectively through the neural network model to obtain processing results of the K-1 frame non-key frame images respectively, and the processing results of the K-1 frame non-key frame images respectively include multiple features of the K-1 frame non-key frame images respectively.
3. The method according to claim 2, characterized in that The first device processes the key frame image through the neural network model to obtain a processing result of the key frame image, including: The first device divides the key frame image into M1×N1 grids, where M1 and N1 are both integers greater than 3, M1 is the number of grids in a row, and N1 is the number of grids in a column; The first device processes the M1×N1 grids in sequence through the neural network model to obtain M1×N1 processing results, each of the M1×N1 processing results including a bit sequence; The first device divides the M1×N1 processing results into P1 feature sets by calculating the correlation of the bit sequences between the M1×N1 processing results, where P1 is an integer greater than 1 and less than M1×N1, and the processing results belonging to the same feature set among the M1×N1 processing results are bit sequence-related processing results, and the processing results belonging to different feature sets among the M1×N1 processing results are bit sequence-independent processing results; Correspondingly, the first device processes the K-1 frames of non-key frame images respectively through the neural network model to obtain respective processing results of the K-1 frames of non-key frame images, including: For the i-th non-key frame image in the K-1 non-key frame images, i is any integer from 2 to K: The first device divides the i-th non-key frame image into Mi×Ni grids, Mi=M1, Ni=N1, Mi is the number of grids in a row, Ni is the number of grids in a column; The first device processes the Mi×Ni grids in sequence through the neural network model to obtain Mi×Ni processing results, each of which includes a bit sequence.
4. The method according to claim 3, characterized in that The first device sends a processing result of the key frame image and a part of the processing results of the K-1 frame non-key frame image to the second device, including: The first device sends the P1 feature sets to the second device; And, when i traverses 2 to K, the first device determines, based on the P1 feature sets, Pi processing results among the Mi×Ni processing results that correspond one-to-one to and are related to the P1 feature sets, and sends the Pi processing results to the second device.
5. The method according to claim 4, characterized in that The first device determines, according to the P1 feature sets, Pi processing results among the Mi×Ni processing results that correspond one-to-one and are related to the P1 feature sets, and sends the Pi processing results to the second device, including: The first device divides the Mi×Ni processing results into Pi feature sets according to the positions of the grids, the position of the grid corresponding to the jth feature set in the Pi feature sets in the i-th non-key frame image is the same as the position of the grid corresponding to the jth feature set in the P1 feature sets in the key frame image, Pi=P1, j is any integer from 1 to Pi, and the grid corresponding to the feature set means that the bit sequence in the feature set is obtained by processing the grid corresponding to the feature set; The first device determines a bit sequence from the jth feature set of the Pi feature sets, and obtains Pi bit sequences when j traverses from 1 to Pi, and the Pi bit sequences are the Pi processing results; The first device sends the Pi bit sequences to the second device.
6. The method according to claim 5, characterized in that The first device determines a bit sequence from the jth feature set of the Pi feature sets, and obtains Pi bit sequences when j traverses from 1 to Pi, including: In the case where i=2, the first device determines a bit sequence in the jth feature set of the Pi feature sets that has the highest correlation with the P1 feature set by sequentially calculating the correlation between each bit sequence in the jth feature set of the Pi feature sets and the P1 feature set, and obtains the Pi bit sequences when j traverses from 1 to Pi; When i>2, the first device determines a bit sequence in the jth feature set of the Pi feature sets that has the second highest correlation with the Pi-1 feature set by successively calculating the correlation between each bit sequence in the jth feature set of the Pi feature sets and the Pi-1 feature set, and when j traverses from 1 to Pi, the Pi bit sequences are obtained, and the Pi-1 feature set corresponds to the i-1th non-key frame image in the K-1 non-key frame images.
7. The method according to claim 5 or 6, characterized in that: The first device sending the Pi bit sequences to the second device includes: For the j-th bit sequence among the Pi bit sequences; The first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, and sends the flipped j-th bit sequence to the second device. When j traverses from 1 to Pi, a total of Pi flipped bit sequences are sent.
8. The method according to claim 7, characterized in that The first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: The first device pads the j-th bit sequence into a bit sequence of length L×L, and converts the bit sequence of length L×L into a matrix of L×L rows and columns; L is an odd number greater than or equal to 3; For any bit on the diagonal of the matrix, the first device flips the value of any bit from 1 to 0, or flips the value of any bit from 0 to 1, to obtain a flipped matrix, where the bits on the diagonal of the matrix are the structural key bits; The first device restores the flipped matrix to a sequence to obtain the flipped j-th bit sequence.
9. The method according to claim 7, characterized in that: The first device flips the structural key bit in the j-th bit sequence to obtain the flipped j-th bit sequence, including: The first device divides the j-th bit sequence into a plurality of sub-bit sequences, and the lengths of the plurality of sub-bit sequences are sequentially increased from 1 to 2; For any sub-bit sequence among the multiple sub-bit sequences, the first device flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 1 to 0, or flips the values of the bits located at both ends of the sequence in any sub-bit sequence from 0 to 1, to obtain multiple flipped sub-bit sequences, where the bits located at both ends of the sub-bit sequence are the structural key bits; The first device concatenates the flipped multiple sub-bit sequences to obtain the flipped j-th bit sequence.
10. An image processing device based on key frame analysis, characterized in that: Applied to a first device, the first device is configured to: The first device acquires K consecutive frames of images, wherein the first frame of the K frames of images is a key frame image, and the K-1 frames of images after the first frame of the K frames of images are K-1 non-key frame images, where K is an integer greater than 2; The first device processes the K frame images through a neural network model to obtain processing results of the key frame images and processing results of the K-1 frame non-key frame images; The first device sends the processing results of the key frame image and part of the processing results of the K-1 frame non-key frame image to the second device, and the part of the processing results is determined based on the processing results of the key frame image. The processing results of the key frame image and the part of the processing results are used to restore the K-1 frame non-key frame image.
Citation Information
Patent Citations
Image processing method and device
CN110189242A
Video semantic segmentation method and device
CN112465826A
Lane line detection method and device, electronic equipment, storage medium and vehicle
CN112560684A
Video coding and decoding method and device and computer equipment
CN113347421A
Traffic scene analysis method and device based on video stream
CN114898243A