Encoding method, decoding method, encoding apparatus, and decoding apparatus
Patent Information
- Application Number
- CN202211169034.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-20
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-09-20
AI Technical Summary
然而图像压缩方法不仅耗时,而且占用较大的带宽,传输较慢,影响机器识别的效率,因此特征域的压缩编码技术被提出
[0010]有益效果是:本申请对待处理对象的基础特征进行编码处理,得到第一码流,并获取与对第一码流进行解码的结果相同的预重建特征,进而对基于基础特征以及预重建特征而得到的残差特征进行编码处理,得到第二码流,使得第一码流携带的是待处理对象的主要特征信息,第二码流携带的是对待处理对象的基础特征进行编码处理、解码处理后的损失特征信息,因此编码设备将第一码流和第二码流发送给解码设备后,解码设备可以利用第二码流对第一码流修正处理,保证重建得到的目标重建特征的准确率,进而减少待处理对象在传输过程中的失真,保证待处理对象的传输质量,例如当编码对象为图像时,通过上述的方式,可以保证服务器在接收到摄像头抓拍的图像后,服务器最终重建的图像与摄像头抓拍的图像之间的误差非常小,进而提高后续基于重建的图像进行视觉任务的准确率。
Smart Images

Figure CN115941950B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of encoding and decoding technology, and in particular relates to an encoding method, a decoding method, an encoding device, and a decoding device. Background Technology
[0002] In traditional machine vision tasks, images captured by a camera are compressed into an image stream using JPEG and then transmitted to a server. The server then decodes the image stream and matches it with images in a database to obtain the recognition result. However, image compression methods are not only time-consuming but also consume a large amount of bandwidth, resulting in slow transmission and affecting the efficiency of machine recognition. Therefore, feature domain compression coding techniques have been proposed.
[0003] Feature domain compression coding technology mainly encodes, transmits, and decodes the extracted features before inputting them into the subsequent network to complete specific machine vision tasks. It can simultaneously reduce bandwidth usage and protect user information privacy. Therefore, research on feature domain compression coding technology is of great significance. Summary of the Invention
[0004] This application provides an encoding method, a decoding method, an encoding device, and a decoding device, which can improve the reconstruction accuracy of decoding and ensure the transmission quality of the object to be processed.
[0005] A first aspect of this application provides an encoding method, the encoding method comprising: acquiring basic features of an object to be processed; performing a first encoding process on the basic features to obtain a first bitstream; acquiring pre-reconstructed features, wherein the pre-reconstructed features are the same as the features obtained by decoding the first bitstream; determining residual features based on the basic features and the pre-reconstructed features; and performing a second encoding process on the residual features to obtain a second bitstream, wherein the target reconstruction features of the object to be processed are obtained by correcting the first bitstream using the second bitstream.
[0006] A second aspect of this application provides a decoding method, the decoding method comprising: receiving a first bitstream and a second bitstream of an object to be processed; the first bitstream and the second bitstream being obtained by the above-described encoding method; modifying the first bitstream using the second bitstream to obtain a modification result; and obtaining the target reconstruction features of the object to be processed based on the modification result.
[0007] A third aspect of this application provides an encoding device, which includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit, respectively. The memory stores program data, and the processor executes the program data in the memory to implement the steps in the above-described encoding method.
[0008] A fourth aspect of this application provides a decoding device, which includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit, respectively. The memory stores program data, and the processor executes the program data in the memory to implement the steps in the above-described decoding method.
[0009] A fifth aspect of this application provides a computer-readable storage medium storing a computer program that can be executed by a processor to implement the steps in the above-described method.
[0010] The beneficial effects are as follows: This application encodes the basic features of the object to be processed to obtain a first bitstream, and obtains the same pre-reconstructed features as the result of decoding the first bitstream. Then, it encodes the residual features obtained based on the basic features and the pre-reconstructed features to obtain a second bitstream. The first bitstream carries the main feature information of the object to be processed, and the second bitstream carries the loss feature information after encoding and decoding the basic features of the object to be processed. Therefore, after the encoding device sends the first bitstream and the second bitstream to the decoding device, the decoding device can use the second bitstream to correct the first bitstream, ensuring the accuracy of the reconstructed target features, thereby reducing the distortion of the object to be processed during transmission and ensuring the transmission quality of the object to be processed. For example, when the encoded object is an image, the above method can ensure that the error between the image finally reconstructed by the server and the image captured by the camera is very small after the server receives the image captured by the camera, thereby improving the accuracy of subsequent visual tasks based on the reconstructed image. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein: Figure 1 This is a flowchart illustrating one embodiment of the coding method of this application; Figure 2 yes Figure 1 A flowchart illustrating step S120; Figure 3 It corresponds Figure 1 Framework diagram of the coding method; Figure 4 yes Figure 2 A flowchart illustrating step S121; Figure 5 yes Figure 4 A flowchart illustrating step S1211 in an application scenario; Figure 6 It corresponds to an application scenario. Figure 2 A schematic diagram of the process of obtaining the first low-dimensional feature in step S121; Figure 7 This corresponds to another application scenario. Figure 2 A schematic diagram of the process of obtaining the first low-dimensional feature in step S121; Figure 8 Is Figure 2 A flowchart illustrating the process of obtaining quantization values before step S122; Figure 9 This is a diagram illustrating the process of a decoding device decoding the first bitstream. Figure 10 This is a schematic diagram of the process of performing comprehensive transformation on the first inverse quantization feature in an application scenario to obtain the pre-reconstructed feature. Figure 11 This is a schematic diagram of the process of performing comprehensive transformation on the first inverse quantization feature to obtain the pre-reconstructed feature in another application scenario; Figure 12 yes Figure 1 A flowchart illustrating step S130; Figure 13 This is a flowchart illustrating one embodiment of the decoding method of this application; Figure 14 This is a schematic diagram of a decoding device obtaining target reconstruction features in an application scenario; Figure 15 This is a schematic diagram of the decoding device obtaining target reconstruction features in another application scenario; Figure 16 This is an overall block diagram of the encoding and decoding methods corresponding to this application; Figure 17 This is a schematic diagram of one embodiment of the encoding device of this application; Figure 18 This is a schematic diagram of another embodiment of the encoding device of this application; Figure 19 This is a schematic diagram of one embodiment of the decoding device of this application; Figure 20 This is a schematic diagram of another embodiment of the decoding device of this application; Figure 21 This is a schematic diagram of one embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0013] It should be noted that the terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0014] First, it should be noted that the encoding method of this application is executed by the encoding device, and the decoding method is executed by the decoding device.
[0015] See Figure 1 , Figure 1 This is a flowchart illustrating one embodiment of the encoding method of this application, which includes: S110: Obtain the basic characteristics of the object to be processed.
[0016] Specifically, the object to be processed can be an image, video, text, or audio, etc. Feature extraction methods can be used to extract features from the object to be processed, obtaining its basic features. Any existing feature extraction method can be used to extract features from the object to be processed; this application does not impose any specific limitations.
[0017] The basic features of the object to be processed can be the original features of the object to be processed, or features after some basic processing of the original features.
[0018] S120: Perform the first encoding process on the basic features to obtain the first bitstream.
[0019] Specifically, since the first bitstream is obtained by performing a first encoding process on the basic features, the first bitstream carries the main feature information of the object to be processed.
[0020] Combination Figure 2In this embodiment, step S120 specifically includes: S121: Analyze and transform the basic features to obtain the first low-dimensional feature, wherein the dimension of the first low-dimensional feature is smaller than the dimension of the basic feature.
[0021] Specifically, the purpose of step S121 is to map the basic features from the high-dimensional space to the low-dimensional space in order to remove information redundancy and obtain compact first low-dimensional features.
[0022] Combination Figure 3 In this embodiment, the encoding device includes a first encoder. After receiving the basic features, the encoding device inputs the basic features into the first encoder and uses the first encoder to generate a first low-dimensional feature.
[0023] See Figure 4 In this embodiment, the process of the first encoder processing the basic features includes: S1211: Obtain the first feature.
[0024] The first feature is obtained based on the basic features, and the dimension of the first feature is smaller than the dimension of the basic features.
[0025] Since the first feature is obtained based on the basic features, the first feature can represent the global information of the basic features.
[0026] In one application scenario, refer to Figure 5 Step S1211 may specifically include: S12111: Perform fully connected processing on the basic features to obtain a second feature with the same dimensions as the basic features.
[0027] Specifically, the purpose of performing fully connected processing on the basic features is to enable the final first feature to better represent the global information of the object to be processed.
[0028] For example, see Figure 6 The 2048-dimensional basic feature (denoted by label 1) is processed by a fully connected layer to generate a new 2048-dimensional second feature (denoted by label 2).
[0029] S12112: Dimensionality reduction is performed on the second feature to obtain the first feature.
[0030] Continue reading Figure 6 The 2048-dimensional second feature is subjected to dimensionality reduction processing, such as fully connected processing or convolution processing, to obtain the 1024-dimensional first feature (denoted by label 3).
[0031] It should be noted that in other application scenarios, the basic features can also be directly reduced in dimensionality to obtain the first feature. For example, in... Figure 7 In application scenarios, the 2048-dimensional basic features (represented by label 1) are processed by, for example, fully connected processing or convolutional processing, to obtain the 1024-dimensional first feature (represented by label 3).
[0032] It should be noted that this application does not limit the specific process of obtaining the first feature.
[0033] S1212: Divide the first feature into multiple first sub-features.
[0034] The first feature is divided into multiple first sub-features according to a preset segmentation rule. For example, the first feature is divided into multiple first sub-features on an average basis to obtain multiple first sub-features with the same dimension, or the first feature is not divided into multiple first sub-features on an unequal basis to obtain multiple first sub-features with different dimensions.
[0035] For example, in Figure 6 In the application scenario, the first feature is divided into two 512-dimensional first sub-features (denoted by labels 31 and 32, respectively); while... Figure 7 In the application scenario, the first feature is divided into a first sub-feature with a dimension of 64 (represented by label 31) and another first sub-feature with a dimension of 960 (represented by label 32).
[0036] S1213: Perform dimensionality reduction processing on at least some of the first sub-features to obtain the corresponding first dimensionality-reduced sub-features.
[0037] Dimensionality reduction can be performed on only some of the first sub-features, or it can be performed on all of the first sub-features.
[0038] It is understandable that, since the dimensionality reduction is performed on at least some of the first sub-features separately, each first sub-feature will not participate in all the dimensionality reduction processes. Therefore, the resulting first dimensionality-reduced sub-features can better represent the local information of the object to be processed.
[0039] S1214: Determine the first low-dimensional feature based on the obtained first reduced-dimensional sub-feature and the first sub-feature that has not undergone dimensionality reduction processing.
[0040] The first low-dimensional feature can be constructed by directly using the obtained first dimensionality-reduced sub-feature and the first sub-feature without dimensionality reduction processing. In this case, the first low-dimensional feature includes the obtained first dimensionality-reduced sub-feature and the first sub-feature without dimensionality reduction processing.
[0041] Alternatively, depending on the requirements for the first low-dimensional feature dimension, the obtained first dimensionality-reduced sub-feature or the first sub-feature that has not undergone dimensionality reduction can be further processed for dimensionality reduction.
[0042] In one application scenario, when step S1213 performs dimensionality reduction processing on each first sub-feature to obtain the first dimensionality-reduced sub-feature corresponding to each first sub-feature, step S1214 specifically includes: dividing each first dimensionality-reduced sub-feature into multiple second sub-features; performing dimensionality reduction processing on each second sub-feature to obtain the corresponding second dimensionality-reduced sub-feature; and determining the first low-dimensional feature based on the obtained second dimensionality-reduced sub-features.
[0043] Specifically, after obtaining the first reduced-dimensionality sub-feature corresponding to each first sub-feature, each first reduced-dimensionality sub-feature is processed in the same way as the first feature. Specifically, each first reduced-dimensionality sub-feature is segmented to obtain the second reduced-dimensionality sub-feature corresponding to each first reduced-dimensionality sub-feature. Then, each second sub-feature is subjected to dimensionality reduction processing to obtain the second reduced-dimensionality sub-feature corresponding to each second sub-feature. Finally, the first low-dimensional feature is determined based on the obtained second reduced-dimensionality sub-feature.
[0044] The method of segmenting the first dimension-reduced sub-feature can be the same as or different from the method of segmenting the first feature. For example, although the first feature is divided into two first sub-features on average, when segmenting the first dimension-reduced sub-feature, it can be divided into two second sub-features on average, or it can be divided into three second sub-features on average, or it can be divided into multiple second sub-features with different dimensions.
[0045] To better understand, combine Figure 6 The application scenarios will be explained further: exist Figure 6 In the application scenario, after performing dimensionality reduction on each first sub-feature, two 256-dimensional first dimensionality-reduced sub-features are obtained (represented by label 4). Then, each first dimensionality-reduced sub-feature is divided equally to obtain four 128-dimensional second sub-features (represented by label 41). Next, each second sub-feature is dimensionality-reduced to obtain a 64-dimensional second dimensionality-reduced sub-feature corresponding to each second sub-feature (represented by label 5). At this point, the sum of the dimensions of all second dimensionality-reduced sub-features is 256.
[0046] When the target dimension of the first low-dimensional feature is 256, the four second low-dimensional features obtained at this time constitute the first low-dimensional feature, and subsequent processing can be performed on each second low-dimensional feature separately.
[0047] However, when the target dimension of the first low-dimensional feature is less than 256, the second dimensionality-reduced sub-feature can be obtained by further dividing it, and the dimensionality of the sub-feature obtained by the division can be further reduced. The previous steps are repeated until the dimension of the first low-dimensional feature obtained finally meets the preset requirements.
[0048] In the above application scenario, after each segmentation operation, the segmented features are reduced in dimensionality, and then the reduced features are segmented again, and the dimensionality is reduced again. This process of segmentation and dimensionality reduction is repeated until the feature dimension reaches the target dimension, and finally the first low-dimensional feature is obtained.
[0049] In another application scenario, when step S1213 performs dimensionality reduction processing on the first sub-features other than the first target sub-feature among the multiple first sub-features to obtain the corresponding first dimensionality-reduced sub-features, step S1214 specifically includes: dividing each first dimensionality-reduced sub-feature into multiple second sub-features; performing dimensionality reduction processing on the obtained multiple second sub-features other than the second target sub-feature to obtain the corresponding second dimensionality-reduced sub-features; and determining the first low-dimensional feature based on the first target sub-feature, the second target sub-feature, and the obtained second dimensionality-reduced sub-features.
[0050] Specifically, in this application scenario, after segmenting the first feature to obtain multiple first sub-features, a first target sub-feature is determined from among the multiple first sub-features according to preset rules (the number of first target sub-features can be one or more). The first target sub-feature is directly used as a part of the first low-dimensional feature. Then, for all first sub-features other than the first target sub-feature, dimensionality reduction processing is performed to obtain the corresponding first dimensionality-reduced sub-features. Then, each first dimensionality-reduced sub-feature is segmented into multiple second sub-features. Then, among the multiple second sub-features obtained, a second target sub-feature is determined (the number can be one or more). Similar to the first target sub-feature, the second target sub-feature is directly used as a part of the first low-dimensional feature. Then, dimensionality reduction processing is performed on the other second sub-features to obtain the corresponding second dimensionality-reduced sub-features. Then, referring to the above steps, the segmentation, dimensionality reduction and other processes are repeated until the dimension of the first low-dimensional feature reaches the target dimension.
[0051] To better understand, combine Figure 7 Application scenarios will be explained: exist Figure 7In the application scenario, we first assume that the target dimension of the first low-dimensional feature is 256. After dividing the first feature into a first sub-feature with a dimension of 64 (denoted by label 31) and another first sub-feature with a dimension of 960 (denoted by label 32), the first sub-feature with a dimension of 64 is determined as the first target sub-feature. Then, the first sub-feature with a dimension of 960 is subjected to dimensionality reduction processing to obtain a second dimensionality-reduced sub-feature with a dimension of 480 (denoted by label 4). The second dimensionality-reduced sub-feature is then divided into a second sub-feature with a dimension of 64 (denoted by label 41) and another second sub-feature with a dimension of 416 (denoted by label 42). Next, the second sub-feature with a dimension of 64 is determined as the second target sub-feature, and the first sub-feature with a dimension of 416 is further divided into two sub-features. The two sub-features are dimensionality reduced to obtain a second dimensionality-reduced sub-feature with a dimension of 208 (denoted by label 5). Then, the second dimensionality-reduced sub-feature is divided into a third sub-feature with a dimension of 64 (denoted by label 51) and another third sub-feature with a dimension of 144 (denoted by label 52). Next, the third sub-feature with a dimension of 64 is determined as the third target sub-feature, and the third sub-feature with a dimension of 144 is dimensionality-reduced to obtain a third dimensionality-reduced sub-feature with a dimension of 64 (denoted by label 6). Finally, the first low-dimensional feature is constructed using the first sub-feature with label 31, the second sub-feature with label 41, the third sub-feature with label 51, and the third dimensionality-reduced sub-feature with label 6. Subsequent processing can be performed on each sub-feature included in the first low-dimensional feature.
[0052] In other words, Figure 7 In the application scenario, each time one of the multiple sub-features obtained from the segmentation is directly used as part of the first low-dimensional feature, while the other sub-features are subjected to dimensionality reduction. Then, the result of the dimensionality reduction is segmented again, and one of the multiple sub-features obtained is directly used as part of the first low-dimensional feature, while the other sub-features are subjected to dimensionality reduction. This process is repeated until the dimension of the first low-dimensional feature reaches the target dimension.
[0053] Understandably, with Figure 6 Compared to application scenarios, Figure 7 It can make full use of hierarchical features.
[0054] The above describes the process of generating the first low-dimensional feature in this embodiment. It should be noted that in other embodiments, any method in the prior art can be used to generate the first low-dimensional feature based on the basic feature.
[0055] S122: Replace each feature value in the first low-dimensional feature with the corresponding quantized value to obtain the first quantized feature.
[0056] It should be noted that when the first low-dimensional feature includes multiple sub-features, steps S122 and S123 need to be executed separately for each sub-feature.
[0057] It is understood that before executing step S122, multiple quantization values need to be determined. These multiple quantization values can be preset according to actual needs, but in this embodiment, refer to... Figure 8 The steps for determining multiple quantization values include: S1221: Determine multiple binary encoding vectors based on a pre-set number of quantization bits.
[0058] In this context, the length of each binary encoded vector is equal to the number of quantization bits.
[0059] Specifically, the number of quantization bits is preset. If the preset number of quantization bits is K, it means that each feature value in the first low-dimensional feature needs to be represented by a binary encoded vector of length K. The length of the binary encoded vector is K, which means that the binary encoded vector includes K values. For ease of explanation, the following explanation will use the number of quantization bits as K.
[0060] Since the values in the binary encoded vectors are only 0 and 1, when the preset number of quantization bits is K, at most 2^K binary encoded vectors of length K can be used to represent the feature values in the first low-dimensional feature. That is, step S1221 determines at most 2^K (2 to the power of K) binary encoded vectors.
[0061] In order to reduce encoding loss, in this embodiment, when the number of quantization bits is K, step S1221 determines 2^K (2 to the power of K) binary encoding vectors of length K. The following description assumes that the number of binary encoding vectors is 2^K.
[0062] For example, when the preset number of quantization bits is 2, four binary encoding vectors of length 2 are determined, namely [0,0], [0,1], [1,0] and [1,1].
[0063] Of course, when the number of quantization bits is K, the number of binary encoded vectors determined in step S1221 can also be less than 2^K. For example, when the preset number of quantization bits is 2, three binary encoded vectors of length 2 are determined, namely [0,0], [0,1] and [1,0].
[0064] S1222: Determine multiple quantization values based on the operation values of the pre-learned basis vectors and each binary encoded vector.
[0065] Among them, multiple quantization values correspond one-to-one with multiple binary code vectors.
[0066] Specifically, the basis vectors are obtained through pre-learning, for example, by pre-training the first quantizer to obtain the basis vectors. Since the basis vectors can be obtained through learning, they can be continuously improved, thereby reducing quantization errors and further improving the accuracy of subsequent reconstructions.
[0067] For each of the binary encoded vectors determined in step S1221, there is a corresponding quantization value. For ease of explanation, the quantization value is denoted as p. i , where i = 1, 2, ..., 2^K.
[0068] Understandably, the purpose of step S1222 is to establish the correspondence between quantization values and binary encoding vectors. It should be noted that different binary encoding vectors correspond to different quantization values; that is, the multiple quantization values determined in step S1222 are different.
[0069] In this embodiment, step S1222 specifically includes: performing operations on the base vector and each binary encoded vector respectively to obtain multiple operation values, wherein the multiple operation values correspond one-to-one with the multiple binary encoded vectors; performing full connection processing on the multiple operation values to obtain multiple quantization values, wherein the multiple quantization values correspond one-to-one with the multiple operation values.
[0070] Specifically, the base vector is first operated on with each binary encoded vector to obtain the operation value corresponding to each binary encoded vector. That is, the multiple operation values correspond one-to-one with the multiple binary encoded vectors. Here, the operation value is denoted as q. i , where i = 1, 2, ..., 2^K.
[0071] Then, to improve the accuracy of subsequent reconstruction, the correlation between all subsequent quantized values is established. A fully connected layer is applied to all calculated values, resulting in multiple quantized values, at which point the quantized value p... i AND operation value q i It is a one-to-one correspondence, but due to the operation value q i There is a one-to-one correspondence between the quantization value p and the binary encoded vector. i There is a one-to-one correspondence between the binary encoded vector and the vector, thus generating the quantization value corresponding to each binary encoded vector.
[0072] It should be noted that, in other implementations, the operation value corresponding to the binary encoded vector can also be determined as the quantization value of the binary encoded vector, that is, the quantization value p in this case. i Equal to the operand q i .
[0073] It should be noted that the quantization value corresponding to the same binary encoded vector is the same during the encoding and decoding processes, as detailed below.
[0074] Since the basis vector and the binary encoded vector need to be operated on, the lengths of the basis vector and the binary encoded vector are equal. In other words, the length of the basis vector is equal to the pre-defined number of quantization bits. For ease of explanation, the basis vector is denoted as V, where V = [v1, v2, ..., v...]. K The binary encoded vector is denoted as B. i , i=1,2,…,2^K,B i =[b i1 b i2 , ..., b iK In one application scenario, B is determined according to the following formula. i The corresponding operand q i : q i =V·B i T = [v1, v2, ..., v K ]·[b i1 b i2 , ..., b iK ] T That is, the inner product of the basis vector and the binary encoded vector is determined as the operation value of the basis vector and the binary encoded vector.
[0075] It should be noted that this application does not limit how the base vector and the binary encoded vector are operated on. For example, in other embodiments, the base vector and the binary encoded vector can be vector-added, and then the average value of the values in the vector obtained by addition can be calculated. Finally, the average value is determined as the operation value corresponding to the binary encoded vector.
[0076] The above describes the process of determining multiple quantization values. The following describes the specific content of step S122: In this embodiment, step S122 specifically includes: obtaining multiple quantization values and obtaining the quantization interval corresponding to each quantization value; determining the quantization interval of each feature value in the first low-dimensional feature as the target quantization interval corresponding to each feature value; and replacing each feature value with the quantization value corresponding to the target quantization interval.
[0077] Specifically, the multiple quantization values obtained here can be determined in the manner described above, or they can be determined through other methods. When the quantization values are determined through other methods, the correspondence between the quantization values and the binary code vectors is predetermined.
[0078] For each quantized value, there is a corresponding quantization range. The quantization range corresponding to the quantized value can be set by the designer according to the requirements or determined according to preset rules. This application does not impose any specific restrictions.
[0079] Specifically, when determining the quantization interval corresponding to each quantization value, it is necessary to ensure that for any feature value in the first low-dimensional feature, it must fall within one of the quantization intervals. At the same time, it is also necessary to ensure that the quantization intervals corresponding to different quantization values do not overlap.
[0080] After obtaining the quantization interval, for each feature value, the quantization interval in which it lies is determined as the corresponding target quantization interval. Then, the feature value is replaced with the quantized value corresponding to the target quantization interval.
[0081] In other words, for any feature value in the first low-dimensional feature, if it is in the quantization space corresponding to a certain quantization value, then the feature value is replaced with that quantization value.
[0082] In one application scenario, the steps to determine the quantization interval corresponding to each quantization value include: sorting multiple quantization values in ascending order, determining the mean of two adjacent quantization values, and constructing the quantization interval corresponding to each quantization value using the mean of each quantization value.
[0083] To facilitate understanding, specific examples are provided below: Assume there are 4 quantized values: p1, p2, p3, and p4. Sorted from smallest to largest, the result is: p1, p2, p3, p4. Then, determine (p1+p2) / 2, (p2+p3) / 2, and (p3+p4) / 2. The quantization interval for quantized value p1 is (-∞, (p1+p2) / 2), for p2 it is ((p1+p2) / 2, (p2+p3) / 2], for p3 it is ((p2+p3) / 2, (p3+p4) / 2], and for p4 it is ((p3+p4) / 2, -∞).
[0084] S123: Replace each quantization value in the first quantization feature with the corresponding binary encoding vector to obtain the first bitstream.
[0085] Specifically, since the quantized values are floating-point data, after the aforementioned steps, each value in the first low-dimensional feature is still floating-point data. In order to obtain the final first bitstream, it is necessary to convert the floating-point data into integer data.
[0086] Since the correspondence between quantization values and binary encoding vectors was established in the aforementioned steps, step S123 directly uses the binary encoding vector corresponding to the quantization value to represent the quantization value, thereby obtaining the first bitstream.
[0087] Continue reading Figure 3 In this embodiment, the encoding device includes a first quantizer. A first low-dimensional feature output from a first encoder is input to the first quantizer, and then the first quantizer outputs a first bitstream. That is, steps S122 and S123 are executed using the first quantizer.
[0088] The above describes the process of performing a first encoding process on the first basic feature to obtain a first bitstream. It should be noted that in other embodiments, any existing technology can also be used to perform the first encoding process on the basic feature to obtain the first bitstream.
[0089] Continue reading Figure 1 The following describes the steps following step S120.
[0090] S130: Obtain pre-reconstruction features.
[0091] The pre-reconstructed features are the same as those obtained by decoding the first bitstream.
[0092] Specifically, the pre-reconstructed features are the same as the features obtained by decoding the first bitstream. However, due to the loss during the encoding and decoding process, the features obtained by decoding the first bitstream have a certain deviation from the basic features. Therefore, the residual features can be determined based on the basic features and the pre-reconstructed features. The residual features represent the loss that exists after encoding and decoding the basic features.
[0093] Since the process of decoding the first bitstream needs to correspond to the process of encoding the basic features, the process of decoding the first bitstream, corresponding to the above encoding process, should include: (a) Replace each binary encoded vector in the first bitstream with the corresponding quantization value to obtain the first inverse quantization feature, wherein the length of the binary encoded vector is equal to the preset number of quantization bits.
[0094] Specifically, the purpose of this step is to convert the first bitstream from integer data to floating-point data. First, the first bitstream needs to be divided into blocks according to a preset number of quantization bits to obtain individual binary encoded vectors in the first bitstream. For example, when the preset number of quantization bits is 2, the length of the binary encoded vector is determined to be two. Then, in the first bitstream, every two values are used as a binary encoded vector in a sequential order from front to back.
[0095] Specifically, when converting integer data to floating-point data, the correspondence between the binary encoded vector and the quantized value must be the same as the correspondence used when converting floating-point data to integer data.
[0096] Corresponding to the above, the process of determining the quantization value corresponding to each binary vector is as follows: based on the pre-learned base vector and the operation value of each binary encoded vector in the first bit stream, the quantization value corresponding to each binary encoded vector is determined.
[0097] Specifically, if the operation value of the base vector and the binary encoded vector is directly used as the quantization value corresponding to the binary encoded vector, then the binary encoded vector in the first bitstream is operated on with the base vector, and the operation result is directly used as the quantization value corresponding to the binary encoded vector.
[0098] However, when the above process first performs operations on the base vector with each binary encoded vector to obtain the operation value corresponding to each binary encoded vector, and then performs a fully connected operation on all operation values to obtain the quantized values corresponding to multiple binary encoded vectors, since not all quantized values may be used when converting the first quantization feature into the first bitstream, the specific process of determining the quantized values corresponding to each binary encoded vector in the first bitstream includes: (1): Determine multiple binary encoding vectors based on the pre-set number of quantization bits.
[0099] This step is exactly the same as step S1221, and the multiple binary encoding vectors determined are also exactly the same as the multiple binary encoding vectors determined in step S1221.
[0100] (2): Perform operations on the basis vector and each binary encoded vector respectively to obtain multiple operation values, wherein the multiple operation values correspond one-to-one with the multiple binary encoded vectors.
[0101] (3): Perform full connection processing on multiple operation values to obtain multiple quantization values, wherein each quantization value corresponds to one of the multiple operation values.
[0102] (4): Next, the binary encoded vector in the first bitstream is operated on with the base vector to obtain the operation value corresponding to the binary encoded vector.
[0103] (5): Based on the correspondence between the quantization value and the operation value determined in step (3), find the quantization value corresponding to the operation value corresponding to the binary encoding vector in the first bit stream, thereby determining the quantization value corresponding to the binary encoding vector in the first bit stream.
[0104] In summary, the process of determining the quantization value corresponding to the binary encoded vector in the first bitstream is determined by the process of determining the quantization value corresponding to the binary encoded vector in the aforementioned step S1222. It is sufficient to ensure that the quantization value corresponding to the same binary encoded vector is the same during both the encoding and decoding processes.
[0105] (b) Perform a comprehensive transformation on the first inverse quantization feature to obtain the pre-reconstructed feature, wherein the dimension of the pre-reconstructed feature is equal to the dimension of the basic feature.
[0106] The process of performing comprehensive transformation on the first inverse quantization feature corresponds to the process of analyzing and transforming the basic feature.
[0107] See Figure 9 When the decoding device needs to decode the first bitstream, it inputs the first bitstream into the first dequantizer, which then outputs the first dequantized feature. This first dequantized feature is then input into the first decoder, which performs a comprehensive transformation on it to obtain the pre-reconstructed feature. The first dequantizer corresponds to the first quantizer, and the first decoder corresponds to the first encoder.
[0108] In one application scenario, when the first inverse quantization feature includes multiple first inverse quantization sub-features, the process of the first decoder performing a comprehensive transformation on the first inverse quantization feature includes: (b1): The first inverse quantization features in the multiple first inverse quantization features are concatenated to obtain the corresponding concatenated features.
[0109] To facilitate understanding, examples are provided below: For the first dequantized sub-feature a and the first dequantized sub-feature b, if in the first low-dimensional features, sub-feature c corresponds to the first dequantized sub-feature a (sub-feature c is processed sequentially to obtain the first dequantized sub-feature a), and sub-feature d corresponds to the first dequantized sub-feature b (sub-feature d is processed sequentially to obtain the first dequantized sub-feature b), then in the process of generating the first low-dimensional features, if feature e is divided into two sub-features and then dimensionality reduction is performed on these two sub-features respectively to obtain sub-feature c and sub-feature d, then the first dequantized sub-feature a and the first dequantized sub-feature b correspond.
[0110] In this case, there can be two corresponding first inverse quantized features or three corresponding first inverse quantized features, depending on the process of generating the first low-dimensional features.
[0111] (b2): Each spliced sub-feature is subjected to dimensionality-upgrading processing to obtain the corresponding first dimensionality-upgraded sub-feature.
[0112] Specifically, the method used to upgrade the dimensions of the spliced sub-features is also determined by the process of generating the first low-dimensional feature.
[0113] Using the example above, if we use a fully connected approach to reduce the dimensionality of the two sub-features after splitting feature e, resulting in sub-features c and d, then after concatenating the first inverse-quantized sub-feature a and the first inverse-quantized sub-feature b, we also use a fully connected approach to increase the dimensionality of the concatenated sub-features.
[0114] (b3): Determine the pre-reconstructed features based on the obtained first-dimensional sub-features.
[0115] To understand steps (b1) to (b3), combine Figure 6 and Figure 10 Explanation: exist Figure 6 In the application scenario, the obtained first low-dimensional feature includes four 64-dimensional second dimensionality-reduced sub-features. Figure 6 (represented by the symbol 5), then correspondingly, in Figure 10 In the application scenario, the first inverse quantization feature includes four 64-dimensional first inverse quantization sub-features (represented by label 301), and the four 64-dimensional first inverse quantization sub-features included in the first inverse quantization feature correspond one-to-one with the four 64-dimensional second dimensionality reduction sub-features included in the first low-dimensional feature.
[0116] and Figure 6 Application scenarios correspond, Figure 10 First, the two corresponding 64-dimensional first inverse quantized sub-features (denoted by 301) are concatenated together, and then dimensionality-upgrading is performed to obtain a 256-dimensional first dimensionality-upgraded sub-feature (denoted by 302). Then, the two obtained first dimensionality-upgraded sub-features are concatenated together and dimensionality-upgraded again to obtain a 1024-dimensional first dimensionality-upgraded feature (denoted by 303). Next, the first dimensionality-upgraded feature is further dimensionality-upgraded to obtain a 2048-dimensional second dimensionality-upgraded feature (denoted by 304). Finally, the 2048-dimensional second dimensionality-upgraded feature is fully connected to obtain a 2048-dimensional pre-reconstructed feature (denoted by 305).
[0117] In another application scenario, if in step S121 above, when generating the first dimensionality-reduced feature, one of the multiple sub-features obtained from the segmentation is directly used as part of the first low-dimensional feature, while the other sub-features are subjected to dimensionality reduction processing, then when generating the pre-reconstructed feature, the following steps (c1)-(c3) are adopted. At this time, the first inverse quantization feature also includes multiple first inverse quantization sub-features. The process of the first decoder performing comprehensive transformation processing includes: (c1): Upgrade the first target inverse feature among multiple first inverse features to obtain the corresponding first upgraded feature.
[0118] To facilitate understanding, examples are provided below: For the first dequantized sub-feature o, if the sub-feature L in the obtained first low-dimensional feature corresponds to the first dequantized sub-feature o (the sub-feature L is processed in sequence to obtain the first dequantized sub-feature o), and in the process of generating the first low-dimensional feature, the sub-feature L is obtained by the last layer of dimensionality reduction processing, then the first dequantized sub-feature o is the first target dequantized sub-feature.
[0119] (c2): The first dimension-upgraded sub-feature is concatenated with the second target inverse sub-feature among multiple first inverse sub-features to obtain the corresponding concatenated sub-feature.
[0120] Using the example above, if, in the obtained first sub-inverse quantization feature m, sub-feature f corresponds to the first sub-inverse quantization feature m (sub-feature f is processed in sequence to obtain the first inverse quantization feature m), and during the generation of the first low-dimensional feature, sub-feature g is segmented to obtain sub-feature f and sub-feature h, and sub-feature h is reduced in dimensionality to obtain sub-feature L, then the first inverse quantization feature m is the second target inverse quantization feature.
[0121] (c3): Determine the pre-reconstruction features based on the splicing sub-features and the first dequantization sub-features among multiple first dequantization sub-features, excluding the first target dequantization sub-features and the second target dequantization sub-features.
[0122] The following is combined Figure 7 as well as Figure 11 Steps (c1) to (c3) are explained below: When using Figure 7The method shown involves dimensionality reduction of the basic features to obtain the first low-dimensional features. These first low-dimensional features include a first sub-feature labeled 31, a second sub-feature labeled 41, a third sub-feature labeled 51, and a third dimensionality-reduced sub-feature labeled 6. The first inverse quantization features include first inverse quantization sub-features labeled 401, 402, 403, and 404. The first inverse quantization sub-feature labeled 401 corresponds to the third dimensionality-reduced sub-feature labeled 6, therefore, the first inverse quantization sub-feature labeled 401 is the first target inverse quantization sub-feature. The first inverse quantization sub-feature labeled 402 corresponds to the third sub-feature labeled 51, therefore, the first inverse quantization sub-feature labeled 402 is the second target inverse quantization sub-feature. Correspondingly, the first inverse quantization sub-feature labeled 403 corresponds to the second sub-feature labeled 41, and the first inverse quantization sub-feature labeled 404 corresponds to the first sub-feature labeled 31. Then, a comprehensive transformation is performed on the first inverse quantization features... During the transformation process, the first inversely quantized sub-feature with 64 dimensions and labeled 401 is upgraded to obtain a first upgraded sub-feature with 144 dimensions and labeled 405. Then, the first upgraded sub-feature (144 dimensions) is concatenated with the first inversely quantized sub-feature (64 dimensions) labeled 402 to obtain a concatenated sub-feature with 208 dimensions. This concatenated sub-feature is then upgraded to obtain a sub-feature with 416 dimensions and labeled 406. Finally, the sub-feature labeled 406 (with 416 dimensions) is... The sub-feature labeled 403 (64 dimensions) is concatenated with the first inverse quantized sub-feature, resulting in a 480-dimensional sub-feature. This sub-feature is then upgraded to a 960-dimensional sub-feature labeled 407. The 960-dimensional sub-feature labeled 407 is then concatenated with the first inverse quantized sub-feature labeled 404 (64 dimensions), resulting in a 1024-dimensional feature. This sub-feature is then upgraded to a pre-reconstructed feature labeled 408 with a 2048-dimensional dimension.
[0123] The above describes the process of decoding the first bitstream to obtain pre-reconstructed features, which includes two parts: generating the first inverse quantization feature based on the first bitstream and performing a comprehensive transformation on the first inverse quantization feature.
[0124] Since the first bitstream is obtained by replacing each quantization value in the first quantization feature with the corresponding binary code vector, the first inverse quantization feature is obtained by replacing each binary code vector in the first bitstream with the corresponding quantization value, which is the same as the first quantization feature.
[0125] As for the encoding device, it is aware of the first quantization feature; therefore, in this embodiment, refer to... Figure 12 Step S130 specifically includes: S131: Perform a comprehensive transformation on the first quantized feature to obtain the pre-reconstructed feature, wherein the dimension of the pre-reconstructed feature is equal to the dimension of the basic feature.
[0126] See Figure 3 At this point, the encoding device includes a first decoder, which receives the first quantization feature from the first quantizer, and then performs a comprehensive transformation on the first quantization feature to output the pre-reconstructed feature.
[0127] The process of performing comprehensive transformation on the first quantization feature is the same as the process of performing comprehensive transformation on the first inverse quantization feature described above, and the process is as follows: In one application scenario, the first quantization feature includes multiple first quantization sub-features. The process of comprehensively transforming the first quantization feature includes: concatenating the corresponding first quantization sub-features from the multiple first quantization sub-features to obtain the corresponding concatenated sub-features; performing dimensionality upscaling on each concatenated sub-feature to obtain the corresponding first dimensionality upscaling sub-features; and determining the pre-reconstruction feature based on the obtained first dimensionality upscaling sub-features.
[0128] In another application scenario, the first quantization feature includes multiple first quantization sub-features. In this case, the process of comprehensively transforming the first quantization feature includes: performing dimensionality increase processing on the first target quantization sub-feature among the multiple first quantization sub-features to obtain the corresponding first dimensionality increase sub-feature; concatenating the first dimensionality increase sub-feature with the second target quantization sub-feature among the multiple first quantization sub-features to obtain the corresponding concatenated sub-feature; and determining the pre-reconstruction feature based on the concatenated sub-feature and the first quantization sub-features among the multiple first quantization sub-features excluding the first target quantization sub-feature and the second target quantization sub-feature.
[0129] The specific process of performing comprehensive transformation on the first quantization feature can be found above and will not be repeated here.
[0130] The above describes in detail the decoding process of the first bitstream and the process by which the encoding device acquires pre-reconstructed features. Please refer to the following sections for further details. Figure 1 The steps following step S130 are described below.
[0131] S140: Determine residual features based on basic features and pre-reconstruction features.
[0132] Specifically, feature-level subtraction operations can be performed on the basic features and the pre-reconstructed features to obtain residual features.
[0133] Continue reading Figure 3 In this embodiment, the basic features and pre-reconstructed features are input into the feature-level subtractor ( Figure 3 In the section (represented by label 101), the feature-level subtractor then outputs the residual feature.
[0134] It is understandable that the information carried by the residual features represents the loss features after the basic features have undergone encoding and decoding processes.
[0135] S150: Perform a second encoding process on the residual features to obtain a second bitstream, wherein the target reconstruction features of the object to be processed are obtained by modifying the first bitstream with the second bitstream.
[0136] Specifically, the second bitstream carries the loss feature information after the basic features have been encoded and decoded.
[0137] The steps for performing the second encoding process on the residual features are basically the same as those for performing the first encoding process on the basic features, with only some parameters differing, such as the feature dimensions being different in the two encoding processes. The process for performing the second encoding process on the residual features can be found in the process for performing the first encoding process on the basic features, and will not be repeated here.
[0138] Continue reading Figure 3 In this embodiment, the encoding device includes a second encoder and a second quantizer. After obtaining the residual features, the residual features are input into the second encoder. Then, the second quantizer receives the output of the second encoder and finally outputs the second bitstream. The structure of the second encoder is basically the same as that of the first encoder, and the structure of the second quantizer is basically the same as that of the first quantizer.
[0139] The first bitstream carries the main feature information of the object to be processed, while the second bitstream carries the loss feature information after encoding and decoding the basic features of the object to be processed. Therefore, after the encoding device sends the first and second bitstreams to the decoding device, the decoding device can use the second bitstream to correct the first bitstream. Then, based on the correction result, the target reconstructed features of the object to be processed are determined, thereby reducing the distortion of the object to be processed during transmission and ensuring the transmission quality of the object to be processed. For example, when the encoded object is an image, the above method can ensure that the error between the image reconstructed by the server and the image captured by the camera is very small after the server receives the image captured by the camera, thereby improving the accuracy of subsequent visual tasks based on the reconstructed image.
[0140] See Figure 13 The following describes the decoding and reconstruction process performed by the decoding device in this application, which includes: S210: Receive the first bitstream and the second bitstream of the object to be processed.
[0141] The first and second bitstreams are obtained by encoding devices. The encoding process of the encoding devices can be found in the above content and will not be repeated here.
[0142] S220: The first bitstream is corrected using the second bitstream to obtain the corrected result.
[0143] S230: Based on the correction results, obtain the target reconstruction features of the object to be processed.
[0144] Specifically, the first bitstream carries the main feature information of the object to be processed, while the second bitstream carries the loss feature information after encoding and decoding the basic features of the object to be processed. Therefore, the second bitstream can be used to correct the first bitstream. Specifically, the information in the second bitstream is used to compensate and enhance the information in the first bitstream, thereby ensuring the accuracy of the reconstructed target features.
[0145] In one application scenario, combined Figure 14 Step S220 specifically includes: replacing each binary encoded vector in the first bitstream with its corresponding quantization value to obtain a first inverse quantization feature, wherein the length of the binary encoded vector is equal to the preset number of quantization bits; replacing each binary encoded vector in the second bitstream with its corresponding quantization value to obtain a second inverse quantization feature; and using the second inverse quantization feature to correct the first quantization feature to obtain a corrected feature. At this time, step S230 specifically includes: performing a comprehensive transformation on the corrected feature to obtain a target reconstruction feature, wherein the dimension of the target reconstruction feature is equal to the dimension of the basic feature.
[0146] Specifically, the process of determining the first inverse quantization feature has been described above; please refer to the details above. The process of determining the second inverse quantization feature is essentially the same as that of determining the first inverse quantization feature; please refer to the relevant content above for details, and will not be repeated here.
[0147] In this process, the second inverse quantization feature can be added to the first quantization feature at the feature level to obtain the corrected feature.
[0148] The process of performing comprehensive transformation on the modified features is the same as the process of performing comprehensive transformation on the first inverse quantization features described above, as detailed above.
[0149] In this application scenario, the decoding device includes a first dequantizer and a second dequantizer. After receiving the first bitstream and the second bitstream, the decoding device inputs the first bitstream into the first dequantizer to obtain the first dequantized feature, and inputs the second bitstream into the second dequantizer to obtain the second dequantized feature. Then, the first and second dequantized features are input into the first adder (denoted by label 102) at the feature level to obtain the corrected feature, and then the corrected feature is input into the second decoder. The second decoder outputs the target reconstructed feature.
[0150] The prerequisite for inputting the first and second inverse quantization features into the first adder at the feature level is that the dimensions of the first and second inverse quantization features are the same. Therefore, if the dimensions of the first and second inverse quantization features are different, they can be processed by fully connected or convolutional methods before inputting them into the first adder at the feature level.
[0151] For example, when the dimension of the second inverse quantization feature is smaller than the dimension of the first inverse quantization feature, the second inverse quantization feature is subjected to dimensionality-up processing such as fully connected or convolution, so that the feature obtained after processing has the same dimension as the first inverse quantization feature. Then, the feature obtained after processing and the first inverse quantization feature are input together into the first adder at the feature level.
[0152] In another application scenario, combined with Figure 15 Step S220 specifically includes: decoding the first bitstream to obtain pre-reconstructed features; decoding the second bitstream to obtain residual reconstruction features; using the residual reconstruction features to correct the pre-reconstructed features to obtain corrected features. At this time, step S230 specifically includes: determining the corrected features as the target reconstruction features.
[0153] Specifically, the process of decoding the first bitstream has been described above, and the process of decoding the second bitstream is the same as that of decoding the first bitstream ...
[0154] In this process, the residual reconstruction features and the pre-reconstruction features can be added at the feature level to obtain the target reconstruction features.
[0155] In this application scenario, the decoding device includes a first dequantizer, a first decoder, a second dequantizer, and a second decoder. After the first bitstream is input to the first dequantizer, the first decoder receives the output of the first dequantizer and outputs pre-reconstructed features. After the second bitstream is input to the second dequantizer, the second decoder receives the output of the second dequantizer and outputs residual reconstructed features. Then, the residual reconstructed features and the pre-reconstructed features are input together into the feature-level second adder (denoted by label 103). The feature-level second adder finally outputs the target reconstructed features.
[0156] To better understand the scheme of this application, the following will be combined with... Figure 16 The encoding and decoding methods of this application are described simultaneously: After acquiring the basic features of the object to be processed, the basic features are input into the first encoder. Then, the first quantizer receives the output of the first encoder and outputs the first bitstream. This process corresponds to... Figure 16 The process within the dashed box 10 is part of the encoding process.
[0157] Simultaneously, the encoding device will input the output of the first encoder into the first decoder to obtain the pre-reconstructed features. Then, the encoding device will input the basic features and the pre-reconstructed features together into the feature-level subtractor (represented by label 101) to obtain the residual features.
[0158] Next, the residual features are input into the second encoder, the output of the second encoder is fed into the second quantizer, and the second quantizer outputs the second bitstream. This process corresponds to... Figure 16 The process of the dashed box 20 is also part of the encoding process.
[0159] After obtaining the first and second bitstreams, the encoding device sends both bitstreams to the decryption device.
[0160] exist Figure 16 In the diagram, dashed boxes 30 and 40 both correspond to the decoding process, while dashed lines 1 and 2 are not selected simultaneously.
[0161] When dashed line 1 is selected, the decoding process of the decoding device is as follows: First, the first bitstream is input into the first dequantizer to obtain the first dequantized feature, and the second bitstream is input into the second dequantizer to obtain the second dequantized feature. Then, the first and second dequantized features are input together into the first adder at the feature level (denoted by label 102). The feature-level adder then outputs the corrected feature, which is then fed into the second decoder. Finally, the second decoder outputs the target reconstructed feature.
[0162] When feeding the first and second inverse quantization features into the first adder of the feature level, it is necessary to ensure that the dimensions of the first and second inverse quantization features are the same. Therefore, if the dimensions of the first and second inverse quantization features are different, the first or second inverse quantization features need to be processed. For details, please refer to the above.
[0163] When dashed line 2 is selected, the decoding process of the decoding device is as follows: The first bitstream is fed into the first dequantizer, and the output of the first dequantizer is fed into the first decoder. The first decoder outputs the pre-reconstructed features.
[0164] The second bitstream is fed into the second dequantizer, and the output of the second dequantizer is fed into the second decoder. The second decoder outputs the residual reconstruction features.
[0165] The pre-reconstructed features and residual reconstruction features are then fed together into the second adder at the feature level (denoted by label 103). Finally, the second adder at the feature level outputs the corrected features, which are the target reconstruction features.
[0166] See Figure 17 , Figure 17 This is a schematic diagram of one embodiment of the encoding device of this application. The encoding device 200 includes a processor 210, a memory 220, and a communication circuit 230. The processor 210 is coupled to the memory 220 and the communication circuit 230 respectively. The memory 220 stores program data. The processor 210 executes the program data in the memory 220 to implement the steps of the encoding method in any of the above embodiments. The detailed steps can be found in the above embodiments and will not be repeated here.
[0167] The encoding device 200 can be any device with algorithm processing capabilities, such as a computer, mobile phone, or camera, and there are no restrictions on its use.
[0168] See Figure 18 , Figure 18 This is a schematic diagram of another embodiment of the encoding device of this application. The encoding device 300 includes a first acquisition module 310, a first encoding module 320, a second acquisition module 330, a determination module 340, and a second encoding module 350.
[0169] The first acquisition module 310 is used to acquire the basic characteristics of the object to be processed.
[0170] The first encoding module 320 is connected to the first acquisition module 310 and is used to perform a first encoding process on the basic features to obtain a first code stream.
[0171] The second acquisition module 330 is connected to the first encoding module 320 and is used to acquire pre-reconstructed features, wherein the pre-reconstructed features are the same as the features obtained by decoding the first bitstream.
[0172] The determination module 340 is connected to both the first acquisition module 310 and the second acquisition module 330, and is used to determine the residual features based on the basic features and the pre-reconstruction features.
[0173] The second encoding module 350 is connected to the determination module 340 and performs a second encoding process on the residual features to obtain a second bitstream. The target reconstruction features of the object to be processed are obtained by correcting the first bitstream with the second bitstream.
[0174] When the encoding device 300 of this application is working, it executes the steps of the encoding method in any of the above embodiments. For details of the steps, please refer to the relevant content above, which will not be repeated here.
[0175] The encoding device 300 can be any device with algorithm processing capabilities, such as a computer, mobile phone, or camera, and there are no restrictions on its use.
[0176] See Figure 19 , Figure 19 This is a schematic diagram of one embodiment of the decoding device of this application. The decoding device 400 includes a processor 410, a memory 420, and a communication circuit 430. The processor 410 is coupled to the memory 420 and the communication circuit 430. The memory 420 stores program data. The processor 410 executes the program data in the memory 420 to implement the steps of the decoding method in any of the above embodiments. The detailed steps can be found in the above embodiments and will not be repeated here.
[0177] Specifically, the decoding device 400 can be any device with algorithm processing capabilities, such as a computer, mobile phone, or camera, without any restrictions.
[0178] See Figure 20 , Figure 20 This is a schematic diagram of another embodiment of the decoding device of this application. The decoding device 500 includes a receiving module 510, a correction module 520, and a determining module 530.
[0179] The receiving module 510 is used to receive the first bitstream and the second bitstream of the object to be processed. The first bitstream and the second bitstream are obtained by any of the encoding methods described above. For details of the methods, please refer to the above content, which will not be repeated here.
[0180] The correction module 520 is connected to the receiving module 510 and is used to correct the first bitstream using the second bitstream to obtain the correction result.
[0181] The determination module 530 is connected to the correction module 520 and is used to obtain the target reconstruction features of the object to be processed based on the correction results.
[0182] When the decoding device 500 is working, it executes the steps of the decoding method in any of the above embodiments. For details of the steps, please refer to the relevant content above, which will not be repeated here.
[0183] Specifically, the decoding device 500 can be any device with algorithm processing capabilities, such as a computer, mobile phone, or camera, without any restrictions.
[0184] See Figure 21 , Figure 21 This is a schematic diagram of one embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 600 stores a computer program 610, which can be executed by a processor to implement the steps in any of the above methods.
[0185] Specifically, the computer-readable storage medium 600 can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or a device that can store the computer program 610. Alternatively, it can be a server that stores the computer program 610, which can send the stored computer program 610 to other devices for execution or run the stored computer program 610 itself.
[0186] The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An encoding method, characterized in that, The method includes: Obtain the basic characteristics of the object to be processed; The basic features are subjected to a first encoding process to obtain a first bitstream; Obtain pre-reconstructed features, wherein the pre-reconstructed features are the same as the features obtained by decoding the first bitstream; Based on the aforementioned basic features and the pre-reconstruction features, the residual features are determined; The residual features are subjected to a second encoding process to obtain a second bitstream, wherein the target reconstruction features of the object to be processed are obtained by the correction result of the first bitstream modified by the second bitstream. The step of performing a first encoding process on the basic features to obtain a first bitstream includes: The basic features are analyzed and transformed to obtain a first low-dimensional feature, wherein the dimension of the first low-dimensional feature is smaller than the dimension of the basic features. Each feature value in the first low-dimensional feature is replaced with the quantized value corresponding to the feature value to obtain the first quantized feature; Each quantization value in the first quantization feature is replaced with the binary encoding vector corresponding to the quantization value to obtain the first bitstream; The step of analyzing and transforming the basic features to obtain the first low-dimensional features includes: Obtain a first feature, wherein the first feature is obtained based on the basic feature, and the dimension of the first feature is smaller than the dimension of the basic feature; The first feature is divided into multiple first sub-features; At least some of the first sub-features are subjected to dimensionality reduction processing to obtain the corresponding first dimensionality-reduced sub-features; The first low-dimensional feature is determined based on the obtained first dimensionality-reduced sub-feature and the first sub-feature that has not undergone dimensionality reduction processing.
2. The method according to claim 1, characterized in that, The step of performing dimensionality reduction processing on at least a portion of the first sub-features to obtain the corresponding first dimensionality-reduced sub-features includes: Dimensionality reduction is performed on each of the plurality of first sub-features except for the first target sub-feature to obtain the corresponding first dimension-reduced sub-features; The step of determining the first low-dimensional feature based on the obtained first dimensionality-reduced sub-feature and the first sub-feature without dimensionality reduction includes: Each of the first dimension-reduced sub-features is divided into multiple second sub-features; Dimensionality reduction is performed on each of the obtained second sub-features other than the second target sub-feature to obtain the corresponding second dimension-reduced sub-features. The first low-dimensional feature is determined based on the first target sub-feature, the second target sub-feature, and the obtained second dimensionality-reduced sub-feature.
3. The method according to claim 1, characterized in that, The step of obtaining the first feature includes: The basic features are processed by a fully connected layer to obtain a second feature with the same dimension as the basic features. The second feature is then subjected to dimensionality reduction processing to obtain the first feature.
4. The method according to claim 1, characterized in that, Before replacing each feature value in the first low-dimensional feature with the quantized value corresponding to the feature value to obtain the first quantized feature, the method further includes: Based on a preset number of quantization bits, multiple binary encoding vectors are determined, wherein the length of each binary encoding vector is equal to the number of quantization bits; Based on the pre-learned basis vectors and the operation values of each binary encoded vector, multiple quantization values are determined, wherein each of the multiple quantization values corresponds one-to-one with the multiple binary encoded vectors.
5. The method according to claim 4, characterized in that, The step of determining multiple quantization values based on the computational values of the pre-learned basis vectors and each of the binary encoded vectors includes: Each of the base vectors is operated on with each of the binary encoded vectors to obtain a plurality of operation values, wherein the plurality of operation values correspond one-to-one with the plurality of binary encoded vectors; A full-connection process is performed on the multiple operational values to obtain multiple quantized values, wherein each of the multiple quantized values corresponds one-to-one with the multiple operational values.
6. The method according to claim 1, characterized in that, The step of replacing each feature value in the first low-dimensional feature with the quantized value corresponding to the feature value to obtain the first quantized feature includes: Obtain multiple quantization values, and obtain the quantization interval corresponding to each quantization value; The quantization interval of each feature value in the first low-dimensional feature is determined as the target quantization interval for each feature value. Each of the aforementioned feature values is replaced with the quantization value corresponding to the target quantization interval.
7. The method according to claim 6, characterized in that, The step of obtaining the quantization interval corresponding to each quantization value includes: After sorting the multiple quantized values in ascending order, the average value of two adjacent quantized values is determined. The quantization interval corresponding to each quantization value is constructed by using the mean value corresponding to each quantization value.
8. The method according to claim 1, characterized in that, The step of obtaining pre-reconstructed features includes: The first quantized feature is subjected to a comprehensive transformation process to obtain the pre-reconstructed feature, wherein the dimension of the pre-reconstructed feature is equal to the dimension of the basic feature.
9. The method according to claim 8, characterized in that, The first quantization feature includes multiple first quantization sub-features; the step of performing a comprehensive transformation on the first quantization feature to obtain the pre-reconstructed feature includes: The first quantization sub-features corresponding to the plurality of first quantization sub-features are concatenated to obtain the corresponding concatenated sub-features; Each of the spliced sub-features is subjected to dimensionality-upgrading processing to obtain the corresponding first dimensionality-upgrading sub-feature; Based on the obtained first raised-dimensional sub-feature, the pre-reconstructed features are determined.
10. The method according to claim 8, characterized in that, The first quantization feature includes multiple first quantization sub-features; the step of performing a comprehensive transformation on the first quantization feature to obtain the pre-reconstructed feature includes: The first target quantization sub-feature among the plurality of first quantization sub-features is subjected to dimensionality-upgrading processing to obtain the corresponding first dimensionality-upgrading sub-feature; The first dimension-upgrading sub-feature is concatenated with the second target quantization sub-feature among the plurality of first quantization sub-features to obtain the corresponding concatenated sub-feature; The pre-reconstruction features are determined based on the splicing sub-features and the first quantization sub-features among the plurality of first quantization sub-features, excluding the first target quantization sub-feature and the second target quantization sub-feature.
11. A decoding method, characterized in that, The method includes: Receive a first bitstream and a second bitstream of the object to be processed; the first bitstream and the second bitstream are obtained by the encoding method according to any one of claims 1-10; The first bitstream is corrected using the second bitstream to obtain the correction result; Based on the correction results, the target reconstruction features of the object to be processed are obtained.
12. The method according to claim 11, characterized in that, The step of correcting the first bitstream using the second bitstream to obtain the correction result includes: Each binary encoded vector in the first bitstream is replaced with its corresponding quantization value to obtain a first inverse quantization feature, wherein the length of the binary encoded vector is equal to the number of quantization bits preset. Each binary encoded vector in the second bitstream is replaced with its corresponding quantization value to obtain a second inverse quantization feature; The first inverse quantization feature is modified using the second inverse quantization feature to obtain the modified feature; The step of obtaining the target reconstruction features of the object to be processed based on the correction result includes: The modified features are subjected to a comprehensive transformation process to obtain the target reconstruction features, wherein the dimension of the target reconstruction features is equal to the dimension of the base features.
13. The method according to claim 12, characterized in that, Before determining the target reconstruction features of the object to be processed based on the first bitstream and the second bitstream, the method further includes: The quantization value corresponding to each binary encoded vector is determined based on the operation value of the pre-learned basis vectors and each binary encoded vector in the first bitstream.
14. The method according to claim 11, characterized in that, The step of correcting the first bitstream using the second bitstream to obtain the correction result includes: The first bitstream is decoded to obtain the pre-reconstructed features; The second bitstream is decoded to obtain residual reconstruction features; The pre-reconstructed features are modified using the residual reconstruction features to obtain the modified features; The step of obtaining the target reconstruction features of the object to be processed based on the correction result includes: The modified feature is determined as the target reconstruction feature.
15. The method according to claim 14, characterized in that, The step of decoding the first bitstream to obtain the pre-reconstructed features includes: Each binary encoded vector in the first bitstream is replaced with its corresponding quantization value to obtain a first inverse quantization feature, wherein the length of the binary encoded vector is equal to the number of quantization bits preset. The first inverse quantization feature is subjected to a comprehensive transformation process to obtain the pre-reconstructed feature.
16. An encoding device, characterized in that, The encoding device includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit respectively. The memory stores program data. The processor executes the program data in the memory to implement the steps of the method as described in any one of claims 1-10.
17. A decoding device, characterized in that, The decoding device includes a processor, a memory, and a communication circuit. The processor is coupled to the memory and the communication circuit respectively. The memory stores program data. The processor executes the program data in the memory to implement the steps of the method as described in any one of claims 11-15.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the steps of the method as described in any one of claims 1-15.
Citation Information
Patent Citations
Deep learning image coding method for man-machine cooperation scene under resources constraint
CN113822954A
Method and apparatus for processing image for machine vision
WO2022075754A1