A coding method to achieve translation and scaling invariance of spatial signals

By randomly sampling and recoding the visual targets and using neural networks to achieve sparse projection, the problem of translation and scaling invariance of visual targets in the prior art is solved, and a sparse encoding method with translation and scaling invariance is realized.

CN113298012BActive Publication Date: 2025-05-02ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110631633.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-07
Publication Date
2025-05-02
Estimated Expiration
2041-06-07

AI Technical Summary

Technical Problem

The prior art is difficult to achieve translation and scaling invariance of visual objects, especially in sparse encoding methods.

Method used

By randomly sampling and recoding the target of the input signal, sparse projection is achieved using a neural network with the unique characteristics of the winner, thereby completing the encoding of the target.

Benefits of technology

The translation and scaling invariance characteristics of visual targets are achieved, and the encoding structure is simple and has bionic characteristics, which is suitable for designing feature encoding special chips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113298012B_ABST
    Figure CN113298012B_ABST
Patent Text Reader

Abstract

The present invention discloses a coding method for achieving translation and scaling invariance of spatial signals. Inspired by the structure of the fruit fly brain, a neural coding method including sampling, recoding, sparse projection and other computing processes is proposed. The coding randomly samples the input signal at the target coordinates, and recodes the data. Then, the recoded data is sparsely projected using a neural network with a winner-only feature, thereby completing the coding method for the target with the characteristics of translation and scaling invariance. Especially for image classification tasks, no matter what kind of translation or scaling operation is performed on the image target, the present invention can obtain higher accuracy and generalization ability using a small amount of training data. The coding structure and calculation designed by the present invention are simple, and its bionic characteristics are important references for designing special chips for feature coding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual target coding, and in particular to a coding method for achieving translation and scaling invariance of spatial signals. Background Art

[0002] Modern deep learning can partially achieve translation invariance of image targets through the maximum pooling of convolution calculations. However, this invariance often relies on a large amount of training data, and the range of translation invariance is small. Biological experiments have found that insects with relatively simple brain structures have the ability of visual perception invariance. At present, the results of neuroscience research have been able to depict the brain structure of insects (such as fruit flies), especially the inherent structure of the insect brain can complete complex cognitive tasks. In addition, neurons in the brain also have the characteristics of sparse coding. However, existing methods have not yet achieved sparse coding methods with translation or scaling invariance characteristics. Summary of the invention

[0003] In view of the problems existing in the prior art, the present invention proposes a coding method for achieving translation and scaling invariance of spatial signals. The present invention randomly samples the target coordinates of the input signal and re-encodes the data. The re-encoded data is then sparsely projected using a neural network with a winner-only feature, thereby completing the encoding of the target.

[0004] The technical solution of the present invention is as follows:

[0005] A coding method for achieving translation and scaling invariance of spatial signals is designed based on the brain structure of insects, including sampling, recoding and sparse projection calculation processes, including the following steps:

[0006] Step 1) Perform random sampling and re-encoding calculations on the target M of the input signal to obtain an intermediate representation C of the target;

[0007] The specific steps are as follows:

[0008] Step 1.1) The signal is represented in the form of a matrix S, where the matrix S contains the target M;

[0009] Step 1.2) The target M consists of m elements, that is, M = [(x1, y1), (x2, y2), … (x m ,y m )],(x i ,y i ) represents the coordinates of the element in the signal matrix S;

[0010] Step 1.3) Compare the number of samples t with the number of elements m; if m is greater than t, directly execute step 1.4); otherwise, first perform an expansion operation on M in the signal matrix so that the total number of elements in the signal matrix is ​​equal to t, that is, expand the number of elements in the signal matrix without changing the target signal M;

[0011] Step 1.4) Randomly sample the elements of the target M t times, where each element (x i ,y i ) are sampled with the same probability, which is 1 / m (the present invention adopts a uniform random sampling method, and the use of other uniform sampling methods will not affect the encoding effect of the present invention); the sampled element I i According to the sampling order, they are placed in the set P, P = [I1, I2, ..., I t ].

[0012] Step 1.5) Recode the elements in set P:

[0013] Step 1.5.1) According to the horizontal coordinate x in P i Sort all elements in the set in ascending or descending order to obtain a new sorted coordinate set:

[0014]

[0015] Among them, s i is the horizontal coordinate x in the set P i Rearrange the subscripts and extract P X All the horizontal coordinates in and the vertical coordinate Then we get two one-dimensional column vectors

[0016] Step 1.5.2) According to the vertical coordinate y in P i Sort all elements in the set in ascending or descending order to obtain a new set of coordinate information after sorting:

[0017]

[0018] Among them, i is the ordinate y in the set P i Rearrange the subscripts and extract P Y All the horizontal coordinates in and the vertical coordinate Then we get two one-dimensional column vectors

[0019] Step 1.6) for the 4 vectors in 1.5.1) and 1.5.2) and Normalization is performed; the (0, 1) standard normalization method is adopted (using other normalization methods does not affect the encoding effect of the present invention). For any coordinate set E, the result of element normalization is shown in Formula 1:

[0020]

[0021] where x new Represents the normalized coordinate result, where x i Represents the original coordinates, Min E Represents the smallest element in the coordinate set, Max E Represents the largest element in the coordinate set;

[0022] After normalization, connect the horizontal coordinate set X and the vertical coordinate set Y to obtain and

[0023] Step 1.7) Code X With Code Y The intermediate representation vector C of the target signal M is:

[0024]

[0025] where q1,q2,…,q 4t Code X With Code Y Elements in

[0026] Step 2) Use a single-layer neural network to perform sparse coding on the intermediate representation C of the target, and finally obtain the feature representation O of the target;

[0027] The specific steps are as follows:

[0028] Step 2.1) Take the intermediate representation vector C of the target as input and record it as a one-dimensional column vector containing 4t neurons: I = [q1, q2, ..., q 4t ] T ;

[0029] Step 2.2) Determine the connection matrix W between the neurons of the input layer I and the neurons of the coding layer Y. The neurons of the input layer are connected to the neurons of the coding layer with probability p=0.1, that is, the connection between them has a probability of 90% to be 0, and non-zero values ​​are randomly assigned according to the standard Gaussian distribution (mean 0, variance 1);

[0030] Step 2.3) The activity of the coding layer neurons is calculated according to Y; among them, whether the coding layer neurons are fired will be determined according to the activity value calculated by Y=WI, and a threshold ε is set. If a neuron in Y has an activity value greater than ε, the neuron will be activated, and the neuron with an activity value less than ε will be in a resting state. The output value of the activated neuron is 1, and the output value of the resting neuron is 0. A vector of 0 or 1 is taken as the feature representation O of the target.

[0031] The beneficial effects of the present invention are as follows: inspired by the brain structure of fruit flies, the present invention proposes a neural coding method including computing processes such as sampling, recoding and sparse projection; the method can realize the coding of visual targets including but not limited to, and the coding method has the characteristics of translation and scaling invariance; the coding structure and calculation designed by the present invention are simple, and its bionic characteristics are an important reference for designing a special chip for feature coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is the overall flow chart of the present invention;

[0033] Figure 2 It is a target feature extraction flow chart of the present invention;

[0034] Figure 3 It is a schematic diagram of feature extraction and encoding of the present invention;

[0035] Figure 4 It is a schematic diagram of target feature extraction of handwritten font images of the present invention. DETAILED DESCRIPTION

[0036] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0037] like Figure 1 As shown, a coding method for realizing translation and scaling invariance of spatial signals, the specific steps are as follows:

[0038] Step 1) Perform random sampling and recoding calculations on the target of the input signal as shown in the attached diagram. Figure 2 As shown;

[0039] Step 1.1) As attached Figure 2 As shown in the reference numeral 1, the information matrix S to be extracted is input;

[0040] Step 1.2) as attached Figure 2 As shown in the number 2 in the figure, the coordinate information of all elements describing the target in the information matrix S to be extracted is added to M, where (x i ,y i ) represents the coordinates of the element in the signal matrix S.

[0041] Step 1.3) as attached Figure 2 As shown in the reference numeral 3, the sampling number t is compared with the number of elements m; if m is greater than t, step 1.4) is directly executed; otherwise, an expansion operation is first performed on M in the signal matrix so that the total number of elements in the signal matrix is ​​equal to t, that is, the number of elements in the signal matrix is ​​expanded without changing the target signal M;

[0042] Step 1.4) as attached Figure 2 As shown in the number 4 in the figure, the elements of the target M are randomly sampled t times, where each element (x i ,y i ) are sampled with the same probability of 1 / m; the sampled element I i According to the sampling order, they are placed in the set P, P = [I1, I2, ..., I t ].

[0043] Step 1.5) as attached Figure 2 As shown in the reference numeral 5 in FIG. 5 , the elements in the set P describing the target coordinate information are recoded (the present invention uses a method of ascending coordinate sorting, and using other recoding methods for the coordinate information does not affect the results of the present invention):

[0044] Step 1.5.1) According to the horizontal coordinate x in P i Sort all elements in the set in ascending order to obtain a new set of sorted coordinates:

[0045]

[0046] Among them, s i is the horizontal coordinate x in the set P i Rearrange the subscripts and extract P X All the horizontal coordinates in and the vertical coordinate Then we get two one-dimensional column vectors

[0047] Step 1.5.2) According to the vertical coordinate y in P i Sort all elements in the set in ascending or descending order to obtain a new set of coordinate information after sorting:

[0048]

[0049] Among them, i is the ordinate y in the set P i Rearrange the subscripts and extract P Y All the horizontal coordinates in and the vertical coordinate Then we get two one-dimensional column vectors

[0050] Step 1.6) as attached Figure 2 As shown in the number 6 in the figure, the recoded coordinate set and After normalization, we can get and

[0051] Step 1.7) as attached Figure 2 As shown in the number 7 in the code, the final Code X With Code Y Connection as a target characterization And connect the Code as input data to the sparse coding neural network.

[0052] Step 2) Sparse coding process as shown in the attached Figure 3 As shown, attached Figure 3 The number 7 in the middle represents the information after the target features are extracted and re-encoded. As the input data of neural network sparse coding, as shown in the attached Figure 3 The input neurons shown in the middle mark 8 are connected to the neurons of the sparse coding layer through the probability matrix W, as shown in the attached figure. Figure 3 The activation state of the neurons in the sparse coding layer shown by number 9 is controlled by the threshold ε. Neurons with input greater than ε will be activated, while other neurons are in a resting state, and the output will be used as the encoding of the target in the information matrix.

[0053] Example:

[0054] The present invention relates to a method for encoding a target including but not limited to vision, and the method has the characteristics of translation invariance and scaling invariance. The encoding process of the present invention is described below by taking the recognition of handwritten numbers as an example. The MNIST data set used in the experiment is a grayscale image of size 28*28, each pixel is an eight-bit byte (0-255), and the data set contains ten numbers from 0 to 9, mainly including 60,000 training images and 10,000 test images (download address: http: / / yann.lecun.com / exdb / mnist / ). The present invention aims to solve the problem of translation and scaling invariance of spatial signals. Therefore, the experiment fills each picture in the MNIST data set with blanks, expands the original image (28×28) into a two-dimensional matrix of 56×56, ensures that the size and shape of the target in the original image remain unchanged, but its position in the image changes randomly. The experiment randomly collects and processes 10,000 images in the MNIST data set as a training data set, and randomly moves the target numbers in the image in the image, and tests and compares it with the traditional CNN image recognition method.

[0055] Using the encoding method of the present invention, for images with the same number but different positions, the same target feature encoding can be obtained (such as the attached Figure 4 As shown, attached Figure 4 The coordinate information of the image is simplified. In actual experiments, the dimension of target division and coordinate extraction is larger than that of the attached Figure 4 The example of the present invention is shown in Figure 1, and the feature extraction process of the image scaling example is the same as that of the translation example), while different numbers have different encoding results. The experimental comparison proves that the present invention has the characteristics of translation invariance and scaling invariance.

[0056] The structures, proportions, sizes, etc. shown in the drawings of the present invention are only used to match the contents disclosed in the specification for people familiar with the technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention, so they have no substantial technical significance. Any modification of the structure, change of the proportion relationship or adjustment of the size, without affecting the effects and purposes that can be achieved by the present invention, should be within the scope of the technical content disclosed by the present invention. The specific values ​​given to the parameters in the present invention are only for the convenience of description, and are not used to limit the scope of the implementation of the present invention. The changes or adjustments in their relative relationships should also be regarded as the scope of the implementation of the present invention without substantially changing the technical content.

[0057] 1) The target sampling and recoding process of the handwritten digital image instance is shown in the attached figure. Figure 4 As shown:

[0058] 1.1) As attached Figure 4 As shown in the reference numerals 10 and 15, the coordinate information set M describing all elements of the target in the input image is extracted.

[0059] M=[(x1,y1),(x2,y2),…(x m ,y m )] (1)

[0060] The dimension size of M is compared with the threshold t=100. If the dimension of M is greater than or equal to 100, t-dimensional target information is uniformly and randomly added to the set P in the original image.

[0061] P=[I1,I2,…,I t ], I t ∈(x m ,y m ) (2)

[0062] 1.2) If the dimension of M is less than 100, the original image is expanded

[0063]

[0064] This formula indicates that B is used to dilate image A, where B is a convolution template or convolution kernel, which can be square or circular in shape. Convolution calculation is performed on template B and image A, and each pixel in the image is scanned. The template element and the binary image element are used to perform an "AND" operation. If both are 0, the target pixel is 0, otherwise it is 1. Thus, the maximum value of the pixel in the area covered by B is calculated, and the pixel value of the reference point is replaced with this value to achieve dilation. For the dilated image, perform step 1.1 again.

[0065] 1.3) As attached Figure 4 As shown in the numbers 11 and 16 in the figure, all elements are sorted according to the horizontal coordinates in the set P. i Sort all elements in the set in ascending order to obtain a new set of coordinate information after sorting As attached Figure 4 As shown in the numbers 12 and 17, P is extracted respectively. X All the horizontal coordinates in and the vertical coordinate and are all (1×100) column vectors.

[0066] 1.4) As attached Figure 4 As shown in the numbers 11 and 16, all elements are sorted according to the ordinate in the set P. i Sort all elements in the set in ascending order to obtain a new set of coordinate information after sorting Then extract P Y All the horizontal coordinates in and the vertical coordinate and are all (1×100) column vectors.

[0067] 1.5) As attached Figure 4 As shown in the numbers 13 and 18 in the figure, the coordinate column vectors obtained in 1.3 and 1.4 are normalized. The elements in are normalized according to the (0,1) standard (using other normalization methods does not affect the encoding effect of the present invention):

[0068]

[0069] 1.6) As attached Figure 4 As shown in the numbers 14 and 19 in , the normalized coordinate set E is connected in the form of a column vector to obtain the input vector Code of the sparse coding neural network, which has a size of 400 dimensions.

[0070]

[0071] 2) Sparsely project the target re-encoded information and finally obtain the feature representation of the target:

[0072] 2.1) The Code target feature information obtained in step 1.6 above will be used as input data, as shown in the attached Figure 3 The input layer I = [q1, q2, ..., q 4t ] T By connecting the matrix W ( Figure 3 The number 8 shows that it is connected to the encoding neurons, and the number of neurons in the input layer I is 4t=400.

[0073] 2.2) As attached Figure 3 As shown by the reference numeral 8 in , the connection matrix W is a 0,1 matrix randomly generated with a probability of 90%. As shown in formula (6), the input data of the sparse coding layer can be expressed as Y=WI.

[0074]

[0075] 2.3) Activation of neurons in the sparse coding layer

[0076] As attached Figure 3 As shown in the figure with reference 9, the number of neurons in the sparse coding layer Y is n=50000. The activation state of the neurons is controlled by the threshold ε. The neurons with input greater than ε=0.8445 will be activated, while the other neurons are in a resting state, and their outputs will be used as the feature code of the information matrix.

[0077] 3) Comparison of experimental results

[0078] For the processed MNIST data set, an experimental comparison was conducted with the traditional CNN method. The present invention can achieve an accuracy rate of 98.88% using only 10,000 images for training, while the accuracy rate of the CNN method is only 10%.

Claims

1. A coding method for achieving translation and scaling invariance of spatial signals, characterized in that: This encoding method obtains the same target feature encoding for images with the same number but different positions, so as to improve the image recognition accuracy; Each image in the MNIST dataset is filled with blanks, and the original image 28×28 is expanded to a 56×56 two-dimensional matrix, ensuring that the size and shape of the target in the original image remain unchanged, but its position in the image is randomly changed; The encoding method includes the following steps: Step 1) Perform random sampling and re-encoding calculations on the target M of the input signal to obtain an intermediate representation C of the target; The target M of the input signal is the coordinate information set M of all elements describing the target in the input image; The specific steps are as follows: Step 1.1) The signal is represented in the form of a matrix S, where the matrix S contains the target M; Step 1.2) The target M consists of m elements, that is, M = [(x1, y1), (x2, y2), … (x m ,y m )],(x i ,y i ) represents the coordinates of the element in the signal matrix S; Step 1.3) Compare the number of samples t with the number of elements m; if m is greater than t, directly execute step 1.4); otherwise, first perform an expansion operation on M in the signal matrix so that the total number of elements in the signal matrix is ​​equal to t, that is, expand the number of elements in the signal matrix without changing the target signal M; Step 1.4) Randomly sample the elements of the target M t times, where each element (x i ,y i ) are sampled with the same probability of 1 / m; the sampled element I i According to the sampling order, they are placed in the set P, P = [I1, I2, ..., I t ]; Step 1.5) Recode the elements in set P: Step 1.5.1) According to the horizontal coordinate x in P i Sort all elements in the set in ascending or descending order to obtain a new coordinate set in ascending order: Among them, s i is the horizontal coordinate x in the set P i Rearrange the subscripts and extract P X All the horizontal coordinates in and the vertical coordinate Then we get two one-dimensional column vectors Step 1.5.2) According to the vertical coordinate y in P i Sort all elements in the set in ascending or descending order to obtain a new coordinate information set in ascending order: Among them, i is the ordinate y in the set P i Rearrange the subscripts and extract P Y All the horizontal coordinates in and the vertical coordinate Then we get two one-dimensional column vectors Step 1.6) for the 4 vectors in 1.5.1) and 1.5.2) and Normalization is performed; using the (0,1) standard normalization method, for any coordinate set E, the result of element normalization is shown in Formula 1: where x new Represents the normalized coordinate result, where x i Represents the original coordinates, Min E Represents the smallest element in the coordinate set, Max E Represents the largest element in the coordinate set; After normalization, connect the horizontal coordinate set X and the vertical coordinate set Y to obtain and Step 1.7) Code X With Code Y The intermediate representation vector C of the target signal M is: where q1,q2,…,q 4t Code X With Code Y Elements in Step 2) Use a single-layer neural network to perform sparse coding on the intermediate representation C of the target, and finally obtain the feature representation O of the target; The specific steps are as follows: Step 2.1) Take the intermediate representation vector C of the target as input and record it as a one-dimensional column vector containing 4t neurons: I = [q1, q2, ..., q 4t ] T ; Step 2.2) Determine the connection matrix W between the neurons of the input layer I and the neurons of the coding layer Y. The neurons of the input layer are connected to the neurons of the coding layer with probability p=0.1, that is, the connection between them has a probability of 90% to be 0, and non-zero values ​​are randomly assigned according to the standard Gaussian distribution; Step 2.3) The activity of the coding layer neurons is calculated according to Y; among them, whether the coding layer neurons are fired will be determined according to the activity value calculated by Y=WI, and a threshold ε is set. If a neuron in Y has an activity value greater than ε, the neuron will be activated, and the neuron with an activity value less than ε will be in a resting state. The output value of the activated neuron is 1, and the output value of the resting neuron is 0. A vector of 0 or 1 is taken as the feature representation O of the target.

Citation Information

Patent Citations

  • An image feature description method based on impulse neural network

    CN109214395A

  • Image classification method based on cluster recurrent neural network

    CN110543888A