Image feature extraction method and device, electronic equipment and medium
By splitting the matrix of the target image into sub-matrixes and inputting it into the self-attention model, the problem of excessive calculation of the self-attention mechanism is solved, and the efficiency of image feature extraction and recognition is improved.
Patent Information
- Application Number
- CN202311783047.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
During the image recognition process, the self-attention mechanism requires multiple matrix calculations, which leads to excessive calculation amount and long time, affecting the efficiency of image feature extraction.
The target matrix corresponding to the target image is split into at least two submatrices, and the input data is determined based on the submatrix and the corresponding weight matrix, and input into the self-attention model to obtain the feature matrix.
By reducing the input data size of the self-attention model and increasing the number of input data, the calculation amount is reduced, the efficiency of image feature extraction is improved, and the efficiency of image recognition is improved.
Smart Images

Figure CN120198674A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information recognition technology, and in particular, to an image feature extraction method, apparatus, electronic device, and medium. Background Art
[0002] With the rapid development of neural networks, neural networks are applied to solve technical problems in more and more scenarios. In information recognition, a neural network model has the capabilities of self-learning, associative storage, and high-speed searching for optimal solutions. It can achieve more abstract and complex feature learning through multi-level feature representations, effectively improving the accuracy and robustness of information recognition. In the feature extraction stage of the image recognition process, the self-attention mechanism can globally perceive the front and back correlation information of the input, improving the model's expression ability and reasoning ability, and is widely used in image feature extraction.
[0003] The self-attention mechanism involves multiple matrix calculations. When the matrix dimension corresponding to the target image is large, the amount of calculation is too large and the time consumption is too long when processed by the self-attention mechanism, affecting the efficiency of image feature extraction. Summary of the Invention
[0004] Embodiments of this application provide an image feature extraction method, apparatus, electronic device, and medium to reduce the amount of calculation in information recognition and thus improve the efficiency of image feature extraction.
[0005] According to one aspect of this application, an image feature extraction method is provided. The method includes:
[0006] Splitting a target matrix corresponding to a target image into at least two sub-matrices;
[0007] For each sub-matrix, determining input data according to the sub-matrix and the corresponding weight matrix;
[0008] Inputting the input data into the self-attention model to obtain a feature matrix, so as to recognize the target image based on the feature matrix.
[0009] According to one aspect of this application, an image feature extraction apparatus is provided. The apparatus includes:
[0010] A splitting module, configured to split a target matrix corresponding to a target image into at least two sub-matrices;
[0011] An input data determination module, configured to determine input data according to each sub-matrix and the corresponding weight matrix for each sub-matrix;
[0012] An input module, configured to input the input data into the self-attention model to obtain a feature matrix, so as to recognize the target image based on the feature matrix.
[0013] According to another aspect of the present application, there is provided an electronic device, which includes:
[0014] at least one processor; and
[0015] a memory information-recognitionally connected to the at least one processor; wherein,
[0016] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the image feature extraction method according to any embodiment of the present application.
[0017] According to another aspect of the present application, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the image feature extraction method according to any embodiment of the present application when executed.
[0018] In the technical solution of the embodiment of the present application, the target matrix corresponding to the target image is split into at least two sub-matrices; for each sub-matrix, input data is determined according to the sub-matrix and the corresponding weight matrix; the input data is input into the self-attention model to obtain a feature matrix, so as to identify the target image based on the feature matrix. The above solution effectively reduces the amount of calculation and improves the efficiency of the self-attention model for image feature extraction, and further improves the efficiency of image recognition, by reducing the size of the input data of the self-attention model and increasing the number of input data, compared with directly calculating the input data of a larger size.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0021] Figure 1 is a flowchart of an image feature extraction method provided in Embodiment 1 of the present application;
[0022] Figure 2 is a flowchart of an image feature extraction method provided in Embodiment 2 of the present application;
[0023] Figure 3It is a flowchart of an image feature extraction method provided in Embodiment 3 of this application;
[0024] Figure 4 It is a schematic diagram of the target matrix provided in Embodiment 3 of this application;
[0025] Figure 5 It is a schematic diagram of the first splitting provided in Embodiment 3 of this application;
[0026] Figure 6 It is a schematic diagram of the second splitting provided in Embodiment 3 of this application;
[0027] Figure 7 It is a schematic diagram of the sub - matrix splicing provided in Embodiment 3 of this application;
[0028] Figure 8 It is a schematic diagram of the feature extraction loop process provided in Embodiment 3 of this application;
[0029] Figure 9 It is a schematic diagram of the structure of an image feature extraction device provided in Embodiment 4 of this application;
[0030] Figure 10 It is a schematic diagram of the structure of an electronic device provided in Embodiment 5 of this application. Detailed implementation manners
[0031] In order to enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0032] It should be noted that the terms "first", "second", "third", "fourth", "actual", "preset", etc. in the specification and claims of this application and the above - mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0033] Embodiment 1
[0034] Figure 1 The flowchart of an image feature extraction method provided in the first embodiment of the present application is applicable to the situation of feature extraction. Typically, the embodiments of the present application are applicable to the situation of feature extraction using the self-attention mechanism. This method can be executed by an image feature extraction device, which can be implemented in the form of hardware and / or software, and the image feature extraction device can be configured in an electronic device. As Figure 1 shown, the method includes:
[0035] S110. Split the target matrix corresponding to the target image into at least two sub-matrices.
[0036] Among them, the target image can be an image collected and updated in real time by an image collector, or a static image transmitted by other devices or stored locally. The target matrix is data in matrix form obtained by converting the target image.
[0037] In the embodiments of the present application, the target matrix is split into at least two sub-matrices, and the number of sub-matrices and the dimensions of each sub-matrix are not limited. It can be split into at least two sub-matrices with the same dimension, or into sub-matrices with different dimensions. Each sub-matrix can be n*n-dimensional or m*n-dimensional.
[0038] In the embodiments of the present application, after converting the target image into a target matrix, the target matrix can be split into at least two sub-matrices, or the target image can be first segmented into at least two sub-data, and then the at least two sub-data are converted into at least two sub-matrices. For example, the target image can be first segmented to obtain at least two sub-images, and then corresponding sub-matrices are formed according to the gray values of the pixels of the at least two sub-images to obtain at least two sub-matrices. Since the image pixels include the gray values of the R, G, and B channels, each sub-image correspondingly forms three sub-matrices. If the size of each sub-matrix is m*n, the scale of the sub-matrices of each image is m*n*c, where c is 3.
[0039] S120. For each sub-matrix, determine the input data according to the sub-matrix and the corresponding weight matrix.
[0040] Among them, the weight matrix is a matrix in the self-attention mechanism used to multiply with the sub-matrix to obtain the query matrix, key matrix, and value matrix, and is obtained through training. The input data includes the input matrix, that is, the query matrix, key matrix, and value matrix, and other input parameters determined according to the input matrix.
[0041] Exemplarily, for each sub-matrix, the sub-matrix is multiplied with three weight matrices respectively to obtain three input matrices, that is, the query matrix, key matrix, and value matrix. Other input parameters are determined according to the dimension of the key matrix.
[0042] In the embodiment of the present application, the dimension of the weight matrix is determined according to the dimension of the sub-matrix, so that the number of rows of the weight matrix is equal to the number of columns of the sub-matrix. If the dimensions of each sub-matrix are different, the dimensions of the weight matrices corresponding to each sub-matrix are also different. If the dimensions of each sub-matrix are the same, the dimensions of the weight matrices corresponding to each sub-matrix are the same, and different sub-matrices may also correspond to the same weight matrix. For example, assuming that sub-matrix A is 4*3-dimensional and sub-matrix B is 6*3-dimensional, the weight matrix a corresponding to sub-matrix A can be 3*4-dimensional, and the weight matrix b corresponding to sub-matrix B can be 3*6-dimensional. If sub-matrix A is 4*4-dimensional and sub-matrix B is also 4*4-dimensional, the weight matrix A corresponding to sub-matrix A is 4*4-dimensional, and the weight matrix b corresponding to sub-matrix B is also 4*4-dimensional. It is also possible that the weight matrices corresponding to both sub-matrix A and sub-matrix B are a or both are b.
[0043] S130. Input the input data into the self-attention model to obtain a feature matrix, so as to identify the target image based on the feature matrix.
[0044] Exemplarily, the input data is input into the self-attention model to identify the features of the target image and obtain a feature matrix. Specifically, the formula of the self-attention model is where Q is the query matrix, K is the key matrix, V is the value matrix, d k is the dimension of the key matrix, and the above are all input data. softmax is a normalization operation. The input data is input into the formula of the self-attention model to obtain a feature matrix.
[0045] In the embodiment of the present application, each input data is obtained by multiplying each sub-matrix by a weight matrix. The target matrix includes at least two sub-matrices, and thus corresponds to at least two input data. After obtaining the feature matrix according to the input data, the feature matrices obtained according to the input data corresponding to each sub-matrix can be concatenated to obtain the feature matrix corresponding to the target matrix, representing the feature matrix of the target image.
[0046] In the embodiment of the present application, assuming that c target matrices of m*n dimension are directly input into the self-attention model for calculation, the amount of calculation generated is 3mnc 2 +(mn) 2 c+(mn) 2 c+mnc 2 =4mnc 2 +2(mn) 2 c. In the embodiment of the present application, assuming that the size of the sub-matrix is h*h, the target matrix is split into a total of sub-matrices, and the amount of calculation generated is Substituting specific data can show that the computational complexity generated by the solution of the embodiment of the present application is effectively reduced.
[0047] In the embodiment of the present application, after obtaining the feature matrix, it can be input into the transformer model to identify the target image.
[0048] The technical solution of the embodiment of the present application splits the target matrix corresponding to the target image into at least two sub-matrices; for each sub-matrix, input data is determined according to the sub-matrix and the corresponding weight matrix; the input data is input into the self-attention model to obtain a feature matrix, so as to identify the target image based on the feature matrix. The above solution effectively reduces the computational complexity and improves the efficiency of the self-attention model for feature extraction of the target image by reducing the size of the input data of the self-attention model and increasing the number of input data, compared with directly calculating the input data of a larger size, thereby improving the image recognition efficiency.
[0049] Embodiment 2
[0050] Figure 2 The flowchart of an image feature extraction method provided in the second embodiment of the present application is based on the above embodiment for optimization. For the solutions not described in detail in the second embodiment of the present application, refer to the above embodiment. As Figure 2 shown, the method of the embodiment of the present application specifically includes the following steps:
[0051] S210. Split the target matrix corresponding to the target image into at least two sub-matrices.
[0052] S220. For each sub-matrix, determine the input data according to the sub-matrix and the corresponding weight matrix.
[0053] S230. Input the input data into the self-attention model to obtain a feature matrix.
[0054] S240. Take the feature matrix as the target matrix, and form a first sub-matrix with a preset number of elements that are the closest and belong to different sub-matrices in the target matrix.
[0055] The feature matrix obtained through S210-S230 is the feature matrix corresponding to at least two sub-matrices. Each feature matrix is independent and does not reflect the correlation information between the sub-matrices, nor does it contain the correlation relationship between the sub-matrix regions in the target matrix, that is, it does not reflect the correlation information between the sub-images corresponding to the sub-matrices. For example, there is a common object in two sub-images. In the embodiment of the present application, the feature matrix is used as the target matrix, and the target matrix is further split to divide the elements belonging to different sub-matrices into the same sub-matrix to learn the correlation information between different sub-matrices.
[0056] Exemplarily, a first sub-matrix is formed by a preset number of elements that are the closest in distance and belong to different sub-matrices in the target matrix, so as to converge the elements belonging to different sub-matrices and learn the correlation information between different sub-matrices. Assume the target matrix is Four sub-matrices with dimensions of 2*2 are obtained in the first splitting. In this splitting, assume the preset number is 2, then the elements that are the closest in distance and belong to different sub-matrices include f 12 and f 13 、f 22 and f 23 、f 32 and f 33 、f 42 and f 43 、f 21 and f 31 、f 22 and f 32 、f 23 and f 33 、f 24 and f 34 . f 12 and f 13 、f 22 and f 23 can be merged to obtain the first sub-matrix. f 32 and f 33 、f 42 and f 43 can also be merged to obtain the first sub-matrix. f 21 and f 31 、f 22 and f 32 can also be merged to obtain the first sub-matrix. f 23 and f 33 、f 24 and f 34 can also be merged to obtain the first sub-matrix. If the preset number is 4, f 22 and f 23 、f 32 and f 33 can be merged to obtain the first sub-matrix.
[0057] S250. The elements that belong to the same sub-matrix and are other than the elements of the first sub-matrix are used as the second sub-matrix.
[0058] Exemplarily, the elements that belong to the same sub-matrix and are other than the elements of the first sub-matrix in the target matrix are used as the second sub-matrix. For example, assume in the target matrix F, f 12 and f 13 、f 22 and f 23 are merged to obtain the first sub-matrix, and f32 and f 33 、f 42 and f 43 are combined to obtain the first sub - matrix, then f 11 and f 21 are combined as the second sub - matrix, and f 14 and f 24 are combined as the second sub - matrix, and f 31 and f 41 are combined as the second sub - matrix, and f 34 and f 44 are combined as the second sub - matrix.
[0059] S260. Determine the input data according to the first sub - matrix, the second sub - matrix and the corresponding weight matrix, and input the input data into the self - attention model to update the feature matrix.
[0060] Exemplarily, determine the first sub - matrix, the second sub - matrix and the corresponding weight matrix, multiply the first sub - matrix and the weight matrix, multiply the second sub - matrix and the weight matrix, determine the input data, input the input data into the self - attention model, and then splice the output data to obtain the updated feature matrix.
[0061] In the process of splitting the target matrix into the first sub - matrix and the second sub - matrix above, elements that meet the conditions can be merged to form the first sub - matrix and the second sub - matrix simultaneously, and then S260 is executed. Or some elements that meet the conditions can be merged to form the first sub - matrix and the second sub - matrix first, and then S260 is executed. Then the remaining elements that meet the conditions are merged to form the first sub - matrix and the second sub - matrix, and then S260 is executed. Traverse and merge a preset number of elements that are closest in distance and belong to different sub - matrices once, and extract the feature matrix to more fully learn the correlation information between elements belonging to different sub - matrices, that is, be able to learn the correlation relationship between adjacent sub - image regions, thereby improving the accuracy of image feature extraction.
[0062] In the embodiment of the present application, after taking the elements other than the elements of the first sub - matrix that belong to the same sub - matrix as the second sub - matrix, the method further includes:
[0063] Splice at least two second sub - matrices so that the size of the new second sub - matrix after splicing is the same as the size of the first sub - matrix.
[0064] Exemplarily, the sizes of the determined first sub-matrix and second sub-matrix may be inconsistent, affecting subsequent calculation efficiency. Therefore, at least two second sub-matrices can be spliced together so that the size of the new second sub-matrix after splicing is the same as that of the first sub-matrix. For example, in the above example, the first sub-matrix is 2*2 dimensional and the second sub-matrix is 1*2 dimensional. Two second sub-matrices can be spliced together to obtain a new 2*2 dimensional second sub-matrix, so that the size of the first sub-matrix is the same as that of the new second sub-matrix, and the first sub-matrix and the second sub-matrix can be processed in parallel during subsequent execution to improve processing efficiency.
[0065] In the embodiments of the present application, the input data includes a query matrix, a key matrix, and a value matrix generated according to the new second sub-matrix and the corresponding weight matrix;
[0066] Inputting the input data into the self-attention model to update the feature matrix includes:
[0067] Inputting the query matrix and the key matrix into the self-attention model to obtain attention scores;
[0068] Subtracting a preset score from the attention scores, normalizing them, and then multiplying them by the value matrix to update the feature matrix.
[0069] Exemplarily, in the above embodiment, at least two second sub-matrices are spliced together to obtain a new second sub-matrix. However, the elements in the new second sub-matrix may not be adjacent in the target matrix, and the correlation is not strong. The adjacent elements formed after splicing together enhance the correlation, which does not conform to the actual situation of the target matrix. During the process of feature extraction of the second sub-matrix, a query matrix, a key matrix, and a value matrix can be formed according to the second sub-matrix and the corresponding weight matrix. According to the formula of the self-attention model, calculate the product of the query matrix and the key matrix, and then divide it by the square root of the dimension of the key matrix to obtain the attention scores. Since the correlation of the elements in the second sub-matrix is not strong, the preset score is subtracted from the attention scores as the actual attention scores of the second sub-matrix. Then, the actual attention scores are normalized so that the normalized actual attention scores are close to 0, which conforms to the correlation characteristics of the elements in the target matrix. The preset score can be determined according to the actual situation, such as 100, 50, 10, etc.
[0070] S270. Identify the target image based on the feature matrix.
[0071] An embodiment of the present application provides an image feature extraction method. The feature matrix is used as the target matrix, and a first sub-matrix is formed by a preset number of elements that are the closest in distance and belong to different sub-matrices in the target matrix; the elements that belong to the same sub-matrix and are other than the elements of the first sub-matrix are used as the second sub-matrix; according to the first sub-matrix, the second sub-matrix, and the corresponding weight matrix, input data is determined, and the input data is input into the self-attention model to obtain a feature matrix. Through the above solution, the elements that were split apart during the previous split can be combined into a sub-matrix again to further extract features, and the correlation information between different image regions can be refined, and the feature information of the image can be obtained more accurately and richly.
[0072] Embodiment III
[0073] Figure 3 The following is a flowchart of an image feature extraction method provided by Embodiment III of the present application. The embodiments of the present application are optimized based on the above embodiments. For the solutions not described in detail in the embodiments of the present application, please refer to the above embodiments. As Figure 3 shown, the method of the embodiment of the present application specifically includes the following steps:
[0074] S310: Make a row splitting line at the midpoint of all rows of the target matrix, and make a column splitting line at the midpoint of all columns of the target matrix, and split the target matrix into four sub-matrices with the same size.
[0075] Exemplarily, as Figure 4 shown and Figure 5 shown, assuming the size of the target matrix is 4*4, then a row splitting line is made at the midpoint of all rows of the target matrix, and a column splitting line is made at the midpoint of all columns, and the target matrix is split into four sub-matrices with the same size, as Figure 5 shown, and the size of each sub-matrix is 2*2. Figure 4 and Figure 5 are just examples, and the actual size is determined according to the size of the target image.
[0076] S320: For each sub-matrix, determine input data according to the sub-matrix and the corresponding weight matrix.
[0077] S330: Input the input data into the self-attention model to obtain a feature matrix.
[0078] S340: Use the feature matrix as the target matrix, make row splitting lines at the quarter and three-quarter points of all rows of the target matrix, and make column splitting lines at the quarter and three-quarter points of all columns of the target matrix, and split the target matrix into nine sub-matrices.
[0079] Exemplarily, the output data of the self-attention model in S330 are concatenated to obtain a feature matrix, and the feature matrix is used as the target matrix. Row splitting lines are made at one quarter and three quarters of all rows of the target matrix, and column splitting lines are made at one quarter and three quarters of all columns of the target matrix, and the target matrix is split into nine sub-matrices, such as Figure 6 shown.
[0080] S350: Use the matrix at the center as the first sub-matrix, and use the other sub-matrices as the second sub-matrix.
[0081] For example, Figure 6 5 in the figure is used as the first sub-matrix, and the other sub-matrices are used as the second sub-matrix to distinguish the sub-matrix with the largest size from the other sub-matrices.
[0082] S360, splicing the second sub-matrices located at the four vertices in the target matrix, splicing the second sub-matrices located above and below the first sub-matrix, and splicing the second sub-matrices located on the left and right of the first sub-matrix, to obtain three second sub-matrices with the same size as the first sub-matrix.
[0083] Exemplarily, the second submatrix may be spliced based on the first submatrix, so that the size of the new second submatrix obtained by splicing is consistent with the size of the first submatrix. Figure 7 As shown, the second submatrix 2 and the second submatrix 8 are spliced to obtain a new second submatrix with the same size as the first submatrix. The second submatrix 4 and the second submatrix 6 are spliced to obtain a new second submatrix with the same size as the first submatrix. The second submatrix 1, 2, 7, and 9 are spliced to obtain a new second submatrix with the same size as the first submatrix. In this case, the first submatrix and the three new second submatrices can be calculated in parallel, which improves processing efficiency.
[0084] Figure 4 It can also be regarded as the target image. Figure 4 is the complete target image, Figure 5 To divide the target image into 2*2 sub-images, each sub-image corresponds to a sub-matrix. Figure 6 The vertical dividing line at one-third of the horizontal position of the target image, the vertical dividing line at two-thirds, the horizontal dividing line at one-third of the vertical position, and the horizontal dividing line at one-third of the vertical position are used to compare the horizontal dividing line. Figure 4 Divide the target image in , and get 9 sub-images. Sub-image 5 corresponds to the first sub-matrix, and the other sub-images correspond to the second sub-matrix. Then translate sub-images 1, 2, and 3 to the bottom of sub-images 7, 8, and 9, and translate sub-images 1, 4, and 7 to the right of sub-images 3, 6, and 9 to get the spliced Figure 7 .exist Figure 7Among them, sub-image 5 corresponds to the first sub-matrix, sub-images 4 and 6 correspond to a second sub-matrix after splicing, sub-images 2 and 8 correspond to a second sub-matrix after splicing, and sub-images 1, 3, 7, and 9 correspond to a second sub-matrix after splicing.
[0085] S370. Determine input data according to the first sub-matrix, the second sub-matrix, and the corresponding weight matrix, and input the input data into the self-attention model to update the feature matrix.
[0086] In the embodiment of the present application, since the second sub-matrix 2 and the second sub-matrix 8, the second sub-matrix 4 and the second sub-matrix 6, and the second sub-matrices 1, 2, 7, and 9 are not adjacent in the target matrix and have no correlation, a preset score is subtracted when calculating the attention score to reduce the attention score corresponding to the second sub-matrix.
[0087] In the embodiment of the present application, after the input data is input into the self-attention model to obtain a feature matrix, the method further includes:
[0088] Taking the feature matrix as the target matrix, continue to execute the step of splitting the target matrix into at least two sub-matrices, determining input data according to the at least two sub-matrices and the corresponding weight matrix, and inputting the input data into the self-attention model to obtain a feature matrix until the preset number of times is looped;
[0089] Among them, the number of at least two sub-matrices obtained by splitting the target matrix during at least two loop executions increases or decreases.
[0090] Exemplarily, in order to expand the field of view to extract features or narrow the field of view to extract features, the target matrix can be split multiple times according to different splitting scales to loop and execute the feature extraction process. Exemplarily, as Figure 8 shown. The target matrix can be first split into 4 sub-matrices, feature extraction is performed through the self-attention mechanism, and then the feature matrix is split into 8 sub-matrices, and so on. The preset number of times of the feature extraction loop is not limited. It is also possible to Figure 8 reverse the process in, first split the target matrix into 32 sub-matrices, perform feature extraction through the self-attention mechanism, and then split the feature matrix into 16 sub-matrices, and so on. Before each self-matrix split, the target matrix can be downsampled to reduce the resolution, improve the processing efficiency, and reduce the calculation amount.
[0091] S380. Identify the target image based on the feature matrix.
[0092] In an embodiment of the present application, a specific application example is lip reading recognition. Still rationally, video samples are prepared in advance, in which the facial and lip images are clear, and the video samples can clearly extract the text information corresponding to the speech. The video samples are processed, and the lip region is located from the image sequence of the video samples by using a face key point detection algorithm, and the lip regions in each image frame are intercepted and saved as a lip image sample set. The lip reading features are extracted from the lip image sample set by using the image feature extraction method in the above embodiment to obtain lip reading features. The lip reading recognition model is trained according to the lip reading features and the text information. During the use of the lip reading recognition model, first process the target video, locate the lip region from the image sequence of the target video by using a face key point detection algorithm, and intercept the lip regions in each image frame to obtain lip reading images. Then, use the image feature extraction method in the above embodiment to extract features from the lip reading images to obtain lip reading features, and input the lip reading features into the lip reading recognition model to obtain the corresponding text.
[0093] The embodiment of the present application provides an image feature extraction method. A row splitting line is made at the midpoint of all rows of the target matrix corresponding to the target image, and a column splitting line is made at the midpoint of all columns of the target matrix, and the target matrix is split into four sub-matrices of the same size; for each sub-matrix, input data is determined according to the sub-matrix and the corresponding weight matrix; the input data is input into the self-attention model to obtain a feature matrix. The feature matrix is used as the target matrix. A row splitting line is made at the quarter and three-quarter points of all rows of the target matrix, and a column splitting line is made at the quarter and three-quarter points of all columns of the target matrix, and the target matrix is split into nine sub-matrices; the matrix at the central position is used as the first sub-matrix, and the other sub-matrices are used as the second sub-matrices. By using the above-mentioned quartering method and nine-partition method to determine sub-matrices for target image feature extraction, the calculation amount can be effectively reduced, and the correlation information between different image regions can be comprehensively extracted, improving the accuracy of image feature extraction.
[0094] Embodiment 4
[0095] Figure 9 It is a schematic structural diagram of an image feature extraction device provided in Embodiment 4 of the present application. The device can execute the image feature extraction method provided in any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. As Figure 9 shown, the device includes:
[0096] A splitting module 410, configured to split the target matrix corresponding to the target image into at least two sub-matrices;
[0097] An input data determination module 420, configured to determine input data for each sub-matrix according to the sub-matrix and the corresponding weight matrix;
[0098] An input module 430 for inputting the input data into the self-attention model to obtain a feature matrix, and for recognizing a target image based on the feature matrix.
[0099] In an embodiment of the present application, the apparatus further includes:
[0100] A first sub-matrix determination module for using the feature matrix as a target matrix, and for forming a first sub-matrix with a preset number of elements that are closest in distance and belong to different sub-matrices in the target matrix;
[0101] A second sub-matrix determination module for using, as a second sub-matrix, elements other than the elements of the first sub-matrix that belong to the same sub-matrix;
[0102] A feature extraction module for determining input data according to the first sub-matrix, the second sub-matrix, and a corresponding weight matrix, and for inputting the input data into the self-attention model to update the feature matrix.
[0103] In an embodiment of the present application, the apparatus further includes:
[0104] A splicing module for splicing at least two second sub-matrices so that the size of the spliced new second sub-matrix is the same as the size of the first sub-matrix.
[0105] In an embodiment of the present application, the input data includes a query matrix, a key matrix, and a value matrix generated according to the new second sub-matrix and a corresponding weight matrix;
[0106] The feature extraction module is specifically configured to:
[0107] Input the query matrix and the key matrix into the self-attention model to obtain attention scores;
[0108] Subtract a preset score from the attention scores, normalize them, and then multiply them by the value matrix to update the feature matrix.
[0109] In an embodiment of the present application, the splitting module 410 is specifically configured to:
[0110] Make a row splitting line at the midpoint of all rows of the target matrix, make a column splitting line at the midpoint of all columns of the target matrix, and split the target matrix into four sub-matrices of the same size;
[0111] The apparatus further includes:
[0112] A sub - matrix splitting module, which is used to draw row splitting lines at the one - quarter and three - quarters positions of all rows of the target matrix, and draw column splitting lines at the one - quarter and three - quarters positions of all columns of the target matrix, so as to split the target matrix into nine sub - matrices;
[0113] A sub - matrix determination module, which is used to take the matrix in the central position as the first sub - matrix and the other sub - matrices as the second sub - matrices.
[0114] In an embodiment of the present application, the device further includes:
[0115] A sub - matrix splicing module, which is used to splice the second sub - matrices located at the four vertices of the target matrix, splice the second sub - matrices located above and below the first sub - matrix, and splice the second sub - matrices located on the left and right sides of the first sub - matrix, so as to obtain three second sub - matrices with the same size as the first sub - matrix.
[0116] In an embodiment of the present application, the device further includes:
[0117] A continue - execution module, which is used to take the feature matrix as the target matrix, and continue to execute the step of splitting the target matrix into at least two sub - matrices, so as to determine input data according to the at least two sub - matrices and the corresponding weight matrix, and input the input data into the self - attention model to obtain the feature matrix, until the preset number of times of loop execution;
[0118] Wherein, the number of at least two sub - matrices obtained by splitting the target matrix during at least two loop executions increases or decreases.
[0119] An image feature extraction device provided by an embodiment of the present application can execute an image feature extraction method provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method.
[0120] Embodiment Five
[0121] Figure 10 Fig. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described herein and / or claimed.
[0122] As Figure 10As shown, the electronic device 10 includes at least one processor 11 and a memory information-recognitionally connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0123] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and an information recognition unit 19, such as a network card, a modem, a wireless information recognition transceiver, etc. The information recognition unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0124] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the image feature extraction method.
[0125] In some embodiments, the image feature extraction method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the information recognition unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the image feature extraction method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the image feature extraction method by any other appropriate means (for example, by means of firmware).
[0126] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0127] The computer programs for implementing the methods of this application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable image feature extraction device, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0128] In the context of this application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0129] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0130] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data information identification (e.g., an information identification network). Examples of information identification networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0131] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through an information identification network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0132] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the information desired by the technical solution of this application can be achieved, and this is not limited herein.
[0133] The above specific implementation manners do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. An image feature extraction method, characterized in that, The method includes: splitting a target matrix corresponding to a target image into at least two sub - matrices; for each sub - matrix, determining input data according to the sub - matrix and the corresponding weight matrix; inputting the input data into the self - attention model to obtain a feature matrix, so as to recognize the target image based on the feature matrix.
2. The method according to claim 1, characterized in that, After inputting the input data into the self - attention model to obtain a feature matrix, the method further includes: taking the feature matrix as the target matrix, and forming a first sub - matrix with a preset number of elements that are closest to each other and belong to different sub - matrices in the target matrix; taking the elements other than the elements of the first sub - matrix that belong to the same sub - matrix as the second sub - matrix; determining input data according to the first sub - matrix, the second sub - matrix and the corresponding weight matrix, and inputting the input data into the self - attention model to update the feature matrix.
3. The method according to claim 2, wherein After taking the elements other than the elements of the first sub - matrix that belong to the same sub - matrix as the second sub - matrix, the method further includes: concatenating at least two second sub - matrices so that the size of the new concatenated second sub - matrix is the same as the size of the first sub - matrix.
4. The method according to claim 3, characterized in that The input data includes a query matrix, a key matrix, and a value matrix generated according to the new second sub - matrix and the corresponding weight matrix; Inputting the input data into the self - attention model to update the feature matrix includes: inputting the query matrix and the key matrix into the self - attention model to obtain attention scores; subtracting a preset score from the attention scores, normalizing them, and then multiplying by the value matrix to update the feature matrix.
5. The method according to claim 3, characterized in that, Splitting the target matrix corresponding to the target image into at least two sub - matrices includes: making a row splitting line at the mid - point of all rows of the target matrix and a column splitting line at the mid - point of all columns of the target matrix, and splitting the target matrix into four sub - matrices of the same size; The determination process of the first sub - matrix and the second sub - matrix includes: making a row splitting line at the one - fourth and three - fourth positions of all rows of the target matrix and a column splitting line at the one - fourth and three - fourth positions of all columns of the target matrix, and splitting the target matrix into nine sub - matrices; taking the matrix at the central position as the first sub - matrix and the other sub - matrices as the second sub - matrices.
6. The method according to claim 5, wherein After taking the matrix at the central position as the first sub - matrix and the other sub - matrices as the second sub - matrices, the method further includes: concatenating the second sub - matrices at the four vertices of the target matrix, concatenating the second sub - matrices above and below the first sub - matrix, and concatenating the second sub - matrices on the left and right sides of the first sub - matrix to obtain three second sub - matrices with the same size as the first sub - matrix.
7. The method according to claim 1, wherein After inputting the input data into the self - attention model to obtain a feature matrix, the method further includes: taking the feature matrix as the target matrix, and continuing to execute the step of splitting the target matrix into at least two sub - matrices, so as to determine input data according to at least two sub - matrices and the corresponding weight matrix, and inputting the input data into the self - attention model to obtain a feature matrix, until the loop is executed a preset number of times; Among them, the number of at least two sub-matrices obtained by splitting the target matrix during at least two loop executions increases or decreases.
8. An image feature extraction device, characterized in that The device includes: A splitting module, configured to split the target matrix corresponding to the target image into at least two sub-matrices; An input data determination module, configured to determine input data for each sub-matrix according to the sub-matrix and the corresponding weight matrix; An input module, configured to input the input data into the self-attention model to obtain a feature matrix, so as to recognize the target image based on the feature matrix.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory information recognition-connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute the image feature extraction method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the image feature extraction method according to any one of claims 1-7 when executed by a processor.