An image data processing method and apparatus

By establishing the correlation matrix and pixel-level spatial position matrix in the image data processing method, the problem of low accuracy of image object detection in the prior art is solved, and higher detection accuracy and affine transformation prediction are achieved.

CN114926666BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210642575.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-08
Publication Date
2025-07-18
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

The existing image-based object detection method has low detection accuracy.

Method used

The correlation matrix establishes the similarity between the pixels in the feature map of the image to be detected and the pixels in the feature map of the target object image, and establishes the spatial position relationship between the pixels to be detected and the similar pixels of the target object image according to the pixel-level spatial position matrix, and finally detects the target object in the image to be detected.

Benefits of technology

Improves the detection accuracy of the target object and can predict the affine transformation of the target object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926666B_ABST
    Figure CN114926666B_ABST
Patent Text Reader

Abstract

The present application provides an image data processing method and related devices. Embodiments of the present application can be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving. The method includes: obtaining an image to be detected and a target object image; extracting features from the image to be detected and the target object image to obtain two feature image data; generating a correlation matrix according to the feature image data; generating a pixel-level spatial position matrix through the correlation matrix; and generating a target object detection frame containing the target object in the image to be detected according to the pixel-level spatial position matrix. Embodiments of the present application provide an image data processing method, which establishes the similarity degree between pixels in the feature map of the image to be detected and pixels in the feature map of the target object image through the correlation matrix, and establishes the spatial position relationship of similar pixels between the image to be detected and the target object image through the pixel-level spatial position matrix, thereby improving the accuracy of detecting the target object from the image to be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly relates to an image data method and apparatus. Background Art

[0002] With the development of technology, object detection technology is used more and more widely. The image-based object detection technology refers to the technology of detecting the target objects included in an image, which is a common image processing method and is widely applied to detection tasks of objects such as commodities, landmarks, and pets.

[0003] Existing image-based object detection methods are mainly divided into two categories. One is the end-to-end object detection method based on a deep neural network, and the other is the object detection method that combines image region prediction and image retrieval. However, both of these methods have the problem of low detection accuracy. Summary of the Invention

[0004] The embodiments of this application provide an image data processing method and related apparatus. First, a correlation matrix is established to represent the similarity degree between the pixels in the feature map of the image to be detected and the pixels in the feature map of the target object image. Then, according to the pixel-level spatial position matrix, the spatial position relationship of the similar pixels between the image to be detected and the target object image is established. Finally, the target object is detected in the image to be detected according to the pixel-level spatial position matrix, improving the detection accuracy of the target object.

[0005] One aspect of this application provides an image data processing method, including:

[0006] Obtain the image to be detected and the target object image;

[0007] Respectively use the image to be detected and the target object image as the input of the feature extraction network in the single-sample detection model, and respectively output a first feature image and a second feature image through the feature extraction network. Among them, the first feature image is generated by the feature extraction network according to the image to be detected, and the first feature image includes K first feature pixels. The second feature image is generated by the feature extraction network according to the target object image, and the second feature image includes L second feature pixels. K is an integer greater than 1, and L is an integer greater than 1;

[0008] Generate a correlation matrix according to the first feature image and the second feature image. Among them, the correlation matrix includes K×L similarity values, and the K×L similarity values represent the similarity degree between K first feature pixels and L second feature pixels;

[0009] Use the correlation matrix as the input of the transformation network in the single-sample detection model, and output a pixel-level spatial position matrix through the transformation network. Here, the pixel-level spatial position matrix includes K×L×2 elements, and the K×L×2 elements represent the corresponding position coordinates of L second feature pixels in the first feature image when any one of the K first feature pixels is used as an anchor point;

[0010] Generate T target object detection frames in the image to be detected according to the pixel-level spatial position matrix. Here, the T target object detection frames include T target objects, and the T confidence values corresponding to the T target object detection frames all meet the confidence threshold, where T is an integer greater than or equal to 0.

[0011] Another aspect of this application provides an image data processing device, including:

[0012] An image acquisition module for acquiring the image to be detected and the target object image;

[0013] A feature extraction module for respectively using the image to be detected and the target object image as the input of the feature extraction network in the single-sample detection model, and respectively outputting a first feature image and a second feature image through the feature extraction network. Here, the first feature image is generated by the feature extraction network according to the image to be detected, and the first feature image includes K first feature pixels. The second feature image is generated by the feature extraction network according to the target object image, and the second feature image includes L second feature pixels, where K is an integer greater than 1 and L is an integer greater than 1;

[0014] A correlation matrix generation module for generating a correlation matrix according to the first feature image and the second feature image. Here, the correlation matrix includes K×L similarity values, and the K×L similarity values represent the similarity degree between the K first feature pixels and the L second feature pixels;

[0015] A pixel-level spatial position matrix generation module for using the correlation matrix as the input of the transformation network in the single-sample detection model, and outputting a pixel-level spatial position matrix through the transformation network. Here, the pixel-level spatial position matrix includes K×L×2 elements, and the K×L×2 elements represent the corresponding position coordinates of L second feature pixels in the first feature image when any one of the K first feature pixels is used as an anchor point;

[0016] A detection frame generation module for generating T target object detection frames in the image to be detected according to the pixel-level spatial position matrix. Here, the T target object detection frames include T target objects, and the T confidence values corresponding to the T target object detection frames all meet the confidence threshold, where T is an integer greater than or equal to 0.

[0017] In another implementation manner of the embodiment of the present application, the K first feature pixels correspond to K detection frames generated with the first feature pixels as anchor points; the image data processing device further includes: the confidence matrix calculation module is used for:

[0018] Resample the correlation matrix according to the pixel-level spatial position matrix to generate a resampled matrix, where the resampled matrix includes U dimensions, and U is an integer greater than or equal to 2;

[0019] Perform average pooling processing on the U dimensions in the resampled matrix to obtain a confidence matrix, where the confidence matrix includes K confidence values, and the K confidence values correspond to the K detection frames.

[0020] In another implementation manner of the embodiment of the present application, the detection frame generation module includes: a detection frame first generation sub-module, which is used for:

[0021] Generate T corresponding grids in the image to be detected according to the pixel-level spatial position matrix, and the corresponding grids are the corresponding grids of the similar pixels of the image to be detected and the target object image;

[0022] Generate T target object detection frames according to the T corresponding grids, where each target object detection frame is the circumscribed rectangle of each corresponding grid.

[0023] In another implementation manner of the embodiment of the present application, the detection frame generation module includes: a detection frame second generation sub-module, which is used for:

[0024] Determine the vertex coordinates of the target object in the image to be detected according to the pixel-level spatial position matrix;

[0025] Generate a target object detection frame according to the vertex coordinates of the target object.

[0026] In another implementation manner of the embodiment of the present application, the feature extraction network includes a convolutional sub-network and a dictionary sub-network, and the dictionary sub-network carries a dictionary matrix; the feature extraction module is further used for:

[0027] Use the image to be detected as the input of the convolutional sub-network, and output a first intermediate matrix through the convolutional sub-network;

[0028] Input the first intermediate matrix into the dictionary sub-network, and perform feature crossing on the first intermediate matrix and the dictionary matrix through the dictionary sub-network to generate a feature matrix of the image to be detected;

[0029] Normalize the feature matrix of the image to be detected to obtain a first feature matrix;

[0030] Generate a first feature image according to the first feature matrix;

[0031] Take the target object image as the input of the convolutional sub-network, and output the second intermediate matrix through the convolutional sub-network;

[0032] Input the second intermediate matrix into the dictionary sub-network, and perform feature crossing between the second intermediate matrix and the dictionary matrix through the dictionary sub-network to generate the target object image feature matrix;

[0033] Normalize the target object image feature matrix to obtain the second feature matrix;

[0034] Generate the second feature image according to the second feature matrix.

[0035] In another implementation manner of the embodiment of the present application, the image data processing device further includes a model training module, and the model training module includes:

[0036] A training image acquisition sub-module, configured to acquire a first training sample image, a second training sample image, and a training object image, where the first training sample image includes T B training object annotation frames, and the T B training object annotation frames include T B training objects, the T B training object annotation frames correspond to T B annotation frame data, the second training sample image does not include training objects, the training object image includes training objects, and T B is an integer greater than or equal to 1;

[0037] A training feature extraction sub-module, configured to respectively take the first training sample image, the second training sample image, and the training object image as the input of the feature extraction network in the single-sample detection model, and respectively output a first training feature image, a second training feature image, and a third training feature image through the feature extraction network, where the first training feature image is generated by the feature extraction network according to the first training sample image, and the first training feature image includes K X1 first training feature pixels, the second training feature image is generated by the feature extraction network according to the second training sample image, and the second training feature image includes K X2 second training feature pixels, the third training feature image is generated by the feature extraction network according to the training object image, and the third training feature image includes L X third training feature pixels, K X1 is an integer greater than 1, K X2 is an integer greater than 1, and L X is an integer greater than 1;

[0038] A training correlation matrix generation sub-module, configured to generate a first training correlation matrix according to the first training feature image and the third training feature image, where the first training correlation matrix includes KX1 ×L X similarity values, K X1 ×L X similarity values are K X1 similarity degrees between a first training feature pixel and L X third training feature pixels;

[0039] The training correlation matrix generation sub-module is also used to generate a second training correlation matrix according to the second training feature image and the third training feature image, where the second training correlation matrix includes K X2 ×L X similarity values, K X2 ×L X similarity values are K X2 similarity degrees between a second training feature pixel and L X third training feature pixels;

[0040] The training pixel-level spatial position matrix generation sub-module is used to take the first training correlation matrix as the input of the transformation network in the single-sample detection model, and output the first training pixel-level spatial position matrix through the transformation network, where the first training pixel-level spatial position matrix includes K X1 ×K X ×2 first training elements, K X1 ×L X ×2 first training elements represent that when any one of the K X1 first training feature pixels is used as an anchor point, L X corresponding position coordinates of the third training feature pixels in the first training feature image;

[0041] The training pixel-level spatial position matrix generation sub-module is also used to take the second training correlation matrix as the input of the transformation network in the single-sample detection model, and output the second training pixel-level spatial position matrix through the transformation network, where the second training pixel-level spatial position matrix includes K X2 ×L X ×2 second training elements, K X2 ×L X ×2 second training elements represent that when any one of the K X2 second training feature pixels is used as an anchor point, L X corresponding position coordinates of the third training feature pixels in the second training feature image;

[0042] The training detection box generation sub-module is used to generate T X1 training object first detection boxes in the first training sample image, where T X1 training object first detection boxes include T X1One training object, T X1 One first detection box corresponding to the training object, T X1 One first training confidence level, T X1 One first detection box corresponding to the training object, T X1 One first detection box data;

[0043] The training detection box generation sub-module is also used to generate T in the second training sample image according to the second training pixel-level spatial position matrix, where X2 One second detection box for the training object, where T X2 One second detection box corresponding to the training object, T X2 One second training confidence level, T X2 One second detection box corresponding to the training object, T X2 One second detection box data;

[0044] The loss result generation sub-module is used to generate the single-sample detection model loss result according to T X1 One first training confidence level, T X1 One first detection box data, T X2 One second training confidence level, T X2 One second detection box data, and T B One annotation box data to generate the single-sample detection model loss result;

[0045] The model training sub-module is used to train the single-sample detection model according to the single-sample detection model loss result.

[0046] In another implementation manner of the embodiment of the present application, the first training sample image corresponds to a first confidence level reference value, and the second training sample image corresponds to a second confidence level reference value; the loss result generation sub-module is further used for:

[0047] According to T X1 One first training confidence level, T X1 One first detection box data, T X2 One second training confidence level, T X2 One second detection box data, and T B One annotation box data to generate the single-sample detection model loss result, including:

[0048] According to T X1 One first training confidence level and the first confidence level reference value to generate the first classification loss result;

[0049] According to T X2 One second training confidence level and the second confidence level reference value to generate the second classification loss result;

[0050] According to T X1 One first detection box data, T X2 One second detection box data, and TB Generate a positioning loss result based on the annotation box data;

[0051] Generate a single-sample detection model loss result based on the first classification loss result, the second classification loss result, and the positioning loss result.

[0052] Another aspect of the present application provides a computer device, including:

[0053] A memory, a transceiver, a processor, and a bus system;

[0054] Wherein, the memory is used to store programs;

[0055] The processor is used to execute the programs in the memory, including executing the methods of the above aspects;

[0056] The bus system is used to connect the memory and the processor, so that the memory and the processor can communicate.

[0057] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a computer, the computer is enabled to execute the methods of the above aspects.

[0058] Another aspect of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.

[0059] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0060] The present application provides an image data processing method and related devices. The method includes: First, obtain an image to be detected and a target object image containing a target object; Next, perform feature extraction on the image to be detected and the target object image respectively to obtain a first feature image and a second feature image; Then, generate a correlation matrix for representing the similarity degree between the pixels in the first feature image and the pixels in the second feature image according to the first feature image and the second feature image; Again, generate a pixel-level spatial position matrix for representing the similarity degree and positional relationship between the pixels in the image to be detected and the pixels in the target object image through the correlation matrix; Finally, generate a target object detection frame containing the target object in the image to be detected according to the pixel-level spatial position matrix. An embodiment of the present application provides an image data processing method. By establishing the similarity degree between the pixels in the feature map of the image to be detected and the pixels in the feature map of the target object image through the correlation matrix, and establishing the spatial position relationship of the similar pixels between the image to be detected and the target object image through the pixel-level spatial position matrix, the accuracy of detecting the target object from the image to be detected is improved; An embodiment of the present application provides an image data processing method, which can detect any target object, and can realize the detection of the target object from the image to be detected only through one target object image. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a schematic diagram of an architecture of an image data processing system provided by an embodiment of the present application;

[0062] Figure 2 It is a flowchart of an image data processing method provided by an embodiment of the present application;

[0063] Figure 3 It is a schematic diagram of the corresponding relationship of the coordinate positions of some similar pixels of the image to be detected and the target object image provided by an embodiment of the present application;

[0064] Figure 4 It is a schematic diagram of generating a target object detection frame in the image to be detected provided by an embodiment of the present application;

[0065] Figure 5 It is a flowchart of an image data processing method provided by another embodiment of the present application;

[0066] Figure 6 It is a flowchart of an image data processing method provided by another embodiment of the present application;

[0067] Figure 7 It is a schematic diagram of the corresponding grid determination process provided by an embodiment of the present application;

[0068] Figure 8 It is a flowchart of an image data processing method provided by another embodiment of the present application;

[0069] Figure 9 Schematic diagram of the vertex coordinate determination process provided for an embodiment of the present application;

[0070] Figure 10 Flowchart of the image data processing method provided for another embodiment of the present application;

[0071] Figure 11 Schematic diagram of the structure of the feature extraction network in the single-sample detection model provided for an embodiment of the present application;

[0072] Figure 12 Flowchart of the image data processing method provided for another embodiment of the present application;

[0073] Figure 13 Flowchart of the image data processing method provided for yet another embodiment of the present application;

[0074] Figure 14 Flowchart of the image data processing method applied to kitten detection provided for an embodiment of the present application;

[0075] Figure 15 Schematic diagram of the structure of the image data processing device provided for an embodiment of the present application;

[0076] Figure 16 Schematic diagram of the structure of the image data processing device provided for another embodiment of the present application;

[0077] Figure 17 Schematic diagram of the structure of the image data processing device provided for yet another embodiment of the present application;

[0078] Figure 18 Schematic diagram of the server structure provided for an embodiment of the present application. Detailed implementation manners

[0079] The embodiments of the present application provide an image data processing method, which establishes the similarity degree between the pixels in the feature map of the image to be detected and the pixels in the feature map of the target object image through a correlation matrix, and establishes the spatial position relationship of the similar pixels between the image to be detected and the target object image through a pixel-level spatial position matrix, thereby improving the accuracy of detecting the target object from the image to be detected.

[0080] The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims, and the above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0081] The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.

[0082] Object detection technology is a technology that is widely used and applied in the field of visual intelligence and can be applied to the detection of target objects in images or the detection of target objects in videos. For example, when positioning a certain commodity on a commodity shelf, the shelf image can be processed by an image-based object detection method to quickly locate the target commodity. Another example is that in an e-commerce live broadcast, when locking and intercepting a live broadcast segment containing a certain commodity, the complete live broadcast video can be processed by a video-based object detection method to lock the video segment containing the target commodity from the complete live broadcast video and then intercept it.

[0083] In view of the problem that the detection accuracy of existing image-based object detection methods is relatively low, the embodiments of this application provide an image data processing method and related device. First, the similarity degree between the pixels in the feature map of the image to be detected and the pixels in the feature map of the target object image is established through a correlation matrix. Then, the spatial position relationship of similar pixels between the image to be detected and the target object image is established according to the pixel-level spatial position matrix. Finally, the target object is detected in the image to be detected according to the pixel-level spatial position matrix, improving the detection accuracy of the target object.

[0084] For ease of understanding, please refer to Figure 1 , Figure 1 which is the application environment diagram of the image data processing method in the embodiments of this application. As Figure 1As shown in the figure, the image data processing method in the embodiment of the present application is applied to an image data processing system. The image data processing system includes: a server and a terminal device; among them, the server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, in-vehicle terminals, etc. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and the embodiment of the present application does not limit this here.

[0085] First, the server obtains the image to be detected sent by the terminal and the target object image containing the target object; secondly, the server respectively extracts features from the image to be detected and the target object image to obtain a first feature image and a second feature image; then, the server generates a correlation matrix for representing the similarity degree of the pixels in the first feature image and the pixels in the second feature image according to the first feature image and the second feature image; again, the server generates a pixel-level spatial position matrix for representing the similarity degree and the position relationship of the pixels in the image to be detected and the pixels in the target object image through the correlation matrix; then, the server generates a target object detection frame containing the target object in the image to be detected according to the pixel-level spatial position matrix; finally, the server sends the image to be detected containing the target object detection frame to the terminal. The terminal displays the image to be detected containing the target object detection frame.

[0086] Next, the image data processing method in the present application will be introduced from the perspective of the server. Please refer to Figure 2 , the image data processing method provided by the embodiment of the present application includes: step S110 to step S150.

[0087] Specifically:

[0088] S110. Obtain the image to be detected and the target object image.

[0089] Among them, the image to be detected includes M first original pixels, the target object image includes N second original pixels, the target object image includes the target object, M is an integer greater than 1, and N is an integer greater than 1.

[0090] It should be noted that the purpose of the embodiments of this application is to detect whether a target object is included in the image to be detected and determine the position of the target object in the image to be detected. The target object image includes at least one target object. A pixel is the smallest unit in an image represented by a sequence of numbers, and each pixel has a definite position and an assigned color value.

[0091] S120: Respectively take the image to be detected and the target object image as the inputs of the feature extraction network in the one-shot detection model, and output a first feature image and a second feature image through the feature extraction network respectively.

[0092] Among them, the first feature image is generated by the feature extraction network according to the image to be detected. The first feature image includes K first feature pixels. The second feature image is generated by the feature extraction network according to the target object image. The second feature image includes L second feature pixels. K is an integer greater than 1, and L is an integer greater than 1.

[0093] It should be noted that the characteristics of the one-shot detection (OSD) model include: First, the category scalability of OSD is stronger and the cost is very small. For a new category to be detected, only one image needs to be provided, and the OSD technology can detect it. Second, the OSD technology can produce more accurate detection results because it fully extracts the image features of the category samples and uses them as priors to detect in the image to be detected, and finally can accurately mine and locate the object regions similar to the category image to be detected. The feature extraction network is a network layer used to perform feature extraction, normalization, etc. on the input image.

[0094] It can be understood that through the feature extraction network in the one-shot detection model, the image to be detected is subjected to feature extraction, normalization, etc. to obtain the first feature image. The first feature image includes K first feature pixels. If the size of the first feature image is h1×w1, then K = h1×w1, where h1 is the height value of the first feature image and w1 is the length value of the first feature image.

[0095] Through the feature extraction network in the one-shot detection model, the target object image is subjected to feature extraction, normalization, etc. to obtain the second feature image. The second feature image includes L second feature pixels. If the size of the second feature image is h2×w2, then L = h2×w2, where h2 is the height value of the second feature image and w2 is the length value of the second feature image.

[0096] S130: Generate a correlation matrix according to the first feature image and the second feature image.

[0097] Among them, the correlation matrix includes K×L similarity values, and the K×L similarity values represent the similarity degrees between K first feature pixels and L second feature pixels.

[0098] It should be noted that the similarity value is a numerical value between 0 and 1, and the higher the value, the greater the similarity degree.

[0099] It can be understood that the first feature matrix corresponding to the first feature image and the second feature matrix corresponding to the second feature image are subjected to matrix multiplication operation, and any one of the matrices is first subjected to matrix transpose operation before the matrix multiplication operation. If the size of the first feature image is h1×w1 and the size of the second feature image is h2×w2, the size of the correlation matrix is h1w1×h2w2.

[0100] S140: Use the correlation matrix as the input of the transformation network in the single-sample detection model, and output a pixel-level spatial position matrix through the transformation network.

[0101] Among them, the pixel-level spatial position matrix includes K×L×2 elements, and the K×L×2 elements represent the corresponding position coordinates of the L second feature pixels in the first feature image when any one of the K first feature pixels is used as an anchor point.

[0102] It should be noted that the transformation network (TransformNet) is used for the corresponding position coordinates of each first feature pixel in the first feature image and each second feature pixel in the second feature image, and its output result is a pixel-level spatial position matrix. The pixel-level spatial position matrix includes K×L×2 elements, and each element includes the predicted coordinate positions between two pixels. The anchor point refers to the target point. When any one of the K first feature pixels is used as an anchor point, the corresponding position coordinates of the L second feature pixels in the first feature image refer to the position coordinates corresponding to the pixel points similar to the L second feature pixels found in the first feature image. If the size of the first feature image is h1×w1, the size of the second feature image is h2×w2, and the size of the correlation matrix is h1w1×h2w2, then the size of the pixel-level spatial position matrix is h1×w1×h2×w2×2.

[0103] S150: Generate T target object detection frames in the image to be detected according to the pixel-level spatial position matrix.

[0104] Among them, the T target object detection frames include T target objects, and the T confidence values corresponding to the T target object detection frames all meet the confidence threshold, and T is an integer greater than or equal to 0.

[0105] It can be understood that T pixel groups similar to the target object pixels in the target object image are found in the image to be detected through the pixel-level spatial position matrix. Based on the T pixel groups representing T target objects, T target object detection frames are generated in the image to be detected.

[0106] For ease of understanding, please refer to Figure 3 and 4 , Figure 3 which is a schematic diagram of the correspondence relationship of the coordinate positions of some similar pixels between the image to be detected and the target object image. Figure 4 is a schematic diagram of generating a target object detection frame in the image to be detected. As Figure 3 shown, 10 is the image to be detected, 20 is the target object image. According to the target object image 20, the target object is "A". The similarity degree between each pixel in the image to be detected 10 and each pixel in the target object image 20 is used to establish the connection between the pixel coordinate positions. The connection between the pixel coordinate positions is an intuitive manifestation form of a pixel-level spatial position matrix; as Figure 4 shown, 2 target object detection frames are generated in the image to be detected according to the connection between the pixel coordinate positions.

[0107] The embodiment of the present application provides an image data processing method. The similarity degree between the pixels in the feature map of the image to be detected and the pixels in the feature map of the target object image is established through the correlation matrix, and the spatial position relationship of the similar pixels between the image to be detected and the target object image is established through the pixel-level spatial position matrix, which improves the accuracy of detecting the target object from the image to be detected. Moreover, by establishing the spatial position relationship of the similar pixels between the image to be detected and the target object image through the pixel-level spatial position matrix, the affine transformation of the target object can also be predicted.

[0108] In an optional embodiment of the image data processing method provided in the corresponding embodiment of the present application, please refer to Figure 2 , K first feature pixels correspond to K detection frames generated with the first feature pixels as anchor points. After step S140, it includes: step S141 to step S142. Specifically: Figure 5 S141. Resample the correlation matrix according to the pixel-level spatial position matrix to generate a resampled matrix.

[0109] wherein, the resampled matrix includes U dimensions, and U is an integer greater than or equal to 2.

[0110] It should be noted that the resampling process mentioned in the embodiment of the present application refers to image resampling.

[0111]

[0112] ​It can be understood that the resampling matrix can be obtained through matrix operations of the pixel-level spatial position matrix and the correlation matrix. If the size of the first feature image is h1×w1, the size of the second feature image is h2×w2, the size of the correlation matrix is h1w1×h2w2, and the size of the pixel-level spatial position matrix is h1×w1×h2×w2×2, then the resampling matrix can be calculated by the following formula:

[0113]

[0114] where k is the abscissa of each pixel in the second feature image, that is, the abscissa of each second feature pixel, l is the ordinate of each pixel in the second feature image, that is, the ordinate of each second feature pixel, p is the abscissa of each pixel in the first feature image, that is, the abscissa of each first feature pixel, q is the ordinate of each pixel in the first feature image, that is, the ordinate of each first feature pixel, c is the correlation matrix, and g is the pixel-level spatial position matrix.

[0115] S142. Perform average pooling on U dimensions in the resampling matrix to obtain a confidence matrix.

[0116] Among them, the confidence matrix includes K confidence values, and the K confidence values correspond to K detection frames.

[0117] It should be noted that performing average pooling on U dimensions in the resampling matrix can be to calculate the average value along these U dimensions, and the confidence value represents the confidence of the detection frame obtained with the position of each pixel in the first feature image as the anchor point.

[0118] It can be understood that then perform average pooling on the p and q dimensions in the resampling matrix that is, calculate the average value along these two dimensions to obtain a confidence matrix.

[0119] The embodiment of the present application provides an image data processing method. The resampling matrix is obtained through the pixel-level spatial position matrix and the correlation matrix, and then the confidence matrix is obtained through the resampling matrix, so as to calculate the confidence of the detection frame obtained with the position of each pixel in the first feature image as the anchor point, and the accuracy of the detection frame containing the target object is improved through the detection frame confidence.

[0120] In an optional embodiment of the image data processing method provided in the corresponding embodiment of the present application, please refer to Figure 2 or Figure 5 In an optional embodiment of the image data processing method provided in the corresponding embodiment of the present application, please refer to Figure 6 , step S150 further includes: step S1501 to step S1503.

[0121] Specifically:

[0122] S1501. Generate T corresponding grids in the image to be detected according to the pixel-level spatial position matrix, where the corresponding grids are the corresponding grids of the similar pixels of the image to be detected and the target object image.

[0123] It can be understood that, first, among the M first original pixels in the image to be detected through the pixel-level spatial position matrix, determine the similar pixels whose similarity degree to the pixels of the target object in the target object image is greater than the similarity threshold; then, determine the associated pixels in each similar pixel whose degree of association with it is greater than the correlation threshold; then, connect each similar pixel with its corresponding associated pixel pairwise to obtain the corresponding grids. Each corresponding grid is an approximate contour of the target object.

[0124] S1503. Generate T target object detection frames according to the T corresponding grids, where each target object detection frame is the circumscribed rectangle of each corresponding grid.

[0125] It can be understood that generate the rectangular target object detection frames according to the circumscribed rectangles of the corresponding grids.

[0126] For easy understanding, please refer to Figure 3 and Figure 7 , Figure 7 is a schematic diagram of the corresponding grid. As Figure 3 shown, in the image 10 to be detected, 2 groups of similar pixels whose similarity degree to the pixels of the target object in the target object image is greater than the similarity threshold are determined. As Figure 7 shown, determine the associated pixels of each similar pixel, connect each similar pixel with its associated pixel pairwise to obtain the corresponding grids, and finally determine the target object detection frames according to the corresponding grids.

[0127] The embodiment of the present application provides an image data processing method, which determines the target object detection frames in the way of corresponding grids, improving the accuracy of generating the target object detection frames.

[0128] In an optional embodiment of the image data processing method provided in the Figure 2 or Figure 5 corresponding embodiments of the present application, please refer to Figure 8 , step S150 further includes: step S1502 to step S1504.

[0129] Specifically:

[0130] S1502. Determine the vertex coordinates of the target object in the image to be detected according to the pixel-level spatial position matrix.

[0131] It can be understood that, first, similar pixels whose pixel similarity with the target object of the target object image is greater than a similarity threshold are determined among the M first original pixels in the image to be detected through the pixel-level spatial position matrix; then, related pixels whose correlation degree with the target object is greater than the correlation threshold are determined in each similar pixel; then, a target object pixel set is obtained based on each similar pixel and its related pixels, and the maximum and minimum values on the x-axis and y-axis are determined in the target object pixel set, thereby obtaining the vertex coordinates of the target object.

[0132] S1504: Generate a target object detection frame according to the vertex coordinates of the target object.

[0133] It can be understood that the target object detection frame is generated according to the maximum values of the coordinate values in the four directions.

[0134] For easier understanding, see Figure 3 and Figure 9 , Figure 9 is a schematic diagram of the vertex coordinate determination process. Figure 3 As shown, two groups of similar pixels whose pixel similarity with the target object of the target object image is greater than the similarity threshold are determined in the image to be detected 10, such as Figure 9 As shown, the coordinates of the pixels corresponding to the maximum and minimum values on the x-axis and y-axis in the target object pixel set are determined, thereby obtaining the vertex coordinates of the target object, and generating a target object detection frame according to the vertex coordinates.

[0135] The embodiment of the present application provides an image data processing method, which determines a target object detection frame by determining the vertex coordinates of the target object, thereby improving the accuracy of generating the target object detection frame.

[0136] In this application Figure 2 In an optional embodiment of the image data processing method provided in the corresponding embodiment, please refer to Figure 10 , the feature extraction network includes a convolutional subnetwork and a dictionary subnetwork, and the dictionary subnetwork carries a dictionary matrix. Step S120 further includes steps S1201 to S1208. Specifically:

[0137] S1201. Use the image to be detected as the input of the convolutional sub-network, and output a first intermediate matrix through the convolutional sub-network.

[0138] It can be understood that the image to be detected is processed by the convolution sub-network to obtain the first intermediate matrix.

[0139] S1203: Input the first intermediate matrix into the dictionary sub-network, perform feature crossover between the first intermediate matrix and the dictionary matrix through the dictionary sub-network, and generate a feature matrix of the image to be detected.

[0140] It should be noted that feature crossing, also known as feature combination, is a method of synthesizing features, which can perform non-linear feature fitting on a multi-dimensional feature dataset.

[0141] It can be understood that performing feature crossing on the first intermediate matrix and the dictionary matrix can be to perform matrix multiplication on the first intermediate matrix and the dictionary matrix to obtain the feature matrix of the image to be detected.

[0142] S1205. Normalize the feature matrix of the image to be detected to obtain the first feature matrix.

[0143] It should be noted that for the convenience of data processing, mapping the data to the range of 0 to 1 for processing is called normalization processing.

[0144] It can be understood that normalizing the feature matrix of the image to be detected does not change the dimension of the matrix, that is, the feature matrix of the image to be detected and the first feature matrix have the same dimension. The normalization processing can be achieved through the softmax function.

[0145] S1207. Generate the first feature image according to the first feature matrix.

[0146] Steps S1201 to S1207 are the process of taking the image to be detected as the input of the feature extraction network in the single-sample detection model and outputting the first feature image through the feature extraction network.

[0147] S1202. Take the target object image as the input of the convolutional sub-network and output the second intermediate matrix through the convolutional sub-network.

[0148] It can be understood that the second intermediate matrix is obtained by processing the target object image through the convolutional sub-network.

[0149] S1204. Input the second intermediate matrix into the dictionary sub-network, and perform feature crossing on the second intermediate matrix and the dictionary matrix through the dictionary sub-network to generate the feature matrix of the target object image.

[0150] It can be understood that performing feature crossing on the second intermediate matrix and the dictionary matrix can be to perform matrix multiplication on the second intermediate matrix and the dictionary matrix to obtain the feature matrix of the target object image.

[0151] S1206. Normalize the feature matrix of the target object image to obtain the second feature matrix.

[0152] It can be understood that normalizing the feature matrix of the target object image does not change the dimension of the matrix, that is, the feature matrix of the target object image and the second feature matrix have the same dimension. The normalization processing can be achieved through the softmax function.

[0153] S1208. Generate a second feature image according to the second feature matrix.

[0154] Steps S1202 to S1208 are the process of using the target object image as the input of the feature extraction network in the single-sample detection model and outputting the second feature image through the feature extraction network.

[0155] For ease of understanding, please refer to Figure 11 , Figure 11 which is the structural schematic diagram of the feature extraction network in the single-sample detection model provided by the embodiments of the present application. The process of processing the image to be detected includes: First, input the image to be detected into the convolutional sub-network and output the first intermediate matrix through the convolutional sub-network; then, input the first intermediate matrix into the dictionary sub-network to perform feature crossing between the first intermediate matrix and the dictionary matrix in the dictionary sub-network to obtain the feature matrix of the image to be detected; then, perform normalization processing on the feature matrix of the image to be detected to obtain the first feature image. The process of processing the target object image includes: First, input the target object image into the convolutional sub-network and output the second intermediate matrix through the convolutional sub-network; then, input the second intermediate matrix into the dictionary sub-network to perform feature crossing between the second intermediate matrix and the dictionary matrix in the dictionary sub-network to obtain the feature matrix of the target object image; then, perform normalization processing on the feature matrix of the target object image to obtain the second feature image. The output matrix of the feature extraction network can be expressed by the following formula:

[0156] f a = f s (FD)D T ;

[0157] where f a is the output matrix of the feature extraction network, f s is the normalization function, F is the output matrix of the convolutional sub-network, and D is the dictionary matrix. For example, f s can be the softmax function.

[0158] The embodiments of the present application provide an image data processing method. In order to enable the single-sample detection model to better learn the feature primitives contained in the input target object image, a dictionary sub-network is introduced into the single-sample detection model, which can exchange a lower computational increment for a stronger generalization ability of the target object, so as to more accurately detect new target objects that have not been seen during training.

[0159] In an optional embodiment of the image data processing method provided in the corresponding embodiment of the present application, please refer to Figure 2 for details. The image data processing method further includes steps S210 to S270. Specifically: Figure 12 ,

[0160] S210. Obtain a first training sample image, a second training sample image, and a training object image.

[0161] Among them, the first training sample image includes M X1 first training original pixels, the first training sample image includes T B training object annotation boxes, the T B training object annotation boxes include T B training objects, the T B training object annotation boxes correspond to T B annotation box data, the second training sample image includes M X2 second training original pixels, the second training sample image does not include training objects, the training object image includes N X third training original pixels, the training object image includes training objects, M X1 is an integer greater than 1, M X2 is an integer greater than 1, N X is an integer greater than 1, T B is an integer greater than or equal to 1.

[0162] It should be noted that the purpose of the embodiments of this application is to train a single-sample detection model by inputting a training object image, a first training sample image (positive sample) containing training objects and training object annotation boxes, and a second training sample image (negative sample) without training objects. The annotation box data includes the coordinate encoding of the annotation box. The training object image includes at least one training object, the first training sample image includes at least one training object, and the second training sample image does not include training objects.

[0163] S220. Respectively use the first training sample image, the second training sample image, and the training object image as the inputs of the feature extraction network in the single-sample detection model, and respectively output a first training feature image, a second training feature image, and a third training feature image through the feature extraction network.

[0164] Among them, the first training feature image is generated by the feature extraction network according to the first training sample image, and the first training feature image includes K X1 first training feature pixels, the second training feature image is generated by the feature extraction network according to the second training sample image, and the second training feature image includes K X2 second training feature pixels, the third training feature image is generated by the feature extraction network according to the training object image, and the third training feature image includes L X third training feature pixels, K X1 is an integer greater than 1, K X2is an integer greater than 1, L X is an integer greater than 1.

[0165] It can be understood that through the feature extraction network in the single-sample detection model, feature extraction, normalization, etc. are performed on the first training sample image to obtain the first training feature matrix, and the first training feature image is generated according to the first training feature matrix. The first training feature image includes K X1 first training feature pixels. If the size of the first training feature image is h X1 ×w X1 , then K X1 = h X1 ×w X1 , h X1 is the height value of the first training feature image, and w X1 is the length value of the first training feature image.

[0166] Through the feature extraction network in the single-sample detection model, feature extraction, normalization, etc. are performed on the second training sample image to obtain the second training feature matrix, and the second training feature image is generated according to the second training feature matrix. The second training feature image includes K X2 second training feature pixels. If the size of the second training feature image is h X2 ×w X2 , then K X2 = h X2 ×w X2 , h X2 is the height value of the second training feature image, and w X2 is the length value of the second training feature image.

[0167] Through the feature extraction network in the single-sample detection model, feature extraction, normalization, etc. are performed on the training object image to obtain the third training feature matrix, and the third training feature image is generated according to the third training feature matrix. The third training feature image includes L X third training feature pixels. If the size of the third training feature image is h X3 ×w X3 , then L X = h X3 ×w X3 , h X3 is the height value of the third training feature image, and w X3 is the length value of the third training feature image.

[0168] S231. Generate the first training correlation matrix according to the first training feature image and the third training feature image.

[0169] Among them, the first training correlation matrix includes K X1 ×L XThe similarity value, K X1 ×L X The similarity value is K X1 The degree of similarity between the first training feature pixels and L X third training feature pixels.

[0170] It should be noted that the similarity value is a value between 0 and 1, and the higher the value, the greater the degree of similarity.

[0171] It can be understood that the first training feature matrix corresponding to the first training feature image and the third training feature matrix corresponding to the third training feature image are subjected to matrix multiplication, and before the matrix multiplication, any one of the matrices is first subjected to matrix transpose. If the size of the first feature image is h X1 ×w X1 , and the size of the third feature image is h X3 ×w X3 , then the size of the correlation matrix is h X1 w X1 ×h X3 w X3 .

[0172] S232. Generate a second training correlation matrix according to the second training feature image and the third training feature image.

[0173] Among them, the second training correlation matrix includes K X2 ×L X similarity values, K X2 ×L X The similarity value is K X2 The degree of similarity between the second training feature pixels and L X third training feature pixels.

[0174] It should be noted that the similarity value is a value between 0 and 1, and the higher the value, the greater the degree of similarity.

[0175] It can be understood that the second training feature matrix corresponding to the second training feature image and the third training feature matrix corresponding to the third training feature image are subjected to matrix multiplication, and before the matrix multiplication, any one of the matrices is first subjected to matrix transpose. If the size of the second feature image is h X2 ×w X2 , and the size of the third feature image is h X3 ×w X3 , then the size of the correlation matrix is h X2 w X2 ×h X3 w X3 .

[0176] S241. Use the first training correlation matrix as the input of the transformation network in the single-sample detection model, and output the first training pixel-level spatial position matrix through the transformation network.

[0177] Among them, the first training pixel-level spatial position matrix includes K X1 ×L X ×2 first training elements, and M X1 ×N X ×2 first training elements represent the corresponding position coordinates of L X1 third training feature pixels in the first training feature image when any one of the K X first training feature pixels is used as the anchor point.

[0178] It should be noted that the transformation network (TransformNet) is used to calculate the corresponding position coordinates of each first training feature pixel in the first training feature image and each third training feature pixel in the third training feature image, and its output result is the first training pixel-level spatial position matrix. The first training pixel-level spatial position matrix includes K X1 ×L X ×2 first training elements, and each first training element includes the predicted coordinate positions between two pixels. The anchor point refers to the target point. When any one of the K X1 first training feature pixels is used as the anchor point, the corresponding position coordinates of L X third training feature pixels in the first training feature image refer to finding the corresponding position coordinates of the pixel points similar to the L X third training feature pixels in the first training feature image. If the size of the first feature image is h X1 ×w X1 , the size of the third feature image is h X3 ×w X3 , the size of the correlation matrix is h X1 w X1 ×h X3 w X3 , then the size of the first training pixel-level spatial position matrix is h X1 ×w X1 ×h X3 ×w X3 ×2.

[0179] S242. Use the second training correlation matrix as the input of the transformation network in the single-sample detection model, and output the second training pixel-level spatial position matrix through the transformation network.

[0180] Among them, the second training pixel-level spatial position matrix includes K X2 ×L X ×2 second training elements, and K X2 ×LX ×2 second training elements represent the coordinates of the corresponding positions of L X2 third training feature pixels in the second training feature image when any one of the K X second training feature pixels is used as an anchor point.

[0181] It should be noted that the TransformNet is used to calculate the coordinates of the corresponding positions of each second training feature pixel in the second training feature image and each third training feature pixel in the third training feature image, and its output result is the second training pixel-level spatial position matrix. The second training pixel-level spatial position matrix includes K X2 ×L X ×2 second training elements, and each second training element includes the predicted coordinate positions between two pixels. An anchor point refers to a target point. When any one of the K X2 second training feature pixels is used as an anchor point, the coordinates of the corresponding positions of L X third training feature pixels in the second training feature image refer to finding the coordinate positions of the pixel points similar to the L X third training feature pixels in the second training feature image. If the size of the second feature image is h X2 ×w X2 , the size of the third feature image is h X3 ×w X3 , the size of the correlation matrix is h X2 w X2 ×h X3 w X3 , then the size of the second training pixel-level spatial position matrix is h X2 ×w X2 ×h X3 ×w X3 ×2.

[0182] S251. Generate T X1 first detection frames of training objects in the first training sample image according to the first training pixel-level spatial position matrix.

[0183] Among them, the T X1 first detection frames of training objects include T X1 training objects, the T X1 first detection frames of training objects correspond to T X1 first training confidence levels, and the T X1 first detection frames of training objects correspond to T X1 first detection frame data, where T X1 is an integer greater than or equal to 1.

[0184] It is understandable that T pixel groups similar to the training object pixels in the training object image are found in the first training sample image through the first training pixel-level spatial position matrix. The T pixel groups represent T training objects, and then T first detection frames for the training objects are generated in the first training sample image, and each first detection frame for the training object corresponds to a first detection frame data. X1 T X1 The T pixel groups represent X1 T training objects, and then T first detection frames for the training objects are generated in the first training sample image, and each first detection frame for the training object corresponds to a first detection frame data. X1 It is understandable that T pixel groups similar to the training object pixels in the training object image are found in the first training sample image through the first training pixel-level spatial position matrix. The T pixel groups represent T training objects, and then T first detection frames for the training objects are generated in the first training sample image, and each first detection frame for the training object corresponds to a first detection frame data.

[0185] S252. Generate T second detection frames for the training objects in the second training sample image according to the second training pixel-level spatial position matrix. X2 T second detection frames for the training objects.

[0186] Among them, the T second detection frames for the training objects correspond to T second training confidence levels, and the T second detection frames for the training objects correspond to T second detection frame data, where T is an integer greater than or equal to 0. X2 The T second detection frames for the training objects correspond to X2 T second training confidence levels, and the T second detection frames for the training objects correspond to X2 T second detection frame data, where T is an integer greater than or equal to 0. X2 Among them, T is an integer greater than or equal to 0. X2 It is understandable that T pixel groups similar to the training object pixels in the training object image are found in the second training sample image through the second training pixel-level spatial position matrix. The T pixel groups represent T training objects, and then T second detection frames for the training objects are generated in the second training sample image, and each second detection frame for the training object corresponds to a second detection frame data.

[0187] It is understandable that T pixel groups similar to the training object pixels in the training object image are found in the second training sample image through the second training pixel-level spatial position matrix. The T pixel groups represent T training objects, and then T second detection frames for the training objects are generated in the second training sample image, and each second detection frame for the training object corresponds to a second detection frame data. X1 T X1 The T pixel groups represent X1 T training objects, and then T second detection frames for the training objects are generated in the second training sample image, and each second detection frame for the training object corresponds to a second detection frame data. X1 It is understandable that T pixel groups similar to the training object pixels in the training object image are found in the second training sample image through the second training pixel-level spatial position matrix. The T pixel groups represent T training objects, and then T second detection frames for the training objects are generated in the second training sample image, and each second detection frame for the training object corresponds to a second detection frame data.

[0188] S260. Generate a single-sample detection model loss result according to T first training confidence levels, T first detection frame data, T second training confidence levels, T second detection frame data, and T annotation frame data. X1 T first training confidence levels, X1 T first detection frame data, X2 T second training confidence levels, X2 T second detection frame data, and B T annotation frame data.

[0189] It is understandable that the loss result is obtained by fitting the trained data with the annotation data.

[0190] S270. Train the single-sample detection model according to the single-sample detection model loss result.

[0191] It is understandable that the single-sample detection model is trained with positive and negative sample images, and the model parameters are adjusted during the training process to make the loss result meet the preset loss result, thus completing the training of the model.

[0192] It should be noted that the above steps S210 to S270 are one training process. One training process requires one training object image, one positive sample image, and two negative sample images. In the actual training process, multiple trainings are required, and the total number of positive sample images and the total number of negative sample images satisfy a 1:2 relationship.

[0193] The embodiment of the present application provides an image data processing method. By adjusting the model parameters during model training, the output result of the target object detection frame in the test is made more accurate.

[0194] In the Figure 12 In an optional embodiment of the image data processing method provided in the corresponding embodiment of the present application, please refer to Figure 13 , step S260 further includes steps S2601 to S2604. Specifically:

[0195] S2601. Generate a first classification loss result according to T X1 first training confidence levels and a first confidence reference value.

[0196] It can be understood that the first classification loss result is the classification loss result of the positive sample. The first classification loss result can be expressed by the following formula:

[0197]

[0198] Among them, is the first classification loss result, m pos is the first confidence reference value, and m pos can be set to 0.7, s1 is the first training confidence level, and max is the maximum value function.

[0199] S2602. Generate a second classification loss result according to T X2 second training confidence levels and a second confidence reference value.

[0200] It can be understood that the second classification loss result is the classification loss result of the negative sample. The second classification loss result can be expressed by the following formula:

[0201]

[0202] Among them, is the second classification loss result, m neg is the second confidence reference value, and m neg can be set to 0.3, s2 is the first training confidence level, and max is the maximum value function.

[0203] S2603. According to T X1 first detection box data, TX2 a second detection box data and T B annotation box data, and generate a positioning loss result.

[0204] It can be understood that the positioning loss result is the comprehensive positioning loss result of positive samples and negative samples. The positioning loss result can be expressed by the following formula:

[0205]

[0206] where l loc (x, y) is the positioning loss result, x i represents the first detection box data or the second detection box data, and y i represents the annotation box data.

[0207] S2604. Generate a single-sample detection model loss result according to the first classification loss result, the second classification loss result, and the positioning loss result.

[0208] It can be understood that by adding the first classification loss result, the second classification loss result, and the positioning loss result, the single-sample detection model loss result is obtained. The single-sample detection model loss result can be expressed by the following formula:

[0209]

[0210] where l is the single-sample detection model loss result, is the first classification loss result, is the second classification loss result, and l loc (x, y) is the positioning loss result.

[0211] The embodiment of the present application provides an image data processing method, which adjusts the model parameters during model training to make the output result of the target object detection box in the test more accurate.

[0212] For the sake of easy understanding, the following will be combined with Figure 14 introduce an image data processing method applied to kitten detection, the purpose of which is to detect the same kitten in the image to be detected 11 as in the target object image 21, Figure 12 The schematic diagram of the image data processing method for kitten detection includes:

[0213] Step 1: Image acquisition.

[0214] Specifically: Acquire the image to be detected and the target object image.

[0215] Among them, the size of the image to be detected is 500×600, and the image to be detected includes 300,000 first original pixels. The size of the target object image is 150×200, and the target object image includes 30,000 second original pixels.

[0216] Step 2: Image feature extraction.

[0217] Specifically: Take the image to be detected as the input of the feature extraction network in the single-sample detection model, and output the first feature image through the feature extraction network. Take the target object image as the input of the feature extraction network in the single-sample detection model, and output the second feature image through the feature extraction network respectively.

[0218] Among them, the first feature image includes 1,000 first feature pixels. The second feature image includes 100 second feature pixels.

[0219] Step 3: Calculate the correlation matrix.

[0220] Specifically: Generate a correlation matrix according to the first feature matrix corresponding to the first feature image and the second feature image corresponding to the second feature image.

[0221] Among them, the correlation matrix includes 1,000×100 similarity values, and the 1,000×100 similarity values represent the similarity degree between 1,000 first feature pixels and 100 second feature pixels.

[0222] Step 4: Calculate the pixel-level spatial position matrix.

[0223] Specifically: Take the correlation matrix as the input of the transformation network in the single-sample detection model, and output the pixel-level spatial position matrix through the transformation network.

[0224] Among them, the pixel-level spatial position matrix includes 1,000×100×2 elements, and the 1,000×100×2 elements represent the corresponding position coordinates of 100 second feature pixels in the first feature image when using 1,000 first original pixels as the anchor points.

[0225] Step 5: Calculate the confidence matrix.

[0226] Specifically: Perform resampling processing on the correlation matrix according to the pixel-level spatial position matrix to generate a resampled matrix; perform average pooling processing on the resampled matrix to obtain the confidence matrix.

[0227] Step 6: Generate the target detection box.

[0228] Specifically: Generate 1 target object detection box in the image to be detected according to the pixel-level spatial position matrix.

[0229] The embodiments of the present application provide an image data processing method. By using a correlation matrix, the similarity degree between the pixels in the feature map of the image to be detected and the pixels in the feature map of the target object image is established. By using a pixel-level spatial position matrix, the spatial position relationship of the similar pixels between the image to be detected and the target object image is established, which improves the accuracy of detecting the target object from the image to be detected. Moreover, by using the pixel-level spatial position matrix to establish the spatial position relationship of the similar pixels between the image to be detected and the target object image, the affine transformation of the target object can also be predicted.

[0230] The following will describe the image data processing device in the present application in detail. Please refer to Figure 15 。 Figure 15 FIG. 7 is a schematic diagram of an embodiment of the image data processing device 100 in the embodiments of the present application. The image data processing device 100 includes:

[0231] An image acquisition module 110, configured to acquire an image to be detected and a target object image. Among them, the image to be detected includes M first original pixels, the target object image includes N second original pixels, the target object image includes a target object, M is an integer greater than 1, and N is an integer greater than 1.

[0232] A feature extraction module 120, configured to respectively input the image to be detected and the target object image into the feature extraction network in a single-sample detection model, and respectively output a first feature image and a second feature image through the feature extraction network. Among them, the first feature image is generated by the feature extraction network according to the image to be detected, the first feature image includes K first feature pixels, the second feature image is generated by the feature extraction network according to the target object image, the second feature image includes L second feature pixels, K is an integer greater than 1, and L is an integer greater than 1.

[0233] A correlation matrix generation module 130, configured to generate a correlation matrix according to the first feature image and the second feature image. Among them, the correlation matrix includes K×L similarity values, and the K×L similarity values represent the similarity degree between K first feature pixels and L second feature pixels.

[0234] A pixel-level spatial position matrix generation module 140, configured to input the correlation matrix into the transformation network in the single-sample detection model, and output a pixel-level spatial position matrix through the transformation network. Among them, the pixel-level spatial position matrix includes M×N×2 elements, and the M×N×2 elements represent the corresponding position coordinates of L second feature pixels in the first feature image when any one of the K first feature pixels is used as an anchor point.

[0235] The detection box generation module 150 is configured to generate T target object detection boxes in the image to be detected according to the pixel-level spatial position matrix, where the T target object detection boxes include T target objects, and the T confidence values corresponding to the T target object detection boxes all satisfy the confidence threshold, and T is an integer greater than or equal to 0.

[0236] The embodiment of the present application provides an image data processing device. By using the correlation matrix, the similarity degree between the pixels in the feature map of the image to be detected and the pixels in the feature map of the target object image is established. By using the pixel-level spatial position matrix, the spatial position relationship of the similar pixels between the image to be detected and the target object image is established, which improves the accuracy of detecting the target object from the image to be detected. Moreover, by using the pixel-level spatial position matrix to establish the spatial position relationship of the similar pixels between the image to be detected and the target object image, the affine transformation of the target object can also be predicted.

[0237] In the Figure 15 corresponding optional embodiment of the image data processing device provided in the embodiment of the present application, please refer to Figure 16 , the K first feature pixels correspond to K detection boxes generated with the first feature pixels as the anchor points. The image data processing device 100 further includes a confidence matrix calculation module 141, configured to:

[0238] Resample the correlation matrix according to the pixel-level spatial position matrix to generate a resampled matrix, where the resampled matrix includes U dimensions, and U is an integer greater than or equal to 2;

[0239] Perform average pooling on the U dimensions in the resampled matrix to obtain a confidence matrix, where the confidence matrix includes K confidence values, and the K confidence values correspond to the K detection boxes.

[0240] The embodiment of the present application provides an image data processing device. By using the pixel-level spatial position matrix and the correlation matrix, a resampled matrix is obtained, and then a confidence matrix is obtained through the resampled matrix, so as to calculate the confidence of the detection boxes obtained with the position of each pixel in the first feature image as the anchor point. The accuracy of the detection boxes containing the target object is improved by the confidence of the detection boxes.

[0241] In the Figure 15 or Figure 16 corresponding optional embodiment of the image data processing device provided in the embodiment of the present application, the detection box generation module 150 includes: a detection box first generation sub-module, configured to:

[0242] Generate T corresponding grids in the image to be detected according to the pixel-level spatial position matrix, and the corresponding grids are the corresponding grids of the similar pixels between the image to be detected and the target object image;

[0243] Generate T target object detection frames according to T corresponding grids, where each target object detection frame is an external rectangle of each corresponding grid.

[0244] The embodiment of the present application provides an image data processing device, which determines the target object detection frame in the way of corresponding grids, improving the accuracy of generating the target object detection frame.

[0245] In the Figure 15 or Figure 16 In an optional embodiment of the image data processing device provided in the corresponding embodiment of the present application, the detection frame generation module 150 includes: The detection frame second generation sub-module is used for:

[0246] Determine the vertex coordinates of the target object in the image to be detected according to the pixel-level spatial position matrix;

[0247] Generate a target object detection frame according to the vertex coordinates of the target object.

[0248] The embodiment of the present application provides an image data processing device, which determines the target object detection frame by determining the vertex coordinates of the target object, improving the accuracy of generating the target object detection frame.

[0249] In the Figure 15 In an optional embodiment of the image data processing device provided in the corresponding embodiment of the present application, the feature extraction network includes a convolutional sub-network and a dictionary sub-network, and the dictionary sub-network carries a dictionary matrix; the feature extraction module 120 is further used for:

[0250] Take the image to be detected as the input of the convolutional sub-network, and output a first intermediate matrix through the convolutional sub-network;

[0251] Input the first intermediate matrix into the dictionary sub-network, and perform feature crossing between the first intermediate matrix and the dictionary matrix through the dictionary sub-network to generate a feature matrix of the image to be detected;

[0252] Normalize the feature matrix of the image to be detected to obtain a first feature matrix;

[0253] Generate a first feature image according to the first feature matrix.

[0254] Take the target object image as the input of the convolutional sub-network, and output a second intermediate matrix through the convolutional sub-network;

[0255] Input the second intermediate matrix into the dictionary sub-network, and perform feature crossing between the second intermediate matrix and the dictionary matrix through the dictionary sub-network to generate a feature matrix of the target object image;

[0256] Normalize the feature matrix of the target object image to obtain a second feature matrix;

[0257] Generate a second feature image according to the second feature matrix.

[0258] The embodiments of the present application provide an image data processing device. In order to enable a one-shot detection model to better learn the feature primitives contained in the input target object image, a dictionary sub-network is introduced into the one-shot detection model, which can exchange a lower computational increment for a stronger generalization ability of the target object, so as to more accurately detect new target objects that have not been seen during training.

[0259] In the Figure 15 corresponding optional embodiment of the image data processing device provided by the embodiment of the present application, please refer to Figure 17 , the image data processing device 100 further includes a model training module 200, and the model training module 200 includes:

[0260] A training image acquisition sub-module, configured to acquire a first training sample image, a second training sample image, and a training object image, where the first training sample image includes M X1 first training original pixels, the first training sample image includes T B training object annotation boxes, T B training object annotation boxes include T B training objects, T B training object annotation boxes correspond to T B annotation box data, the second training sample image includes M X2 second training original pixels, the second training sample image does not include training objects, the training object image includes N X third training original pixels, the training object image includes training objects, M X1 is an integer greater than 1, M X2 is an integer greater than 1, N X is an integer greater than 1;

[0261] A training feature extraction sub-module, configured to respectively use the first training sample image, the second training sample image, and the training object image as inputs to the feature extraction network in the one-shot detection model, and respectively output a first training feature image, a second training feature image, and a third training feature image through the feature extraction network, where the first training feature image is generated by the feature extraction network according to the first training sample image, and the first training feature image includes K X1 first training feature pixels, the second training feature image is generated by the feature extraction network according to the second training sample image, and the second training feature image includes K X2 second training feature pixels, the third training feature image is generated by the feature extraction network according to the training object image, and the third training feature image includes L XA third training feature pixel, K X1 is an integer greater than 1, K X2 is an integer greater than 1, L X is an integer greater than 1;

[0262] The training correlation matrix generation sub-module is used to generate a first training correlation matrix according to the first training feature image and the third training feature image, wherein the first training correlation matrix includes K X1 ×L X similarity values, K X1 ×L X The similarity values of ×L are K X1 the similarity degree between K first training feature pixels and L X third training feature pixels;

[0263] The training correlation matrix generation sub-module is also used to generate a second training correlation matrix according to the second training feature image and the third training feature image, wherein the second training correlation matrix includes K X2 ×L X similarity values, K X2 ×L X The similarity values of ×L are K X2 the similarity degree between K second training feature pixels and L X third training feature pixels;

[0264] The training pixel-level spatial position matrix generation sub-module is used to take the first training correlation matrix as the input of the transformation network in the single-sample detection model, and output the first training pixel-level spatial position matrix through the transformation network, wherein the first training pixel-level spatial position matrix includes K X1 ×L X ×2 first training elements, K X1 ×L X ×2 first training elements indicate that when any one of the K first training feature pixels is used as an anchor point, L X1 the corresponding position coordinates of the L third training feature pixels in the first training feature image; X

[0265] The training pixel-level spatial position matrix generation sub-module is also used to take the second training correlation matrix as the input of the transformation network in the single-sample detection model, and output the second training pixel-level spatial position matrix through the transformation network, wherein the second training pixel-level spatial position matrix includes K X2 ×L X ×2 second training elements, K X2 ×L X ×2 second training elements indicate that when any one of the K second training feature pixels is used as an anchor point, L X2 the L third training feature pixels; X ​The corresponding position coordinates of the third training feature pixels in the second training feature image;

[0266] A training detection box generation sub-module, configured to generate T X1 first detection boxes of training objects in the first training sample image according to the first training pixel-level spatial position matrix, where T X1 first detection boxes of training objects include T X1 training objects, and T X1 first detection boxes of training objects correspond to T X1 first training confidence levels, and T X1 first detection boxes of training objects correspond to T X1 first detection box data;

[0267] The training detection box generation sub-module is further configured to generate T X2 second detection boxes of training objects in the second training sample image according to the second training pixel-level spatial position matrix, where T X2 second detection boxes of training objects correspond to T X2 second training confidence levels, and T X2 second detection boxes of training objects correspond to T X2 second detection box data;

[0268] A loss result generation sub-module, configured to generate a single-sample detection model loss result according to T X1 first training confidence levels, T X1 first detection box data, T X2 second training confidence levels, T X2 second detection box data, and T B annotation box data;

[0269] A model training sub-module, configured to train the single-sample detection model according to the single-sample detection model loss result.

[0270] The embodiment of the present application provides an image data processing device, which adjusts the model parameters during model training, so that the output result of the target object detection box in the test is more accurate.

[0271] In an optional embodiment of the image data processing device provided in the corresponding embodiment of the present application, the first training sample image corresponds to a first confidence reference value, and the second training sample image corresponds to a second confidence reference value; the loss result generation sub-module is further configured to: Figure 17 According to T

[0272] According to T X1 first training confidence levels, T X1 first detection box data, T X2 second training confidence levels, TX2 a second detection box data and T B annotation box data, and generate a single-sample detection model loss result, including:

[0273] According to T X1 first training confidence levels and a first confidence reference value, generate a first classification loss result;

[0274] According to T X2 second training confidence levels and a second confidence reference value, generate a second classification loss result;

[0275] According to T X1 first detection box data, T X2 second detection box data and T B annotation box data, generate a localization loss result;

[0276] According to the first classification loss result, the second classification loss result, and the localization loss result, generate a single-sample detection model loss result.

[0277] The embodiment of the present application provides an image data processing device, which adjusts model parameters during model training, so that the output result of the target object detection box in the test is more accurate.

[0278] Figure 18 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server 300 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the server 300.

[0279] The server 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or, one or more operating systems 341, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM, FreeBSD TM and so on

[0280] In the above embodiments, the steps executed by the server can be based on the Figure 18 server structure shown

[0281] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0282] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0283] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0284] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0285] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0286] In the above, the above embodiments are only used to illustrate the technical solution of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. An image data processing method, characterized in that, Including: Obtain the image to be detected and the target object image; Respectively use the image to be detected and the target object image as the inputs of the feature extraction network in the single-sample detection model, and respectively output a first feature image and a second feature image through the feature extraction network. Among them, the first feature image is generated by the feature extraction network according to the image to be detected, the first feature image includes K first feature pixels, the second feature image is generated by the feature extraction network according to the target object image, the second feature image includes L second feature pixels, K is an integer greater than 1, and L is an integer greater than 1; Generate a correlation matrix according to the first feature image and the second feature image. Among them, the correlation matrix includes K×L similarity values, and the K×L similarity values represent the similarity degree between K first feature pixels and L second feature pixels; Use the correlation matrix as the input of the transformation network in the single-sample detection model, and output a pixel-level spatial position matrix through the transformation network. Among them, the pixel-level spatial position matrix includes K×L×2 elements, and the K×L×2 elements represent the corresponding position coordinates of L second feature pixels in the first feature image when any one of the K first feature pixels is used as an anchor point; Generate T target object detection frames in the image to be detected according to the pixel-level spatial position matrix. Among them, the T target object detection frames include T target objects, and the T confidence values corresponding to the T target object detection frames all meet the confidence threshold, and T is an integer greater than or equal to 1; Among them, the feature extraction network includes a convolutional sub-network and a dictionary sub-network, and the dictionary sub-network carries a dictionary matrix; The step of respectively using the image to be detected and the target object image as the inputs of the feature extraction network in the single-sample detection model, and respectively outputting a first feature image and a second feature image through the feature extraction network includes: Use the image to be detected as the input of the convolutional sub-network, and output a first intermediate matrix through the convolutional sub-network; Input the first intermediate matrix into the dictionary sub-network, and perform feature crossing between the first intermediate matrix and the dictionary matrix through the dictionary sub-network to generate a feature matrix of the image to be detected; Perform normalization processing on the feature matrix of the image to be detected to obtain the first feature matrix; Generate the first feature image according to the first feature matrix; Use the target object image as the input of the convolutional sub-network, and output a second intermediate matrix through the convolutional sub-network; Input the second intermediate matrix into the dictionary sub-network, and perform feature crossing between the second intermediate matrix and the dictionary matrix through the dictionary sub-network to generate a feature matrix of the target object image; Perform normalization processing on the feature matrix of the target object image to obtain the second feature matrix; Generate the second feature image according to the second feature matrix.

2. The image data processing method according to claim 1, wherein The K first feature pixels correspond to K detection frames generated with the first feature pixel as an anchor point; After outputting the pixel-level spatial position matrix through the transformation network, the following steps are further included: Resample the correlation matrix according to the pixel-level spatial position matrix to generate a resampled matrix, where the resampled matrix includes u dimensions, and U is an integer greater than or equal to 2; Perform average pooling on the U dimensions in the resampled matrix to obtain a confidence matrix, where the confidence matrix includes K confidence values, and the K confidence values correspond to K detection frames.

3. The image data processing method according to any one of claims 1-2, characterized in that The step of generating T target object detection frames in the image to be detected according to the pixel-level spatial position matrix includes: Generate T corresponding grids in the image to be detected according to the pixel-level spatial position matrix, where the corresponding grids are the corresponding grids of the similar pixels of the image to be detected and the target object image; Generate T target object detection frames according to the T corresponding grids, where each target object detection frame is the circumscribed rectangle of each corresponding grid.

4. The image data processing method according to any one of claims 1-2, characterized in that The step of generating T target object detection frames in the image to be detected according to the pixel-level spatial position matrix includes: Determine the vertex coordinates of the target object in the image to be detected according to the pixel-level spatial position matrix; Generate the target object detection frame according to the vertex coordinates of the target object.

5. The image data processing method according to claim 1, wherein The method further includes: Obtain a first training sample image, a second training sample image, and a training object image, where the first training sample image includes T B training object annotation boxes, and the T B training object annotation boxes include T B training objects, and the T B training object annotation boxes correspond to T B annotation box data. The second training sample image does not include the training objects, and the training object image includes the training objects. T B is an integer greater than or equal to 1; Respectively take the first training sample image, the second training sample image, and the training object image as the inputs of the feature extraction network in the single-sample detection model, and respectively output a first training feature image, a second training feature image, and a third training feature image through the feature extraction network. Among them, the first training feature image is generated by the feature extraction network according to the first training sample image, and the first training feature image includes K X1 first training feature pixels. The second training feature image is generated by the feature extraction network according to the second training sample image, and the second training feature image includes K X2 second training feature pixels. The third training feature image is generated by the feature extraction network according to the training object image, and the third training feature image includes L X third training feature pixels. K X1 is an integer greater than 1, K X2 is an integer greater than 1, and L X is an integer greater than 1; Generate a first training correlation matrix according to the first training feature image and the third training feature image, where the first training correlation matrix includes K X1 ×L X similarity values, and the K X1 ×L X similarity values represent the similarity degrees between K X1 first training feature pixels and L X third training feature pixels; Generate a second training correlation matrix according to the second training feature image and the third training feature image, where the second training correlation matrix includes K X2 ×L X similarity values, and the K X2 ×L X similarity values represent the similarity between K X2 second training feature pixels and L X third training feature pixels; Use the first training correlation matrix as the input of the transformation network in the single-sample detection model, and output the first training pixel-level spatial position matrix through the transformation network, where the first training pixel-level spatial position matrix includes K X1 ×L X ×2 first training elements, and K X1 ×L X ×2 first training elements represent the corresponding position coordinates of L X1 third training feature pixels in the first training feature image when any one of the K X first training feature pixels is used as the anchor point; Use the second training correlation matrix as the input of the transformation network in the single-sample detection model, and output a second training pixel-level spatial position matrix through the transformation network, where the second training pixel-level spatial position matrix includes K X2 ×L X ×2 second training elements, and K X2 ×L X ×2 second training elements represent the corresponding position coordinates of L X2 third training feature pixels in the second training feature image when any one of the K X second training feature pixels is used as an anchor point; Generate T training object first detection boxes in the first training sample image according to the first training pixel-level spatial position matrix, where X1 the T X1 training object first detection boxes include T X1 training objects, the T X1 training object first detection boxes correspond to T X1 first training confidence levels, the T X1 training object first detection boxes correspond to T X1 first detection box data, where T X1 is an integer greater than or equal to 1; Generate T training object second detection boxes in the second training sample image according to the second training pixel-level spatial position matrix, where X2 the T X2 training object second detection boxes correspond to T X2 second training confidences, and the T X2 training object second detection boxes correspond to T X2 second detection box data, where X2 is an integer greater than or equal to 1; According to T X1 the first training confidence levels, T X1 the first detection box data, T X2 the second training confidence levels, T X2 the second detection box data, and T B the labeled box data, generate the loss result of the single-sample detection model; Train the single-sample detection model according to the loss result of the single-sample detection model.

6. The image data processing method according to claim 5, wherein The first training sample image corresponds to a first confidence reference value, and the second training sample image corresponds to a second confidence reference value; The said according to T X1 of the said first training confidence, T X1 of the said first detection box data, T X2 of the second training confidence, T X2 of the second detection box data, and T B of the annotation box data, generate the loss result of the single-sample detection model, including: According to T X1 generate a first classification loss result based on the first training confidence level and the first confidence level reference value; According to T X2 generate a second classification loss result based on the second training confidence and the second confidence reference value; According to T X1 of the first detection box data, T X2 of the second detection box data, and T B of the annotation box data, generate a positioning loss result; Generate the loss result of the single-sample detection model according to the first classification loss result, the second classification loss result, and the localization loss result.

7. An image data processing device, characterized in that, It includes: An image acquisition module for acquiring an image to be detected and a target object image; A feature extraction module for respectively taking the image to be detected and the target object image as inputs of a feature extraction network in a single-sample detection model, and respectively outputting a first feature image and a second feature image through the feature extraction network, where the first feature image is generated by the feature extraction network according to the image to be detected, the first feature image includes K first feature pixels, the second feature image is generated by the feature extraction network according to the target object image, the second feature image includes L second feature pixels, K is an integer greater than 1, and L is an integer greater than 1; A correlation matrix generation module for generating a correlation matrix according to the first feature image and the second feature image, where the correlation matrix includes K×L similarity values, and the K×L similarity values represent the similarity degree between K first feature pixels and L second feature pixels; A pixel-level spatial position matrix generation module, configured to use the correlation matrix as the input of a transformation network in the single-sample detection model, and output a pixel-level spatial position matrix through the transformation network, where the pixel-level spatial position matrix includes K×L×2 elements, and the K×L×2 elements represent the corresponding position coordinates of L second feature pixels in the first feature image when any one of the K first feature pixels is used as an anchor point; A detection box generation module, configured to generate T target object detection boxes in the image to be detected according to the pixel-level spatial position matrix, where the T target object detection boxes include T target objects, and the T confidence values corresponding to the T target object detection boxes all satisfy the confidence threshold, and T is an integer greater than or equal to 0; Wherein, the feature extraction network includes a convolutional sub-network and a dictionary sub-network, and the dictionary sub-network carries a dictionary matrix; The feature extraction module is further configured to: Use the image to be detected as the input of the convolutional sub-network, and output a first intermediate matrix through the convolutional sub-network; Input the first intermediate matrix into the dictionary sub-network, and perform feature crossing on the first intermediate matrix and the dictionary matrix through the dictionary sub-network to generate a feature matrix of the image to be detected; Perform normalization processing on the feature matrix of the image to be detected to obtain the first feature matrix; Generate the first feature image according to the first feature matrix; Use the target object image as the input of the convolutional sub-network, and output a second intermediate matrix through the convolutional sub-network; Input the second intermediate matrix into the dictionary sub-network, and perform feature crossing on the second intermediate matrix and the dictionary matrix through the dictionary sub-network to generate a feature matrix of the target object image; Perform normalization processing on the feature matrix of the target object image to obtain the second feature matrix; Generate the second feature image according to the second feature matrix.

8. The device according to claim 7, characterized in that, The K first feature pixels correspond to K detection boxes generated with the first feature pixels as anchors; The apparatus further includes a confidence matrix calculation module, and the confidence matrix calculation module is specifically configured to: Perform resampling processing on the correlation matrix according to the pixel-level spatial position matrix to generate a resampling matrix, where the resampling matrix includes U dimensions, and U is an integer greater than or equal to 2; Perform average pooling processing on the U dimensions in the resampling matrix to obtain a confidence matrix, where the confidence matrix includes K confidence values, and the K confidence values correspond to the K detection boxes.

9. The device according to any one of claims 7-8, characterized in that, The detection box generation module includes a first detection box generation sub-module, and the first detection box generation sub-module is used for: Generate T corresponding grids in the image to be detected according to the pixel-level spatial position matrix, and the corresponding grids are the corresponding grids of similar pixels of the image to be detected and the target object image; Generate T target object detection boxes according to the T corresponding grids, where each target object detection box is the circumscribed rectangle of each corresponding grid.

10. The device according to any one of claims 7-8, characterized in that, The detection box generation module includes a second generation sub-module, and the second generation sub-module is configured to: Determine the vertex coordinates of the target object in the image to be detected according to the pixel-level spatial position matrix; Generate a detection box for the target object according to the vertex coordinates of the target object.

11. The device according to claim 7, characterized in that, The apparatus further includes a model training module, and the model training module includes: A training image acquisition sub-module, configured to acquire a first training sample image, a second training sample image, and a training object image, wherein the first training sample image includes T B training object annotation boxes, and the T B training object annotation boxes include T B training objects. The T B training object annotation boxes correspond to T B annotation box data. The second training sample image does not include the training object, and the training object image includes the training object, where T B is an integer greater than or equal to 1; The training feature extraction sub-module is configured to use the first training sample image, the second training sample image, and the training object image as inputs to the feature extraction network in the single-sample detection model respectively, and output a first training feature image, a second training feature image, and a third training feature image through the feature extraction network respectively. Among them, the first training feature image is generated by the feature extraction network according to the first training sample image, and the first training feature image includes K X1 first training feature pixels, the second training feature image is generated by the feature extraction network according to the second training sample image, and the second training feature image includes K X2 second training feature pixels, the third training feature image is generated by the feature extraction network according to the training object image, and the third training feature image includes L X third training feature pixels, K X1 is an integer greater than 1, K X2 is an integer greater than 1, L X is an integer greater than 1; A training correlation matrix generation sub-module, configured to generate a first training correlation matrix according to the first training feature image and the third training feature image, where the first training correlation matrix includes K X1 ×L X similarity values, and the K X1 ×L X similarity values represent the similarity between K X1 first training feature pixels and L X third training feature pixels; The training-related matrix generation sub-module is further configured to generate a second training-related matrix according to the second training feature image and the third training feature image, where the second training-related matrix includes K X2 ×L X similarity values, and the K X2 ×L X similarity values represent the similarity degrees between K X2 second training feature pixels and L X third training feature pixels; Train a pixel-level spatial position matrix generation sub-module, which is used to take the first training correlation matrix as the input of the transformation network in the single-sample detection model, and output a first training pixel-level spatial position matrix through the transformation network, where the first training pixel-level spatial position matrix includes K X1 ×L X ×2 first training elements, and K X1 ×L X ×2 first training elements represent the corresponding position coordinates of L X1 third training feature pixels in the first training feature image when any one of the K X first training feature pixels is used as an anchor point; Train a pixel-level spatial position matrix generation sub-module, which is used to take the second training correlation matrix as the input of the transformation network in the single-sample detection model, and output a second training pixel-level spatial position matrix through the transformation network, where the second training pixel-level spatial position matrix includes K X2 ×L X ×2 second training elements, and K X2 ×L X ×2 second training elements represent the corresponding position coordinates of L X2 third training feature pixels in the second training feature image when any one of the K X second training feature pixels is used as an anchor point; A training detection box generation sub-module, configured to generate T training object first detection boxes in the first training sample image according to the first training pixel-level spatial position matrix, where X1 the T X1 training object first detection boxes include T X1 training objects, the T X1 training object first detection boxes correspond to T X1 first training confidence levels, the T X1 training object first detection boxes correspond to T X1 first detection box data, where T X1 is an integer greater than or equal to 1; A training detection box generation sub-module, configured to generate T training object second detection boxes in the second training sample image according to the second training pixel-level spatial position matrix, where X2 T X2 of the training object second detection boxes correspond to T X2 second training confidences, and T X2 of the training object second detection boxes correspond to T X2 second detection box data, where X2 is an integer greater than or equal to 1; A loss result generation sub-module, configured to generate a loss result of a single-sample detection model according to T X1 of the first training confidence levels, T X1 of the first detection box data, T X2 of the second training confidence levels, T X2 of the second detection box data, and T B of the labeled box data A model training sub-module, configured to train the single-sample detection model according to the single-sample detection model loss result.

12. The apparatus according to claim 11, wherein The first training sample image corresponds to a first confidence reference value, and the second training sample image corresponds to a second confidence reference value; The loss result generation sub-module is further configured to: According to T X1 generate a first classification loss result based on the first training confidence level and the first confidence level reference value; According to T X2 generate a second classification loss result based on the second training confidence and the second confidence reference value; According to T X1 pieces of the first detection box data, T X2 pieces of the second detection box data, and T B pieces of annotation box data, generate a positioning loss result; Generate the single-sample detection model loss result according to the first classification loss result, the second classification loss result, and the positioning loss result.

13. A computer device, characterized in that, Comprising: A memory, a transceiver, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is configured to execute the programs in the memory, including executing the image data processing method according to any one of claims 1 to 6; The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.

14. A computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the image data processing method according to any one of claims 1 to 6.

15. A computer program product, comprising a computer program, characterized in that, The computer program is executed by the processor to perform the image data processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • An image matching deep learning method and system

    CN109635824A

  • Multiple targets-tracking method and apparatus, device and storage medium

    US20190005657A1