Model training method, target recognition method, device, equipment and storage medium

By using the features of positive sample images as prior information in the self-attention memory neural network layer, the problem of insufficient discrimination ability of the model for residual targets in industrial quality inspection is solved, and a more efficient prediction effect is achieved.

CN114663687BActive Publication Date: 2025-05-23BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210255817.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-05-23
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

In industrial quality inspection scenarios, it is difficult for existing models to effectively distinguish positive and negative samples, resulting in insufficient discrimination ability of the model on residual targets, affecting the prediction effect.

Method used

By introducing a storage module into the self-attention memory neural network layer, the features of the positive sample image containing non-residual targets are used as prior information to perform feature mapping and fusion, and the model's detection ability of residual targets is improved.

Benefits of technology

It improves the ability to identify the residual targets of the model, and improves the prediction effect and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114663687B_ABST
    Figure CN114663687B_ABST
Patent Text Reader

Abstract

The present application proposes a model training method, a target recognition method, an apparatus, a device and a storage medium, wherein the method includes: dividing a sample image into blocks to obtain a plurality of first sub-blocks; extracting features from the plurality of first sub-blocks respectively to obtain sub-image features corresponding to the plurality of first sub-blocks; inputting each sub-image feature into the self-attention memory neural network layer in the recognition model, and using the attention mechanism to perform feature mapping according to the similarity between each sub-image feature and the corresponding target image feature, to obtain the mapping features corresponding to each first sub-block; fusing the mapping features of the plurality of first sub-blocks to obtain the fused features; using the prediction layer in the recognition model to perform target prediction on the fused features to obtain the predicted annotation information; and training the recognition model according to the difference between the predicted annotation information and the actual annotation information included in the sample image. Thus, the model's ability to distinguish defective targets can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a model training method, an object recognition method, a device, a device and a storage medium. Background Art

[0002] In a wide range of industrial production scenarios, such as the 3C, machinery manufacturing, semiconductor and electronics, chemical, pharmaceutical and other industries, the quality inspection of industrial products (referred to as industrial quality inspection for short) is an essential link. Among them, the main content involved in industrial quality inspection is the detection of appearance defects of products, including the detection of surface assembly, printing, shape and other defects.

[0003] Benefiting from the wide application of deep learning methods, a quality inspection model can be used to complete general recognition tasks in industrial quality inspection scenarios (such as classification, positioning, segmentation, etc. of defective products or defective areas), so as to replace traditional manual visual inspection and improve productivity, competitiveness and quality inspection accuracy. In order to improve the prediction effect of the model, how to implement the training of the model is very important. Summary of the Invention

[0004] The present application aims to solve at least one of the technical problems in the related art to some extent.

[0005] The present application provides a model training method, an object recognition method, a device, a device and a storage medium, so as to store the features of positive sample images containing non-defective objects through a self-attention memory neural network layer, which can provide prior information of positive sample images for the recognition model, and detect defective objects according to the prior information, which can improve the discrimination ability of the recognition model for defective objects, thereby improving the prediction effect of the model.

[0006] The first aspect of the present application provides a model training method, including:

[0007] Obtain a sample image, and divide the sample image into blocks to obtain a plurality of first sub-image blocks;

[0008] Extract features from each of the plurality of first sub-image blocks to obtain sub-image features corresponding to the plurality of first sub-image blocks;

[0009] Input the sub-image features corresponding to each first sub-image block into a self-attention memory neural network layer in a recognition model, and perform feature mapping using an attention mechanism according to the similarity between the sub-image features of each first sub-image block and the corresponding target image features, to obtain mapping features corresponding to each first sub-image block; wherein, the target image features are the image features of each second sub-image block divided from a positive sample image containing non-defective objects, and are the image features matching the sub-image features of the corresponding first sub-image block;

[0010] fusing the mapping features of the plurality of first sub-image blocks to obtain a fused feature;

[0011] Using the prediction layer in the recognition model, performing target prediction on the fused features to obtain prediction labeling information;

[0012] The recognition model is trained according to the difference between the predicted labeling information and the actual labeling information included in the sample image.

[0013] The second aspect of the present application provides a target recognition method, including:

[0014] Acquire an image to be detected, and divide the image to be detected into blocks to obtain multiple sub-blocks;

[0015] Extracting features from the plurality of sub-image blocks respectively to obtain sub-image features corresponding to the plurality of sub-image blocks;

[0016] Inputting the sub-image features corresponding to each of the sub-image blocks into the self-attention memory neural network layer in the recognition model to output the mapping features corresponding to each of the sub-image blocks; wherein the recognition model is trained using the method described in the embodiment of the first aspect of the present application;

[0017] Fusing the mapping features of the plurality of sub-image blocks to obtain a fused feature;

[0018] The prediction layer in the recognition model is used to perform target prediction on the fused features to obtain a recognition result of the target.

[0019] The third aspect of the present application provides a model training device, including:

[0020] An acquisition module, used for acquiring a sample image;

[0021] A segmentation module, used for segmenting the sample image into blocks to obtain a plurality of first sub-blocks;

[0022] An extraction module, used for performing feature extraction on the plurality of first sub-image blocks respectively, so as to obtain sub-image features corresponding to the plurality of first sub-image blocks;

[0023] An input module, used for inputting the sub-image features corresponding to each of the first sub-image blocks into the self-attention memory neural network layer in the recognition model, so as to perform feature mapping using the attention mechanism according to the similarity between the sub-image features of each of the first sub-image blocks and the corresponding target image features, and obtain the mapping features corresponding to each of the first sub-image blocks; wherein the target image features are image features of each of the second sub-image blocks divided from the positive sample image containing non-defective targets, which match the sub-image features of the corresponding first sub-image blocks;

[0024] A fusion module, used for fusing the mapping features of the plurality of first sub-blocks to obtain a fusion feature;

[0025] A prediction module, used to use the prediction layer in the recognition model to perform target prediction on the fusion feature to obtain prediction labeling information;

[0026] A training module is used to train the recognition model according to the difference between the predicted annotation information and the actual annotation information included in the sample image.

[0027] The fourth aspect of the present application provides a target recognition device, including:

[0028] An acquisition module, used for acquiring an image to be detected;

[0029] A segmentation module, used for segmenting the image to be detected into blocks to obtain a plurality of sub-blocks;

[0030] An extraction module, used for performing feature extraction on the plurality of sub-image blocks respectively, so as to obtain sub-image features corresponding to the plurality of sub-image blocks;

[0031] An input module, used for inputting the sub-image features corresponding to each of the sub-image blocks into the self-attention memory neural network layer in the recognition model, so as to output the mapping features corresponding to each of the sub-image blocks; wherein the recognition model is trained by the device as described in the embodiment of the third aspect of the present application;

[0032] A fusion module, used for fusing the mapping features of the plurality of sub-blocks to obtain a fusion feature;

[0033] The prediction module is used to use the prediction layer in the recognition model to perform target prediction on the fusion features to obtain the recognition result of the target.

[0034] The fifth aspect embodiment of the present application proposes a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the model training method proposed in the first aspect embodiment of the present application, or implements the target recognition method proposed in the second aspect embodiment of the present application.

[0035] The sixth aspect embodiment of the present application proposes a non-temporary computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the model training method proposed in the first aspect embodiment of the present application, or implements the target recognition method proposed in the second aspect embodiment of the present application.

[0036] The seventh aspect embodiment of the present application proposes a computer program product. When the instructions in the computer program product are executed by a processor, the model training method proposed in the first aspect embodiment of the present application is executed, or the target recognition method proposed in the second aspect embodiment of the present application is executed.

[0037] One embodiment of the present application has at least the following advantages or beneficial effects:

[0038] A plurality of first sub-blocks are obtained by dividing a sample image into blocks; feature extraction is performed on each of the plurality of first sub-blocks to obtain sub-image features corresponding to the plurality of first sub-blocks; the sub-image features corresponding to each first sub-block are input into a self-attention memory neural network layer in a recognition model, and feature mapping is performed using an attention mechanism according to the similarity between the sub-image features of each first sub-block and the corresponding target image features to obtain mapping features corresponding to each first sub-block; wherein the target image features are image features of each second sub-block divided from a positive sample image containing a non-defective target, which match the sub-image features of the corresponding first sub-block; the mapping features of the plurality of first sub-blocks are fused to obtain fused features; a prediction layer in the recognition model is used to perform target prediction on the fused features to obtain predicted annotation information; and the recognition model is trained according to the difference between the predicted annotation information and the actual annotation information included in the sample image. Therefore, by storing the features of positive sample images containing non-defective targets in the self-attention memory neural network layer, it is possible to provide the recognition model with prior information of the positive sample images, so as to detect the defective targets based on the prior information, thereby improving the recognition model's ability to distinguish defective targets and thus improving the model's prediction effect.

[0039] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0041] Figure 1 A flowchart of the model training method provided in Example 1 of the present application;

[0042] Figure 2 A flowchart of the model training method provided in Example 2 of the present application;

[0043] Figure 3 A flowchart of the model training method provided in Example 3 of the present application;

[0044] Figure 4A flowchart of the model training method provided in Example 4 of the present application;

[0045] Figure 5 This is a schematic diagram of the structure of the recognition model in the embodiment of the present application;

[0046] Figure 6 A flowchart of a target recognition method provided in Embodiment 5 of the present application;

[0047] Figure 7 This is a schematic diagram of the structure of the model training device provided in Example 6 of the present application;

[0048] Figure 8 This is a schematic diagram of the structure of the model training device provided in Example 7 of the present application;

[0049] Fig. 9 A block diagram of an exemplary computer device suitable for implementing embodiments of the present application is shown. DETAILED DESCRIPTION

[0050] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0051] Products that undergo industrial quality inspection usually have the following two characteristics: (1) There is a huge disparity in the distribution of samples of defective and non-defective products, that is, there are a large number of non-defective products and a small number of defective products; (2) The visual feature templates of the products are relatively fixed and simple (the reason is that the position of the quality inspection camera used for image acquisition is fixed, the shooting environment remains basically unchanged, and the shooting target / product appearance is unified on the same production line).

[0052] Thanks to the widespread application of deep learning methods, general image recognition tasks (classification, positioning, segmentation, etc. of defective products or defective areas) can be completed using existing machine vision detection algorithms. For example, the representative residual network (Residual Neural Network), by deepening the number of layers of convolutional neural networks and proposing residual structures, can extract richer feature information while solving the problem of easy degradation during deep neural network training. Another example is the deep self-attention model (Vision Transformer) that has replaced the convolutional neural network layer in recent years. This model extracts features from the input image, divides it into several sub-blocks and flattens it, and uses the self-attention module to establish long-distance dependencies and global connections between sub-blocks. It also has weights that dynamically adapt to input changes, and has a wider receptive field than models based on convolutional neural networks. For general image classification tasks, the performance of the deep self-attention model exceeds that of the convolutional neural network-based model.

[0053] At present, intelligent industrial quality inspection tasks based on computer vision are usually deployed in the above-mentioned general detection algorithm based on convolutional neural network, replacing traditional manual visual inspection, improving productivity, competitiveness and quality inspection accuracy.

[0054] However, the design of general detection algorithms usually ignores the disparity between positive and negative samples in the above-mentioned industrial quality inspection, as well as the single nature of visual feature patterns. Existing general image recognition, detection, and segmentation models are trained on general data sets, with high sample complexity and a wide variety of image features that need to be processed. However, industrial quality inspection tasks have a single sample size and rely more on "positive and negative samples". If each sample is trained as an independent individual, it is difficult to directly model the difference between "positive and negative samples". If the model is not allowed to compare the fixed features of defective products with those of the corresponding positive samples during the quality inspection of defective products, the model lacks prior knowledge of the samples, thereby reducing the model's ability to discriminate defective products. In actual industrial quality inspection scenarios, for example, there are two images containing wire mesh, one of which is a non-defective image (positive sample), and the other is a defective image with three sections of bent wire (negative sample). Through the self-attention mechanism of the deep self-attention model, all wires will be listed as observation objects. In contrast, since the three sections of bent wire in the negative sample only occupy a small area and the degree of bending is not high, the model will mistakenly identify the negative sample as a positive sample when identifying the negative sample. The reason is that the three defective areas in the negative sample are closer to the positive sample than the obvious defective product, which makes the model's judgment of the sample easily confused. In addition, for the entire image, the features of the defective area are not significant, which will cause the feature vectors of the positive and negative samples in the model to be close, making it difficult for the model to distinguish between positive and negative samples.

[0055] In summary, the lack of a mechanism to distinguish and process negative sample features affects the model's ability to judge negative sample features, thereby reducing the accuracy of the quality inspection model's prediction results. Even after replacing the convolutional neural network with a deep self-attention neural network, these problems still cannot be fundamentally solved. The core of the visual industrial quality inspection task is still the summary of prior knowledge of the task to be inspected and the comparative learning of positive and negative samples.

[0056] Therefore, in response to the above problems, this application mainly proposes a model training method to solve the problem of learning negative sample features in industrial quality inspection scenarios. That is, in view of the lack of consideration of the characteristics of industrial quality inspection data sets in the prior art (the product model is relatively fixed, and a large number of samples are non-defective products), this application introduces a storage module in the self-attention memory neural network layer to provide prior information for the input, thereby improving the model's ability to distinguish defective products.

[0057] That is, in this application, in order to make the model have better discrimination ability for defective products, a self-attention memory neural network layer with the function of integrating the features of prior non-defective targets (such as non-defective products) can be proposed. That is, the self-attention memory neural network layer integrates the features of positive sample images containing non-defective targets, which can effectively utilize the inherent characteristics of the quality inspection workpiece to provide prior information and complete the visual industrial quality inspection tasks.

[0058] The following describes the model training method, target recognition method, device, equipment and medium of the embodiments of the present application with reference to the accompanying drawings.

[0059] Figure 1 A flowchart of the model training method provided in Example 1 of the present application.

[0060] The embodiment of the present application takes the model training method being configured in a model training device as an example. The model training device can be applied to any electronic device so that the electronic device can perform a model training function.

[0061] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, a mobile terminal, a server, etc. The mobile terminal can be, for example, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, and other hardware devices with various operating systems, touch screens and / or display screens.

[0062] like Figure 1 As shown, the model training method may include the following steps:

[0063] Step 101: obtain a sample image, and divide the sample image into blocks to obtain a plurality of first sub-blocks.

[0064] In an embodiment of the present application, the sample image may be an image collected online, for example, the sample image may be collected online through web crawler technology, or the sample image may be an image collected offline, or the sample image may be an image collected in real time, or the sample image may be an artificially synthesized image, or the sample image may be an image obtained from an existing test set or training set, etc. The embodiment of the present application does not impose any restrictions on this.

[0065] In the embodiment of the present application, there may be multiple sample images, and each sample image may be annotated with annotation information, which is recorded as actual annotation information in the present application.

[0066] As an example, the recognition model is applied to a classification scenario or a classification task for illustrative explanation, and the actual annotation information may include the category of each target in the sample image.

[0067] For example, taking the application of the recognition model to the classification task in the industrial quality inspection scenario as an example, the sample image can be an image including the object to be inspected (such as the quality inspection product), the target in the sample image can be a defective area or a defective product, and the category of the target can be the category of the defective area or defective product. For example, when the object to be inspected is a mobile phone, the category of the target can include: no defect, scratches, depressions, black spots, white spots and other categories. For another example, when the object to be inspected is a road, the category of the target can include: no defect, cracks, protrusions, depressions and other categories.

[0068] As another example, taking the recognition model applied to a detection scenario or a detection task as an example, the actual annotation information may include the category of each target in the sample image and the prediction box containing each target (the prediction box may include location information).

[0069] For example, taking the application of the recognition model to the detection task in the industrial quality inspection scenario, the sample image can be an image including the object to be detected, the target in the sample image can be a defective area or a defective product, the category of the target can be a defective area or a category of defective products, and the prediction box containing the target can be a prediction box containing the defective area.

[0070] In the embodiment of the present application, after obtaining the sample image, the sample image can be divided into blocks to obtain multiple sub-blocks, which are referred to as first sub-blocks in the present application. For example, the sample image can be divided into n regions of the same size to obtain n first sub-blocks.

[0071] Step 102 : extracting features from the plurality of first sub-image blocks respectively to obtain sub-image features corresponding to the plurality of first sub-image blocks.

[0072] In an embodiment of the present application, for each first sub-image block, feature extraction can be performed on the first sub-image block based on a feature extraction algorithm to obtain image features corresponding to the first sub-image block, which are referred to as sub-image features in the present application.

[0073] In a possible implementation of the embodiment of the present application, in order to improve the accuracy and reliability of the feature extraction result, feature extraction can be performed on each first sub-block based on deep learning technology to obtain sub-image features corresponding to each first sub-block. For example, a mainstream backbone network (such as a residual network (ResNet), a DarkNet network (an open source neural network framework written in C and CUDA), etc.) can be used to extract features from the first sub-block to obtain sub-image features corresponding to each first sub-block.

[0074] In step 103, the sub-image features corresponding to each first sub-image block are input into the self-attention memory neural network layer in the recognition model, so as to perform feature mapping using the attention mechanism according to the similarity between the sub-image features of each first sub-image block and the corresponding target image features, and obtain the mapping features corresponding to each first sub-image block.

[0075] The target image feature is an image feature that matches the sub-image feature of the corresponding first sub-image block among the image features of each second sub-image block divided from the positive sample image containing the non-defective target.

[0076] In an embodiment of the present application, the positive sample image may be a sample image containing non-defective targets. For example, taking the application of the recognition model in an industrial quality inspection scenario as an example, the positive sample image may be an image containing non-defective products.

[0077] In an embodiment of the present application, the self-attention memory neural network layer in the recognition model can store image features of multiple second sub-blocks, that is, the positive sample image can be divided into blocks to obtain each second sub-block, and features of the second sub-blocks can be extracted to obtain image features of each second sub-block, so that the extracted image features of each second sub-block can be stored in the self-attention memory neural network layer.

[0078] In an embodiment of the present application, for each first sub-block, the sub-image features of the first sub-block can be matched with the image features of each second sub-block in the self-attention memory neural network layer, and the image features of the second sub-block that match the sub-image features of the first sub-block can be used as the target image features corresponding to the first sub-block.

[0079] In an embodiment of the present application, for each first sub-image block, an attention mechanism can be used to perform feature mapping on the sub-image features of the first sub-image block based on the similarity between the sub-image features of the first sub-image block and the corresponding target image features to obtain the mapping features corresponding to the first sub-image block.

[0080] Step 104: Fusing the mapping features of the plurality of first sub-image blocks to obtain a fused feature.

[0081] In the embodiment of the present application, mapping features of multiple first sub-blocks may be fused to obtain fused features.

[0082] As an example, the mapping features of the multiple first sub-image blocks may be spliced ​​according to the positions of the multiple first sub-image blocks in the sample image to obtain the fused features.

[0083] As another example, a fusion algorithm may be used to fuse the mapping features of multiple first sub-image blocks to obtain a fused feature.

[0084] As another example, mapping features of multiple first sub-image blocks may be spliced ​​according to their positions in the sample image to obtain spliced ​​features, and the spliced ​​features may be input into a convolutional layer to be fused to obtain the fused features.

[0085] Step 105, using the prediction layer in the recognition model, performs target prediction on the fused features to obtain predicted labeling information.

[0086] In an embodiment of the present application, the prediction layer in the recognition model can be used to perform target prediction on the fused features to obtain predicted annotation information.

[0087] As a possible implementation method, the recognition model is applied to a classification scenario or classification task for exemplary explanation, the prediction layer may be a FC (Fully Connected layers), and the FC in the recognition model may be used to predict the category of the target on the mapping feature to obtain the predicted annotation information of the sample image. The predicted annotation information may include the category to which the target in the sample image belongs.

[0088] It is understandable that the sample image may include at least one target. For example, there may be multiple defective areas in the sample image. Therefore, the predicted labeling information and the actual labeling information may include the category to which the at least one target belongs.

[0089] As another possible implementation method, the recognition model is applied to the detection scene or detection task for illustrative explanation. The prediction layer may include two branches, each of which may include multiple layers of convolutional layers, that is, each branch may be obtained by connecting multiple layers of convolutional layers in series. One of the branches can be used to predict the category of the target on the mapped features to obtain the category to which the target in the sample image belongs, and the other branch can be used to perform regression prediction of the target on the mapped features to obtain a prediction box containing the target.

[0090] Similarly, the sample image may include at least one target. For example, there may be multiple defective areas in the sample image. Therefore, the predicted annotation information and the actual annotation information may include at least one prediction box and the category to which the target in each prediction box belongs.

[0091] Step 106: training the recognition model according to the difference between the predicted annotation information and the actual annotation information included in the sample image.

[0092] In the embodiment of the present application, the difference between the predicted annotation information and the actual annotation information included in the sample image can be determined, and the recognition model can be trained according to the difference. For example, the recognition model can be trained according to the difference to minimize the difference, that is, the model parameters of the recognition model can be adjusted according to the difference to minimize the difference.

[0093] For example, a target loss function can be generated based on the above differences, and the recognition model can be trained based on the value of the target loss function to minimize the value of the target loss function, wherein the value of the target loss function is positively correlated with the above differences, that is, the smaller the difference, the smaller the value of the target loss function, and conversely, the larger the difference, the larger the value of the target loss function.

[0094] It should be noted that the above only uses the termination condition of model training as the minimization of the value of the objective loss function as an example. In actual application, other termination conditions can also be set. For example, the termination condition can also be that the number of training times reaches a set threshold, etc. This application does not impose any restrictions on this.

[0095] The model training method of the embodiment of the present application obtains multiple first sub-blocks by dividing the sample image into blocks; extracts features from the multiple first sub-blocks respectively to obtain sub-image features corresponding to the multiple first sub-blocks; inputs the sub-image features corresponding to each first sub-block into the self-attention memory neural network layer in the recognition model, and uses the attention mechanism to perform feature mapping according to the similarity between the sub-image features of each first sub-block and the corresponding target image features to obtain mapping features corresponding to each first sub-block; wherein the target image feature is an image feature that matches the sub-image feature of the corresponding first sub-block among the image features of each second sub-block divided from a positive sample image containing a non-defective target; fuses the mapping features of the multiple first sub-blocks to obtain a fused feature; uses the prediction layer in the recognition model to perform target prediction on the fused feature to obtain predicted annotation information; and trains the recognition model according to the difference between the predicted annotation information and the actual annotation information included in the sample image. Therefore, by storing the features of positive sample images containing non-defective targets in the self-attention memory neural network layer, it is possible to provide the recognition model with prior information of the positive sample images, so as to detect the defective targets based on the prior information, thereby improving the recognition model's ability to distinguish defective targets and thus improving the model's prediction effect.

[0096] In order to clearly illustrate how the self-attention memory neural network layer is used in the present application to perform feature mapping on the sub-image features of each first sub-block, this embodiment provides another model training method.

[0097] Figure 2 A flow chart of the model training method provided in Example 2 of the present application.

[0098] like Figure 2 As shown, the model training method may include the following steps:

[0099] Step 201: obtain a sample image, and divide the sample image into blocks to obtain a plurality of first sub-blocks.

[0100] Step 202 : extracting features from the plurality of first sub-image blocks respectively to obtain sub-image features corresponding to the plurality of first sub-image blocks.

[0101] The execution process of steps 201 to 202 can refer to the execution process of the above embodiment, and will not be described in detail here.

[0102] Step 203, obtaining multiple positive example image features stored in the self-attention memory neural network layer in the recognition model, wherein the multiple positive example image features are obtained after feature extraction of each second sub-block obtained after dividing the positive sample image into blocks.

[0103] In the embodiments of the present application, the explanation of the positive sample image can refer to the above embodiments and will not be repeated here.

[0104] In an embodiment of the present application, the positive sample image can be divided into blocks to obtain multiple second sub-blocks, and feature extraction is performed on the multiple second sub-blocks to obtain multiple positive image features, and the multiple positive image features are stored in the self-attention memory neural network layer.

[0105] Step 204 : Determine, from the plurality of positive example image features, target image features that match the sub-image features of each first sub-image block.

[0106] In an embodiment of the present application, multiple positive image features stored in the self-attention memory neural network layer in the recognition model can be obtained, and target image features that match the sub-image features of each first sub-block can be determined from the multiple positive image features.

[0107] In a possible implementation of the embodiment of the present application, for each first sub-block, the similarity between the sub-image feature of the first sub-block and multiple positive image features can be determined, and the positive image feature corresponding to the highest similarity is used as the target image feature that matches the sub-image feature of the first sub-block.

[0108] As an example, the number of marked first sub-blocks is n, the number of positive image features is m, and the sub-image feature of the jth first sub-block is q j , 1≤j≤n, the feature of the i-th positive image is p i , 1≤i≤m, then we can calculate q j With p i The cosine similarity of j The most relevant target image feature can be determined by the following formula (1) j Matching or most relevant target image features:

[0109] m j = argmax 1≤i≤m cosine(q j ,p i ),q' j =p mj ; (1)

[0110] Among them, q' j Indicates that j Matching target image features.

[0111] Step 205 , based on the similarity between the sub-image features of each first sub-image block and the corresponding target image features, feature mapping is performed on the sub-image features of each first sub-image block using an attention mechanism to obtain mapping features corresponding to each first sub-image block.

[0112] In an embodiment of the present application, the sub-image features of each first sub-image block can be feature mapped using an attention mechanism based on the similarity between the sub-image features of each first sub-image block and the corresponding target image features to obtain the mapping features corresponding to each first sub-image block.

[0113] Step 206: Fusing the mapping features of the plurality of first sub-image blocks to obtain a fused feature.

[0114] Step 207, using the prediction layer in the recognition model, performs target prediction on the fused features to obtain predicted labeling information.

[0115] Step 208: training the recognition model according to the difference between the predicted annotation information and the actual annotation information included in the sample image.

[0116] The execution process of steps 206 to 207 can refer to the execution process of the above embodiment, and will not be described in detail here.

[0117] The model training method of the embodiment of the present application can provide the recognition model with prior information of the positive sample images containing non-defective targets by storing the features of the positive sample images through the self-attention memory neural network layer, so as to detect the defective targets based on the prior information, thereby improving the recognition model's ability to distinguish defective targets, thereby improving the prediction effect of the model.

[0118] In order to clearly explain how the present application uses an attention mechanism to perform feature mapping on the sub-image features of each first sub-image block based on the similarity between the sub-image features of each first sub-image block and the corresponding target image features, this embodiment provides another model training method.

[0119] Figure 3 This is a flow chart of the model training method provided in Example 3 of the present application.

[0120] like Figure 3 As shown, the model training method may include the following steps:

[0121] Step 301: obtain a sample image, and divide the sample image into blocks to obtain a plurality of first sub-blocks.

[0122] Step 302 : extracting features from the plurality of first sub-image blocks respectively to obtain sub-image features corresponding to the plurality of first sub-image blocks.

[0123] Step 303, obtaining multiple positive example image features stored in the self-attention memory neural network layer, wherein the multiple positive example image features are obtained by performing feature extraction on each second sub-block obtained after dividing the positive sample image into blocks.

[0124] Step 304 : Determine, from the plurality of positive example image features, target image features that match the sub-image features of each first sub-image block.

[0125] The execution process of steps 301 to 304 may refer to the execution process of any of the above embodiments, and will not be described in detail here.

[0126] Step 305 : for each first sub-image block, determine the key-value feature corresponding to the first sub-image block according to the matched target image feature and the sub-image features of the plurality of first sub-image blocks.

[0127] For example, the sub-image feature q for the i-th first sub-block is i , 1≤i≤n, n is the number of the first sub-blocks, assuming that i The matching target image feature is q' i , then q i The corresponding key-value feature V can be:

[0128] V = {q 1 ,…,q n}∪{q' i}; (2)

[0129] Step 306 : determining an intermediate feature according to the similarity between the sub-image feature of the first sub-image block and the corresponding target image feature.

[0130] In the embodiment of the present application, the intermediate feature corresponding to the first sub-image block can be determined according to the similarity between the sub-image feature of the first sub-image block and the corresponding target image feature. Still taking the above example, the intermediate feature can be in, Performs a signed vector square root operation.

[0131] Step 307, normalize the inner product of the intermediate feature and the key value feature to obtain the attention weight.

[0132] In the embodiment of the present application, the inner product of the intermediate feature and the key feature can be normalized to obtain the attention weight. Still taking the above example, the attention value can be: Among them, softmax is the activation function and d is the vector dimension of the sub-image feature.

[0133] Step 308: weight the key-value features according to the attention weights to obtain mapping features corresponding to the first sub-block.

[0134] In an embodiment of the present application, the key-value features may be weighted according to the attention weight to obtain the mapping features corresponding to the first sub-block.

[0135] For example, the mapping feature corresponding to the i-th first sub-block may be determined according to the following formula (3):

[0136]

[0137] Among them, Attention(q i ) represents the mapping feature corresponding to the i-th first sub-block.

[0138] In summary, the attention mechanism not only considers the correlation between the sub-image features of the currently calculated first sub-block and the sub-image features of other first sub-blocks, but also considers the correlation between the sub-image features of the currently calculated first sub-block and the corresponding target image features, that is, the greater the above correlation, the greater the attention weight. In this way, the recognition model can capture important information in the image and improve the prediction effect of the model.

[0139] Step 309 , merging the mapping features of the multiple first sub-image blocks to obtain a fused feature.

[0140] Step 310, using the prediction layer in the recognition model, performs target prediction on the fused features to obtain predicted labeling information.

[0141] Step 311 : training the recognition model according to the difference between the predicted annotation information and the actual annotation information included in the sample image.

[0142] The execution process of steps 309 to 311 may refer to the execution process of any of the above embodiments, and will not be described in detail here.

[0143] The model training method of the embodiment of the present application realizes feature mapping of the sub-image features of each first sub-block through the attention mechanism, which can enable the recognition model to capture important information in the image and improve the prediction effect of the model.

[0144] In a possible implementation of the embodiment of the present application, the features of the positive sample images can also be dynamically updated according to the sample images in the training process to ensure that the image features of the positive sample images are effectively stored by the self-attention memory neural network layer. Figure 4 , the above process is explained in detail.

[0145] Figure 4 This is a flow chart of the model training method provided in Example 4 of the present application.

[0146] like Figure 4 As shown, the model training method may include the following steps:

[0147] Step 401: obtain a sample image, and divide the sample image into blocks to obtain a plurality of first sub-blocks.

[0148] Step 402 : extracting features from the plurality of first sub-image blocks respectively to obtain sub-image features corresponding to the plurality of first sub-image blocks.

[0149] Step 403, obtaining multiple positive example image features stored in the self-attention memory neural network layer, wherein the multiple positive example image features are obtained by extracting features from each second sub-block obtained after dividing the positive sample image into blocks.

[0150] The execution process of steps 401 to 403 may refer to the execution process of any of the above embodiments, and will not be described in detail here.

[0151] Step 404 : for each first sub-image block, determine the similarity between the sub-image feature of the first sub-image block and features of a plurality of positive example images.

[0152] In the embodiment of the present application, for each first sub-image block, the similarity between the sub-image feature of the first sub-image block and multiple positive image features can be calculated. For example, the cosine similarity between the sub-image feature of the first sub-image block and multiple positive image features can be calculated.

[0153] Step 405 : Determine the weight between the sub-image feature of the first sub-image block and the multiple positive example image features according to the similarity between the sub-image feature of the first sub-image block and the multiple positive example image features.

[0154] In an embodiment of the present application, for each first sub-image block, the weight between the sub-image feature of the first sub-image block and multiple positive image features can be determined according to the similarity between the sub-image feature of the first sub-image block and multiple positive image features.

[0155] As an example, for the sub-image feature q of the j-th first sub-image block j , 1≤j≤n, the q j and the i-th positive image feature p i The weights between can be:

[0156]

[0157] Among them, v i,j for q j With p i , 1≤i≤m, and m is the number of positive image features stored in the self-attention memory neural network layer.

[0158] Furthermore, we can also i,j Perform standardization to obtain the standardized weights:

[0159]

[0160] Step 406 : for each positive example image feature, weight the sub-image features of the plurality of first sub-image blocks according to the weight between the positive example image feature and the sub-image features of the plurality of first sub-image blocks to obtain a weighted image feature.

[0161] In an embodiment of the present application, for each positive example image feature, the sub-image features of multiple first sub-image blocks can be weighted according to the weight between the positive example image feature and the sub-image features of multiple first sub-image blocks to obtain a weighted image feature.

[0162] As an example, for the i-th positive image feature p i , the corresponding weighted image features can be: or

[0163] Step 407: Update the positive example image features according to the weighted image features to obtain updated positive example image features.

[0164] In the embodiment of the present application, for each positive example image feature, the positive example image feature can be updated according to the corresponding weighted image feature to obtain an updated positive example image feature.

[0165] As an example, for the i-th positive image feature p i , the p can be expressed by the following formula i To update:

[0166] or,

[0167] Wherein, f in formula (6) represents the L2 regularization operation.

[0168] It should be noted that, when the sample image obtained in step 401 is a positive sample image (an image containing a non-defective target), steps 404 to 407 may be performed, and when the sample image obtained in step 401 is a negative sample image (an image containing a defective target), steps 404 to 407 may not be performed. Alternatively, considering the large disparity in the ratio of positive sample images to negative sample images, and the extremely unbalanced distribution of positive sample images and negative sample images, it can be ensured that the vast majority of the features stored in the self-attention memory neural network layer correspond to features related to the positive sample image, that is, regardless of whether the sample image obtained in step 401 is a positive sample image or a negative sample image, steps 404 to 407 may be performed, and the present application does not impose any limitation on this.

[0169] Step 408 : Determine, from the updated plurality of positive example image features, target image features that match the sub-image features of each first sub-image block.

[0170] Step 409 , based on the similarity between the sub-image features of each first sub-image block and the corresponding target image features, feature mapping is performed on the sub-image features of each first sub-image block using an attention mechanism to obtain mapping features corresponding to each first sub-image block.

[0171] Step 410: Fusing the mapping features of the plurality of first sub-image blocks to obtain a fused feature.

[0172] Step 411, using the prediction layer in the recognition model, performs target prediction on the fused features to obtain predicted labeling information.

[0173] Step 412: training the recognition model according to the difference between the predicted annotation information and the actual annotation information included in the sample image.

[0174] The execution process of steps 408 to 412 may refer to the execution process of any of the above embodiments, and will not be described in detail here.

[0175] As an example, the structure of the recognition model can be as follows Figure 5 As shown, the recognition model may include multiple layers of self-attention memory neural network layers. Before training the recognition model using sample images containing the object to be detected (such as quality inspection products), the sample images may be randomly flipped, scaled, cropped, and other data enhancement operations that can improve the generalization ability of the model. Afterwards, the sample images may be divided into n regions of the same size and input into the self-attention memory neural network layer in the recognition model.

[0176] Specifically, considering that in industrial quality inspection scenarios, the vast majority of samples are positive sample images and there are only a very small number of negative sample images (images containing defective targets or defective products), the sub-image features of n sub-blocks can be compared with the positive image features stored in the self-attention memory neural network layer to determine the positive image features that are most similar to each sub-block. At the same time, by utilizing the large difference in the distribution of positive sample images and negative sample images, a large number of similar image features can be clustered and updated to ensure that the features of the positive sample images are effectively stored.

[0177] The self-attention memory neural network layer may include a storage operation module and a self-attention operation module. The storage operation module mainly involves the following two groups of operation operations: update and query.

[0178] First, update: In order to update the positive image features stored in the self-attention memory neural network layer, for the sub-image features of each sub-block in the sample image, the positive image features that match the sub-image features can be queried, and then the positive image features stored in the self-attention memory neural network layer can be further corrected by combining the sub-image features with the weighted judgment results. In this way, the positive image features stored in the self-attention memory neural network layer can be adjusted accordingly according to the sub-blocks in the sample image to achieve the effect of memory learning. Specifically, the positive image features p stored in the self-attention memory neural network layer can be i and the sub-image feature q in the sample image j Calculate the cosine similarity and then normalize it to get q j With p i The weight between:

[0179]

[0180] Among them, 1≤j≤n, n is the number of sub-blocks in the sample image, 1≤i≤m, m is the number of positive image features stored in the self-attention memory neural network layer.

[0181] Furthermore, the weight v of all queries i,j After re-standardization, the standardized weights are obtained:

[0182]

[0183] Finally, the n sub-image features in the sample image can be fused into the positive image features stored in the self-attention memory neural network layer to obtain the updated positive image features:

[0184]

[0185] Second, query: For each sub-block in the sample image, the positive image features that are most similar to the sub-block can be queried. Specifically, each sub-image feature q can be calculated j With all the updated positive image features p i The cosine similarity between j The most relevant positive image feature is used as the target image feature q' j :

[0186]

[0187] Since the ratio of positive sample images to negative sample images is very different in industrial quality inspection tasks, according to the weight calculation method of formula (4), in the process of updating the positive sample image features, the extremely uneven distribution of positive sample images and negative sample images ensures that the positive sample image features stored in the self-attention memory neural network layer mostly correspond to features related to the positive sample images. Finally, in the query process, for both positive sample images and negative sample images, only one positive sample image feature that is most similar to it is returned as the query result, ensuring the correlation between the result returned by each query and the sub-image feature of the corresponding sub-block.

[0188] Among them, the self-attention operation module: the sub-image features q of all sub-blocks of the sample image 1 ,…,q n , and the most similar positive image feature q' 1 ,…,q' n In this application, for the sub-image feature q of the sub-image block in the sample image, i , the range of self-attention operation can be set to V = {q 1 ,…,q n}∪{q' i}, where {q 1 ,…,q n The self-attention operation in} describes the correlation between the sub-image features of the currently calculated sub-block and the sub-image features of other sub-blocks. i Before the self-attention operation with V, it will first be compared with q' i The multiplication operation is performed to characterize the correlation and difference between the sub-image features of the currently calculated sub-block and the corresponding target image features. Specifically, for the sub-image features of any sub-block, the self-attention operation process can be shown as follows:

[0189]

[0190] By using formula (3), the corresponding self-attention output result can be calculated for the sub-image feature of each sub-block, which is denoted as the mapping feature Attention(q i ), the mapping features corresponding to each sub-block can be input into the next self-attention memory neural network layer.

[0191] After feature mapping or feature transformation of multiple layers of self-attention memory neural network layers, the mapping features corresponding to each sub-block output by the last layer of self-attention memory neural network layers can be obtained. Thus, the feature information of the positive sample images in the entire image training data set and the feature vectors (i.e., mapping features) after the mutual correlation of each sub-block in the current sample image can be comprehensively considered. The mapping features corresponding to each sub-block output by the last layer of self-attention memory neural network layers can be directly input into the loss function of tasks such as defective product detection / defective area detection / defective area segmentation for end-to-end neural network training.

[0192] By using the above method, the recognition model is trained on industrial datasets such as SDNET2018, KolektorSDD, and TIG_Aluminium, and the performance is comparable to that of the standard self-attention model and deep convolutional neural network model when only 50% of the training data is used. The difference between positive and negative sample images in industrial quality inspection scenarios can be effectively mined, the dependence of model training on the amount of labeled data can be effectively reduced, and the model development cycle and cost can be greatly reduced.

[0193] It should be noted that visual industrial quality inspection is a very important part of intelligent manufacturing and an indispensable part of the new generation of intelligent supply chain. Traditional visual industrial quality inspection requires a lot of manpower and financial resources, is costly, and the quality of quality inspection is uncontrollable. Although the visual industrial quality inspection technology based on deep learning can replace manual quality inspection tasks to a certain extent by using powerful computing power support, the training of the visual industrial quality inspection model based on deep learning requires a large amount of labeled data. This is mainly because the existing industrial quality inspection model cannot deeply explore the difference information of positive and negative samples in industrial quality inspection scenarios.

[0194] The recognition model proposed in this application, which includes a multi-layer self-attention memory neural network layer, integrates the memory network with the self-attention network. When applied to image feature learning in the field of industrial quality inspection or other image classification / detection / segmentation tasks, it can effectively utilize the disparity between the number of positive sample images and negative sample images in industrial quality inspection tasks, adaptively record the image features in the positive sample images in the entire training data, and compare / associate with the input image features. While improving the model performance, it greatly reduces the dependence of model training on data labeling, greatly shortens the model development cycle, and reduces the data labeling cost.

[0195] The above are various embodiments corresponding to the training method of the recognition model. The present application also proposes an application method of the recognition model, that is, the recognition model is used for target recognition.

[0196] Figure 6This is a flowchart of the target recognition method provided in Example 5 of the present application.

[0197] like Figure 6 As shown, the target recognition method may include the following steps:

[0198] Step 601: obtain an image to be detected, and divide the image to be detected into blocks to obtain multiple sub-blocks.

[0199] In an embodiment of the present application, the image to be detected may be an image collected online, for example, the image to be detected may be collected online through web crawler technology, or the image to be detected may be an image collected offline, or the image to be detected may be an image collected in real time, or the image to be detected may be an artificially synthesized image, or the image to be detected may be an image obtained from an existing test set, etc. The embodiment of the present application does not impose any restrictions on this.

[0200] In the embodiment of the present application, after the image to be detected is acquired, the image to be detected may be divided into blocks to obtain a plurality of sub-blocks. For example, the image to be detected may be divided into n regions of the same size to obtain n sub-blocks.

[0201] Step 602 : extracting features from the plurality of sub-image blocks respectively to obtain sub-image features corresponding to the plurality of sub-image blocks.

[0202] In an embodiment of the present application, for each sub-block, feature extraction can be performed on the sub-block based on a feature extraction algorithm to obtain image features corresponding to the sub-block, which are referred to as sub-image features in the present application.

[0203] Step 603, input the sub-image features corresponding to each sub-block into the self-attention memory neural network layer of the recognition model to output the mapping features corresponding to each sub-block.

[0204] Among them, the recognition model is based on the above Figures 1 to 4 It is obtained by training the model training method shown in any embodiment. It should be noted that the above explanation of the model training method embodiment is also applicable to this embodiment, and its implementation principle is similar, which will not be repeated here.

[0205] In the embodiment of the present application, the sub-image features corresponding to each sub-image block can be input into the self-attention memory neural network layer of the recognition model, so that the self-attention memory neural network layer outputs the mapping features corresponding to each sub-image block. That is, the self-attention memory neural network layer can use the attention mechanism to perform feature mapping on the sub-image features of the corresponding sub-image block according to the similarity between the sub-image features of each sub-image block and the corresponding target image features, and obtain the mapping features corresponding to each sub-image block.

[0206] Step 604: Fusing the mapping features of multiple sub-blocks to obtain a fused feature.

[0207] As an example, mapping features of multiple sub-image blocks may be concatenated according to positions of the multiple sub-image blocks in the image to be detected to obtain fused features.

[0208] As another example, a fusion algorithm may be used to fuse mapping features of multiple sub-blocks to obtain fused features.

[0209] As another example, mapping features of multiple sub-image blocks may be spliced ​​according to their positions in the sample image to obtain spliced ​​features, and the spliced ​​features may be input into a convolutional layer to be fused to obtain the fused features.

[0210] Step 605, using the prediction layer in the recognition model, performs target prediction on the fused features to obtain the recognition result of the target.

[0211] As a possible implementation method, the recognition model is applied to a classification scenario or classification task for exemplary explanation, the prediction layer may be a FC (Fully Connected layers), and the FC in the recognition model may be used to predict the category of the target on the mapping feature to obtain the recognition result of the target. The recognition result may include the category to which the target in the image to be detected belongs.

[0212] For example, taking the application of the recognition model to the classification task in the industrial quality inspection scenario as an example, the image to be detected can be an image including the object to be detected, the target in the image to be detected can be a defective area or a defective product, and the category of the target can be the category of the defective area or defective product. For example, when the object to be detected is a mobile phone, the category of the target can include: no defect, scratches, depressions, black spots, white spots and other categories. For another example, when the object to be detected is a road, the category of the target can include: no defect, cracks, protrusions, depressions and other categories.

[0213] As another possible implementation, the recognition model is applied to a detection scene or detection task for exemplary explanation. The prediction layer may include two branches, each of which may include multiple convolutional layers, that is, each branch may be obtained by connecting multiple convolutional layers in series. One of the branches may be used to predict the category of the target on the mapped features to obtain the category to which the target in the image to be detected belongs, and another branch may be used to perform regression prediction of the target on the mapped features to obtain a prediction box containing the target. In other words, the recognition result may include the category to which the target in the image to be detected belongs, and the prediction box containing the target.

[0214] For example, taking the application of the recognition model to the detection task in the industrial quality inspection scenario, the sample image can be an image including the object to be detected, the target in the sample image can be a defective area, the category of the target can be the category of the defective area, and the prediction box containing the target can be a prediction box containing the defective area.

[0215] The target recognition method of the embodiment of the present application obtains an image to be detected, and divides the image to be detected into blocks to obtain multiple sub-blocks; extracts features from the multiple sub-blocks respectively to obtain sub-image features corresponding to the multiple sub-blocks; inputs the sub-image features corresponding to each sub-block into the self-attention memory neural network layer in the recognition model to output the mapping features corresponding to each sub-block; fuses the mapping features of the multiple sub-blocks to obtain fused features; uses the prediction layer in the recognition model to perform target prediction on the fused features to obtain the target recognition result. Therefore, based on deep learning technology, target prediction for the image to be detected can improve the accuracy and reliability of the prediction results.

[0216] With the above Figures 1 to 4 Corresponding to the model training method provided in the embodiment, the present application also provides a model training device. Figures 1 to 4 The model training method provided in the embodiment corresponds to the embodiment, so the implementation method of the model training method is also applicable to the model training device provided in the embodiment of the present application, and will not be described in detail in the embodiment of the present application.

[0217] Figure 7 This is a structural diagram of the model training device provided in Example 6 of the present application.

[0218] like Figure 7 As shown, the model training device 700 may include: an acquisition module 710, a segmentation module 720, an extraction module 730, an input module 740, a fusion module 750, a prediction module 760 and a training module 770.

[0219] The acquisition module 710 is used to acquire a sample image.

[0220] The segmentation module 720 is used to segment the sample image into blocks to obtain a plurality of first sub-blocks.

[0221] The extraction module 730 is used to perform feature extraction on the multiple first sub-image blocks respectively to obtain sub-image features corresponding to the multiple first sub-image blocks.

[0222] The input module 740 is used to input the sub-image features corresponding to each first sub-image block into the self-attention memory neural network layer of the recognition model, so as to perform feature mapping using the attention mechanism according to the similarity between the sub-image features of each first sub-image block and the corresponding target image features, so as to obtain the mapping features corresponding to each first sub-image block; wherein the target image features are image features of the second sub-image blocks divided from the positive sample image containing non-defective targets, which match the sub-image features of the corresponding first sub-image blocks.

[0223] The fusion module 750 is used to fuse the mapping features of multiple first sub-blocks to obtain a fused feature.

[0224] The prediction module 760 is used to use the prediction layer in the recognition model to perform target prediction on the fused features to obtain prediction labeling information.

[0225] The training module 770 is used to train the recognition model according to the difference between the predicted annotation information and the actual annotation information included in the sample image.

[0226] In a possible implementation of the embodiment of the present application, the input module 740 may include:

[0227] An acquisition unit is used to acquire multiple positive example image features stored in the self-attention memory neural network layer, wherein the multiple positive example image features are obtained after feature extraction of each second sub-block obtained after dividing the positive sample image into blocks.

[0228] The determination unit is used to determine, from a plurality of positive example image features, target image features that match the sub-image features of each first sub-image block.

[0229] A mapping unit is used to perform feature mapping on the sub-image features of each first sub-image block using an attention mechanism according to the similarity between the sub-image features of each first sub-image block and the corresponding target image features, so as to obtain mapping features corresponding to each first sub-image block.

[0230] In a possible implementation of an embodiment of the present application, the determination unit is specifically used to: determine, for each first sub-image block, the similarity between the corresponding sub-image feature and multiple positive image features; and use the positive image feature corresponding to the highest similarity as the target image feature that matches the sub-image feature of the first sub-image block.

[0231] In a possible implementation manner of an embodiment of the present application, the mapping unit is specifically used to: for each first sub-block, determine the key-value feature corresponding to the first sub-block according to the matched target image feature and the sub-image features of multiple first sub-blocks; determine the intermediate feature according to the similarity between the sub-image feature of the first sub-block and the corresponding target image feature; normalize the inner product of the intermediate feature and the key-value feature to obtain an attention weight; and weight the key-value feature according to the attention weight to obtain the mapping feature corresponding to the first sub-block.

[0232] In a possible implementation manner of an embodiment of the present application, the determination unit is further used to determine, for each first sub-block, the similarity between the sub-image feature of the first sub-block and multiple positive image features, and determine the weight between the sub-image feature of the first sub-block and the multiple positive image features based on the similarity between the sub-image feature of the first sub-block and the multiple positive image features.

[0233] The input module 740 may further include:

[0234] The weighting unit is used to weight the sub-image features of the multiple first sub-image blocks according to the weight between the positive image feature and the sub-image features of the multiple first sub-image blocks for each positive image feature, so as to obtain a weighted image feature.

[0235] The updating unit is used to update the positive example image feature according to the weighted image feature to obtain the updated positive example image feature.

[0236] In a possible implementation of the embodiment of the present application, the prediction module 760 is specifically used to: use the fully connected layer in the prediction layer to predict the category of the target for the fused features to obtain the category to which the target belongs.

[0237] In a possible implementation of an embodiment of the present application, the prediction module 760 is specifically used to: use the first branch in the prediction layer to perform target category prediction on the fused features to obtain the category to which the target belongs; use the second branch in the prediction layer to perform regression prediction on the fused features to obtain a prediction box containing the target.

[0238] The model training device of the embodiment of the present application obtains multiple first sub-blocks by dividing the sample image into blocks; extracts features from the multiple first sub-blocks respectively to obtain sub-image features corresponding to the multiple first sub-blocks; inputs the sub-image features corresponding to each first sub-block into the self-attention memory neural network layer in the recognition model, and uses the attention mechanism to perform feature mapping according to the similarity between the sub-image features of each first sub-block and the corresponding target image features to obtain mapping features corresponding to each first sub-block; wherein the target image feature is the image feature of each second sub-block divided from the positive sample image containing non-defective targets, which matches the sub-image feature of the corresponding first sub-block; fuses the mapping features of the multiple first sub-blocks to obtain fused features; uses the prediction layer in the recognition model to perform target prediction on the fused features to obtain predicted annotation information; and trains the recognition model according to the difference between the predicted annotation information and the actual annotation information included in the sample image. Therefore, by storing the features of positive sample images containing non-defective targets in the self-attention memory neural network layer, it is possible to provide the recognition model with prior information of the positive sample images, so as to detect the defective targets based on the prior information, thereby improving the recognition model's ability to distinguish defective targets and thus improving the model's prediction effect.

[0239] With the above Figure 6 Corresponding to the target recognition method provided in the embodiment, the present application also provides a target recognition device. Since the model training device provided in the embodiment of the present application is similar to the above Figure 6 The target recognition method provided in the embodiment corresponds to the target recognition method, so the implementation method of the target recognition method is also applicable to the target recognition device provided in the embodiment of the present application, and will not be described in detail in the embodiment of the present application.

[0240] Figure 8 This is a schematic diagram of the structure of the target recognition device provided in Example 7 of the present application.

[0241] like Figure 8 As shown, the model training device 800 may include: an acquisition module 810, a segmentation module 820, an extraction module 830, an input module 840, a fusion module 850 and a prediction module 860.

[0242] The acquisition module 810 is used to acquire the image to be detected.

[0243] The segmentation module 820 is used to divide the image to be detected into blocks to obtain multiple sub-blocks.

[0244] The extraction module 830 is used to perform feature extraction on the multiple sub-image blocks respectively to obtain sub-image features corresponding to the multiple sub-image blocks.

[0245] The input module 840 is used to input the sub-image features corresponding to each sub-block into the self-attention memory neural network layer of the recognition model to output the mapping features corresponding to each sub-block. Figure 7 The device described in the embodiment is trained.

[0246] The fusion module 850 is used to fuse the mapping features of multiple sub-blocks to obtain a fused feature.

[0247] The prediction module 860 is used to use the prediction layer in the recognition model to perform target prediction on the fused features to obtain the recognition result of the target.

[0248] The model training device of the embodiment of the present application obtains the image to be detected and divides the image to be detected into blocks to obtain multiple sub-blocks; extracts features from the multiple sub-blocks respectively to obtain sub-image features corresponding to the multiple sub-blocks; inputs the sub-image features corresponding to each sub-block into the self-attention memory neural network layer in the recognition model to output the mapping features corresponding to each sub-block; fuses the mapping features of the multiple sub-blocks to obtain the fused features; uses the prediction layer in the recognition model to perform target prediction on the fused features to obtain the recognition result of the target. Therefore, based on the deep learning technology, the target prediction of the image to be detected can improve the accuracy and reliability of the prediction results.

[0249] In order to implement the above embodiments, the present application also proposes a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the model training method proposed in any of the above embodiments of the present application, or implements the target recognition method proposed in the above embodiments of the present application.

[0250] In order to implement the above embodiments, the present application also proposes a non-temporary computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the model training method proposed in any of the aforementioned embodiments of the present application, or implements the target recognition method proposed in the aforementioned embodiments of the present application.

[0251] In order to implement the above embodiments, the present application also proposes a computer program product. When the instructions in the computer program product are executed by a processor, the model training method proposed in any of the aforementioned embodiments of the present application is executed, or the target recognition method proposed in the aforementioned embodiments of the present application is implemented.

[0252] Fig. 9 A block diagram of an exemplary computer device suitable for implementing embodiments of the present application is shown. Fig. 9 The computer device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0253] like Fig. 9 As shown, the computer device 12 is in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).

[0254] The bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor or a local bus using any of a variety of bus structures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus and Peripheral Component Interconnection (PCI) bus.

[0255] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0256] The memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Fig. 9 not shown, usually called a "hard drive"). Although Fig. 9Not shown, a disk drive for reading and writing a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing a removable non-volatile optical disk (e.g., a compact disc read only memory (CD-ROM), a digital versatile disc read only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present application.

[0257] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described herein.

[0258] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface 22. In addition, the computer device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the computer device 12 via a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0259] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the methods mentioned in the above embodiments.

[0260] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0261] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0262] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0263] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0264] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0265] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0266] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0267] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A model training method, It is characterized in that The method comprises the following steps: Acquire a sample image, and divide the sample image into blocks to obtain a plurality of first sub-blocks; Performing feature extraction on the plurality of first sub-image blocks respectively to obtain sub-image features corresponding to the plurality of first sub-image blocks; Inputting the sub-image features corresponding to each of the first sub-image blocks into the self-attention memory neural network layer in the recognition model, so as to perform feature mapping using the attention mechanism according to the similarity between the sub-image features of each of the first sub-image blocks and the corresponding target image features, and obtain the mapping features corresponding to each of the first sub-image blocks; wherein the target image features are image features of each of the second sub-image blocks divided from the positive sample image containing non-defective targets, which match the sub-image features of the corresponding first sub-image blocks; fusing the mapping features of the plurality of first sub-image blocks to obtain a fused feature; Using the prediction layer in the recognition model, performing target prediction on the fused features to obtain prediction labeling information; The recognition model is trained according to the difference between the predicted labeling information and the actual labeling information included in the sample image.

2. The method according to claim 1, It is characterized in that The step of inputting the sub-image features corresponding to each of the first sub-image blocks into the self-attention memory neural network layer in the recognition model, and performing feature mapping using an attention mechanism according to the similarity between the sub-image features of each of the first sub-image blocks and the corresponding target image features, to obtain mapping features corresponding to each of the first sub-image blocks, includes: Acquire multiple positive example image features stored in the self-attention memory neural network layer, wherein the multiple positive example image features are obtained by performing feature extraction on each second sub-image block obtained after dividing the positive sample image into blocks; Determining, from the plurality of positive example image features, target image features that match the sub-image features of each of the first sub-image blocks; According to the similarity between the sub-image features of each of the first sub-image blocks and the corresponding target image features, the sub-image features of each of the first sub-image blocks are feature mapped using an attention mechanism to obtain mapping features corresponding to each of the first sub-image blocks.

3. The method according to claim 2, It is characterized in that The step of respectively determining target image features that match the sub-image features of each of the first sub-image blocks from the plurality of positive image features comprises: For each of the first sub-image blocks, determining a similarity between a corresponding sub-image feature and a plurality of the positive example image features; The positive example image feature corresponding to the highest similarity is used as the target image feature that matches the sub-image feature of the first sub-image block.

4. The method according to claim 2, It is characterized in that The method of performing feature mapping on the sub-image features of each of the first sub-image blocks using an attention mechanism according to the similarity between the sub-image features of each of the first sub-image blocks and the corresponding target image features to obtain mapping features corresponding to each of the first sub-image blocks includes: For each of the first sub-image blocks, determining a key-value feature corresponding to the first sub-image block according to the matched target image feature and the sub-image features of the plurality of first sub-image blocks; Determining an intermediate feature according to a similarity between a sub-image feature of the first sub-image block and a corresponding target image feature; The inner product of the intermediate feature and the key value feature is normalized to obtain an attention weight; The key-value features are weighted according to the attention weights to obtain mapping features corresponding to the first sub-block.

5. The method according to claim 2, It is characterized in that After acquiring the plurality of positive image features stored in the self-attention memory neural network layer, the method further comprises: For each of the first sub-image blocks, determining a similarity between a sub-image feature of the first sub-image block and a plurality of features of the positive example images; Determining a weight between the sub-image feature of the first sub-image block and the multiple positive example image features according to the similarity between the sub-image feature of the first sub-image block and the multiple positive example image features; For each of the positive example image features, weighting the sub-image features of the plurality of first sub-image blocks according to a weight between the positive example image feature and the sub-image features of the plurality of first sub-image blocks to obtain a weighted image feature; The positive example image feature is updated according to the weighted image feature to obtain the updated positive example image feature.

6. The method according to any one of claims 1 to 5, in, The using the prediction layer in the recognition model to perform target prediction on the fusion feature to obtain prediction labeling information includes: The fully connected layer in the prediction layer is used to predict the category of the target for the fused features to obtain the category to which the target belongs.

7. The method according to any one of claims 1 to 5, in, The using the prediction layer in the recognition model to perform target prediction on the fusion feature to obtain prediction labeling information includes: Using the first branch in the prediction layer, predicting the category of the target on the fused features, and obtaining the category to which the target belongs; The second branch in the prediction layer is used to perform regression prediction of the target on the fused features to obtain a prediction box containing the target.

8. A target recognition method, It is characterized in that The method comprises the following steps: Acquire an image to be detected, and divide the image to be detected into blocks to obtain multiple sub-blocks; Performing feature extraction on the plurality of sub-image blocks respectively to obtain sub-image features corresponding to the plurality of sub-image blocks; Inputting the sub-image features corresponding to each of the sub-image blocks into the self-attention memory neural network layer in the recognition model to output the mapping features corresponding to each of the sub-image blocks; wherein the recognition model is trained using the method according to any one of claims 1 to 7; Fusing the mapping features of the plurality of sub-image blocks to obtain a fused feature; The prediction layer in the recognition model is used to perform target prediction on the fused features to obtain a recognition result of the target.

9. A model training device, It is characterized in that The device comprises: An acquisition module, used for acquiring a sample image; A segmentation module, used for segmenting the sample image into blocks to obtain a plurality of first sub-blocks; An extraction module, used for performing feature extraction on the plurality of first sub-image blocks respectively, so as to obtain sub-image features corresponding to the plurality of first sub-image blocks; An input module, used for inputting the sub-image features corresponding to each of the first sub-image blocks into the self-attention memory neural network layer in the recognition model, so as to perform feature mapping using the attention mechanism according to the similarity between the sub-image features of each of the first sub-image blocks and the corresponding target image features, and obtain the mapping features corresponding to each of the first sub-image blocks; wherein the target image features are image features of each of the second sub-image blocks divided from the positive sample image containing non-defective targets, which match the sub-image features of the corresponding first sub-image blocks; A fusion module, used for fusing the mapping features of the plurality of first sub-blocks to obtain a fusion feature; A prediction module, used to use the prediction layer in the recognition model to perform target prediction on the fusion feature to obtain prediction labeling information; A training module is used to train the recognition model according to the difference between the predicted annotation information and the actual annotation information included in the sample image.

10. The device according to claim 9, It is characterized in that The input module comprises: An acquisition unit is used to acquire a plurality of positive example image features stored in the self-attention memory neural network layer, wherein the plurality of positive example image features are obtained by performing feature extraction on each second sub-image block obtained after dividing the positive sample image into blocks; a determining unit, configured to respectively determine, from the plurality of positive example image features, target image features that match the sub-image features of each of the first sub-image blocks; A mapping unit is used to perform feature mapping on the sub-image features of each of the first sub-image blocks using an attention mechanism according to the similarity between the sub-image features of each of the first sub-image blocks and the corresponding target image features, so as to obtain mapping features corresponding to each of the first sub-image blocks.

11. The device according to claim 10, It is characterized in that The determining unit is specifically configured to: For each of the first sub-image blocks, determining a similarity between a corresponding sub-image feature and a plurality of the positive example image features; The positive example image feature corresponding to the highest similarity is used as the target image feature that matches the sub-image feature of the first sub-image block.

12. The device according to claim 10, It is characterized in that The mapping unit is specifically used for: For each of the first sub-image blocks, determining a key-value feature corresponding to the first sub-image block according to the matched target image feature and the sub-image features of the plurality of first sub-image blocks; Determining an intermediate feature according to a similarity between a sub-image feature of the first sub-image block and a corresponding target image feature; The inner product of the intermediate feature and the key value feature is normalized to obtain an attention weight; The key-value features are weighted according to the attention weights to obtain mapping features corresponding to the first sub-block.

13. The device according to claim 10, It is characterized in that The determining unit is further configured to determine, for each of the first sub-image blocks, a similarity between a sub-image feature of the first sub-image block and a plurality of features of the positive example images, and determine a weight between the sub-image feature of the first sub-image block and the plurality of features of the positive example images according to the similarity between the sub-image feature of the first sub-image block and the plurality of features of the positive example images; The input module further includes: a weighting unit, configured to weight the sub-image features of the plurality of first sub-image blocks according to a weight between the positive example image feature and the sub-image features of the plurality of first sub-image blocks for each of the positive example image features, so as to obtain a weighted image feature; An updating unit is used to update the positive example image feature according to the weighted image feature to obtain the updated positive example image feature.

14. The device according to any one of claims 9 to 13, in, The prediction module is specifically used for: The fully connected layer in the prediction layer is used to predict the category of the target for the fused features to obtain the category to which the target belongs.

15. The device according to any one of claims 9 to 13, in, The prediction module is specifically used for: Using the first branch in the prediction layer, predicting the category of the target on the fused features, and obtaining the category to which the target belongs; The second branch in the prediction layer is used to perform regression prediction of the target on the fused features to obtain a prediction box containing the target.

16. A target recognition device, It is characterized in that The device comprises: An acquisition module, used for acquiring an image to be detected; A segmentation module, used for segmenting the image to be detected into blocks to obtain a plurality of sub-blocks; An extraction module, used for performing feature extraction on the plurality of sub-image blocks respectively, so as to obtain sub-image features corresponding to the plurality of sub-image blocks; An input module, used for inputting the sub-image features corresponding to each of the sub-image blocks into the self-attention memory neural network layer in the recognition model to output the mapping features corresponding to each of the sub-image blocks; wherein the recognition model is trained using the device according to any one of claims 9 to 15; A fusion module, used for fusing the mapping features of the plurality of sub-blocks to obtain a fusion feature; The prediction module is used to use the prediction layer in the recognition model to perform target prediction on the fusion features to obtain the recognition result of the target.

17. A computer device, It is characterized in that The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented, or the method according to claim 8 is implemented.

18. A non-transitory computer-readable storage medium having stored thereon a computer program, It is characterized in that When the program is executed by a processor, the program implements the method according to any one of claims 1 to 7, or the method according to claim 8.

19. A computer program product, It is characterized in that When the instructions in the computer program product are executed by a processor, the method according to any one of claims 1 to 7 is executed, or the method according to claim 8 is executed.

Citation Information

Patent Citations

  • Model training method, device, image recognition method, device, equipment and medium

    CN113902007A

  • Image feature extraction method and device, and storage medium

    WO2016192213A1