Image encoder determination method, and related device
In product quality inspection scenarios with fewer defective products, self-supervised training of the image encoder by using the correlation of multiple scanned images of the same object under different lighting parameters, solving the problem of high time and cost of labeling defect samples, and achieving the effect of reducing the training cost of image defect detection models.
Patent Information
- Application Number
- PCT/CN2024/117049
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-07
- Filing Date
- 2024-09-05
- Publication Date
- 2025-05-22
AI Technical Summary
In the product quality inspection scenario where there are fewer defective products, the time and cost of labeling defect samples are high, resulting in an increase in the training cost of image defect detection models.
By utilizing the correlation between multiple scanned images of the same object under different lighting parameters, the reconstruction model of the image encoder and the reconstruction network are adopted to perform self-supervised training to optimize the image encoder and make its feature expression capabilities stronger.
It reduces the number of labeling of defect samples, saves labeling time and cost, and reduces the training cost of image defect detection models, making this method suitable for product quality inspection scenarios with fewer defective products.
Smart Images

Figure CN2024117049_22052025_PF_FP_ABST
Abstract
Description
A method for determining an image encoder and related devices
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on October 7, 2023, application number 202311285085.8, and application name “A method for determining an image encoder and related devices”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a determination technology of an image encoder. Background Art
[0003] With the rapid development of artificial intelligence, product quality inspection means: first, scanning and imaging the product to be inspected to obtain a scanned image of the product to be inspected, and then using visual algorithms to perform automated defect detection on the scanned image of the product to be inspected.
[0004] In related technologies, labeling personnel usually label a certain number of defect samples from multiple scanned images corresponding to multiple inspected products to train an initial detection model to obtain an image defect detection model, and then use the image defect detection model to perform defect detection on the scanned images of the products to be inspected.
[0005] However, the above method is difficult to label a certain number of defect samples from multiple scanned images corresponding to multiple quality-inspected products. It requires a lot of labeling time and cost, resulting in high training costs for image defect detection models, making it difficult to apply to product quality inspection scenarios with fewer defective products.
[0006] Summary of the Invention
[0007] In order to solve the above technical problems, the present application provides a method for determining an image encoder and related devices, which reduce the number of annotations of defective samples and save a lot of annotation time and annotation costs, so that a large number of normal samples can be combined for subsequent training to obtain an image defect detection model with stronger feature expression capabilities, thereby reducing the training cost of the image defect detection model and making it suitable for product quality inspection scenarios with fewer defective products.
[0008] The embodiments of this application disclose the following technical solutions:
[0009] In one aspect, an embodiment of the present application provides a method for determining an image encoder, the method being performed by a computer device, and the method comprising:
[0010] performing image encoding on the first sample image using an image encoder in the initial reconstruction model to obtain first image block codes corresponding to a plurality of first image blocks in the first sample image, and obtaining second image block codes corresponding to a plurality of second image blocks in the second sample image, wherein the second image block codes corresponding to the plurality of second image blocks are obtained by performing image encoding on the plurality of second image blocks using a pre-trained encoder; the first sample image and the second sample image are multiple scanned images of a first object under different illumination parameters;
[0011] According to the first image block codes respectively corresponding to the multiple first image blocks, performing code prediction on the multiple second image blocks in the second sample image through the reconstruction network in the initial reconstruction model to obtain first prediction codes respectively corresponding to the multiple second image blocks;
[0012] Performing model training on the initial reconstruction model according to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and the loss function of the initial reconstruction model to obtain a first reconstruction model;
[0013] The image encoder in the first reconstruction model is determined as the image encoder in the initial detection model; the initial detection model is used to train the image defect detection model.
[0014] On the other hand, an embodiment of the present application provides a determination device for an image encoder, the device being deployed on a computer device, the device comprising: an encoding unit, a prediction unit, a training unit, and a determination unit;
[0015] The encoding unit is configured to perform image encoding on the first sample image using an image encoder in the initial reconstruction model to obtain first image block codes corresponding to a plurality of first image blocks in the first sample image, and to obtain second image block codes corresponding to a plurality of second image blocks in the second sample image, wherein the second image block codes corresponding to the plurality of second image blocks are obtained by performing image encoding on the plurality of second image blocks using a pre-trained encoder; the first sample image and the second sample image are multiple scanned images of the first object under different illumination parameters;
[0016] The prediction unit is configured to perform coding prediction on a plurality of second image blocks in the second sample image through a reconstruction network in the initial reconstruction model according to the first image block codes respectively corresponding to the plurality of first image blocks, to obtain first predicted codes respectively corresponding to the plurality of second image blocks;
[0017] The training unit is configured to perform model training on the initial reconstruction model according to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and the loss function of the initial reconstruction model to obtain a first reconstruction model;
[0018] The determining unit is used to determine the image encoder in the first reconstruction model as the image encoder in the initial detection model; the initial detection model is used to train the image defect detection model.
[0019] In another aspect, an embodiment of the present application provides a computer device, comprising a processor and a memory.
[0020] The memory is used to store a computer program and transmit the computer program to the processor;
[0021] The processor is configured to execute the method described in any one of the preceding aspects according to instructions in the computer program.
[0022] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which is used to store a computer program. When the computer program is run on a computer device, the computer device executes the method described in any one of the above aspects.
[0023] On the other hand, an embodiment of the present application provides a computer program product, including a computer program, which, when executed on a computer device, enables the computer device to execute the method described in any one of the aforementioned aspects.
[0024] As can be seen from the above technical solution, first, a first sample image is input into the image encoder of the initial reconstruction model for image encoding. First image block codes corresponding to multiple first image blocks in the first sample image are output, and second image block codes corresponding to multiple second image blocks in the second sample image are obtained. The second image block codes corresponding to the multiple second image blocks are obtained by performing image encoding on the multiple second image blocks by a pre-trained encoder. They are accurate encoding results of the second image blocks and can serve as supervisory signals for training the initial reconstruction model. The first sample image and the second sample image are multiple scanned images of a first object under different illumination parameters. Multiple scanned images of the same object under different illumination parameters are correlated. Therefore, the correlations can be mined by reconstructing the image block codes. To this end, the first image block codes corresponding to the multiple first image blocks are input into the reconstruction network of the initial reconstruction model. By mining the correlations between the multiple scanned images, the first image block codes corresponding to the multiple first image blocks are used to predict the codes for the multiple second image blocks in the second sample image, and the first predicted codes corresponding to the multiple second image blocks are output. The first predictive code is the result of predictive coding. Therefore, the initial reconstruction model can be trained using the first predictive codes corresponding to the plurality of second image blocks, the second image block codes corresponding to the plurality of second image blocks, and the loss function of the initial reconstruction model to obtain the first reconstruction model. This optimizes the image encoder in the initial reconstruction model, making the image encoder in the first reconstruction model more capable of expressing features.
[0025] The image encoder in the first reconstruction model is then used as the image encoder in the initial detection model for training the image defect detection model. This approach uses the image encoder in the first reconstruction model, obtained through self-supervised training, as the image encoder in the initial detection model, reducing the number of annotations required for defect samples. Subsequently, by combining a large number of scanned images of normal objects as normal samples, the initial detection model can be trained to produce an image defect detection model with enhanced feature expression capabilities.
[0026] Based on this, this method uses the characteristic that multiple scanned images of the same object under different lighting parameters are correlated, reconstructs image block encoding to mine the correlation, and optimizes the image encoder in the reconstruction model, so that the image encoder has stronger feature expression capabilities; the optimized image encoder is used in the detection model to reduce the number of labeled defect samples, saving a lot of labeling time and labeling costs, so that it can be combined with a large number of normal samples for subsequent training to obtain an image defect detection model with stronger feature expression capabilities, reducing the training cost of the image defect detection model, and thus being suitable for product quality inspection scenarios with fewer defective products. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technical members in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0028] FIG1 is a schematic diagram of a system architecture of a method for determining an image encoder provided by an embodiment of the present application;
[0029] FIG2 is a flowchart of a method for determining an image encoder provided by an embodiment of the present application;
[0030] FIG3 is a schematic diagram of multiple image blocks in a scanned image of an object corresponding to multiple points under one lighting parameter provided by an embodiment of the present application;
[0031] FIG4 is a schematic diagram of performing image encoding on a scanned image of an object by using an image encoder in an initial reconstruction model, to obtain multiple image block encodings corresponding to multiple image blocks in the scanned image of the object, provided by an embodiment of the present application;
[0032] FIG5 is a structural diagram of an initial encoder provided in an embodiment of the present application;
[0033] FIG6 is a schematic diagram of a pre-trained encoder obtained based on training of an initial encoder and an initial decoder according to an embodiment of the present application;
[0034] FIG7 is a schematic diagram of a multi-stage cascade detector provided in an embodiment of the present application;
[0035] FIG8 is a schematic diagram of output data of defect detection on a scanned image of a product to be inspected using an image defect detection model according to an embodiment of the present application;
[0036] FIG9 is a structural diagram of a determination device of an image encoder provided in an embodiment of the present application;
[0037] FIG10 is a structural diagram of a server provided in an embodiment of the present application;
[0038] FIG11 is a structural diagram of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] The embodiments of the present application are described below with reference to the accompanying drawings.
[0040] Currently, intelligent product quality inspection is being achieved by using visual algorithms to automatically detect defects in scanned images of products undergoing quality inspection. Specifically, labelers first annotate a certain number of defect samples from multiple scanned images of inspected products to train an initial detection model, generating an image defect detection model. The image defect detection model then performs defect detection on the scanned images of the products undergoing quality inspection.
[0041] However, in product quality inspection scenarios with fewer defective products, such as industrial quality inspection scenarios, it is difficult for the above method to label a certain number of defect samples from multiple scanned images corresponding to multiple inspected products. It requires a lot of labeling time and labeling costs, resulting in high training costs for image defect detection models, and thus high quality inspection costs for intelligent product quality inspection.
[0042] An embodiment of the present application provides a method for determining an image encoder, which utilizes the characteristic that multiple scanned images of the same object under different lighting parameters are correlated, reconstructs image block encoding to explore the correlation, and optimizes the image encoder in the reconstruction model, so that the image encoder has stronger feature expression capabilities; the optimized image encoder is used in the detection model, reducing the number of labeled defect samples, saving a lot of labeling time and labeling costs, so that a large number of normal samples can be combined for subsequent training to obtain an image defect detection model with stronger feature expression capabilities, reducing the training cost of the image defect detection model, and thus being suitable for product quality inspection scenarios with fewer defective products.
[0043] Next, the system architecture of the method for determining an image encoder will be introduced. Referring to FIG1 , FIG1 is a schematic diagram of the system architecture of a method for determining an image encoder provided in an embodiment of the present application, wherein the system architecture includes a server 100 for executing the method for determining an image encoder.
[0044] The server 100 performs image encoding on the first sample image through the image encoder in the initial reconstruction model to obtain multiple first image block codes corresponding to the multiple first image blocks in the first sample image.
[0045] As an example, the image encoder is Encoder1, the first sample image is x1, and the first image block is Patch1; the server 100 inputs x1 into Encoder1 in the initial reconstruction model for image encoding, and outputs the first image block codes corresponding to multiple Patches1 in x1. The first image block code can be represented by z1, that is, multiple z1s are obtained.
[0046] The server 100 uses the reconstruction network in the initial reconstruction model to predict the encoding of multiple second image blocks in the second sample image based on the multiple first image block encodings to obtain first predicted encodings corresponding to the multiple second image blocks respectively; the first sample image and the second sample image are multiple scanned images of the first object under different lighting parameters.
[0047] As an example, the second sample image is x2, and the second image block is Patch2. Based on the above example, the server 100 inputs multiple z1 into the reconstruction network in the initial reconstruction model, performs coding prediction on multiple Patch2 in x2, and outputs the first prediction codes corresponding to the multiple Patch2s. The first prediction codes can be represented by z1. p Indicates that multiple z p .
[0048] The server 100 performs model training on the initial reconstruction model according to the first prediction codes corresponding to the multiple second image blocks, the second image block codes corresponding to the multiple second image blocks, and the loss function of the initial reconstruction model to obtain a first reconstruction model; the multiple second image block codes are obtained by performing image encoding on the multiple second image blocks through a pre-trained encoder.
[0049] As an example, the pre-trained encoder is Encoder2. Based on the above example, the server 100 inputs x2 into Encoder2 for image encoding, and outputs the second image block encoding corresponding to multiple Patch2 in x2. The second image block encoding can be represented by z2, that is, multiple z2 are obtained; through multiple z p , multiple z2 and the loss function of the initial reconstruction model, and train the initial reconstruction model to obtain a first reconstruction model.
[0050] The server 100 determines the image encoder in the first reconstruction model as the image encoder in the initial detection model for training the image defect detection model; the initial detection model is used to train the image defect detection model.
[0051] As an example, based on the above example, the server 100 determines Encoder 1 in the first reconstruction model as Encoder 1 in the initial detection model for training the image defect detection model.
[0052] That is, based on the correlation between multiple scanned images of the same object under different illumination parameters, the image block codes are reconstructed to mine the correlation, and the initial reconstruction model is self-supervisedly trained to obtain a first reconstruction model. Specifically, this method uses multiple unlabeled scanned images of the same object under different illumination parameters to optimize the image encoder in the initial reconstruction model, making the image encoder in the first reconstruction model more capable of expressing features. Using the image encoder in the first reconstruction model obtained through self-supervised training as the image encoder in the initial detection model can reduce the number of annotations for defective samples. Subsequently, combining a large number of scanned images of normal objects as normal samples can train the initial detection model to obtain an image defect detection model with even stronger feature expression capabilities. Based on this, this method utilizes the correlation between multiple scanned images of the same object under different illumination parameters, reconstructs the image block codes to mine the correlation, and optimizes the image encoder in the reconstruction model to obtain even stronger feature expression capabilities. Using the optimized image encoder in the detection model reduces the number of annotations for defective samples, saving significant annotation time and cost. This allows subsequent training with a large number of normal samples to obtain an image defect detection model with even stronger feature expression capabilities, reducing the training cost of the image defect detection model, making it suitable for product quality inspection scenarios with fewer defective products.
[0053] It should be noted that in the embodiments of the present application, the computer device may be a server or a terminal, and the method provided in the embodiments of the present application may be executed by the terminal or the server alone, or by the terminal and the server in combination. The embodiment corresponding to FIG1 is mainly described by taking the server executing the method provided in the embodiments of the present application as an example.
[0054] Furthermore, when the method provided in the embodiments of the present application is executed solely by a terminal, its execution method is similar to the embodiment corresponding to FIG1 , primarily with the server being replaced by the terminal. Furthermore, when the method provided in the embodiments of the present application is executed in conjunction with a terminal and a server, steps that need to be displayed on the front-end interface can be executed by the terminal, while steps that require background computing and do not need to be displayed on the front-end interface can be executed by the server.
[0055] The terminal may be, but is not limited to, a smartphone, tablet computer, laptop computer, desktop computer, intelligent voice interaction device, vehicle-mounted terminal, or aircraft. The server may be, but is not limited to, an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. The terminal and server may be connected directly or indirectly via wired or wireless communication, which is not limited in this application. For example, the terminal and server may be connected via a network, which may be a wired or wireless network.
[0056] Among them, the embodiment of the present application can automatically determine the image encoder through artificial intelligence technology.
[0057] In addition, the embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, audio and video, assisted driving, etc.
[0058] Next, the method for determining an image encoder provided by an embodiment of the present application will be described in detail with reference to the accompanying drawings, using the example of a server executing the method provided by an embodiment of the present application. Referring to FIG. 2 , FIG. 2 is a flow chart of a method for determining an image encoder provided by an embodiment of the present application, the method comprising:
[0059] S201: Perform image encoding on the first sample image through the image encoder in the initial reconstruction model to obtain first image block codes corresponding to the multiple first image blocks in the first sample image, and obtain second image block codes corresponding to the multiple second image blocks in the second sample image, where the second image block codes corresponding to the multiple second image blocks are obtained by performing image encoding on the multiple second image blocks by a pre-trained encoder; the first sample image and the second sample image are multiple scanned images of the first object under different lighting parameters.
[0060] In related technologies, labelers label a certain number of defect samples from multiple scanned images corresponding to multiple inspected products, train an initial detection model to obtain an image defect detection model, and then use the image defect detection model to perform defect detection on the scanned images of the products to be inspected. However, in product quality inspection scenarios with fewer defective products, such as industrial quality inspection scenarios, this method is difficult to label a certain number of defect samples from multiple scanned images corresponding to multiple inspected products. This requires a lot of labeling time and cost, resulting in high training costs for the image defect detection model, and thus high quality inspection costs for intelligent product quality inspection.
[0061] Therefore, in an embodiment of the present application, considering that in a scanning imaging scenario, different illumination parameters are usually configured for shooting quality inspection products, any one of the different illumination parameters includes dozens of different points, and the same point has a certain correlation on multiple scanned images of the same quality inspection product under different illumination parameters; based on this, in order to solve the above technical problems, first, a reconstruction model including an image encoder and a reconstruction network can be constructed. For the same quality inspection product, the scanned images under one or more illumination parameters are first encoded by the image encoder into image block encodings of multiple image blocks in the scanned image, and then the reconstruction network predicts the predicted encodings of multiple image blocks in the scanned image under another illumination parameter. The predicted encoding under another illumination parameter, the image block encoding under another illumination parameter and the loss function of the initial reconstruction model are used to self-supervise the reconstruction model to mine the correlation and optimize the image encoder in the reconstruction model, so that the image encoder has stronger feature expression capabilities; wherein, the image block encoding under another illumination parameter is obtained by encoding the scanned image under another illumination parameter through a pre-trained encoder. Then, the optimized image encoder is used in the detection model to reduce the number of labeled defect samples, saving a lot of labeling time and labeling costs, so that a large number of normal samples can be combined for subsequent training to obtain an image defect detection model with stronger feature expression capabilities, reducing the training cost of the image defect detection model, and making it suitable for product quality inspection scenarios with fewer defective products.
[0062] Scanning imaging can include two scanning methods: area scanning and line scanning. Different scanning methods produce different scanned images. Area scanning can refer to scanning using an area scan camera to obtain a corresponding scanned image. In this case, the scanned image can be called an area scan image. That is, the first sample image and the second sample image can be multiple area scan images of the first object under different lighting parameters. Line scanning can refer to scanning using a line scan camera to obtain a corresponding scanned image. In this case, the scanned image can be called a line scan image. In other words, the first sample image and the second sample image can be multiple line scan images of the first object under different lighting parameters.
[0063] It should be noted that different lighting parameters can be formed in a variety of ways. In one possible implementation, different lighting parameters can be achieved by different light source hardware. For example, an object is illuminated under different light source hardware, thereby obtaining scanned images under different lighting parameters. In another possible implementation, different lighting parameters can be formed by the same light source hardware under different lighting modes. For example, a light source hardware has multiple lighting modes, and each lighting mode corresponds to different lighting parameters. Thus, by adjusting different lighting modes, scanned images under different lighting parameters can be obtained.
[0064] Based on the above description, first, a reconstruction model including an image encoder and a reconstruction network is constructed as an initial reconstruction model. On the basis of the quality inspection product belonging to the object, the scanned image of the first object under one or more lighting parameters is used as the first sample image; the first sample image is input into the image encoder in the initial reconstruction model for image encoding, and the first image block codes corresponding to the multiple first image blocks in the first sample image are output.
[0065] Reconstruction is a technique that processes, calculates, and restores a 2D image of an object to 3D data, ultimately reconstructing a realistic 3D model of the object on a computer. Therefore, in embodiments of the present application, a reconstruction model may refer to a neural network model used to reconstruct 2D graphics, including, for example, an initial reconstruction model to be trained and a trained first reconstruction model.
[0066] The initial reconstruction model may include an image encoder and a reconstruction network. The image encoder is used to encode each image block in the sample image input to the initial reconstruction model to obtain a corresponding image block code. The reconstruction network is used to explore the correlation between multiple scanned images under different lighting parameters, so that the first predicted codes corresponding to the multiple second image blocks in the second sample image can be predicted based on the first image block codes corresponding to the multiple first image blocks in the first sample image.
[0067] It should be noted that the embodiments of the present application do not limit the network structure of the image encoder and the reconstruction network. For example, the image encoder may include a convolution layer, a pooling layer, and a fully connected layer. Of course, the image encoder may also be similar to the network structure of the subsequent initial encoder, and the embodiments of the present application do not limit this. The reconstruction network may include, for example, a deconvolution layer, an upsampling layer, and a fully connected layer.
[0068] It is understood that in the embodiments of the present application, the first object can be various objects. Since the image defect detection model trained in the embodiments of the present application is used to detect defects in quality inspection products, and quality inspection products are quality inspection objects, the first object can be the first quality inspection object to better adapt to defect detection scenarios.
[0069] In practical applications, the first sample image is first divided into a plurality of first image blocks, and then the plurality of first image blocks are image encoded by an image encoder to obtain first image block codes corresponding to the plurality of first image blocks.
[0070] Image coding refers to mapping an image to a low-dimensional representation, which can be a vector or matrix. This low-dimensional representation typically offers better interpretability and stronger expressiveness. Image coding is a broader concept that encompasses the entire process of converting an image into a lower-dimensional representation. First-block coding refers to a low-dimensional representation of the image block features of the first image block, enabling better interpretation and expression of the image block features of the first image block.
[0071] The above S201 encodes the scanned image of the first object under one or more illumination parameters, that is, the first sample image, into first image block codes corresponding to multiple first image blocks in the first sample image through an image encoder, thereby providing image block coding data for subsequent reconstruction of the image block codes to explore the correlation between the multiple scanned images of the first object under different illumination parameters.
[0072] As an example of the above S201, the image encoder is Encoder1, the first sample image is x1, and the first image block is Patch1; x1 is input into Encoder1 in the initial reconstruction model for image encoding, and the first image block encodings corresponding to multiple Patches1 in x1 are output, that is, multiple z1s are output.
[0073] Refer to Figure 3, which is a schematic diagram of multiple image blocks in a scanned image of an object corresponding to multiple points under one lighting parameter provided by an embodiment of the present application; wherein, one lighting parameter includes 6×6 points, that is, 36 points, and correspondingly, each scanned image of each object needs to be divided into 36 image blocks, that is, P001, P002,..., P036; based on this, multiple Patches 1 can be 36 Patches 1.
[0074] Refer to Figure 4, which is a schematic diagram of an embodiment of the present application, in which an image encoder in an initial reconstruction model is used to perform image encoding on a scanned image of an object, thereby obtaining image block encodings corresponding to multiple image blocks in the scanned image. For each scanned image of each object, the scanned image is divided into multiple image blocks, and based on the image block embedding vectors and position embedding vectors corresponding to the multiple image blocks, the image block embedding vectors and position embedding vectors are input into the image encoder in the initial reconstruction model for image encoding, thereby obtaining image block encodings corresponding to the multiple image blocks. Based on this, the image block embedding vectors and position embedding vectors corresponding to the multiple Patches 1 are input into Encoder 1 in the initial reconstruction model for image encoding, and z1 corresponding to the multiple Patches 1 is output.
[0075] The model structure of the pre-trained encoder may be the same as the model structure of the image encoder in the initial reconstruction model.
[0076] In practical applications, the second sample image is first divided into a plurality of second image blocks, and then image encoding is performed on the plurality of second image blocks using a pre-trained encoder to obtain second image block codes corresponding to the plurality of second image blocks. The second image block code is a low-dimensional representation of the image block features of the second image blocks, which is used to better interpret and express the image block features of the second image blocks.
[0077] It should be noted that, in the embodiment of the present application, obtaining the second image block codes corresponding to the plurality of second image blocks in the second sample image can include multiple methods. One method is to obtain the second image block codes corresponding to the plurality of second image blocks in advance using a pre-trained encoder and store them. In this way, when executing S201, the second image block codes corresponding to the plurality of second image blocks can be directly read from the storage space, thereby improving acquisition efficiency.
[0078] Another method may be to obtain second image block codes corresponding to multiple second image blocks respectively through a pre-trained encoder when executing S201, thereby performing image encoding in real time according to current actual needs, and obtaining second image block codes corresponding to multiple second image blocks respectively that better meet the needs.
[0079] S202: According to the first image block codes respectively corresponding to the plurality of first image blocks, a reconstruction network in the initial reconstruction model is used to perform coding prediction on the plurality of second image blocks in the second sample image to obtain first prediction codes respectively corresponding to the plurality of second image blocks.
[0080] In the embodiment of the present application, after executing S201 to obtain first image block codes corresponding to the plurality of first image blocks, considering the correlation between the plurality of scanned images of the first object under different illumination parameters, in order to exploit the correlation to optimize the image encoder in the initial reconstruction model and enhance the feature expression capability of the image encoder in the first reconstruction model, the image block codes may be reconstructed. Specifically, based on a scanned image of the first object under another illumination parameter as a second sample image, the plurality of first image block codes are input into a reconstruction network in the initial reconstruction model, code prediction is performed on the plurality of second image blocks in the second sample image, and first predicted codes corresponding to the plurality of second image blocks are output.
[0081] Predictive coding refers to predicting a low-dimensional representation of an image through an image reconstruction mechanism. This involves predicting the information currently to be encoded based on information that has already been encoded. In this embodiment of the present application, a low-dimensional representation of a second image block is predicted based on the encoding of a first image block. The first predictive coding refers to the predicted low-dimensional representation of multiple second image blocks in the second sample image.
[0082] When the scanned image of the first object under one or more illumination parameters and the scanned image of the first object under another illumination parameter are three scanned images of the first object under three illumination parameters, the two scanned images of the first object under any two illumination parameters among the three scanned images of the first object under the three illumination parameters are used as first sample images, and the scanned image of the first object under another illumination parameter among the three scanned images of the first object under the three illumination parameters is used as a second sample image; wherein the two scanned images of the first object under any two illumination parameters can be repeatedly sampled to implement a three-channel input of the first sample image into an image encoder in an initial reconstruction model.
[0083] Based on the above S201, the above S202 encodes multiple first image blocks under one or more illumination parameters and predicts the first predicted codes corresponding to multiple second image blocks under another illumination parameter through a reconstruction network, thereby realizing reconstructed image block encoding, and providing predicted coding data for the subsequent mining of the correlation between multiple scanned images of the first object under different illumination parameters, so as to obtain the first reconstruction model through self-supervised training of the initial reconstruction model.
[0084] As an example of the above S202, the second sample image is x2, and the second image block is Patch2. Based on the example of the above S201, multiple first image block codes z1 are input into the reconstruction network in the initial reconstruction model, and multiple Patch2 in x2 are predicted, and the first prediction codes corresponding to the multiple Patch2 are output, that is, multiple z p .
[0085] S203: Perform model training on the initial reconstruction model according to the first prediction codes respectively corresponding to the multiple second image blocks, the second image block codes respectively corresponding to the multiple second image blocks, and the loss function of the initial reconstruction model to obtain a first reconstruction model.
[0086] In an embodiment of the present application, after executing S202 to predict and obtain first prediction codes corresponding to the plurality of second image blocks, in order to explore the correlation between the plurality of scanned images of the first object under different illumination parameters and optimize the image encoder in the initial reconstruction model so as to enhance the feature expression capability of the image encoder in the first reconstruction model, the initial reconstruction model may be trained in a self-supervised manner. That is, after inputting the second sample image into a pre-trained encoder for image encoding and outputting second image block codes corresponding to the plurality of second image blocks in the second sample image, the initial reconstruction model is trained using the plurality of first prediction codes, the plurality of second image block codes, and the loss function of the initial reconstruction model to obtain the first reconstruction model.
[0087] Among them, the loss function of the initial reconstruction model is used to measure the difference between each first prediction code and the corresponding second image block code; model training refers to parameter adjustment of the model parameters of the initial reconstruction model; the first reconstruction model refers to the initial reconstruction model after the model training is completed, and the end condition of the model training refers to the model training convergence of the initial reconstruction model or the model training times of the initial reconstruction model reach the maximum training times.
[0088] The above S203 mines the correlation between each first prediction code and the corresponding second image block code through the loss function of the initial reconstruction model, so as to mine the correlation between multiple scanned images of the first object under different lighting parameters, realize self-supervised training of the initial reconstruction model, optimize the image encoder in the initial reconstruction model, and make the feature expression ability of the image encoder in the first reconstruction model stronger, so as to provide an image encoder for the subsequent construction of the initial detection model for training the image defect detection model.
[0089] As an example of the above S203, the pre-trained encoder is Encoder2. Based on the example of the above S202, multiple second image blocks Patch2 in the second sample image x2 are input to Encoder2 for image encoding, and multiple second image block codes corresponding to each of Patch2 are output, that is, multiple z2 are output. p , multiple z2 and the loss function of the initial reconstruction model, and train the initial reconstruction model to obtain a first reconstruction model.
[0090] S204: Determine the image encoder in the first reconstruction model as the image encoder in the initial detection model for training the image defect detection model; the initial detection model is used to train the image defect detection model.
[0091] In an embodiment of the present application, after executing S203 training to obtain the first reconstruction model, considering that the image encoder in the first reconstruction model has stronger feature expression capabilities, the initial detection model is used to train the image defect detection model, and the image encoder in the first reconstruction model is determined as the image encoder in the initial detection model. This can reduce the number of annotations of defective samples and save a lot of annotation time and annotation costs, so that a large number of scanned images of normal quality inspection objects can be combined as normal samples to train the initial detection model to obtain an image defect detection model with stronger feature expression capabilities, thereby reducing the training cost of the image defect detection model and making it suitable for product quality inspection scenarios with fewer defective products.
[0092] The above S204 uses the image encoder in the first reconstruction model obtained by self-supervised training as the image encoder in the initial detection model, which can reduce the number of annotations of defect samples. Subsequently, by combining a large number of scanned images of normal objects as normal samples, the initial detection model can be trained to obtain an image defect detection model with stronger feature expression capabilities.
[0093] As an example of the above S204 , based on the example of the above S203 , the image encoder Encoder 1 in the first reconstruction model is determined as the Encoder 1 in the initial detection model for training the image defect detection model.
[0094] As can be seen from the above technical solution, first, a first sample image is input into the image encoder of the initial reconstruction model for image encoding. First image block codes corresponding to multiple first image blocks in the first sample image are output, and second image block codes corresponding to multiple second image blocks in the second sample image are obtained. The second image block codes corresponding to the multiple second image blocks are obtained by performing image encoding on the multiple second image blocks by a pre-trained encoder. They are accurate encoding results of the second image blocks and can serve as supervisory signals for training the initial reconstruction model. The first sample image and the second sample image are multiple scanned images of a first object under different illumination parameters. Multiple scanned images of the same object under different illumination parameters are correlated. Therefore, the correlations can be mined by reconstructing the image block codes. To this end, the first image block codes corresponding to the multiple first image blocks are input into the reconstruction network of the initial reconstruction model. By mining the correlations between the multiple scanned images, the first image block codes corresponding to the multiple first image blocks are used to predict the codes for the multiple second image blocks in the second sample image, and the first predicted codes corresponding to the multiple second image blocks are output. The first predictive code is the result of predictive coding. Therefore, the initial reconstruction model can be trained using the first predictive codes corresponding to the plurality of second image blocks, the second image block codes corresponding to the plurality of second image blocks, and the loss function of the initial reconstruction model to obtain the first reconstruction model. This optimizes the image encoder in the initial reconstruction model, making the image encoder in the first reconstruction model more capable of expressing features.
[0095] The image encoder in the first reconstruction model is then used as the image encoder in the initial detection model for training the image defect detection model. This approach uses the image encoder in the first reconstruction model, obtained through self-supervised training, as the image encoder in the initial detection model, reducing the number of annotations required for defect samples. Subsequently, by combining a large number of scanned images of normal objects as normal samples, the initial detection model can be trained to produce an image defect detection model with enhanced feature expression capabilities.
[0096] Based on this, this method uses the characteristic that multiple scanned images of the same object under different lighting parameters are correlated, reconstructs image block encoding to mine the correlation, and optimizes the image encoder in the reconstruction model, so that the image encoder has stronger feature expression capabilities; the optimized image encoder is used in the detection model to reduce the number of labeled defect samples, saving a lot of labeling time and labeling costs, so that it can be combined with a large number of normal samples for subsequent training to obtain an image defect detection model with stronger feature expression capabilities, reducing the training cost of the image defect detection model, and thus being suitable for product quality inspection scenarios with fewer defective products.
[0097] In the above embodiment, when S203 is specifically implemented, the loss function of the initial reconstruction model can be a cross-entropy loss function; based on this, first, multiple first prediction codes and multiple second image block codes are substituted into the cross-entropy loss function, and the first prediction probability of each first prediction code being the corresponding second image block code is calculated; then, considering that the training direction of the initial reconstruction model is to make the multiple first prediction codes close to the corresponding multiple second image block codes, the initial reconstruction model is trained to obtain the first reconstruction model by maximizing the multiple first prediction probabilities. Therefore, the present application provides a possible implementation method, in which the loss function of the initial reconstruction model is a cross-entropy loss function; S203 includes the following S2031-S2032 (not shown in the figure):
[0098] S2031: Determine a first prediction probability that each first prediction code is the corresponding second image block code based on the first prediction codes respectively corresponding to the multiple second image blocks, the second image block codes respectively corresponding to the multiple second image blocks, and a cross-entropy loss function.
[0099] S2032: Performing model training on the initial reconstruction model with the goal of maximizing the plurality of first prediction probabilities to obtain a first reconstruction model.
[0100] S2031-S2032 accurately explores the correlations between multiple scanned images of the first object under different illumination parameters by calculating the first prediction probability that the first prediction code is the corresponding second image block code. By maximizing the multiple first prediction probabilities, the initial reconstruction model is trained in a training direction that ensures that the multiple first prediction codes are close to the corresponding multiple second image block codes, accurately implementing self-supervised training of the initial reconstruction model and accurately optimizing the image encoder in the initial reconstruction model, thereby enhancing the feature expression capability of the image encoder in the first reconstruction model.
[0101] As an example of the above S2031-S2032, based on the above S203 example, multiple first prediction codes z p And multiple second image block codes z2 are substituted into the cross entropy loss function to calculate each z pThe first predicted probability of z2 is represented by p1, that is, multiple p1s are obtained. By maximizing the multiple p1s, the initial reconstruction model is trained to obtain a first reconstruction model.
[0102] In the above embodiment, the pre-trained encoder is pre-trained. Considering that the scanned image of the object is obtained by clearly photographing the object, that is, the scanned image of the object has a high resolution and has a certain amount of pixel redundancy. In order to reduce the pixel redundancy of the scanned image, the pre-trained encoder can be trained by mapping the image blocks in the scanned image to discrete codes and reconstructing the scanned image based on the discrete codes. Based on this, the steps for obtaining the pre-trained encoder include: first, using the scanned image of the second object as the third sample image, inputting the third sample image into the initial encoder for image encoding, and outputting third image block features corresponding to multiple third image blocks in the third sample image; second, discretizing the multiple third image block features into multiple third image block codes using multiple preset discrete codes, that is, the multiple third image block codes belong to multiple preset discrete codes; then, inputting the third image block codes corresponding to the multiple third image block features into the initial decoder to reconstruct the third sample image, and outputting a reconstructed sample image of the third sample image; finally, using the reconstructed sample image, the third sample image, and the loss function between the initial encoder and the initial decoder, model training is performed on the initial encoder and the multiple preset discrete codes to obtain the pre-trained encoder. Therefore, the present application provides a possible implementation method, in which the steps of obtaining the pre-trained encoder include the following S1-S4 (not shown in the figure):
[0103] S1: performing image encoding on a third sample image through an initial encoder to obtain third image block features respectively corresponding to a plurality of third image blocks in the third sample image; the third sample image is a scanned image of the second object.
[0104] See Figure 5, which is a structural diagram of an initial encoder provided in an embodiment of the present application; wherein the initial encoder is an encoder based on ViT as the basic network, and the ViT is composed of 6 transformer encoders, each transformer encoder including a normalization layer (Norm layer), a multi-head attention layer (Muliti-Head Attention layer), a normalization layer (Norm layer) and a multilayer perceptron (Multilayer Perceptron, MLP).
[0105] The second object can be various objects. Since the image defect detection model trained in the embodiment of the present application is used to perform defect detection on quality inspection products, which are quality inspection objects, the second object can be a second quality inspection object in order to be more suitable for defect detection scenarios.
[0106] Similar to the first sample image and the second sample image, the third sample image may be an area scan image of the second object or a line scan image of the second object.
[0107] S2: Determine third image block codes corresponding to multiple third image block features respectively from multiple preset discrete codes; the third image block codes corresponding to the multiple third image block features respectively belong to the multiple preset discrete codes.
[0108] Each preset discrete code is a code consisting of multiple integer values, and the dimension of each preset discrete code is the same as the dimension of the initial encoder output data, that is, the dimension of each preset discrete code is the same as the dimension of each third image block feature.
[0109] S3: encoding the third image blocks corresponding to the features of the plurality of third image blocks respectively, and reconstructing the third sample image through an initial decoder to obtain a reconstructed sample image.
[0110] S4: Based on the reconstructed sample image, the third sample image, and the loss function of the initial encoder and the initial decoder, model training is performed on the initial encoder and multiple preset discrete codes to obtain a pre-trained encoder.
[0111] The above-mentioned S1-S4 uses multiple preset discrete codes to map multiple third image blocks in the third sample image through the initial encoder to multiple third image block codes belonging to the multiple preset discrete codes, and reconstructs the third sample image through the initial decoder based on the multiple third image block codes. This training method optimizes the initial encoder and the multiple preset discrete codes, improves the training speed and training effect of the initial encoder, enables the pre-trained encoder to reduce pixel redundancy in the scanned image, and thus improves the training speed and training effect of the initial reconstruction model. In addition, it can avoid the overfitting problem of the first reconstruction model obtained by training the initial reconstruction model.
[0112] As an example of the above S1-S4, see Figure 6, which is a schematic diagram of a pre-trained encoder obtained based on the initial encoder and initial decoder training provided in an embodiment of the present application, the initial encoder is encoder, the initial decoder is decoder, the third sample image is x3, the third image block is Patch3, and multiple preset discrete codes are E = [e1, e2, ..., e K ], based on the above S201 example, x3 is input into the encoder for image encoding, and multiple third image block features corresponding to multiple Patch3 in x3 are output as multiple v3; through E=[e1, e2, ..., e K ] Discretize multiple v3 into multiple z3, that is, multiple z3 belongs to E = [e1, e2, ..., eK ]; multiple z3 are input into the decoder to reconstruct the image of x3, and the reconstructed sample image of x3 is output as x3'; through x3', x3, and the loss function of encoder and decoder, the encoder and E = [e1, e2, ..., e K ] Perform model training to obtain the pre-trained encoder Encoder2, that is, Encoder2 is the encoder after training.
[0113] Since E=[e1,e2,…,e K ] The process of discretizing multiple v3 into multiple z3 does not support back propagation. When training the model in back propagation, stop training by E=[e1, e2, ..., e K ] The model parameters of the process of discretizing multiple v3 into multiple z3 are trained directly through x3', x3, and the loss function of encoder and decoder, and the model parameters of the process of inputting x3 into the encoder for image encoding and outputting multiple v3 corresponding to multiple patches in x3, as well as E = [e1, e2, ..., e K ].
[0114] Among them, when S2 is specifically implemented, among multiple preset discrete codes, the third image block codes corresponding to multiple third image block features can be determined by nearest neighbor search; specifically, for each third image block feature, the similarity between the third image block feature and each preset discrete code is first calculated to obtain multiple similarities between the third image block feature and the multiple preset discrete codes; then, the preset discrete code corresponding to the maximum similarity among the multiple similarities is used as the third image block code corresponding to the third image block feature. Therefore, the present application provides a possible implementation method, S2 includes the following S21-S22 (not shown in the figure):
[0115] S21: For each third image block feature, perform similarity calculation based on the third image block feature and a plurality of preset discrete codes to obtain a similarity between the third image block feature and each of the plurality of preset discrete codes.
[0116] S22: Determine the preset discrete code corresponding to the maximum similarity as the third image block code corresponding to the third image block feature.
[0117] The above S21-S22, based on multiple preset discrete codes, discretizes multiple third image block features into multiple third image block codes through a nearest neighbor search method, which can more accurately map multiple third image blocks in the third sample image to multiple third image block codes belonging to multiple preset discrete codes, so as to facilitate the subsequent optimization of the initial encoder and multiple preset discrete codes, so that the pre-trained encoder can reduce the pixel redundancy of the scanned image and provide more accurate image block coding data.
[0118] As an example of the above S21-S22, based on the above S1-S4 examples, for each third image block feature v3, calculate v3 and multiple preset discrete codes E = [e1, e2, ..., e K ] Each preset discrete code e i The similarity between v3 and e is a positive integer, i = 1, 2, ..., K, and the similarity between v3 and e is obtained. i Multiple similarities between them; then the e corresponding to the maximum similarity among the multiple similarities i , as the third image block code z3 corresponding to v3. Among them, v3 and e i The similarity between v3 and e i The distance between them is expressed as ||v3-e i ||2, then the e corresponding to the maximum similarity among multiple similarities i where i=argmin i ||v3-e i ||2.
[0119] Among them, when S4 is specifically implemented, the loss function of the initial encoder and the initial decoder can be a cross-entropy loss function; based on this, first, the reconstructed sample image and the third sample image are substituted into the cross-entropy loss function, and the second prediction probability that the reconstructed sample image is the third sample image is calculated; then, considering that the training direction of the initial encoder and the initial decoder is to make the reconstructed sample image close to the third sample image, by maximizing the second prediction probability, the initial encoder and multiple preset discrete codes are model trained to obtain a pre-trained encoder. Therefore, the present application provides a possible implementation method, in which the loss function of the initial encoder and the initial decoder is a cross-entropy loss function; S4 includes the following S41-S42 (not shown in the figure):
[0120] S41: Determine a second prediction probability that the reconstructed sample image is the third sample image according to the reconstructed sample image, the third sample image, and the cross entropy loss function.
[0121] S22: Perform model training on the initial encoder and multiple preset discrete codes with the goal of maximizing the second prediction probability to obtain a pre-trained encoder.
[0122] The above S41-S42 accurately mines the correlation between the multiple third image block codes belonging to multiple preset discrete codes obtained by mapping multiple third image blocks in the third sample image and the third sample image by calculating the second prediction probability that the reconstructed sample image is the third sample image; by maximizing the second prediction probability, the initial encoder and the multiple preset discrete codes are trained in the training direction of making the reconstructed sample image close to the third sample image, and the initial encoder and the multiple preset discrete codes are accurately optimized, so that the pre-trained encoder can reduce the pixel redundancy of the scanned image on the basis of accurate feature expression.
[0123] As an example of the above S41-S42, based on the above S1-S4 examples, the reconstructed sample image x3' and the third sample image x3 are substituted into the cross entropy loss function, and the second prediction probability of x3' being x3 is calculated as p2; by maximizing p2, the initial encoder encoder and multiple preset discrete encodings E = [e1, e2, ..., e K ] Perform model training to obtain the pre-trained encoder Encoder2.
[0124] In the above embodiment, corresponding to S1-S4, in the specific implementation of S201, in order to reduce pixel redundancy of the second sample image so as to reduce the prediction difficulty when the multiple first image block codes are subsequently executed in S202 and the first prediction codes corresponding to the multiple second image blocks are predicted by the reconstruction network, it is necessary to use the multiple trained preset discrete codes to map the multiple second image blocks in the second sample image to multiple second image block codes belonging to the multiple trained preset discrete codes through a pre-trained encoder, and then encode the multiple first image blocks in the first sample image to multiple first image block features as multiple first image block codes through the image encoder in the initial reconstruction model. Specifically, the first sample image is first input into the image encoder in the initial reconstruction model for image encoding, and the first image block features corresponding to the multiple first image blocks in the first sample image are output; then, the multiple second image blocks in the second sample image are input into the pre-trained encoder for image encoding, and the second image block features corresponding to the multiple second image blocks are output. The multiple second image block features are discretized into multiple second image block codes through the multiple trained preset discrete codes, that is, the multiple second image block codes belong to the multiple trained preset discrete codes. Therefore, the present application provides a possible implementation method, wherein the multiple first image block codes are multiple first image block features, and the multiple second image block codes are multiple pre-set discrete codes after training; S201 includes the following S2010 (not shown in the figure): using the image encoder in the initial reconstruction model, image encoding is performed on the first sample image to obtain multiple first image block features corresponding to the multiple first image blocks. Correspondingly, the steps of obtaining the multiple second image block codes include the following S5-S6 (not shown in the figure):
[0125] S5: performing image encoding on the plurality of second image blocks using a pre-trained encoder to obtain second image block features corresponding to the plurality of second image blocks respectively;
[0126] S6: Determine second image block codes corresponding to the plurality of second image block features respectively from the plurality of preset discrete codes after training.
[0127] The dimension of each preset discrete code after training is the same as the dimension of each first image block feature, and the dimension of each preset discrete code after training is the same as the dimension of each second image block feature.
[0128] The above-mentioned S5-S6 uses the trained multiple preset discrete codes to map the multiple second image blocks in the second sample image to multiple second image block codes belonging to the trained multiple preset discrete codes through the pre-trained encoder, which can reduce the pixel redundancy of the second sample image and provide image block coding data for subsequently reducing the prediction difficulty of predicting the multiple first prediction codes corresponding to the multiple second image blocks through the reconstruction network through the multiple first image block codes.
[0129] As an example of the above S2010, S5-S6, based on the above S201 and the above S1-S4 examples, the first sample image x1 is input into the Encoder1 in the initial reconstruction model for image encoding, and the multiple first image block features corresponding to the multiple first image blocks Patch1 in x1 are output as multiple v1; the multiple second image blocks Patch2 in the second sample image x2 are input into the pre-trained encoder for image encoding, and the multiple second image block features corresponding to the multiple Patch2 are output as multiple v2; through the trained E=[e1, e2,…, e K ] Discretize multiple v2 into multiple z2, that is, multiple z2 belong to the trained E = [e1, e2, ..., e K ]. The multiple v1s are the multiple first image block codes z1 corresponding to the multiple Patches1 in x1.
[0130] In the above embodiment, corresponding to the above S2010, S5-S6, when S203 is specifically implemented, the loss function of the initial reconstruction model can be a cross-entropy loss function; based on this, first, the first prediction codes corresponding to the multiple second image blocks and the second image block codes corresponding to the multiple second image blocks are respectively substituted into the cross-entropy loss function, and the first prediction probability of each first prediction code being the corresponding second image block code is calculated; then, on the basis that the multiple first image block codes belong to the multiple preset discrete codes after training, the training direction of the initial reconstruction model is further refined so that the multiple first prediction codes are close to the corresponding multiple second image block codes; therefore, for each first prediction code, it is first determined whether the first prediction code belongs to the multiple preset discrete codes after training. If so, the preset coefficient associated with the first prediction probability corresponding to the first prediction code is determined to be 1. Then, by maximizing the first prediction probability associated with the preset coefficient of 1, the initial reconstruction model is trained to obtain the first reconstruction model. That is, the present application provides a possible implementation method, in which the loss function of the initial reconstruction model is a cross-entropy loss function; S203 includes the following S2033-S2035 (not shown in the figure):
[0131] S2033: Determine a first prediction probability that each first prediction code is the corresponding second image block code based on the first prediction codes respectively corresponding to the multiple second image blocks, the second image block codes respectively corresponding to the multiple second image blocks, and a cross-entropy loss function.
[0132] S2034: For each first predicted code, if the first predicted code belongs to a plurality of preset discrete codes after training, determine that a preset coefficient associated with a first prediction probability corresponding to the first predicted code is 1.
[0133] S2035: With the goal of maximizing the first prediction probability associated with a preset coefficient of 1, the initial reconstruction model is trained to obtain a first reconstruction model.
[0134] The above-described S2033-S2035 accurately explores the correlation between multiple scanned images of the same object under different illumination parameters by calculating multiple first prediction probabilities that the multiple first prediction codes correspond to the corresponding multiple second image block codes. If the multiple first prediction codes belong to multiple pre-trained preset discrete codes, the initial reconstruction model is trained in a training direction that maximizes the first prediction probabilities corresponding to the multiple first prediction codes so that the multiple first prediction codes are close to the corresponding multiple second image block codes, thereby further accurately achieving self-supervised training of the initial reconstruction model, further optimizing the image encoder in the initial reconstruction model, and enhancing the feature expression capability of the image encoder in the first reconstruction model.
[0135] As an example of the above S2033-S2035, based on the above S203 and the above S2011-S2012 examples, multiple first prediction codes z p And multiple second image block codes z2 are substituted into the cross entropy loss function to calculate each z p is the first predicted probability of the corresponding z2, thereby obtaining multiple p1. p , first determine z p Whether it belongs to the multiple preset discrete codes E=[e1, e2, ..., e K ], if so, z p The preset coefficient associated with the corresponding p1 is determined to be 1. Then, by maximizing multiple p1s associated with the preset coefficient of 1, the initial reconstruction model is trained to obtain a first reconstruction model.
[0136] Among them, when S2034 is specifically implemented, in order to reduce the difficulty of determining whether the first prediction code belongs to the multiple preset discrete codes after training, corresponding preset discrete identifiers can be configured for the multiple preset discrete codes after training; based on this, for each first prediction code, there is no need to determine whether the first prediction code itself belongs to the multiple preset discrete codes after training, but rather to determine whether the first prediction code corresponds to any preset discrete identifier among the multiple preset discrete identifiers. If so, it means that the first prediction code belongs to the multiple preset discrete codes after training, and the preset coefficient associated with the first prediction probability corresponding to the first prediction code can be determined to be 1. Therefore, the present application provides a possible implementation method, S2034 includes the following S7-S8 (not shown in the figure):
[0137] S7: Obtain preset discrete identifiers corresponding to the multiple preset discrete codes after training.
[0138] S8: For each first prediction code, if the first prediction code corresponds to any preset discrete identifier among a plurality of preset discrete identifiers, determine that a preset coefficient associated with a first prediction probability corresponding to the first prediction code is 1.
[0139] The above S7-S8 judge whether the first predicted code corresponds to any of the multiple preset discrete identifiers based on the multiple preset discrete identifiers corresponding to the multiple preset discrete coding configurations after training, instead of judging whether the first predicted code belongs to the multiple preset discrete codes after training. The judgment operation is simpler and more convenient, thereby reducing the difficulty of judgment and improving the training speed of the initial reconstruction model.
[0140] As an example of the above S7-S8, based on the above S2034 example, multiple preset discrete codes E=[e1, e2, ..., e K ] corresponds to multiple preset discrete identifiers C = [1, 2, ..., K], for each zp , first determine z p Does it correspond to any preset discrete identifier in C = [1, 2, ..., K]? If so, set z p The corresponding preset coefficient associated with p1 is determined to be 1.
[0141] Based on the above description, the formal expression of the loss function of the initial reconstruction model can be as follows:
[0142] Where m×n represents the number of the first prediction codes, j is a positive integer, and y j represents the first prediction probability of the jth first prediction code being the corresponding jth second image block code, c j Indicates the coding identifier corresponding to the jth first prediction code, ∏(c j =C) represents the y corresponding to the jth first prediction code j When the jth first prediction code corresponds to any preset discrete identifier in C=[1, 2, ..., K], Π(c j =C) = 1; otherwise, ∏(c j =C)=0.
[0143] In addition, in an embodiment of the present application, in order to improve the training speed of the initial reconstruction model, after executing S201 to obtain multiple first image block codes corresponding to multiple first image blocks, it is not necessary to input all of the multiple first image block codes into the reconstruction network in the initial reconstruction model to perform coding prediction on multiple second image blocks. Instead, some of the first image block codes corresponding to some of the first image blocks can be input into the reconstruction network in the initial reconstruction model to perform coding prediction on some of the second image blocks, thereby reducing the number of coding predictions and improving the training speed of the initial reconstruction model.
[0144] In a specific implementation, a plurality of first image blocks are first randomly sampled to obtain a portion of the first image blocks, that is, a first number of first image blocks, where the first number is less than the number of blocks in the plurality of first image blocks. The first image block codes corresponding to the first number of first image blocks are then input into the reconstruction network in the initial reconstruction model, and coding prediction is performed on the second number of second image blocks, and the first prediction codes corresponding to the second number of second image blocks are output. The second number of second image blocks corresponds to the first number of first image blocks. Correspondingly, the initial reconstruction model is subsequently trained using the second number of first prediction codes, the second image block codes corresponding to the second number of second image blocks, and the loss function of the initial reconstruction model to obtain the first reconstruction model.
[0145] Therefore, the present application provides a possible implementation method, and the method also includes S9 (not shown in the figure): randomly sampling multiple first image blocks to obtain a first number of first image blocks, and the first number is less than the number of blocks of the multiple first image blocks. Correspondingly, S202 includes S2021 (not shown in the figure): through the reconstruction network in the initial reconstruction model, according to the first image block codes corresponding to the first number of first image blocks, the second number of second image blocks are encoded and predicted to obtain the first prediction codes corresponding to the second number of second image blocks. S203 includes S2036 (not shown in the figure): according to the first prediction codes corresponding to the second number of second image blocks, the second image block codes corresponding to the second number of second image blocks, and the loss function of the initial reconstruction model, the initial reconstruction model is trained to obtain a first reconstruction model.
[0146] As an example of the above S9, S2021 and S2036, the first number is s1, the second number is s2, and both s1 and s2 are positive integers. Based on the above S201 example, s1 is less than the number of blocks of the multiple first image blocks Patch1, s2 is less than the number of blocks of the multiple second image blocks Patch2, and s1 Patch1 corresponds to s2 Patch2. Randomly sample the multiple Patches1 to obtain s1 Patches1, input the s1 first image block codes z1 corresponding to the s1 Patches1 into the reconstruction network in the initial reconstruction model, perform code prediction on the s2 Patches2, and output the s2 first prediction codes z corresponding to the s2 Patches2. p . By s2z p , s2 second image block codes z2 corresponding to s2 Patches2, and the loss function of the initial reconstruction model, and the initial reconstruction model is trained to obtain the first reconstruction model.
[0147] Furthermore, in an embodiment of the present application, during the execution of steps S201-S203 above, the initial reconstruction model is used to encode multiple first image blocks in the first sample image into multiple first image block codes. First prediction codes corresponding to the multiple second image blocks are predicted from the multiple first image block codes to reconstruct the image block codes of the multiple second image blocks, thereby exploiting the correlation between the first sample image and the second sample image. Based on the self-supervised training of the initial reconstruction model to obtain a first reconstruction model, in order to fully exploit the correlation between the first sample image and the second sample image and further optimize the image encoder in the initial reconstruction model to enhance the feature expression capability of the image encoder in the first reconstruction model, the first reconstruction model may further be used to encode the multiple second image blocks into multiple fourth image block codes. Second prediction codes corresponding to the multiple first image blocks are predicted from the multiple fourth image block codes to reconstruct the image block codes of the multiple first image blocks. This fully exploits the correlation between the first sample image and the second sample image, and the self-supervised training of the first reconstruction model to obtain a second reconstruction model, thereby enhancing the feature expression capability of the image encoder in the second reconstruction model compared to the image encoder in the first reconstruction model. Correspondingly, the image encoder in the second reconstruction model is more suitable for constructing an initial detection model to train the image defect detection model than the image encoder in the first reconstruction model.
[0148] In a specific implementation, first, multiple second image blocks are input into the image encoder in the first reconstruction model for image encoding, and the fourth image block codes corresponding to the multiple second image blocks are output. Secondly, the fourth image block codes corresponding to the multiple second image blocks are input into the reconstruction network in the first reconstruction model, and the encoding prediction is performed on the multiple first image blocks, and the second prediction codes corresponding to the multiple first image blocks are output. Then, the first reconstruction model is trained to obtain a second reconstruction model through the multiple second prediction codes, the multiple fifth image block codes corresponding to the multiple second image blocks and the loss function of the first reconstruction model, wherein the multiple fifth image block codes are obtained by inputting the multiple second image blocks into the pre-trained encoder for image encoding. Finally, the image encoder in the second reconstruction model is determined as the image encoder in the initial detection model. Therefore, the present application provides a possible implementation method, and the method also includes the following S10-S12 (not shown in the figure):
[0149] S10: Perform image encoding on multiple second image blocks through the image encoder in the first reconstruction model to obtain fourth image block codes corresponding to the multiple second image blocks, and obtain fifth image block codes corresponding to the multiple first image blocks. The fifth image block codes corresponding to the multiple first image blocks are obtained by performing image encoding on the multiple first image blocks through a pre-trained encoder.
[0150] S11: According to the fourth image block codes respectively corresponding to the plurality of second image blocks, encoding prediction is performed on the plurality of first image blocks through a reconstruction network in a first reconstruction model to obtain second prediction codes respectively corresponding to the plurality of first image blocks.
[0151] S12: According to the second prediction codes respectively corresponding to the multiple first image blocks, the fifth image block codes respectively corresponding to the multiple first image blocks and the loss function of the first reconstruction model, the first reconstruction model is trained to obtain a second reconstruction model.
[0152] Correspondingly, S204 includes S2041 (not shown in the figure): determining the image encoder in the second reconstruction model as the image encoder in the initial detection model.
[0153] In summary, in the embodiments of the present application, the specific architecture of the initial detection model is not limited. The initial detection model can be a multi-stage cascade detector, or an end-to-end set prediction-based detector, etc. Referring to Figure 7, Figure 7 is a schematic diagram of a multi-stage cascade detector provided in an embodiment of the present application; wherein, B0 represents the detection frame of the first stage, H1 represents the detection network of the second stage, C1 represents the classification result of the second stage, and B1 represents the detection frame of the second stage; H2 represents the detection network of the third stage, C2 represents the classification result of the third stage, and B2 represents the detection frame of the third stage; H3 represents the detection network of the fourth stage, C3 represents the classification result of the fourth stage, and B3 represents the detection frame of the fourth stage.
[0154] Refer to Figure 8, which is a schematic diagram of output data for defect detection on a scanned image of a product to be inspected through an image defect detection model provided by an embodiment of the present application; wherein, the product to be inspected has defects, and the image encoder in the initial detection model is determined through the above embodiment. After the image defect detection model is trained, the scanned image of the product to be inspected is input into the image defect detection model for defect detection, and a defect detection box in the scanned image of the product to be inspected is output.
[0155] It should be noted that, based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods.
[0156] Based on the determination method of the image encoder provided in the embodiment corresponding to FIG2 , an embodiment of the present application further provides a determination device for an image encoder. Referring to FIG9 , FIG9 is a structural diagram of a determination device for an image encoder provided in an embodiment of the present application. The determination device 900 for an image encoder includes: an encoding unit 901, a prediction unit 902, a training unit 903, and a determination unit 904.
[0157] An encoding unit 901 is configured to perform image encoding on a first sample image using an image encoder in an initial reconstruction model to obtain first image block codes corresponding to a plurality of first image blocks in the first sample image, and to obtain second image block codes corresponding to a plurality of second image blocks in the second sample image, wherein the second image block codes corresponding to the plurality of second image blocks are obtained by performing image encoding on the plurality of second image blocks using a pre-trained encoder; the first sample image and the second sample image are multiple scanned images of a first object under different illumination parameters;
[0158] A prediction unit 902 is configured to perform coding prediction on a plurality of second image blocks in a second sample image using a reconstruction network in an initial reconstruction model according to the first image block codes respectively corresponding to the plurality of first image blocks, to obtain first predicted codes respectively corresponding to the plurality of second image blocks;
[0159] A training unit 903 is configured to perform model training on the initial reconstruction model based on the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and a loss function of the initial reconstruction model to obtain a first reconstruction model;
[0160] The determining unit 904 is configured to determine the image encoder in the first reconstruction model as the image encoder in the initial detection model used for training the image defect detection model; the initial detection model is used for training the image defect detection model.
[0161] In one possible implementation, the loss function of the initial reconstruction model is a cross entropy loss function; the training unit 903 is specifically configured to:
[0162] Determining a first prediction probability that each first prediction code is a corresponding second image block code according to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and a cross-entropy loss function;
[0163] The initial reconstruction model is trained with the goal of maximizing multiple first prediction probabilities to obtain a first reconstruction model.
[0164] In a possible implementation, the training unit 903 is further configured to:
[0165] performing image encoding on the third sample image by the initial encoder to obtain third image block features corresponding to a plurality of third image blocks in the third sample image; the third sample image is a scanned image of the second object;
[0166] Determining third image block codes corresponding to the plurality of third image block features respectively from a plurality of preset discrete codes; the third image block codes corresponding to the plurality of third image block features respectively belonging to the plurality of preset discrete codes;
[0167] Reconstructing the third sample image through an initial decoder according to the third image block codes corresponding to the plurality of third image block features to obtain a reconstructed sample image;
[0168] According to the reconstructed sample image, the third sample image, and the loss function of the initial encoder and the initial decoder, model training is performed on the initial encoder and multiple preset discrete codes to obtain a pre-trained encoder.
[0169] In a possible implementation, the determining unit 904 is further configured to:
[0170] For each third image block feature, performing similarity calculation based on the third image block feature and a plurality of preset discrete codes to obtain a similarity between the third image block feature and each of the plurality of preset discrete codes;
[0171] The preset discrete code corresponding to the maximum similarity is determined as the third image block code corresponding to the third image block feature.
[0172] In one possible implementation, the loss function of the initial encoder and the initial decoder is a cross entropy loss function; the training unit 903 is further specifically configured to:
[0173] Determining a second prediction probability that the reconstructed sample image is the third sample image based on the reconstructed sample image, the third sample image, and the cross entropy loss function;
[0174] The initial encoder and multiple preset discrete codes are trained with the goal of maximizing the second prediction probability to obtain a pre-trained encoder.
[0175] In one possible implementation, the first image block codes corresponding to the plurality of first image blocks are first image block features, and the second image block codes corresponding to the plurality of second image blocks are trained preset discrete codes. The encoding unit 901 is specifically configured to:
[0176] Performing image encoding on the first sample image by using an image encoder in the initial reconstruction model to obtain first image block features corresponding to the plurality of first image blocks;
[0177] The encoding unit 901 is further specifically configured to:
[0178] Performing image encoding on the plurality of second image blocks by using a pre-trained encoder to obtain second image block features corresponding to the plurality of second image blocks respectively;
[0179] Among the multiple preset discrete codes after training, second image block codes corresponding to the multiple second image block features are determined.
[0180] In one possible implementation, the loss function of the initial reconstruction model is a cross entropy loss function; the training unit 903 is specifically configured to:
[0181] Determining a first prediction probability that each first prediction code is a corresponding second image block code according to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and a cross-entropy loss function;
[0182] For each first predicted code, if the first predicted code belongs to a plurality of preset discrete codes after training, determining a preset coefficient associated with a first prediction probability corresponding to the first predicted code to be 1;
[0183] With the goal of maximizing the first prediction probability associated with a preset coefficient of 1, the initial reconstruction model is trained to obtain a first reconstruction model.
[0184] In a possible implementation, the determining unit 904 is further configured to:
[0185] Obtain preset discrete identifiers corresponding to the plurality of preset discrete codes after training;
[0186] For each first prediction code, if the first prediction code corresponds to any one of a plurality of preset discrete identifiers, a preset coefficient associated with a first prediction probability corresponding to the first prediction code is determined to be 1.
[0187] In a possible implementation, the apparatus further includes: a sampling unit;
[0188] a sampling unit, configured to randomly sample the plurality of first image blocks to obtain a first number of first image blocks; the first number being smaller than a number of the plurality of first image blocks;
[0189] The prediction unit 902 is specifically configured to:
[0190] According to the first image block codes respectively corresponding to the first number of first image blocks, a reconstruction network in the initial reconstruction model is used to perform code prediction on the second number of second image blocks to obtain first prediction codes respectively corresponding to the second number of second image blocks, where the second number of second image blocks corresponds to the first number of first image blocks;
[0191] The training unit 903 is specifically configured to:
[0192] The initial reconstruction model is trained according to the first prediction codes corresponding to the second number of second image blocks, the second image block codes corresponding to the second number of second image blocks, and the loss function of the initial reconstruction model to obtain a first reconstruction model.
[0193] In a possible implementation, the encoding unit 901 is further configured to:
[0194] performing image encoding on the plurality of second image blocks using an image encoder in the first reconstruction model to obtain fourth image block codes respectively corresponding to the plurality of second image blocks, and obtaining fifth image block codes respectively corresponding to the plurality of first image blocks, wherein the fifth image block codes respectively corresponding to the plurality of first image blocks are obtained by performing image encoding on the plurality of first image blocks using a pre-trained encoder;
[0195] The prediction unit 902 is further configured to:
[0196] According to the fourth image block codes respectively corresponding to the plurality of second image blocks, encoding prediction is performed on the plurality of first image blocks by a reconstruction network in the first reconstruction model to obtain second prediction codes respectively corresponding to the plurality of first image blocks;
[0197] The training unit 903 is further configured to:
[0198] training the first reconstruction model according to the second prediction codes respectively corresponding to the plurality of first image blocks, the fifth image block codes respectively corresponding to the plurality of first image blocks, and the loss function of the first reconstruction model to obtain a second reconstruction model;
[0199] The determining unit 904 is specifically configured to:
[0200] The image encoder in the second reconstruction model is determined as the image encoder in the initial detection model.
[0201] As can be seen from the above technical solution, first, a first sample image is input into the image encoder of the initial reconstruction model for image encoding. First image block codes corresponding to multiple first image blocks in the first sample image are output, and second image block codes corresponding to multiple second image blocks in the second sample image are obtained. The second image block codes corresponding to the multiple second image blocks are obtained by performing image encoding on the multiple second image blocks by a pre-trained encoder. They are accurate encoding results of the second image blocks and can serve as supervisory signals for training the initial reconstruction model. The first sample image and the second sample image are multiple scanned images of a first object under different illumination parameters. Multiple scanned images of the same object under different illumination parameters are correlated. Therefore, the correlations can be mined by reconstructing the image block codes. To this end, the first image block codes corresponding to the multiple first image blocks are input into the reconstruction network of the initial reconstruction model. By mining the correlations between the multiple scanned images, the first image block codes corresponding to the multiple first image blocks are used to predict the codes for the multiple second image blocks in the second sample image, and the first predicted codes corresponding to the multiple second image blocks are output. The first predictive code is the result of predictive coding. Therefore, the initial reconstruction model can be trained using the first predictive codes corresponding to the plurality of second image blocks, the second image block codes corresponding to the plurality of second image blocks, and the loss function of the initial reconstruction model to obtain the first reconstruction model. This optimizes the image encoder in the initial reconstruction model, making the image encoder in the first reconstruction model more capable of expressing features.
[0202] The image encoder in the first reconstruction model is then used as the image encoder in the initial detection model for training the image defect detection model. Using the image encoder in the first reconstruction model, obtained through self-supervised training, as the image encoder in the initial detection model reduces the number of annotations required for defect samples. Subsequently, by combining a large number of scanned images of normal objects as normal samples, the initial detection model can be trained to produce an image defect detection model with enhanced feature expression capabilities.
[0203] Based on this, the device uses the characteristic that multiple scanned images of the same object under different lighting parameters are correlated, reconstructs the image block code to explore the correlation, and optimizes the image encoder in the reconstruction model, so that the image encoder has stronger feature expression capabilities; the optimized image encoder is used in the detection model to reduce the number of labeled defect samples, saving a lot of labeling time and labeling costs, so that it can be combined with a large number of normal samples for subsequent training to obtain an image defect detection model with stronger feature expression capabilities, reducing the training cost of the image defect detection model, and thus being suitable for product quality inspection scenarios with fewer defective products.
[0204] The embodiment of the present application also provides a computer device, which can be a server. See Figure 10, which is a structural diagram of a server provided in an embodiment of the present application. The server 1000 may have relatively large differences due to different configurations or performances, and may include one or more processors, such as CPU 1022, and memory 1032, one or more storage media 1030 (such as one or more massive storage devices) storing application programs 1042 or data 1044. Among them, the memory 1032 and the storage medium 1030 can be temporary storage or persistent storage. The program stored in the storage medium 1030 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1022 can be configured to communicate with the storage medium 1030 and execute a series of instruction operations in the storage medium 1030 on the server 1000.
[0205] The server 1000 may also include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input and output interfaces 1058, and / or one or more operating systems 1041, such as Windows Server 2003. TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM etc.
[0206] In this embodiment, the central processing unit 1022 in the server 1000 can execute the methods provided in various optional implementations of the above embodiments.
[0207] The computer device provided in the embodiment of the present application can also be a terminal, see Figure 11, which is a structural diagram of a terminal provided in the embodiment of the present application. Taking the terminal as a smartphone as an example, the smartphone includes components such as a radio frequency (RF) circuit 1110, a memory 1120, an input unit 1130, a display unit 1140, a sensor 1150, an audio circuit 1160, a wireless fidelity (WiFi) module 1170, a processor 1180, and a power supply 11120. The input unit 1130 may include a touch panel 1131 and other input devices 1132, the display unit 1140 may include a display panel 1141, and the audio circuit 1160 may include a speaker 1161 and a microphone 1162. Those skilled in the art will understand that the smartphone structure shown in Figure 11 does not constitute a limitation on smartphones, and may include more or fewer components than shown, or combine certain components, or arrange components differently.
[0208] The memory 1120 can be used to store software programs and modules. The processor 1180 executes the various functional applications and data processing of the smartphone by running the software programs and modules stored in the memory 1120. The memory 1120 may mainly include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created based on the use of the smartphone (such as audio data, a phone book, etc.). In addition, the memory 1120 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0209] Processor 1180 is the control center of the smartphone, connecting all components of the smartphone using various interfaces and circuits. It executes software programs and / or modules stored in memory 1120 and accesses data stored in memory 1120 to perform various smartphone functions and process data. Optionally, processor 1180 may include one or more processing units. Preferably, processor 1180 integrates an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 1180.
[0210] In this embodiment, the processor 1180 in the smartphone may execute the methods provided in various optional implementations of the above embodiments.
[0211] According to one aspect of the present application, a computer-readable storage medium is provided, which is used to store a computer program. When the computer program is run on a computer device, the computer device executes the methods provided in various optional implementations of the above embodiments.
[0212] According to one aspect of the present application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in various optional implementations of the above-described embodiments.
[0213] The descriptions of the processes or structures corresponding to the above figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.
[0214] The terms "first", "second" etc. in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, the process, method, system, product or equipment comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0215] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0216] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0217] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0218] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store computer programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a RAM, a magnetic disk or an optical disk.
[0219] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, ordinary technical members in this field should understand that they can still modify the technical solutions recorded in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for determining an image encoder, the method being executed by a computer device, the method comprising: The first sample image is image-encoded by an image encoder in an initial reconstruction model to obtain first image block codes respectively corresponding to a plurality of first image blocks in the first sample image, and second image block codes respectively corresponding to a plurality of second image blocks in the second sample image are obtained, wherein the second image block codes respectively corresponding to the plurality of second image blocks are obtained by image-encoding the plurality of second image blocks by a pre-trained encoder; the first sample image and the second sample image are a plurality of scanned images of a first object under different illumination parameters; According to the first image block codes respectively corresponding to the multiple first image blocks, encoding prediction is performed on the multiple second image blocks in the second sample image through the reconstruction network in the initial reconstruction model to obtain the first prediction codes respectively corresponding to the multiple second image blocks; According to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and the loss function of the initial reconstruction model, the initial reconstruction model is trained to obtain a first reconstruction model; The image encoder in the first reconstruction model is determined as the image encoder in the initial detection model; the initial detection model is used to train the image defect detection model.
2. According to the method of claim 1, the loss function of the initial reconstruction model is a cross entropy loss function; the initial reconstruction model is trained according to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks and the loss function of the initial reconstruction model to obtain the first reconstruction model, comprising: Determine, according to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and the cross entropy loss function, a first prediction probability that each first prediction code is a corresponding second image block code; The initial reconstruction model is trained with the goal of maximizing the plurality of first prediction probabilities to obtain the first reconstruction model.
3. According to the method of claim 1 or 2, the step of obtaining the pre-trained encoder comprises: Performing image encoding on the third sample image by using an initial encoder to obtain third image block features respectively corresponding to a plurality of third image blocks in the third sample image; The third sample image is a scanned image of the second object; Determine, from a plurality of preset discrete codes, third image block codes corresponding to the plurality of third image block features respectively; the third image block codes corresponding to the plurality of third image block features respectively belong to the plurality of preset discrete codes; According to the third image block codes respectively corresponding to the plurality of third image block features, reconstructing the third sample image through an initial decoder to obtain a reconstructed sample image; According to the reconstructed sample image, the third sample image, and the loss function of the initial encoder and the initial decoder, model training is performed on the initial encoder and the multiple preset discrete codes to obtain the pre-trained encoder.
4. The method according to claim 3, wherein determining the third image block codes corresponding to the plurality of third image block features respectively from a plurality of preset discrete codes comprises: For each third image block feature, perform similarity calculation according to the third image block feature and the plurality of preset discrete codes to obtain a similarity between the third image block feature and each of the plurality of preset discrete codes; The preset discrete code corresponding to the maximum similarity is determined as the third image block code corresponding to the third image block feature.
5. The method according to claim 3 or 4, wherein the loss function of the initial encoder and the initial decoder is a cross entropy loss function; and performing model training on the initial encoder and the plurality of preset discrete codes according to the reconstructed sample image, the third sample image, and the loss function of the initial encoder and the initial decoder to obtain the pre-trained encoder comprises: Determining, according to the reconstructed sample image, the third sample image and the cross entropy loss function, a second prediction probability that the reconstructed sample image is the third sample image; The initial encoder and the plurality of preset discrete codes are model trained with the goal of maximizing the second prediction probability to obtain the pre-trained encoder.
6. The method according to any one of claims 3 to 5, wherein the first image block codes respectively corresponding to the plurality of first image blocks are first image block features, and the second image block codes respectively corresponding to the plurality of second image blocks belong to a plurality of preset discrete codes after training; and the image encoding is performed on the first sample image by the image encoder in the initial reconstruction model to obtain the first image block codes respectively corresponding to the plurality of first image blocks in the first sample image, comprising: Performing image encoding on the first sample image by using an image encoder in the initial reconstruction model to obtain first image block features respectively corresponding to the plurality of first image blocks; The obtaining of second image block codes respectively corresponding to a plurality of second image blocks in the second sample image comprises: Performing image encoding on the plurality of second image blocks by using the pre-trained encoder to obtain second image block features respectively corresponding to the plurality of second image blocks; Among the plurality of preset discrete codes after training, second image block codes corresponding to a plurality of second image block features are determined.
7. According to the method of claim 6, the loss function of the initial reconstruction model is a cross entropy loss function; the initial reconstruction model is trained according to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and the loss function of the initial reconstruction model to obtain the first reconstruction model, comprising: Determine, according to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and the cross entropy loss function, a first prediction probability that each first prediction code is a corresponding second image block code; For each first prediction code, if the first prediction code belongs to the plurality of preset discrete codes after the training, determining that a preset coefficient associated with a first prediction probability corresponding to the first prediction code is 1; With the goal of maximizing the first prediction probability associated with the preset coefficient being 1, model training is performed on the initial reconstruction model to obtain the first reconstruction model.
8. The method according to claim 7, wherein for each first prediction code, if the first prediction code belongs to the plurality of preset discrete codes after training, determining that the preset coefficient associated with the first prediction probability corresponding to the first prediction code is 1 comprises: Obtaining preset discrete identifiers corresponding to the plurality of preset discrete codes after the training; For each first prediction code, if the first prediction code corresponds to any preset discrete identifier among the plurality of preset discrete identifiers, it is determined that the preset coefficient associated with the first prediction probability corresponding to the first prediction code is 1.
9. The method according to any one of claims 1 to 8, further comprising: Randomly sampling the multiple first image blocks to obtain a first number of first image blocks; The first number is smaller than the number of blocks of the plurality of first image blocks; The step of encoding the first image blocks respectively corresponding to the plurality of first image blocks and predicting the encoding of the plurality of second image blocks in the second sample image through the reconstruction network in the initial reconstruction model to obtain the first prediction encodings respectively corresponding to the plurality of second image blocks includes: According to the first image block codes respectively corresponding to the first number of first image blocks, encoding prediction is performed on the second number of second image blocks through the reconstruction network in the initial reconstruction model to obtain first prediction codes respectively corresponding to the second number of second image blocks, and the second number of second image blocks corresponds to the first number of first image blocks; The method of performing model training on the initial reconstruction model according to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and the loss function of the initial reconstruction model to obtain the first reconstruction model comprises: According to the first prediction codes respectively corresponding to the second number of second image blocks, the second image block codes respectively corresponding to the second number of second image blocks, and the loss function of the initial reconstruction model, the initial reconstruction model is trained to obtain the first reconstruction model.
10. The method according to any one of claims 1 to 9, further comprising: Performing image encoding on the plurality of second image blocks by using an image encoder in the first reconstruction model to obtain fourth image block codes respectively corresponding to the plurality of second image blocks, and obtaining fifth image block codes respectively corresponding to the plurality of first image blocks, wherein the fifth image block codes respectively corresponding to the plurality of first image blocks are obtained by performing image encoding on the plurality of first image blocks by using the pre-trained encoder; According to the fourth image block codes respectively corresponding to the plurality of second image blocks, encoding prediction is performed on the plurality of first image blocks through a reconstruction network in the first reconstruction model to obtain second prediction codes respectively corresponding to the plurality of first image blocks; According to the second prediction codes respectively corresponding to the plurality of first image blocks, the fifth image block codes respectively corresponding to the plurality of first image blocks, and the loss function of the first reconstruction model, the first reconstruction model is trained to obtain a second reconstruction model; The step of determining the image encoder in the first reconstruction model as the image encoder in the initial detection model comprises: The image encoder in the second reconstruction model is determined as the image encoder in the initial detection model.
11. A device for determining an image encoder, the device being deployed on a computer device, the device comprising: Encoding unit, prediction unit, training unit and determination unit; The encoding unit is used to perform image encoding on the first sample image through the image encoder in the initial reconstruction model to obtain first image block codes respectively corresponding to multiple first image blocks in the first sample image, and obtain second image block codes respectively corresponding to multiple second image blocks in the second sample image, wherein the second image block codes respectively corresponding to the multiple second image blocks are obtained by performing image encoding on the multiple second image blocks respectively through a pre-trained encoder; the first sample image and the second sample image are multiple scanned images of the first object under different illumination parameters; The prediction unit is used to perform coding prediction on a plurality of second image blocks in the second sample image through a reconstruction network in the initial reconstruction model according to the first image block codes respectively corresponding to the plurality of first image blocks, so as to obtain first prediction codes respectively corresponding to the plurality of second image blocks; The training unit is used to perform model training on the initial reconstruction model according to the first prediction codes respectively corresponding to the multiple second image blocks, the second image block codes respectively corresponding to the multiple second image blocks, and the loss function of the initial reconstruction model to obtain a first reconstruction model; The determining unit is used to determine the image encoder in the first reconstruction model as the image encoder in the initial detection model; The initial detection model is used to train the image defect detection model.
12. The apparatus according to claim 11, wherein the loss function of the initial reconstruction model is a cross entropy loss function; and the training unit is specifically configured to: Determine, according to the first prediction codes respectively corresponding to the plurality of second image blocks, the second image block codes respectively corresponding to the plurality of second image blocks, and the cross entropy loss function, a first prediction probability that each first prediction code is a corresponding second image block code; The initial reconstruction model is trained with the goal of maximizing the plurality of first prediction probabilities to obtain the first reconstruction model.
13. A computer device, comprising a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1 to 10 according to the instructions in the computer program.
14. A computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and when the computer program is executed on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 10.
15. A computer program product, comprising a computer program, which enables a computer device to execute the method according to any one of claims 1 to 10 when the computer program is run on the computer device.