Image reconstruction method, image detection method, device, equipment, vehicle and medium

By acquiring and merging the first coding vector of low-dimensionality, the problem of large amount of calculation of variable component quantization autoencoder is solved, and the image reconstruction efficiency is improved.

CN119991668AActive Publication Date: 2025-05-13BYD CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510466456.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The variable component quantization autoencoder has a large amount of calculation during image reconstruction, resulting in low reconstruction efficiency.

Method used

By acquiring a plurality of first coded vectors of the original image, the dimension of each coded vector is lower than that of the second coded vector, which contains the global features of the original image. These encoded vectors are then quantized and merged, and merged vectors are generated for image reconstruction.

Benefits of technology

The calculation amount in the quantization stage is reduced, and the efficiency of image reconstruction is improved by merging vectors to contain the global features of the original image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991668A_ABST
    Figure CN119991668A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image reconstruction method and device, an image detection method and device, equipment, a vehicle and a medium, a plurality of first coding vectors of an original image are obtained according to the original image, the dimension corresponding to each first coding vector is lower than the dimension of a second coding vector, and the second coding vector comprises the global feature of the original image. Therefore, the calculation amount of the quantization stage can be reduced, and since each first coding vector at least has a part of different features, after the quantization vectors are obtained according to the first coding vectors and the quantization vectors corresponding to the first coding vectors are merged, global features of the original image can be contained, and the quantization efficiency is improved. Therefore, image reconstruction can be realized based on the combined vector, and the image reconstruction efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to an image reconstruction method, an image detection method, a device, a equipment, a vehicle and a medium. Background Art

[0002] In the process of industrial production and manufacturing, image anomaly detection is an important means to achieve product quality control. The reconstruction method is widely used in image anomaly detection. The reconstruction method refers to reconstructing the original image to obtain a reconstructed image, and determining whether the original image is abnormal based on the difference between the reconstructed image and the original image.

[0003] Currently, the original image can be reconstructed through the Vector Quantized Variational Auto-Encoder (VQVAE).

[0004] However, the variational quantization autoencoder has a large amount of calculation in the process of reconstructing the original image, resulting in low reconstruction efficiency. Summary of the invention

[0005] Embodiments of the present application provide an image reconstruction method, an image detection method, an apparatus, a device, a vehicle and a medium to improve reconstruction efficiency.

[0006] In a first aspect, an embodiment of the present application provides an image reconstruction method, comprising:

[0007] Obtaining a plurality of first coding vectors of the original image according to the original image, wherein each of the first coding vectors has at least partially different features, a dimension corresponding to each of the first coding vectors is lower than a dimension of a second coding vector, and the second coding vector includes a global feature of the original image;

[0008] Obtain a quantization vector according to the first encoding vector;

[0009] Merging the quantized vectors corresponding to each of the first coding vectors to obtain a merged vector;

[0010] The original image is reconstructed according to the merged vector to obtain a reconstructed image.

[0011] In some optional implementations, obtaining a plurality of first encoding vectors of the original image according to the original image includes:

[0012] The features of the original image are extracted respectively by using multiple encoders in the reconstruction model, and the extracted features are encoded to obtain multiple first encoding vectors of the original image.

[0013] In some optional implementations, obtaining a quantization vector according to the first encoding vector includes:

[0014] The similarity between the first encoding vector and the codebook vector is calculated by a quantizer in the reconstruction model, and a quantization vector corresponding to the first encoding vector is determined according to the similarity.

[0015] In some optional implementations, reconstructing the original image according to the merged vector to obtain a reconstructed image includes:

[0016] The original image is reconstructed according to the merged vector by a decoder in the reconstruction model to obtain a reconstructed image.

[0017] In some optional embodiments, the decoder is established based on a first reconstruction loss function and a second reconstruction loss function, wherein the first reconstruction loss function is a loss function corresponding to an abnormally annotated area of ​​a training image, and the second reconstruction loss function is a loss function corresponding to an unannotated area of ​​the training image, and the first reconstruction loss function is greater than the second reconstruction loss function.

[0018] In some optional implementations, obtaining a plurality of first encoding vectors of the original image according to the original image includes:

[0019] Obtaining a feature vector of the original image;

[0020] The feature vector is encoded to obtain a plurality of first encoding vectors of the original image.

[0021] In some optional implementations, encoding the feature vector to obtain a plurality of first encoding vectors of the original image includes:

[0022] The feature vectors are encoded respectively by multiple encoders in the reconstruction model to obtain multiple first encoding vectors of the original image.

[0023] In a second aspect, the present application provides an image detection method, comprising:

[0024] Acquire a reconstructed image reconstructed by the image reconstruction method described in the first aspect;

[0025] It is determined whether the original image is abnormal according to the reconstructed image.

[0026] In some optional implementations, determining whether the original image is abnormal according to the reconstructed image includes:

[0027] Comparing the reconstructed image with the original image to obtain a comparison result;

[0028] It is determined whether the original image is abnormal according to the comparison result.

[0029] In some optional implementations, comparing the reconstructed image with the original image to obtain a comparison result includes:

[0030] The reconstructed image and the original image are compared using an image segmentation model to obtain a comparison result.

[0031] In a third aspect, the present application provides an image reconstruction device, comprising:

[0032] an encoding module, configured to obtain a plurality of first encoding vectors of the original image according to the original image, wherein each of the first encoding vectors has at least partially different features, and a dimension corresponding to each of the first encoding vectors is lower than a dimension of a second encoding vector, and the second encoding vector includes a global feature of the original image;

[0033] A quantization module, used for obtaining a quantization vector according to the first encoding vector;

[0034] A merging module, used for merging the quantization vectors corresponding to each of the first encoding vectors to obtain a merged vector;

[0035] A reconstruction module reconstructs the original image according to the merged vector to obtain a reconstructed image.

[0036] In a fourth aspect, the present application provides an image detection device, comprising:

[0037] An acquisition module, used to acquire a reconstructed image reconstructed by the image reconstruction method described in the first aspect;

[0038] A determination module is used to determine whether the original image is abnormal according to the reconstructed image.

[0039] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0040] The memory stores computer-executable instructions;

[0041] The processor executes the computer-executable instructions stored in the memory, so that the processor performs the above method.

[0042] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-readable storage medium is stored computer-executable instructions, and the computer-executable instructions are used to implement the above method when executed by a processor.

[0043] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0044] The image reconstruction method, image detection method, apparatus, device, vehicle and medium provided in the embodiments of the present application obtain multiple first coding vectors of the original image based on the original image. Since the dimension corresponding to each first coding vector is lower than the dimension of the second coding vector, the second coding vector includes the global features of the original image, thereby reducing the amount of calculation in the quantization stage. Moreover, since each first coding vector has at least partially different features, after obtaining the quantization vector based on the first coding vector and merging the quantization vectors corresponding to each first coding vector, the global features of the original image can be included, thereby enabling image reconstruction based on the merged vector, thereby improving image reconstruction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0046] Figure 1 A schematic diagram of the process of image reconstruction method provided in this application;

[0047] Figure 2 Schematic diagram of image reconstruction provided for this application;

[0048] Figure 3 A schematic diagram of the process of the image detection method provided in this application;

[0049] Figure 4 Schematic diagram of image detection provided for this application;

[0050] Figure 5 The architecture diagram of the image segmentation model provided for this application;

[0051] Figure 6 A schematic diagram of the structure of the image reconstruction device provided in this application;

[0052] Figure 7 A schematic diagram of the structure of the image detection device provided in this application;

[0053] Figure 8 A schematic diagram of the structure of the electronic device provided in this application.

[0054] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0055] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0056] In the process of industrial production and manufacturing, image anomaly detection is an important means to achieve product quality control. The reconstruction method is widely used in image anomaly detection. The reconstruction method reconstructs the original image through a pre-trained model to obtain a reconstructed image. Since the pre-trained model is trained based on normal samples, it can capture the characteristics of normal samples. If the abnormal sample is reconstructed, the abnormal characteristics cannot be well captured, which increases the difference between the reconstructed image and the original image. Therefore, when the difference between the reconstructed image and the original image is large, the original image is considered abnormal.

[0057] For example, the original image can be reconstructed by a variational auto-encoder (VAE). A variational auto-encoder is a reconstruction (also called generation) model that can extract the latent features of the input data and represent these latent features in the form of encoding vectors (also called latent vectors), which can then be used to reconstruct the original data or generate new data samples. A variational auto-encoder includes an encoder and a decoder. The encoder maps the input data to a low-dimensional latent space. In the process of mapping high-dimensional input data to a low-dimensional space, the encoder attempts to retain the key information in the input data while removing redundancy and noise. The decoder then reconstructs data similar to the input data based on the latent vectors in the latent space.

[0058] However, when the variational autoencoder uses a relatively low-dimensional latent space to represent image data, it is easy to cause the details of complex images to be lost, making it difficult to accurately reconstruct complex images.

[0059] For example, the original image can also be reconstructed through a variational quantized autoencoder (VQVAE). The variational quantized autoencoder is a reconstruction model that combines the variational autoencoder and vector quantization technology. The variational quantized autoencoder includes not only an encoder and a decoder, but also a quantizer. The encoder maps the original image to a latent space and outputs a coding vector. The quantizer can map the coding vector to the nearest neighbor codebook vector in the codebook to obtain a discrete quantized vector. The decoder is based on the discrete quantized vector to better capture and represent discrete features in the image, such as edges, textures, etc., so that an image similar to the original image can be reconstructed.

[0060] However, the variational quantization autoencoder needs to map the encoding vector to the nearest neighbor codebook vector in the codebook, which requires a lot of calculation and results in low reconstruction efficiency.

[0061] To this end, the present application proposes an image reconstruction method, which obtains multiple first coding vectors of the original image based on the original image. Since the dimension corresponding to each first coding vector is lower than the dimension of the second coding vector, the second coding vector includes the global features of the original image, thereby reducing the amount of calculation in the quantization stage. Moreover, since each first coding vector has at least partially different features, after the quantization vectors corresponding to each first coding vector are merged, the global features of the original image can be included, thereby realizing image reconstruction.

[0062] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0063] Figure 1 A schematic diagram of the process of image reconstruction method provided in this application, such as Figure 1 As shown, the method includes:

[0064] S101. Obtain multiple first coding vectors of the original image based on the original image, wherein there are at least partially different features between each of the first coding vectors, and the dimension corresponding to each first coding vector is lower than the dimension of the second coding vector, and the second coding vector includes the global features of the original image.

[0065] For example, images usually exist in the form of pixels. Since images are usually composed of a large number of pixels, images can be regarded as high-dimensional data points. The encoding vector is converted into a low-dimensional representation by the original image, which may involve compression and denoising to extract various feature information such as texture, color, shape, etc. of the original image.

[0066] For example, the original image can be mapped to a low-dimensional continuous latent space, which includes multiple points, each of which corresponds to a coding vector, representing a compressed representation of the feature. Therefore, by performing feature extraction and encoding on the original image, the original image can be converted into a low-dimensional coding vector.

[0067] In some embodiments, the method of the present application can be performed by reconstructing the model.

[0068] In some optional embodiments, the reconstruction model may include an encoder, through which the original image is feature extracted, and the extracted features are encoded to obtain a second encoding vector. Since the original image is feature extracted and encoded by an encoder, the second encoding vector may include global features of the original image. The global features here can be understood as including all key information of the original image, and the key information of the original image may include important attributes such as the main content, structure, texture, and color of the image.

[0069] In some optional embodiments, the reconstruction module may include multiple encoders, such as Figure 2 As shown, multiple encoders can be used to extract features from the original image respectively, and after encoding the extracted features, multiple first encoding vectors of the original image can be obtained. After each encoder extracts and encodes the features of the original image, a first encoding vector of the original image can be obtained, and multiple first encoding vectors of the original image can be obtained through multiple encoders. Since the original image is processed by multiple encoders respectively, the feature information of the original image corresponding to each encoder can be at least partially different, so there are at least partially different features between each first encoding vector, so that the dimension of the first encoding vector output by each encoder can be reduced. Accordingly, the dimension of the first encoding vector can be lower than the dimension of the second encoding vector.

[0070] For example, when the number of encoders is two, the dimension of the first encoding vector output by each encoder may be half the dimension of the second encoding vector.

[0071] For example, the number of encoders may be determined according to actual conditions, such as two or three.

[0072] For example, the first encoding vector may include multiple vectors of the same dimension, each of which may focus on different aspects or features of the data. For example, one vector may capture the color information of the image, while another vector may capture the shape information of the image.

[0073] By way of example, the reconstruction model may be, for example, a variational quantized autoencoder.

[0074] In some embodiments, considering that the receptive field of the reconstruction model is limited, it is difficult to capture the long-term dependencies in the long sequence data when processing long sequence data. Therefore, the feature vector of the original image can be obtained first, and the feature vector can be encoded to obtain multiple first encoding vectors of the original image to solve the problem that the reconstruction model may not be able to process long sequence data.

[0075] In some optional implementations, the original image may be feature extracted by a deep learning model to obtain a feature vector of the original image. The deep learning module may include, for example, a Transformer architecture, which can capture long-distance dependencies and local features in an image through techniques such as a self-attention mechanism and positional encoding.

[0076] In some optional embodiments, after obtaining the feature vector of the original image, the feature vector can be encoded separately by multiple encoders in the reconstruction model to obtain multiple first encoding vectors of the original image, thereby reducing the dimension of the first encoding vector output by each encoder.

[0077] S102. Obtain a quantization vector according to the first encoding vector.

[0078] It should be noted that the quantization vector can better capture and represent discrete features in the image, such as edges, textures, etc. Therefore, in this step, the quantization vector is obtained according to the first encoding vector.

[0079] In some embodiments, the reconstruction model includes a quantizer, and the similarity between the first coding vector and the codebook vector is calculated by the quantizer in the reconstruction model, and the quantization vector corresponding to the first coding vector is determined based on the similarity, thereby obtaining a quantization vector corresponding to the distance of the first coding vector.

[0080] Specifically, the similarity between each first encoding vector and the codebook vector is calculated to determine the quantization vector corresponding to each first encoding vector.

[0081] For example, taking a first coding vector as an example, the similarity between the first coding vector and each codebook vector in the codebook can be calculated, and the codebook vector with the greatest similarity to the first coding vector is used as the quantization vector corresponding to the first coding vector, such as Figure 2 shown.

[0082] The codebook is generated during the training process of the reconstruction model. During the training process of the reconstruction model, the reconstruction model needs to learn how to encode the input data (such as an image) into discrete representations in the latent space and reconstruct the input data through these discrete representations. The codebook is a set of discrete vectors used to represent features in the latent space.

[0083] S103: Merge the quantization vectors corresponding to each first encoding vector to obtain a merged vector.

[0084] Since each first coding vector has at least partially different features, the quantization vectors corresponding to each first coding vector are merged in this step to obtain a merged vector, such as Figure 2 As shown, the merged vector can contain the global characteristics of the original image, which facilitates the subsequent more accurate reconstruction of the original image.

[0085] In some optional implementations, multiple first encoding vectors may be concatenated or superimposed along a specific dimension.

[0086] For example, multiple first encoders can be spliced ​​or superimposed in the channel direction. In deep learning, such as models such as variational quantization autoencoders, image data is usually decomposed into multiple channels, each channel representing a different feature or color component of the image, such as the three color channels of RGB (Red Green Blue) images. The encoding vector is a high-dimensional representation extracted by the encoder network based on these channel features. Merging the encoding vectors in the channel direction means splicing these high-dimensional representations along the channel dimension to form a new vector containing more feature information, which helps the reconstruction model better capture and utilize the multi-scale and multi-level features in the image, thereby improving the quality of image generation or encoding.

[0087] For example, two encoding vectors z1 and z2, which have shapes (batch_size, height, width, channels1) and (batch_size, height, width, channels2), can be concatenated along the channel dimension to form a new encoding vector with shape (batch_size, height, width, channels1+channels2).

[0088] S104: Reconstruct the original image according to the merged vector to obtain a reconstructed image.

[0089] In some embodiments, the reconstruction model includes a decoder, and the decoder in the reconstruction model reconstructs the original image according to the merged vector to obtain a reconstructed image, such as Figure 2shown.

[0090] For example, the decoder is able to convert a low-dimensional discrete representation (quantized vector) into a high-dimensional raw data space (such as an image). For example, the decoder receives the quantized vector and processes it through a series of neural network layers, which gradually expand the low-dimensional latent representation into a high-dimensional output. In image reconstruction tasks, the last few layers of the decoder usually use deconvolution operations to generate outputs of the same size as the input image.

[0091] In some optional embodiments, the decoder is established based on a first reconstruction loss function and a second reconstruction loss function, wherein the first reconstruction loss function is a loss function corresponding to the abnormally annotated area of ​​the training image, and the second reconstruction loss function is a loss function corresponding to the non-annotated area of ​​the training image. The first reconstruction loss function is greater than the second reconstruction loss function. By reducing the reconstruction loss of the abnormally annotated area and increasing the reconstruction loss of the non-annotated area, the original image can be reconstructed more accurately and the reconstruction effect can be improved.

[0092] For example, the training image can be first input into a pre-established reconstruction model, and the reconstructed image output by the reconstruction model can be displayed on an interactive user interface (UI). Then, the reconstructed image displayed in the user interface is manually annotated to indicate which areas are normal (i.e., non-annotated areas) and which areas are abnormal (i.e., abnormal annotated areas). Subsequently, the reconstruction model can be adjusted based on the manually annotated data. For areas annotated as normal (non-annotated areas), the model will try to reduce the reconstruction loss to better match these normal features. For areas annotated as abnormal (abnormal annotated areas), the reconstruction model may try to increase the reconstruction loss so that these abnormal features can be more accurately identified in the future. Through this combination of supervised and unsupervised training, the reconstruction model can gradually learn the information that is considered abnormal by humans, thereby improving the accuracy and robustness of the reconstruction model for anomaly detection.

[0093] The image reconstruction method provided by the embodiment of the present application obtains multiple first coding vectors of the original image based on the original image. Since the dimension corresponding to each first coding vector is lower than the dimension of the second coding vector, the second coding vector includes the global features of the original image, thereby reducing the amount of calculation in the quantization stage. Moreover, since each first coding vector has at least partially different features, after the quantization vectors corresponding to each first coding vector are merged, the global features of the original image can be included, thereby realizing image reconstruction.

[0094] Figure 3 A schematic diagram of the process of the image detection method provided in this application, such as Figure 3 As shown, the method includes:

[0095] S201: Obtain a reconstructed image reconstructed by the above reconstruction method.

[0096] S202: Determine whether the original image is abnormal based on the reconstructed image.

[0097] In some optional embodiments, after acquiring the reconstructed image, the reconstructed image and the original image are compared to obtain a comparison result, and whether the original image is abnormal is determined based on the comparison result, thereby achieving abnormality detection of the original image.

[0098] In some optional implementations, the reconstructed image and the original image may be compared using an image segmentation model to obtain a comparison result, so as to more clearly display the difference area between the original image and the reconstructed image.

[0099] For example, the image segmentation model can be a UnetPP architecture, which combines large-scale and small-scale information through skip connections and downsampling, so that small and large objects can be effectively distinguished to ensure the accuracy of segmentation. The input is the original image and the reconstructed image, and the output is a single-channel image, such as Figure 4 In the right figure, black represents the area with difference and white represents the area without difference. Compared with the actual difference area, Figure 4 The left image in the figure shows the difference area more clearly.

[0100] The UnetPP architecture adopts an encoder-decoder structure and introduces multi-scale feature fusion and denser skip connections, such as Figure 5 As shown in the figure, through the synergy of components such as convolution, upsampling, downsampling and skip connection, the image features are effectively extracted and fused, thereby improving the accuracy of image segmentation and reducing errors.

[0101] Convolution refers to extracting local features (such as edges, textures, etc.) from the input data through sliding calculation of the convolution kernel. Each convolution kernel only focuses on the local area of ​​the input data. Upsampling is used in the decoder part. The upsampling operation is used to gradually restore the spatial resolution of the feature map to make it close to the size of the input image. Through upsampling, the network can retain more detail information, which helps to more accurately identify different areas in the image. Downsampling is used in the encoder part. The downsampling operation extracts higher-level features by reducing the spatial dimension of the feature map. Downsampling helps to reduce video memory and computation, while increasing the receptive field, allowing the network to extract features over a larger image range, thereby improving segmentation accuracy. Skip connection refers to directly connecting the feature map in the encoder to the corresponding layer in the decoder, realizing multi-scale feature fusion. Through skip connections, the network can fuse the feature map in the downsampling process during the upsampling process, thereby restoring more detail information and reducing segmentation errors.

[0102] In some optional embodiments, the reconstruction of normal images also has errors. In order to prevent the reconstruction errors of normal areas from affecting the judgment of abnormal areas, after acquiring the reconstructed image, the reconstructed image can be processed by a noise filter. The noise filter can identify and process the anomalies or noise in the reconstructed image, thereby being able to more accurately determine whether the original image is abnormal.

[0103] For example, a noise filter can be trained using labeled data, where the labels indicate which regions are normal and which are abnormal, so that the filter can learn how to effectively identify and handle abnormal regions.

[0104] In some optional implementations, some image comparison tools may be used to compare the reconstructed image with the original image.

[0105] The image detection method provided in the embodiment of the present application determines whether the original image is abnormal by reconstructing the image.

[0106] Figure 6 A schematic diagram of the structure of the image reconstruction device provided in this application, such as Figure 6 As shown, the image reconstruction device 10 provided in this embodiment includes:

[0107] The encoding module 11 is used to obtain a plurality of first encoding vectors of the original image according to the original image, wherein each of the first encoding vectors has at least partially different features, and the dimension corresponding to each of the first encoding vectors is lower than the dimension of the second encoding vector, and the second encoding vector includes the global features of the original image;

[0108] A quantization module 12, configured to obtain a quantization vector according to the first encoding vector;

[0109] A merging module 13, configured to merge the quantization vectors corresponding to each first encoding vector to obtain a merged vector;

[0110] The reconstruction module 14 reconstructs the original image according to the merged vector to obtain a reconstructed image.

[0111] In some optional implementations, the encoding module 11 is specifically used to extract features from the original image using multiple encoders in the reconstruction model, and encode the extracted features to obtain multiple first encoding vectors of the original image.

[0112] In some optional implementations, the quantization module 12 is specifically configured to calculate the similarity between the first coding vector and the codebook vector through a quantizer in the reconstruction model, and determine the quantization vector corresponding to the first coding vector according to the similarity.

[0113] In some optional implementations, the reconstruction module 14 is specifically configured to reconstruct the original image according to the merged vector through a decoder in the reconstruction model to obtain a reconstructed image.

[0114] In some optional embodiments, the decoder is established based on a first reconstruction loss function and a second reconstruction loss function, the first reconstruction loss function is a loss function corresponding to the abnormal annotated area of ​​the training image, the second reconstruction loss function is a loss function corresponding to the non-annotated area of ​​the training image, and the first reconstruction loss function is greater than the second reconstruction loss function.

[0115] In some optional implementations, the encoding module 11 is specifically used to obtain a feature vector of the original image; and encode the feature vector to obtain a plurality of first encoding vectors of the original image.

[0116] In some optional implementations, the encoding module 11 is specifically configured to encode the feature vectors respectively through a plurality of encoders in the reconstruction model to obtain a plurality of first encoding vectors of the original image.

[0117] The image reconstruction device provided in this embodiment can execute the image reconstruction method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail in this embodiment.

[0118] Figure 7 The schematic diagram of the structure of the image detection device provided in this application is as follows: Figure 7 As shown, the image detection device 20 provided in this embodiment includes:

[0119] An acquisition module 21, used for acquiring a reconstructed image reconstructed by the above-mentioned reconstruction method;

[0120] The determination module 22 is used to determine whether the original image is abnormal according to the reconstructed image.

[0121] In some optional implementations, the determination module 22 is specifically configured to compare the reconstructed image with the original image to obtain a comparison result; and determine whether the original image is abnormal based on the comparison result.

[0122] In some optional implementations, the determination module 22 is specifically configured to compare the reconstructed image and the original image using an image segmentation model to obtain a comparison result.

[0123] The image detection device provided in this embodiment can execute the image detection method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail in this embodiment.

[0124] Figure 8 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 8As shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the electronic device 50 also includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus.

[0125] In a specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that at least one processor 501 executes the above method.

[0126] The specific implementation process of the processor 501 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0127] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0128] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.

[0129] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0130] The present application also provides a vehicle, comprising the above-mentioned electronic device.

[0131] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0132] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0133] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. An image reconstruction method, characterized in that: include: Obtaining a plurality of first coding vectors of the original image according to the original image, wherein each of the first coding vectors has at least partially different features, a dimension corresponding to each of the first coding vectors is lower than a dimension of a second coding vector, and the second coding vector includes a global feature of the original image; Obtain a quantization vector according to the first encoding vector; Merging the quantized vectors corresponding to each of the first coding vectors to obtain a merged vector; The original image is reconstructed according to the merged vector to obtain a reconstructed image.

2. The method according to claim 1, characterized in that The step of obtaining a plurality of first encoding vectors of the original image according to the original image comprises: The features of the original image are extracted respectively by using multiple encoders in the reconstruction model, and the extracted features are encoded to obtain multiple first encoding vectors of the original image.

3. The method according to claim 1, characterized in that: Obtaining a quantization vector according to the first encoding vector includes: The similarity between the first coding vector and the codebook vector is calculated by a quantizer in the reconstruction model, and a quantization vector corresponding to the first coding vector is determined according to the similarity.

4. The method according to claim 1, characterized in that The reconstructing the original image according to the merged vector to obtain a reconstructed image includes: The original image is reconstructed according to the merged vector by a decoder in the reconstruction model to obtain a reconstructed image.

5. The method according to claim 4, characterized in that The decoder is established based on a first reconstruction loss function and a second reconstruction loss function, wherein the first reconstruction loss function is a loss function corresponding to an abnormally annotated area of ​​a training image, and the second reconstruction loss function is a loss function corresponding to a non-annotated area of ​​the training image, and the first reconstruction loss function is greater than the second reconstruction loss function.

6. The method according to claim 1, characterized in that The step of obtaining a plurality of first encoding vectors of the original image according to the original image comprises: Obtaining a feature vector of the original image; The feature vector is encoded to obtain a plurality of first encoding vectors of the original image.

7. The method according to claim 6, characterized in that The encoding of the feature vector to obtain a plurality of first encoding vectors of the original image includes: The feature vectors are encoded respectively by multiple encoders in the reconstruction model to obtain multiple first encoding vectors of the original image.

8. An image detection method, characterized in that: include: Acquire a reconstructed image reconstructed by the image reconstruction method according to any one of claims 1 to 7; It is determined whether the original image is abnormal according to the reconstructed image.

9. The method according to claim 8, characterized in that The determining whether the original image is abnormal according to the reconstructed image includes: Comparing the reconstructed image with the original image to obtain a comparison result; It is determined whether the original image is abnormal according to the comparison result.

10. The method according to claim 9, characterized in that The comparing the reconstructed image and the original image to obtain a comparison result includes: The reconstructed image and the original image are compared using an image segmentation model to obtain a comparison result.

11. An image reconstruction device, characterized in that: include: an encoding module, configured to obtain a plurality of first encoding vectors of the original image according to the original image, wherein each of the first encoding vectors has at least partially different features, and a dimension corresponding to each of the first encoding vectors is lower than a dimension of a second encoding vector, and the second encoding vector includes a global feature of the original image; A quantization module, used for obtaining a quantization vector according to the first encoding vector; A merging module, used for merging the quantization vectors corresponding to each of the first encoding vectors to obtain a merged vector; A reconstruction module reconstructs the original image according to the merged vector to obtain a reconstructed image.

12. An image detection device, characterized in that: include: An acquisition module, used for acquiring a reconstructed image reconstructed by the image reconstruction method according to any one of claims 1 to 7; A determination module is used to determine whether the original image is abnormal according to the reconstructed image.

13. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 10.

14. A vehicle, characterized in that: An electronic device comprising the electronic device described in claim 13.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 10 when executed by a processor.

16. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 10 when being executed by a processor.

Citation Information

Patent Citations

  • Sparse grayscale image encoding and decoding method and system based on reconstruction residouble errors

    CN111343458A

  • Feature coding and decoding method, coding and decoding device training method, device and medium

    CN116012662A

  • Image compression system, image processing method, encoding and decoding method and electronic equipment

    CN118158425A

  • Abnormality determination device and abnormality determination method

    JP2018005773A