Image reconstruction method and device and storage medium

By training a target reconstruction network and utilizing a feature extraction module, an asymmetric attention module, and a quality-aware gating module to process industrial images, the problem of low-resolution image quality is solved, high-frequency feature enhancement and high-quality image reconstruction are achieved, meeting the needs of industrial inspection and recognition.

CN121481867APending Publication Date: 2026-02-06CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511462256.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

In industrial settings, the acquired images often suffer from insufficient resolution and suffer from problems such as high noise, motion blur, and compression artifacts, resulting in low image quality and making it difficult to perform anomaly detection and target recognition.

Method used

The trained target reconstruction network is used to infer from the initial image. The target reconstruction network includes a feature extraction module, an asymmetric attention module, a quality-aware gating module, and a reconstruction module connected in sequence. By training on the first and second historical image sets, the attention map and nonlinear activation function are adjusted using prior features to achieve the mapping from low-resolution image to high-resolution image.

Benefits of technology

It improves the high-frequency features of images, achieves high-quality image reconstruction under unknown degradation conditions, and enhances the performance of anomaly detection and target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121481867A_ABST
    Figure CN121481867A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data processing, and discloses an image reconstruction method and device and a storage medium, and the method comprises the steps: carrying out the reasoning of an initial image through employing a trained target reconstruction network, and obtaining a reconstructed second image, the target reconstruction network comprises a feature extraction module, an asymmetric attention module, a quality perception gating module and a reconstruction module which are connected in sequence, and the second image comprises more high-frequency image features of the first target object than the initial image, the asymmetric attention module is obtained by training a first historical image set and a second historical image set through an original reconstruction network, the quality perception gating module is obtained by training the first historical image set and the second historical image set, the first historical image set is a set of large-view images including a second target object, and the second historical image set is a set of large-view images including a second target object. The second historical image set is a set of a plurality of small-view images including the second target object, and the target reconstruction network can realize the reconstruction of the high-quality image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data processing technology, and provides a method, apparatus and storage medium for reconstructing images. Background Technology

[0002] Currently, high-resolution images are widely used and play a very important role in industrial scenarios such as intelligent manufacturing, defect detection, and equipment maintenance. However, in the industrial production process, due to various limitations (such as the hardware conditions of industrial cameras, ambient lighting conditions, etc.), the acquired images often have insufficient resolution. In addition, the images also have a variety of degradation problems such as high noise, motion blur, and compression artifacts, resulting in low image quality and making it difficult to achieve industrial goals such as anomaly detection and target recognition. Summary of the Invention

[0003] This application provides a method, apparatus, and storage medium for reconstructing images, thereby improving the quality of image reconstruction.

[0004] The specific technical solution provided in this application is as follows: In a first aspect, embodiments of this application provide a method for reconstructing an image, including: The trained target reconstruction network is used to reason about the initial image to obtain the reconstructed second image. The target reconstruction network includes a feature extraction module, an asymmetric attention module, a quality-aware gating module, and a reconstruction module connected in sequence. The second image has more high-frequency image features of the first target object than the initial image. Among them, the asymmetric attention module is obtained by adjusting the attention map of the original reconstruction network during the training process of the first historical image set and the second historical image set using the first prior features, and the quality perception gating module is obtained by adjusting the nonlinear activation function of the original reconstruction network during the training process of the first historical image set and the second historical image set using the first prior features. The first historical image set is a set of large field-of-view images including the second target object, and the second historical image set is a set of multiple small field-of-view images including the second target object. The first prior feature is obtained by training the first historical image set using the first prior feature extraction network. The high-frequency image features of the second target object represented by the first prior feature in the first historical image set include the high-frequency image features of the second target object represented by the second prior feature in the second historical image set. The second prior feature is obtained by training the second historical image set using the second prior feature extraction network.

[0005] Optionally, the target reconstruction network is trained in the following way: The original feature extraction module included in the original reconstruction network is used to extract features from each of the first historical images in the first historical image set to obtain the original features; The original asymmetric attention module included in the original reconstruction network is used to train the original features and the first prior features to obtain attention features; The original quality-aware gating module included in the original reconstruction network is used to train the attention features and the first prior features to obtain the quality features; The original reconstruction network is trained using the original reconstruction modules included in the original reconstruction network to obtain the trained historical images. The training is considered to have converged when the total loss value of the historical images obtained in this round meets the convergence condition after multiple rounds of iteration.

[0006] Optionally, the original asymmetric attention module included in the original reconstruction network is used to train the original features and the first prior features to obtain attention features, including: The original features are trained using the original asymmetric attention module included in the original reconstruction network to obtain value vectors and key vectors; The query vector is obtained by training the first prior feature using the original asymmetric attention module; Perform matrix multiplication on the query vector and the transposed key vector to obtain the attention map, and perform matrix multiplication on the attention map and the value vector to obtain the attention features.

[0007] Optionally, the attention features and first prior features are trained using the original quality-aware gating module included in the original reconstruction network to obtain quality features, including: The attention features and the first prior features are input into the original quality-aware gating module included in the original reconstruction network; The first prior features are averaged in the spatial dimension to obtain the quality-aware gating factor. The nonlinear activation function is modulated using a quality-aware gating factor; The quality features are obtained by performing matrix addition on the attention features using the modulated nonlinear activation function.

[0008] Optionally, the first prior features are trained in the following way: The first historical image set includes a large-view image of the second target object, which is then input into the first prior feature extraction network for training. The first high-frequency image features of the second target object included in the large field-of-view image are extracted using a first prior feature extraction network. For different locations of the second target object represented in the large field-of-view image, the first high-frequency image features and the second prior features are fused to obtain the first prior features.

[0009] Optionally, the second prior feature is trained in the following way: Multiple small-field images containing the second target object from the second historical image set are input into the second prior feature extraction network for training. The second high-frequency image features of the second target object included in each small field-of-view image are extracted using the second prior feature extraction network, and the second high-frequency image features are provided as the second prior features to the first prior feature extraction network.

[0010] Secondly, embodiments of this application also provide a system for reconstructing images, including: a cloud server, wherein the cloud server integrates any of the above-mentioned image reconstruction methods; The cloud server is configured to generate a reconstructed second image after performing an image reconstruction method based on an initial image; The cloud server was also configured to detect the second image to obtain target information.

[0011] Optionally, it may also include: data acquisition devices, industrial control devices, edge computing devices, and human-computer interaction devices; Both the data acquisition device and the industrial control device are connected to the edge computing device, which is connected to the cloud server, and the cloud server is also connected to the human-computer interaction device. The data acquisition device is configured to acquire a first image and send the first image to the edge computing device; The industrial control device is configured to send control timing signals to the edge computing device based on the acquired control signals; The edge computing device is configured to preprocess the first image under the control of the control timing, obtain an initial image after preprocessing, and send the initial image to the cloud server; The human-computer interaction device is configured to display target information and trigger defect alarms.

[0012] Thirdly, a cloud server includes: Memory, used to store executable instructions; A processor for reading and executing executable instructions stored in memory to implement the method as described in any of the first aspects.

[0013] Fourthly, a computer-readable storage medium, when instructions in the storage medium are executed by a processor, enables the processor to perform the method described in any of the first aspects above.

[0014] The beneficial effects of this application are as follows: In summary, the embodiments of this application provide a method, apparatus, and storage medium for reconstructing an image. The method includes: using a trained target reconstruction network to infer from an initial image to obtain a reconstructed second image. The target reconstruction network includes a feature extraction module, an asymmetric attention module, a quality-aware gating module, and a reconstruction module connected sequentially. The second image contains more high-frequency image features of a first target object than the initial image. The asymmetric attention module is obtained by adjusting the attention map of the original reconstruction network during training on a first historical image set and a second historical image set using first prior features. The quality-aware gating module is obtained by adjusting the nonlinear activation of the original reconstruction network during training on the first historical image set and the second historical image set using first prior features. The function yields a first historical image set, which is a collection of large-field-of-view images including the second target object, and a second historical image set, which is a collection of multiple small-field-of-view images including the second target object. The first prior feature is obtained by training the first historical image set using the first prior feature extraction network. The high-frequency image features of the second target object represented by the first prior feature in the first historical image set include the high-frequency image features of the second target object represented by the second prior feature in the second historical image set. The second prior feature is obtained by training the second historical image set using the second prior feature extraction network. The above-mentioned method adaptively learns the mapping from low-resolution images to high-resolution images through deep learning, that is, the target reconstruction network can achieve high-quality image reconstruction under unknown degradation conditions.

[0015] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the system architecture for image reconstruction in an embodiment of this application; Figure 2 This is a schematic diagram illustrating the determination of the first prior feature and the second prior feature in an embodiment of this application; Figure 3 This is a schematic diagram of an image reconstruction process in an embodiment of this application; Figure 4 This is a schematic diagram of the target reconstruction network trained in the embodiments of this application; Figure 5This is a schematic diagram of the process of obtaining attention features based on the first prior feature in an embodiment of this application; Figure 6 This is a schematic diagram of the process for obtaining quality characteristics through nonlinear functions in an embodiment of this application; Figure 7 This is a schematic diagram of the logical architecture of an image reconstruction system according to an embodiment of this application; Figure 8 This is a schematic diagram of the physical architecture of the cloud server in the embodiments of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0018] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in sequences other than those illustrated or described herein.

[0019] First, the technical terms used in this application are explained as follows: Artificial Intelligence (AI): A science and technology that simulates human intelligence.

[0020] Deep Learning (DL): A machine learning technique based on deep neural networks.

[0021] Computer Vision (CV): Artificial intelligence technology that analyzes and understands images and videos.

[0022] Printed Circuit Board (PCB): A basic carrier used for assembling and electrically connecting electronic components.

[0023] Programmable Logic Controller (PLC): A specialized computing device for industrial automation control, used to perform logical operations, timing control, and signal processing on mechanical equipment in the production process.

[0024] The preferred embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0025] See Figure 1 As shown, this embodiment includes at least one cloud server 600, which integrates an image reconstruction method. That is, in this embodiment, the image reconstruction method is primarily executed on the cloud server 600 side, as detailed below.

[0026] In this embodiment of the application, a specific process for reconstructing an image is as follows: The trained target reconstruction network is used to reason about the initial image to obtain the reconstructed second image. The target reconstruction network includes a feature extraction module, an asymmetric attention module, a quality-aware gating module, and a reconstruction module connected in sequence. The second image contains more high-frequency image features of the first target object than the initial image.

[0027] In this embodiment, artificial intelligence and deep learning methods are first used to obtain a target reconstruction network for industrial processes. This target reconstruction network includes a feature extraction module, an asymmetric attention module, a quality-aware gating module, and a reconstruction module connected in sequence. Unlike related technologies, the asymmetric attention module and the quality-aware gating module also reference prior features to achieve a mapping from low-resolution pixels to high-resolution pixels in the image. Thus, in industrial processes, after inference on the initial image using the trained target reconstruction network, a reconstructed second image can be obtained. Furthermore, the second image contains more high-frequency image features of the first target object than the initial image. In other words, the image after inference by the target reconstruction network has more high-resolution pixels than the image before inference, thereby enabling the inferred image to better achieve industrial goals such as anomaly detection and target recognition. It should be noted that the aforementioned first target object is the target that needs to be detected and identified in the industrial environment during the inference process, such as a machine or a pipeline. To detect and identify the first target object, the first target object is usually a large-field-of-view image including its entire outline.

[0028] Next, the origins of the aforementioned feature extraction module, asymmetric attention module, quality-aware gating module, and reconstruction module will be introduced as follows.

[0029] The asymmetric attention module is obtained by adjusting the attention map of the original reconstruction network during the training process of the first historical image set and the second historical image set using the first prior features.

[0030] In related technologies, attention modules are usually trained directly on a set of historical images including the target object. However, in this embodiment, the set of historical images used for training is first subdivided into a first set of historical images and a second set of historical images.

[0031] The aforementioned first historical image set is a collection of wide-field-of-view images including the second target object. It should be noted that the second target object is the target that needs to be detected and identified in an industrial environment during the training process. The second target object may be the same as or different from the first target object. In order to perform comprehensive detection and identification of the second target object, each image in the aforementioned first historical image set is usually a wide-field-of-view image including the overall outline of the second target object. Due to the limited hardware conditions of industrial cameras and the shooting environment, the resolution of the aforementioned wide-field-of-view images is usually low. Consequently, the details of the second target object shown in the wide-field-of-view images may not be clear.

[0032] The aforementioned second historical image set is a collection of multiple small-field-of-view images including the second target object. Similarly, the second target object is the target that needs to be detected and identified in an industrial environment during the training process. This second target object may or may not be the same as the first target object. To perform detailed detection and identification of the second target object, each image in the aforementioned second historical image set is typically a small-field-of-view image including different details of the second target object. Under the same industrial camera and shooting environment, the resolution of these small-field-of-view images is usually higher, meaning they contain more high-frequency image features. Compared to large-field-of-view images, the details of the second target object shown in the small-field-of-view images are clearer.

[0033] In this embodiment, the setting of a first historical image set and a second historical image set enables compensation for the high-frequency image features of the second target object. Specifically, high-resolution pixels in a small-field-of-view image are mapped to a large-field-of-view image through knowledge distillation. In the specific implementation, a second prior feature extraction network is first used to train the second historical image set to obtain second prior features. Then, a first prior feature extraction network is used to train the first historical image set to obtain first prior features. During the training process of the first prior feature extraction network on the first historical image set, the second prior features are applied to the training process of the second prior feature extraction network through knowledge distillation. This ensures that the high-frequency image features of the second target object represented by the first prior features in the first historical image set include the high-frequency image features of the second target object represented by the second prior features in the second historical image set, thus achieving the conversion from low-resolution to high-resolution images.

[0034] The following section will explain in detail how the first and second prior features are trained.

[0035] The aforementioned second prior feature is obtained through training in the following manner: [1] Multiple small-field images of the second target object from the second historical image set are input into the second prior feature extraction network for training.

[0036] See Figure 2 As shown, the second prior feature extraction network in this embodiment includes a reshaping layer, a linear layer, a first normalization layer, a multi-head attention layer, a second normalization layer, a multilayer perceptron, and a pooling layer connected in sequence. During training, after obtaining multiple small-field images including the second target object, these multiple small-field images are referred to as the second historical image set, and each small-field image in the second historical image set is input into the second prior feature extraction network one by one.

[0037] [2] The second high-frequency image features of the second target object included in each small field-of-view image are extracted using the second prior feature extraction network, and the second high-frequency image features are provided as the second prior features to the first prior feature extraction network.

[0038] When the small field-of-view image is input into the second prior feature extraction network, it is processed sequentially by the reshaping layer, linear layer, first normalization layer, multi-head attention layer, second normalization layer, multilayer perceptron, and pooling layer to extract features from the small field-of-view image. The steps of the first normalization layer, multi-head attention layer, second normalization layer, and multilayer perceptron processing the features output by the linear layer need to be repeated N times to obtain the high-resolution information of the second target object included in each small field-of-view image, that is, the second high-frequency image features.

[0039] Since the large-field-of-view images in the first historical image set and the small-field-of-view images in the second historical image set describe the second target object from different perspectives, the aforementioned second high-frequency image features correspond to a certain part of the large-field-of-view image. Based on this, the second high-frequency image features are provided as second prior features to the first prior feature extraction network so as to map the aforementioned second high-frequency image features to the large-field-of-view image, thereby enabling the large-field-of-view image obtained after training to include more high-resolution pixels.

[0040] The first prior feature mentioned above is obtained through training in the following way: (1) Input the large-view image of the second target object in the first historical image set into the first prior feature extraction network for training.

[0041] First, it should be noted that the reference... Figure 2As shown, the first prior feature extraction network has the same structure as the second prior feature extraction network, that is, the first prior feature extraction network is also composed of a reshaping layer, a linear layer, a first normalization layer, a multi-head attention layer, a second normalization layer, a multilayer perceptron and a pooling layer connected in sequence.

[0042] During training, after obtaining multiple large-field-of-view images including the second target object, these multiple large-field-of-view images are referred to as the first historical image set, and each large-field-of-view image in the first historical image set is input into the first prior feature extraction network.

[0043] (2) Use the first prior feature extraction network to extract the first high-frequency image features of the second target object included in the large field of view image.

[0044] When a large field-of-view image is input into the first prior feature extraction network, the above-mentioned reshaping layer, linear layer, first normalization layer, multi-head attention layer, second normalization layer, multilayer perceptron, and pooling layer sequentially process the large field-of-view image for feature extraction. The steps of the first normalization layer, multi-head attention layer, second normalization layer, and multilayer perceptron processing the features output by the linear layer need to be repeated N times to obtain the high-resolution information of the second target object included in each large field-of-view image, i.e., the first high-frequency image features.

[0045] (3) For different positions of the second target object represented in the large field of view image, the first high-frequency image features and the second prior features are fused to obtain the first prior features.

[0046] Considering that the second high-frequency image features are provided as the second prior features to the first prior feature extraction network, and that the large field-of-view image obtains its own first high-frequency image features through the first prior feature extraction network, but the aforementioned first high-frequency image features and second high-frequency image features may be different parts of the second target object, based on this, during the implementation process, the aforementioned first high-frequency image features and second high-frequency image features are analyzed separately to clarify which position of the second target object is represented by the first high-frequency image features and second high-frequency image features respectively.

[0047] When the positions of the second target objects represented by the first high-frequency image features and the second high-frequency image features are the same, either the first high-frequency image features or the second high-frequency image features can be used to represent the position. When the positions of the second target objects represented by the first high-frequency image features and the second high-frequency image features are not the same, the position is represented by the first high-frequency image feature or the second high-frequency image feature with higher resolution at that position. For example, the low-resolution position in the large field-of-view image is replaced with the second prior feature, thereby realizing the fusion of the first high-frequency image features and the second prior feature in the large field-of-view image, thus obtaining the first prior feature. In the implementation process, the limit of the above fusion can be limited by the distillation loss function, specifically by the following formula (1), so that the high-resolution pixels in the small field-of-view image are mapped to the large field-of-view image through knowledge distillation.

[0048] Formula (1)

[0049] in, This represents the second prior feature. Indicates the first prior feature. This represents the number of features in the second prior feature and the first prior feature. The number of channels for the second prior feature and the first prior feature is represented by , i represents the horizontal index of the second prior feature and the first prior feature, and j represents the vertical index of the second prior feature and the first prior feature.

[0050] During implementation, after obtaining the first prior features, the first prior features are added to the training process of the attention map in the original reconstruction network. In addition, the training process also incorporates the features of the first historical image set, thus obtaining a more complete asymmetric attention module.

[0051] The quality-aware gating module is obtained by adjusting the nonlinear activation function of the original reconstruction network during the training process of the first historical image set and the second historical image set using the first prior features.

[0052] In related technologies, quality perception modules are usually trained directly on a set of historical images including the target object. However, in this embodiment, after obtaining the first prior features, the first prior features are added to the training process of the nonlinear activation function in the original reconstruction network. In addition, the training process also incorporates the features of the first historical image set, thereby obtaining a more complete quality perception gating module.

[0053] The following section details how the target reconstruction network is trained. (See attached document.) Figure 3 As shown, the target reconstruction network is trained in the following way: Step 101: Use the original feature extraction module included in the original reconstruction network to extract features from each of the first historical images in the first historical image set to obtain the original features.

[0054] To obtain the target reconstruction network, during training, after obtaining multiple first historical images that include the second target object, these multiple first historical images are referred to as the first historical image set. Here, the first historical images are large-view images that include the second target object. First, the original feature extraction module included in the original reconstruction network is used to extract features from each first historical image, thereby obtaining multiple original features. The specific methods of feature extraction are not limited here.

[0055] Step 102: Train the original features and the first prior features using the original asymmetric attention module included in the original reconstruction network to obtain attention features.

[0056] In related technologies, the original asymmetric attention module includes only one linear layer, which generates attention features based on the original features and uses the attention features to train the attention map.

[0057] See Figure 4 As shown, the original asymmetric attention module in this embodiment includes two linear layers, Layer 1 and Layer 2, executed in parallel. After obtaining the original features, the original features and the first prior features obtained above are input into the original asymmetric attention module included in the original reconstruction network for training. The attention features are generated by Linear Layer 1 and Linear Layer 2. (See also...) Figure 4 and Figure 5 As shown, the steps for generating attention features specifically include: Step 1021: Train the original features using the original asymmetric attention module included in the original reconstruction network to obtain the value vector and key vector.

[0058] After the original features are input into the original asymmetric attention module included in the original reconstruction network, the original features are trained by the linear layer 1 included in the original asymmetric attention module to obtain the value vector and the key vector.

[0059] Step 1022: Train the first prior feature using the original asymmetric attention module to obtain the query vector.

[0060] After the first prior feature is input into the original asymmetric attention module included in the original reconstruction network, the first prior feature is trained by the second linear layer included in the original asymmetric attention module to obtain the query vector.

[0061] Step 1023: Perform matrix multiplication on the query vector and the transposed key vector to obtain the attention map, and perform matrix multiplication on the attention map and the value vector to obtain the attention features.

[0062] During training, the key vector is transposed and matrix multiplication is performed on the query vector and the transposed key vector to obtain the attention map. After obtaining the attention map, matrix multiplication is performed on the attention map and the value vector to obtain the attention features. The attention features can be calculated using the following formula (2). The attention features include attention weights and corresponding high-frequency image features. The attention weights can guide the model to focus on key information in the image.

[0063] Formula (2)

[0064] in, This represents the softmax operation. The attention feature is represented by Q, where Q represents the query vector, K represents the key vector, V represents the value vector, and d represents the number of channels for the query vector, key vector, and value vector. Note that the number of channels for the query vector, key vector, and value vector are all the same.

[0065] Step 103: Use the original quality-aware gating module included in the original reconstruction network to train the attention features and the first prior features to obtain the quality features.

[0066] See Figure 4 and Figure 6 As shown, the specific steps for obtaining quality features through the above training include: Step 1031: Input the attention features and the first prior features into the original quality-aware gating module included in the original reconstruction network.

[0067] During training, after obtaining the attention features, the attention features and the aforementioned first prior features are further provided to the connected original quality perception gating module. The original quality perception gating module first extracts features from the attention features through the feature extraction unit, and then provides the extracted attention features to the nonlinear activation function.

[0068] Step 1032: Perform average pooling on the first prior feature in the spatial dimension to obtain the quality-aware gating factor.

[0069] Meanwhile, in order to better train the nonlinear activation function, the first prior features are averaged in the spatial dimension to obtain the quality-aware gating factor.

[0070] Step 1033: Modulate the nonlinear activation function using a quality-aware gating factor.

[0071] After obtaining the quality-aware gating factor, the aforementioned quality-aware gating factor can reflect the difference between the feature distribution and the target distribution, and modulate the nonlinear activation function through the quality-aware gating factor.

[0072] Step 1034: Apply the modulated nonlinear activation function to perform matrix addition on the attention features to obtain the quality features.

[0073] During training, the attention features are matrix-added using a modulated nonlinear activation function, and the result of the above operation is further extracted by feature extraction unit two to obtain quality features.

[0074] Specifically, the quality characteristics are obtained through the following formula (3).

[0075] Formula (3)

[0076] in, Indicates the first prior feature. Indicates quality characteristics, This represents the feature extraction layer in the quality-aware gating module. Represents a non-linear activation function. This represents the quality-perceived gating factor.

[0077] Step 104: Use the original reconstruction modules included in the original reconstruction network to train the quality features and obtain the trained historical images. After multiple rounds of iteration, the training of the original reconstruction network is considered to have converged when the total loss value of the historical images obtained in this round meets the convergence condition.

[0078] During training, after obtaining the quality features, the quality features are input into the original reconstruction module included in the original reconstruction network. The original reconstruction module is used to perform multiple rounds of iterative training on the quality features. After each round of iteration, a trained historical image is obtained. When the total loss value of the historical image obtained after a certain round of iteration meets the convergence condition, for example, if it is less than a certain preset threshold, the original reconstruction network training that meets the convergence condition is determined to be converged, and the converged original reconstruction network is determined as the target reconstruction network.

[0079] See Figure 4 As shown, after obtaining the target reconstruction network, the target reconstruction network can be used to infer the acquired initial image including the first target object. Here, the initial image is a large field-of-view image including the first target object, so that the reconstructed second image is obtained after inference. Compared with the initial image, the second image includes more high-frequency image features of the first target object.

[0080] Based on the same inventive concept, this application provides a system for reconstructing images, including a cloud server 600, which integrates any of the above-mentioned image reconstruction methods.

[0081] The cloud server 600 is configured to generate a reconstructed second image after performing a method to reconstruct the image based on the initial image.

[0082] In this embodiment of the application, the above-mentioned image reconstruction method is integrated into the cloud server 600. After the initial image is collected and input into the cloud server 600, the cloud server 600 uses the above-mentioned image reconstruction method to infer the initial image, thereby obtaining the reconstructed second image.

[0083] The cloud server 600 is also configured to detect the second image and obtain target information.

[0084] During implementation, after obtaining the second image, the cloud server 600 can further detect the second image, such as anomaly detection in industrial processes and target recognition, to obtain target information.

[0085] In addition, for the purpose of industrial automation, see [reference needed]. Figure 7 As shown, the system for reconstructing images in this embodiment of the application also includes: a data acquisition device 601, an industrial control device 602, an edge computing device 603, and a human-computer interaction device 604.

[0086] Both the data acquisition device 601 and the industrial control device 602 are connected to the edge computing device 603, which is connected to the cloud server 600. The cloud server 600 is also connected to the human-computer interaction device 604.

[0087] In order to achieve image acquisition, reconstruction and display, the data acquisition device 601 and the industrial control device 602 in the above system are both electrically connected to the edge computing device 603, the edge computing device 603 is electrically connected to the cloud server 600, and the cloud server 600 is also electrically connected to the human-computer interaction device 604, thereby facilitating image transmission, etc.

[0088] The data acquisition device 601 is configured to acquire a first image and send the first image to the edge computing device 603.

[0089] During implementation, the data acquisition device 601 can be an industrial camera, sensor, video camera, etc. The data acquisition device 601 is mainly used to acquire a large field-of-view image, i.e., a first image, of the first target object in the industrial environment, and further send the first image to the edge computing device 603 for processing.

[0090] The industrial control device 602 is configured to send control timing signals to the edge computing device 603 based on the acquired control signals.

[0091] The industrial control device 602 can be a PLC, a computer, or a transmission mechanism such as a cylinder. During implementation, the industrial control device 602 is used to collect control signals, such as timing signals and pulse signals, and generate a control timing sequence based on the control signals. The control timing sequence is then sent to the edge computing device 603 to control the frequency at which the edge computing device 603 processes the initial image.

[0092] The edge computing device 603 is configured to preprocess the first image under the control of the control timing, obtain an initial image after preprocessing, and send the initial image to the cloud server 600.

[0093] During implementation, the edge computing device 603 preprocesses the input first image under the control of the aforementioned timing sequence. This preprocessing specifically includes image data reception, image caching, image processing, timing synchronization, and edge inference. After the preprocessing, the edge computing device 603 converts the first image into an initial image and sends it to the cloud server 600 for reconstruction. The cloud server 600 then detects the reconstructed second image to obtain the target information.

[0094] The human-computer interaction device 604 is configured to display target information and trigger defect alarms.

[0095] During implementation, in order to enable relevant managers to understand the real-time status of a target object in the industrial process, the human-machine interaction device 604 can also display the target information on the screen in real time and trigger a defect alarm based on the target information.

[0096] Based on the same inventive concept, see [reference] Figure 8 As shown, this application embodiment provides a cloud server, including: a memory 801 for storing executable instructions; and a processor 802 for reading and executing the executable instructions stored in the memory, and executing any of the methods described in the first aspect above.

[0097] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor, enables the processor to perform the method described in any of the first aspects above.

[0098] In summary, the embodiments of this application provide a method, apparatus, and storage medium for reconstructing an image. The method includes: using a trained target reconstruction network to infer from an initial image to obtain a reconstructed second image. The target reconstruction network includes a feature extraction module, an asymmetric attention module, a quality-aware gating module, and a reconstruction module connected sequentially. The second image contains more high-frequency image features of a first target object than the initial image. The asymmetric attention module is obtained by adjusting the attention map of the original reconstruction network during training on a first historical image set and a second historical image set using first prior features. The quality-aware gating module is obtained by adjusting the nonlinear activation of the original reconstruction network during training on the first historical image set and the second historical image set using first prior features. The function yields a first historical image set, which is a collection of large-field-of-view images including the second target object, and a second historical image set, which is a collection of multiple small-field-of-view images including the second target object. The first prior feature is obtained by training the first historical image set using the first prior feature extraction network. The high-frequency image features of the second target object represented by the first prior feature in the first historical image set include the high-frequency image features of the second target object represented by the second prior feature in the second historical image set. The second prior feature is obtained by training the second historical image set using the second prior feature extraction network. The above-mentioned method adaptively learns the mapping from low-resolution images to high-resolution images through deep learning, that is, the target reconstruction network can achieve high-quality image reconstruction under unknown degradation conditions.

[0099] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program product systems. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product system implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0100] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program product systems according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0101] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0102] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0103] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for reconstructing an image, characterized in that, The method includes: The trained target reconstruction network is used to reason about the initial image to obtain the reconstructed second image. The target reconstruction network includes a feature extraction module, an asymmetric attention module, a quality-aware gating module, and a reconstruction module connected in sequence. The second image has more high-frequency image features of the first target object included in the initial image. The asymmetric attention module is obtained by adjusting the attention map of the original reconstruction network during the training process of the first historical image set and the second historical image set using the first prior features. The quality-aware gating module is obtained by adjusting the nonlinear activation function of the original reconstruction network during the training process of the first historical image set and the second historical image set using the first prior features. The first historical image set is a set of large-field-of-view images including the second target object, and the second historical image set is a set of multiple small-field-of-view images including the second target object. Wherein, the first prior feature is obtained by training the first historical image set using a first prior feature extraction network, and the high-frequency image features of the second target object represented by the first prior feature in the first historical image set include the high-frequency image features of the second target object represented by the second prior feature in the second historical image set, and the second prior feature is obtained by training the second historical image set using a second prior feature extraction network.

2. The method as described in claim 1, characterized in that, The target reconstruction network is trained in the following way: The original feature extraction module included in the original reconstruction network is used to extract features from each of the first historical images included in the first historical image set to obtain the original features; The original features and the first prior features are trained using the original asymmetric attention module included in the original reconstruction network to obtain attention features; The attention features and the first prior features are trained using the original quality-aware gating module included in the original reconstruction network to obtain quality features; The quality features are trained using the original reconstruction module included in the original reconstruction network to obtain trained historical images. The training of the original reconstruction network is considered to have converged when the total loss value of the historical images obtained in this round meets the convergence condition after multiple rounds of iteration.

3. The method as described in claim 2, characterized in that, The process of training the original features and the first prior features using the original asymmetric attention module included in the original reconstruction network to obtain attention features includes: The original features are trained using the original asymmetric attention module included in the original reconstruction network to obtain value vectors and key vectors; The query vector is obtained by training the first prior feature using the original asymmetric attention module. A matrix multiplication operation is performed on the query vector and the transposed key vector to obtain an attention map, and a matrix multiplication operation is performed on the attention map and the value vector to obtain the attention feature.

4. The method as described in claim 2, characterized in that, The process of training the attention features and the first prior features using the original quality-aware gating module included in the original reconstruction network to obtain quality features includes: The attention features and the first prior features are input into the original quality-aware gating module included in the original reconstruction network; The first prior feature is averaged and pooled in the spatial dimension to obtain the quality-aware gating factor. The nonlinear activation function is modulated using the quality-aware gating factor; The quality features are obtained by performing matrix addition on the attention features using a modulated nonlinear activation function.

5. The method as described in claim 1, characterized in that, The first prior feature is obtained through training in the following manner: The first historical image set includes a large-view image of the second target object, which is then input into the first prior feature extraction network for training. The first prior feature extraction network is used to extract the first high-frequency image features of the second target object included in the large field-of-view image; For different positions of the second target object represented in the large field-of-view image, the first high-frequency image features and the second prior features are fused to obtain the first prior features.

6. The method as described in claim 1, characterized in that, The second prior feature is obtained through training in the following manner: Multiple small-field images including the second target object from the second historical image set are input into the second prior feature extraction network for training. The second prior feature extraction network is used to extract the second high-frequency image features of the second target object included in each of the small field-of-view images, and the second high-frequency image features are provided as the second prior features to the first prior feature extraction network.

7. A system for reconstructing an image, characterized in that, include: A cloud server, wherein the image reconstruction method as described in any one of claims 1 to 6 is integrated; The cloud server is configured to generate a reconstructed second image after performing the image reconstruction method based on an initial image; The cloud server is also configured to detect the second image to obtain target information.

8. The system as described in claim 7, characterized in that, Also includes: Data acquisition devices, industrial control devices, edge computing devices, and human-computer interaction devices; The data acquisition device and the industrial control device are both connected to the edge computing device, the edge computing device is connected to the cloud server, and the cloud server is also connected to the human-computer interaction device. The data acquisition device is configured to acquire a first image and send the first image to the edge computing device; The industrial control device is configured to send control timing signals to the edge computing device based on the acquired control signals; The edge computing device is configured to preprocess the first image under the control of the control timing, obtain the initial image after preprocessing, and send the initial image to the cloud server; The human-computer interaction device is configured to display the target information and trigger a defect alarm.

9. A cloud server, characterized in that, include: Memory, used to store executable instructions; A processor for reading and executing executable instructions stored in the memory to implement the method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor, the processor is able to perform the method as described in any one of claims 1 to 6.