Brightness enhancement method and device of power battery x-ray image and electronic equipment

CN117689594BActive Publication Date: 2026-09-25SUNWODA MOBILITY ENERGY TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311763359.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2026-09-25
Estimated Expiration
2043-12-19

AI Technical Summary

Technical Problem

[0003]本申请提供了一种动力电池X射线图像的亮度增强方法、装置以及电子设备,以解决因动力电池X射线图像整体亮度较低而影响Overhang区域尺寸测量任务的准确性的技术问题

Benefits of technology

[0016]在本申请实施例中,采用了将动力电池X射线图像的训练数据集输入至原始编码器,以使上述原始编码器对上述训练数据集的特征向量进行提取,得到对应的第一特征向量,其中,上述训练数据集中包括两种不同亮度的原始X射线图像;使用注意力机制对上述第一特征向量进行处理,得到第二特征向量并将上述第二特征向量输入至原始解码器,以使上述原始解码器对上述第二特征向量进行解码,得到对应的目标图像,其中,上述目标图像的目标位置的明亮程度大于对应的上述原始X射线图像的对应位置的明亮程度,上述目标位置为上述目标图像的电芯区域;使用上述训练数据集对上述原始编码器进行训练,上述目标图像对上述原始解码器进行训练,以调整上述原始编码器中的参数和上述原始解码器中的参数,得到训练后的目标编码器和目标解码器;通过上述目标编码器、上述注意力机制以及上述目标解码器对待增强X射线图像进行图像亮度增强处理,得到对应的亮度增强图像的方法,由于在上述方法中,通过构建原始编码器和原始解码器,将数据集输入至原始编码器得到第一特征向量,再将第一特征向量融合注意力机制并输入至原始解码器,得到电芯区域的明亮程度得到增强的目标图像,最后使用训练数据集和目标图像对原始编码器和原始解码器进行训练,得到训练后的目标解码器和目标解码器,在调整待增强X射线图像的亮度值时,使用训练后的目标解码器、注意力机制以及目标解码器对待增强X射线图像的电芯区域进行明亮程度的调整,从而实现了在有效提升X射线图像的整体亮度之后,使得X射线图像更直观,可以减轻人工对极片Overhang区域进行目检的难度目的,进而解决了因动力电池X射线图像整体亮度较低而影响Overhang区域尺寸测量任务的准确性的技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117689594B_ABST
    Figure CN117689594B_ABST
Patent Text Reader

Abstract

The application relates to a brightness enhancement method and device of a power battery X-ray image and electronic equipment. The method comprises the following steps: inputting a training data set of a power battery X-ray image into an original encoder to enable the original encoder to extract a feature vector of the training data set, and obtaining a corresponding first feature vector; using an attention mechanism to process the first feature vector, obtaining a second feature vector, inputting the second feature vector into an original decoder, and obtaining a corresponding target image; using the training data set to train the original encoder, using the target image to train the original decoder, and obtaining a trained target encoder and a target decoder; and performing image brightness enhancement processing on a to-be-enhanced X-ray image through the target encoder, the attention mechanism and the target decoder, and obtaining a corresponding brightness enhancement image. The application solves the technical problem that the accuracy of an Overhang region size measurement task is affected due to the low overall brightness of a power battery X-ray image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, and electronic device for enhancing the brightness of X-ray images of power batteries. Background Technology

[0002] To ensure the safety of power batteries and prevent potential risks such as lithium plating and short circuits, the negative electrode sheet should cover the positive electrode sheet under the protection of the separator. This means the external dimensions of the negative electrode sheet should exceed the corresponding range of the positive electrode sheet. The area where the negative electrode sheet is slightly longer than the positive electrode sheet is called the overhang. Dimensional inspection of the overhang area is a crucial step in power battery production, typically performed using non-destructive X-ray imaging. However, during X-ray imaging, insufficient X-ray radiation dose, low X-ray tube voltage, or a thick cell under inspection can lead to reduced overall image brightness, weakened electrode contrast, and blurred detail contours. These issues negatively impact the efficiency and accuracy of subsequent overhang area dimensional measurements. Summary of the Invention

[0003] This application provides a method, apparatus, and electronic device for enhancing the brightness of X-ray images of power batteries, in order to solve the technical problem that the low overall brightness of X-ray images of power batteries affects the accuracy of Overhang region size measurement tasks.

[0004] In a first aspect, this application provides a method for enhancing the brightness of X-ray images of a power battery, comprising: inputting a training dataset of X-ray images of a power battery into an original encoder, so that the original encoder extracts feature vectors from the training dataset to obtain a corresponding first feature vector, wherein the training dataset includes two original X-ray images with different brightness; processing the first feature vector using an attention mechanism to obtain a second feature vector and inputting the second feature vector into an original decoder, so that the original decoder decodes the second feature vector to obtain a corresponding target image, wherein the brightness of a target location in the target image is greater than the brightness of a corresponding location in the original X-ray image, and the target location is the cell region of the target image; training the original encoder using the training dataset and training the original decoder using the target image to adjust the parameters in the original encoder and the original decoder, thereby obtaining a trained target encoder and target decoder; and performing image brightness enhancement processing on the X-ray image to be enhanced using the target encoder, the attention mechanism, and the target decoder to obtain a corresponding brightness-enhanced image.

[0005] Secondly, this application provides a brightness enhancement device for X-ray images of power batteries, comprising: a first input module, configured to input a training dataset of X-ray images of power batteries into an original encoder, so that the original encoder extracts feature vectors from the training dataset to obtain a corresponding first feature vector, wherein the training dataset includes two original X-ray images with different brightness; a second input module, configured to process the first feature vector using an attention mechanism to obtain a second feature vector and input the second feature vector into an original decoder, so that the original decoder decodes the second feature vector to obtain a corresponding target image, wherein the brightness of a target location in the target image is greater than the brightness of a corresponding location in the original X-ray image, and the target location is the cell region of the target image; a training module, configured to train the original encoder using the training dataset and train the original decoder using the target image to adjust the parameters in the original encoder and the original decoder, thereby obtaining a trained target encoder and target decoder; and a first processing module, configured to perform image brightness enhancement processing on the X-ray image to be enhanced using the target encoder, the attention mechanism, and the target decoder to obtain a corresponding brightness-enhanced image.

[0006] As an optional example, the above-mentioned device further includes: a second processing module, used to replace the fully connected layers of a pre-trained preset neural network model with 1×1 convolutional layers before inputting the training dataset of the power battery X-ray image into the original encoder, to obtain the above-mentioned original encoder.

[0007] As an optional example, the above apparatus further includes: a first building module for building and training the attention mechanism before processing the first feature vector using the attention mechanism; and a second building module for building the original decoder, wherein the original decoder consists of transposed convolutions.

[0008] As an optional example, the second input module includes: a first acquisition unit, used to acquire feature maps of each layer of the original encoder using the attention mechanism described above; a first processing unit, used to concatenate the feature maps of each layer along the channels, and perform downsampling and global average pooling to obtain a third feature vector with the same dimension as the first feature vector; a second processing unit, used to normalize the third feature vector to obtain a fourth feature vector; and a first calculation unit, used to multiply the fourth feature vector by the first feature vector point by point to obtain the second feature vector.

[0009] As an optional example, the training module includes: a construction unit for constructing a first discriminator and a second discriminator, wherein the first discriminator consists of a five-layer neural network for overall obfuscating the training dataset and the target image, and the second discriminator consists of a four-layer small neural network for local obfuscating the training dataset and the target image; a second calculation unit for performing inversion operations on the first adversarial optimization objective of the first discriminator and the second adversarial optimization objective of the second discriminator to obtain a first loss function and a second loss function; a third calculation unit for calculating a total loss function based on the first loss function and the second loss function; and a training unit for training the original encoder and the original decoder using the total loss function to obtain a trained target decoder and a target decoder.

[0010] As an optional example, the training module further includes: a second acquisition unit, configured to acquire a first data size of the training dataset and a second data size of the target image before performing an inversion operation on the first adversarial optimization objective of the first discriminator and the second adversarial optimization objective of the second discriminator to obtain a first loss function and a second loss function; a third acquisition unit, configured to acquire a first sample of the training dataset and a second sample of the target image; and a fourth calculation unit, configured to calculate the first adversarial optimization objective based on the first data size, the second data size, the first sample, and the second sample.

[0011] As an optional example, the training module further includes: a fourth acquisition unit, configured to acquire a first data size of the training dataset and a second data size of the target image before acquiring a third sample of the first local patch of the training dataset and a fourth sample of the second local patch of the target image; a fifth acquisition unit, configured to acquire a third sample of the first local patch of the training dataset and a fourth sample of the second local patch of the target image; and a fifth calculation unit, configured to calculate the second adversarial optimization objective based on the first data size, the second data size, the third sample, and the fourth sample.

[0012] As an optional example, the training module further includes: a first selection unit, configured to select multiple neighborhoods of the image center point of the training dataset to obtain the first local patch before obtaining the third sample of the first local patch of the training dataset and the fourth sample of the second local patch of the target image; and a second selection unit, configured to select multiple neighborhoods of the image center point of the target image to obtain the second local patch.

[0013] As an optional example, the third computation unit includes: a first processing subunit, configured to formally define a third loss function corresponding to the original decoder and the first discriminator based on the first loss function, and formally define a fourth loss function corresponding to the original decoder and the second discriminator based on the second loss function, wherein the gradients of the third loss function and the fourth loss function are less than 1; a second processing subunit, configured to formally define an overall difference loss function based on a first maximum mean difference, and formally define a local difference loss function based on a second maximum mean difference; and a third processing subunit, configured to weightedly sum the third loss function, the fourth loss function, the overall difference loss function, and the local difference loss function to obtain the total loss function, wherein the overall difference loss function and the local difference loss function are regularization terms.

[0014] Thirdly, this application provides a storage medium storing a computer program, wherein the computer program is executed by a processor to perform the above-described method for enhancing the brightness of X-ray images of power batteries.

[0015] Fourthly, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-described method for enhancing the brightness of a power battery X-ray image through the computer program.

[0016] In this embodiment, a training dataset of X-ray images of a power battery is input into a raw encoder, which extracts feature vectors from the training dataset to obtain a first feature vector. The training dataset includes two raw X-ray images with different brightness levels. An attention mechanism is used to process the first feature vector to obtain a second feature vector, which is then input into a raw decoder to decode the second feature vector and obtain a target image. The target image has a target location with a brightness greater than the corresponding location in the original X-ray image; the target location is the battery cell region of the target image. The training dataset is used to train the raw encoder, and the target image is used to train the raw decoder, adjusting the parameters in both the raw encoder and decoder to obtain a trained target encoder and decoder. The target encoder, attention mechanism, and target image are then used to train the target image. The decoder performs image brightness enhancement processing on the X-ray image to be enhanced, obtaining the corresponding brightness-enhanced image. In this method, an original encoder and decoder are constructed. The dataset is input into the original encoder to obtain a first feature vector. This first feature vector is then fused with an attention mechanism and input into the original decoder to obtain the enhanced target image of the cell region's brightness. Finally, the original encoder and decoder are trained using the training dataset and the target image to obtain the trained target decoder and target decoder. When adjusting the brightness value of the X-ray image to be enhanced, the trained target decoder, attention mechanism, and target decoder are used to adjust the brightness of the cell region in the X-ray image. This effectively improves the overall brightness of the X-ray image, making it more intuitive and reducing the difficulty of manual visual inspection of the electrode overhang region. This solves the technical problem of the low overall brightness of the power battery X-ray image affecting the accuracy of overhang region size measurement. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0020] Figure 1 This is a flowchart of an optional method for enhancing the brightness of an X-ray image of a power battery according to an embodiment of this application;

[0021] Figure 2 This is a flowchart illustrating the implementation of an optional brightness enhancement method for X-ray images of power batteries according to an embodiment of this application.

[0022] Figure 3 This is a schematic diagram of an optional method for enhancing the brightness of an X-ray image of a power battery according to an embodiment of this application;

[0023] Figure 4 This is a sample brightness enhancement effect diagram of an optional brightness enhancement method for X-ray images of power batteries according to an embodiment of this application;

[0024] Figure 5 This is a schematic diagram of an optional brightness enhancement device for X-ray images of a power battery according to an embodiment of this application;

[0025] Figure 6 This is a schematic diagram of an optional electronic device according to an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0028] According to a first aspect of the embodiments of this application, a method for enhancing the brightness of an X-ray image of a power battery is provided, optionally, as follows: Figure 1As shown, the above method includes:

[0029] S102, the training dataset of X-ray images of the power battery is input into the original encoder so that the original encoder can extract the feature vector of the training dataset to obtain the corresponding first feature vector. The training dataset includes two original X-ray images with different brightness.

[0030] S104, the first feature vector is processed using an attention mechanism to obtain a second feature vector, and the second feature vector is input into the original decoder so that the original decoder decodes the second feature vector to obtain the corresponding target image. The brightness of the target position in the target image is greater than the brightness of the corresponding position in the original X-ray image. The target position is the cell region of the target image.

[0031] S106, Use the training dataset to train the original encoder and the target image to train the original decoder, so as to adjust the parameters in the original encoder and the original decoder, and obtain the trained target encoder and target decoder.

[0032] S108 performs image brightness enhancement processing on the X-ray image to be enhanced through a target encoder, attention mechanism, and target decoder to obtain the corresponding brightness-enhanced image.

[0033] Optionally, the size detection of the region where the negative electrode sheet is slightly longer than the positive electrode sheet is a crucial step in the production of power batteries, typically performed using non-destructive X-ray imaging. However, during X-ray imaging, insufficient X-ray radiation dose, low X-ray tube voltage, or a thick cell under inspection can lead to problems such as reduced overall image brightness, weakened electrode contrast, and blurred detail contours. These issues negatively impact the efficiency and accuracy of subsequent overhang region size measurement tasks. Therefore, in this embodiment, unsupervised adversarial generative processing is used. The trained target decoder, attention mechanism, and real-time generation of corresponding brightness-enhanced images by the target decoder not only effectively improve overall image brightness but also effectively suppress drastic fluctuations in local image brightness. Specifically, existing power battery X-ray images stored in memory are acquired and used to construct a training dataset X. I The training set contains two different brightness levels of raw X-ray images: one darker and one with normal brightness. A raw encoder E(·) is constructed to acquire the X-ray images from the input training dataset. I The mapping to the implicit space yields the corresponding first eigenvector E(X). I ). (The training dataset X) I The input is fed into the original encoder E(·) to obtain the corresponding first feature vector E(X). IAn attention mechanism A(·) and a raw decoder G(·) are constructed. The attention mechanism A(·) acts as a constraint on the raw decoder, enabling it to focus more on key features while ignoring some useless information. The introduction of the attention mechanism can avoid generating unnatural or distorted parts in the enhanced image. The attention mechanism is fused with the first feature vector to obtain a second feature vector. The second feature vector has a higher dimension than the first feature vector. The second feature vector fused with the attention mechanism is input into the raw decoder, thereby generating a target image with adjusted brightness for the corresponding battery cell region while retaining high-quality feature information. Based on the target image, the brightness of the target image is adjusted from the training dataset X. I Samples are extracted from the image and trained iteratively multiple times. After reaching the maximum number of iterations or model convergence, the trained target decoder, attention mechanism, and target decoder are used to adjust the brightness of the cell region in the X-ray image to be enhanced, thereby enhancing the brightness of the X-ray image to be enhanced.

[0034] Optionally, in this embodiment, by constructing an original encoder and an original decoder, the dataset is input into the original encoder to obtain a first feature vector, and then the first feature vector is fused with an attention mechanism and input into the original decoder to obtain a target image. Finally, the original encoder, attention mechanism, and original decoder are trained using the training dataset and the target image to obtain a trained target decoder, attention mechanism, and target decoder. When enhancing the brightness value of the X-ray image to be enhanced, the trained target decoder, attention mechanism, and target decoder are used to enhance the brightness. This achieves the goal of making the X-ray image more intuitive after effectively improving the overall brightness of the X-ray image, reducing the difficulty of manually inspecting the overhang area of ​​the electrode, and thus solving the technical problem that the low overall brightness of the X-ray image of the power battery affects the accuracy of the overhang area size measurement task.

[0035] As an optional example, the above method further includes the following steps before inputting the training dataset of X-ray images of the power battery into the original encoder:

[0036] The original encoder is obtained by replacing the fully connected layers of the pre-trained preset neural network model with 1×1 convolutional layers.

[0037] Optionally, in this embodiment, an original encoder E(·) is constructed. Specifically, the original encoder E(·) consists of a pre-trained VGG-16 network. The VGG (Visual Geometry Group) deep convolutional neural network architecture uses consecutive small convolutional kernels and pooling layers to construct a deep neural network. The VGG16 network is composed of multiple convolutional layers and pooling layers stacked alternately, and finally, a fully connected layer is used for classification. This network initially has three fully connected layers, which are not used in this embodiment; therefore, the fully connected layers are removed and replaced with 1×1 convolutional layers. The original encoder E(·) is used to obtain data from the input training dataset X. I The mapping to the implicit space yields the corresponding first eigenvector E(X). I ).

[0038] As an alternative example, the above method also includes the following before processing the first feature vector using an attention mechanism:

[0039] Build and train the attention mechanism;

[0040] Construct the original decoder, which consists of transposed convolutions.

[0041] Optionally, in this embodiment, an attention mechanism A(·) and a raw decoder G(·) are constructed. Specifically, the attention mechanism A(·) acts as a constraint on the raw decoder G(·), enabling it to focus more intently on key features while ignoring some useless information. The introduction of the attention mechanism can prevent the generation of unnatural or distorted parts in the enhanced image. The raw decoder consists of five layers of transposed convolutions. Transposed convolution, also known as "deconvolution," represents the inverse process of convolution. It can restore the image size before convolution based on the kernel size and the output size, rather than restoring the original values. Thus, after inputting the second feature vector fused with the attention mechanism into the raw decoder, a target image with corresponding brightness enhancement can be generated while retaining high-quality feature information.

[0042] As an optional example, processing the first feature vector using an attention mechanism to obtain the second feature vector includes:

[0043] Use an attention mechanism to obtain the feature map of each layer of the original encoder;

[0044] The feature maps of each layer are concatenated along the channels, and then downsampled and global average pooling is performed to obtain a third feature vector with the same dimension as the first feature vector.

[0045] The third eigenvector is normalized to obtain the fourth eigenvector;

[0046] The second eigenvector is obtained by multiplying the fourth eigenvector by the first eigenvector point by point.

[0047] Optionally, in this embodiment, the attention mechanism fuses the first feature vector to obtain the second feature vector. Specifically, the attention mechanism obtains the feature maps of each layer of the original encoder E(·), concatenates the feature maps along the channels, and then performs downsampling and global average pooling to obtain the third feature vector A(X). I The first feature vector E(X) output by the original encoder E(·) is compared with the first feature vector E(X). I The dimensions are the same. Furthermore, the obtained third feature vector A(X) is... I Normalize the normalized fourth feature vector and then compare it with the first feature vector E(X) output by the original encoder E(·). I The Hadamard product is performed point-by-point, allowing different weights to be assigned to features at different positions, thus filtering features of varying importance. This yields a third feature vector that incorporates the attention mechanism, which serves as the input to the original decoder G(·). The Hadamard product is a type of matrix operation. If A = (aij) and B = (bij) are two matrices of the same order, and cij = aij × bij, then matrix C = (cij) is called the Hadamard product of A and B, or the fundamental product.

[0048] As an optional example, the original encoder is trained using the training dataset, and the original decoder is trained using the target images to adjust the parameters in the original encoder and decoder, resulting in the trained target encoder and target decoder, including:

[0049] Construct a first discriminator and a second discriminator. The first discriminator consists of a five-layer neural network and is used to confuse the training dataset and the target image as a whole. The second discriminator consists of a four-layer small neural network and is used to confuse the training dataset and the target image locally.

[0050] The first adversarial optimization objective of the first discriminator and the second adversarial optimization objective of the second discriminator are inverted to obtain the first loss function and the second loss function;

[0051] The total loss function is calculated based on the first loss function and the second loss function.

[0052] The original encoder and decoder are trained using the total loss function to obtain the trained target encoder and target decoder.

[0053] Optionally, in this embodiment, the original encoder and decoder are trained using a training dataset and a target image to obtain a trained target decoder and target decoder. Specifically, a first discriminator D1(·) is constructed, consisting of a five-layer neural network, used to overall confuse the real input image and the brightness-enhanced image. Based on the distribution differences, a first adversarial optimization objective for the first discriminator D1(·) is formally defined. A second discriminator D2(·) is constructed, consisting of a four-layer small neural network, used to locally confuse the real input image and the brightness-enhanced image. Similarly, based on the distribution differences, a second adversarial optimization objective for the second discriminator D2(·) is formally defined. To achieve the first adversarial optimization objective, it is necessary to inverse the first adversarial optimization objective to obtain the first loss function L of the first discriminator D1(·). D1 Similarly, to achieve the second adversarial optimization objective, it is necessary to inverse the second adversarial optimization objective to obtain the second loss function L of the second discriminator D2(·). D2 Based on the first loss function L D1 Second loss function L D2 Calculate the total loss function L total Based on the total loss function L total From the training dataset X I Samples are extracted from the original encoder E(·) and the original decoder G(·) are trained repeatedly through multiple iterations. After reaching the maximum number of iterations or the model converges, the trained target decoder and target decoder are obtained.

[0054] As an optional example, before inverting the first adversarial optimization objective of the first discriminator and the second adversarial optimization objective of the second discriminator to obtain the first loss function and the second loss function, the above method further includes:

[0055] Obtain the first data size of the training dataset and the second data size of the target images;

[0056] Obtain the first sample from the training dataset and the second sample from the target image;

[0057] Based on the first data size, the second data size, the first sample, and the second sample, the first adversarial optimization objective is calculated.

[0058] Optionally, in this embodiment, a first adversarial optimization objective is calculated based on the training dataset and the target image to measure the overall data distribution distance between the training dataset and the target image. Specifically, the training dataset X is obtained. I Data scale and generating enhanced image sets X O Data scale and the sample x for the corresponding dataset i x o Based on the above parameters, a first adversarial optimization objective is calculated to measure the overall data distribution distance between the training dataset and the target image.

[0059] As an optional example, before inverting the first adversarial optimization objective of the first discriminator and the second adversarial optimization objective of the second discriminator to obtain the first loss function and the second loss function, the above method further includes:

[0060] Obtain the first data size of the training dataset and the second data size of the target images;

[0061] Obtain the third sample of the first local patch in the training dataset and the fourth sample of the second local patch in the target image;

[0062] The second adversarial optimization objective is calculated based on the first data size, the second data size, the third sample, and the fourth sample.

[0063] Optionally, in this embodiment, a second adversarial optimization objective is calculated based on the first local patch and the second local patch to measure the overall data distribution distance between the first local patch and the second local patch. Specifically, the training dataset X is obtained. I Data scale and generate image set X O Data scale and the sample patch x obtained after local sampling and cropping of the corresponding dataset. i ′、x o Based on the above parameters, a second adversarial optimization objective is calculated to determine the data distribution distance between the real input local patch and the brightness-enhanced local patch.

[0064] As an optional example, before obtaining the third sample of the first local patch in the training dataset and the fourth sample of the second local patch in the target image, the above method further includes:

[0065] Multiple neighborhoods of the center point of the image in the training dataset are selected to obtain the first local patch;

[0066] Multiple neighborhoods of the center point of the target image are selected to obtain the second local patch.

[0067] Optionally, in this embodiment, the first loss function L is obtained by inverting the first adversarial optimization objective of the first discriminator D1(·) and the second adversarial optimization objective of the second discriminator D2(·). D1 Second loss function L D2Previously, a first adversarial optimization objective was needed to calculate, based on the training dataset and the target image, to measure the overall data distribution distance between the training dataset and the target image. Local sampling was performed on the real input image and the target image in the training dataset. The four neighborhoods of the image center point were selected, and four local patches were obtained, which were then used to obtain the first local patch and the second local patch. A second adversarial optimization objective was then calculated based on the first and second local patches to measure the overall data distribution distance between them.

[0068] As an optional example, the total loss function calculated based on the first loss function and the second loss function includes:

[0069] According to the formal definition of the first loss function, the third loss function corresponding to the first discriminator of the original decoder is defined, and according to the formal definition of the second loss function, the fourth loss function corresponding to the second discriminator of the original decoder is defined, wherein the gradients of the third loss function and the fourth loss function are less than 1;

[0070] The overall difference loss function can be formally defined based on the first maximum mean difference, and the local difference loss function can be formally defined based on the second maximum mean difference.

[0071] The total loss function is obtained by weighting and superimposing the third loss function, the fourth loss function, the overall difference loss function, and the local difference loss function, where the overall difference loss function and the local difference loss function are regularization terms.

[0072] Optionally, in this embodiment, the first loss function L D1 Second loss function L D2 Only one term in the equation relates to the original encoder E(·), which is responsible for generating the brightness-enhanced image. Furthermore, the optimization objective of the original encoder E(·) is opposite to that of the first discriminator D1(·) and the second discriminator D2(·), and there is no gradient vanishing. Therefore, a third loss function L corresponding to the original encoder E(·) and the first discriminator D1(·) and the second discriminator D2(·) can be formally defined. G1 and the fourth loss function L G2 Furthermore, during training, the third loss function L mentioned above needs to be adjusted. G1 and the fourth loss function L G1 The gradient is constrained to satisfy the Lipschitz continuity condition, ensuring that the norm of the gradient is less than 1. To prevent overall or local distortion in the generated target image, further constraints are needed. The maximum mean difference is used to measure the similarity between the generated target image and the real input image. This represents an unbiased estimate of the maximum mean difference measure function, where the similarity is... Based on the regenerable kernel Hilbert space H and kernel mapping function The calculated kernel mapping function A radial basis function kernel is used. Based on the maximum mean difference, the overall difference loss function L can be formally defined. MMD1 The four-neighbor local patches obtained by local sampling of the real input image and the target image are constrained. Similarly, the maximum mean difference is used to measure the similarity between the generated brightness-enhanced local patch and the real input local patch. Similarly, based on the maximum mean difference, the local difference loss function L can be formally defined. MMD2 The obtained third loss function L G1 Fourth loss function L G2 And the overall difference loss function L MMD1 Local difference loss function L MMD2 The total loss function L is obtained by weighting and summing the terms as regularization terms. total , used for gradient backpropagation to train the original encoder E(·) and the original decoder G(·).

[0073] To illustrate with an example, this application relates to a brightness enhancement method for X-ray images of power batteries, addressing issues such as low overall image brightness, decreased electrode contrast and unclear edges, and potential information loss leading to difficulty in detecting minor anomalies. The specific implementation process is as follows: Figure 2 As shown:

[0074] 1. Acquire existing X-ray images of power batteries from the memory and construct a training dataset X. I The training set contains two different levels of original X-ray images. Specifically, because the overhang region inside the power battery cannot be directly detected after hot pressing, X-ray fluoroscopic imaging is required to obtain detection data. Before training the network, the training dataset needs to be constructed, which should include X-ray images under two different brightness conditions: darker and normal brightness.

[0075] 2. Construct the original encoder E(·). The original encoder E(·) consists of a pre-trained VGG-16 network. Since this network initially has three fully connected layers, which are not used in this embodiment, the fully connected layers are removed and replaced with 1×1 convolutional layers. The original encoder E(·) is used to obtain the input training dataset X. I The mapping to the implicit space yields the corresponding reduced-dimensional first eigenvector E(X). I ).

[0076] 3. Construct and train an attention mechanism A(·). The attention mechanism A(·) acts as a constraint on the original decoder G(·), enabling it to focus more intently on key features while ignoring useless information. The introduction of the attention mechanism avoids generating unnatural or distorted parts in the enhanced image. Specifically, for example... Figure 3 As shown, the attention mechanism obtains the feature maps of each layer of the original encoder E(·), concatenates the feature maps along the channels, and then performs downsampling and global average pooling to obtain the feature vector A(X). I The first feature vector E(X) output by the original encoder E(·) is compared with the first feature vector E(X). I The dimensions are the same. Furthermore, for the obtained feature vector A(X)... I Normalization is performed, and the normalized feature vector is multiplied point by point by the first feature vector output by the original encoder E(·) using the Hadamard product. This allows different weights to be assigned to features at different positions, and features of different importance to be selected. This results in a high-dimensional second feature vector that incorporates the attention mechanism, which is then used as the input to the original decoder G(·).

[0077] 4. Construct the original decoder G(·), which consists of five transposed convolutional layers. Input the second feature vector of the attention mechanism into the original decoder G(·) to generate the corresponding brightness-enhanced image while retaining high-quality feature information.

[0078] 5. Construct the first discriminator D1(·), which consists of a five-layer neural network. This discriminator is used to confuse the real input image with the brightness-enhanced image. Based on the distribution differences, the first adversarial optimization objective of the first discriminator D1(·) is formally defined as:

[0079]

[0080] In formula (1), W1 is used to measure the overall data distribution distance between the real input image and the target image, where, These represent the input training dataset X respectively. I and generating enhanced image sets X O First data size and second data size, x i x o These are samples from the corresponding dataset.

[0081] To achieve the first adversarial optimization objective, it is necessary to invert formula (1), that is, to obtain the first loss function L of the final first discriminator D1(·). D1 The formal definition is:

[0082]

[0083] 6. To prevent local distortion in the generated enhanced image, local sampling is performed on the real input image and the target image. The four neighborhoods of the center point of the image are selected and cropped to obtain 4 local patches.

[0084] A second discriminator D2(·) is constructed, consisting of a four-layer small neural network. It is defined to confuse the local real input image with the brightness-enhanced image. Similarly, based on the distribution difference, the second adversarial optimization objective of the second discriminator D2(·) is formally defined as follows:

[0085]

[0086] In formula (3), W2 is used as a metric function to calculate the data distribution distance between the real input local patch and the brightness-enhanced local patch, where They represent the training dataset X respectively. I and generate image set X O Data size, x i ′、x o ′ represents a sample patch of the corresponding dataset that has undergone local sampling.

[0087] To achieve the second adversarial optimization objective, it is necessary to invert formula (3), that is, to obtain the second loss function L of the final second discriminator D2(·). D2 The formal definition is:

[0088]

[0089] Furthermore, the first loss function L D1 Second loss function L D2 Only one term in the equation relates to the original decoder G(·), which is responsible for generating the brightness-enhanced image. The optimization objective of the original decoder G(·) is opposite to that of the first discriminator D1(·) and the second discriminator D2(·), and there is no gradient vanishing. Therefore, a third loss function L corresponding to the original decoder G(·) and the first discriminator D1(·) and the second discriminator D2(·) can be formally defined. G1 and the fourth loss function L G2 :

[0090]

[0091]

[0092] Furthermore, during the training process, the aforementioned third loss function L needs to be adjusted. G1 and the fourth loss function L G2 The gradient is constrained to satisfy the Lipshitz continuity condition, that is, the norm of the gradient is ensured to be less than 1 through gradient clipping.

[0093] Although the generated enhanced image and the original input image are confused as much as possible after adversarial training, additional restrictions need to be placed on the original decoder G(·) used to generate the enhanced image in order to prevent overall or local distortion of the generated target image.

[0094] Specifically, an appropriate metric function is selected to calculate the distribution distance between the generated enhanced image and the original input image, and this distribution distance is ensured to be small during training. In this embodiment, the maximum mean difference metric function is chosen. This metric function is a non-parametric method that does not require any prior assumptions about the data distribution; the calculation depends only on the mean and covariance of the samples. A kernel mapping function is mainly used to map two probability distributions to the same feature space and calculate the difference between their means in that space. The choice of kernel function should consider factors such as computational efficiency and the smoothness of the kernel function. The main idea of ​​the maximum mean difference metric is that if two distributions are completely identical, their moments of any order should be the same; if they are not identical, the moment with the largest difference is used as the standard to measure their difference. Specifically, given any two probability distributions P and Q, the distribution distance is calculated as follows:

[0095]

[0096] In formula (7), H represents the regenerating kernel Hilbert space. The kernel mapping function maps two distributions to a high-dimensional feature space, resulting in the corresponding second eigenvector representation.

[0097] Furthermore, the similarity between the generated target image and the real input image is measured by the maximum mean difference:

[0098]

[0099] In the above formula (8) S in formula (7) H An unbiased estimate. Let H denote the maximum mean difference measure function, and let H denote the reproducing kernel Hilbert space. This represents the kernel mapping function; in this embodiment, the radial basis kernel function is selected.

[0100] Based on the maximum mean difference, the overall difference loss function L can be formally defined. MMD1 :

[0101]

[0102] Furthermore, it is necessary to constrain the four-neighbor local patches after local sampling of the real input image and the target image. Similarly, the similarity between the generated brightness-enhanced local patch and the real input local patch is measured by the maximum mean difference.

[0103]

[0104] Based on the maximum mean difference, the local difference loss function L can be formally defined. MMD2 :

[0105]

[0106] The obtained third loss function L G1 Fourth loss function L G2 And the overall difference loss function L MMD1 Local difference loss function L MMD2 The total loss function L is obtained by weighting and summing the terms as regularization terms. total , used for gradient backpropagation to train the original encoder E(·) and the original decoder G(·), defining the total loss function L in each round. total for:

[0107] L total =(L G1 +L G2 )+β(L MMD1 +L MMD2 (12)

[0108] In formula (12), β represents the loss function balancing hyperparameter, which is set to 1-5 according to the relative weight relationship of each loss function.

[0109] 7. From the training dataset X I Samples are extracted from the image, and after repeated iterations of training, once the maximum number of iterations is reached or the model converges, the trained target decoder, target attention mechanism, and target decoder are used to enhance the brightness of the X-ray image to be enhanced.

[0110] Optionally, such as Figure 4 The image shown is an illustration of the sample brightness enhancement effect. Figure 4 The first column shows the original input image to be enhanced. After the steps in this embodiment, the final image with enhanced brightness is obtained, as shown below. Figure 4 As shown in the second column.

[0111] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0112] According to another aspect of the embodiments of this application, a brightness enhancement device for X-ray images of power batteries is also provided, such as... Figure 5 As shown, it includes:

[0113] The first input module 502 is used to input the training dataset of X-ray images of the power battery into the original encoder so that the original encoder can extract the feature vector of the training dataset to obtain the corresponding first feature vector. The training dataset includes two original X-ray images with different brightness.

[0114] The second input module 504 is used to process the first feature vector using an attention mechanism to obtain a second feature vector and input the second feature vector into the original decoder so that the original decoder decodes the second feature vector to obtain the corresponding target image. The brightness of the target position in the target image is greater than the brightness of the corresponding position in the original X-ray image. The target position is the cell region of the target image.

[0115] Training module 506 is used to train the original encoder using the training dataset and the original decoder using the target image to adjust the parameters in the original encoder and the original decoder, so as to obtain the trained target encoder and target decoder.

[0116] The first processing module 508 is used to perform image brightness enhancement processing on the X-ray image to be enhanced through a target encoder, an attention mechanism, and a target decoder to obtain a corresponding brightness-enhanced image.

[0117] Optionally, the size detection of the region where the negative electrode sheet is slightly longer than the positive electrode sheet is a crucial step in the production of power batteries, typically performed using non-destructive X-ray imaging. However, during X-ray imaging, insufficient X-ray radiation dose, low X-ray tube voltage, or a thick cell under inspection can lead to problems such as reduced overall image brightness, weakened electrode contrast, and blurred detail contours. These issues negatively impact the efficiency and accuracy of subsequent overhang region size measurement tasks. Therefore, in this embodiment, unsupervised adversarial generative processing is used. The trained target decoder, attention mechanism, and real-time generation of corresponding brightness-enhanced images by the target decoder not only effectively improve overall image brightness but also effectively suppress drastic fluctuations in local image brightness. Specifically, existing power battery X-ray images stored in memory are acquired and used to construct a training dataset X. I The training set contains two different brightness levels of raw X-ray images: one darker and one with normal brightness. A raw encoder E(·) is constructed to acquire the X-ray images from the input training dataset. I The mapping to the implicit space yields the corresponding first eigenvector E(X). I ). (The training dataset X) I The input is fed into the original encoder E(·) to obtain the corresponding first feature vector E(X). I An attention mechanism A(·) and a raw decoder G(·) are constructed. The attention mechanism A(·) acts as a constraint on the raw decoder, enabling it to focus more on key features while ignoring some useless information. The introduction of the attention mechanism can avoid generating unnatural or distorted parts in the enhanced image. The attention mechanism is fused with the first feature vector to obtain a second feature vector. The second feature vector has a higher dimension than the first feature vector. The second feature vector fused with the attention mechanism is input into the raw decoder, thereby generating a target image with adjusted brightness for the corresponding battery cell region while retaining high-quality feature information. Based on the target image, the brightness of the target image is adjusted from the training dataset X. I Samples are extracted from the image and trained iteratively multiple times. After reaching the maximum number of iterations or model convergence, the trained target decoder, target attention mechanism, and target decoder are used to adjust the brightness of the cell region in the X-ray image to be enhanced, thereby enhancing the brightness of the X-ray image to be enhanced.

[0118] Optionally, in this embodiment, by constructing an original encoder and an original decoder, the dataset is input into the original encoder to obtain a first feature vector, and then the first feature vector is fused with an attention mechanism and input into the original decoder to obtain a target image. Finally, the original encoder, attention mechanism, and original decoder are trained using the training dataset and the target image to obtain a trained target decoder, target attention mechanism, and target decoder. When enhancing the brightness value of the X-ray image to be enhanced, the brightness value is adjusted using the trained target decoder, target attention mechanism, and target decoder. This achieves the goal of making the X-ray image more intuitive after effectively improving the overall brightness of the X-ray image, reducing the difficulty of manually inspecting the overhang area of ​​the electrode, and thus solving the technical problem that the low overall brightness of the power battery X-ray image affects the accuracy of the overhang area size measurement task.

[0119] As an optional example, the above-described apparatus further includes:

[0120] The second processing module is used to replace the fully connected layers of the pre-trained preset neural network model with 1×1 convolutional layers before inputting the training dataset of the power battery X-ray images into the original encoder, so as to obtain the original encoder.

[0121] Optionally, in this embodiment, an original encoder E(·) is constructed. Specifically, the original encoder E(·) consists of a pre-trained VGG-16 network. The VGG (Visual Geometry Group) deep convolutional neural network architecture uses consecutive small convolutional kernels and pooling layers to construct a deep neural network. The VGG16 network is composed of multiple convolutional layers and pooling layers stacked alternately, and finally, a fully connected layer is used for classification. This network initially has three fully connected layers, which are not used in this embodiment; therefore, the fully connected layers are removed and replaced with 1×1 convolutional layers. The original encoder E(·) is used to obtain data from the input training dataset X. I The mapping to the implicit space yields the corresponding first eigenvector E(X). I ).

[0122] As an optional example, the above-described apparatus further includes:

[0123] The first building module is used to build and train the attention mechanism before processing the first feature vector using the attention mechanism;

[0124] The second building block is used to build the original decoder, which consists of five layers of transposed convolutions.

[0125] Optionally, in this embodiment, an attention mechanism A(·) and a raw decoder G(·) are constructed. Specifically, the attention mechanism A(·) acts as a constraint on the raw decoder G(·), enabling it to focus more intently on key features while ignoring some useless information. The introduction of the attention mechanism can prevent the generation of unnatural or distorted parts in the enhanced image. The raw decoder consists of five layers of transposed convolutions. Transposed convolution, also known as "deconvolution," represents the inverse process of convolution. It can restore the image size before convolution based on the kernel size and the output size, rather than restoring the original values. Thus, after inputting the second feature vector fused with the attention mechanism into the raw decoder, a target image with corresponding brightness enhancement can be generated while retaining high-quality feature information.

[0126] As an optional example, the second input module includes:

[0127] The first acquisition unit is used to acquire feature maps of each layer of the original encoder using an attention mechanism;

[0128] The first processing unit is used to concatenate the feature maps of each layer along the channels, and perform downsampling and global average pooling to obtain a third feature vector with the same dimension as the first feature vector.

[0129] The second processing unit is used to normalize the third feature vector to obtain the fourth feature vector.

[0130] The first calculation unit is used to multiply the fourth eigenvector by the first eigenvector point by point to obtain the second eigenvector.

[0131] Optionally, in this embodiment, the attention mechanism fuses the first feature vector to obtain the second feature vector. Specifically, the attention mechanism obtains the feature maps of each layer of the original encoder E(·), concatenates the feature maps along the channels, and then performs downsampling and global average pooling to obtain the third feature vector A(X). I The first feature vector E(X) output by the original encoder E(·) is compared with the first feature vector E(X). I The dimensions are the same. Furthermore, the obtained third feature vector A(X) is... I Normalize the normalized fourth feature vector and then compare it with the first feature vector E(X) output by the original encoder E(·). IThe Hadamard product is performed point-by-point, allowing different weights to be assigned to features at different positions, thus filtering features of varying importance. This yields a third feature vector that incorporates the attention mechanism, which serves as the input to the original decoder G(·). The Hadamard product is a type of matrix operation. If A = (aij) and B = (bij) are two matrices of the same order, and cij = aij × bij, then matrix C = (cij) is called the Hadamard product of A and B, or the fundamental product.

[0132] As an optional example, the training module includes:

[0133] The building unit is used to build a first discriminator and a second discriminator. The first discriminator consists of a five-layer neural network and is used to confuse the training dataset and the target image as a whole. The second discriminator consists of a four-layer small neural network and is used to confuse the training dataset and the target image locally.

[0134] The second calculation unit is used to perform inversion operations on the first adversarial optimization objective of the first discriminator and the second adversarial optimization objective of the second discriminator to obtain the first loss function and the second loss function;

[0135] The third calculation unit is used to calculate the total loss function based on the first loss function and the second loss function;

[0136] The training unit is used to train the original encoder and decoder using the total loss function to obtain the trained target decoder and target decoder.

[0137] Optionally, in this embodiment, the original encoder and decoder are trained using a training dataset and a target image to obtain trained target decoders. Specifically, a first discriminator D1(·) is constructed, consisting of a five-layer neural network, used to overall confuse the real input image and the brightness-enhanced image. Based on the distribution differences, a first adversarial optimization objective for the first discriminator D1(·) is formally defined. A second discriminator D2(·) is constructed, consisting of a four-layer small neural network, used to locally confuse the real input image and the brightness-enhanced image. Similarly, based on the distribution differences, a second adversarial optimization objective for the second discriminator D2(·) is formally defined. To achieve the first adversarial optimization objective, it is necessary to inverse the first adversarial optimization objective to obtain the first loss function L of the first discriminator D1(·). D1 Similarly, to achieve the second adversarial optimization objective, it is necessary to inverse the second adversarial optimization objective to obtain the second loss function L of the second discriminator D2(·). D2 Based on the first loss function L D1 Second loss function LD2 Calculate the total loss function L total Based on the total loss function L total From the training dataset X I Samples are extracted from the original encoder E(·) and the original decoder G(·) are trained repeatedly through multiple iterations. After reaching the maximum number of iterations or the model converges, the trained target decoder and target decoder are obtained.

[0138] As an optional example, the training module also includes:

[0139] The second acquisition unit is used to acquire the first data size of the training dataset and the second data size of the target image before performing the inversion operation on the first adversarial optimization objective of the first discriminator and the second adversarial optimization objective of the second discriminator to obtain the first loss function and the second loss function.

[0140] The third acquisition unit is used to acquire the first sample of the training dataset and the second sample of the target image;

[0141] The fourth calculation unit is used to calculate the first adversarial optimization objective based on the first data scale, the second data scale, the first sample, and the second sample.

[0142] Optionally, in this embodiment, a first adversarial optimization objective is calculated based on the training dataset and the target image to measure the overall data distribution distance between the training dataset and the target image. Specifically, the training dataset X is obtained. I Data scale and generating enhanced image sets X O Data scale and the sample x for the corresponding dataset i x o Based on the above parameters, a first adversarial optimization objective is calculated to measure the overall data distribution distance between the training dataset and the target image.

[0143] As an optional example, the training module also includes:

[0144] The fourth acquisition unit is used to acquire the first data scale of the training dataset and the second data scale of the target image before acquiring the third sample of the first local patch of the training dataset and the fourth sample of the second local patch of the target image.

[0145] The fifth acquisition unit is used to acquire the third sample of the first local patch of the training dataset and the fourth sample of the second local patch of the target image;

[0146] The fifth calculation unit is used to calculate the second adversarial optimization objective based on the first data scale, the second data scale, the third sample, and the fourth sample.

[0147] Optionally, in this embodiment, a second adversarial optimization objective is calculated based on the first local patch and the second local patch to measure the overall data distribution distance between the first local patch and the second local patch. Specifically, the training dataset X is obtained. I Data scale and generate image set X O Data scale and the sample patch x obtained after local sampling and cropping of the corresponding dataset. i ′、x o Based on the above parameters, a second adversarial optimization objective is calculated to determine the data distribution distance between the real input local patch and the brightness-enhanced local patch.

[0148] As an optional example, the training module also includes:

[0149] The first selection unit is used to select multiple neighborhoods of the center point of the image in the training dataset to obtain the first local patch.

[0150] The second selection unit is used to select multiple neighborhoods of the center point of the target image to obtain a second local patch.

[0151] Optionally, in this embodiment, the first loss function L is obtained by inverting the first adversarial optimization objective of the first discriminator D1(·) and the second adversarial optimization objective of the second discriminator D2(·). D1 Second loss function L D2 Previously, a first adversarial optimization objective was needed to calculate, based on the training dataset and the target image, to measure the overall data distribution distance between the training dataset and the target image. Local sampling was performed on the real input image and the target image in the training dataset. The four neighborhoods of the image center point were selected, and four local patches were obtained, which were then used to obtain the first local patch and the second local patch. A second adversarial optimization objective was then calculated based on the first and second local patches to measure the overall data distribution distance between them.

[0152] As an optional example, the third computational unit includes:

[0153] The first processing subunit is used to formally define the third loss function corresponding to the original decoder and the first discriminator according to the first loss function, and to formally define the fourth loss function corresponding to the original decoder and the second discriminator according to the second loss function, wherein the gradients of the third loss function and the fourth loss function are less than 1;

[0154] The second processing subunit is used to formally define the overall difference loss function based on the first maximum mean difference, and to formally define the local difference loss function based on the second maximum mean difference.

[0155] The third processing subunit is used to weight and superimpose the third loss function, the fourth loss function, the overall difference loss function, and the local difference loss function to obtain the total loss function, where the overall difference loss function and the local difference loss function are regularization terms.

[0156] Optionally, in this embodiment, the first loss function L D1 Second loss function L D2 Only one term in the equation relates to the original encoder E(·), which is responsible for generating the brightness-enhanced image. Furthermore, the optimization objective of the original encoder E(·) is opposite to that of the first discriminator D1(·) and the second discriminator D2(·), and there is no gradient vanishing. Therefore, a third loss function L corresponding to the original encoder E(·) and the first discriminator D1(·) and the second discriminator D2(·) can be formally defined. G1 and the fourth loss function L G2 Furthermore, during training, the third loss function L mentioned above needs to be adjusted. G1 and the fourth loss function L G1 The gradient is constrained to satisfy the Lipschitz continuity condition, ensuring that the norm of the gradient is less than 1. To prevent overall or local distortion in the generated target image, further constraints are needed. The maximum mean difference is used to measure the similarity between the generated target image and the real input image. This represents an unbiased estimate of the maximum mean difference measure function, where the similarity is... Based on the regenerable kernel Hilbert space H and kernel mapping function The calculated kernel mapping function A radial basis function kernel is used. Based on the maximum mean difference, the overall difference loss function L can be formally defined. MMD1 The four-neighbor local patches obtained by local sampling of the real input image and the target image are constrained. Similarly, the maximum mean difference is used to measure the similarity between the generated brightness-enhanced local patch and the real input local patch. Similarly, based on the maximum mean difference, the local difference loss function L can be formally defined. MMD2 The obtained third loss function L G1 Fourth loss function L G2 And the overall difference loss function L MMD1 Local difference loss function L MMD2 The total loss function L is obtained by weighting and summing the terms as regularization terms.total , used for gradient backpropagation to train the original encoder E(·) and the original decoder G(·).

[0157] For other examples of this embodiment, please refer to the examples above, which will not be repeated here.

[0158] Figure 6 This is a schematic diagram of an optional electronic device according to an embodiment of this application, such as... Figure 6 As shown, it includes a processor 602, a communication interface 604, a memory 606, and a communication bus 608. The processor 602, communication interface 604, and memory 606 communicate with each other via the communication bus 608.

[0159] Memory 606 is used to store computer programs;

[0160] When processor 602 executes a computer program stored in memory 606, it performs the following steps:

[0161] The training dataset of X-ray images of power batteries is input into the original encoder so that the original encoder can extract the feature vector of the training dataset to obtain the corresponding first feature vector. The training dataset includes two original X-ray images with different brightness.

[0162] The first feature vector is processed using an attention mechanism to obtain a second feature vector, which is then input into the original decoder so that the original decoder can decode the second feature vector to obtain the corresponding target image. The brightness of the target position in the target image is greater than the brightness of the corresponding position in the original X-ray image, and the target position is the cell region of the target image.

[0163] The original encoder is trained using the training dataset, and the original decoder is trained using the target image to adjust the parameters in the original encoder and the original decoder, resulting in the trained target encoder and target decoder.

[0164] The target encoder, attention mechanism, and target decoder are used to perform image brightness enhancement processing on the X-ray image to be enhanced, resulting in the corresponding brightness-enhanced image.

[0165] Optionally, in this embodiment, the communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6The symbol is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0166] The memory may include RAM, or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0167] As an example, the memory 606 described above may include, but is not limited to, the first input module 502, the second input module 504, the training module 506, and the first processing module 508 from the brightness enhancement device for the X-ray image of the power battery described above. Furthermore, it may include, but is not limited to, other module units from the brightness enhancement device for the X-ray image of the power battery described above, which will not be elaborated upon in this example.

[0168] The processors mentioned above can be general-purpose processors, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; they can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0169] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.

[0170] Those skilled in the art will understand that Figure 6 The structure shown is for illustrative purposes only. The device that implements the above-mentioned method for enhancing the brightness of X-ray images of power batteries can be a terminal device, such as a smartphone (e.g., an Android phone, an iOS phone), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 6 This does not limit the structure of the aforementioned electronic devices. For example, the electronic device may also include components that are more... Figure 6 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 6 The different configurations shown.

[0171] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, ROM, RAM, disk or optical disk, etc.

[0172] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, which, when executed by a processor, performs the steps in the above-described method for enhancing the brightness of X-ray images of power batteries.

[0173] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0174] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0175] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0176] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0177] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.

[0178] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0179] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0180] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for enhancing the brightness of X-ray images of power batteries, characterized in that, include: The training dataset of X-ray images of power batteries is input into the original encoder so that the original encoder can extract the feature vector of the training dataset to obtain the corresponding first feature vector. The training dataset includes two original X-ray images with different brightness. The first feature vector is processed using an attention mechanism to obtain a second feature vector, and the second feature vector is input into the original decoder so that the original decoder decodes the second feature vector to obtain the corresponding target image. The brightness of the target position in the target image is greater than the brightness of the corresponding position in the original X-ray image. The target position is the cell region of the target image. The original encoder is trained using the training dataset, and the original decoder is trained using the target image to adjust the parameters in the original encoder and the original decoder, thereby obtaining the trained target encoder and target decoder. The target encoder, the attention mechanism, and the target decoder perform image brightness enhancement processing on the X-ray image to be enhanced, resulting in a corresponding brightness-enhanced image.

2. The method according to claim 1, characterized in that, The method further includes, before inputting the training dataset of X-ray images of the power battery into the original encoder: The original encoder is obtained by replacing the fully connected layers of the pre-trained preset neural network model with 1×1 convolutional layers.

3. The method according to claim 1, characterized in that, Before processing the first feature vector using the attention mechanism, the method further includes: Construct and train the attention mechanism; Construct the original decoder, wherein the original decoder consists of transposed convolutions.

4. The method according to claim 1, characterized in that, The step of processing the first feature vector using an attention mechanism to obtain the second feature vector includes: The attention mechanism is used to obtain the feature map of each layer of the original encoder; The feature maps of each layer are concatenated along the channels, and then downsampled and global average pooling is performed to obtain a third feature vector with the same dimension as the first feature vector. The third feature vector is normalized to obtain the fourth feature vector; The second feature vector is obtained by multiplying the fourth feature vector by the first feature vector point by point.

5. The method according to claim 1, characterized in that, The step of training the original encoder using the training dataset and training the original decoder using the target image to adjust the parameters in the original encoder and the original decoder, thereby obtaining the trained target encoder and target decoder, includes: Construct a first discriminator and a second discriminator, wherein the first discriminator consists of a five-layer neural network for overall obfuscating the training dataset and the target image, and the second discriminator consists of a four-layer small neural network for local obfuscating the training dataset and the target image; The first adversarial optimization objective of the first discriminator and the second adversarial optimization objective of the second discriminator are inverted to obtain the first loss function and the second loss function; The total loss function is calculated based on the first loss function and the second loss function; The original encoder and the original decoder are trained using the total loss function to obtain the trained target encoder and target decoder.

6. The method according to claim 5, characterized in that, Before performing inversion operations on the first adversarial optimization objective of the first discriminator and the second adversarial optimization objective of the second discriminator to obtain the first loss function and the second loss function, the method further includes: Obtain the first data size of the training dataset and the second data size of the target image; Obtain the first sample from the training dataset and the second sample from the target image; The first adversarial optimization objective is calculated based on the first data scale, the second data scale, the first sample, and the second sample.

7. The method according to claim 5, characterized in that, Before performing inversion operations on the first adversarial optimization objective of the first discriminator and the second adversarial optimization objective of the second discriminator to obtain the first loss function and the second loss function, the method further includes: Obtain the first data size of the training dataset and the second data size of the target image; Obtain the third sample of the first local patch of the training dataset and the fourth sample of the second local patch of the target image; The second adversarial optimization objective is calculated based on the first data scale, the second data scale, the third sample, and the fourth sample.

8. The method according to claim 7, characterized in that, Before acquiring the third sample of the first local patch in the training dataset and the fourth sample of the second local patch in the target image, the method further includes: The first local patch is obtained by selecting multiple neighborhoods of the center point of the image in the training dataset; The second local patch is obtained by selecting multiple neighborhoods of the center point of the target image.

9. The method according to claim 5, characterized in that, The step of calculating the total loss function based on the first loss function and the second loss function includes: According to the formal definition of the first loss function, the original encoder corresponds to the third loss function and the first discriminator. According to the formal definition of the second loss function, the original decoder corresponds to the second discriminator. The gradients of the third loss function and the fourth loss function are less than 1. The overall difference loss function can be formally defined based on the first maximum mean difference, and the local difference loss function can be formally defined based on the second maximum mean difference. The total loss function is obtained by weighting and superimposing the third loss function, the fourth loss function, the overall difference loss function, and the local difference loss function, wherein the overall difference loss function and the local difference loss function are regularization terms.

10. A brightness enhancement device for X-ray images of a power battery, characterized in that, include: The first input module is used to input the training dataset of X-ray images of power batteries into the original encoder, so that the original encoder can extract the feature vector of the training dataset to obtain the corresponding first feature vector. The training dataset includes two original X-ray images with different brightness. The second input module is used to process the first feature vector using an attention mechanism to obtain a second feature vector and input the second feature vector into the original decoder so that the original decoder decodes the second feature vector to obtain the corresponding target image, wherein the brightness of the target position in the target image is greater than the brightness of the corresponding position in the original X-ray image, and the target position is the cell region of the target image; The training module is used to train the original encoder using the training dataset and to train the original decoder using the target image, so as to adjust the parameters in the original encoder and the parameters in the original decoder to obtain the trained target encoder and target decoder. The first processing module is used to perform image brightness enhancement processing on the X-ray image to be enhanced through the target encoder, the attention mechanism, and the target decoder to obtain the corresponding brightness-enhanced image.

11. A computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to perform the method described in any one of claims 1 to 9.

12. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 9 through the computer program.

Citation Information

Patent Citations

  • Depth field adaptive image classification method based on prototype network

    CN114611617A

  • Systems and methods with integrated gaming engines and smart contracts

    US20230201722A1