Multi-view model training method and apparatus, and storage medium
Patent Information
- Application Number
- CN202310133942.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-07
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-02-07
AI Technical Summary
[0004]本申请提供一种多视图模型训练方法、装置及存储介质,能够解决识别异常数据的准确性较低的问题
Smart Images

Figure CN118470450B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a multi-view model training method, apparatus and storage medium. Background Technology
[0002] Currently, with the increasing demand for automation and intelligent upgrading of industrial production lines, automated inspection of product quality is becoming increasingly important. Existing technologies utilize deep learning algorithms for supervised training with large amounts of labeled data to classify and detect various anomalies.
[0003] However, the above methods require a large amount of data for supervised training of deep learning models. But the frequency of abnormal data in industrial products is often low, and supervised data is usually difficult to obtain and collect. As a result, the amount of data is often insufficient to support a sufficient supervised classification training method to train deep neural networks to identify abnormal data. Thus, the accuracy of identifying abnormal data is low. Summary of the Invention
[0004] This application provides a multi-view model training method, apparatus, and storage medium, which can solve the problem of low accuracy in identifying abnormal data.
[0005] To achieve the above objectives, this application adopts the following technical solution:
[0006] In a first aspect, this application provides a multi-view model training method, which includes: acquiring a target image, wherein the target image is an image of an industrial component; performing reverse encoding on the target image based on a latent space and a reconstruction network to obtain a first latent space representation library corresponding to the latent space, wherein the first latent space representation library includes at least one first latent space representation, wherein the at least one latent space representation is used to characterize at least two first feature information with different weights corresponding to the target image; wherein, if the first latent space representation library has the highest matching degree with N second feature information libraries, determining whether there is a defect in the industrial component corresponding to the target image based on the first latent space representation library, wherein the N second feature information libraries are obtained through N preset neural networks, wherein the weights corresponding to the N preset neural networks are different, and N is an integer greater than or equal to 2; wherein, the latent space is used to indicate the consistency and complementarity among the N second feature information libraries.
[0007] Based on the above technical solution, the multi-view model training method provided in this application embodiment can fuse information from different views (i.e., N second feature information libraries) through latent space representation, thereby enabling the multi-view model to learn the completeness and structure of the feature information. Moreover, through reverse encoding, the multi-view model does not need to be trained with a large amount of supervised data. Instead, it compares the first latent space representation library with N second feature information libraries (i.e., positive samples). Thus, when the matching degree is the highest, it determines whether there is a defect in the industrial part corresponding to the target image based on the first latent space representation library. This avoids the problem of insufficient supervised data preventing the model from being fully trained, which leads to poor accuracy in identifying abnormal data. In this way, the accuracy of the multi-view model training device in identifying abnormal data is improved.
[0008] In the first possible implementation of the first aspect, before performing inverse encoding on the target image based on the latent space and the reconstruction network to obtain the first latent space representation library corresponding to the latent space, the method further includes: randomly initializing at least one second latent space representation of the latent space and the network parameters of the reconstruction network; performing inverse encoding on the target image based on the latent space and the reconstruction network to obtain the first latent space representation library corresponding to the latent space includes: updating at least one second latent space representation and the network parameters according to N second feature information libraries to obtain the first latent space representation library.
[0009] In a second possible implementation of the first aspect, the reconstruction network includes at least one reconstruction network layer, with different network parameters in each layer; updating at least one second latent space representation based on N second feature information libraries to obtain a first latent space representation library includes: inputting at least one first latent space representation into at least one reconstruction network layer to obtain N third feature information libraries, wherein the N third feature information libraries correspond one-to-one with the N second feature information libraries, and each third feature information library has a different weight; processing each third feature information in the N third feature information libraries based on the N second feature information libraries to obtain the first latent space representation library; wherein the network parameters in each reconstruction network layer are different.
[0010] In the third possible implementation of the first aspect, after obtaining the first latent space representation library, the method further includes: performing alternating iterative processing on the first latent space representation library and network parameters according to the objective function to obtain the converged first latent space representation library and converged network parameters.
[0011] In the fourth possible implementation of the first aspect, when the matching degree between the first latent space representation library and N second feature information libraries is the highest, determining whether the industrial component corresponding to the target image has a defect based on the first latent space representation library includes: acquiring third feature information in the first latent space representation library and acquiring fourth feature information in each of the second feature information libraries; processing the third feature information based on first other feature information to obtain a first metric distribution between the third feature information and the first other feature information, wherein the first other feature information is feature information other than the third feature information in the first latent space representation library, and the first metric distribution is used to characterize the similarity between the third feature information and the first other feature information; processing the fourth feature information based on second other feature information to obtain a second metric distribution between the fourth feature information and the second other feature information, wherein the second other feature information is feature information other than the third feature information in each of the second feature information libraries, and the second metric distribution is used to characterize the similarity between the fourth feature information and the second other feature information; and determining whether the industrial component corresponding to the target image has a defect based on the first latent space representation library when the difference between the first metric distribution and the second metric distribution is less than a preset threshold.
[0012] In the fifth possible implementation of the first aspect, after obtaining the target image, the method further includes: inputting the target image into N preset neural networks respectively to obtain N second feature information databases, wherein the weights of the N preset neural networks are different, and the N preset neural networks correspond one-to-one with the N second feature information databases.
[0013] Secondly, this application provides a multi-view model training device, comprising: an acquisition unit, a processing unit, and a determination unit. The acquisition unit is used to acquire a target image, wherein the target image is an image of an industrial component. The processing unit is used to perform reverse encoding on the target image based on a latent space and a reconstruction network to obtain a first latent space representation library corresponding to the latent space. The first latent space representation library includes at least one first latent space representation, which is used to characterize at least two first feature information with different weights corresponding to the target image. The determination unit is used to determine whether the industrial component corresponding to the target image has a defect, based on the first latent space representation library, when the matching degree between the first latent space representation library and N second feature information libraries is the highest. The N second feature information libraries are obtained through N preset neural networks, and the weights corresponding to the N preset neural networks are different, where N is an integer greater than or equal to 2; wherein the latent space is used to indicate the consistency and complementarity between at least one first feature information.
[0014] In a first possible implementation of the second aspect, the aforementioned multi-view model training device further includes an initialization unit. The initialization unit is used to randomly initialize at least one second latent space representation of the latent space and the network parameters of the reconstruction network before performing inverse encoding on the target image based on the latent space and the reconstruction network to obtain the first latent space representation library corresponding to the latent space. The processing unit is specifically used to update at least one second latent space representation and the network parameters according to N second feature information libraries to obtain the first latent space representation library.
[0015] In a second possible implementation of the second aspect, the aforementioned reconstruction network includes at least one reconstruction network layer, with different network parameters in each reconstruction network layer; the processing unit is specifically used to input at least one first latent space representation into at least one reconstruction network layer to obtain N third feature information libraries, the N third feature information libraries corresponding one-to-one with N second feature information libraries, and the weights corresponding to each third feature information library being different; and to process each third feature information in the N third feature information libraries according to the N second feature information libraries to obtain the first latent space representation library.
[0016] In a third possible implementation of the second aspect, the processing unit is further configured to, after obtaining the first latent space representation library, perform alternating iterative processing on the first latent space representation library and network parameters according to the objective function, to obtain the converged first latent space representation library and converged network parameters.
[0017] In the fourth possible implementation of the second aspect, the determining unit is specifically used to acquire third feature information in a first latent space representation library and fourth feature information in each second feature information library; and to process the third feature information according to first other feature information to obtain a first metric distribution between the third feature information and the first other feature information, wherein the first other feature information is feature information other than the third feature information in the first latent space representation library, and the first metric distribution is used to characterize the similarity between the third feature information and the first other feature information; and to process the fourth feature information according to second other feature information to obtain a second metric distribution between the fourth feature information and the second other feature information, wherein the second other feature information is feature information other than the third feature information in each second feature information library, and the second metric distribution is used to characterize the similarity between the fourth feature information and the second other feature information; and if the difference between the first metric distribution and the second metric distribution is less than a preset threshold, to determine whether there is a defect in the industrial component corresponding to the target image according to the first latent space representation library.
[0018] In the fifth possible implementation of the second aspect, the processing unit is further configured to, after acquiring the target image, input the target image into N preset neural networks respectively to obtain N second feature information databases, wherein the weights corresponding to the N preset neural networks are different, and the N preset neural networks correspond one-to-one with the N second feature information databases respectively.
[0019] Thirdly, this application provides a multi-view model training apparatus, which includes: a processor and a communication interface; the communication interface and the processor are coupled, and the processor is used to run computer programs or instructions to implement the multi-view model training method as described in the first aspect and any possible implementation of the first aspect.
[0020] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a terminal, cause the terminal to perform the multi-view model training method as described in the first aspect and any possible implementation thereof.
[0021] Fifthly, embodiments of this application provide a computer program product containing instructions that, when run on a multi-view model training device, cause the multi-view model training device to execute the multi-view model training method as described in the first aspect and any possible implementation thereof.
[0022] In a sixth aspect, embodiments of this application provide a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run computer programs or instructions to implement the multi-view model training method as described in the first aspect and any possible implementation thereof.
[0023] Specifically, the chip provided in this application embodiment also includes a memory for storing computer programs or instructions. Attached Figure Description
[0024] Figure 1 A flowchart illustrating a multi-view model training method provided in this application embodiment;
[0025] Figure 2 This is a schematic diagram of the structure of a multi-view model training system provided in an embodiment of this application;
[0026] Figure 3 A schematic diagram illustrating an example of a multi-view model training method provided in this application embodiment;
[0027] Figure 4 A schematic diagram of the structure of a multi-view model training device provided in an embodiment of this application;
[0028] Figure 5A schematic diagram of another multi-view model training device provided in this application embodiment;
[0029] Figure 6 This is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation
[0030] The multi-view model training method, apparatus, and storage medium provided in the embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0031] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0032] The terms "first" and "second," etc., used in the specification and drawings of this application are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of objects.
[0033] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.
[0034] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0035] In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0036] Currently, with the emergence of development models such as Industry 4.0, digitalization, and intelligent manufacturing, the demand for automation and intelligent upgrading of industrial production lines is constantly increasing. Automated inspection of product quality is an indispensable link, such as detecting abnormalities like appearance damage and displacement. Most industries, such as semiconductors and textiles, rely mainly on human experience for product quality inspection. With the development of machine vision, some anomaly detection algorithms have emerged. For example, traditional vision algorithms use operators from open-source libraries to achieve anomaly detection; deep learning algorithms use a large amount of labeled data for supervised training to classify and detect various anomalies; in the field of unsupervised anomaly detection, some algorithms use pre-trained networks to extract features, then build a database at the feature level. The established normal sample database and abnormal samples are compared through metric learning to obtain an anomaly score map, thereby achieving anomaly detection and localization.
[0037] However, the methods mentioned above rely on human experience to inspect product quality, such as in the semiconductor and textile industries. Prolonged work can lead to fatigue among inspectors, resulting in low efficiency and inspection errors. Traditional visual algorithms utilize operators from open-source libraries for anomaly detection, but these methods are unsuitable for industrial applications due to high computational resource consumption, long processing times, and the fact that some algorithms are not commercially viable. Supervised training of deep learning models requires a large amount of data, but anomaly data from industrial products often occurs infrequently, making it difficult to obtain and collect. The amount of data is often insufficient to support adequate supervised classification training for deep neural networks to identify anomalies. Unsupervised anomaly detection algorithms can learn complete image information of the object to be detected using features extracted from one or more pre-trained networks, achieving anomaly detection and localization through differential score map calculation, even in the absence of anomaly data. However, they do not fully utilize features, merely concatenating or performing simple calculations. The consistency and complementarity of features extracted from multiple pre-trained networks are not fully explored, and the feature fusion method is unscientific, resulting in low accuracy in identifying anomalies.
[0038] To address the problem of low accuracy in identifying anomalous data in existing technologies, this application provides a multi-view model training method. The multi-view model training device acquires a target image and performs inverse encoding on the target image based on latent space and a reconstruction network to obtain a first latent space representation library. Then, when the matching degree between the first latent space representation library and N second feature information libraries is the highest, the method determines whether the industrial component corresponding to the target image has a defect based on the first latent space representation library. In this scheme, information from different views (i.e., at least one latent space representation) can be fused through latent space representation, allowing the multi-view model to learn the completeness and structure of the representational feature information. Moreover, by using inverse encoding, the multi-view model does not require a large amount of supervised data for training. Instead, it compares the first latent space representation library with N second feature information libraries (i.e., positive samples), and determines whether the industrial component corresponding to the target image has a defect based on the first latent space representation library when the matching degree is the highest. This avoids the problem of insufficient supervised data preventing adequate model training and thus improving the accuracy of the multi-view model training device in identifying anomalous data.
[0039] like Figure 1 The diagram shows a flowchart of a multi-view model training method provided in an embodiment of this application. The method includes the following steps S101 to S103:
[0040] S101, The multi-view model training device acquires the target image.
[0041] In this embodiment of the application, the target image is an image of an industrial component.
[0042] Optionally, in this embodiment of the application, the target image can be one or more.
[0043] For example, the target image described above can be an image of a screw or a screw cap.
[0044] Optionally, in this embodiment of the application, the multi-view model training device may include a camera, thereby acquiring target images through the camera.
[0045] Optionally, in the embodiments of this application, the camera may include at least one of the following: a panoramic camera, a zoom camera, a macro camera, a wide-angle camera, and an ultra-wide-angle camera.
[0046] Optionally, in this embodiment of the application, the multi-view model training device can acquire target images sent by other devices (such as mobile phones and cameras) to obtain target images.
[0047] Optionally, in this embodiment of the application, the multi-view model training device can obtain the target image via SMS or a preset instant messaging application.
[0048] S102. The multi-view model training device performs reverse encoding on the target image based on the latent space and reconstruction network to obtain the first latent space representation library corresponding to the latent space.
[0049] In this embodiment of the application, the first latent space representation library includes at least one first latent space representation, which is used to characterize at least two first feature information with different weights corresponding to the target image. The latent space is used to indicate the consistency and complementarity among N second feature information libraries.
[0050] In this embodiment of the application, the multi-view model training device can obtain a first latent space representation library corresponding to the target image by reconstructing the network. Each latent space representation in the first latent space representation library can cover the effective information of at least two first feature information with different weights.
[0051] Optionally, in the embodiments of this application, the above-mentioned effective information may include at least one of the following: image distribution information of the target image, pixel information of the target image, color temperature information of the target image, and saturation information of the target image.
[0052] S103. When the matching degree between the first latent space representation library and the N second feature information libraries is the highest, the multi-view model training device determines whether there is a defect in the industrial component corresponding to the target image based on the first latent space representation library.
[0053] In this embodiment of the application, the above-mentioned N second feature information databases are obtained through N preset neural networks, and the weights corresponding to the N preset neural networks are different, where N is an integer greater than or equal to 2.
[0054] It is understandable that determining whether the industrial component corresponding to the target image has defects is considered abnormal data for determining the target image.
[0055] Optionally, in this embodiment of the application, when the multi-view model training device determines that the industrial component corresponding to the target image has a defect, the multi-view model training device can send the target image and the location information of the defect in the target image to the maintenance personnel to prompt the maintenance personnel to carry out repairs.
[0056] Optionally, in this embodiment of the application, step S103 can be specifically implemented by the following steps S103a to S103d:
[0057] S103a, the multi-view model training device acquires the third feature information in the first latent space representation library and acquires the fourth feature information in each second feature information library.
[0058] It should be noted that the third feature information mentioned above is any feature information in the first latent space representation library; the fourth feature information mentioned above is any feature information in each second feature information library.
[0059] In this embodiment of the application, the multi-view model training device can randomly acquire third feature information and fourth feature information.
[0060] S103b: The multi-view model training device processes the third feature information based on the first other feature information to obtain a first metric distribution between the third feature information and the first other feature information.
[0061] In this embodiment of the application, the aforementioned first other feature information refers to feature information other than the third feature information in the first latent space representation library, and the first metric distribution is used to characterize the similarity between the third feature information and the first other feature information.
[0062] In this embodiment of the application, for any third feature information in the first latent space representation library, the first other feature information is taken as the supervision information. Then, the distance between the third feature information and the first other feature information is measured to obtain the first metric distribution.
[0063] For example, in the first latent space representation library, for any R m (i.e., the third feature information), take k R values at the corresponding positions. i (i.e., the first other feature information) is used as supervisory information to measure R. m With each R i The distance R is obtained m The metric distribution P(R) m (i.e., the first metric distribution).
[0064] It should be noted that the metric distribution P(R) of the reconstructed complete latent space representation is... m The metric distribution should be consistent with the original features obtained by each preset neural network. (i.e., the second metric distribution below) should be as consistent as possible.
[0065] S103c, the multi-view model training device processes the fourth feature information based on the second other feature information to obtain the second metric distribution between the fourth feature information and the second other feature information.
[0066] In this embodiment of the application, the aforementioned second other feature information is feature information other than the third feature information in each second feature information database, and the second metric distribution is used to characterize the similarity between the fourth feature information and the second other feature information.
[0067] In this embodiment of the application, in each second feature information database, for any fourth feature information, the second other feature information is taken as supervision information, and then the distance between the fourth feature information and the second other feature information is measured to obtain the second metric distribution.
[0068] For example, in each second feature information database, for any (i.e., the fourth feature information), take any other k. (i.e., the second other feature information) is used as supervisory information for measurement. With each The distance between them is obtained metric distribution (i.e., the second metric distribution).
[0069] It should be noted that the embodiments of this application obtain the first metric distribution and the second metric distribution through the KL divergence metric.
[0070] S103d, when the difference between the first metric distribution and the second metric distribution is less than a preset threshold, the multi-view model training device determines whether there is a defect in the industrial component corresponding to the target image based on the first latent space representation library.
[0071] In this embodiment of the application, the multi-view model training device can calculate the difference between the first metric distribution and the second metric distribution, and thus determine whether there is a defect in the industrial component corresponding to the target image based on the first latent space representation library when the difference is small.
[0072] In this embodiment, by introducing context loss, the latent space representation is constrained so that the constructed latent space representation and each original feature have similar distributions, thereby effectively balancing the contribution of each view (i.e., N second feature information bases) to the latent space representation.
[0073] Optionally, in this embodiment of the application, before step S102 above, the multi-view model training method provided in this embodiment of the application further includes the following step S201, and the above step S102 can be specifically implemented by the following step S102a:
[0074] S201. The multi-view model training device randomly initializes at least one second latent space representation of the latent space and reconstructs the network parameters of the network.
[0075] For example, the view model training device can randomly initialize the latent space representation. and randomly initialize and reconstruct network parameters The latent space represents a network structure that shares a common structure, which is reconstructed using N views (i.e., the N third feature information bases described below).
[0076] S102a: The multi-view model training device updates at least one second latent space representation and network parameters based on N second feature information databases to obtain a first latent space representation database.
[0077] In this embodiment, the multi-view model training device can update the randomly initialized second latent space representation and the network parameters of the reconstructed network based on N second feature information databases, thereby obtaining the first latent space representation database.
[0078] In this embodiment, the multi-view model training device adjusts the network parameters of the randomly initialized second latent space representation and the reconstructed network through N second feature information databases, without the need to introduce a large amount of supervised data to train the multi-view model, thus improving the efficiency of the multi-view model training device in identifying abnormal data.
[0079] Optionally, in this embodiment of the application, the reconstructed network includes at least one layer of the reconstructed network, and the network parameters in each layer of the at least one layer of the reconstructed network are different; the above step S102 can be specifically implemented by the following steps S301 and S302:
[0080] S301, the multi-view model training device inputs at least one first latent space representation into at least one layer of reconstruction network to obtain N third feature information databases.
[0081] In this embodiment of the application, the above-mentioned N third feature information databases correspond one-to-one with N second feature information databases, and the weights of each of the N third feature information databases are different.
[0082] For example, a multi-view model training device can represent the latent space representation R. m The input is fed into at least one layer of the reconstruction network, thereby obtaining different views corresponding to the target image (i.e., N third feature information libraries, each of which corresponds to one view).
[0083] S302. The multi-view model training device processes each third feature information in the N third feature information databases based on the N second feature information databases to obtain the first latent space representation database.
[0084] For example, a multi-view model training device can represent the latent space representation R. m The input is fed into at least one layer of the reconstruction network. To supervise information, the latent space representation R is made... m It can fit the information from each view very well. Similarly, any other patch... i The corresponding latent space representation R can be obtained through reverse encoding. i (m,i∈[1,M]).
[0085] In this embodiment, the multi-view model training device can obtain N third feature information libraries corresponding to N second feature information libraries through reverse encoding. Thus, based on the N second feature information libraries, the N third feature information libraries can be processed to obtain the first latent space representation library, thereby improving the accuracy of the first latent space representation library.
[0086] This application provides a multi-view model training method. The multi-view model training device can acquire a target image, and then perform reverse encoding on the target image based on latent space and a reconstruction network to obtain a first latent space representation library. Then, when the matching degree between the first latent space representation library and N second feature information libraries is the highest, the method determines whether the industrial component corresponding to the target image has a defect based on the first latent space representation library. In this scheme, information from different views (i.e., N second feature information libraries) can be fused through latent space representation, allowing the multi-view model to learn the completeness and structure of the feature information representation. Furthermore, by using reverse encoding, the multi-view model does not require a large amount of supervised data for training. Instead, it compares the first latent space representation library with the N second feature information libraries (i.e., positive samples), and determines whether the industrial component corresponding to the target image has a defect based on the first latent space representation library when the matching degree is the highest. This avoids the problem of insufficient supervised data preventing adequate model training and resulting in poor accuracy in identifying abnormal data. Thus, the accuracy of the multi-view model training device in identifying abnormal data is improved.
[0087] Optionally, in this embodiment of the application, after step S102 above, the multi-view model training method provided in this embodiment of the application further includes the following step S401:
[0088] S401: The multi-view model training device performs alternating iterative processing on the first latent space representation library and network parameters according to the objective function to obtain the converged first latent space representation library and converged network parameters.
[0089] For example, with Minimize the target function and iteratively update the training latent space representation. and reconstructing network parameters That is, fixed latent space representation Reconstruct network parameters Train until the network stabilizes and converges; then fix and reconstruct the network parameters. Representing the latent space Training continues until the model stabilizes and converges. After multiple iterations and updates, the model converges overall, yielding the latent space representation. in and Let λ represent the losses of the reconstruction module (i.e., at least one layer of the reconstruction network) and the context constraint module (i.e., the first and second metric distributions mentioned above), respectively. n This represents the balance hyperparameter between the reconstruction loss and the losses of n KL divergences. The training objective should be to minimize the overall loss function.
[0090] Optionally, in this embodiment of the application, after step S101 above, the multi-view model training method provided in this embodiment of the application further includes the following step S501:
[0091] S501, the multi-view model training device inputs the target image into N preset neural networks respectively to obtain N second feature information databases.
[0092] In this embodiment, the weights corresponding to the N preset neural networks are different, and the N preset neural networks correspond one-to-one with the N second feature information databases.
[0093] Optionally, in the embodiments of this application, the preset neural network can be any of the following: a ResNet neural network or a Vit neural network.
[0094] For example, for each target image, it is input into N (N>=2) different pre-trained networks (such as ResNet / Vit networks pre-trained with ImageNet). For each network branch, M ordered patches (i.e., feature information) can be obtained. This represents the m-th patch block obtained from the n-th network.
[0095] In this embodiment, the multi-view model training device can obtain different view representations of the target image through multiple different pre-trained networks, and then obtain a first latent space representation library based on these different view representations through reverse encoding. In this way, the comprehensiveness of the multi-view model training device in identifying abnormal data is improved.
[0096] For example, the multi-view model training method provided in this application will be further explained and illustrated below through specific examples, such as... Figure 2 As shown, this application embodiment provides a multi-view model training system 10, which includes a pre-training feature extraction module 11, a reconstruction module 12, and a context constraint module 13.
[0097] The pre-trained feature extraction module 11 is used to input each target image into N (N>=2) different pre-trained networks (such as ResNet / Vid networks pre-trained with ImageNet). For each network branch, M ordered patch blocks can be obtained. This represents the m-th patch block obtained from the n-th network.
[0098] Reconstruction module 12 is used for reconstructing n patches extracted from n networks. m A block can be considered as a patch. m Given n views, the reconstruction module aims to obtain a common latent space representation R. m The latent space representation encompasses the effective information of each view, and can fully express the consistency and complementarity between the views. The latent space representation is obtained by inverse encoding from the observed data. The model framework is as follows: The latent space representation R... m The input is fed into the N-way reconstruction network. To supervise information, the latent space representation R is made... m It can fit the information from each view very well. Similarly, any other patch... i The corresponding latent space representation R can be obtained through reverse encoding. i (m,i∈[1,M]).
[0099] Context constraint module 13 is used during training, where input images are processed in batches, with each batch containing p images. For each network branch, each image corresponds to a feature library. It also corresponds to a latent space representation library. Each network branch, for the entire batch, contains p*M feature representations and a latent space representation. Wherein, for any... Choose any k other ones As supervisory information, measurement With each The distance is obtained metric distribution Similarly, for any R m Take k R values at the corresponding positions i As supervisory information, the metric R m With each R i The distance R is obtained m The metric distribution P(R) m The metric distribution P(R) of the reconstructed complete latent space representation. m The metric distribution should be consistent with the original feature distribution obtained from each network. Keep it as consistent as possible.
[0100] loss function
[0101] in and Let λ represent the losses of the reconstruction module and the context constraint module, respectively.n This represents the balance hyperparameter between the reconstruction loss and the losses of n KL divergences. The training objective should be to minimize the overall loss function.
[0102] Specifically, such as Figure 3 As shown, the training method for the multi-view model is as follows:
[0103] Randomly initialize the hidden space representation Randomly initialize and rebuild network parameters Each m corresponds to N reconstruction networks, sharing a common network structure, that is, using N views for reconstruction.
[0104] To minimize The objective function is to iteratively update the training latent space representation. and reconstructing network parameters That is, fixed latent space representation Reconstruct network parameters Train until the network stabilizes and converges; then fix and reconstruct the network parameters. Representing the latent space Training continues until the model stabilizes and converges. After multiple iterations and updates, the model converges overall, yielding the latent space representation.
[0105] This application provides a multi-view model training method. The multi-view model training device can acquire a target image, and then perform reverse encoding on the target image based on latent space and a reconstruction network to obtain a first latent space representation library. Then, when the matching degree between the first latent space representation library and N second feature information libraries is the highest, the method determines whether the industrial component corresponding to the target image has a defect based on the first latent space representation library. In this scheme, information from different views (i.e., N second feature information libraries) can be fused through latent space representation, allowing the multi-view model to learn the completeness and structure of the feature information representation. Furthermore, by using reverse encoding, the multi-view model does not require a large amount of supervised data for training. Instead, it compares the first latent space representation library with the N second feature information libraries (i.e., positive samples), and determines whether the industrial component corresponding to the target image has a defect based on the first latent space representation library when the matching degree is the highest. This avoids the problem of insufficient supervised data preventing adequate model training and resulting in poor accuracy in identifying abnormal data. Thus, the accuracy of the multi-view model training device in identifying abnormal data is improved.
[0106] This application embodiment can divide the multi-view model training device into functional modules or functional units according to the above method examples. For example, each function can be divided into its own functional modules or functional units, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module or functional unit. The module or unit division in this application embodiment is illustrative and represents only one logical functional division; other division methods may be used in actual implementation.
[0107] like Figure 4 The diagram shown is a structural schematic of a multi-view model training device provided in an embodiment of this application. The device includes: an acquisition unit 201, a processing unit 202, and a determination unit 203.
[0108] The acquisition unit 201 is used to acquire a target image, which is an image of an industrial component. The processing unit 202 is used to perform reverse encoding on the target image based on the latent space and a reconstruction network to obtain a first latent space representation library corresponding to the latent space. The first latent space representation library includes at least one first latent space representation, which is used to characterize at least two first feature information with different weights corresponding to the target image. The determination unit 203 is used to determine whether the industrial component corresponding to the target image has a defect, based on the first latent space representation library, when the matching degree between the first latent space representation library and N second feature information libraries is the highest. The N second feature information libraries are obtained through N preset neural networks, and the weights corresponding to the N preset neural networks are different, where N is an integer greater than or equal to 2. The latent space is used to indicate the consistency and complementarity between at least one first feature information.
[0109] In one possible implementation, the multi-view model training device further includes an initialization unit. The initialization unit is used to randomly initialize at least one second latent space representation and the network parameters of the reconstruction network before performing reverse encoding on the target image based on the latent space and the reconstruction network to obtain the first latent space representation library corresponding to the latent space. The processing unit 202 is specifically used to update at least one second latent space representation and the network parameters according to N second feature information libraries to obtain the first latent space representation library.
[0110] In one possible implementation, the reconstruction network includes at least one reconstruction network layer, with different network parameters in each layer; the processing unit 202 is specifically used to input at least one first latent space representation into at least one reconstruction network layer to obtain N third feature information libraries, the N third feature information libraries correspond one-to-one with N second feature information libraries, and the weights corresponding to each third feature information library are different; and to process each third feature information in the N third feature information libraries according to the N second feature information libraries to obtain the first latent space representation library.
[0111] In one possible implementation, the processing unit 202 is further configured to, after obtaining the first latent space representation library, perform alternating iterative processing on the first latent space representation library and network parameters according to the objective function, so as to obtain the converged first latent space representation library and converged network parameters.
[0112] In one possible implementation, the determining unit 203 is specifically used to acquire third feature information from a first latent space representation library and fourth feature information from each second feature information library; and to process the third feature information according to first other feature information to obtain a first metric distribution between the third feature information and the first other feature information, wherein the first other feature information is feature information other than the third feature information in the first latent space representation library, and the first metric distribution is used to characterize the similarity between the third feature information and the first other feature information; and to process the fourth feature information according to second other feature information to obtain a second metric distribution between the fourth feature information and the second other feature information, wherein the second other feature information is feature information other than the third feature information in each second feature information library, and the second metric distribution is used to characterize the similarity between the fourth feature information and the second other feature information; and, if the difference between the first metric distribution and the second metric distribution is less than a preset threshold, to determine whether there is a defect in the industrial component corresponding to the target image according to the first latent space representation library.
[0113] In one possible implementation, the processing unit 202 is further configured to, after acquiring the target image, input the target image into N preset neural networks respectively to obtain N second feature information databases, wherein the weights corresponding to the N preset neural networks are different, and the N preset neural networks correspond one-to-one with the N second feature information databases respectively.
[0114] This application provides a multi-view model training device that integrates information from different views (i.e., N second feature information bases) through latent space representation. This allows the multi-view model to learn the completeness and structure of the feature information. Furthermore, by using reverse encoding, the multi-view model does not require a large amount of supervised data for training. Instead, it compares the first latent space representation base with the N second feature information bases (i.e., positive samples). Based on the first latent space representation base, it determines whether there are defects in the industrial parts corresponding to the target image when the matching degree is highest. This avoids the problem of insufficient supervised data preventing the model from being adequately trained and resulting in poor accuracy in identifying abnormal data. Thus, the accuracy of the multi-view model training device in identifying abnormal data is improved.
[0115] When implemented in hardware, the acquisition unit 201 described in this embodiment can be integrated onto the communication interface, and the processing unit 202 and the determination unit 203 can be integrated onto the processor. Specific implementation methods are as follows: Figure 5 As shown.
[0116] Figure 5 A schematic diagram of another possible structure of the multi-view model training device involved in the above embodiments is shown. This multi-view model training device includes a processor 302 and a communication interface 303. The processor 302 is used to control and manage the operation of the multi-view model training device, for example, executing the steps performed by the processing unit 202 and the determining unit 203, and / or performing other processes of the technology described herein. The communication interface 303 is used to support communication between the multi-view model training device and other network entities, for example, executing the steps performed by the acquiring unit 201. The multi-view model training device may also include a memory 301 and a bus 304. The memory 301 is used to store the program code and data of the multi-view model training device.
[0117] The memory 301 may be a memory in a multi-view model training device, and the memory may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk or solid-state drive; the memory may also include a combination of the above types of memory.
[0118] The processor 302 described above can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0119] Bus 304 can be an Extended Industry Standard Architecture (EISA) bus, etc. Bus 304 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0120] Figure 6 This is a schematic diagram of the structure of chip 170 provided in an embodiment of this application. Chip 170 includes one or more (including two) processors 1710 and communication interfaces 1730.
[0121] Optionally, the chip 170 also includes a memory 1740, which may include read-only memory and random access memory, and provides operation instructions and data to the processor 1710. A portion of the memory 1740 may also include non-volatile random access memory (NVRAM).
[0122] In some implementations, memory 1740 stores elements such as execution modules or data structures, or subsets thereof, or extended sets thereof.
[0123] In this embodiment of the application, the corresponding operation is executed by calling the operation instructions stored in the memory 1740 (the operation instructions can be stored in the operating system).
[0124] The processor 1710 described above can implement or execute various exemplary logic blocks, units, and circuits described in conjunction with the disclosure of this application. The processor can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, units, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0125] The memory 1740 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk or solid-state drive; the memory may also include combinations of the above types of memory.
[0126] The Bus 1720 can be an Extended Industry Standard Architecture (EISA) bus, etc. The Bus 1720 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The symbol is represented by only one line, but this does not mean that there is only one bus or one type of bus.
[0127] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0128] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the multi-view model training method described in the above method embodiments.
[0129] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the multi-view model training method in the method flow shown in the above method embodiments.
[0130] The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires; a portable computer disk drive; a hard disk drive; a random access memory (RAM); a read-only memory (ROM); an erasable programmable read-only memory (EPROM); a register; a hard disk drive; an optical fiber; a portable compact disk read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination thereof; or any other form of computer-readable storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). In the embodiments of this application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0131] Embodiments of the present invention provide a computer program product containing instructions that, when executed on a computer, cause the computer to perform the multi-view model training method described in the above method embodiments.
[0132] Since the multi-view model training device, computer-readable storage medium, and computer program product in the embodiments of the present invention can be applied to the above method, the technical effects obtained can also be referred to the above method embodiments. The embodiments of the present invention will not be repeated here.
[0133] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0134] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0135] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0136] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A multi-view model training method, characterized in that, The method includes: Acquire a target image, wherein the target image is an image of an industrial component; Based on the latent space and reconstruction network, the target image is inversely encoded to obtain a first latent space representation library corresponding to the latent space. The first latent space representation library includes at least one first latent space representation, and the at least one latent space representation is used to characterize at least two first feature information with different weights corresponding to the target image. Obtain the third feature information from the first latent space representation library, and obtain the fourth feature information from each of the N second feature information libraries; the N second feature information libraries are obtained through N preset neural networks, and the weights of the N preset neural networks are different, where N is an integer greater than or equal to 2; Based on the first other feature information, the third feature information is processed to obtain a first metric distribution between the third feature information and the first other feature information. The first other feature information is feature information in the first latent space representation library other than the third feature information. The first metric distribution is used to characterize the similarity between the third feature information and the first other feature information. Based on the second other feature information, the fourth feature information is processed to obtain a second metric distribution between the fourth feature information and the second other feature information. The second other feature information is feature information other than the third feature information in each second feature information database. The second metric distribution is used to characterize the similarity between the fourth feature information and the second other feature information. If the difference between the first metric distribution and the second metric distribution is less than a preset threshold, the industrial component corresponding to the target image is determined to have a defect based on the first latent space representation library. The latent space is used to indicate the consistency and complementarity among the N second feature information bases.
2. The method according to claim 1, characterized in that, Before performing inverse encoding on the target image based on the latent space and reconstruction network to obtain the first latent space representation library corresponding to the latent space, the method further includes: Randomly initialize at least one second hidden space representation of the hidden space and the network parameters of the reconstructed network; The method of inverse encoding the target image based on the latent space and reconstruction network to obtain the first latent space representation library corresponding to the latent space includes: Based on the N second feature information databases, the at least one second latent space representation and the network parameters are updated to obtain the first latent space representation database.
3. The method according to claim 2, characterized in that, The reconstruction network includes at least one reconstruction network layer, with different network parameters in each layer; the step of updating the at least one second latent space representation based on the N second feature information databases to obtain the first latent space representation database includes: The at least one first latent space representation is input into the at least one layer of the reconstruction network to obtain N third feature information libraries. The N third feature information libraries correspond one-to-one with the N second feature information libraries, and each third feature information library has a different weight. Based on the N second feature information databases, each third feature information in the N third feature information databases is processed to obtain the first latent space representation database.
4. The method according to claim 2, characterized in that, After obtaining the first latent space representation library, the method further includes: Based on the objective function, the first latent space representation library and the network parameters are iteratively processed alternately to obtain the converged first latent space representation library and the converged network parameters.
5. The method according to claim 1, characterized in that, After acquiring the target image, the method further includes: The target image is input into N preset neural networks to obtain N second feature information databases. The weights of the N preset neural networks are different, and the N preset neural networks correspond one-to-one with the N second feature information databases.
6. A multi-view model training device, characterized in that, The device includes: an acquisition unit, a processing unit, and a determination unit; The acquisition unit is used to acquire a target image, wherein the target image is an image of an industrial component; The processing unit is used to perform reverse encoding on the target image based on the latent space and the reconstruction network to obtain a first latent space representation library corresponding to the latent space. The first latent space representation library includes at least one first latent space representation, and the at least one latent space representation is used to characterize at least two first feature information with different weights corresponding to the target image. The determining unit is configured to acquire third feature information from the first latent space representation library and fourth feature information from each of the N second feature information libraries; the N second feature information libraries are obtained through N preset neural networks, each with different weights, and N is an integer greater than or equal to 2; the third feature information is processed according to first other feature information to obtain a first metric distribution between the third feature information and the first other feature information, wherein the first other feature information is feature information in the first latent space representation library other than the third feature information, and the first metric distribution is used to characterize the similarity between the third feature information and the first other feature information; the fourth feature information is processed according to second other feature information to obtain a second metric distribution between the fourth feature information and the second other feature information, wherein the second other feature information is feature information in each of the second feature information libraries other than the third feature information, and the second metric distribution is used to characterize the similarity between the fourth feature information and the second other feature information; if the difference between the first metric distribution and the second metric distribution is less than a preset threshold, the industrial component corresponding to the target image is determined to have a defect according to the first latent space representation library; The latent space is used to indicate the consistency and complementarity among the N second feature information bases.
7. The apparatus according to claim 6, characterized in that, The multi-view model training device also includes an initialization unit; The initialization unit is used to randomly initialize at least one second latent space representation of the latent space and the network parameters of the reconstruction network before performing reverse encoding on the target image based on the latent space and the reconstruction network to obtain the first latent space representation library corresponding to the latent space. The processing unit is specifically used to update the at least one second latent space representation and the network parameters according to the N second feature information databases to obtain the first latent space representation database.
8. The apparatus according to claim 7, characterized in that, The reconstruction network includes at least one reconstruction network layer, with different network parameters in each layer. The processing unit is specifically used to input the at least one first latent space representation into the at least one reconstruction network layer to obtain N third feature information libraries, which correspond one-to-one with the N second feature information libraries, and each third feature information library has a different weight. The processing unit is also used to process each third feature information in the N third feature information libraries according to the N second feature information libraries to obtain the first latent space representation library.
9. The apparatus according to claim 7, characterized in that, The processing unit is further configured to, after obtaining the first latent space representation library, perform alternating iterative processing on the first latent space representation library and the network parameters according to the objective function, to obtain the converged first latent space representation library and the converged network parameters.
10. The apparatus according to claim 6, characterized in that, The processing unit is further configured to, after acquiring the target image, input the target image into N preset neural networks respectively to obtain the N second feature information databases, wherein the weights corresponding to the N preset neural networks are different, and the N preset neural networks correspond one-to-one with the N second feature information databases.
11. A multi-view model training device, characterized in that, include: A processor and a communication interface; the communication interface is coupled to the processor, the processor being used to run computer programs or instructions to implement the multi-view model training method as described in any one of claims 1-5.
12. A computer-readable storage medium storing instructions, characterized in that, When the computer executes the instruction, the computer performs the multi-view model training method as described in any one of claims 1-5.
Citation Information
Patent Citations
Missing multi-view data classification method and system
CN110543916A
Unsupervised defect detection method based on quantization auto-encoder
CN115375604A